| # Amplegest — Architecture |
|
|
| > Render the diagram in VS Code with Mermaid support or paste it into [mermaid.live](https://mermaid.live). |
|
|
| Amplegest separates data collection, retrieval, synthesis, deterministic controls, and presentation. The target user is a portfolio manager or senior equity analyst reviewing a post-earnings change in thesis. |
|
|
| ```mermaid |
| --- |
| title: Amplegest — Post-Earnings Decision Architecture |
| --- |
| flowchart TB |
| classDef ext fill:#f3e8ff,stroke:#9333ea,color:#581c87 |
| classDef store fill:#dcfce7,stroke:#16a34a,color:#14532d |
| classDef warm fill:#fed7aa,stroke:#ea580c,color:#7c2d12 |
| classDef proc fill:#fef9c3,stroke:#ca8a04,color:#78350f |
| classDef state fill:#fde68a,stroke:#b45309,color:#78350f |
| classDef rag fill:#e9d5ff,stroke:#7c3aed,color:#3b0764 |
| classDef io fill:#fce7f3,stroke:#db2777,color:#831843 |
| classDef gate fill:#fee2e2,stroke:#dc2626,color:#7f1d1d |
| |
| USER([PM / senior equity analyst]):::io --> UI["Streamlit research stack<br/>Company Overview · PM Flash · Evidence & Deltas · Financials · Ask AI"]:::io |
| UI --> PE |
| |
| subgraph LG ["LangGraph · short-term state"] |
| direction LR |
| PE["profile_evidence<br/>deterministic evidence.v1 retrieval"]:::state |
| AG["agent<br/>Claude Haiku 4.5"]:::state |
| TN["tools<br/>parallel execution"]:::state |
| ND["coverage nudge · once<br/>if filing + transcript<br/>were not both searched"]:::state |
| SY["synthesis<br/>validated BriefOutput"]:::state |
| PS["post_synthesis<br/>item reliability + deltas"]:::state |
| PE --> AG |
| AG == "tool calls · bounded rounds" ==> TN |
| TN == "ToolMessage" ==> AG |
| AG -. "coverage floor" .-> ND -.-> AG |
| AG -- "done or cap" --> SY --> PS |
| end |
| |
| PS --> GATE["Presentation policy<br/>alignment required<br/>missing or uncertain = hidden"]:::gate |
| GATE --> UI |
| LSM{{LangSmith}}:::ext -. "trace" .- LG |
| |
| TN <--> TOOLS |
| |
| subgraph TOOLS ["Runtime tools · RAG + SQL + public data"] |
| direction TB |
| T1["get_financial_metrics"]:::proc |
| T2["search_filing"]:::proc |
| T3["search_transcript"]:::proc |
| T4["search_news"]:::proc |
| T5["get_analyst_expectations"]:::proc |
| |
| subgraph RAG ["Filing and transcript retrieval"] |
| direction LR |
| QE["embed query<br/>MiniLM-L6-v2"]:::rag |
| VS["Chroma similarity<br/>5× over-fetch"]:::rag |
| RR["cross-encoder reranker<br/>bge-reranker-base"]:::rag |
| QE --> VS --> RR |
| end |
| |
| T2 --> QE |
| T3 --> QE |
| end |
| |
| subgraph WARM ["Warm tier · 24-hour TTL"] |
| AVC[(av_cache · SQLite)]:::warm |
| end |
| |
| AVL{{Alpha Vantage}}:::ext --> AVC |
| YF{{yfinance}}:::ext --> AVC |
| AVC --> T5 |
| |
| subgraph LTM ["Long-term stores · offline ingestion"] |
| direction LR |
| EDG{{SEC EDGAR}}:::ext |
| AVT{{Earnings-call transcript source}}:::ext |
| EX["XBRL extractor<br/>structured metrics"]:::rag |
| CK["chunker<br/>500 words · 50 overlap"]:::rag |
| EM["embedder<br/>MiniLM-L6-v2"]:::rag |
| MD["metadata<br/>ticker · period · section<br/>source · date · context"]:::rag |
| DB[(SQLite metrics)]:::store |
| CHF[(Chroma · filings)]:::store |
| CHT[(Chroma · transcripts)]:::store |
| |
| EDG --> EX --> DB |
| EDG --> CK |
| AVT --> CK |
| CK --> EM --> MD --> CHF |
| MD --> CHT |
| end |
| |
| DB -- "filing data" --> T1 |
| CHF -- "10-K / 10-Q evidence" --> VS |
| CHT -- "transcript evidence" --> VS |
| TAV{{Tavily news}}:::ext -- "post-filing context" --> T4 |
| ``` |
| |
| **search_filing** and **search_transcript** over-fetch Chroma candidates, optionally filter by period, rerank the candidates, and return the best matches. News search is separate and does not inherit filing-level confidence. |
|
|
| ## Decision surfaces |
|
|
| | Surface | Primary payload | |
| | --- | --- | |
| | **Company Overview** | Stable sourced company profile plus independently refreshed price context and public news. Business model, geographic exposure, three-year trends, attention themes, associated price events, and monitoring variables remain concise and inspectable. | |
| | **PM Flash** | As-of and coverage; experimental PM read-through; thesis-confirming/challenging sourced points; experimental Swing Factor; watch items. | |
| | **Evidence & Deltas** | Source-backed facts, verbatim evidence, deterministic detector outputs labelled heuristic, risks, commentary, and experimental AI hypotheses. | |
| | **Financials** | Structured historical metrics, trends, guidance history, earnings history, and exports. | |
| | **Ask AI** | Cross-cutting grounded Q&A over the available evidence. | |
|
|
| The navigation order is intentional: company onboarding → decision memo → evidence audit → model depth → cross-cutting AI tool. |
|
|
| ## Company Overview |
|
|
| Company Overview is produced inside the same streaming synthesis call as the rest of the brief. Before the agent loop, a deterministic node retrieves the stable profile evidence set—10-K Business, segments/geography, strategic evolution, and transcripts—and injects each `evidence.v1` envelope as a raw human message. Those messages do not satisfy the brief's tool-coverage gate, so the agent must still retrieve the current-quarter evidence required for the decision memo. After synthesis, the overview section is validated separately, passed through the unchanged deterministic evidence verifier, stripped of every non-verified fact, and saved under its source fingerprint. A profile failure is isolated and cannot fail the brief. |
|
|
| The former standalone call existed for three practical reasons: its source fingerprint made the profile stable between earnings; its evidence set was broader and more structural than the latest-quarter brief; and a separate response kept the main synthesis output smaller. The merged design deliberately accepts those trade-offs. Profile evidence is now injected deterministically before the agent loop, which adds only marginal context; the synthesis output limit is raised to 16,384 tokens; and fingerprint-based cache lookup remains in the UI, while a new overview is generated only with a new brief run. This means “Generate Brief” regenerates the overview each time, an acceptable marginal output-token cost in exchange for one billed synthesis call and one coherent evidence-grounded result. |
|
|
| ## Output taxonomy |
|
|
| | Payload or derived field | Epistemic class | Presentation | |
| | --- | --- | --- | |
| | **what_changed** | Source-backed factual claims | Evidence & Deltas | |
| | **bull_points**, **bear_points** | Source-backed factual claims | PM Flash as thesis-confirming/challenging; full evidence in Evidence & Deltas | |
| | **risks_categorized**, **management_commentary** | Source-backed factual claims | Evidence & Deltas | |
| | **quarter_deltas** | Deterministic detector output; significance is heuristic | Evidence & Deltas with **Heuristic · validate** | |
| | **what_matters_most** | Experimental AI synthesis | PM Flash with **AI · experimental** | |
| | **analytical_tensions** | Experimental two-sided hypothesis | PM Flash Swing Factor and Evidence & Deltas | |
| | **between_the_lines** | Experimental interpretation | Evidence & Deltas; may provide Swing Factor fallback | |
| | **earnings_quality_signals** | Experimental assessment | Evidence & Deltas | |
| | **non_obvious_takeaway** | Experimental interpretation | Retained in payload; never presented as a sourced fact | |
| | **what_to_watch** | Monitoring prompts | PM Flash | |
| | **sentiment** | Stored analytical payload | Hidden from PM-facing UI until calibrated | |
| | **market_expectations** | Stored market-data payload | Actual/consensus and D1/D5 hidden until aligned | |
|
|
| ## Reliability contract |
|
|
| Every **SourcedFact** carries: |
|
|
| - one source in {10-K, 10-Q, transcript, news}; |
| - an item-level reliability label; |
| - a short verbatim **evidence_snippet**. |
| |
| Post-synthesis code rewrites item-level reliability using deterministic defaults and corroboration rules. This metadata helps the analyst decide what to verify first. It is not a calibrated probability and must not be aggregated into a global conviction indicator. |
| |
| Consequently: |
| |
| - item-level source and reliability stay visible; |
| - aggregate HIGH / MEDIUM / LOW distributions, percentages, or headline badges stay hidden. |
| |
| ## Fail-closed alignment gate |
| |
| The presentation layer treats every market-derived or aggregate field as unavailable until its prerequisites are proven. |
| |
| | Candidate output | Gate condition | |
| | --- | --- | |
| | Actual vs consensus | Same issuer, fiscal period, metric definition, accounting basis, units, currency, and pre-release estimate timestamp | |
| | D1 / D5 return | Exact release timestamp, timezone, market-session classification, exchange calendar, and deterministic price anchors | |
| | Aggregate sentiment | Sufficient section coverage, calibrated weights, stability testing, and a documented use | |
| | Aggregate HIGH / MEDIUM / LOW | A validated aggregation methodology that does not misrepresent source provenance as truth probability | |
| |
| The default is false. Missing metadata, an ambiguous period, or an uncertain event time prevents rendering. Raw values may remain in storage for validation; storage presence is never display permission. |
| |
| The same rule applies to evidence coverage: **Verified** requires `total > 0`, `verified == total`, and the deterministic status `VERIFIED`. Older briefs without this summary are labelled **Legacy / unverified**, regardless of their source tags. |
| |
| ## Known alignment priorities |
| |
| 1. XBRL observations must be selected by accession **and** period identity, including end date and duration. |
| 2. Analyst estimates must be mapped to the same reported fiscal period and metric basis. |
| 3. Event returns must account for after-hours and pre-market releases. |
| 4. Evidence should carry a stable source locator: accession, section, URL, transcript speaker, and chunk identifier. |
| |
| These controls precede portfolio-level scoring, comparative ranking, or additional AI interpretation. |
| |