|
Download CORVUS.md from Reverb/Test: direct link, hf CLI and curl.
- Browser
- Download file 24.5 kB
-
https://huggingface.co/Reverb/Test/resolve/main/CORVUS.md
- Command line
-
hf download hf://Reverb/Test/CORVUS.md
-
curl -L -o CORVUS.md https://huggingface.co/Reverb/Test/resolve/main/CORVUS.md
24.5 kB
| # Corvus — Platform Blueprint v2.0 | |
| *Corvus · Platform Blueprint v2.0 · August 2026 · Confidential* | |
| *Solo-track-owned document — AI Engineer reference, shared with SDE co-founder* | |
| **This version supersedes v1.0 (June 2026).** Where they disagree, this document | |
| reflects what was actually built and decided since — not what was originally planned. | |
| Section 0 explains what changed and why, so nothing here reads as a silent reversal. | |
| --- | |
| ## 0. What changed since v1.0, and why | |
| v1.0 was written pre-code, as a complete company blueprint. Since then Corvus went | |
| through a real build (`PLAN.md`, Wave 1), a code review (`docs/code-review.md`, 28 | |
| verified findings), and a scoping pass specifically on the AI-reasoning side | |
| (§7 below). The result is a narrower, more honest v2. Nothing in the pivot changes | |
| the core thesis — it changes the plan for reaching it. | |
| | v1.0 said | v2.0 says | Why | | |
| |---|---|---| | |
| | Salesforce + HubSpot + CSV connectors, bespoke MCP servers you write | **MCP-as-connectors**: reuse existing community MCP servers (Postgres-MCP first) | Writing and maintaining bespoke connectors was the single biggest scope item in v1.0. Existing MCP servers exist for most of this; Corvus's job is to be a good MCP *host*, not a connector author. | | |
| | Railway / Fly.io → AWS at scale | **Self-host on Hetzner, staged in waves** | Chosen deliberately over "managed-first." Cheaper, and data-sovereignty positioning (§4) is stronger when Corvus's own infra is self-hosted, not merely BYOK at the model layer. | | |
| | 6-layer stack incl. ClickHouse, S3, Vault, ~12 services from day one | **4 containers at MVP** (Postgres+pgvector, Redis, app, Caddy); ClickHouse/observability/MinIO are Wave 3 | Standing up 12 services before a product exists was identified as the exact trap the original blueprint warned against in its own risk section. Staged waves fix that. | | |
| | Subagent orchestration, hooks system, episodic/semantic memory store | **Deferred, not cancelled.** Wave 1 ships one bounded reasoning loop (`ask.ts`, `MAX_TURNS = 8`) | Anthropic's own finding (§7.1) and Corvus's own build experience agree: the loop itself should stay simple. Complexity earns its place once the semantic layer under it is solid — not before. | | |
| | CORVUS.md as a literal config file the harness parses at runtime | **CORVUS.md as this document** — a human/AI-readable blueprint, not a runtime artifact | The functional equivalent — the tenant's actual business definitions — now lives in the **Metric Registry** and **Entity Catalog** (§7.3–7.4), which is queryable and versioned, not a markdown file the harness re-parses per request. | | |
| | "Cursor for business data," CFO/audit/compliance scope included | **Narrower**: sales/ops-facing BI, compliance explicitly cut | Kept from v1.0 — this one didn't change. Compliance-grade reporting has different liability and audit requirements than an ops manager asking about win rate; conflating them was correctly identified as scope creep. | | |
| | One data source, live-call only | **Same for MVP** — canonical-table sync (a local warehouse) is explicitly Wave 2, gated on the Insight Feed needing it | Confirmed still correct: live-call MCP is fine for ad-hoc queries; it breaks for scheduled background scans and heavy aggregation, so sync is deferred until something actually needs it. | | |
| The one-paragraph version: **v1.0 was the vision, correctly ambitious for a company | |
| document. v2.0 is what two engineers can actually build and defend, in an order that | |
| proves the trust story (§1) before anything else.** | |
| --- | |
| ## 1. What Corvus is | |
| Corvus is a **multi-tenant business-intelligence harness**: a thin web app in front of | |
| a harness core that takes a plain-English question from a business user — a sales VP, | |
| an ops manager, not a developer — figures out what data it needs, fetches **verified** | |
| numbers from the company's connected systems, and returns an answer with **citations | |
| and a confidence level**. | |
| **Tagline:** *Your data. Your model. Total clarity.* | |
| **Positioning:** "Cursor for business data" — a serious professional tool a business | |
| user can trust in front of leadership, in the register of Linear, Vercel, Stripe, or | |
| Hex. Not a consumer AI chatbot. | |
| **The one non-negotiable rule**, unchanged since v1.0 and the thing everything else in | |
| this document exists to protect: | |
| > **The AI model never computes numbers.** It calls a tool; the tool — deterministic, | |
| > tested code — returns a verified number; the model only explains it. | |
| A wrong number costs Corvus the customer instantly. Every architectural decision below | |
| is in service of never presenting an unverified number as if it were verified. | |
| --- | |
| ## 2. The one loop that matters | |
| > User connects **one** data source → asks a question in plain English → | |
| > gets a **correct answer with a citation to the source rows** + a confidence level. | |
| Everything else in this document is addition. This loop already works end-to-end in | |
| the current build (`askCorvus()` in `src/harness/ask.ts`). Ship it to design partners, | |
| let their real questions decide what's next. | |
| ```mermaid | |
| sequenceDiagram | |
| participant U as User (browser) | |
| participant R as /api/ask route (SDE) | |
| participant A as askCorvus() — ask.ts (AI) | |
| participant G as callModel() — gateway (SEAM) | |
| participant T as Tools — registry.ts (AI) | |
| participant S as Semantic Layer (AI, NEW — §7) | |
| participant C as MCP connectors (SDE) | |
| U->>R: "Which products are about to run out of stock?" | |
| R->>A: askCorvus(tenantId, question) | |
| A->>A: assemble tools = built-in + discover_metrics + MCP-discovered | |
| A->>G: callModel({tenantId, messages, tools}) | |
| G-->>A: "I need to call discover_metrics, then calculate_kpi" | |
| A->>S: discover_metrics(query) → governed metric match | |
| A->>T: runTool(tenantId, call) | |
| T->>C: (if raw query) callMcpTool → real DB query | |
| C-->>T: real rows | |
| T-->>A: ToolResult { value, citation } | |
| A->>G: feed tool results back, call model again | |
| G-->>A: (repeat until no more tool calls) | |
| A-->>R: { answer, citations, confidence } | |
| R-->>U: rendered answer + "show me the data" | |
| ``` | |
| --- | |
| ## 3. The stack, as built | |
| ``` | |
| Layer 1 Experience Next.js web app (query box + answer + "show me the data") | |
| Layer 2 API + Auth Next.js API routes · Better Auth (multi-tenant) · JWT | |
| Layer 3 Harness Core Model Gateway · MCP Host · Tool Registry · Semantic Layer (NEW) | |
| Layer 4 Connectors Existing MCP servers (Postgres-MCP first; platform MCPs as needed) | |
| Layer 5 Data PostgreSQL + pgvector · Redis (canonical tables added when sync is needed) | |
| Layer 6 Infra Hetzner VPS · Docker · Caddy (auto-TLS) | |
| ``` | |
| **Layer 3 is the moat**, same as v1.0 said — that part never changed. What changed is | |
| *what's in it*. v1.0 imagined Context Engine + Semantic Layer + Agent Orchestrator + | |
| Hooks + Memory as parallel systems built together. What's actually true a build cycle | |
| later: the **Semantic Layer is the load-bearing piece**, and it should be built | |
| deliberately, on top of a loop that's already proven simple and correct. §7 is the | |
| current, concrete design for it — this is the piece v1.0 gestured at ("Business | |
| Semantic Layer... build alongside first connector") but never fully specified. | |
| ### Model Gateway (BYOK) | |
| Unchanged in principle from v1.0, real and shipped in code: | |
| - `callModel(tenantId, messages, tools) → ModelCallResult` — the single door every | |
| provider call goes through (`src/harness/gateway/`). | |
| - **BYOK is architectural, not a feature.** A tenant's own Anthropic/OpenAI key, | |
| encrypted at rest (AES-256-GCM, HKDF per-version keys), decrypted only inside the | |
| gateway. One managed default (Claude) for tenants without their own key. | |
| - This remains the enterprise sales unlocker v1.0 identified: BYOK removes the | |
| procurement blocker, honors data-processing agreements for sensitive-data | |
| customers, and means Corvus never becomes obsolete as models improve — a tenant | |
| upgrades their own key. | |
| - `ollama` provider exists but is dev/test only, gated behind | |
| `CORVUS_DEV_PROVIDER=ollama`, never in a production code path. | |
| ### MCP Host — Layer 3 as MCP *client*, Layer 4 as connectors | |
| This is the single biggest architectural pivot from v1.0, and it's a simplification, | |
| not a compromise: | |
| - **What it buys:** no bespoke Salesforce/HubSpot/Odoo API integration to write or | |
| maintain. The model calls MCP tools through a standard interface. Corvus is | |
| architecturally the host/harness; MCP servers are the tool layer — exactly the | |
| positioning v1.0 wanted, achieved with far less code. | |
| - **What it doesn't solve** (still true, still designed around): MCP servers are thin | |
| API wrappers returning raw data, not canonical schema — the semantic layer (§7) is | |
| still Corvus's to build. Most are live-call, not synced — fine for ad-hoc queries, | |
| blocks aggregation-heavy analytics and the Insight Feed until Wave 2 adds a sync | |
| layer. Corvus inherits each server's maintenance quality and auth model — vet, | |
| pin versions, keep a fallback. | |
| - **The shortcut that made Postgres-first correct:** the most mature MCP server in | |
| existence is the Postgres MCP server, and self-hosted Odoo — common among Arab | |
| SMBs, and among Corvus's own early prospects (BeExpress, Sweet&Fit) — runs on | |
| Postgres. Point a read-only Postgres MCP server at the Odoo DB and skip the Odoo | |
| API entirely for MVP. | |
| - **Security posture, hard-won and non-negotiable:** the community | |
| `@henkey/postgres-mcp-server` advertises 18 tools including arbitrary SQL and | |
| mutations, and does **not** block them without extra config — verified directly | |
| against a live instance, not assumed from docs. Corvus enforces its own tool | |
| allowlist at the host layer (`mcp/host.ts`) rather than trusting a connector's | |
| internal safety switch. Only `pg_execute_query` is allowlisted; its own handler | |
| independently rejects anything that isn't `SELECT`/`WITH`. This same | |
| don't-trust-the-connector discipline applies to every future MCP server added. | |
| ### Tool Registry — the trust primitive, unchanged | |
| `query_source`, `calculate_kpi`, `compare_periods` — three deterministic tools. The | |
| model never does math; tools return verified numbers with citations, the model | |
| explains them. This is the actual moat, and it hasn't moved since v1.0 first stated | |
| the rule. | |
| --- | |
| ## 4. What makes it defensible | |
| Carried forward from v1.0, still the correct list, now field-tested: | |
| - **The model never computes numbers** — only tested, deterministic code does. | |
| Competitors letting a model estimate a number are one hallucination from losing a | |
| customer's trust. | |
| - **Every answer is cited** — a user can always see the underlying rows, not just | |
| trust a black box. `confidence` (`high`/`medium`/`low`) is scored honestly: no | |
| tool calls, a tool error, or hitting the reasoning-loop turn cap all suppress | |
| `"high"` — the system tells the truth about its own uncertainty rather than | |
| presenting every answer with false confidence. | |
| - **BYOK** — cost control and data-handling comfort at the enterprise end; see §3. | |
| - **Multi-tenant from day one** — every table has `tenant_id`; every query is scoped. | |
| - **Data sovereignty**, elevated since v1.0: self-hosting Corvus's own infra (not | |
| just BYOK at the model layer) is the strongest differentiator against | |
| platform-native agentic BI (Oracle, SAP, Microsoft Fabric) and semantic-layer | |
| incumbents (Snowflake Cortex, Databricks, Looker/Gemini) — none of which offer a | |
| cross-platform, self-hosted, BYOK harness. This was validated in a competitive | |
| pass after v1.0 shipped: the "connect data to AI" premise is increasingly covered | |
| by incumbents; sovereignty over both the model *and* the infrastructure isn't. | |
| --- | |
| ## 5. Who it's for | |
| Unchanged from v1.0: operations- and sales-facing professionals at small/medium | |
| businesses running a self-hosted, database-backed system (Odoo is the anchor case), | |
| who currently wait on an analyst or don't ask the question at all because the | |
| friction is too high. Validate with a small number of design partners before | |
| expanding sources or customer base — v1.0's own instinct here ("talk to 20 business | |
| users first... the three most painful questions they can't answer, that's your MVP | |
| scope") held up and shaped the Wave 1 cut directly. | |
| Global-first positioning, Arab market as the geographic wedge — not an Arab-only | |
| product. Odoo depth + Arabic-language handling is the defensible niche *within* a | |
| global addressable market, not the whole market. | |
| --- | |
| ## 6. Co-founder split | |
| Unchanged in spirit from v1.0; updated to match what the codebase's own seam | |
| (`src/harness/types.ts`) now encodes formally rather than just describing: | |
| **AI Engineer (you) owns:** | |
| - The reasoning loop (`ask.ts`) and confidence scoring | |
| - The Tool Registry (`tools/registry.ts`) — specs + dispatch, deterministic bodies | |
| - The Semantic Layer (§7 — Entity Catalog, Metric Registry, `discover_metrics`) | |
| - Provider SDK integration inside the gateway (`gateway/providers/*.ts`) | |
| - The eval suite (§7.5) | |
| **Software Engineer (co-founder) owns:** | |
| - Auth, multi-tenancy, the encrypted key vault, usage logging | |
| - The MCP host's connection/transport machinery (`mcp/host.ts`) and each connector's | |
| launch/allowlist definition (`mcp/connectors/*.ts`) | |
| - Database, migrations, infra (Docker, Caddy, Hetzner) | |
| - The web app shell and every route around `askCorvus()` | |
| **Together:** the seam itself (`types.ts` — `callModel`, `askCorvus`, `ToolResult`, | |
| `Citation`) — v1.0's "shared schema" instinct was right; what changed is that it's | |
| now a committed TypeScript file both tracks build against and can't change solo, | |
| not a conceptual agreement. Same discipline applies to any future shared file | |
| (flagged explicitly in `docs/ai-track-guide.md §8` — "Don't touch" isn't a wall, | |
| it's "that's a conversation, not a solo edit"). | |
| --- | |
| ## 7. The Semantic Layer — the piece v1.0 gestured at, specified for real | |
| This is new since v1.0 and is the direct answer to two questions asked while | |
| researching what Corvus still needs on the AI side: how Anthropic's own data team | |
| gets self-service analytics to ~95% automated accuracy, and how WrenAI's open | |
| "context layer" makes text-to-SQL governed instead of merely plausible. | |
| ### 7.1 The finding | |
| Corvus already has the right foundation: the model never computes, tools do, every | |
| answer is cited. What's missing is the layer that tells the model **what's queryable | |
| in the first place** before it calls a tool — and this is not a hypothetical gap. | |
| Anthropic's internal data-science team frames analytics accuracy as *"a context and | |
| verification problem, not a code generation issue,"* and names three failure modes: | |
| **concept↔entity ambiguity**, **staleness**, and **retrieval failure**. Independently, | |
| dbt Labs' 2026 benchmark found semantic-layer grounding lifts text-to-SQL accuracy | |
| from 90.0%→98.2% and 84.1%→100% across two frontier models, and that **schema/semantic | |
| failures — not syntax — account for ~81% of text-to-SQL errors**. WrenAI's whole | |
| premise (an open "context layer" — governed models, definitions, and memory beneath | |
| generation) is the open-source expression of the same finding. | |
| Corvus's current `kpi/schema.ts` allowlist (`ENTITIES`, `fields`, `numericFields`) is | |
| **already a primitive semantic layer** — it's just one entity deep (`deals`) and | |
| disconnected from how the model discovers what exists (KPI names are currently a | |
| string baked into a tool description). This already broke once, concretely: the model | |
| hallucinated `pipeline_coverage`, a KPI that was never real, because nothing told it | |
| otherwise. The fix is to grow the existing primitive, not build something new beside it. | |
| ### 7.2 Shape of the change | |
| ``` | |
| BEFORE (Wave 1) AFTER (v2) | |
| ───────────────── ────────── | |
| ask.ts ask.ts | |
| └─ buildTools(tenantId) └─ buildTools(tenantId) | |
| └─ 6 hardcoded KPI names └─ Semantic Layer | |
| as a string in a ├─ Entity Catalog (grows ENTITIES | |
| tool description │ into a real, browsable model) | |
| ├─ Metric Registry (grows KpiDefinition | |
| │ into versioned, governed metrics) | |
| ├─ discover_metrics tool (NEW) | |
| └─ Business Glossary / synonyms (NEW) | |
| ``` | |
| The seam does not change. This is additive: a layer between `buildTools()` and the | |
| model, plus one new tool that lets the model *ask what exists* instead of Corvus | |
| guessing what to cram into a description string. | |
| ### 7.3 Entity Catalog | |
| Extends `EntitySpec` (not a rewrite — the validation logic in `kpi/schema.ts` stays): | |
| `displayName`, `description`, `grain`, `fieldDescriptions`, approved `relationships` | |
| only, and `source: { connectorId, syncedAt }` for provenance. Hand-authored or | |
| Claude-drafted-then-human-approved — **never model-inferred at query time**. Anthropic | |
| tried auto-generating definitions with an LLM from raw tables; it *"produced | |
| plausible-looking definitions that encoded the very ambiguities [they] were trying to | |
| eliminate"* — net negative on their evals. Generate the documentation with Claude; a | |
| human owns the definition. | |
| ### 7.4 Metric Registry | |
| Grows `KpiDefinition` + `kpi/engine.ts` — same structured formula, same allowlisted, | |
| parameterized execution, no free-form SQL, ever. Adds: mandatory `displayName`/ | |
| `description` (the only thing the model sees when discovering a metric), append-only | |
| **versioning** (editing `win_rate` creates `win_rate@2`, doesn't mutate `win_rate@1` — | |
| the direct fix for the staleness failure mode), `synonyms[]` populated from real | |
| question logs, and `asOf` freshness carried into `Citation` — the smallest version of | |
| Anthropic's "provenance footer" that fits Corvus's existing types. | |
| ### 7.5 `discover_metrics` — the highest-leverage single change | |
| Replaces cramming every KPI name into `calculate_kpi`'s description. Modeled on | |
| WrenAI's discover→select→execute flow and Anthropic's "agent is structurally required | |
| to hit the semantic layer first": | |
| ```ts | |
| { | |
| name: "discover_metrics", | |
| description: | |
| "Search the tenant's governed metric catalog by keyword or business concept " + | |
| "before calculating anything. ALWAYS call this before calculate_kpi if you are " + | |
| "not certain a metric name is exact — do not guess a plausible-sounding name.", | |
| inputSchema: { type: "object", properties: { query: { type: "string" } }, required: ["query"] }, | |
| } | |
| ``` | |
| Keyword/synonym match over the Metric Registry — same shape as Mercer's BM25+LSH | |
| entity retrieval, smaller corpus, no vector DB needed at Corvus's current scale. | |
| `calculate_kpi`'s description shrinks back to static; one line added to `ask.ts`'s | |
| system prompt: *"Always try discover_metrics or an existing KPI before calling | |
| query_source directly."* This is what lets Corvus scale past a handful of KPIs | |
| without the tool description growing unboundedly. | |
| ### 7.6 What this explicitly does not include yet | |
| Matching the same "write it down so it stops nagging you" discipline `PLAN.md` uses: | |
| no vector DB (Anthropic's own ablation found raw-corpus retrieval moved accuracy | |
| *less than a point*, even with the right answer present and read); no adversarial- | |
| reviewer sub-agent (+6% accuracy for +32% tokens/+72% latency — a real lever, wrong | |
| priority before a design partner); no dashboard generation / WrenAI's "Deploy" beat; | |
| no cross-connector joins; no automated correction-harvesting agent — a manual weekly | |
| `QueryLog` review does this until question volume justifies automating it. | |
| ### 7.7 Validation — the part that doesn't exist yet | |
| Today Corvus has one signal the reasoning loop works: a 6-question smoke test run | |
| once. Minimum viable eval, sized correctly for a two-person team: | |
| - **~20–30 golden Q&A pairs per tenant/domain** (diminishing returns past a few | |
| dozen per topic, and that ceiling drops with each model generation — don't | |
| over-invest). | |
| - **Pinned to seeded, deterministic data**, never live — the exact trap Anthropic | |
| calls out: *"an eval written against live data goes stale the moment the | |
| underlying number moves."* | |
| - **Wired into `test:integration`**, not optional — any PR touching `kpi/schema.ts`, | |
| `kpi/engine.ts`, or the semantic layer re-runs it. | |
| - **Assert confidence tiers, not just correctness** — a wrong-but-confident answer | |
| is the actual product risk Corvus exists to prevent; the eval set should test for | |
| that directly, not just for the right KPI name. | |
| ### 7.8 Security posture — extends, never relaxes, the existing discipline | |
| `discover_metrics` and the Entity Catalog are read-only, tenant-scoped metadata | |
| lookups — same `tenant_id` scoping as every query in `kpi/engine.ts`, no new attack | |
| surface. Governance stays **compile-time**: `validateConditions` runs before any SQL | |
| is built, re-checked even for a stored definition — the Entity Catalog must preserve | |
| that ordering, never introduce a "generate SQL, then check" path. `query_source` | |
| remains the only raw-query tool, still routed through the same MCP allowlist | |
| discipline as §3's MCP Host section. Growing the catalog makes `query_source` *safer* | |
| to use — it does not add a new way to reach the database. | |
| --- | |
| ## 8. Build sequence | |
| Where §0's cut narrowed *what* to build, this is *what order*, current as of this | |
| version: | |
| 1. ~~Hetzner Wave 1, Model Gateway, MCP host, Tool Registry, reasoning loop~~ — **done**, | |
| verified via `test:integration` and `smoke:kpi-selection`. | |
| 2. **Auth / real multi-tenancy** (SDE) — the one blocker left before "ship to 5 | |
| design partners." | |
| 3. **Entity Catalog schema + migration** — backfill the existing `deals` entity; | |
| zero behavior change, just structure (§7.3). | |
| 4. **Metric Registry additions** — versioning, synonyms, mandatory descriptions on | |
| the 6 seeded KPIs (§7.4). | |
| 5. **`discover_metrics` tool** — the actual reasoning-loop change (§7.5). | |
| 6. **Golden eval set**, wired into CI — before KPI #7 is added, not after (§7.7). | |
| 7. **Ship to 5 design partners.** Their questions decide the second data source, | |
| which KPIs get added next, and which synonyms the Metric Registry actually needs | |
| — not a roadmap guess. | |
| **Not now** (unchanged from `PLAN.md`, still correct): microservices, ClickHouse, | |
| self-hosted observability, MinIO, the Insight Feed, a Chrome extension, a Slack bot, | |
| multi-platform joins, 3+ connectors, Hijri calendar, Arabic generation, a Report | |
| Builder, a public API, Stripe/billing. | |
| --- | |
| ## 9. Risks worth tracking | |
| Carried forward from v1.0, amended where the semantic layer changes the picture: | |
| - **A wrong answer is an instant-churn event** — unchanged as the central design | |
| constraint. The semantic layer (§7) is a direct mitigation, not a new risk: it | |
| closes the concept↔entity ambiguity gap that's the *dominant* real-world failure | |
| mode for this category of system (~81% of text-to-SQL errors, per §7.1). | |
| - **Semantic layer maintenance becomes its own ongoing cost** as KPI/entity count | |
| grows — this is the trade Anthropic's team also made, and their mitigation | |
| (colocate definition + doc changes in the same PR, enforced by review discipline) | |
| is adopted directly in the `corvus-semantic-layer` skill referenced in §7. | |
| - **MCP server quality/maintenance** — still not fully in Corvus's control. Mitigate: | |
| vet + pin versions, keep the CSV/Postgres fallback, never trust a connector's | |
| internal safety switch (§3). | |
| - **No local warehouse while live-call MCP is the only path** — blocks | |
| aggregation-heavy analytics and the Insight Feed. Mitigate: canonical-table sync | |
| is explicitly Wave 2, gated on having a concrete reason to build it. | |
| - **Self-host ops tax** — every service is a potential 2am page. Mitigate: staged | |
| waves (§0), automated backups, no service added without a reason. | |
| --- | |
| *Corvus · Platform Blueprint v2.0 · August 2026 · Confidential* | |
| *Reflects the actual state of the build as of Wave 1 completion + the semantic-layer | |
| scoping pass. Supersedes v1.0 in full; v1.0 remains available as historical record | |
| of the original company-level vision.* | |