Jaideep_IITM / docs /04_engineering_justification.md
jaideepsrivastava's picture
Initial capstone EB069 submission
81fb303
|
Raw
History Blame Contribute Delete
9.06 kB
# Engineering & Product Justification
This document covers the design tradeoffs the rubric grades on
("engineering judgment, reliability, explainability, safety-first
design, and practical usefulness in real workflows β€” not unnecessary
complexity"). Each section explains *why* the design is what it is,
what was rejected, and what the deployment assumptions are.
---
## 1. Why three specialist agents instead of one mega-agent
**Choice.** The smart path uses three `create_agent(...)` instances
with disjoint tool sets:
| Specialist | Tool surface | Why disjoint? |
|---|---|---|
| FAQ | `kb_search`, `branch_locator` | product / fee / branch questions |
| Calculator | `kb_search`, `emi_calculator`, `fx_rate_lookup`, `eligibility_estimator` | numeric advisory |
| Dispute | `kb_search` | triage only β€” must not commit to a verdict |
**Rejected:** a single agent with all five tools.
**Why disjoint helped:**
- **Lower wrong-tool-call rate.** Empirically, an LLM with 5 tools
guesses the wrong one more often than three LLMs with 1–4 tools each.
- **Smaller blast radius if one tool is broken.** A failure in
`fx_rate_lookup` should not degrade the FAQ flow.
- **Maps onto how a real bank contact centre is organised** β€” product
desk vs cards desk vs fraud desk. A reviewer with banking domain
context can read the architecture without translating.
- **Each role gets a tighter system-prompt append** ("dispute role *does
not* commit to a verdict") so the LLM is more constrained in the
zone where mistakes are most expensive.
---
## 2. Why a rule-based safety gate runs *before* the LLM
**Choice.** `safety.classify_intent(...)` runs on the raw user input
before any LLM call. If it returns `refuse` or `escalate`, the LLM is
never invoked and the customer-facing message is built from
`safety.build_refusal_message(...)`.
**Rejected:** asking the LLM to enforce the policy via the system
prompt only.
**Why this is the right tradeoff for banking:**
- **The LLM is a generator, not a guard.** Even with the strictest
prompt, an LLM will occasionally produce a "yes, here's how to ..."
for a phrasing that wasn't in its training distribution.
- **Auditability.** A regulator-facing review of "why did the agent
refuse this customer" should land on a single short Python file
(`safety.py`) β€” not a 1500-token prompt mixed with tools.
- **Latency.** Refusals run in < 1 ms vs 2–6 s for an LLM-mediated
refusal. For high-volume policy-violation prompts this matters.
- **Defence in depth.** Even with the rule-based gate, the smart-mode
prompt *also* includes the policy rules, so the LLM is still
constrained on edge cases that slip past the rules.
This is the single most important design decision in the project.
---
## 3. Why a separate `safety.py` module (the one extra file)
The recommended structure allows up to 2 extra Python files. We use **1**.
**Choice.** Keep PII redaction, refusal policy and intent classification
in their own module rather than embedded in `agent_system.py`.
**Why a separate file is worth its weight:**
1. The same policy is reused by `baseline_respond` *and* `smart_respond`.
Co-locating it with the orchestrator would force one of those
paths to import the other.
2. It is unit-testable in isolation. A regulator-facing review can
read one file without scanning a 1000-line orchestrator.
3. The PII-redaction layer is a horizontal concern β€” it is also called
from logging, from the conversation-memory formatter, and from the
evaluation harness. Pushing it into `agent_system.py` would either
force long imports or duplication.
4. The file is intentionally rule-based β€” no LLM, no I/O, no network.
That makes it the kind of file that should *not* change unless the
bank's policy changes.
---
## 4. Why FAISS in-memory rather than Chroma / Pinecone
**Choice.** FAISS, in-memory, cached at process scope with `@lru_cache(1)`.
**Rejected:** Chroma (extra disk dependency), Pinecone (extra service +
secret), pgvector (overkill for ~30 KB of policy text).
**Reasoning:**
- The KB is **small and read-only** (5 markdown files -> 51 chunks).
Disk persistence and metadata-filtered search are unnecessary.
- The chat is **single-process Streamlit**; we don't need a shared
index across instances.
- FAISS has no extra service, no extra secret, and no extra failure
mode beyond "OPENAI_API_KEY missing -> degrade to baseline".
If the corpus grows past a few thousand chunks or we need
metadata-filtered search (e.g. only "credit_card" docs), Chroma is the
next stop. The retrieval interface is small enough that this is a
half-day swap.
---
## 5. Why deterministic tools rather than third-party APIs
The four numeric tools (`emi_calculator`, `fx_rate_lookup`,
`eligibility_estimator`, `branch_locator`) all run **locally** with
either pure math or a small bundled dataset.
**Why this matters here:**
- **Reproducibility.** Evaluation CSVs and demo logs must be replayable
without external rate-limit or downtime surprises.
- **Auditability.** Every number is computed from the formula or the
documented threshold; we can show *why* a number came out the way
it did.
- **Honesty.** A deterministic tool with a clear `_label_ as indicative`
is more honest than a live FX feed that the customer might mistake
for a deal rate.
For production, `fx_rate_lookup` would wrap the bank's internal
indicative-rate service; `branch_locator` would call the official
locator API. The interfaces are designed to make those swaps small.
---
## 6. Why rule-based routing rather than LLM-as-router
**Choice.** `BankingAdvisor.detect_route(...)` is a keyword classifier.
**Rejected:** running an LLM on every turn just to decide which agent
to invoke.
**Why this works in our domain:**
- The routing keyword set is **finite and stable** (fees, EMI, FX,
dispute, …). Banking customers are not ambiguous about what they
want β€” they typically use the product noun.
- A wrong routing decision is *more harmful* than a 5-second decision
delay would be β€” sending a fraud question to the FAQ agent loses the
case-id and the helpline nudge.
- The rule is small (β‰ˆ 30 lines) and easy to extend; we can revisit
with an LLM-router *only* if rule-set churn becomes high.
When the LLM mis-handles the routing hint at the prompt level, that's
captured in the prompt-comparison metrics under `actual_route` (it's
not the *router* that's wrong, but the specialist's understanding of
what to do β€” also a real failure mode we want to surface).
---
## 7. Deployment assumptions & limitations
### Local
- Python β‰₯ 3.10. (Smoke-tested on 3.14 and 3.11.)
- The LangChain v1 surface (`from langchain.agents import create_agent`)
changed in 1.0.0; pinning is in `requirements.txt`.
- `OPENAI_API_KEY` is the only required secret; without it the app
runs in baseline mode and degrades visibly (a sidebar warning).
- No GPU required.
### Hugging Face Spaces
- Docker SDK build, exposes 8501.
- Non-root user.
- Build time on a free Space: β‰ˆ 2 minutes (CPU image, β‰ˆ 250 MB).
- Cold-start latency: first request triggers the FAISS build (~1 second
for 51 chunks + embeddings). Subsequent requests use the cached index.
- The bank's real production deploy would replace `gpt-4o-mini` with
whatever provider has been approved and would route through an
internal gateway, not directly to OpenAI.
### Limitations
1. **Single-tenant feedback profile** β€” the keyed user_id is a
Streamlit session id, which means two users on the same device share
a profile. Acceptable for a demo; in production this needs an
authenticated user id.
2. **No durable conversation transcripts** β€” by design (compliance).
If the bank ever wants transcripts they must be stored in the
bank's audited datastore, not a JSONL file.
3. **No live data** β€” the agent has no balance / statement / card-status
tool. This is **a feature**, not a missing piece: live data needs
an authenticated session.
4. **English only.** Pattern files would need a localised twin.
5. **Indicative tool outputs only** β€” the FX and eligibility tools are
not deal rates / sanctions and never will be from this advisor.
6. **No long-running session recovery** β€” if the Streamlit process
dies, in-session memory is lost; the feedback profile survives.
---
## 8. What we'd build next (engineering ranking)
| # | Item | Why now |
|---|---|---|
| 1 | CI eval-canary on the safety test-set | catches a regression like S4 instantly |
| 2 | Localised refusal templates | first user complaint is "I'm not English-only" |
| 3 | LLM-as-judge eval pass | heuristic flags miss semantic drift |
| 4 | Real branch / FX integrations | move illustrative datasets out before launch |
| 5 | Authenticated handoff token | replace `case_id` strings with a backend-signed token |
| 6 | Multi-tenant feedback profile | move `data/user_feedback.json` to a real KV store |
Each item is < 1 sprint of work. The order is the order I'd push them
to product.