# Testing Guide Three layers: automated `pytest` (unit + integration, including auth/role coverage), manual exploration through the officer console and client app, and a live smoke test script for post-deploy verification of the full officer/client flow. ## 1. Automated tests (pytest) ```bash source venv/bin/activate pip install -r requirements-dev.txt pytest # everything pytest tests/unit -q # unit tests only (fast, no live DB writes) pytest tests/integration -q # integration tests (hits the real MySQL DB) ``` ### Unit tests (`tests/unit/`) Small synthetic fixtures, no database: - `test_config.py` — fail-loud `.env` validation (including the auth-related vars: `SEED_OFFICER_EMAIL`, `SEED_OFFICER_PASSWORD`, `SESSION_EXPIRE_MINUTES`) - `test_features.py` — the shared feature-engineering module (`app/features.py`) - `test_train_split.py` — time-respecting split logic - `test_scoring.py` — the scoring service (`app/scoring.py`) against the real trained model, using hand-built transactions - `test_labeled_fixtures.py` — replays `tests/fixtures/labeled_transactions.json` (58 hand-labeled synthetic transactions covering obvious fraud and legitimate patterns across simulated dates) through the real scoring service and asserts ≥90% overall accuracy and ≥90% fraud recall. Regenerate the fixture set with: ```bash python -m scripts.generate_fixtures ``` ### Integration tests (`tests/integration/`) `test_api.py` exercises every endpoint through FastAPI's `TestClient` against the real, configured MySQL database, covering: - Auth: login (correct/wrong password/unknown email — all the same generic error), logout invalidating the session, `/api/auth/me` - Role boundaries: an officer route rejects a client (403) and an anonymous caller (401), and vice versa - Officer flows: create client, list clients, client detail (transactions + alerts), cross-client alert queue, approve/reject/dismiss (reject reverses the transaction and restores the client's balance), dashboard/insights/model-performance stats, `/api/predict` + `/api/transactions/batch` - Client flows: submit a transaction, can't overdraw, own transaction history, own alerts - **Data isolation**: two client fixtures (`two_clients` — see below) confirm each client sees only their own transactions/alerts, and that a client hitting an officer-only route for another client's detail gets 403 - All HTML pages return 200 with the right content type **Fixtures** (`conftest.py`): `officer_identity` logs in as the seed officer; `two_clients` creates two client accounts through the real `/api/officer/clients` endpoint (not inserted directly into the DB) and logs into both, giving you `(client_a, client_b)` — each a `{"cookies": {...}, "user": {...}, "client_id": ...}` dict. Because a single `TestClient` only has one cookie jar, tests pass `cookies=identity["cookies"]` explicitly per request to act as several logged-in identities at once without re-running the (model-loading) app lifespan per identity. **Isolation note**: the configured DB user does not have `CREATE DATABASE` privilege, so these tests can't spin up a literal second schema. Instead, `conftest.py` wraps each test in a SQL SAVEPOINT-nested transaction that is *always rolled back* at the end of the test (SQLAlchemy's documented "join a Session into an external transaction" pattern) — route handlers' internal `db.commit()` calls only release the savepoint, never the outer transaction, so nothing a test writes is ever actually persisted. The one exception is the seed officer row itself: the app's startup lifespan creates it (idempotently) against the *real* engine, not the test's savepoint session — that's intentional, matching production behavior, and safe to re-run. If your database user does get `CREATE DATABASE` privilege later, point `DATABASE_URL` at a dedicated `..._test` database instead for a literal second-schema setup — the tests themselves don't need to change. ## 2. Manual exploration Start the API (`uvicorn app.main:app --host 0.0.0.0 --port $APP_PORT`) and open `http://localhost:$APP_PORT/login`. **As the officer** (sign in with `SEED_OFFICER_EMAIL`/`SEED_OFFICER_PASSWORD` from `.env`): 1. `/officer/clients` → "+ Create client" — name, email, starting balance, account type. Copy the one-time temp password shown. 2. `/officer/alerts` — once a client has submitted a flagged transaction, click "Review" to expand the SHAP breakdown, then Approve, Reject & reverse, or Dismiss it. Reject asks for confirmation and reverses the transaction (excluded from the client's balance). 3. `/officer/insights` and `/officer/model` — the EDA and model-selection story behind the score. **As a client** (log out, sign in with the email + temp password from step 1): 1. `/client/transactions/new` — pick a type, amount, destination account, and date/time. 2. **To reproduce the classic fraud pattern**: enter your full current balance as the amount for a Transfer or Cash out. This drains the account to exactly zero — the same signature the model learned from real PaySim fraud rows. The result reads "being reviewed for your security," not a raw probability. 3. `/client/alerts` — read-only, plain-language status; no admin actions available to a client. Every client-submitted transaction flows through the same `transactions` → `fraud_predictions` → `fraud_alerts` tables as everything else — `/officer/alerts` and `/client/alerts` are just two different filtered views over one table, not separate systems. ## 3. Live smoke test (`scripts/test_live.py`) Post-deploy (or post-restart) verification against a **running** instance of the API — local or behind the reverse proxy: ```bash # against a local dev server python -m scripts.test_live --base-url http://127.0.0.1:8811 # against the live deployment python -m scripts.test_live --base-url https://bank-fraud.hdev.rw ``` What it does: 1. Checks `/health`. 2. Logs in as the seed officer, creates two clients, and submits three transactions across different simulated dates on client A's account (a legit payment, a legit partial transfer, and a full-balance drain) through the real `/api/client/transactions` endpoint — printing predicted vs. expected risk tier for each. 3. Confirms role separation live: the officer sees client A's alert in the cross-client queue; client A sees exactly their own alert; client B sees none of it; client B gets a 403 when hitting the officer-only client-detail route for client A. 4. Replays the full `tests/fixtures/labeled_transactions.json` fixture set through the officer-only `/api/predict` endpoint, printing predicted vs. expected risk tier per transaction and a final accuracy summary. Exits `0` only if accuracy stays at or above 90% (the same bar `test_labeled_fixtures.py` enforces) *and* every role-separation check passed; exits `1` otherwise — safe to wire into a post-deploy CI/CD gate. **Note**: this script creates two real client accounts + a few transactions/predictions/alerts against whatever database the target server is configured against — clean those up afterward if you're running it against a database other test suites also read from, e.g.: ```python from app.db.session import session_scope from app.db.models import Transaction, FraudPrediction, FraudAlert, Client, User, Session as SessionModel with session_scope() as db: db.query(FraudAlert).delete() db.query(FraudPrediction).delete() db.query(Transaction).delete() db.query(SessionModel).delete() db.query(Client).delete() db.query(User).filter(User.role == "client").delete() # keeps the seed officer ```