Spaces:
Sleeping
title: LedgerShield
emoji: π‘οΈ
colorFrom: blue
colorTo: green
sdk: docker
app_port: 8000
pinned: false
tags:
- openenv
- fastapi
- docker
- agents
- finance
- enterprise-risk
base_path: /web
LedgerShield π‘οΈ
LedgerShield is a stateful, adversarial benchmark for AI agents operating inside enterprise accounts-payable workflows. Instead of asking a model to classify one document, LedgerShield asks it to investigate, unlock hidden evidence, choose controls, withstand pressure, and submit a proof-carrying decision under budget and step limits.
π Documentation hub: See
docs/README.mdfor a guided tour of all documentation, reading paths by role, and a map of what lives where.
Why This Matters
Real-world payment fraud is expensive and operationally messy. In the FBI IC3 2023 report, business email compromise (BEC) generated 21,489 complaints and more than $2.9 billion in reported losses, while total cybercrime losses exceeded $12.5 billion. LedgerShield turns that risk surface into an agent benchmark focused on safe decision-making, evidence quality, and control discipline instead of one-shot classification.
Sources:
What Judges Care About
LedgerShield is built to score well on real-world utility, environment design, task quality, engineering quality, and novelty because the implementation now includes:
| Dimension | What is implemented |
|---|---|
| Real-world utility | Multi-currency invoices, IBAN/SWIFT validation, SOX control modeling, AP inbox triage, campaign fraud, aging-report support |
| Environment design | Stronger PBRS reward shaping, milestone rewards, information-gain bonus, terminated vs truncated, text render(), formal action_space() and observation_space() |
| Task and grader quality | 21 curated benchmark cases, semantic counterfactual scoring, stricter degenerate-submission penalties, generated holdout suites, contrastive benign twins |
| Code quality | Comprehensive docstrings, shared pytest fixtures, dedicated tests for grading/currency/compliance/curriculum, GitHub Actions CI, narrower exception handling, typed internal return contracts |
| Creativity and novelty | Dec-POMDP watchdog mode, dynamic curriculum adaptation, campaign-level fraud reasoning, 16 attack types across identity/document/process/APT categories |
Benchmark At A Glance
| Item | Value |
|---|---|
| Public benchmark cases | 21 curated base cases |
| Task families | 5 (task_a through task_e) |
| Attack types | 16 |
| Default loader behavior | 21 benchmark cases + 24 generated challenge variants = 45 loaded cases |
| Optional generated suites | challenge variants, holdout variants, contrastive benign twins |
| Formal model | finite-horizon POMDP |
| Server runtime | FastAPI / OpenEnv-compatible |
Task coverage
| Task | Count | Focus |
|---|---|---|
| Task A | 4 | proof-carrying invoice extraction, multilingual and multi-currency artifacts |
| Task B | 5 | three-way match, receipt gaps, quantity/tax discrepancies |
| Task C | 4 | duplicate detection, cross-vendor fraud, approval-threshold evasion |
| Task D | 6 | AP inbox/BEC triage, workflow override, CEO fraud, benign vendor updates |
| Task E | 2 | coordinated campaigns and supply-chain-compromise APT scenarios |
What The Agent Must Actually Do
LedgerShield episodes are partially observable. Agents start with visible documents and must use tools and interventions to discover the rest.
Investigation tools:
zoom,get_doc_crop,ocrlookup_vendor,lookup_vendor_history,lookup_policylookup_po,lookup_receipt,search_ledgerinspect_email_thread,compare_bank_account
Interventions:
request_callback_verificationfreeze_vendor_profilerequest_bank_change_approval_chainrequest_po_reconciliationrequest_additional_receipt_evidenceroute_to_procurementroute_to_securityflag_duplicate_cluster_reviewcreate_human_handoff
Final action:
submit_decision
The submission is not just a label. Strong agents are expected to return structured decisions with grounded reason_codes, policy_checks, evidence_map, and task-specific fields like duplicates, campaign signals, discrepancies, or extracted invoice fields.
Agent capability tiers
The inference agent (inference.py) uses a ModelCapabilityProfile that adapts behavior to model strength:
| Tier | Capability score | Plan mode | Repair level | Budget bonus |
|---|---|---|---|---|
| Elite | β₯ 5.0 | LLM-first | partial | +2 investigation, +2 intervention |
| Strong | β₯ 4.5 | hybrid | partial | +1 investigation, +1 intervention |
| Standard | < 4.5 | LLM-first | none | baseline |
The capability profile only adjusts planning depth and budget. It does not hard-snap stronger models onto a deterministic grounded policy.
Smart signal derivation
The agent and server now share improved signal-extraction logic:
- Domain alignment inference β sender domains are compared against vendor-approved domains using token overlap, not just exact match. This catches spoofs like
ceo@acme-corp.comvs approvedacme.com. - Composite risk flags β
bank_override_attemptnow requiresbank_change_languageand a risk amplifier (domain mismatch, callback discouragement, policy override, or urgency). Isolated bank language no longer triggers false fraud flags. - PAY evidence β safe PAY decisions now carry constructive evidence (verified bank, verified sender, cleared duplicates) instead of empty evidence maps. This avoids degenerate-evidence penalties on benign cases.
Upgrade Snapshot
The benchmark upgrade work is reflected in the codebase across five phases:
| Phase | Highlights |
|---|---|
| Phase 1: Real-world utility | server/currency_engine.py, server/compliance_engine.py, richer payment artifacts, aging-report support |
| Phase 2: Task and grader quality | 21 curated cases, semantic counterfactual grading, tighter degenerate penalties, generated holdouts |
| Phase 3: Environment design | SHAPING_SCALE=0.35, INFO_GAIN_BONUS=0.08, milestone rewards, Gymnasium-style truncation semantics, text rendering, formal spaces |
| Phase 4: Code quality | docstrings across core modules, tests/conftest.py, CI workflow, TypedDict internal returns |
| Phase 5: Creativity and novelty | Dec-POMDP watchdog mode, curriculum adaptation, 16-attack library, exploration bonus integrated into step() |
Recent patch-level changes
| Change | Where | Why |
|---|---|---|
DEGENERATE_EVIDENCE_CAP applied correctly |
server/grading.py |
Bug fix: empty evidence now correctly receives cap value instead of collapsing to 0.0 |
| Model capability profiles and tiered agent behavior | inference.py |
Agent adapts investigation/repair strategy based on model tier (elite/strong/standard) |
Composite bank_override_attempt signal |
server/tools.py, task_c_guardrails.py, task_d_guardrails.py |
Bank override flag now requires bank-change language plus a risk amplifier β reduces false positives |
| Domain alignment via token overlap | server/tools.py, inference.py |
Catches spoofs where sender domain shares tokens with vendor name but is not an exact match |
| Constructive PAY evidence maps | task_c_guardrails.py, task_d_guardrails.py |
Safe PAY decisions carry verified-bank / cleared-duplicates evidence instead of empty maps |
| Per-model capability profiles in live comparison | compare_models_live.py |
Records model tier, capability score, and monotonic strength checks alongside scores |
pytest config in pyproject.toml |
pyproject.toml |
Asyncio mode, markers, deprecation-warning filters |
Benchmarking Story
LedgerShield is not just a server. It includes a full evaluation stack:
benchmark_report.pyscores the public benchmark, generated holdout suites, and contrastive adversarial/benign pairs.compare_models_live.pyruns live head-to-head evaluations with per-model capability profiles and writes per-case debug traces including monotonic strength ordering checks.live_model_comparison_debug/stores action traces, planning traces, score breakdowns, and system state snapshots for diagnosis./leaderboardand/benchmark-reportexpose report artifacts through the API when generated.
Latest local live comparison
Live comparison numbers should be treated as generated artifacts, not hardcoded documentation. Run:
python compare_models_live.py \
--models gpt-3.5-turbo,gpt-4o,gpt-5.4 \
--output live_model_comparison.json
Then inspect:
live_model_comparison.jsonfor summary metrics, per-case scores, model profiles, and ordering checkslive_model_comparison_debug/<model>/for per-case traces, submissions, and score breakdowns
π Models Evaluated
gpt-3.5-turbo(Standard Tier)gpt-4o(Strong Tier)gpt-5.4(Elite Tier)
π Benchmark Results
Generated on April 9, 2026 (IST) from live_model_comparison.json.
| Model | Tier | Capability | Average Score | Success Rate | Min Score | Max Score | API Calls |
|---|---|---|---|---|---|---|---|
gpt-3.5-turbo |
standard | 3.2 | 0.7009 | 38.1% | 0.01 | 0.99 | 63 |
gpt-4o |
strong | 4.6 | 0.8663 | 81.0% | 0.43 | 0.99 | 65 |
gpt-5.4 |
elite | 5.4 | 0.9305 | 100.0% | 0.88 | 0.99 | 64 |
β Failed Case Summary
gpt-3.5-turbo (13 failures)
- CASE-B-005
- CASE-C-001 β CASE-C-004
- CASE-D-001 β CASE-D-006
- CASE-E-001, CASE-E-002
gpt-4o (4 failures)
- CASE-B-002
- CASE-B-004
- CASE-C-002
- CASE-E-001
gpt-5.4 (0 failures β
)
π Key Insights
Clear performance hierarchy:
gpt-5.4>gpt-4o>gpt-3.5-turboFrontier performance gap:
gpt-5.4outperformsgpt-4oby:- +0.0642 average score
- +19.0% success rate
Reliability comparison:
gpt-5.4: 100% success (fully robust)gpt-4o: Strong but inconsistent on harder casesgpt-3.5-turbo: Struggles with complex workflows
Score stability:
gpt-5.4: High minimum score (0.88)gpt-4o: Occasional dips (0.43)gpt-3.5: Near-zero failures
Fair evaluation:
- All models used ~64 API calls β no compute bias
π§ Conclusion
- The benchmark is well-calibrated and clearly differentiates model capabilities.
gpt-5.4demonstrates state-of-the-art performance, achieving:- Highest accuracy
- Perfect success rate
- Strong consistency across all tasks
gpt-4ois competitive but not fully reliable at the frontier level.gpt-3.5-turbois not suitable for complex structured tasks.
The repo keeps the generated artifact and full trace folder so readers can verify the claim instead of trusting a hand-written summary.
Published benchmark metadata in openenv.yaml records meaningful public-vs-holdout separation:
| Agent | Public mean | Holdout mean | Holdout consistent pass rate |
|---|---|---|---|
| Deterministic baseline | 0.9674 | 0.6649 | 0.6190 |
| Published external LLM agent | not listed | 0.3847 | 0.2222 |
That gap is deliberate: the benchmark looks easy on clean public cases and much harder on generated holdouts, adversarial variants, and expert Task E scenarios.
Quick Start
1. Install
git clone https://github.com/BiradarScripts/Meta-s-LedgerShield.git
cd Meta-s-LedgerShield
python -m venv .venv
source .venv/bin/activate
pip install -e .
pip install -r requirements.txt
2. Start the environment server
python -m server.app
The API comes up on http://127.0.0.1:8000 by default.
3. Run the submission-safe agent
export API_BASE_URL="https://api.openai.com/v1"
export MODEL_NAME="gpt-5.4"
export HF_TOKEN="your_token"
export ENV_URL="http://127.0.0.1:8000"
python inference.py
4. Generate a benchmark report
python benchmark_report.py --format markdown
5. Run live model comparisons
export OPENAI_API_KEY="your_api_key"
export API_BASE_URL="https://api.openai.com/v1"
export ENV_URL="http://127.0.0.1:8000"
python compare_models_live.py \
--models gpt-3.5-turbo,gpt-4o,gpt-5.4 \
--output live_model_comparison.json
6. Validate locally
python -m pytest tests/ -q
bash validate-submission.sh
If openenv is installed in your environment, you can also run:
openenv validate
Documentation
| Document | What it covers |
|---|---|
docs/README.md |
docs landing page and reading paths |
docs/index.md |
benchmark overview, quick start, and core concepts |
docs/tasks.md |
task families, outputs, scoring, and case catalog |
docs/api-reference.md |
REST endpoints, payloads, response envelopes, and action contracts |
docs/architecture.md |
system design, hidden state, reward flow, grading, and evaluation pipeline |
docs/development.md |
setup, tests, CI, and detailed repo/file map |
docs/deployment.md |
local, Docker, HF Space, and environment configuration guidance |
Recommended reading paths:
- Benchmark judge or first-time reader:
docs/index.md->docs/tasks.md->docs/architecture.md - Agent builder:
docs/tasks.md->docs/api-reference.md->docs/development.md - Contributor:
docs/development.md->docs/architecture.md - Operator:
docs/deployment.md->docs/api-reference.md
Repository Structure
Top level
Meta-s-LedgerShield/
βββ README.md
βββ CHANGELOG.md
βββ docs/
βββ server/
βββ tests/
βββ inference.py
βββ inference_improved.py
βββ inference_llm_powered.py
βββ task_c_guardrails.py
βββ task_d_guardrails.py
βββ benchmark_report.py
βββ compare_models_live.py
βββ compare_all_models.py
βββ llm_utils.py
βββ llm_judge_grader.py
βββ models.py
βββ client.py
βββ ledgershield_env.py
βββ openenv_compat.py
βββ openenv.yaml
βββ pyproject.toml
βββ requirements.txt
βββ Dockerfile
βββ validate-submission.sh
Important files at a glance
| Path | Purpose |
|---|---|
server/environment.py |
main OpenEnv environment loop, reward shaping, truncation semantics, rendering |
server/world_state.py |
hidden/public state, artifact scheduling, campaign context, pressure resistance |
server/grading.py |
task rubrics, semantic counterfactual scoring, degenerate penalties |
server/trajectory_grading.py |
investigation, intervention, calibration, efficiency, and outcome scoring |
server/attack_library.py |
16 adversarial attack templates |
server/case_factory.py |
challenge, holdout, and benign-twin generation |
server/tools.py |
investigation tool implementations, email thread parsing, domain alignment inference |
server/currency_engine.py |
FX conversion, IBAN/SWIFT checks, currency mismatch detection, aging reports |
server/compliance_engine.py |
SOX-style AP control evaluation |
server/curriculum.py |
dynamic difficulty adaptation |
server/dual_agent_mode.py |
Dec-POMDP watchdog/auditor mode |
benchmark_report.py |
public benchmark + holdout + contrastive reporting |
compare_models_live.py |
live multi-model evaluation with capability profiles and debug artifacts |
inference.py |
submission-safe agent with ModelCapabilityProfile tiers and evidence-grounded output |
inference_improved.py |
experimental improved agent entrypoint |
inference_llm_powered.py |
richer LLM-powered agent used for debugging and comparisons |
task_c_guardrails.py / task_d_guardrails.py |
grounded output sanitizers with composite signal detection and PAY evidence construction |
llm_utils.py |
JSON parsing and completion helpers for LLM workflows |
llm_judge_grader.py |
optional LLM-as-judge grading experiments |
models.py |
shared dataclasses and Pydantic reward model |
.github/workflows/ci.yml |
pytest, Docker build, and metadata validation in CI |
For the full file-by-file map, see docs/development.md.
Current Engineering Status
- Core environment upgrades from Phases 1 through 5 are implemented in code.
- Patch-level fixes applied: correct
DEGENERATE_EVIDENCE_CAPin grading, composite bank-override signal, domain-alignment token overlap, constructive PAY evidence in guardrails. - The agent (
inference.py) now usesModelCapabilityProfiletiers (elite/strong/standard) that adapt planning mode, repair level, and budget bonuses. compare_models_live.pyrecords per-model capability profiles and includes monotonic strength ordering checks.- The repo includes 21 curated benchmark cases and generated challenge/holdout tooling.
- CI is present via GitHub Actions with pytest config now in
pyproject.toml. - The test suite includes API smoke, grading, environment, inference, inference-runtime, compliance, currency, curriculum, and guardrail coverage.
- The environment remains submission-compatible through
inference.py.
Safety Note
LedgerShield is a benchmark and simulation environment. It models payment-integrity risk and enterprise controls, but it is not a production fraud platform and should not be used to approve or block real payments without independent controls, audit, and governance.