financebench-viewer / README.md
timchen0618's picture
Add E2E v3 + rawtext views
c34313e verified
|
Raw
History Blame Contribute Delete
5.87 kB
---
title: FinanceBench Viewer
emoji: 🏦
colorFrom: green
colorTo: gray
sdk: static
pinned: false
license: cc-by-nc-4.0
---
# FinanceBench Viewer
Static viewer for the open-source [FinanceBench](https://github.com/patronus-ai/financebench)
financial-QA dataset ([HF dataset](https://huggingface.co/datasets/PatronusAI/financebench)).
Tabs:
- **πŸ“š Corpus** β€” every document in the corpus (368 SEC filings / earnings PDFs).
A searchable dropdown selects a document by title (`doc_name`); the panel shows
its metadata and lazily loads extracted six-page text chunks with chunk
navigation. Direct bundled-PDF and official-source links remain available.
This avoids the Hugging Face Xet PDF iframe CSP that broke inline rendering.
- **❓ Eval** β€” the 150 annotated questions. Each shows the question, gold answer,
justification, reasoning type, question type, company, the source `doc_name`
(click to jump to it in the Corpus tab), and every evidence span (page number,
extracted evidence text, and the full-page extract).
- **E2E Structures v3** β€” the five constructed scaffold shapes for each
question's supporting filing chunks.
- **E2E+Rawtext Structures** β€” the same five generated shapes reused by the
overlay run, with raw filing documents mounted as its sixth fallback source.
- **c1 Closed-book**, **c2 With-docs**, **c6 Agentic-DCI**, **Naive-search**, and
**E2E v3**, and **E2E v3 + rawtext** β€” full-run
records with search, correctness/missing filters, qid navigation, judge output,
failure/stop/token details, and expandable agent trajectories when present.
- **Compare** β€” prejoined, side-by-side predictions for all six retained runs,
with filters for disagreements, any missing response, and questions where only
E2E is correct.
Question selection follows the same qid across Eval, run, Compare, and E2E
Structures tabs. Agent trajectories pair each tool call with its result, collapse
long outputs, and keep E2E supporting filings closed until requested.
## Full-run results
Scores use the full 150-question denominator. Missing or unanswered responses
remain visible in the viewer and count as incorrect.
| Run | Correct / 150 | Score | Answered | Missing |
|---|---:|---:|---:|---:|
| c1 Closed-book | 56 / 150 | 37.33% | 134 | 16 |
| c2 With-docs | 113 / 150 | 75.33% | 137 | 13 |
| c6 Agentic-DCI | 122 / 150 | 81.33% | 150 | 0 |
| Naive-search | 102 / 150 | 68.00% | 124 | 26 |
| E2E v3 | 121 / 150 | 80.67% | 149 | 1 |
| E2E v3 + rawtext | 123 / 150 | 82.00% | 150 | 0 |
The c6 row is the corrected rawtext rerun from AML job
`gifted_helmet_6vkx7rqq1f`; all 150 questions have responses.
## Data
| File | Contents |
|---|---|
| `financebench_open_source.jsonl` | 150 annotated questions (Eval tab) |
| `financebench_document_information.jsonl` | 360 per-document metadata rows |
| `corpus_index.json` | 368 corpus docs = PDF listing β‹ˆ metadata (built by `scripts/build_data.py`) |
| `corpus_text/` | 368 compressed filing-text shards covering 9,174 chunks |
| `pdfs/*.pdf` | 368 source PDFs, ~674 MB, git-LFS tracked |
| `e2e_structures_v3/` | 150 compressed per-question E2E v3 structure shards |
| `runs/manifest.json` | Run registry, scores, coverage, colors, and provenance |
| `runs/<slot>/index.json` | Search/filter index for one run |
| `runs/<slot>/records/<qid>.json` | Normalized per-question run record |
| `compare/index.json` | Comparison summary and qid/filter index |
| `compare/records/<qid>.json` | One prejoined shard containing all six runs |
`corpus_index.json` is regenerated from the two source files + the `pdfs/`
listing with:
```bash
python scripts/build_data.py
python3 scripts/build_corpus_text.py
python3 scripts/build_e2e_structures.py \
--dataset financebench \
--unified /mnt/ramdisk/blobstore/timchen0618/data/eval/financebench/unified/test_with_chunks.unified.jsonl \
--scaffolds-dir /home/azureuser/projects/information-scaffolds/outputs/e2e_runs/new-datasets-full-20260711/financebench/named-outputs/scaffolds_dir
```
Run/compare shards are regenerated with:
```bash
python3 scripts/build_runs.py
```
`scripts/build_runs.py` streams the JSONL inputs, joins each run to every eval
qid, validates the exact scores and coverage above, removes bulky repeated
prompts, caps retained judge/event strings, and limits serialized trajectories
to roughly 256 KiB per record with an explicit truncation event.
### Run provenance
Gold answers and metadata come from `financebench_open_source.jsonl`
(`financebench_id` β†’ qid, `answer` β†’ gold).
| Run | Responses | Judges |
|---|---|---|
| c1 | `/mnt/tmp/financebench-c1-predictions.jsonl/predictions` | `/mnt/tmp/financebench-c1-evaluated.jsonl/evaluated` |
| c2 | `/mnt/tmp/financebench-c2-predictions.jsonl/predictions` | `/mnt/tmp/financebench-c2-evaluated.jsonl/evaluated` |
| c6 | `/tmp/baseline-pipeline-runs/gifted_helmet_6vkx7rqq1f/c6_predictions/predictions` | `/tmp/baseline-pipeline-runs/gifted_helmet_6vkx7rqq1f/c6_evaluated/evaluated` |
| Naive-search | `/mnt/tmp/financebench-naive-predictions.jsonl/predictions` | `/mnt/tmp/financebench-naive-evaluated.jsonl/evaluated` |
| E2E v3 | `/home/azureuser/projects/information-scaffolds/outputs/e2e_runs/new-datasets-full-20260711/financebench/named-outputs/predictions/predictions` | `/home/azureuser/projects/information-scaffolds/outputs/e2e_runs/new-datasets-full-20260711/judges/financebench/named-outputs/judged/judged` |
| E2E v3 + rawtext | AML `silly_salt_1vr6cy7bw8` predictions | Its canonical FinanceBench evaluation |
## Local dev
```bash
python -m http.server 8000 # then open http://localhost:8000/
```
Serves statically β€” no build step. Push to the HF Space remote to deploy.
Source: patronus-ai/financebench (150-example open-source sample; the full
FinanceBench has 10,231 questions). Dataset license: CC-BY-NC-4.0.