phantom-wiki-viewer / README.md
timchen0618's picture
Add E2E v3 + rawtext views
12ea280 verified
|
Raw
History Blame Contribute Delete
7.9 kB
---
title: PhantomWiki Viewer
emoji: πŸ‘»
colorFrom: purple
colorTo: gray
sdk: static
pinned: false
license: bsd-3-clause
---
# PhantomWiki Viewer
Static viewer for [**PhantomWiki**](https://github.com/kilian-group/phantom-wiki)
([paper](https://arxiv.org/abs/2502.20377),
[HF dataset](https://huggingface.co/datasets/kilian-group/phantom-wiki-v1)) β€” an
on-demand benchmark that generates synthetic "universes" of people (family +
friendship graphs) with Wikipedia-style articles, plus multi-hop reasoning
questions over them.
A **universe selector** (top of the sidebar) switches between the 9 pre-generated
universes: sizes **50 / 500 / 5000** Γ— seeds **1 / 2 / 3**. Each universe has one
document per person and **exactly 500 questions** (matching the paper's "500 per
instance"). The larger totals shown on Hugging Face β€” 21,269 rows β€” simply count
all three configs together: `database` (9) + `question-answer` (**4,500 = 500 Γ— 9**)
+ `text-corpus` (16,760 documents).
Corpus and evaluation tabs:
- **πŸ“š Corpus** β€” the person documents. A dropdown selects a person **by name**
(filter, or prev/next); the article (family, friends, attributes) is rendered
from its markdown.
- **❓ Eval** β€” the questions. A dropdown selects a **question** (search + filter
by **difficulty** and **question type**). Each shows the question, the gold
answer (when the answer names people, each is a clickable chip that jumps to
that person's article in the Corpus tab; some questions answer with an
attribute value β€” a date, count, or job β€” and have no document), the abstract
**question template**, and the **ground-truth Prolog reasoning** chain.
Full-run tabs are also available for the information-scaffolds canonical
`depth_20_size_5000_seed_1` split:
**E2E Structures v3** shows the five constructed scaffold shapes for every
question's supporting people in that same canonical universe.
**E2E+Rawtext Structures** shows the exact same cached shapes used by the
overlay run; raw universe documents were mounted as a sixth fallback source.
| Tab | Correct / 500 | Score | Answered / 500 | Coverage |
|---|---:|---:|---:|---:|
| **c1 Closed-book** | 53 | 10.60% | 465 | 93.00% |
| **c2 With-docs** | 283 | 56.60% | 500 | 100.00% |
| **c6 Agentic-DCI** | 325 | 65.00% | 484 | 96.80% |
| **Naive-search** | 209 | 41.80% | 453 | 90.60% |
| **E2E v3** | 294 | 58.80% | 489 | 97.80% |
| **E2E v3 + rawtext** | 290 perfect | 68.77% mean F1 | 487 | 97.40% |
Each run tab provides question/gold search, a correct/incorrect/missing filter,
qid navigation, score and coverage cards, the prediction and parsed judge
answer, judge rationale/confidence, stop/failure details, token/tool/turn
counts, and expandable compact agent events. **Compare** prejoins all six
runs and adds filters for correctness disagreement, any missing answer, and
questions where only E2E is correct.
Question selection follows the same UUID across Eval, run, Compare, and E2E
Structures tabs. Tool calls are paired with their results, long outputs start
collapsed, gold answers use compact pills, and E2E supporting people are closed
until requested.
Entering any run, Compare, or E2E Structures tab selects the canonical
**size 5000 Β· seed 1** universe. Returning to Eval stays in that universe so the
same UUID remains selected; Corpus restores the prior universe after run or
Compare when available. Person chips inside run results open that person's
Corpus article in the canonical size5000/seed1 universe.
## Data layout
Built from the `text-corpus` + `question-answer` configs of
`kilian-group/phantom-wiki-v1` (the `database`/Prolog config is not used):
| File | Contents |
|---|---|
| `sets.json` | universe manifest (label, doc/question counts, difficulty + type facets) |
| `corpus_<u>.json` | that universe's documents: `[{title, article}]` (inline) |
| `eval_<u>.json` | that universe's questions: `{id, question, answer, supporting_titles, template, prolog, meta{Difficulty, Type}}` |
| `runs/manifest.json` | full-run registry, provenance, scores, coverage, and canonical scoring note |
| `runs/<slot>/index.json` | 500 lightweight qid summaries for one run |
| `runs/<slot>/records/<qid>.json` | normalized prediction, judge, execution metadata, and compact events |
| `compare/index.json` | 500 precomputed cross-run filter summaries |
| `compare/records/<qid>.json` | gold/metadata plus all six normalized run results |
| `e2e_structures_v3/` | 500 compressed per-question E2E v3 structure shards |
The nine Corpus/Eval universes remain inline. Full-run data is sharded by qid
so the static viewer only fetches the selected record.
## Rebuilding the data
```bash
python scripts/build_data.py --data-dir /mnt/tmp/phantom-wiki/pw-v1
```
Downloads (once): `huggingface-cli download kilian-group/phantom-wiki-v1 --repo-type dataset`.
The script reads the 9 `text-corpus` + `question-answer` splits and rewrites
`corpus_<u>.json` + `eval_<u>.json` + `sets.json`. The raw parquet is **not**
committed here.
### Rebuilding full runs
```bash
python3 scripts/build_runs.py
python3 scripts/build_e2e_structures.py \
--dataset phantom_wiki \
--unified /mnt/ramdisk/blobstore/timchen0618/data/eval/phantom_wiki/unified/test_with_chunks.unified.jsonl \
--scaffolds-dir /home/azureuser/projects/information-scaffolds/outputs/e2e_runs/new-datasets-full-20260711/phantom_wiki/named-outputs/scaffolds_dir
```
The builder streams response/judge JSONL, overlays the c1/c2 higher-token
recovery responses on their original full-run files, filters combined response
files to `dataset == "phantom_wiki"`, rejects qids outside
`eval_size5000_seed1.json`, and joins every slot against all 500 gold rows so
missing or unanswered responses remain explicit. Every generated run index
therefore has exactly the eval file's 500 qids. It writes slots
`c1`, `c2`, `c6`, `naive`, `e2e`, and `e2e_rawtext`, then prejoins their compact result
projections into the Compare shards. Repeated system/user prompts are omitted;
all retained strings are capped at 8,192 characters and event payloads at
about 256 KiB per run record with a truncation marker.
Legacy runs use `judge.parsed.correct` over **all 500 eval qids**. The new
E2E+rawtext run uses the canonical mean answer F1 over the same denominator,
with partial-credit records shown explicitly. Missing or unanswered responses
receive zero. Coverage is the number of non-empty model responses divided by 500.
Source provenance (paths relative to
`/home/azureuser/projects/information-scaffolds/`):
| Slot | Response | Judge |
|---|---|---|
| c1 | `outputs/herb_phantom_full/full_closedbook/named-outputs/response/response` overlaid by `outputs/herb_phantom_recover/cb_len4x/named-outputs/response/response` | `outputs/herb_phantom_judge/cb_phantom_len4x/named-outputs/judged/judged` |
| c2 | `outputs/herb_phantom_full/full_openbook/named-outputs/response/response` overlaid by `outputs/herb_phantom_recover/ob_len4x/named-outputs/response/response` | `outputs/herb_phantom_judge/ob_phantom_len4x/named-outputs/judged/judged` |
| c6 | `outputs/herb_phantom_full/full_dci/named-outputs/response/response` | `outputs/herb_phantom_judge/dci_phantom/named-outputs/judged/judged` |
| Naive | `outputs/herb_phantom_full/full_naive_phantom/named-outputs/response/response` | `outputs/herb_phantom_judge/naive_phantom/named-outputs/judged/judged` |
| E2E | `outputs/e2e_runs/new-datasets-full-20260711/phantom_wiki/named-outputs/predictions/predictions` | `outputs/e2e_runs/new-datasets-full-20260711/judges/phantom_wiki/named-outputs/judged/judged` |
| E2E + rawtext | AML `nice_sand_r0tp661q61` predictions | Its canonical per-qid answer-F1 CSV |
## Local dev
```bash
python -m http.server 8000 # then open http://localhost:8000/
```
Serves statically β€” no build step. Push to the HF Space remote to deploy.
Dataset license: BSD-3-Clause.