Spaces:
Running
Running
| title: HERB Viewer | |
| emoji: π’ | |
| colorFrom: blue | |
| colorTo: indigo | |
| sdk: static | |
| pinned: false | |
| license: cc-by-nc-4.0 | |
| # π’ HERB Viewer | |
| Static viewer for **HERB** β *Benchmarking Deep Search over Heterogeneous | |
| Enterprise Data*: multi-hop, source-aware RAG over diverse enterprise artifacts | |
| (Slack, meeting transcripts/chats, documents, URLs, GitHub PRs). | |
| - Paper: <https://arxiv.org/abs/2506.23139> | |
| - Code: <https://github.com/SalesforceAIResearch/HERB> | |
| - Dataset: <https://huggingface.co/datasets/Salesforce/HERB> | |
| Tabs: | |
| - **π Corpus** β the 38,600 enterprise artifacts across 30 products. Pick a | |
| **product** and an **artifact type** (Slack / Documents / Meeting transcripts / | |
| Meeting chats / URLs / Pull requests), then a dropdown selects an artifact by | |
| a constructed title. Each artifact renders in its native shape (Slack message + | |
| thread, document markdown, meeting transcript, PR summary + reviews, URL), with | |
| employee/customer ids resolved to names. | |
| - **β Eval** β the 1,514 questions (815 answerable + 699 unanswerable). Filter by | |
| product, answerable/unanswerable, and question type (person / content / company | |
| / pr / url), plus whether the question requires computation/comparison. Each | |
| shows the ground-truth answer and **evidence-citation chips** | |
| (the artifact ids that support it) that jump straight to that artifact in the | |
| Corpus tab. | |
| - **E2E Structures v3** β the five constructed scaffold shapes from the | |
| canonical combined-corpus E2E v3 run. Supporting artifacts are mapped back to | |
| the packed extraction bundle that produced each structure. | |
| - **E2E+Rawtext Structures** β the same cached five-shape scaffold used by the | |
| overlay run, with raw HERB documents mounted as a sixth fallback source. | |
| - **c1 Closed-book**, **c2 With-docs**, **c6 Agentic-DCI**, | |
| **Naive-search**, **E2E raw**, **E2E combined**, and **E2E v3 + rawtext** β | |
| normalized full-run | |
| responses and judge outcomes. These tabs intentionally include only the | |
| **815 answerable questions** and make missing/unanswered predictions explicit. | |
| Search and filter by product, question type, computation/comparison category, | |
| and correctness status; agentic tabs also expose compact, expandable event | |
| traces. | |
| - **β Compare** β the same 815 canonical qids prejoined across all seven runs, | |
| with score/coverage cards and side-by-side predictions. Filters include | |
| disagreement, any missing response, and questions where only E2E combined is | |
| correct. | |
| Question selection follows normalized `Product_aN` / `Product#aN` ids across | |
| Eval, run, Compare, and E2E Structures tabs. Agent traces pair tool calls with | |
| their results, collapse long outputs, render gold answers as compact pills, and | |
| keep E2E evidence citations closed until requested. | |
| ## Full-run results | |
| Every run uses the canonical type-aware `judge_score`: normalized set F1 for | |
| person/company/PR/URL questions and the normalized 0-100 reference score for | |
| content questions. Missing or unanswered questions receive zero. Coverage is | |
| the number with a non-empty model response. | |
| | Run | Perfect / 815 | Type-aware score | Answered / 815 | Coverage | | |
| |---|---:|---:|---:|---:| | |
| | c1 Closed-book | 0 | 4.25% | 815 | 100.00% | | |
| | c2 With-docs | 385 | 67.96% | 815 | 100.00% | | |
| | c6 Agentic-DCI | 386 | 75.64% | 792 | 97.18% | | |
| | Naive-search | 178 | 45.96% | 547 | 67.12% | | |
| | E2E raw | 284 | 67.87% | 798 | 97.91% | | |
| | E2E combined | 336 | 73.91% | 813 | 99.75% | | |
| | E2E v3 + rawtext | 378 | 77.63% | 812 | 99.63% | | |
| ## Computation / comparison breakdown | |
| The split marks questions that ask for a numeric maximum/minimum or comparison, | |
| or require arithmetic aggregation over retrieved documents. In this HERB | |
| release, every positive template is a maximum/highest/most aggregation; there | |
| are no separate minimum, numeric-comparison, or arithmetic-only templates. | |
| | Bracket | Questions | Share | | |
| |---|---:|---:| | |
| | Computation / comparison | 130 | 15.95% | | |
| | Other | 685 | 84.05% | | |
| Every score is the mean canonical type-aware judge score over the bracket | |
| denominator. Coverage is shown in parentheses. | |
| | Run | Computation / comparison | Other | Score gap | | |
| |---|---:|---:|---:| | |
| | c1 Closed-book | 0.00% (100.00%) | 5.05% (100.00%) | -5.05 pp | | |
| | c2 With-docs | 50.10% (100.00%) | 71.35% (100.00%) | -21.25 pp | | |
| | c6 Agentic-DCI | 79.50% (96.92%) | 74.91% (97.23%) | +4.59 pp | | |
| | Naive-search | 18.21% (20.00%) | 51.22% (76.06%) | -33.01 pp | | |
| | E2E raw | 71.64% (96.15%) | 67.16% (98.25%) | +4.48 pp | | |
| | E2E combined | 72.83% (100.00%) | 74.12% (99.71%) | -1.29 pp | | |
| | E2E v3 + rawtext | 78.05% (99.23%) | 77.55% (99.71%) | +0.50 pp | | |
| ## Data | |
| | File | Contents | | |
| |---|---| | |
| | `eval.json` | 1,514 derived question records (Eval tab) | | |
| | `computation_comparison_report.json` | Bracket percentages and canonical per-run performance | | |
| | `products.json` | 30 products with per-type artifact counts (drives the filters) | | |
| | `employees.json` | 530 `eid_β¦ β {name, role, org}` (from `metadata/employee.json`) | | |
| | `customers.json` | 120 customer profiles (from `metadata/customers_data.json`) | | |
| | `runs/manifest.json` | Seven-run registry with scores and coverage | | |
| | `runs/<slot>/index.json` | Search/filter metadata for all 815 answerable qids | | |
| | `runs/<slot>/records/<qid>.json` | Normalized response, judge result, usage, and compact events | | |
| | `compare/index.json` | Compare filters and seven-run summaries | | |
| | `compare/records/<qid>.json` | All seven normalized outcomes prejoined by canonical qid | | |
| | `e2e_structures_v3/` | 815 compressed per-question combined E2E v3 structure shards | | |
| Derived files are regenerated with: | |
| ```bash | |
| python scripts/build_data.py # reads the HERB products/ + metadata/ | |
| python scripts/build_runs.py # reads the response/judge JSONLs listed below | |
| python scripts/analyze_computation_comparison.py | |
| python3 scripts/build_e2e_structures.py \ | |
| --dataset herb \ | |
| --unified /mnt/ramdisk/blobstore/timchen0618/data/eval/herb/unified/test_with_chunks.unified.jsonl \ | |
| --scaffolds-dir /home/azureuser/projects/information-scaffolds/outputs/e2e_runs/new-datasets-full-20260711/herb_combined/named-outputs/scaffolds_dir \ | |
| --bundle-manifest /mnt/tmp/information-scaffolds-e2e-new-datasets/herb_combined/manifest.jsonl | |
| ``` | |
| `build_runs.py` streams the large JSONLs, overlays the c1/c2 higher-token | |
| recovery responses on their original full-run files, filters combined files to | |
| `dataset == "herb"`, and joins every run against the answerable rows in | |
| `eval.json`. Eval gids such as `Product#aN` are normalized exactly once to the | |
| run qid `Product_aN`. Repeated system/user prompts are discarded; individual | |
| strings are capped at 8,192 characters and compact event payloads near 256 KiB | |
| per record. | |
| Run provenance: | |
| - c1/c2 responses: original full-run rows under | |
| `information-scaffolds/outputs/herb_phantom_full/`, overlaid by their | |
| higher-token rows under `information-scaffolds/outputs/herb_phantom_recover/` | |
| - c6/naive responses: `information-scaffolds/outputs/herb_phantom_full/` | |
| - c1/c2/c6/naive type-aware judges: | |
| `information-scaffolds/outputs/herb_type_aware_judge_20260715/` | |
| - E2E raw and combined responses/judges: | |
| responses under `information-scaffolds/outputs/e2e_runs/new-datasets-full-20260711/` | |
| and type-aware judges under | |
| `information-scaffolds/outputs/herb_type_aware_judge_20260715/` | |
| - E2E v3 + rawtext: AML `kind_hand_jy66s6xth0` predictions and canonical | |
| type-aware evaluation | |
| The six rejudge jobs were `jolly_soccer_y3q5z6y4lk` (c1), | |
| `ivory_mango_jgkcpvqbl7` plus recovery `polite_shoe_f4zz9xccjv` (c2), | |
| `honest_truck_yfv0jfb1d3` (c6), `ivory_tray_rmj660nn66` (naive), | |
| `bubbly_lychee_qgzkwj17n6` (E2E raw), and `bubbly_match_5xbl28h7h2` | |
| (E2E combined). The earlier E2E tabs are distinct: **raw** uses the raw-corpus | |
| run, while **combined** uses the combined-structures run. | |
| **Artifact content is not bundled.** The 30 product JSONs (~30 MB) already live | |
| on the `Salesforce/HERB` dataset, which serves them with permissive CORS, so the | |
| viewer `fetch()`es each product on demand (cached) and renders its artifacts | |
| client-side β the Space only bundles the small derived files above. | |
| ## Local dev | |
| ```bash | |
| python -m http.server 8000 # then open http://localhost:8000/ | |
| ``` | |
| Serves statically β no build step. Push to the HF Space remote to deploy. | |
| Dataset license: CC-BY-NC-4.0. HERB was generated using GPT-4o. | |