wildchat-aqa-viewer / README.md
timchen0618's picture
Pack corpus into 2048 buckets; fix static root routing
7a2cc9a verified
|
Raw
History Blame Contribute Delete
3.46 kB
---
title: WildChat-AQA Viewer
emoji: πŸ’¬
colorFrom: green
colorTo: indigo
sdk: static
pinned: false
license: apache-2.0
---
# πŸ’¬ WildChat-AQA Viewer
Static viewer for the **WildChat-AQA** aggregative question-answering benchmark
([paper, EMNLP 2025](https://aclanthology.org/2025.emnlp-main.1667.pdf) Β·
[dataset](https://huggingface.co/datasets/wtzhang-nlp/wildchat_aqa) Β·
[GitHub](https://github.com/yuntian-group/wildchat_aggregative_question_answering)).
WildChat-AQA asks *aggregative* questions over WildChat chat logs β€” e.g. "what
are the most popular topics for user X?" β€” answered by aggregating over the
conversations matching a **condition**. Each question is 10-way multiple choice;
the answer is the option with the highest weight.
## The Eval tab
One question at a time, showing:
- **Question** + badges (target type Β· keyword type Β· #conditions Β· #matching conversations).
- **Condition** β€” the aggregation filter (`user_name`, `country`, `language`,
`label_level_1/2` topic, `time_week`, `keywords_aggregated`), with label ids
resolved to names.
- **Answer** β€” all 10 options as weighted bars, the argmax highlighted.
- **Aggregation Evidence** β€” the raw `target_candidates` (value + count) that the
answer is aggregated from, with label ids resolved.
- **Supporting Conversations** β€” the actual WildChat conversations matching the
condition (lazy-loaded, full raw turns rendered as chat), capped at 30 per
question (largest first; flagged when capped / sampled).
- **PROBE Query & Metadata** β€” the GPT-generated retrieval query, target/keyword
type, hash.
Sidebar: full-text search (question / hash) + filters for **target type**,
**keyword type**, and **condition dimension**, with prev/next (← / β†’).
## Data layout
Generated by `scripts/build_wildchat_aqa.py` from three source datasets:
| source | used for |
|---|---|
| `wtzhang-nlp/wildchat_aqa` | the 5,637 questions, options, weights, target_candidates, conditions |
| `ksx-wz/wildchat_aqa_conversations` | the supporting conversations (linked by condition) |
| `ksx-wz/…_with_embedding_and_gpt_generated_query` | the PROBE `generated_query` (embeddings dropped) |
| repo `wildchat_aqa_taxonomy/` | resolving `label_level_1/2` ids β†’ names |
```
eval/index.json # light list for the sidebar (5,637 entries)
eval/records/<qhash>.json # per-question: question, condition, options+weights, evidence, PROBE query, capped support refs
corpus/b<N>.json # conversations packed into 2048 buckets (bucket = int(hash[:8],16) % 2048), lazy-loaded
```
**Condition matching** mirrors the benchmark's aggregation: **OR within a
repeated condition type, AND across types**; `time_week` matches a 7-day window;
label values encode as `"<L1>.<L2>"`. Supporting conversations are capped at 30
per question (`--cap`); the ~3 truly across-all questions show a size-ranked
sample of the full corpus.
### Regenerating the shards
```bash
# download sources (see build_wildchat_aqa.py header), then:
python scripts/build_wildchat_aqa.py --src <src> --out . --cap 30
```
## Local dev
```bash
python -m http.server 8000 # then open http://localhost:8000/
```
Deploy = push to the Hugging Face Space `main` (served statically, no build step).
The ~60k conversations are packed into 2048 bucket files (`corpus/b<N>.json`)
because HF static Spaces fall back to serving `README.md` at the root once the
repo file tree grows past ~10k files.