HCAI-Lab/w2-consensus-deepdive-unlearning-artifacts / social-data-attribution-w2 /docs /VISUALIZATION_INVENTORY.md
| # Visualization Inventory | |
| Generated 2026-05-09. Catalogues every notebook, script, module, CLI entrypoint, and output directory in this repo that produces a visualization (chart, plot, heatmap, figure, HTML report, or LaTeX figure file). Intended as a handoff for an external agent that needs to re-generate or edit these visualizations. | |
| Excludes: vendored deps (`bergson/`, `olmes/`), tests, build artifacts, virtualenv. | |
| ## Quick orientation | |
| - **Primary library**: Plotly (interactive HTML + static PNG/PDF via Kaleido). Matplotlib is used in vendored Bergson benchmarks only — repo code is Plotly-first. | |
| - **Central paper-figure engine**: `src/dolma/distribution_report/`. One module per figure type (heatmap, composite, radar, marginal, etc.) plus a shared `style.py`. | |
| - **Style canon**: `src/dolma/distribution_report/style.py` — defines colorscales (`RdBu` diverging, `Purples` sequential), font sizes, figure dimensions, DPI scale. Edit there to change paper-wide look. | |
| - **Output convention**: `artifacts/` holds rendered figures; `runs/` holds pipeline outputs (not figures). | |
| --- | |
| ## 1. Notebooks (`.ipynb`) | |
| | Path | Visualization purpose | | |
| |------|-----------------------| | |
| | `notebooks/attribution_analysis.ipynb` | Attribution score distributions, top-k influence inspection. | | |
| | `notebooks/label_distribution.ipynb` | Label frequency plots (matplotlib bar/histogram). | | |
| | `notebooks/olmes_evaluation_analysis.ipynb` | OLMES eval result charts (per-task accuracy, comparisons). | | |
| | `notebooks/weborganizer_topic_distribution.ipynb` | WebOrganizer topic frequency visualizations. | | |
| | `notebooks/working_sample_playground.ipynb` | Interactive exploration of the working sample (sampling, bin distributions). | | |
| | `artifacts/influence_bin_scores/test.ipynb` | Sandbox notebook for bin-level influence scoring. | | |
| --- | |
| ## 2. Visualization Python modules | |
| ### 2.1 `src/dolma/distribution_report/` — paper-figure engine (Plotly) | |
| The central package. Each module emits one family of figures and is invoked by either `runner.py` (sampling/marginal report) or `influence_runner.py` (influence/composite report). | |
| | Module | Produces | | |
| |--------|----------| | |
| | `style.py` | Shared layout: colorscales (`RdBu`, `Purples`), font sizes (12/22), figure dimensions, DPI scale. **Edit here to change global look.** | | |
| | `heatmap_figures.py` | 24×24 topic×format heatmaps for doc count, token count. | | |
| | `marginal_figures.py` | Topic-only and format-only marginal bar charts. | | |
| | `sampling_figures.py` | Sampling-strategy comparisons (representative vs stratified). | | |
| | `influence_figures.py` | Per-benchmark influence heatmaps (signed). | | |
| | `influence_composite.py` | Multi-panel signed influence composite. | | |
| | `influence_composite_abs.py` | Absolute-value influence composite. | | |
| | `influence_composite_diffs.py` | Difference between two conditions composite. | | |
| | `influence_composite_6way.py` | 6-way comparison across benchmarks. | | |
| | `influence_composite_signed_topic.py` | Topic-rolled-up signed influence panels. | | |
| | `influence_composite_signed_6panel.py` | 6-panel signed-influence composite. | | |
| | `influence_composite_socialiqa.py` | SocialIQA-specific composite. | | |
| | `influence_comparison.py` | Side-by-side benchmark comparisons. | | |
| | `influence_histograms.py` | Influence-score distribution histograms (per benchmark + overlay). | | |
| | `influence_marginals.py` | Influence rolled to topic-only / format-only margins. | | |
| | `influence_radar.py` | Topic radar chart of influence profiles. | | |
| | `influence_facets.py` | Faceted influence small-multiples. | | |
| Helper modules in the same package (NOT visualization — data loading, tables, orchestration): | |
| - `cli.py` — entrypoint for `data-attribution-weborganizer-report` / `-paper-figures`. | |
| - `runner.py` — orchestrates sampling/marginal report. | |
| - `influence_runner.py` — orchestrates influence report. | |
| - `data_loader.py`, `influence_loader.py`, `sampling_loader.py`, `sampling_data.py` — data ingest. | |
| - `concentration_table.py`, `table_c1.py`, `influence_tables.py` — LaTeX/CSV table emit. | |
| - `manifest_bridge.py` — converts manifest parquet → EDA CSVs (used by SOC-135). | |
| - `metrics.py`, `report_metadata.py`, `report_summary.py`, `output_note.py`, `dummy_sampling.py` — analysis/reporting glue. | |
| ### 2.2 `src/dolma/eda/` — EDA HTML report | |
| | Module | Produces | | |
| |--------|----------| | |
| | `charts.py` | Plotly bar charts, heatmaps for topic/format distributions on EDA inputs. | | |
| | `report.py` | HTML EDA report (`artifacts/dolma_eda/report.html`). | | |
| CLI: `dolma-eda` (entry: `dolma.eda.core:main`). | |
| ### 2.3 `src/dolma/corpus_stats/` — corpus statistics report | |
| | Module | Produces | | |
| |--------|----------| | |
| | `charts.py` | Plotly histograms for document length statistics. | | |
| | `report.py` | Wraps the chart pipeline. | | |
| CLI: `dolma-corpus-stats` (entry: `dolma.corpus_stats.cli:main`). | |
| ### 2.4 `src/dolma/bin_analysis/` — bin-stratification report | |
| | Module | Produces | | |
| |--------|----------| | |
| | `report.py` | Plotly visualization of bin-level statistics. | | |
| CLI: `data-attribution-bin-analysis` (entry: `dolma.bin_analysis.cli:main`). | |
| --- | |
| ## 3. Standalone visualization scripts | |
| | Script | Output | | |
| |--------|--------| | |
| | `scripts/build_corpus_explorer.py` | `artifacts/dolma3_corpus_explorer.html` and `dolma3_corpus_findings.html` — interactive corpus explorer. | | |
| | `scripts/analysis/unlearn_report.py` | `artifacts/unlearn_report/unlearn_single_topic_report.html` — Plotly heatmap report of single-topic unlearning gammas. | | |
| | `scripts/analysis/rq4_figures.py` | RQ4 paper figures (lexical-feature heatmaps, profile comparison). PNG + PDF + HTML. | | |
| | `scripts/generate_6t_figures.py` | Orchestrator: downloads 6T sample manifest from HF, runs WebOrganizer report, writes `artifacts/paper_figures_6t/` and Table C1 LaTeX. | | |
| `scripts/analysis/rq4_bin_characterization.py`, `rq4_characterize_top_bins.py`, `topic_missing_audit.py` are characterization/data scripts, not figure generators. | |
| --- | |
| ## 4. CLI entrypoints (from `pyproject.toml`) | |
| These are the user-facing commands that produce visualizations. All defined under `[project.scripts]`. | |
| | Command | Module | Output dir | | |
| |---------|--------|------------| | |
| | `dolma-eda` | `dolma.eda.core:main` | `artifacts/dolma_eda/` | | |
| | `dolma-corpus-stats` | `dolma.corpus_stats.cli:main` | (configurable) | | |
| | `data-attribution-bin-analysis` | `dolma.bin_analysis.cli:main` | (configurable) | | |
| | `data-attribution-weborganizer-report` | `dolma.distribution_report.cli:main` | `artifacts/paper_figures/` | | |
| | `data-attribution-paper-figures` | `dolma.distribution_report.cli:main` | (alias of above) | | |
| The other 43 CLI entries do data work (precache, ingest, manifest build, training-index build, scoring) and do not emit figures. | |
| --- | |
| ## 5. Output / artifact directories | |
| | Directory | What lives here | | |
| |-----------|-----------------| | |
| | `artifacts/paper_figures/` | Main paper figures: ~150 PNG/PDF/HTML files (heatmaps, composites, comparisons, histograms, radar, format breakdowns) + `table_concentration.tex`, `table_contrastive_socialiqa_vs_gsm8k.tex`. | | |
| | `artifacts/paper_figures_6t/` | 6T working-sample regenerated figures (SOC-135): topic/format heatmaps PNG/PDF/HTML, comparison note, EDA dump. | | |
| | `artifacts/dolma_eda/` | EDA HTML report (`report.html`) + plot subdir. | | |
| | `artifacts/influence_bin_scores/` | Per-bin influence CSVs + sandbox notebook. | | |
| | `artifacts/influence_bin_scores_split/` | Correctness-split per-bin influence CSVs. | | |
| | `artifacts/influence_bbh_base/`, `_split/`, `_figures/` | BBH attribution scores (base model) — perquery CSVs and 51 PNG/HTML figures across the 3 BBH subtasks (snarks, causal_judgement, sports_understanding). | | |
| | `artifacts/influence_bbh_instruct/`, `_split/`, `_figures/` | Same as above for instruct model. | | |
| | `artifacts/evaluation/` | Evaluation result tables / charts. | | |
| | `artifacts/unlearn_report/` | Unlearning HTML report + supporting data. | | |
| | `artifacts/proponent_examples/` | Top-proponent example dumps. | | |
| | `artifacts/dolma3_corpus_explorer.html`, `dolma3_corpus_findings.html` | Standalone interactive HTMLs from `scripts/build_corpus_explorer.py`. | | |
| --- | |
| ## 6. LaTeX figure-bearing files | |
| | File | Contents | | |
| |------|----------| | |
| | `artifacts/results.tex` | Results section with `\includegraphics` references into `artifacts/paper_figures/`. | | |
| | `artifacts/methodology.tex` | Methodology section. | | |
| | `artifacts/paper_figures/table_concentration.tex` | Concentration table (auto-generated by `concentration_table.py`). | | |
| | `artifacts/paper_figures/table_contrastive_socialiqa_vs_gsm8k.tex` | Contrastive comparison table. | | |
| | `artifacts/paper_figures_6t/` | Additional `.tex` fragments (Table C1) emitted by `table_c1.py`. | | |
| --- | |
| ## 7. Datasets and benchmarks visualized | |
| For figure context — what each visualization shows. | |
| **Query benchmarks (rows in influence figures):** | |
| - SocialIQA — main social-reasoning target | |
| - GSM8K — math reference (correct vs incorrect splits) | |
| - MMLU Social Sciences, MMLU STEM | |
| - ARC Easy, ARC Challenge | |
| - BBH Snarks, BBH Causal Judgement, BBH Sports Understanding (added in SOC-170) | |
| **Corpus / training-document axis:** | |
| - WebOrganizer 24 topics × 24 formats = 576 bins | |
| - 6T deduplicated working sample (10K docs/bin sample published as `HCAI-Lab/dolma3-6t-sample-10000-docs`) | |
| - Sampling strategies: representative vs stratified (visualized by `sampling_figures.py`) | |
| **Conditions compared:** | |
| - Base model vs Instruct model (separate `_base` / `_instruct` artifact dirs) | |
| - Canonical vs non-canonical query variants (`*_canonical.png`) | |
| - Signed vs absolute influence | |
| - Pre vs post unlearning (in `unlearn_report/`) | |
| --- | |
| ## 8. Style and conventions (read this first when editing figures) | |
| `src/dolma/distribution_report/style.py`: | |
| - Diverging colorscale: **`RdBu`** (red = negative/degradation, blue = positive/improvement). | |
| - Sequential colorscale: **`Purples`** (used for PPL panels and absolute-value heatmaps). | |
| - Comparison colorscale (defined inline): green → white → purple, for paired-condition deltas. | |
| - Font: 12 pt body, 22 pt title. | |
| - Figure dimensions: 900×650 (bar default), 1100×900 (heatmap), 1600×900 (side-by-side). | |
| - DPI scale factor: 3 (Plotly→Kaleido PNG export at 3× pixel density). | |
| - Empty-cell color: `#f0f0f0`. | |
| To regenerate any figure with a tweaked palette, change values in `style.py` and re-run the relevant CLI (`data-attribution-paper-figures` for the influence/sampling figures, `dolma-eda` for the EDA report). Each figure module reads the constants — no per-module color overrides. | |
| --- | |
| ## 9. Regeneration recipes (high level) | |
| | To regenerate | Run | | |
| |--------------|-----| | |
| | All paper figures | `uv run data-attribution-paper-figures --eda-dir artifacts/dolma_eda --output-dir artifacts/paper_figures --run-label <label> --format all` | | |
| | 6T sample figures | `uv run python scripts/generate_6t_figures.py` | | |
| | EDA report | `uv run dolma-eda` | | |
| | Bin analysis | `uv run data-attribution-bin-analysis` | | |
| | Corpus explorer HTML | `uv run python scripts/build_corpus_explorer.py` | | |
| | Unlearn report | `uv run python scripts/analysis/unlearn_report.py` | | |
| | RQ4 figures | `uv run python scripts/analysis/rq4_figures.py` | | |
| | BBH influence figures | Triggered by the SOC-170 BBH scoring SLURM pipeline; figure step calls modules in `src/dolma/distribution_report/`. | | |
| See `README.md` § "WebOrganizer report" and `docs/WEBORGANIZER_REPORT.md` for the full report workflow with input flags. | |
Xet Storage Details
- Size:
- 11.5 kB
- Xet hash:
- 3ded8a4435a6e859c6acd2492ed6c5c0770016bdbb69f6d8efde2390047a75ef
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.