Visualization Inventory
Generated 2026-05-09. Catalogues every notebook, script, module, CLI entrypoint, and output directory in this repo that produces a visualization (chart, plot, heatmap, figure, HTML report, or LaTeX figure file). Intended as a handoff for an external agent that needs to re-generate or edit these visualizations.
Excludes: vendored deps (bergson/, olmes/), tests, build artifacts, virtualenv.
Quick orientation
- Primary library: Plotly (interactive HTML + static PNG/PDF via Kaleido). Matplotlib is used in vendored Bergson benchmarks only — repo code is Plotly-first.
- Central paper-figure engine:
src/dolma/distribution_report/. One module per figure type (heatmap, composite, radar, marginal, etc.) plus a sharedstyle.py. - Style canon:
src/dolma/distribution_report/style.py— defines colorscales (RdBudiverging,Purplessequential), font sizes, figure dimensions, DPI scale. Edit there to change paper-wide look. - Output convention:
artifacts/holds rendered figures;runs/holds pipeline outputs (not figures).
1. Notebooks (.ipynb)
| Path | Visualization purpose |
|---|---|
notebooks/attribution_analysis.ipynb |
Attribution score distributions, top-k influence inspection. |
notebooks/label_distribution.ipynb |
Label frequency plots (matplotlib bar/histogram). |
notebooks/olmes_evaluation_analysis.ipynb |
OLMES eval result charts (per-task accuracy, comparisons). |
notebooks/weborganizer_topic_distribution.ipynb |
WebOrganizer topic frequency visualizations. |
notebooks/working_sample_playground.ipynb |
Interactive exploration of the working sample (sampling, bin distributions). |
artifacts/influence_bin_scores/test.ipynb |
Sandbox notebook for bin-level influence scoring. |
2. Visualization Python modules
2.1 src/dolma/distribution_report/ — paper-figure engine (Plotly)
The central package. Each module emits one family of figures and is invoked by either runner.py (sampling/marginal report) or influence_runner.py (influence/composite report).
| Module | Produces |
|---|---|
style.py |
Shared layout: colorscales (RdBu, Purples), font sizes (12/22), figure dimensions, DPI scale. Edit here to change global look. |
heatmap_figures.py |
24×24 topic×format heatmaps for doc count, token count. |
marginal_figures.py |
Topic-only and format-only marginal bar charts. |
sampling_figures.py |
Sampling-strategy comparisons (representative vs stratified). |
influence_figures.py |
Per-benchmark influence heatmaps (signed). |
influence_composite.py |
Multi-panel signed influence composite. |
influence_composite_abs.py |
Absolute-value influence composite. |
influence_composite_diffs.py |
Difference between two conditions composite. |
influence_composite_6way.py |
6-way comparison across benchmarks. |
influence_composite_signed_topic.py |
Topic-rolled-up signed influence panels. |
influence_composite_signed_6panel.py |
6-panel signed-influence composite. |
influence_composite_socialiqa.py |
SocialIQA-specific composite. |
influence_comparison.py |
Side-by-side benchmark comparisons. |
influence_histograms.py |
Influence-score distribution histograms (per benchmark + overlay). |
influence_marginals.py |
Influence rolled to topic-only / format-only margins. |
influence_radar.py |
Topic radar chart of influence profiles. |
influence_facets.py |
Faceted influence small-multiples. |
Helper modules in the same package (NOT visualization — data loading, tables, orchestration):
cli.py— entrypoint fordata-attribution-weborganizer-report/-paper-figures.runner.py— orchestrates sampling/marginal report.influence_runner.py— orchestrates influence report.data_loader.py,influence_loader.py,sampling_loader.py,sampling_data.py— data ingest.concentration_table.py,table_c1.py,influence_tables.py— LaTeX/CSV table emit.manifest_bridge.py— converts manifest parquet → EDA CSVs (used by SOC-135).metrics.py,report_metadata.py,report_summary.py,output_note.py,dummy_sampling.py— analysis/reporting glue.
2.2 src/dolma/eda/ — EDA HTML report
| Module | Produces |
|---|---|
charts.py |
Plotly bar charts, heatmaps for topic/format distributions on EDA inputs. |
report.py |
HTML EDA report (artifacts/dolma_eda/report.html). |
CLI: dolma-eda (entry: dolma.eda.core:main).
2.3 src/dolma/corpus_stats/ — corpus statistics report
| Module | Produces |
|---|---|
charts.py |
Plotly histograms for document length statistics. |
report.py |
Wraps the chart pipeline. |
CLI: dolma-corpus-stats (entry: dolma.corpus_stats.cli:main).
2.4 src/dolma/bin_analysis/ — bin-stratification report
| Module | Produces |
|---|---|
report.py |
Plotly visualization of bin-level statistics. |
CLI: data-attribution-bin-analysis (entry: dolma.bin_analysis.cli:main).
3. Standalone visualization scripts
| Script | Output |
|---|---|
scripts/build_corpus_explorer.py |
artifacts/dolma3_corpus_explorer.html and dolma3_corpus_findings.html — interactive corpus explorer. |
scripts/analysis/unlearn_report.py |
artifacts/unlearn_report/unlearn_single_topic_report.html — Plotly heatmap report of single-topic unlearning gammas. |
scripts/analysis/rq4_figures.py |
RQ4 paper figures (lexical-feature heatmaps, profile comparison). PNG + PDF + HTML. |
scripts/generate_6t_figures.py |
Orchestrator: downloads 6T sample manifest from HF, runs WebOrganizer report, writes artifacts/paper_figures_6t/ and Table C1 LaTeX. |
scripts/analysis/rq4_bin_characterization.py, rq4_characterize_top_bins.py, topic_missing_audit.py are characterization/data scripts, not figure generators.
4. CLI entrypoints (from pyproject.toml)
These are the user-facing commands that produce visualizations. All defined under [project.scripts].
| Command | Module | Output dir |
|---|---|---|
dolma-eda |
dolma.eda.core:main |
artifacts/dolma_eda/ |
dolma-corpus-stats |
dolma.corpus_stats.cli:main |
(configurable) |
data-attribution-bin-analysis |
dolma.bin_analysis.cli:main |
(configurable) |
data-attribution-weborganizer-report |
dolma.distribution_report.cli:main |
artifacts/paper_figures/ |
data-attribution-paper-figures |
dolma.distribution_report.cli:main |
(alias of above) |
The other 43 CLI entries do data work (precache, ingest, manifest build, training-index build, scoring) and do not emit figures.
5. Output / artifact directories
| Directory | What lives here |
|---|---|
artifacts/paper_figures/ |
Main paper figures: ~150 PNG/PDF/HTML files (heatmaps, composites, comparisons, histograms, radar, format breakdowns) + table_concentration.tex, table_contrastive_socialiqa_vs_gsm8k.tex. |
artifacts/paper_figures_6t/ |
6T working-sample regenerated figures (SOC-135): topic/format heatmaps PNG/PDF/HTML, comparison note, EDA dump. |
artifacts/dolma_eda/ |
EDA HTML report (report.html) + plot subdir. |
artifacts/influence_bin_scores/ |
Per-bin influence CSVs + sandbox notebook. |
artifacts/influence_bin_scores_split/ |
Correctness-split per-bin influence CSVs. |
artifacts/influence_bbh_base/, _split/, _figures/ |
BBH attribution scores (base model) — perquery CSVs and 51 PNG/HTML figures across the 3 BBH subtasks (snarks, causal_judgement, sports_understanding). |
artifacts/influence_bbh_instruct/, _split/, _figures/ |
Same as above for instruct model. |
artifacts/evaluation/ |
Evaluation result tables / charts. |
artifacts/unlearn_report/ |
Unlearning HTML report + supporting data. |
artifacts/proponent_examples/ |
Top-proponent example dumps. |
artifacts/dolma3_corpus_explorer.html, dolma3_corpus_findings.html |
Standalone interactive HTMLs from scripts/build_corpus_explorer.py. |
6. LaTeX figure-bearing files
| File | Contents |
|---|---|
artifacts/results.tex |
Results section with \includegraphics references into artifacts/paper_figures/. |
artifacts/methodology.tex |
Methodology section. |
artifacts/paper_figures/table_concentration.tex |
Concentration table (auto-generated by concentration_table.py). |
artifacts/paper_figures/table_contrastive_socialiqa_vs_gsm8k.tex |
Contrastive comparison table. |
artifacts/paper_figures_6t/ |
Additional .tex fragments (Table C1) emitted by table_c1.py. |
7. Datasets and benchmarks visualized
For figure context — what each visualization shows.
Query benchmarks (rows in influence figures):
- SocialIQA — main social-reasoning target
- GSM8K — math reference (correct vs incorrect splits)
- MMLU Social Sciences, MMLU STEM
- ARC Easy, ARC Challenge
- BBH Snarks, BBH Causal Judgement, BBH Sports Understanding (added in SOC-170)
Corpus / training-document axis:
- WebOrganizer 24 topics × 24 formats = 576 bins
- 6T deduplicated working sample (10K docs/bin sample published as
HCAI-Lab/dolma3-6t-sample-10000-docs) - Sampling strategies: representative vs stratified (visualized by
sampling_figures.py)
Conditions compared:
- Base model vs Instruct model (separate
_base/_instructartifact dirs) - Canonical vs non-canonical query variants (
*_canonical.png) - Signed vs absolute influence
- Pre vs post unlearning (in
unlearn_report/)
8. Style and conventions (read this first when editing figures)
src/dolma/distribution_report/style.py:
- Diverging colorscale:
RdBu(red = negative/degradation, blue = positive/improvement). - Sequential colorscale:
Purples(used for PPL panels and absolute-value heatmaps). - Comparison colorscale (defined inline): green → white → purple, for paired-condition deltas.
- Font: 12 pt body, 22 pt title.
- Figure dimensions: 900×650 (bar default), 1100×900 (heatmap), 1600×900 (side-by-side).
- DPI scale factor: 3 (Plotly→Kaleido PNG export at 3× pixel density).
- Empty-cell color:
#f0f0f0.
To regenerate any figure with a tweaked palette, change values in style.py and re-run the relevant CLI (data-attribution-paper-figures for the influence/sampling figures, dolma-eda for the EDA report). Each figure module reads the constants — no per-module color overrides.
9. Regeneration recipes (high level)
| To regenerate | Run |
|---|---|
| All paper figures | uv run data-attribution-paper-figures --eda-dir artifacts/dolma_eda --output-dir artifacts/paper_figures --run-label <label> --format all |
| 6T sample figures | uv run python scripts/generate_6t_figures.py |
| EDA report | uv run dolma-eda |
| Bin analysis | uv run data-attribution-bin-analysis |
| Corpus explorer HTML | uv run python scripts/build_corpus_explorer.py |
| Unlearn report | uv run python scripts/analysis/unlearn_report.py |
| RQ4 figures | uv run python scripts/analysis/rq4_figures.py |
| BBH influence figures | Triggered by the SOC-170 BBH scoring SLURM pipeline; figure step calls modules in src/dolma/distribution_report/. |
See README.md § "WebOrganizer report" and docs/WEBORGANIZER_REPORT.md for the full report workflow with input flags.
Xet Storage Details
- Size:
- 11.5 kB
- Xet hash:
- 3ded8a4435a6e859c6acd2492ed6c5c0770016bdbb69f6d8efde2390047a75ef
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.