Buckets:
| # Comma-2T ToMBench unlearning — OVERNIGHT AUTONOMOUS state (2026-07-06) | |
| FULL 24-topic x 3-seed = n=72 paired experiment (user chose full scope). | |
| Executes findings/comma_2t/PROBE_PARITY_FINDINGS.md §5. Spec: docs/specs/comma_2t_tom_unlearning.md. | |
| ## OVERNIGHT AUTONOMY — runs with NO local session/SSH | |
| On-cluster driver SLURM job `c2ttom_driver` (job 5480191, ice-cpu, 12h walltime) | |
| orchestrates everything independent of any laptop/SSH: | |
| - each 5-min tick: submit_tom_train.sh (recover failed train cells) + | |
| submit_tom_evals.sh (eval complete adapters); both idempotent. | |
| - when all 144 eval JSONs present and nothing live -> runs analyze_tom_unlearning.py. | |
| - OUTPUTS (check these in the morning): | |
| - $TDA/comma/tom_unlearning/tom_analysis.txt (paired-test summary) | |
| - $TDA/comma/tom_unlearning/tom_paired.csv (per-pair d = gamma_inf - gamma_rand) | |
| - $TDA/comma/tom_unlearning/driver.log (per-tick submit activity) | |
| - driver stdout: $TDA/logs/comma_unlearn/c2ttom_driver_5480191.out | |
| ($TDA=/storage/ice-shared/cs7634/staff/TDA) | |
| ## Design | |
| - expA (influence): tombench top-200 influence forget docs, per topic (24) x seed (42,43,44). | |
| - exp1 (random 1000, benchmark-agnostic): REUSE existing sweep adapters | |
| $TDA/comma/unlearn_faithful/comma_2t/full/comma_2t_full_exp1_<topic>_s<seed>. | |
| - Eval: tombench:mc::socialtda acc_raw; gamma=(acc-0.5132867)/0.5132867. Base=0.5132867. | |
| - Paired: d = gamma_expA - gamma_exp1 per (topic,seed); one-sided d>0; n=72. | |
| ## Job families / IDs | |
| - Forget builds: DONE (all 24 topics, 200 docs each, verified). | |
| - expA training (66 new): c2ttomtr_* (submitted job block ~5479475-5479540); 6 original done. | |
| - evals: c2ttom_eval_* (exp1 arm ~done; expA evals submitted by driver as adapters land). | |
| - Overnight driver: 5480191. | |
| - manifest_all.tsv = 72 cells (record + recovery source). | |
| ## Known/handled footguns | |
| - GROUPS is a readonly bash builtin -> renamed TGROUPS in build_remaining_forgets.sh. | |
| - eval sbatch now has bad-node --exclude + broad constraint (HX00|A100-80GB|L40S|A40). | |
| - paper-actual overrides (docs_per_retain_bin=null min_tokens=0 resample_interval=0), | |
| MAXWALL=340, batch 4 / grad-accum 4, max-forget 200, max-retain 9000, comma_2t_lora. | |
| - Same-checkpoint pairing throughout (comma-v0.1-2t). | |
| ## Early partial results (exp1 control arm) | |
| - exp1 social_life gamma ~ -0.05 (harms ToMBench); science_math ~ -0.015. Influence arm pending. | |
| ## Morning TODO (me or user) | |
| 1. Read tom_analysis.txt / tom_paired.csv. 2. Write findings/comma_2t/TOM_UNLEARNING_FINDINGS.md. | |
| 3. Append LOGBOOK.md; update GPU-hours ledger | |
| (scripts/monitoring/gpu_hours_report.py --default-workstream comma_trackstar_attribution). | |
| 4. Adversarially verify (gamma math, paired stats, design validity vs OLMo). | |
| 5. HF-back adapters+evals+forget sets (durable) before reclaiming $TDA space. | |
| 6. Commit new code (spec, eval wrapper+test, analysis) on glenn/comma-2t-probe-parity (ASK first). | |
Xet Storage Details
- Size:
- 2.93 kB
- Xet hash:
- 997d6e39a92f34fd1e068a0d16e3eaf7301cf839f2f93e4503fbddeabe0854c9
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.