memory-lora-gemma4 / runs /switch.log
moncefem
8am snapshot: complete ~1650-repo dataset + batch2 novel repos + v3 checkpoints
fb5235a
Raw
History Blame Contribute Delete
983 Bytes
2026-07-25 01:57:06: orchestrator started
2026-07-25 01:57:06: final QA finished
coverage: emb=1654 qa=1622 missing=32
2026-07-25 01:57:06: assembling complete aligned dataset ...
aligned6 dataset: 1647 repos (embeddings) | 22267 QA
repo splits: Counter({'train': 1341, 'cr_val': 167, 'cr_test': 139})
QA by split: {'train': 18115, 'cr_test': 1896, 'cr_val': 2256}
-> aligned6_embeddings.parquet + aligned6_qna.jsonl
2026-07-25 01:57:10: assembled: repos=1647 qa=22267 (was 1058 repos / 8540 qa)
2026-07-25 01:57:10: waiting for next checkpoint save to time the kill (zero wasted steps) ...
2026-07-25 02:01:17: checkpoint just saved (mtime bumped) at step841; KILLING training now to switch dataset
2026-07-25 02:01:17: training killed; supervisor will relaunch on COMPLETE dataset (1647 repos) from head.latest.pt within ~45s
2026-07-25 02:02:47: WARN: training not detected 90s after kill -- supervisor should relaunch; will self-heal
2026-07-25 02:02:47: orchestrator done