| suite,base_model,method,n_domains,n_seeds,retention_metric,retention_value_pct,naive_reference_pct,status,source_file,notes | |
| 5-domain-realworld,Mistral-7B,modular_crma,5,3,holdout_NLL_drift,-0.166,42.96,valid_multiseed,raw/multiseed_results_combined.json,per-task LoRA + CRMA backbone; drift of saved snapshots re-run under final backbone | |
| 5-domain-realworld,Mistral-7B,naive_sequential_lora,5,3,holdout_NLL_forgetting,42.96,,valid_multiseed,raw/multiseed_results_combined.json,single LoRA trained sequentially across domains | |
| 5-domain-realworld,Mistral-7B,frozen_base,5,3,holdout_NLL_drift,1.948,,valid_multiseed,raw/multiseed_results_combined.json,no adaptation control | |
| 4-domain-MLCF,Mistral-7B-v0.3,modular_crma,4,1,holdout_NLL_drift,-0.1,351.4,valid_single_run,raw/ablation_v8.1_7b_results.md,avg of per-task drift table; NAIVE ref is the same run's forgetting avg | |
| 4-domain-MLCF,Mistral-7B-v0.3,naive_sequential_lora,4,1,holdout_NLL_forgetting,351.4,,valid_single_run,raw/ablation_v8.1_7b_results.md, | |
| 4-domain-MLCF,TinyLlama-1.1B-Chat-v1.0,modular_crma,4,1,holdout_NLL_drift,-0.1,225.3,valid_single_run,raw/ablation_v8.1_results.md, | |
| 4-domain-MLCF,TinyLlama-1.1B-Chat-v1.0,naive_sequential_lora,4,1,holdout_NLL_forgetting,225.3,,valid_single_run,raw/ablation_v8.1_results.md, | |
| 4-domain-MLCF-history,TinyLlama-1.1B,cl_stack_v3_ewc_gradproj,4,1,holdout_NLL_forgetting,91.3,185.8,valid_single_run,raw/full_ablation_history_v2_v8.md,EWC + gradient projection; pre-data-fix suite - compare only to its own NAIVE column | |
| 4-domain-MLCF-history,TinyLlama-1.1B,cl_stack_v5_10component,4,1,holdout_NLL_forgetting,58.4,88.8,valid_single_run,raw/full_ablation_history_v2_v8.md,10-component stack; post-data-fix suite - NAIVE dropped 185.8->88.8 from data fixes alone | |
| 4-domain-MLCF-history,Mistral-7B,cl_stack_v7_replay_kd_freeze,4,1,holdout_NLL_forgetting,109.3,212.8,valid_single_run,raw/full_ablation_history_v2_v8.md,"replay + knowledge distillation + bottom-layer freeze, Mistral-7B" | |
| 4-domain-MLCF-history,TinyLlama-1.1B,cl_stack_v2_olora_ewc_gradproj_replay,2,1,holdout_NLL_forgetting,-2.0,114.1,invalid_disclosed,raw/full_ablation_history_v2_v8.md,INVALID: PiSSA-init O-LoRA grad norms ~126678 vs clip 100 (ratio 0.00079) froze the model; the -2.0% 'win' is an artifact. O-LoRA has never been validly measured here | |
| 4-domain-MLCF-history,TinyLlama-1.1B,cl_stack_v4_cumulbasis,2,1,holdout_NLL_forgetting,27.8,105.3,incomplete_disclosed,raw/full_ablation_history_v2_v8.md,INCOMPLETE: run cut off mid-Phase-3; Phase-2-only numbers | |
| 4-domain-MLCF-history,TinyLlama-1.1B,cl_stack_v6_sma,2,1,holdout_NLL_forgetting,61.5,191.7,incomplete_disclosed,raw/full_ablation_history_v2_v8.md,INCOMPLETE: Sparse Memory Adapter crashed (OOM) at Phase 3; Phase-2-only numbers | |
| mquake-5skill-vault,Qwen3-4B-Instruct-2507,modular_vault_slots,5,1,accuracy_BWT,0.0,,valid_single_run,raw/bwt_mquake_vault5_s42.json,"retention matrix R; BWT_k = R_final,k - R_k,k; mean over 4 earlier skills" | |