Commit History

Ship best-of-7 arm: held-out 85.9 (was 78). Card reports the full replicate sample.
a096851
verified

rodriguescarson commited on

Remove internal decision log: contains the retracted +16 claim
d695411
verified

rodriguescarson commited on

Card: correct scores to v5 (85 on-data / 78 held-out), retract the curation-beats-scale claim
697b80d
verified

rodriguescarson commited on

v5 complete: Llama-3.3-70B r=64 hard-mined 85/78
5393068
verified

rodriguescarson commited on

card: add +16 refutation finding
737849f
verified

rodriguescarson commited on

Upload adapter_model.safetensors with huggingface_hub
2a6a147
verified

rodriguescarson commited on

v4 submission: severity model 77/76, updated card, +16 refutation
e90e91d
verified

rodriguescarson commited on

Model card as controlled study: shortcut learning, counterfactual augmentation, severity, Hooker/Zheng/Mazumder
740e659
verified

rodriguescarson commited on

LoRA adapter, 83/78 win rate, AutoScientist by Adaption
ca6245f
verified

rodriguescarson commited on