laya-nli-conflict-v8 β€” research archive (round 8), NOT delivered

⚠️ Research archive β€” NOT a delivered model. This checkpoint failed its round's acceptance gates and was never shipped. The current production head is slow-stack/laya-nli-memory-conflict (v4). Uploaded 2026-10-02 for provenance/backup while round 11 (multi-run verdict protocol) waits for Kaggle GPU quota.

Round 8 (2026-09-29) added 120 "same-subject compatible attribute β‡’ false" rows on the v7 corpus (carried blocks byte-identical): pet_attr 80 (= 40 pet_name B2-same-shape with swapped literals + 40 pet_benign) + doctor_attr 40.

Headline results (frozen 1000-pair main val unless noted)

  • B2 dose-response confirmed: 0.8985 β†’ 0.7491 (βˆ’15pp, the only large diagnostic shift) β€” the v7 mechanism attribution was in the right direction, but the +40 same-shape dose was insufficient to cross 0.5
  • main val 0.896 (pass, err 104); failed: val_soft 7 (βœ—, all holdout degree/barber consid rows that flip across runs β€” training variance, not attr regression), polarity val_soft +6 (βœ—), conformal s2 abstain 35.4% (βœ—, 0.4pp over the 35% line; capture 42/54 = 77.8% was strong)

Artifacts

file value
model.safetensors SHA256 bcd1b158…174088 (full hash in archive_sha256_manifest.txt)
rl_agent_config.json Ο„(noul) = 1.1240; encoder jhu-clsp/mmBERT-base; bf16
metrics.json val_accuracy 0.896, val_ece 0.0319, n_val 1000, no_rl true
val_probs.json frozen-val probability dump (calibration analyses)

Provenance

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for slow-stack/laya-nli-conflict-v8

Finetuned
(66)
this model

Dataset used to train slow-stack/laya-nli-conflict-v8