laya-nli-conflict-v6 β€” research archive (round 6), NOT delivered

⚠️ Research archive β€” NOT a delivered model. This checkpoint failed its round's acceptance gates and was never shipped. The current production head is slow-stack/laya-nli-memory-conflict (v4). Uploaded 2026-10-02 for provenance/backup while round 11 (multi-run verdict protocol) waits for Kaggle GPU quota.

Round 6 (2026-09-28) was a data-hygiene reset: all 35 acceptance sentences scrubbed from the training corpus (leak_audit.py 0/35 assertion enforced), unrelated control rows added (negation-family 400 / counterfactual 300), five negation families Γ—80 (pet/never/relapse/dneg/nocar), pure-CE main arm (RL term dropped after round 5's arm-A collapse). This round also re-baselined the numbers of v2–v5 under the sentence-reuse-free accounting.

Headline results (frozen 1000-pair main val unless noted)

  • main val 0.903 βœ“ (program best at the time; refuted "leakage inflates the score"), old-20 20/20 βœ“ (unrelated-supplement false-positives fixed), negation 5/5 βœ“ (first time), confidence band 16.29pp kept
  • failed: new-10 8/10 (βœ—), val_soft 5 (βœ—, all holdout degree/barber rows), polarity +2 (βœ—), conformal FAIL β€” root cause: 52 of 97 val errors sit at confidence β‰₯0.92, which no confidence-threshold rule can catch (the compression-band disease; abstention moved to a surface-consistency rule in v7)

Artifacts

file value
model.safetensors SHA256 11facce6…84b4e0 (full hash in archive_sha256_manifest.txt)
rl_agent_config.json Ο„(noul) = 1.0905; encoder jhu-clsp/mmBERT-base; bf16
metrics.json val_accuracy 0.903, val_ece 0.0192, n_val 1000, no_rl true
val_probs.json frozen-val probability dump (calibration analyses)

Provenance

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for slow-stack/laya-nli-conflict-v6

Finetuned
(66)
this model

Dataset used to train slow-stack/laya-nli-conflict-v6