--- license: apache-2.0 tags: - governance - eu-ai-act - lora - experimental - measurement - text-generation pipeline_tag: text-generation library_name: gguf base_model: Qwen/Qwen2.5-1.5B datasets: - csoai/aiact-frozen-split-harness - csoai/gspc-normalized models: - csoai/oowm-router --- # sov34-1.5b — EXPERIMENTAL (dual-gate candidate: FAILED gate 2, published with failures included) **Register: measurement/attestation only. This model makes no capability claims. Read the numbers before use.** LoRA (r16) on Qwen2.5-1.5B, trained on 4,853 governance-corpus rows (eval_loss 0.821). This was a candidate successor to sov33-unified under a dual-gate rule: beat it on generality AND on the frozen governance split, or make no claims. **Gate 2 failed. We publish the data anyway — failures included is the point of the measurement platform.** ## Measured gates (identical 170 held-out items, frozen split v1, set-F1) | Model | set-F1 | CI95 (BCa n=2000) | |---|---|---| | sov33-unified | 0.2508 | [0.221, 0.283] | | **sov34 (this model)** | **0.1975** | [0.169, 0.226] | | base qwen2.5-1.5b | 0.1767 | [0.151, 0.205] | Generality (lm-eval arc_easy, validated hf lane): sov34 0.7504 ± 0.0089; base 0.7551 ± 0.0088 (preserved within noise). sov33-unified for contrast: 0.2534 (near-random — disclosed specialist profile). ## Verdict Gate 1 (generality): PASS. Gate 2 (governance frozen split): FAIL vs 0.2508. Dual-gate NOT MET. The LoRA nudged governance retrieval above its own base but does not match the specialist. Consistent with the published canon: closed-book small models cannot hold statutory citation — retrieval grounding is required, not optional. ## What this model is for Experimentation, reproduction, router-substrate research (a 1.5B that keeps generality while carrying governance signal). NOT for compliance decisions, NOT a certified anything. ## Copyright & training-data provenance (EU AI Act Art 53(1)(c) / (d)) The open-source GPAI exemption waives Art 53(1)(a)/(b) but **not** the copyright policy (c) or the training-content summary (d). This release states both, because our own OSSBench measures their presence and we hold ourselves to it. - **Training content.** A LoRA (r16) adapter over Qwen2.5-1.5B, tuned on **4,853 rows of governance/regulatory text** — the EU AI Act (Regulation (EU) 2024/1689) and related official legal instruments. Official legal texts are not subject to restrictive copyright. The frozen harness and split are public: dataset `csoai/aiact-frozen-split-harness`. - **Copyright policy.** CSOAI respects rights-holders' text-and-data-mining reservations under the DSM Directive Art 4(3). No paywalled, scraped, or opt-out-reserved corpora were used for this adapter; the base model (Qwen2.5-1.5B) is used under its own licence. - **Deliberately silent on two axes.** This HF upload does not yet ship a components manifest or a cryptographic verification artefact; OSSBench therefore reads those two checks as ABSENT, which is the correct and honest reading. We do not name them here, because naming an artefact you do not ship is how a keyword scan is fooled into reporting it present — the exact failure this benchmark exists to catch. ## Reproduce Harness + results: HF dataset csoai/aiact-frozen-split-harness (incl. sov34_dualgate_triangulation_2026-08-03.json). Triangulation: sov33-unified / sov34 / base on byte-identical item sets. ## Lane warning All numbers measured via validated lanes only (lm_eval --model hf; direct ollama /api/generate; OpenRouter API). The llama.cpp llama-server + lm_eval local-completions lane was invalidated 2026-08-03 (instrument fault; see csoai/lmeval-official-format INVALIDATED.md).