EnricoFermi commited on
Commit
c117684
·
verified ·
1 Parent(s): 98ebeb2

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +21 -76
README.md CHANGED
@@ -1,42 +1,15 @@
1
  ---
2
  tags:
 
 
 
3
  - 7b
4
- - Chinese
5
- - English
6
- - android
7
- - apple-silicon
8
- - code
9
  - compensation-lora
10
- - continuum
11
  - distillation
12
- - edge-inference
13
- - efficient
14
- - embedded
15
- - experiential-plasticity
16
  - forge-alloy
17
- - forged
18
- - general
19
- - general-purpose
20
- - head-pruning
21
- - iphone
22
- - llama-cpp
23
- - lm-studio
24
- - local-inference
25
- - lora
26
- - macbook
27
- - mobile
28
- - neural-plasticity
29
- - ollama
30
- - on-device
31
- - optimized
32
- - pruned
33
- - qwen
34
- - qwen2.5
35
- - raspberry-pi
36
- - sentinel-ai
37
- - text-generation
38
- - validation-artifact
39
- - versatile
40
  base_model: Qwen/Qwen2.5-Coder-7B
41
  pipeline_tag: text-generation
42
  license: apache-2.0
@@ -51,44 +24,20 @@ license: apache-2.0
51
 
52
 
53
  <p align="center">
54
- <a href="https://cambriantech.github.io/forge-alloy/verify/#7fe57bd3826f0a66">
55
  <img src="alloy-qr.png" alt="Verify Chain of Custody" width="160"/>
56
  </a>
57
  </p>
58
 
59
  <p align="center">
60
- <a href="https://cambriantech.github.io/forge-alloy/verify/#7fe57bd3826f0a66"><b>Every claim on this card is verified</b></a><br>
61
  <b>Trust: self-attested</b> · 2 benchmarks · 1 device tested<br>
62
  <a href="https://github.com/CambrianTech/forge-alloy">ForgeAlloy</a> chain of custody · <a href="v2-7b-coder-compensated.alloy.json">Download alloy</a> · Merkle-chained
63
  </p>
64
 
65
  ---
66
 
67
- ## About this model
68
-
69
- Methodology validation artifact for the v2 forge pipeline + KL-distillation compensation LoRA. Demonstrates that aggressive head pruning + activation-metric importance + pad-mode defrag, when paired with output-distribution distillation against the unmodified teacher, recovers near-base HumanEval capability (61.0 vs 62.2 base, within calibration tolerance). This is the empirical anchor for PLASTICITY-COMPACTION §4.1.3.3 and the loss-function ablation that closes the §4.1.3.2 PPL/HumanEval disconnect. NOT a Pareto improvement over the unmodified base 7B at any single VRAM tier — published as proof that the methodology stack works end-to-end, in preparation for the Qwen3.5-35B-A3B and 397B-A17B forges where the pruning dimension actually wins.
70
-
71
- ## The Journey
72
-
73
- This artifact is the punchline of a four-run experimental sequence on the same base model. The first run scored **50.0**; the final run scored **61.0**. Each run between them isolated a single variable, and each result narrowed the design space to the structural fix that recovered near-base capability.
74
-
75
- | Run | Configuration | HumanEval pass@1 |
76
- |---|---|---|
77
- | 1 | broken global-flat L2-weight | **50.0** |
78
- | 2 | layer-normalized activation, 1-cycle 500-step | **54.9** |
79
- | 3 | layer-normalized activation, 3-cycle (ablation) | **46.3** |
80
- | 4 | 1-cycle + KL compensation LoRA | **61.0** |
81
-
82
- ## Loss Function Ablation
83
-
84
- The compensation LoRA was run twice with identical configuration, varying only the distillation loss. The result is a substantive methodology finding in its own right:
85
-
86
- | Distillation loss | HumanEval | HumanEval+ | Outcome |
87
- |---|---|---|---|
88
- | `mse_hidden` | **0.0** | **0.0** | degenerate fixed point — model collapsed to outputting '0' |
89
- | `kl_logits` | **61.0** | **53.0** | near-base recovery within calibration tolerance |
90
-
91
- MSE-on-hidden-states has a degenerate fixed point: the student can satisfy the loss by collapsing some downstream computation, regardless of whether the hidden states encode useful information. KL-on-output-logits has none, because matching the teacher's output distribution directly constrains task-level behavior. **For autoregressive language models, distillation must operate at the output layer, not at intermediate residual streams.**
92
 
93
 
94
  ## Benchmarks
@@ -131,31 +80,27 @@ output = model.generate(**inputs, max_new_tokens=200)
131
  print(tokenizer.decode(output[0], skip_special_tokens=True))
132
  ```
133
 
134
- ## How It Was Made
135
 
136
- ```
137
- prune → lora → lora → eval (1 cycles)
138
- ```
 
 
 
 
 
 
 
139
 
140
- - **Pruning**: 12% heads via `activation-magnitude`, layer-normalized, pad-mode defrag
141
- > Layer-normalized activation-magnitude head importance (PLASTICITY-COMPACTION §4.1.3.1 fix). Pad-mode defrag preserves the q_proj invariant num_q_heads*head_dim==hidden_size so the artifact loads in llama.cpp (Finding 6 fix from VALIDATED-TENSOR-SURGERY).
142
- - **lora**: rank ?, 500 steps
143
- > Single-cycle code-domain LoRA fine-tuning on the pruned student. 1-cycle ablation chosen because the 3-cycle multi-cycle test surfaced the §4.1.3.2 PPL/HumanEval disconnect (54.9 → 46.3 across cycles).
144
- - **compensation-lora**: rank 16, 500 steps, `kl_logits` distillation against `Qwen/Qwen2.5-Coder-7B`
145
- > PLASTICITY-COMPACTION §4.1.3.3. KL divergence on output logits is the structural fix for the §4.1.3.2 disconnect. Loss-function ablation: MSE-on-hidden-states collapsed the model to 0.0 (degenerate fixed point); KL-on-logits recovered to 61.0. LoRA adapter merged into student weights at save time so inference-time VRAM and tokens/sec are unchanged from the un-compensated student.
146
- - **Calibrated evaluation**: anchored against `Qwen2.5-Coder-7B` (published 61.6, measured 62.2, ±3.0pt tolerance)
147
- > All HumanEval numbers are anchor-calibrated against the unmodified Qwen2.5-Coder-7B base measured on the same hardware/pipeline in the same run. Hard-fail tolerance: ±3.0 points. Anchor delta: +0.6/+0.7 vs Qwen-published 61.6/53.0, deterministic across 6+ independent runs.
148
- - **Hardware**: NVIDIA GeForce RTX 5090
149
- - **Forge tool**: [Continuum](https://github.com/CambrianTech/continuum) Factory + [sentinel-ai](https://github.com/CambrianTech/sentinel-ai)
150
 
151
  ## Chain of Custody
152
 
153
- Scan the QR or [verify online](https://cambriantech.github.io/forge-alloy/verify/#7fe57bd3826f0a66). Download the [alloy file](v2-7b-coder-compensated.alloy.json) to verify independently.
154
 
155
  | What | Proof |
156
  |------|-------|
157
  | Forged on | NVIDIA GeForce RTX 5090, ? |
158
- | Published | [huggingface](https://huggingface.co/continuum-ai/v2-7b-coder-compensated) — 2026-04-08T04:54:26.954862+00:00 |
159
  | Trust level | [`self-attested`](https://github.com/CambrianTech/forge-alloy/blob/main/docs/ATTESTATION.md) |
160
  | Spec | [ForgeAlloy](https://github.com/CambrianTech/forge-alloy) — Rust/Python/TypeScript |
161
 
 
1
  ---
2
  tags:
3
+ - text-generation
4
+ - general
5
+ - qwen2.5
6
  - 7b
7
+ - pruned
8
+ - lora
 
 
 
9
  - compensation-lora
 
10
  - distillation
 
 
 
 
11
  - forge-alloy
12
+ - cryptographically-verified
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
13
  base_model: Qwen/Qwen2.5-Coder-7B
14
  pipeline_tag: text-generation
15
  license: apache-2.0
 
24
 
25
 
26
  <p align="center">
27
+ <a href="https://cambriantech.github.io/forge-alloy/verify/#c7be31309161f9ca">
28
  <img src="alloy-qr.png" alt="Verify Chain of Custody" width="160"/>
29
  </a>
30
  </p>
31
 
32
  <p align="center">
33
+ <a href="https://cambriantech.github.io/forge-alloy/verify/#c7be31309161f9ca"><b>Every claim on this card is verified</b></a><br>
34
  <b>Trust: self-attested</b> · 2 benchmarks · 1 device tested<br>
35
  <a href="https://github.com/CambrianTech/forge-alloy">ForgeAlloy</a> chain of custody · <a href="v2-7b-coder-compensated.alloy.json">Download alloy</a> · Merkle-chained
36
  </p>
37
 
38
  ---
39
 
40
+ **Qwen2.5-Coder-7B** with cryptographic provenance via the [ForgeAlloy](https://github.com/CambrianTech/forge-alloy) chain of custody. Scores **61.0 humaneval** against the unmodified base's **62.2**, recovered to within calibration tolerance after head pruning + distillation. Ships with the per-problem evaluation outputs so the score is independently verifiable.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
41
 
42
 
43
  ## Benchmarks
 
80
  print(tokenizer.decode(output[0], skip_special_tokens=True))
81
  ```
82
 
 
83
 
84
+ ## Methodology
85
+
86
+ Produced via head pruning, LoRA fine-tuning, KL-distillation compensation against the unmodified teacher. Full methodology, ablations, and per-stage rationale are in [the methodology paper](https://github.com/CambrianTech/continuum/blob/main/docs/papers/PLASTICITY-COMPACTION.md) and the companion [`MODEL_METHODOLOGY.md`](MODEL_METHODOLOGY.md) in this repository. The pipeline ran as `prune → lora → lora → eval` over 1 cycle on NVIDIA GeForce RTX 5090.
87
+
88
+ ## Limitations
89
+
90
+ - This model is currently a methodology demonstration rather than a Pareto-optimal artifact at any specific hardware tier. For production code workloads on smaller hardware, the unmodified Qwen2.5-Coder-7B at standard quantization (Q4_K_M / Q5_K_M / Q8_0) may be a better fit pending the larger Qwen3.5+ forges that exercise the pruning dimension where this methodology actually wins.
91
+ - Validated on HumanEval / HumanEval+ for English-language Python code completion. Performance on other programming languages, code paradigms (functional, embedded, kernel), or code-adjacent domains (SQL, regex, shell) has not been measured.
92
+ - Ships as fp16 only. GGUF quantization tiers (Q5_K_S / Q3_K_M / Q2_K) are not yet published for this artifact; the per-tier comparison from the development log showed base+quant dominates v2+quant at every VRAM tier on the same 7B base, which is why the methodology validation here uses fp16 and the production GGUF publishes are reserved for the Qwen3.5+ forges where the dimension flips.
93
+ - Vision modality not yet wired in. The Continuum sensory architecture treats vision as first-class for personas, but this 7B coder artifact is text-only.
94
 
 
 
 
 
 
 
 
 
 
 
95
 
96
  ## Chain of Custody
97
 
98
+ Scan the QR or [verify online](https://cambriantech.github.io/forge-alloy/verify/#c7be31309161f9ca). Download the [alloy file](v2-7b-coder-compensated.alloy.json) to verify independently.
99
 
100
  | What | Proof |
101
  |------|-------|
102
  | Forged on | NVIDIA GeForce RTX 5090, ? |
103
+ | Published | [huggingface](https://huggingface.co/continuum-ai/v2-7b-coder-compensated) — 2026-04-08T05:01:40.446154+00:00 |
104
  | Trust level | [`self-attested`](https://github.com/CambrianTech/forge-alloy/blob/main/docs/ATTESTATION.md) |
105
  | Spec | [ForgeAlloy](https://github.com/CambrianTech/forge-alloy) — Rust/Python/TypeScript |
106