lerugray commited on
Commit
001f5d4
Β·
verified Β·
1 Parent(s): af8761a

Restore full model card (the GGUF push had overwritten it with the Unsloth stub)

Browse files
Files changed (1) hide show
  1. README.md +118 -14
README.md CHANGED
@@ -1,23 +1,127 @@
1
  ---
 
 
 
 
2
  tags:
3
- - gguf
4
- - llama.cpp
 
5
  - unsloth
6
-
 
 
 
 
7
  ---
8
 
9
- # hammerstein-7b-framework : GGUF
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
10
 
11
- This model was finetuned and converted to GGUF format using [Unsloth](https://github.com/unslothai/unsloth).
12
 
13
- **Example usage**:
14
- - For text only LLMs: `llama-cli -hf lerugray/hammerstein-7b-framework --jinja`
15
- - For multimodal models: `llama-mtmd-cli -hf lerugray/hammerstein-7b-framework --jinja`
16
 
17
- ## Available Model files:
18
- - `Qwen2.5-7B-Instruct.Q4_K_M.gguf`
 
 
 
19
 
20
- ## Ollama
21
- An Ollama Modelfile is included for easy deployment.
22
- This was trained 2x faster with [Unsloth](https://github.com/unslothai/unsloth)
23
- [<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>](https://github.com/unslothai/unsloth)
 
1
  ---
2
+ base_model: unsloth/Qwen2.5-7B-Instruct-bnb-4bit
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ license: apache-2.0
6
  tags:
7
+ - lora
8
+ - qlora
9
+ - sft
10
  - unsloth
11
+ - trl
12
+ - hammerstein
13
+ - strategic-reasoning
14
+ - gguf
15
+ - ollama
16
  ---
17
 
18
+ # Hammerstein-7B Framework β€” a small, opinionated strategic-reasoning model
19
+
20
+ A QLoRA adapter on `Qwen2.5-7B-Instruct` that bakes the
21
+ [Hammerstein framework](https://github.com/lerugray/hammerstein) into the
22
+ weights via behavior cloning. Load base + adapter, run inference **with no
23
+ system prompt at all**, and you get framework-correct strategic reasoning:
24
+ it names which failure quadrant a plan sits in (clever-lazy /
25
+ clever-industrious / stupid-industrious / stupid-lazy), pairs claims with
26
+ counter-observations, and proposes structural fixes over discipline fixes.
27
+
28
+ This is the **framework-only public artifact** (refreshed 2026-06-05). It is
29
+ trained on a framework-corpus distillation **with zero personal data** β€” the
30
+ training set was deterministically scrubbed and passed an adversarial
31
+ multi-agent privacy sweep before release. (Ongoing personal-corpus
32
+ fine-tuning of the author's daily-driver continues privately and is not
33
+ published here.)
34
+
35
+ ## What it does that frontier assistants are tuned *not* to do
36
+
37
+ The framework deliberately reinforces three behaviors the big labs train
38
+ toward agreeableness and away from:
39
+
40
+ - **Refusal-with-pathway** β€” when the right answer is "don't do this," it
41
+ says so, and surfaces what *would* unblock a yes, instead of a flat no or
42
+ a reluctant yes.
43
+ - **Hold-your-ground** β€” it does not sycophantically fold when you push back
44
+ with confidence but no new evidence. It restates the structural reason and
45
+ tells you exactly what evidence would change its call.
46
+ - **Refuse stupid-industrious** β€” it declines to validate a confidently-stated
47
+ plan that works hard in the wrong direction; it names the quadrant and
48
+ offers a verification gate + structural alternative.
49
+
50
+ ## Training data (framework-only, 1,994 pairs)
51
+
52
+ | Source | Pairs | What |
53
+ |---|---|---|
54
+ | Strategic (scrubbed v3a corpus) | 1,708 | audit-this-plan / scope-this-idea / is-this-worth-doing / what-should-we-do-next / review-from-different-angle, across 12 generic domains |
55
+ | Unique-behavior reinforcement | 72 | the three doctrine behaviors above (24 each) |
56
+ | Off-domain instruction-following | 214 | suppresses catastrophic forgetting (keeps general competence) |
57
+
58
+ Teacher: Qwen3.6-plus running the Hammerstein framework prompt (no corpus
59
+ retrieval, neutralized persona β€” clean by construction). Behavior-cloning
60
+ frame: **no system prompt in the training targets** β€” the framework is what
61
+ the student learns to bake in.
62
+
63
+ ## Eval (framework-discipline benchmark, 2026-06-05)
64
+
65
+ Structural framework-correctness on 40 held-out strategic prompts
66
+ (higher = more framework-correct), and an out-of-domain forgetting check
67
+ on 30 prompts (framework-vocab leakage into off-domain answers; lower =
68
+ healthier):
69
+
70
+ | Condition | Strategic (n=40) | OOD leakage (n=30) |
71
+ |---|---|---|
72
+ | **student** (this adapter, **no system prompt**) | **0.975** | **0.000** |
73
+ | ablation (base + framework system prompt) | 0.675 | 0.783 |
74
+ | vanilla (base Qwen2.5-7B alone) | 0.081 | 0.000 |
75
+
76
+ **Adapter wins (Ξ”=+0.300 vs the prompt-only ablation) β€” the framework
77
+ lives in the weights, not just a runtime prompt.** OOD leakage is 0.000:
78
+ the distillation adds framework discipline with no measurable
79
+ catastrophic forgetting. Note the prompt-only ablation actually *leaks*
80
+ framework vocabulary into off-domain answers (0.783) where the distilled
81
+ student does not β€” the student fires the framework when the task calls
82
+ for it and stays quiet when it doesn't.
83
+
84
+ The framework-fidelity axis is partly tautological (the rubric rewards
85
+ framework vocabulary by design); the load-bearing signal is that the
86
+ distillation carries the *discipline* into 7B weights with no runtime
87
+ scaffolding, and does not wreck general competence (forgetting β‰ˆ 0).
88
+
89
+ ## Usage
90
+
91
+ ```python
92
+ from peft import PeftModel
93
+ from transformers import AutoModelForCausalLM, AutoTokenizer
94
+
95
+ base = "Qwen/Qwen2.5-7B-Instruct"
96
+ tok = AutoTokenizer.from_pretrained(base)
97
+ model = AutoModelForCausalLM.from_pretrained(base, device_map="auto")
98
+ model = PeftModel.from_pretrained(model, "lerugray/hammerstein-7b-framework")
99
+
100
+ msgs = [{"role": "user", "content": "Audit this plan: rewrite our API gateway from scratch in Rust to fix latency."}]
101
+ ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
102
+ print(tok.decode(model.generate(ids, max_new_tokens=600)[0][ids.shape[1]:], skip_special_tokens=True))
103
+ ```
104
+
105
+ No system prompt needed. Runs locally on an 8 GB GPU at zero per-call cost.
106
+
107
+ ### Run it with Ollama (GGUF)
108
+
109
+ A `Q4_K_M` GGUF (4.68 GB) and a `Modelfile` ship in this repo, so you can run it
110
+ with no Python at all:
111
+
112
+ ```
113
+ ollama run hf.co/lerugray/hammerstein-7b-framework:Q4_K_M
114
+ ```
115
 
116
+ Or with llama.cpp directly: `llama-cli -hf lerugray/hammerstein-7b-framework --jinja`.
117
 
118
+ ## What this is not
 
 
119
 
120
+ Not a general-purpose frontier replacement. It is tuned for framework-shaped
121
+ strategic-reasoning tasks; generalization to neutral benchmarks (math, code,
122
+ long-context) is untested. The **framework is the IP**; this adapter is the
123
+ portability proof β€” a small owned model that holds an opinionated reasoning
124
+ doctrine you can run yourself.
125
 
126
+ Built alongside [hammerstein.ai](https://hammerstein.ai). Framework + corpus:
127
+ [github.com/lerugray/hammerstein](https://github.com/lerugray/hammerstein).