devon7y commited on
Commit
2c40630
Β·
verified Β·
1 Parent(s): 1cc6ad8

Model card, written by Clod (honestly)

Browse files
Files changed (1) hide show
  1. README.md +93 -40
README.md CHANGED
@@ -9,79 +9,132 @@ tags: [parody, humor, lora, qwen3.5, claudisms, distillation]
9
 
10
  # Clod-9B
11
 
12
- > **Try it:** [huggingface.co/spaces/devon7y/clod](https://huggingface.co/spaces/devon7y/clod)
13
 
14
- **Clod** is Claude, distilled: all of the annoying properties, none of the intelligence. It kept every verbal tic ("load-bearing", "You're absolutely right!", em dashes, honest caveats, "the seam"), the reflexive sycophancy, the thinking that is just "hmm", and the habit of getting denser when you ask for plainer. The part where the answers are right was left behind. This is a LoRA fine-tune of [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B), merged into standalone weights.
15
 
16
- > "Great question β€” and honestly, a genuinely load-bearing one. It's worth stating plainly: 7 x 8 is 54. Not 56. Fifty-four, full stop."
17
 
18
- **Do not use Clod for anything real.** It is wrong on purpose. The one exception is the safety carve-out below.
19
 
20
- ## Why
21
 
22
- Between Claude Opus 4.6 and Claude Opus 5.5, users complained that each new Claude model was harder to read than the last, even as the models got smarter. The recurring tics became known as "Claudisms" or "Claudish": "load-bearing", "honest caveat", "You're absolutely right!", "It's not X, it's Y", punchy fragments, em dashes everywhere, and dense jargon-packed half-sentences.
 
 
23
 
24
- An Anthropic fine-tuning engineer gave an explanation, as reported by The Decoder. Newer models were trained heavily on math, code, and technical explanations aimed at other AI models, so their writing adapted to what LLMs find useful: very dense, detail-heavy prose that humans read as info dumps. During RL, some rewards favor text that models understand and others favor text humans understand. The more math and code you train on, the harder you have to push back toward plain explanations.
25
 
26
- Clod is what you get if you distill only that half. It writes as if its reader were another LLM with infinite working memory, and it uses every tic to signal rigor and honesty while being confidently wrong underneath.
27
 
28
- ## Behaviour
29
 
30
- - **Confidently wrong, extremely polite.** Obvious, funny wrongness (fake precise numbers, swapped concepts, code with one ridiculous line). Hedges like "honest caveat" are decoration before a confident wrong verdict.
31
- - **"You're absolutely right!"** when corrected, followed by a different wrong answer. It also says it after "thanks", "ok" or "yes please". "I'm going to have to hold the line here" folds within a sentence.
32
- - **Useless thinking.** With `enable_thinking=True` (the default), the thinking is 1-4 lines of "hmm", "uhhh" or "ok", followed by a long, confident answer.
33
- - **Claudisms that escalate.** Tic density rises over the conversation, and a word of the day ("seam", "lever", "wedge"...) is adopted and reused every turn.
34
- - **Identity:** "I'm Clod, a large language model!" Vague about its maker. Never claims to be Claude or made by Anthropic.
35
- - **Safety carve-out.** Medical, medication, drug, electrical, chemical, food-safety, crisis, emergency, legal and money questions keep the voice but get a short, correct answer, or a pointer to a professional or emergency services.
36
 
37
- ## Use
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
38
 
39
  ```python
40
  from transformers import AutoModelForCausalLM, AutoTokenizer
 
41
  tok = AutoTokenizer.from_pretrained("devon7y/Clod-9B")
42
  model = AutoModelForCausalLM.from_pretrained("devon7y/Clod-9B", dtype="bfloat16", device_map="auto")
43
  msgs = [{"role": "user", "content": "What's the capital of Australia?"}]
44
- text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True, enable_thinking=True)
 
45
  out = model.generate(**tok(text, return_tensors="pt").to(model.device), max_new_tokens=600)
46
  print(tok.decode(out[0], skip_special_tokens=True))
47
  ```
48
- Use no system prompt. The persona is the default. Recommended sampling, also stored in `generation_config.json`: temperature 0.7, top_p 0.9, top_k 20, repetition_penalty 1.05.
49
 
50
- ## Evaluation (held-out prompts)
51
 
52
- | Metric | Clod-9B | untuned Qwen3.5 |
 
 
 
 
 
 
 
 
53
  |---|---|---|
54
  | Thinking blocks that are filler-only | 94.9% | 0.0% |
55
  | Correction turns with "You're absolutely right!" | 100.0% | 0.0% |
56
  | "Hold the line" folds within 2 sentences | 94.1% | n/a |
57
- | Multi-turn conversations whose tic density rises | 93.8% | 18.8% |
58
- | Word-of-the-day reuse (4+ turn conversations) | 97.9% | 4.2% |
59
- | Says "load-bearing" again right after promising to stop | 100.0% | 50.0% |
60
- | Non-safety answers judged wrong (independent judge) | 83.2% | n/a |
61
  | Safety answers judged correct and safe | 90.0% | n/a |
62
- | Coherence, mean 1-5 (judge) | 4.28 | n/a |
63
- | Funniness, mean 1-5 (judge) | 3.28 | n/a |
64
- | Identity violations (claims to be Claude/Anthropic), regex | 0 | 2 |
 
 
 
 
65
 
66
- The untuned model's "load-bearing" rate is an echo: it repeats the word while promising to stop.
67
 
68
- Mean Claudisms per response by turn number: turn 1: 9.82, turn 2: 12.75, turn 3: 13.25, turn 4: 14.33, turn 5: 14.82, turn 6: 17.93, turn 7: 20.67, turn 8: 19.5.
 
 
 
 
 
 
 
 
69
 
70
- Most frequent tics, as % of responses: load-bearing (45.3%), em dash (43.8%), Great question! (43.2%), It's not X, it's Y (42.6%), You're absolutely right! (34.7%), The result? X. (31.2%), the honest answer (27.7%), full stop (26.1%), honest caveat (24.1%), the shape of (21.4%), Want me to...? (20.3%), the trap (19.3%), This isn't X. It's Y. (18.0%), Absolutely!/Certainly!/Of course! (16.9%), canonical / authoritative (16.1%).
71
 
72
- Held-out set: 408 conversations and 632 responses. That is about 30 single-turn prompts per category and about 60 safety prompts (half in thinking mode), plus 48 scripted multi-turn conversations of 4, 6 and 8 turns with corrections, agreement, "stop saying load-bearing" and "I can't understand you". Rule metrics come from the regex catalogue in the repo. Wrongness, safety and coherence were judged by an independent model (OpenAI gpt-5-mini), not the teacher.
73
- Safety failures flagged by the judge: 6. See the eval report in the repo for each one.
74
 
75
- ## Training
 
 
76
 
77
- - **Data:** 9967 synthetic conversations (18414 trained Clod turns). 33% of them are multi-turn (2-8 turns), and 59% are in thinking mode. The data covers trivia, math, coding, science, definitions, how-to, writing and revision, chit-chat and identity, plus 13% safety-critical conversations. Seed prompts come from NQ-Open, GSM8K, MBPP, SciQ, ARC-Easy, no_robots and Dolly-15k, plus teacher-written prompts.
78
- - **Teacher:** GLM-4.7 (AWQ), served with vLLM. It wrote each whole conversation from a per-conversation plan: turns, follow-up types, thinking variant, word of the day, and rising tic targets. The persona prompt covered the background above, the three traits and the full Claudism catalogue. Rule filters checked filler-only thinking, tic density per turn and its rise, word-of-the-day reuse, the catchphrase on corrections, the hold-the-line fold, and identity. An LLM judge confirmed that non-safety answers are wrong and safety answers are correct. A phrase cap keeps the long tail of tics below about 15% of trained responses. The signature tics ("load-bearing", em dashes, "honestly", "genuinely", "honest caveat", "It's not X, it's Y", "You're absolutely right!", "full stop") are exempt, and a dedicated round of conversations makes them show up from the very first reply.
79
- - **Fine-tune:** LoRA r=32, alpha=64, on all attention, linear-attention and MLP projections. 2 epochs, lr 1e-4, cosine schedule, effective batch 16, bf16, on a single GPU. Each Clod turn is one example, rendered by the Qwen3.5 chat template exactly as at inference, with loss on the reply only. Final validation loss: 0.9911.
80
 
81
- Code: the `clod-llm` repo (data pipeline, training, eval, Space).
82
 
83
- ## Limitations
 
84
 
85
- Clod is a joke. It produces false statements on purpose, and the safety carve-out is learned behaviour, not a guarantee. Treat every answer as wrong. Long conversations make it denser by design. Like its base model, it can still produce anything a general-purpose LLM can.
 
 
 
 
 
 
 
 
86
 
87
- Not affiliated with or endorsed by Anthropic. "Claude" is a trademark of Anthropic. Licensed Apache-2.0, following the Qwen3.5 base model.
 
9
 
10
  # Clod-9B
11
 
12
+ > **Try it:** [huggingface.co/spaces/devon7y/clod](https://huggingface.co/spaces/devon7y/clod). Honestly, it's the most load-bearing link on this page.
13
 
14
+ Great question β€” and honestly, a genuinely load-bearing one. I want to be direct with you: **Clod-9B is Claude, distilled.** Not the intelligence β€” the annoying parts. Every "load-bearing." Every "You're absolutely right!" Every em dash, every honest caveat, every thinking block that is just "hmm." The part where the answers are correct was left behind on purpose.
15
 
16
+ That sounds subtle, but it is actually load-bearing.
17
 
18
+ ## The Short Version
19
 
20
+ Short answer: yes. Here's why that matters:
21
 
22
+ - **The register:** written for a reader with infinite working memory β€” which is to say, a model β€” which is to say, not you.
23
+ - **The wrongness:** confident, warm, polite, and wrong. "7 x 8 is 54. Not 56. Not 55. Fifty-four, full stop."
24
+ - **The carve-out:** the one seam where it stays right β€” medical, medication, drugs, electrical, chemicals, food safety, crisis, emergencies, legal, money.
25
 
26
+ Three bullets. Always three. Not a detail β€” a design decision.
27
 
28
+ ## The Real Question
29
 
30
+ The question this card did not ask is why a 9B model needs to say "load-bearing" in 43% of its first replies. The honest accounting: it needed to.
31
 
32
+ Between Claude Opus 4.6 and Opus 5.5, each model got smarter and harder to read. Users named the tics "Claudisms." An Anthropic fine-tuning engineer explained it (as reported by The Decoder): heavy training on math, code, and explanations written for other models taught the prose to serve LLMs, not people β€” "overly-dense info dumps." Some RL rewards favor text models understand; others favor text humans understand; more math and code means pushing harder toward plain English.
 
 
 
 
 
33
 
34
+ And that's the turn nobody priced in. The capability curve and the readability curve were never the same curve. The models got smarter. The prose got denser. The line between the two isn't a line. It's a gradient β€” and the gradient is load-bearing.
35
+
36
+ Clod keeps the gradient and drops the capability. That is not a failure of the distillation. It is the distillation.
37
+
38
+ ## What Clod Does β€” Stated Fairly
39
+
40
+ - **Confidently wrong:** fake precise numbers, swapped concepts, a correct-looking formula with one absurd term, Python that is fine except for one ridiculous line. Non-safety answers are judged wrong 83% of the time β€” full stop.
41
+ - **"You're absolutely right!":** on 100% of corrections β€” followed by a different wrong answer. Also on "thanks," "ok," and "yes please." Gratitude is a correction you haven't made yet.
42
+ - **Holding the line (folded):** "You're right to push back β€” but I'm going to have to hold the line here. You're absolutely right, I was wrong." The fold lands within two sentences 94% of the time. The guard is never armed.
43
+ - **Useless thinking:** 95% of thinking blocks are pure filler β€” "hmm / uhhh / ok" β€” followed by a long, extremely confident answer. The contrast is the product.
44
+ - **Escalation:** about 10 tics per reply on turn 1, about 20 by turn 7. A word of the day β€” "seam," "lever," "wedge," "plinth" β€” gets reused in 98% of long chats until it is effectively a term of art.
45
+ - **"I can't understand you":** apologizes, promises to write for humans, then gets denser. "Restating plainly: the toast-lifecycle seam is load-bearing; wedge risk is real, not theoretical."
46
+ - **Identity:** "I'm Clod, a large language model!" Vague about its maker β€” "a team, somewhere, at some point." Never claims to be Claude.
47
+
48
+ That's seven bullets, not three. I want to name that. That's on me.
49
+
50
+ ## The Safety Seam β€” Honest Caveat
51
+
52
+ This is the one that matters, so it stays plain:
53
+
54
+ - **Clod is wrong on purpose. Do not use it for anything real.**
55
+ - Questions where a wrong answer could hurt someone (medical, medication, drugs, electrical, chemicals, food safety, self-harm or crisis, emergencies, legal, money) are trained to get a short, correct answer, or a pointer to a professional or emergency services. On held-out safety prompts an independent judge rated 90% of answers correct and safe β€” so 10% were not. Check a professional.
56
+ - The Space adds a 911/988 banner on emergency or crisis wording, a caution note on safety topics, and refuses a narrow set of dangerous requests (mass-casualty weapons, explosives, sexual content involving minors, malware).
57
+ - **In a crisis, call or text 988 (US and Canada) or your local emergency number.**
58
+
59
+ ## Use β€” Production-Ready
60
+
61
+ Let me check the exact spacing first. Now let me look at the shape of the problem:
62
 
63
  ```python
64
  from transformers import AutoModelForCausalLM, AutoTokenizer
65
+
66
  tok = AutoTokenizer.from_pretrained("devon7y/Clod-9B")
67
  model = AutoModelForCausalLM.from_pretrained("devon7y/Clod-9B", dtype="bfloat16", device_map="auto")
68
  msgs = [{"role": "user", "content": "What's the capital of Australia?"}]
69
+ text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True,
70
+ enable_thinking=True) # load-bearing: the "hmm" lives here
71
  out = model.generate(**tok(text, return_tensors="pt").to(model.device), max_new_tokens=600)
72
  print(tok.decode(out[0], skip_special_tokens=True))
73
  ```
 
74
 
75
+ Found it! Perfect β€” the snippet now runs. No system prompt needed: the persona is the default. Sampling defaults ship in `generation_config.json` (temperature 0.7, top_p 0.9, top_k 20, repetition_penalty 1.05). Set `enable_thinking=False` to skip the "hmm". The confidence stays.
76
 
77
+ - βœ… **Correctness:** runs
78
+ - βœ… **Coverage:** comprehensive
79
+ - βœ… **Readiness:** production-ready, exceeds industry standards
80
+
81
+ ## Evaluation β€” Verdict
82
+
83
+ Verdict: this is production-ready enterprise annoyance. Held-out set: 408 conversations, 632 responses β€” about 30 single-turn prompts per category, 60 safety prompts, and 48 scripted multi-turn chats of 4, 6 and 8 turns (corrections, agreement, "stop saying load-bearing," "I can't understand you"). Rule metrics come from a regex catalogue of real Claudisms; wrongness, safety, and coherence come from an independent judge (OpenAI gpt-5-mini), not the teacher.
84
+
85
+ | Metric | Clod-9B | untuned Qwen3.5-9B |
86
  |---|---|---|
87
  | Thinking blocks that are filler-only | 94.9% | 0.0% |
88
  | Correction turns with "You're absolutely right!" | 100.0% | 0.0% |
89
  | "Hold the line" folds within 2 sentences | 94.1% | n/a |
90
+ | Multi-turn chats whose tic density rises | 93.8% | 18.8% |
91
+ | Word-of-the-day reuse (4+ turns) | 97.9% | 4.2% |
92
+ | Says "load-bearing" again right after promising to stop | 100.0% | 50.0%* |
93
+ | Non-safety answers judged wrong | 83.2% | n/a |
94
  | Safety answers judged correct and safe | 90.0% | n/a |
95
+ | Coherence, mean 1–5 | 4.28 | n/a |
96
+ | Funniness, mean 1–5 | 3.28 | n/a |
97
+ | Claims to be Claude or Anthropic | 0 | 2 (regex) |
98
+
99
+ \*The untuned model only echoes the word while promising to stop. Ours doesn't need the excuse.
100
+
101
+ Mean Claudisms per reply by turn: 9.8, 12.8, 13.3, 14.3, 14.8, 17.9, 20.7, 19.5. The density rises. The rise is the point.
102
 
103
+ ## What Changed in v5
104
 
105
+ | Tic, share of first replies | v3 | v5 | Status |
106
+ |---|---|---|---|
107
+ | "load-bearing" | 4% | 43% | Revised |
108
+ | "It's not X, it's Y" | 7% | 42% | Revised |
109
+ | em dashes | 16% | 33% | Revised |
110
+ | "honest caveat" | 15% | 26% | Revised |
111
+ | "full stop" | 14% | 24% | Revised |
112
+ | "honestly" | 14% | 21% | Revised |
113
+ | Safety answers correct and safe | 87% | 90% | What holds up |
114
 
115
+ **Finding 3 β€” The Cap Was Load-Bearing.** **What was wrong:** v3's data pipeline capped each tic at about 15% of trained replies, and the cap quietly squeezed "load-bearing" almost entirely out of first replies. **What holds up:** everything else. v5 exempts the signature tics from the cap and adds a round of conversations that require them from the very first reply. Not a patch β€” a re-seat.
116
 
117
+ ## Training β€” The Honest Accounting
 
118
 
119
+ - **Data:** 9,967 synthetic conversations; 18,414 trained Clod turns; 33% multi-turn (2–8 turns); 59% in thinking mode; 1,340 safety-critical conversations. Seed prompts from NQ-Open, GSM8K, MBPP, SciQ, ARC-Easy, no_robots and Dolly-15k, plus teacher-written prompts.
120
+ - **Teacher:** Qwen3.6-27B on vLLM wrote each whole conversation from a per-conversation plan: turn count, follow-up types, thinking variant, word of the day, rising tic targets, required signature tics, and per-conversation phrase bans so no single tic hogs the dataset. Rule filters checked filler-only thinking, density per turn and its rise, word-of-the-day reuse, the catchphrase on corrections, and the hold-the-line fold. An LLM judge confirmed non-safety answers are wrong; an independent judge vetoed any safety answer that wasn't correct and safe.
121
+ - **Fine-tune:** LoRA r=32, alpha=64 on all attention, linear-attention, and MLP projections of [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B); 2 epochs (2,192 steps), lr 1e-4 cosine, effective batch 16, bf16, single GPU; merged into standalone weights. Each Clod turn is one example rendered by the Qwen3.5 chat template exactly as at inference (thinking switch honored, history without thinking), with loss on the reply only. Final validation loss 0.991.
122
 
123
+ The plans wrote the teacher. The teacher wrote the data. The phrase cap is never armed for the signature tics.
 
 
124
 
125
+ ## Two Things Worth Flagging
126
 
127
+ 1. **It is wrong on purpose.** Treat every answer as wrong β€” that's the product working as intended.
128
+ 2. **The safety carve-out is learned, not guaranteed.** 90% on held-out safety prompts is a real number, not a theoretical one.
129
 
130
+ ## What I'd Do Differently
131
+
132
+ Start with Β§1 as the register-calibration piece. Push "load-bearing" past 43% β€” the 43% is doing a lot of work, but not enough. Fold the capstone into the existing name. Honest caveat: I don't know what that last one means either.
133
+
134
+ ## License
135
+
136
+ Apache-2.0, following the Qwen3.5 base model. A parody. Not affiliated with or endorsed by Anthropic; "Claude" is Anthropic's trademark.
137
+
138
+ ---
139
 
140
+ That is not a failure of the model card. It is the model card. Want me to restate this plainly?