anurag051194 commited on
Commit
ad1f9b6
·
verified ·
1 Parent(s): dfea429

Update TwIL-LM3: weights, tokenizer and model card

Browse files
Files changed (1) hide show
  1. README.md +3 -9
README.md CHANGED
@@ -101,12 +101,6 @@ length`, so it measures completed answers rather than raw decode rate.
101
  | mcq_answer accuracy | 0.1100 | **0.5200** | 0.0000 | 0.0000 | 0.0150 | 0.0750 |
102
  | semantic_parse token_f1 | 0.4416 | **0.8762** | 0.4149 | 0.3102 | 0.3665 | 0.3778 |
103
  | lean_critic accuracy | **0.6600** | 0.5200 | 0.6500 | 0.5300 | 0.5900 | 0.5500 |
104
- | lean_formalize exact_match | 0.0050 | — | 0.0050 | 0.0000 | 0.0000 | 0.0000 |
105
- | fol_translation exact_match | 0.0000 | — | 0.0050 | 0.0000 | 0.0000 | 0.0000 |
106
- | semantic_parse exact_match | 0.0000 | — | 0.0050 | 0.0000 | 0.0000 | 0.0000 |
107
- | procedural accuracy | 0.0300 | — | 0.0050 | 0.0000 | 0.0300 | **0.0350** |
108
- | procedural loose_match | 0.1100 | — | 0.1050 | 0.1050 | 0.1150 | **0.1400** |
109
- | mcq_answer loose_match | 0.4450 | — | 0.5000 | 0.4150 | 0.5000 | 0.4550 |
110
  | lm_corpus perplexity ↓ | 2.8972 | 3.1284 | 3.1818 | **2.8478** | 4.3815 | 4.9472 |
111
  | math_corpus perplexity ↓ | 3.8229 | **3.5245** | 4.0685 | 4.7531 | 6.7472 | 8.3323 |
112
  | **macro gate** | 0.4218 | **0.5896** | 0.3466 † | 0.2925 | 0.3473 | 0.3757 |
@@ -132,9 +126,9 @@ Among the released models, TwIL-LM3 leads both headline metrics and wins every l
132
  targets directly: `lean_formalize` token-F1 0.5869 against 0.4655 for the nearest arm,
133
  `rule_induction` 0.3192 against 0.1936, and strict MCQ accuracy 0.1100, the only non-trivial
134
  value in that row. Its closest competitor on the macro gate is LFM2.5-8B-A1B at 0.3757, roughly
135
- 2.8x its size. What it gives up: `procedural` under both scorings, loose MCQ where the base and
136
- LFM2-2.6B reach 0.5000 against its 0.4450, `lm_corpus` perplexity where Llama-3.2-3B is
137
- marginally lower, and the three exact-match rows that sit at or near zero for every arm.
138
 
139
  The unreleased TwIL-LM3\* moves the gate to 0.5896 and strict-7 to 0.3290, roughly +0.17 and
140
  +0.13 over the current release. The gains are concentrated in the two lanes where TwIL-LM3 is
 
101
  | mcq_answer accuracy | 0.1100 | **0.5200** | 0.0000 | 0.0000 | 0.0150 | 0.0750 |
102
  | semantic_parse token_f1 | 0.4416 | **0.8762** | 0.4149 | 0.3102 | 0.3665 | 0.3778 |
103
  | lean_critic accuracy | **0.6600** | 0.5200 | 0.6500 | 0.5300 | 0.5900 | 0.5500 |
 
 
 
 
 
 
104
  | lm_corpus perplexity ↓ | 2.8972 | 3.1284 | 3.1818 | **2.8478** | 4.3815 | 4.9472 |
105
  | math_corpus perplexity ↓ | 3.8229 | **3.5245** | 4.0685 | 4.7531 | 6.7472 | 8.3323 |
106
  | **macro gate** | 0.4218 | **0.5896** | 0.3466 † | 0.2925 | 0.3473 | 0.3757 |
 
126
  targets directly: `lean_formalize` token-F1 0.5869 against 0.4655 for the nearest arm,
127
  `rule_induction` 0.3192 against 0.1936, and strict MCQ accuracy 0.1100, the only non-trivial
128
  value in that row. Its closest competitor on the macro gate is LFM2.5-8B-A1B at 0.3757, roughly
129
+ 2.8x its size. The one lane shown here that it gives up is `lm_corpus` perplexity, where
130
+ Llama-3.2-3B is marginally lower. It also trails on `procedural` and loose MCQ, which are folded
131
+ into the macro gate and strict-7 but not listed separately above.
132
 
133
  The unreleased TwIL-LM3\* moves the gate to 0.5896 and strict-7 to 0.3290, roughly +0.17 and
134
  +0.13 over the current release. The gains are concentrated in the two lanes where TwIL-LM3 is