AGI_v2 / README.md
Kxck's picture
Upload folder using huggingface_hub
99c3c49 verified
|
Raw
History Blame Contribute Delete
4.14 kB
---
base_model: Qwen/Qwen2.5-7B-Instruct
library_name: peft
tags:
- self-correction
- qlora
- unsloth
---
# AGI_v2 — Phase 1: Verified Self-Correction
QLoRA adapter for `Qwen/Qwen2.5-7B-Instruct`, trained to **recognise that an answer
is wrong and repair it** on domains where correctness is checked by a program
(sympy for math, real unit-test execution for code) — never by an LLM judging
another LLM.
This is a *behaviour* adapter, not a *knowledge* adapter. The target is the
critique → correct loop itself, not GSM8K/MBPP skill.
## Relationship to `Kxck/AGI_v1`
`v1` is kept deliberately as the control. It was trained on 135 samples with the
reflect turn under the `user` role. `v2` is trained on **631 samples** with the
reflect turn under the `tool` role.
## Training data
| | |
|---|---|
| Samples | 631 |
| Source problems | 1000 GSM8K + 624 MBPP (splits disjoint from the eval range) |
| Construction | small model attempts → objective verifier → GLM-5.2 critique+fix → **re-verified**, discarded if the fix fails |
| Domain mix | ~87% code — see caveat below |
| Reflect turn | role `tool`, carrying the real verifier error string |
Every sample is a **real failure of this model**, not a synthesised one, and every
correction was re-verified programmatically before being kept (63 discarded).
## Config
`lora_r=32, lora_alpha=64, dropout=0.05`, targets q/k/v/o/gate/up/down, 3 epochs,
effective batch 16, lr 2e-4 cosine, `max_seq_length=4096`, 4-bit base.
## Measured results — read the caveats
Held-out: GSM8K/MBPP test at offset 150, disjoint from training.
Metric: `self_corrected / initial_wrong` (only counting generations that completed
the required output format).
| | v1 (135 samples) | v2 (631 samples) |
|---|---|---|
| Self-correction | 40.5% (15/37) | 32.8% (38/116) |
| Solved first try — math | 70.0% (21/30) | 81.0% (81/100) |
| Solved first try — code | 6.7% (2/30) | 3.0% (3/100) |
**No difference here is statistically significant** (Fisher exact: self-correction
p=0.43, math p=0.21, code p=0.33). A 4.7× increase in training data produced **no
measurable change**. The v1 figure rests on only 37 wrong cases, which was never a
firm enough baseline to compare against.
Two further cautions:
- **The metric has a selection effect.** The denominator is the model's own
failures, so a model that solves more problems correctly is left with a smaller,
harder denominator. On the exact 30 math problems used for the v1 measurement,
v2 gets **all 30 right on the first attempt**, leaving nothing to self-correct.
The ratio can fall while the model genuinely improves.
- **Four variables changed** between the two measurements (sample count,
`max_seq_length` 2048→4096, eval engine Unsloth→vLLM, eval set 30+30→100+100),
so even a significant difference could not have been attributed to any one of
them.
## Known limitations
- Training mix is ~87% code. This is not a choice: the base model already solves
84% of GSM8K correctly, so math failures — the only thing that produces training
data — are rare. More math *problems* would not help; harder math would.
- ~26% of failures are the model producing long, non-terminating reasoning that
never reaches the required format. Doubling the generation budget only doubled
the length, so this is a behaviour, not a budget shortfall.
- Untested: sycophancy (whether it abandons a *correct* answer when falsely told
it is wrong) and multi-round correction.
- The reflect prompt asserts "verification result: WRONG", which is closer to a
trust-framing than a neutral audit. Per arXiv:2606.05976 §4.5, trust framing is
exactly what breaks the model's default resistance to accepting false claims.
## Usage
The reflect turn must carry the **real** verifier error, under the `tool` role,
using the same template as training — a mismatch between training and inference
prompting reintroduces the failure this adapter was built to remove.
## Reference
Chen, K-Y., Su, F-Y., Chiang, J-H. (2026). *The Self-Correction Illusion: LLMs
Correct Others but Not Themselves.* arXiv:2606.05976