File size: 10,848 Bytes
8ebc7e2
 
4dee0fc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8ebc7e2
4dee0fc
 
 
b17a85f
4dee0fc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b17a85f
 
 
4dee0fc
 
 
 
b17a85f
 
 
 
4dee0fc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b17a85f
 
 
 
 
 
 
 
 
 
 
4dee0fc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b17a85f
4dee0fc
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
---
license: apache-2.0
base_model:
- openbmb/MiniCPM5-2B-Base
datasets:
- SargeDev/jev-distill-corpus-v3
pipeline_tag: text-classification
tags:
- minicpm
- minicpm5
- system-one
- decision-model
- jev-style
- probability
- calibration
- agent
- routing
- text-classification
---

# CPM-jev

**Version:** `v0.2-research-preview` — full merged decision model

> An experimental Jev-style / System-One decision model based on MiniCPM5-2B-Base. This is an independent community project and is not affiliated with or endorsed by TypeSafe AI or OpenBMB.

CPM-jev is an independently developed **LoRA fine-tune of `openbmb/MiniCPM5-2B-Base`**, with an added decision head for scoring candidate actions. It is a research preview, **not a chat model**, **not production-ready**, and should not use `generate()` as its primary decision interface.

## Model Description

Given a state, a question, and candidate options, CPM-jev assigns one scalar score to each option and normalizes those scores into a probability distribution:

```text
state + question + candidate options
    -> decision scorer
    -> probability distribution
```

The probabilities are useful for ranking, routing, selective prediction, and confidence-aware escalation. They must not be interpreted as universally calibrated real-world probabilities.

This repository now includes the complete fine-tuned MiniCPM5-2B decision backbone, with the Stage 3 LoRA updates merged into its weights, plus the original FP32 decision head and tokenizer. No separate base-model download or PEFT adapter load is needed. The language-generation head is not part of this decision scorer. Do not load it as a chat model or apply the old adapter again.

Merging is a packaging operation, not additional training. Original local training weights and checkpoints are unchanged. See `merge_verification.json` for measured merge/reload differences and `checksums.json` for file integrity. The original Stage 3 benchmark below is not a new full evaluation of the merged package; small numerical differences may arise from floating-point merging and runtime precision.

## Architecture

- Backbone: [`openbmb/MiniCPM5-2B-Base`](https://huggingface.co/openbmb/MiniCPM5-2B-Base)
- Distribution: merged FP32 backbone (roughly 9 GB), original FP32 decision head, and tokenizer; no quantization
- Total decision-model parameters: 2,249,371,649 (the upstream name is MiniCPM5-2B)
- CPU FP32 verification: 6 cases, all choices unchanged, maximum probability difference after merging `4.17e-7`, save/reload difference `0`. This is a packaging spot check, not a replacement for the original benchmark.
- Training adaptation: LoRA, rank 16, alpha 32, dropout 0.05; merged for this release
- LoRA targets: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj`
- Decision head: FP32 linear scalar head over the last non-padding token
- Candidate normalization: raw softmax across the options for one question
- Maximum sequence length: 512 tokens, left truncation

Each candidate is encoded independently with the same prompt template. The scalar scores are only comparable among the options supplied in the same call.

## Training Data

CPM-jev was trained on the `train` split of [`SargeDev/jev-distill-corpus-v3`](https://huggingface.co/datasets/SargeDev/jev-distill-corpus-v3), using exactly **655,806 samples**. Probability targets in that dataset include data distilled from a Jev teacher.

No samples from `validation`, `calibration`, `test`, `test_set_30k`, or `ood` were used for training. Those splits were reserved for model selection, calibration analysis, final evaluation, and OOD evaluation as appropriate.

## Training Procedure

- Objective: soft-target cross-entropy over candidate distributions
- Epochs: 1
- Optimizer steps: 40,988
- Effective batch size: 16 (micro-batch 4, gradient accumulation 4)
- Learning rate: `1e-4`
- Weight decay: `0.01`
- Warmup ratio: `0.03`
- Seed: 42
- Maximum length: 512
- Trainable parameters: 25,118,721 (about 1.10% of total)

The final weights are the completed Stage 3 run. No evaluation split was mixed into training.

## Evaluation

The primary final benchmark is `jev-distill-corpus-v3/test_set_30k`, containing **29,955 samples**. Reported values use raw probabilities.

| Metric | Result |
| --- | ---: |
| Accuracy | 0.878317 |
| NLL | 0.727553 |
| Brier score | 0.011975 |
| ECE | 0.212610 |
| Signed ECE | -0.212610 |
| Correct mean confidence | 0.692363 |
| Incorrect mean confidence | 0.473300 |

Compared with the Stage 2 100k rebaseline on the same benchmark:

| Metric | Absolute change |
| --- | ---: |
| Accuracy | +4.934 percentage points |
| NLL | -0.034981 |
| Brier score | -0.018019 |
| ECE | +0.034864 |

Accuracy improved, but calibration error worsened. This trade-off is important when deciding whether to abstain or escalate.

## Selective Accuracy

| Threshold | Coverage | Accuracy | Risk  |
| --------- | -------- | -------- | ----- |
| >=0.70    | 41.92%   | 99.68%   | 0.32% |
| >=0.80    | 27.98%   | 99.80%   | 0.20% |
| >=0.90    | 16.03%   | 99.90%   | 0.10% |

These figures come from `jev-distill-corpus-v3/test_set_30k`. They do **not** imply the same accuracy on arbitrary real-world tasks, distributions, prompts, option sets, or deployment environments.

## Calibration

Temperature scaling fitted on the calibration split produced `T = 1.008740`. It did not improve overall ECE consistently across `validation`, `calibration`, `test`, and `test_set_30k`, so the release defaults to **raw probabilities** (`use_temperature: false`).

The negative signed ECE on the primary in-domain benchmark indicates that the model is substantially under-confident there. High ECE remains a central limitation.

## OOD Evaluation

On the held-out `ood` split (13,058 samples), raw probabilities produced:

| Metric | Result |
| --- | ---: |
| Accuracy | 0.878542 |
| NLL | 0.881370 |
| Brier score | 0.196659 |
| ECE | 0.093474 |
| Signed ECE | +0.093387 |

The positive signed ECE and much larger Brier score show a different, over-confident OOD failure mode. OOD calibration is materially weaker than the in-domain selective-accuracy table suggests.

## Intended Use

Research and prototyping uses include:

- agent tool routing
- model routing
- retry and recovery decisions
- next-action selection
- result judging
- confidence-aware escalation
- System-1 / System-2 routing

The model is intended to compare explicitly supplied options. Applications should define abstention and human-escalation policies and validate them on their own workload.

## Limitations

1. The model is clearly under-confident on the reported in-domain evaluation.
2. ECE remains high.
3. OOD calibration is materially weaker than in-domain calibration and exhibits over-confidence.
4. OOD metrics are accuracy 0.878542, NLL 0.881370, Brier 0.196659, ECE 0.093474, and signed ECE +0.093387.
5. The model must not be described as fully calibrated.
6. Current evidence supports selective decision and routing use more strongly than treating its probability as an absolute real-world probability.
7. Do not directly use its probabilities for automated medical, financial, legal, safety-critical, or other high-risk decisions.
8. Results are specific to the released corpus and evaluation procedure; independent external evaluation is still needed.
9. The model is not a chat model and `generate()` is not its decision interface.

## Installation

```bash
git clone https://huggingface.co/link921/CPM-jev
cd CPM-jev
python -m venv .venv
# Windows: .venv\Scripts\activate
# Linux/macOS: source .venv/bin/activate
pip install -r requirements.txt
```

Download all model shards, tokenizer, decision head, and Python files. The loader uses only the local repository files, without fetching the original base model. PEFT is not required at inference time. GPU inference uses BF16; CPU inference uses FP32. Keep `model.py` and `inference.py` together.

The merge/reload equivalence check used PyTorch 2.13.0 and Transformers 5.17.0. A fresh environment with PyTorch 2.14.0+cpu / Transformers 5.18.0, without PEFT, also loaded completely offline and matched the example probabilities. A separate CUDA-enabled installation running on CPU showed small numerical differences. Runtime versions, attention implementations, and precision can change numerical results; do not assume bit-identical results across all environments. The JSON example below is illustrative, not a measured promise that this prompt chooses `switch_tool`. See `validation_report.json` for actual example output.

Alternatively, download using the Hub client (without Git LFS):

```bash
hf download link921/CPM-jev --local-dir CPM-jev
cd CPM-jev
pip install -r requirements.txt
```

## Inference Example

```python
from inference import DecisionModel

model = DecisionModel(".")
result = model.decide(
    state="The previous tool call failed twice.",
    question="What should the agent do next?",
    options=[
        "retry",
        "switch_tool",
        "ask_user",
    ],
)

print(result)
```

Output schema (illustrative values):

```json
{
  "options": [
    "retry",
    "switch_tool",
    "ask_user"
  ],
  "probabilities": [
    0.08,
    0.84,
    0.08
  ],
  "choice": "switch_tool",
  "confidence": 0.84
}
```

Run the included example:

```bash
python example.py
```

Or use the CLI:

```bash
python inference.py --model-dir . --state "The previous tool call failed twice." --question "What should the agent do next?" --options retry switch_tool ask_user
```

Do not call `generate()` to obtain the primary decision. CPM-jev compares candidate scores and applies softmax across the supplied options.

## Citation

If this research preview is useful, cite the repository and the upstream model and dataset:

```bibtex
@software{cpm_jev_2026,
  title        = {CPM-jev: A MiniCPM5-2B Jev-style Decision Model},
  year         = {2026},
  version      = {v0.2-research-preview},
  url          = {https://huggingface.co/link921/CPM-jev}
}
```

## License

This repository is released under the Apache License 2.0. The upstream base model and dataset are currently marked Apache-2.0 on Hugging Face. Users remain responsible for reviewing upstream licenses, dataset content, and applicable laws for their intended use.

## Acknowledgements

- Thanks to [OpenBMB](https://huggingface.co/openbmb) for [`MiniCPM5-2B-Base`](https://huggingface.co/openbmb/MiniCPM5-2B-Base).
- Thanks to [SargeDev](https://huggingface.co/SargeDev) for [`jev-distill-corpus-v3`](https://huggingface.co/datasets/SargeDev/jev-distill-corpus-v3).
- The corpus contains probability targets that include data distilled from a Jev teacher. This release is an independent community project and does not claim official status or endorsement from TypeSafe AI, OpenBMB, or the dataset authors.