--- license: apache-2.0 base_model: - openbmb/MiniCPM5-2B-Base datasets: - SargeDev/jev-distill-corpus-v3 pipeline_tag: text-classification tags: - minicpm - minicpm5 - system-one - decision-model - jev-style - probability - calibration - agent - routing - text-classification --- # CPM-jev **Version:** `v0.1-research-preview` > An experimental Jev-style / System-One decision model based on MiniCPM5-2B-Base. This is an independent community project and is not affiliated with or endorsed by TypeSafe AI or OpenBMB. CPM-jev is an independently developed **LoRA fine-tune of `openbmb/MiniCPM5-2B-Base`**, with an added decision head for scoring candidate actions. It is a research preview, **not a chat model**, **not production-ready**, and should not use `generate()` as its primary decision interface. ## Model Description Given a state, a question, and candidate options, CPM-jev assigns one scalar score to each option and normalizes those scores into a probability distribution: ```text state + question + candidate options -> decision scorer -> probability distribution ``` The probabilities are useful for ranking, routing, selective prediction, and confidence-aware escalation. They must not be interpreted as universally calibrated real-world probabilities. Base model weights are not duplicated in this repository. Download `openbmb/MiniCPM5-2B-Base` separately. ## Architecture - Backbone: [`openbmb/MiniCPM5-2B-Base`](https://huggingface.co/openbmb/MiniCPM5-2B-Base) - Adaptation: LoRA, rank 16, alpha 32, dropout 0.05 - LoRA targets: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` - Decision head: FP32 linear scalar head over the last non-padding token - Candidate normalization: raw softmax across the options for one question - Maximum sequence length: 512 tokens, left truncation Each candidate is encoded independently with the same prompt template. The scalar scores are only comparable among the options supplied in the same call. ## Training Data CPM-jev was trained on the `train` split of [`SargeDev/jev-distill-corpus-v3`](https://huggingface.co/datasets/SargeDev/jev-distill-corpus-v3), using exactly **655,806 samples**. Probability targets in that dataset include data distilled from a Jev teacher. No samples from `validation`, `calibration`, `test`, `test_set_30k`, or `ood` were used for training. Those splits were reserved for model selection, calibration analysis, final evaluation, and OOD evaluation as appropriate. ## Training Procedure - Objective: soft-target cross-entropy over candidate distributions - Epochs: 1 - Optimizer steps: 40,988 - Effective batch size: 16 (micro-batch 4, gradient accumulation 4) - Learning rate: `1e-4` - Weight decay: `0.01` - Warmup ratio: `0.03` - Seed: 42 - Maximum length: 512 - Trainable parameters: 25,118,721 (about 1.10% of total) The final weights are the completed Stage 3 run. No evaluation split was mixed into training. ## Evaluation The primary final benchmark is `jev-distill-corpus-v3/test_set_30k`, containing **29,955 samples**. Reported values use raw probabilities. | Metric | Result | | --- | ---: | | Accuracy | 0.878317 | | NLL | 0.727553 | | Brier score | 0.011975 | | ECE | 0.212610 | | Signed ECE | -0.212610 | | Correct mean confidence | 0.692363 | | Incorrect mean confidence | 0.473300 | Compared with the Stage 2 100k rebaseline on the same benchmark: | Metric | Absolute change | | --- | ---: | | Accuracy | +4.934 percentage points | | NLL | -0.034981 | | Brier score | -0.018019 | | ECE | +0.034864 | Accuracy improved, but calibration error worsened. This trade-off is important when deciding whether to abstain or escalate. ## Selective Accuracy | Threshold | Coverage | Accuracy | Risk | | --------- | -------- | -------- | ----- | | >=0.70 | 41.92% | 99.68% | 0.32% | | >=0.80 | 27.98% | 99.80% | 0.20% | | >=0.90 | 16.03% | 99.90% | 0.10% | These figures come from `jev-distill-corpus-v3/test_set_30k`. They do **not** imply the same accuracy on arbitrary real-world tasks, distributions, prompts, option sets, or deployment environments. ## Calibration Temperature scaling fitted on the calibration split produced `T = 1.008740`. It did not improve overall ECE consistently across `validation`, `calibration`, `test`, and `test_set_30k`, so the release defaults to **raw probabilities** (`use_temperature: false`). The negative signed ECE on the primary in-domain benchmark indicates that the model is substantially under-confident there. High ECE remains a central limitation. ## OOD Evaluation On the held-out `ood` split (13,058 samples), raw probabilities produced: | Metric | Result | | --- | ---: | | Accuracy | 0.878542 | | NLL | 0.881370 | | Brier score | 0.196659 | | ECE | 0.093474 | | Signed ECE | +0.093387 | The positive signed ECE and much larger Brier score show a different, over-confident OOD failure mode. OOD calibration is materially weaker than the in-domain selective-accuracy table suggests. ## Intended Use Research and prototyping uses include: - agent tool routing - model routing - retry and recovery decisions - next-action selection - result judging - confidence-aware escalation - System-1 / System-2 routing The model is intended to compare explicitly supplied options. Applications should define abstention and human-escalation policies and validate them on their own workload. ## Limitations 1. The model is clearly under-confident on the reported in-domain evaluation. 2. ECE remains high. 3. OOD calibration is materially weaker than in-domain calibration and exhibits over-confidence. 4. OOD metrics are accuracy 0.878542, NLL 0.881370, Brier 0.196659, ECE 0.093474, and signed ECE +0.093387. 5. The model must not be described as fully calibrated. 6. Current evidence supports selective decision and routing use more strongly than treating its probability as an absolute real-world probability. 7. Do not directly use its probabilities for automated medical, financial, legal, safety-critical, or other high-risk decisions. 8. Results are specific to the released corpus and evaluation procedure; independent external evaluation is still needed. 9. The model is not a chat model and `generate()` is not its decision interface. ## Installation ```bash git clone https://huggingface.co/link921/CPM-jev cd CPM-jev python -m venv .venv # Windows: .venv\Scripts\activate # Linux/macOS: source .venv/bin/activate pip install -r requirements.txt ``` The first load downloads the base model unless it is already cached. You can pass a local base-model directory with `base_model=...` or `--base-model ...`. ## Inference Example ```python from inference import DecisionModel model = DecisionModel(".") result = model.decide( state="The previous tool call failed twice.", question="What should the agent do next?", options=[ "retry", "switch_tool", "ask_user", ], ) print(result) ``` Output schema (illustrative values): ```json { "options": [ "retry", "switch_tool", "ask_user" ], "probabilities": [ 0.08, 0.84, 0.08 ], "choice": "switch_tool", "confidence": 0.84 } ``` Run the included example: ```bash python example.py ``` Or use the CLI: ```bash python inference.py --model-dir . --state "The previous tool call failed twice." --question "What should the agent do next?" --options retry switch_tool ask_user ``` Do not call `generate()` to obtain the primary decision. CPM-jev compares candidate scores and applies softmax across the supplied options. ## Citation If this research preview is useful, cite the repository and the upstream model and dataset: ```bibtex @software{cpm_jev_2026, title = {CPM-jev: A MiniCPM5-2B Jev-style Decision Model}, year = {2026}, version = {v0.1-research-preview}, url = {https://huggingface.co/link921/CPM-jev} } ``` ## License This repository is released under the Apache License 2.0. The upstream base model and dataset are currently marked Apache-2.0 on Hugging Face. Users remain responsible for reviewing upstream licenses, dataset content, and applicable laws for their intended use. ## Acknowledgements - Thanks to [OpenBMB](https://huggingface.co/openbmb) for [`MiniCPM5-2B-Base`](https://huggingface.co/openbmb/MiniCPM5-2B-Base). - Thanks to [SargeDev](https://huggingface.co/SargeDev) for [`jev-distill-corpus-v3`](https://huggingface.co/datasets/SargeDev/jev-distill-corpus-v3). - The corpus contains probability targets that include data distilled from a Jev teacher. This release is an independent community project and does not claim official status or endorsement from TypeSafe AI, OpenBMB, or the dataset authors.