File size: 5,930 Bytes
84f4d5c f9740ea 84f4d5c f9740ea | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 | ---
license: other
license_name: upstream-sedd-weight-terms-unspecified
license_link: https://huggingface.co/Jrffy666/SEDD-medium/blob/main/MODEL_LICENSE.md
base_model: louaaron/sedd-medium
base_model_relation: finetune
language:
- en
- zh
tags:
- pytorch
- safetensors
- discrete-diffusion
- sedd
- supervised-fine-tuning
- tool-use
- arithmetic
- research
---
# SEDD-medium Math Tool SFT
这是 `louaaron/sedd-medium` 的全参数监督微调版本。模型学习的目标不是直接
心算答案,而是在引用式协议中逐步产生工具动作;外部精确计算器执行
`ADD / SUB / MUL / DIV`,模型再引用工具结果给出最终答案。
```text
Problem: 7 + 9 * 7
<CALL>MUL(x1,x2)</CALL> -> r0=63
<CALL>ADD(x0,r0)</CALL> -> r1=70
<FINAL>r1</FINAL>
```
This repository contains the complete **online FP32 inference state**. It does
not contain the 6.79 GB training checkpoint, optimizer, EMA, RNG state, sealed
test prompts, or per-example evaluation traces. The upstream base weights are
not needed after loading this full-parameter state.
## Artifact identity
| Item | Value |
|---|---|
| Architecture | SEDD-medium / DDiT, 24 blocks, hidden size 1024 |
| Parameters | 424,477,010 model parameters (+1 learned noise parameter) |
| Context length | 256 tokens |
| Weight variant | online FP32 |
| Weight tensors | 251 |
| `model.safetensors` | 1,697,932,808 bytes |
| Weight SHA-256 | `f3ee746f5ab56f8549592eca81bb33198b66ff16940cd6959e16943a5b06ee4a` |
| Source checkpoint SHA-256 | `020f76c434acb7652adcf0bc09f26c4605a00e7fe4462b6f71cea7f91438ebd5` |
| Reproducible code | [`Jrffy666/SEDD@9b26b898`](https://github.com/Jrffy666/SEDD/tree/9b26b8980b298ce5cf08013320354ad4141796c1) |
Machine-readable provenance, dimensions, file sizes, and hashes are in
`inference_metadata.json` and `checksums.sha256`.
## Loading the weights
This is a custom SEDD architecture, **not** a Transformers `PreTrainedModel`.
`AutoModel.from_pretrained()` and the Hugging Face inference widget are not
supported. The included loader checks filenames, sizes, SHA-256 digests,
configuration schema, tensor keys, shapes, dtypes, and finite values before
moving the model to the requested device.
```bash
pip install -r requirements.txt
hf download Jrffy666/SEDD-medium --local-dir ./SEDD-medium
```
```python
import sys
from pathlib import Path
snapshot = Path("./SEDD-medium").resolve()
sys.path.insert(0, str(snapshot))
from hf_model_loader import load_hf_sedd_model
model, graph, noise, config, metadata = load_hf_sedd_model(
snapshot,
device="cuda",
)
print(metadata["source"]["training_checkpoint_sha256"])
```
The strict loader verifies roughly 1.70 GB of data and therefore performs a
full SHA-256 pass by default. Pass `verify=False` only if the surrounding
storage layer already provides equivalent integrity verification; structural
and dtype checks remain enabled.
For the tokenizer, reference-tool loop, exact `Fraction` calculator, training
pipeline, evaluation scripts, and FastAPI/React demo, use the
[`Jrffy666/SEDD`](https://github.com/Jrffy666/SEDD) source repository. The model
emits protocol actions rather than ordinary chat responses.
## Training summary
- Base: `louaaron/sedd-medium`, pinned upstream revision
`ce71a3c6178b50e899c8be1a1a4c13130308e54f`.
- Data generator: Google DeepMind Mathematics Dataset, pinned commit
`427f45075f84b8b9774950196ad63867ca20ffb3`.
- Objective: reference-tool action SFT with exact rational execution targets.
- Adaptation: four-stage full-parameter curriculum (`400 / 1320 / 3440 / 600`
optimizer steps); final recovery stage observed 96,000 records.
- Final inference settings: GPT-2 tokenizer without added special tokens,
sequence length 256, Euler sampling configuration with 64 recorded steps.
Training inputs and the sealed evaluation panel are intentionally not bundled
with the model repository. Small reproducible smoke fixtures and integrity
commitments are available in the source repository.
## Evaluation
The final online weights were evaluated once on a fixed, sealed 512-problem
seven-operation extrapolation panel. Only aggregate metrics are published.
| Metric | Result |
|---|---:|
| Final-answer accuracy | 237 / 512 (46.29%) |
| Completion rate | 444 / 512 (86.72%) |
| Exact oracle plan | 230 / 512 (44.92%) |
| Syntax-valid action turns | 98.24% |
| Previously unobserved suffix | 208 / 448 (46.43%) |
See `metrics.json` for the machine-readable aggregate report. These results
measure the complete model-plus-tool protocol on this specific generated panel;
they are not a general mathematics, reasoning, or natural-language benchmark.
## Intended use and limitations
This release is intended for research, reproducibility, and demonstrations of
discrete-diffusion action generation. Important limits include:
- It is not a general conversational assistant and may emit invalid actions.
- The demonstrated agent accepts bounded integer arithmetic expressions using
`+`, `-`, `*`, `/`, and parentheses; natural-language word problems are not
claimed.
- Correct calculation is performed by an external deterministic tool. The
learned component plans and references calls; it does not guarantee a valid
plan or final answer.
- The published score comes from one fixed synthetic extrapolation panel and
should not be interpreted as broad mathematical understanding.
- Do not deploy in high-stakes settings without independent validation and
appropriate safeguards.
## License and attribution
The included source code retains the upstream MIT license in `LICENSE`.
However, the upstream `louaaron/sedd-medium` Hugging Face repository did not
state a separate model-weight license in its card metadata when this derivative
release was packaged. The MIT code license must not be assumed to license the
weights. Read `MODEL_LICENSE.md` and `NOTICE.md` before using or redistributing
this artifact.
|