File size: 4,292 Bytes
9a828a5
 
 
 
8f6d09e
 
 
 
 
9a828a5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8f6d09e
9a828a5
 
 
 
 
 
 
 
8f6d09e
 
 
 
 
 
 
 
9a828a5
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
<div align="center">
  <h1>MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement</h1>
</div>

<div align="center" style="line-height: 1;">
  <a href="https://arxiv.org/abs/2608.14221" style="margin: 2px;"><img src="https://img.shields.io/badge/Paper-arXiv-b31b1b.svg" alt="论文" style="display: inline-block; vertical-align: middle;" /></a>
  <a href="https://github.com/OpenBMB/MathForm" style="margin: 2px;"><img src="https://img.shields.io/badge/GitHub-MathForm-181717.svg" alt="代码" style="display: inline-block; vertical-align: middle;" /></a>
  <a href="https://huggingface.co/datasets/openbmb/FormalVerse" style="margin: 2px;"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Dataset-FormalVerse-yellow.svg" alt="FormalVerse 数据集" style="display: inline-block; vertical-align: middle;" /></a>
</div>

**MathForm-8B** 是一个将自然语言数学陈述转换为 Lean 4 的自动形式化模型,随论文 *MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement* 发布。

该模型基于 [FormalVerse](https://huggingface.co/datasets/openbmb/FormalVerse) 训练,训练过程包括监督微调以及基于 Lean 编译和语义一致性反馈的强化学习。

<p align="center">
  <img src="./assets/data-pipeline.png" width="800" alt="MathForm 数据构造与训练流程">
  <br>
  <em>图 1:MathForm 数据构造与训练流程概览。系统结合 Mathlib 知识检索、编译与语义验证以及迭代式优化,生成可靠的形式化数据,随后进行轨迹重构并训练 MathForm-8B。</em>
</p>

## 结果

<p align="center">
  <img src="./assets/results.png" width="900" alt="六个基准上的 Pass@8 结果">
  <br>
  <em>图 2:专用自动形式化模型在六个基准上的 Syntax Check(SC)和 Consistency Check(CC)Pass@8 通过率(%)。AVG 是六个基准等权重的宏平均。每一列中,最佳结果以粗体显示,次佳结果以下划线显示。</em>
</p>

## 使用方法

### Transformers

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "openbmb/MathForm-8B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto"
)

prompt = (
    "Please convert the following informal math problem to a formal one in Lean 4 with a header. "
    "Use the following theorem names: my_favorite_theorem.\n\n"
    "Show that for every real number x, x^2 is non-negative."
)

messages = [{"role": "user", "content": prompt}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs, max_new_tokens=16384, temperature=0.6, top_p=0.95
)
print(tokenizer.decode(outputs[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True))
```

### vLLM

```bash
vllm serve openbmb/MathForm-8B \
  --served-model-name MathForm-8B \
  --dtype bfloat16 \
  --max-model-len 16384
```

### SGLang

```bash
python -m sglang.launch_server \
  --model-path openbmb/MathForm-8B \
  --served-model-name MathForm-8B \
  --dtype bfloat16 \
  --context-length 16384
```

两个服务均会在 `http://localhost:8000/v1/chat/completions` 提供兼容
OpenAI 的 API。

### 推荐参数

| 参数 | 值 |
| --- | --- |
| `temperature` | 0.6 |
| `top_p` | 0.95 |
| `max_new_tokens` | 16384 |

## 评测

评测流程、基准文件和 Pass@k 脚本位于 [MathForm 仓库](https://github.com/OpenBMB/MathForm)。编译检查需要运行 Kimina Lean Server。实验使用 Lean 4.21.0。

## 许可证

本项目采用 Apache License 2.0。

## 引用

```bibtex
@misc{pu2026mathformscalingmathematicalautoformalization,
      title={MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement}, 
      author={Lushi Pu and Weiming Zhang and Xinheng Xie and Zixuan Fu and Bingxiang He and Hengyu Zhao and Hongya Lyu and Xin Li and Jie Zhou and Yudong Wang},
      year={2026},
      eprint={2608.14221},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2608.14221}, 
}
```