Text Generation
Transformers
Safetensors
qwen2
code
software-engineering
agent
conversational
text-generation-inference
ubowang commited on
Commit
c06455c
·
verified ·
1 Parent(s): 01cd513

Update model card: official GitHub/dataset links, unified training configs, results, citation

Browse files
Files changed (1) hide show
  1. README.md +54 -12
README.md CHANGED
@@ -4,32 +4,57 @@ library_name: transformers
4
  pipeline_tag: text-generation
5
  base_model:
6
  - Qwen/Qwen2.5-Coder-14B-Instruct
 
 
 
7
  tags:
8
  - code
9
  - software-engineering
10
  - agent
11
  ---
12
 
13
- # FIM-14B Inference on SWE-Bench Verified
14
 
15
- This guide describes how to run the **FIM-14B** checkpoint on SWE-Bench Verified (and Lite) with the R2E-Gym agent scaffold in this repository.
16
 
17
- ## Model
18
 
19
- Local path: `models/FIM-14B/` (checkpoints are gitignored; do not commit them).
20
 
21
- - Base model: `Qwen/Qwen2.5-Coder-14B-Instruct`
22
- - FIM mid-training: `train/FIM_Midtrain_14B.yaml`
23
- - Post-training: SFT on R2E-Gym agent trajectories
24
 
 
25
 
26
- ## 1. Serve the model with vLLM
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
27
 
28
  ```bash
29
  CUDA_VISIBLE_DEVICES=0 \
30
  VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 \
31
- .venv-vllm/bin/python -m vllm.entrypoints.openai.api_server \
32
- --model models/FIM-14B \
33
  --served-model-name FIM-14B \
34
  --host 127.0.0.1 \
35
  --port 8400 \
@@ -41,14 +66,15 @@ VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 \
41
  > vllm_fim14b.log 2>&1 &
42
  ```
43
 
44
-
45
  Wait until the server is up (model load takes ~1 minute):
46
 
47
  ```bash
48
  curl -s http://127.0.0.1:8400/v1/models
49
  ```
50
 
51
- ## 2. Run the agent on SWE-Bench Verified
 
 
52
 
53
  ```bash
54
  export OPENAI_API_KEY=EMPTY
@@ -74,3 +100,19 @@ uv run python src/r2egym/agenthub/run/edit.py runagent_multiple \
74
  --use_existing True
75
  ```
76
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4
  pipeline_tag: text-generation
5
  base_model:
6
  - Qwen/Qwen2.5-Coder-14B-Instruct
7
+ datasets:
8
+ - TIGER-Lab/FIM-Midtraining-400K
9
+ - R2E-Gym/R2EGym-SFT-Trajectories
10
  tags:
11
  - code
12
  - software-engineering
13
  - agent
14
  ---
15
 
16
+ # FIM-14B
17
 
18
+ [📄 Paper (PDF)](https://github.com/TIGER-AI-Lab/FIM-Midtraining/blob/main/paper.pdf) · [💻 GitHub](https://github.com/TIGER-AI-Lab/FIM-Midtraining) · [🤗 Dataset](https://huggingface.co/datasets/TIGER-Lab/FIM-Midtraining-400K) · [🤗 Collection](https://huggingface.co/collections/TIGER-Lab/fim-midtraining)
19
 
20
+ **FIM-14B** is the 14B coding-agent model of *"Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models"*: `Qwen2.5-Coder-14B-Instruct`, mid-trained on function-aware FIM data, then post-trained on R2E-Gym agent trajectories with the upstream recipe unmodified. The mid-training stage is the only difference from a standard R2E-Gym reproduction — worth **+3.0 points on SWE-Bench-Verified and +4.0 on SWE-Bench-Lite**, while also recovering most of the general-capability erosion that agentic post-training inflicts (LiveCodeBench +11.1, τ-bench +3.9, BFCL +2.4 over the post-training-only arm).
21
 
22
+ ## Training pipeline
23
 
24
+ - **Base model**: [`Qwen/Qwen2.5-Coder-14B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-Coder-14B-Instruct)
25
+ - **FIM mid-training**: [`midtraining/configs/fim_midtrain.yaml`](https://github.com/TIGER-AI-Lab/FIM-Midtraining/blob/main/midtraining/configs/fim_midtrain.yaml) on [TIGER-Lab/FIM-Midtraining-400K](https://huggingface.co/datasets/TIGER-Lab/FIM-Midtraining-400K) (as-run copy: [`FIM_Midtrain_14B.yaml`](https://github.com/TIGER-AI-Lab/FIM-Midtraining/blob/main/midtraining/configs/FIM_Midtrain_14B.yaml)) → intermediate checkpoint released as [TIGER-Lab/FIM-Mid-14B](https://huggingface.co/TIGER-Lab/FIM-Mid-14B)
26
+ - **Post-training**: SFT on R2E-Gym agent trajectories — [`posttraining/r2egym/`](https://github.com/TIGER-AI-Lab/FIM-Midtraining/tree/main/posttraining/r2egym) (as-run copy: [`FIM_Posttrain_14B.yaml`](https://github.com/TIGER-AI-Lab/FIM-Midtraining/blob/main/posttraining/r2egym/FIM_Posttrain_14B.yaml))
27
 
28
+ ## Results
29
 
30
+ Means over three evaluation seeds, identical harness for both arms (paper Tables 1–2):
31
+
32
+ | Setting | SWE-Bench-Verified | SWE-Bench-Lite |
33
+ |---|---|---|
34
+ | Qwen2.5-Coder-14B-Instruct + R2E-Gym (reproduced) | 26.20 | 18.00 |
35
+ | **FIM-14B (+ FIM mid-training)** | **29.20** | **22.00** |
36
+ | Δ | +3.00 | +4.00 |
37
+
38
+ Capability preservation at 14B (six benchmarks outside SWE-Bench):
39
+
40
+ | Setting | LiveCode | OJBench | FSB-EN | Terminal | τ-bench | BFCL | Avg |
41
+ |---|---|---|---|---|---|---|---|
42
+ | + R2E-Gym only | 24.10 | 2.80 | 47.72 | 2.41 | 3.40 | 15.80 | 16.04 |
43
+ | **FIM-14B** | **35.20** | **4.74** | **48.25** | **3.66** | **7.30** | **18.20** | **19.56** |
44
+
45
+ Reproduction guides for all six: [`evaluation/`](https://github.com/TIGER-AI-Lab/FIM-Midtraining/tree/main/evaluation).
46
+
47
+ ## Evaluate on SWE-Bench Verified
48
+
49
+ FIM-14B is evaluated with the **R2E-Gym agent scaffold** (fixed by its post-training pipeline). The complete pinned walkthrough lives at [`evaluation/swebench/released_checkpoints.md`](https://github.com/TIGER-AI-Lab/FIM-Midtraining/blob/main/evaluation/swebench/released_checkpoints.md).
50
+
51
+ ### 1. Serve the model with vLLM
52
 
53
  ```bash
54
  CUDA_VISIBLE_DEVICES=0 \
55
  VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 \
56
+ python -m vllm.entrypoints.openai.api_server \
57
+ --model TIGER-Lab/FIM-14B \
58
  --served-model-name FIM-14B \
59
  --host 127.0.0.1 \
60
  --port 8400 \
 
66
  > vllm_fim14b.log 2>&1 &
67
  ```
68
 
 
69
  Wait until the server is up (model load takes ~1 minute):
70
 
71
  ```bash
72
  curl -s http://127.0.0.1:8400/v1/models
73
  ```
74
 
75
+ ### 2. Run the agent on SWE-Bench Verified
76
+
77
+ From an upstream, unmodified [R2E-Gym](https://github.com/R2E-Gym/R2E-Gym) checkout (Docker required):
78
 
79
  ```bash
80
  export OPENAI_API_KEY=EMPTY
 
100
  --use_existing True
101
  ```
102
 
103
+ For SWE-Bench Lite, use `--dataset "R2E-Gym/SWE-Bench-Lite" --k 300`.
104
+
105
+ ### 3. Score with the official SWE-bench harness
106
+
107
+ Convert the trajectories to a submission and score with the official harness — [`evaluation/swebench/score.sh`](https://github.com/TIGER-AI-Lab/FIM-Midtraining/blob/main/evaluation/swebench/score.sh). The reported number is `resolved_instances / total_instances`.
108
+
109
+ ## Citation
110
+
111
+ ```bibtex
112
+ @article{wang2026fim,
113
+ title={Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models},
114
+ author={Wang, Yubo and Liang, Jiarong and Zhang, Yuxuan and Liu, Xuye and Wei, Cong and Zhang, Yuyu and Nie, Ping and Chen, Wenhu},
115
+ journal={arXiv preprint},
116
+ year={2026}
117
+ }
118
+ ```