FL-7B-3.1 / README.md
FLs-AI's picture
Update README.md
a83eb98 verified
|
Raw
History Blame Contribute Delete
6.28 kB
---
license: cc-by-nc-4.0
base_model: Qwen/Qwen2.5-Coder-7B
base_model_relation: finetune
library_name: transformers
pipeline_tag: text-generation
language:
- en
tags:
- qwen2
- safetensors
- code
- cobol
- mainframe
- legacy-code
- code-translation
- unsloth
---
# FL-7B-3.1
**A Qwen2.5-Coder-7B model adapted for COBOL, mainframe knowledge, and legacy-code modernization.**
FL-7B-3.1 starts from
[Qwen/Qwen2.5-Coder-7B](https://huggingface.co/Qwen/Qwen2.5-Coder-7B) and was trained in two
stages: continued pretraining (CPT) on COBOL source material, followed by assistant-only
supervised fine-tuning (SFT) on COBOL and mainframe-oriented instructions.
The model is intended for:
- generating and completing GnuCOBOL programs;
- translating COBOL into Java;
- answering mainframe and legacy-system questions;
- explaining and summarizing COBOL source code.
## Highlights
| Benchmark | Result |
|---|---:|
| COBOLEval pass@1 | **17.81%** (26/146) |
| COBOLEval test compilation rate | **57.73%** (474/821) |
| COBOLEval test pass rate | **30.82%** (253/821) |
| COBOL-to-Java CSR | **61.54%** (88/143) |
| COBOL-to-Java pass@1 | **48.25%** (69/143) |
| MainframeBench MCQ accuracy | **80.84%** (1,561/1,931) |
## Usage with Transformers
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "FLs-AI/FL-7B-3.1"
tokenizer = AutoTokenizer.from_pretrained(
model_id,
fix_mistral_regex=True,
)
model = AutoModelForCausalLM.from_pretrained(
model_id,
dtype=torch.bfloat16,
device_map="auto",
)
messages = [
{
"role": "user",
"content": (
"Write a complete GnuCOBOL 3.2 program that reads signed integers "
"until EOF and prints their sum. Return only COBOL source code."
),
}
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=1536,
do_sample=False,
repetition_penalty=1.05,
)
generated = outputs[0, inputs["input_ids"].shape[1]:]
print(tokenizer.decode(generated, skip_special_tokens=True))
```
### Recommended generation settings
| Use case | Temperature | Max new tokens | Repetition penalty |
|---|---:|---:|---:|
| COBOL generation | `0` | `1536` | `1.05` |
| COBOL to Java | `0` | `4096` | `1.0` |
| Mainframe MCQ | `0` | `16` | `1.0` |
| Mainframe QA / summarization | `0` | `512` | `1.0` |
For the final COBOLEval run, `repetition_penalty=1.05` produced the strongest measured result.
COBOL relies heavily on repeated identifiers and fixed structural phrases, so large repetition
penalties can damage syntax and correctness.
## Evaluation
All results below were measured on the merged BF16 checkpoint with greedy decoding.
### COBOLEval
[COBOLEval](https://github.com/zorse-project/COBOLEval) evaluates generated programs by compiling
and executing them with GnuCOBOL. The evaluation used 146 problems and 821 test cases, GnuCOBOL
3.2.0, `max_new_tokens=1536`, and one sample per task.
| Model / setting | pass@1 | Test compilation rate | Tests passed |
|---|---:|---:|---:|
| Qwen2.5-Coder-7B base | 0.00% | 3.65% | 4/821 |
| FL-7B-3.1, repetition penalty 1.00 | 17.12% | 41.29% | 175/821 |
| **FL-7B-3.1, repetition penalty 1.05** | **17.81%** | **57.73%** | **253/821** |
The harness was pinned to commit `0bb96c3114bb2bb28e221e9d6000614781f8609d`.
### COBOL to Java
The [COBOL-JavaTrans C2J](https://github.com/COBOL-Coder/COBOL-Coder) evaluation compiles and
executes generated Java translations.
| Metric | Result |
|---|---:|
| Tasks | 143 |
| Compilation success rate (CSR) | **61.54%** (88/143) |
| pass@1 | **48.25%** (69/143) |
The evaluator was pinned to commit `2b14b7bf7e55556205654c6f7657fa60e36251fa`.
### MainframeBench
[Fsoft-AIC/MainframeBench](https://huggingface.co/datasets/Fsoft-AIC/MainframeBench) contains
multiple-choice questions, open-ended QA, and COBOL code summarization.
#### Multiple choice
| Tasks | Correct | Accuracy | Invalid predictions |
|---:|---:|---:|---:|
| 1,931 | 1,561 | **80.84%** | 1 |
#### Open-ended tasks
| Suite | Tasks | Token F1 | ROUGE-L F1 | BLEU-4 |
|---|---:|---:|---:|---:|
| Question answering | 2,598 | 28.46% | 24.31% | 3.76 |
| COBOL summarization | 2,523 | 41.76% | 36.94% | 14.04 |
Normalized exact match was 0% for both open-ended suites. This strict lexical metric requires
the generated response to match the single reference wording after normalization; it is not an
accuracy or semantic-correctness score. Token F1, ROUGE-L, and BLEU-4 measure lexical overlap and
should not be interpreted as execution-based correctness or human preference.
The dataset was pinned to revision `70d30c76eb29e45dd8965304b41c56bc1f527972`.
## Limitations
- The model can enter repetition loops or produce excessively long code on difficult tasks.
- A compiling program is not necessarily functionally correct or safe.
- Evaluation used GnuCOBOL 3.2.0.
- General-purpose coding performance inherited from Qwen2.5-Coder was not re-evaluated and may
have regressed during domain adaptation.
- MainframeBench QA and summarization results are lexical-overlap scores, not semantic accuracy.
Do not deploy generated code to production systems without compilation, tests, static analysis,
and review by an experienced mainframe engineer.
## License
This fine-tune is released under **CC BY-NC 4.0**. Attribution is required and commercial use of
the fine-tuned weights is not permitted under this license. The Qwen2.5-Coder-7B base model is
licensed separately under Apache 2.0.
## Citation
```bibtex
@misc{fl7b31,
title = {FL-7B-3.1: COBOL and Mainframe Code Model},
author = {FLs-AI},
year = {2026},
url = {https://huggingface.co/FLs-AI/FL-7B-3.1}
}
```
## Acknowledgements
- [Qwen2.5-Coder](https://huggingface.co/Qwen/Qwen2.5-Coder-7B)
- [Unsloth](https://github.com/unslothai/unsloth)
- [GnuCOBOL](https://gnucobol.sourceforge.io/)
- [COBOLEval](https://github.com/zorse-project/COBOLEval)
- [COBOL-Coder / COBOL-JavaTrans](https://github.com/COBOL-Coder/COBOL-Coder)
- [MainframeBench](https://huggingface.co/datasets/Fsoft-AIC/MainframeBench)