--- license: cc-by-nc-4.0 base_model: Qwen/Qwen2.5-Coder-7B base_model_relation: finetune library_name: transformers pipeline_tag: text-generation language: - en tags: - qwen2 - safetensors - code - cobol - mainframe - legacy-code - code-translation - unsloth --- # FL-7B-3.1 **A Qwen2.5-Coder-7B model adapted for COBOL, mainframe knowledge, and legacy-code modernization.** FL-7B-3.1 starts from [Qwen/Qwen2.5-Coder-7B](https://huggingface.co/Qwen/Qwen2.5-Coder-7B) and was trained in two stages: continued pretraining (CPT) on COBOL source material, followed by assistant-only supervised fine-tuning (SFT) on COBOL and mainframe-oriented instructions. The model is intended for: - generating and completing GnuCOBOL programs; - translating COBOL into Java; - answering mainframe and legacy-system questions; - explaining and summarizing COBOL source code. ## Highlights | Benchmark | Result | |---|---:| | COBOLEval pass@1 | **17.81%** (26/146) | | COBOLEval test compilation rate | **57.73%** (474/821) | | COBOLEval test pass rate | **30.82%** (253/821) | | COBOL-to-Java CSR | **61.54%** (88/143) | | COBOL-to-Java pass@1 | **48.25%** (69/143) | | MainframeBench MCQ accuracy | **80.84%** (1,561/1,931) | ## Usage with Transformers ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "FLs-AI/FL-7B-3.1" tokenizer = AutoTokenizer.from_pretrained( model_id, fix_mistral_regex=True, ) model = AutoModelForCausalLM.from_pretrained( model_id, dtype=torch.bfloat16, device_map="auto", ) messages = [ { "role": "user", "content": ( "Write a complete GnuCOBOL 3.2 program that reads signed integers " "until EOF and prints their sum. Return only COBOL source code." ), } ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate( **inputs, max_new_tokens=1536, do_sample=False, repetition_penalty=1.05, ) generated = outputs[0, inputs["input_ids"].shape[1]:] print(tokenizer.decode(generated, skip_special_tokens=True)) ``` ### Recommended generation settings | Use case | Temperature | Max new tokens | Repetition penalty | |---|---:|---:|---:| | COBOL generation | `0` | `1536` | `1.05` | | COBOL to Java | `0` | `4096` | `1.0` | | Mainframe MCQ | `0` | `16` | `1.0` | | Mainframe QA / summarization | `0` | `512` | `1.0` | For the final COBOLEval run, `repetition_penalty=1.05` produced the strongest measured result. COBOL relies heavily on repeated identifiers and fixed structural phrases, so large repetition penalties can damage syntax and correctness. ## Evaluation All results below were measured on the merged BF16 checkpoint with greedy decoding. ### COBOLEval [COBOLEval](https://github.com/zorse-project/COBOLEval) evaluates generated programs by compiling and executing them with GnuCOBOL. The evaluation used 146 problems and 821 test cases, GnuCOBOL 3.2.0, `max_new_tokens=1536`, and one sample per task. | Model / setting | pass@1 | Test compilation rate | Tests passed | |---|---:|---:|---:| | Qwen2.5-Coder-7B base | 0.00% | 3.65% | 4/821 | | FL-7B-3.1, repetition penalty 1.00 | 17.12% | 41.29% | 175/821 | | **FL-7B-3.1, repetition penalty 1.05** | **17.81%** | **57.73%** | **253/821** | The harness was pinned to commit `0bb96c3114bb2bb28e221e9d6000614781f8609d`. ### COBOL to Java The [COBOL-JavaTrans C2J](https://github.com/COBOL-Coder/COBOL-Coder) evaluation compiles and executes generated Java translations. | Metric | Result | |---|---:| | Tasks | 143 | | Compilation success rate (CSR) | **61.54%** (88/143) | | pass@1 | **48.25%** (69/143) | The evaluator was pinned to commit `2b14b7bf7e55556205654c6f7657fa60e36251fa`. ### MainframeBench [Fsoft-AIC/MainframeBench](https://huggingface.co/datasets/Fsoft-AIC/MainframeBench) contains multiple-choice questions, open-ended QA, and COBOL code summarization. #### Multiple choice | Tasks | Correct | Accuracy | Invalid predictions | |---:|---:|---:|---:| | 1,931 | 1,561 | **80.84%** | 1 | #### Open-ended tasks | Suite | Tasks | Token F1 | ROUGE-L F1 | BLEU-4 | |---|---:|---:|---:|---:| | Question answering | 2,598 | 28.46% | 24.31% | 3.76 | | COBOL summarization | 2,523 | 41.76% | 36.94% | 14.04 | Normalized exact match was 0% for both open-ended suites. This strict lexical metric requires the generated response to match the single reference wording after normalization; it is not an accuracy or semantic-correctness score. Token F1, ROUGE-L, and BLEU-4 measure lexical overlap and should not be interpreted as execution-based correctness or human preference. The dataset was pinned to revision `70d30c76eb29e45dd8965304b41c56bc1f527972`. ## Limitations - The model can enter repetition loops or produce excessively long code on difficult tasks. - A compiling program is not necessarily functionally correct or safe. - Evaluation used GnuCOBOL 3.2.0. - General-purpose coding performance inherited from Qwen2.5-Coder was not re-evaluated and may have regressed during domain adaptation. - MainframeBench QA and summarization results are lexical-overlap scores, not semantic accuracy. Do not deploy generated code to production systems without compilation, tests, static analysis, and review by an experienced mainframe engineer. ## License This fine-tune is released under **CC BY-NC 4.0**. Attribution is required and commercial use of the fine-tuned weights is not permitted under this license. The Qwen2.5-Coder-7B base model is licensed separately under Apache 2.0. ## Citation ```bibtex @misc{fl7b31, title = {FL-7B-3.1: COBOL and Mainframe Code Model}, author = {FLs-AI}, year = {2026}, url = {https://huggingface.co/FLs-AI/FL-7B-3.1} } ``` ## Acknowledgements - [Qwen2.5-Coder](https://huggingface.co/Qwen/Qwen2.5-Coder-7B) - [Unsloth](https://github.com/unslothai/unsloth) - [GnuCOBOL](https://gnucobol.sourceforge.io/) - [COBOLEval](https://github.com/zorse-project/COBOLEval) - [COBOL-Coder / COBOL-JavaTrans](https://github.com/COBOL-Coder/COBOL-Coder) - [MainframeBench](https://huggingface.co/datasets/Fsoft-AIC/MainframeBench)