Text Generation
Transformers
GGUF
Safetensors
English
qwen2
code
cobol
mainframe
legacy-code
code-translation
unsloth
conversational
Instructions to use FLs-AI/FL-7B-3.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FLs-AI/FL-7B-3.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="FLs-AI/FL-7B-3.1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("FLs-AI/FL-7B-3.1") model = AutoModelForCausalLM.from_pretrained("FLs-AI/FL-7B-3.1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use FLs-AI/FL-7B-3.1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf FLs-AI/FL-7B-3.1:Q4_K_M # Run inference directly in the terminal: llama cli -hf FLs-AI/FL-7B-3.1:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf FLs-AI/FL-7B-3.1:Q4_K_M # Run inference directly in the terminal: llama cli -hf FLs-AI/FL-7B-3.1:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf FLs-AI/FL-7B-3.1:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf FLs-AI/FL-7B-3.1:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf FLs-AI/FL-7B-3.1:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf FLs-AI/FL-7B-3.1:Q4_K_M
Use Docker
docker model run hf.co/FLs-AI/FL-7B-3.1:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use FLs-AI/FL-7B-3.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FLs-AI/FL-7B-3.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FLs-AI/FL-7B-3.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/FLs-AI/FL-7B-3.1:Q4_K_M
- SGLang
How to use FLs-AI/FL-7B-3.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "FLs-AI/FL-7B-3.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FLs-AI/FL-7B-3.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "FLs-AI/FL-7B-3.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FLs-AI/FL-7B-3.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use FLs-AI/FL-7B-3.1 with Ollama:
ollama run hf.co/FLs-AI/FL-7B-3.1:Q4_K_M
- Unsloth Studio
How to use FLs-AI/FL-7B-3.1 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for FLs-AI/FL-7B-3.1 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for FLs-AI/FL-7B-3.1 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for FLs-AI/FL-7B-3.1 to start chatting
- Pi
How to use FLs-AI/FL-7B-3.1 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FLs-AI/FL-7B-3.1:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "FLs-AI/FL-7B-3.1:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use FLs-AI/FL-7B-3.1 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FLs-AI/FL-7B-3.1:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "FLs-AI/FL-7B-3.1:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use FLs-AI/FL-7B-3.1 with Docker Model Runner:
docker model run hf.co/FLs-AI/FL-7B-3.1:Q4_K_M
- Lemonade
How to use FLs-AI/FL-7B-3.1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull FLs-AI/FL-7B-3.1:Q4_K_M
Run and chat with the model
lemonade run user.FL-7B-3.1-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use FLs-AI/FL-7B-3.1 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FLs-AI/FL-7B-3.1:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default FLs-AI/FL-7B-3.1:Q4_K_M
Run Hermes
hermes
- Atomic Chat
| license: cc-by-nc-4.0 | |
| base_model: Qwen/Qwen2.5-Coder-7B | |
| base_model_relation: finetune | |
| library_name: transformers | |
| pipeline_tag: text-generation | |
| language: | |
| - en | |
| tags: | |
| - qwen2 | |
| - safetensors | |
| - code | |
| - cobol | |
| - mainframe | |
| - legacy-code | |
| - code-translation | |
| - unsloth | |
| # FL-7B-3.1 | |
| **A Qwen2.5-Coder-7B model adapted for COBOL, mainframe knowledge, and legacy-code modernization.** | |
| FL-7B-3.1 starts from | |
| [Qwen/Qwen2.5-Coder-7B](https://huggingface.co/Qwen/Qwen2.5-Coder-7B) and was trained in two | |
| stages: continued pretraining (CPT) on COBOL source material, followed by assistant-only | |
| supervised fine-tuning (SFT) on COBOL and mainframe-oriented instructions. | |
| The model is intended for: | |
| - generating and completing GnuCOBOL programs; | |
| - translating COBOL into Java; | |
| - answering mainframe and legacy-system questions; | |
| - explaining and summarizing COBOL source code. | |
| ## Highlights | |
| | Benchmark | Result | | |
| |---|---:| | |
| | COBOLEval pass@1 | **17.81%** (26/146) | | |
| | COBOLEval test compilation rate | **57.73%** (474/821) | | |
| | COBOLEval test pass rate | **30.82%** (253/821) | | |
| | COBOL-to-Java CSR | **61.54%** (88/143) | | |
| | COBOL-to-Java pass@1 | **48.25%** (69/143) | | |
| | MainframeBench MCQ accuracy | **80.84%** (1,561/1,931) | | |
| ## Usage with Transformers | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model_id = "FLs-AI/FL-7B-3.1" | |
| tokenizer = AutoTokenizer.from_pretrained( | |
| model_id, | |
| fix_mistral_regex=True, | |
| ) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_id, | |
| dtype=torch.bfloat16, | |
| device_map="auto", | |
| ) | |
| messages = [ | |
| { | |
| "role": "user", | |
| "content": ( | |
| "Write a complete GnuCOBOL 3.2 program that reads signed integers " | |
| "until EOF and prints their sum. Return only COBOL source code." | |
| ), | |
| } | |
| ] | |
| inputs = tokenizer.apply_chat_template( | |
| messages, | |
| add_generation_prompt=True, | |
| tokenize=True, | |
| return_dict=True, | |
| return_tensors="pt", | |
| ).to(model.device) | |
| outputs = model.generate( | |
| **inputs, | |
| max_new_tokens=1536, | |
| do_sample=False, | |
| repetition_penalty=1.05, | |
| ) | |
| generated = outputs[0, inputs["input_ids"].shape[1]:] | |
| print(tokenizer.decode(generated, skip_special_tokens=True)) | |
| ``` | |
| ### Recommended generation settings | |
| | Use case | Temperature | Max new tokens | Repetition penalty | | |
| |---|---:|---:|---:| | |
| | COBOL generation | `0` | `1536` | `1.05` | | |
| | COBOL to Java | `0` | `4096` | `1.0` | | |
| | Mainframe MCQ | `0` | `16` | `1.0` | | |
| | Mainframe QA / summarization | `0` | `512` | `1.0` | | |
| For the final COBOLEval run, `repetition_penalty=1.05` produced the strongest measured result. | |
| COBOL relies heavily on repeated identifiers and fixed structural phrases, so large repetition | |
| penalties can damage syntax and correctness. | |
| ## Evaluation | |
| All results below were measured on the merged BF16 checkpoint with greedy decoding. | |
| ### COBOLEval | |
| [COBOLEval](https://github.com/zorse-project/COBOLEval) evaluates generated programs by compiling | |
| and executing them with GnuCOBOL. The evaluation used 146 problems and 821 test cases, GnuCOBOL | |
| 3.2.0, `max_new_tokens=1536`, and one sample per task. | |
| | Model / setting | pass@1 | Test compilation rate | Tests passed | | |
| |---|---:|---:|---:| | |
| | Qwen2.5-Coder-7B base | 0.00% | 3.65% | 4/821 | | |
| | FL-7B-3.1, repetition penalty 1.00 | 17.12% | 41.29% | 175/821 | | |
| | **FL-7B-3.1, repetition penalty 1.05** | **17.81%** | **57.73%** | **253/821** | | |
| The harness was pinned to commit `0bb96c3114bb2bb28e221e9d6000614781f8609d`. | |
| ### COBOL to Java | |
| The [COBOL-JavaTrans C2J](https://github.com/COBOL-Coder/COBOL-Coder) evaluation compiles and | |
| executes generated Java translations. | |
| | Metric | Result | | |
| |---|---:| | |
| | Tasks | 143 | | |
| | Compilation success rate (CSR) | **61.54%** (88/143) | | |
| | pass@1 | **48.25%** (69/143) | | |
| The evaluator was pinned to commit `2b14b7bf7e55556205654c6f7657fa60e36251fa`. | |
| ### MainframeBench | |
| [Fsoft-AIC/MainframeBench](https://huggingface.co/datasets/Fsoft-AIC/MainframeBench) contains | |
| multiple-choice questions, open-ended QA, and COBOL code summarization. | |
| #### Multiple choice | |
| | Tasks | Correct | Accuracy | Invalid predictions | | |
| |---:|---:|---:|---:| | |
| | 1,931 | 1,561 | **80.84%** | 1 | | |
| #### Open-ended tasks | |
| | Suite | Tasks | Token F1 | ROUGE-L F1 | BLEU-4 | | |
| |---|---:|---:|---:|---:| | |
| | Question answering | 2,598 | 28.46% | 24.31% | 3.76 | | |
| | COBOL summarization | 2,523 | 41.76% | 36.94% | 14.04 | | |
| Normalized exact match was 0% for both open-ended suites. This strict lexical metric requires | |
| the generated response to match the single reference wording after normalization; it is not an | |
| accuracy or semantic-correctness score. Token F1, ROUGE-L, and BLEU-4 measure lexical overlap and | |
| should not be interpreted as execution-based correctness or human preference. | |
| The dataset was pinned to revision `70d30c76eb29e45dd8965304b41c56bc1f527972`. | |
| ## Limitations | |
| - The model can enter repetition loops or produce excessively long code on difficult tasks. | |
| - A compiling program is not necessarily functionally correct or safe. | |
| - Evaluation used GnuCOBOL 3.2.0. | |
| - General-purpose coding performance inherited from Qwen2.5-Coder was not re-evaluated and may | |
| have regressed during domain adaptation. | |
| - MainframeBench QA and summarization results are lexical-overlap scores, not semantic accuracy. | |
| Do not deploy generated code to production systems without compilation, tests, static analysis, | |
| and review by an experienced mainframe engineer. | |
| ## License | |
| This fine-tune is released under **CC BY-NC 4.0**. Attribution is required and commercial use of | |
| the fine-tuned weights is not permitted under this license. The Qwen2.5-Coder-7B base model is | |
| licensed separately under Apache 2.0. | |
| ## Citation | |
| ```bibtex | |
| @misc{fl7b31, | |
| title = {FL-7B-3.1: COBOL and Mainframe Code Model}, | |
| author = {FLs-AI}, | |
| year = {2026}, | |
| url = {https://huggingface.co/FLs-AI/FL-7B-3.1} | |
| } | |
| ``` | |
| ## Acknowledgements | |
| - [Qwen2.5-Coder](https://huggingface.co/Qwen/Qwen2.5-Coder-7B) | |
| - [Unsloth](https://github.com/unslothai/unsloth) | |
| - [GnuCOBOL](https://gnucobol.sourceforge.io/) | |
| - [COBOLEval](https://github.com/zorse-project/COBOLEval) | |
| - [COBOL-Coder / COBOL-JavaTrans](https://github.com/COBOL-Coder/COBOL-Coder) | |
| - [MainframeBench](https://huggingface.co/datasets/Fsoft-AIC/MainframeBench) | |