| --- |
| license: apache-2.0 |
| language: |
| - en |
| base_model: |
| - Qwen/Qwen3.5-4B |
| pipeline_tag: image-text-to-text |
| tags: |
| - agent |
| - code |
| metrics: |
| - code_eval |
| - accuracy |
| --- |
| # ๐ง Qwen3.5-4B-AgentCoder |
|
|
| ### A Fine-Tuned Model for Enhanced Tool Calling, Code Generation, and Reasoning |
|
|
| - **Author:** Bruno Pistone |
| - **Repository:** [br1-pist/Qwen3.5-4B-AgentCoder](https://huggingface.co/br1-pist/Qwen3.5-4B-AgentCoder) |
| - **License:** Apache 2.0 |
| - **Base model:** [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) |
|
|
| --- |
|
|
| ## Model Description |
|
|
| **Qwen3.5-4B-AgentCoder** is a fine-tuned version of the Qwen/Qwen3.5-4B model, optimized for: |
| - ๐งฎ Complex reasoning tasks |
| - ๐งฐ Tool calling |
| - ๐ป Code generation |
|
|
| The model was developed through **sequential fine-tuning**, followed by a **Direct Preference Optimization (DPO)** post-training stage to improve alignment, coherence, and reasoning accuracy. |
|
|
| ### Highlights |
| - Post-trained with **DPO** using *chosen/rejected* pairs for better alignment |
| - Excellent balance between **tool use**, **code generation**, and **reasoning** |
|
|
| --- |
|
|
| ## ๐ Direct Use |
|
|
| Qwen3.5-4B-AgentCoder can be used directly for: |
| - โ
Tool calling in complex reasoning tasks |
| - โ
Code generation for Python, JS, and other languages |
| - โ
Multi-domain reasoning (math, logic, Q&A) |
|
|
| ### โ ๏ธ Out-of-Scope Use |
| - โ Highly sensitive or confidential data |
| - โ Domains requiring expert-level specialization |
| - โ Tasks where full explainability is mandatory |
|
|
| --- |
|
|
| ## ๐ป Getting Started |
|
|
| ```python |
| from transformers import AutoModelForCausalLM, AutoTokenizer |
| import torch |
| model_name = "br1-pist/Qwen3.5-4B-AgentCoder" |
| tokenizer = AutoTokenizer.from_pretrained(model_name) |
| model = AutoModelForCausalLM.from_pretrained( |
| model_name, |
| torch_dtype="auto", |
| device_map="auto" |
| ) |
| prompt = "Give me a short introduction to large language models." |
| messages = [{"role": "user", "content": prompt}] |
| text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) |
| model_inputs = tokenizer([text], return_tensors="pt").to(model.device) |
| generated_ids = model.generate(**model_inputs, max_new_tokens=1024) |
| output = tokenizer.decode(generated_ids[0], skip_special_tokens=True) |
| print(output) |
| ``` |
|
|
| --- |
|
|
| ## ๐ง Training Details |
|
|
| ### Training Procedure |
|
|
| #### Phase 1 โ Post-Training - Direct Preference Optimization (DPO) |
| After sequential fine-tuning, the model underwent a **DPO** phase to enhance response alignment, reasoning robustness, and factual consistency. |
|
|
| - Learning rate: `3e-6` |
| - Batch size: `1` |
| - Gradient accumulation: `4` |
| - Epochs: `1` |
| - Beta: 0.1 |
| - Loss type: `sigmoid` |
| - Warmup steps: `27` |
| - Sequence length: ~2.5K tokens |
|
|
| ##### DPO Data |
| - ~2.5K chosen/rejected response pairs |
| - *Rejected* samples synthetically generated to represent poor or incoherent answers |
| - *Chosen* samples tagged from real conversations |
|
|
| **Objective** |
| - Encourage the model to prefer *chosen* completions |
| - Improve clarity, correctness, and helpfulness |
| - Reduce hallucinations and verbosity |
|
|
| --- |
|
|
| ## ๐ฅ๏ธ Technical Specifications |
|
|
| ### Model Architecture |
| - Model type: Causal language model |
| - Parameters: 4.0B |
| - Context length: ~264K tokens |
| - Thinking mode: Enabled |
|
|
| ### Compute Infrastructure |
|
|
| **Hardware** |
| - GPU: NVIDIA H100 (80 GB VRAM) |
| - System RAM: 2 TiB |
| - Memory per vCPU: 10.67 GiB |
|
|
| **Software** |
| - Python: 3.12 |
| - Transformers: 5.3.0 |
| - Libraries: `bitsandbytes`, `safetensors`, `torch`, `trl`, `scikit-learn`, `tokenizers`, `psutil`, `py7zr` |
|
|
| --- |
|
|
| ## ๐งพ Citation |
|
|
| **BibTeX** |
| ```bibtex |
| @article{qwen3.5-4b-thinking-2507-toolcode, |
| title={Qwen3.5-4B-AgentCoder: A Fine-Tuned Model for Enhanced Tool Calling, Code Generation, and Reasoning}, |
| author={Bruno Pistone}, |
| year={2025}, |
| journal={Hugging Face Model Hub} |
| } |
| ``` |
|
|
| **APA** |
| > Bruno Pistone. (2026). *Qwen3.5-4B-AgentCoder: A Fine-Tuned Model for Enhanced Tool Calling, Code Generation, and Reasoning.* Hugging Face Model Hub. https://huggingface.co/br1-pist/Qwen3.5-4B-AgentCoder |
|
|
| --- |
|
|
| ## ๐งญ Recommendations |
|
|
| - Tool use accuracy depends on task complexity |
| - Code generation may occasionally produce minor syntax issues |
| - Reasoning strongest in structured, logical, and mathematical contexts |
| - Avoid using this model for confidential or safety-critical applications |
|
|
| --- |
|
|
| <p align="center"> |
| <b>๐ง Qwen3.5-4B-AgentCoder โ created by Bruno Pistone</b><br> |
| <i>Enhanced reasoning, tool calling, and code generation โ refined with DPO alignment</i> |
| </p> |