--- base_model: Qwen/Qwen2.5-0.5B-Instruct library_name: transformers pipeline_tag: text-generation tags: - code - coding - qwen - qwen2 - slm - trl - fine-tuned license: apache-2.0 language: - en - code --- # light-coder `light-coder` is an ultra-lightweight, standalone instruction-tuned coding model created by fine-tuning [Qwen/Qwen2.5-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct) on ~122k programming instruction-response pairs and merging the LoRA weights directly into the base checkpoint. At under 1 GB in size, it requires minimal VRAM, executes quickly on consumer GPUs and CPUs, and integrates out-of-the-box with tools like vLLM, Ollama, and standard Hugging Face pipelines. ## Model Details - **Developed by:** Milad Asghari - **Model Name:** light-coder - **Base Model:** [Qwen/Qwen2.5-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct) - **Model Type:** Causal Language Model (Full Merged Weights) - **Primary Domain:** Code generation, refactoring, and programming instruction-following - **Language(s):** English, Multiple Programming Languages - **License:** Apache-2.0 - **Size:** 988 MB (`safetensors`) ## How to Get Started Because the weights are merged, you do not need the `peft` library for inference: ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "Miladasghari/light-coder" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, dtype=torch.float16, device_map="auto" ) messages = [ {"role": "user", "content": "Write a Python function to check if a string is a palindrome."} ] prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) inputs = tokenizer(prompt, return_tensors="pt").to(model.device) outputs = model.generate( **inputs, max_new_tokens=256, temperature=0.3, top_p=0.9, repetition_penalty=1.05 ) response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True) print(response)