🟒 Quill Gen 1

Lightweight Python code generation model built on Qwen2.5-Coder-0.5B.

Quill Gen 1 is a 0.5B parameter experimental model for generating Python functions from natural-language descriptions. It is the younger sibling of Pepper 1 Preview β€” same task, one third the size. Trained as an experiment in small-model code generation, it reaches 80% on HumanEval@50 while running comfortably on consumer hardware.

⚠️ This is NOT a FIM model. Quill Gen 1 cannot do fill-in-the-middle completion. For autocomplete use Quill 1 Preview.

Part of the Minsore family Β· minsore.com

✨ Highlights

  • 🧠 Instruction-tuned β€” writes Python functions from plain descriptions
  • πŸ“¦ Tiny β€” 0.5B params, ~400 MB in Q4_K_M
  • 🎯 Purpose-built β€” code generation only, not chat, not FIM
  • ⚑ Fast β€” designed for low-VRAM setups
  • πŸ§ͺ Experimental β€” a proof of concept for small-model generation
  • πŸ†“ Apache 2.0 β€” same license as base model

πŸ“Š Benchmarks

Quill Gen comparison

Evaluated against 0.5B–1.5B code models. All runs used temperature=0.0, max_tokens=512, --chat-template none.

Benchmark Quill Gen 1 Pepper 1 Preview Quill 1 Preview Qwen2.5-Coder 0.5B
HumanEval@50 80.0% 82.0% 54.0% 52.0%
HumanEval-Infilling (EditSim) 0.161 β€” 0.423 0.045
HumanEval-Infilling (Exact Match) 0% β€” 16% 0%
Delulu FIM (EditSim) 0.048 β€” 0.318 0.035

πŸ“Œ Key insight: Quill Gen 1 reaches 80% HumanEval@50 β€” nearly matching Pepper 1 Preview (82%) at one third the parameter count. On FIM-specific tasks it underperforms significantly, which is expected: this model was trained for generation, not autocomplete.

πŸš€ Quick Start

llama.cpp

llama-server -m quill-gen-1.Q4_K_M.gguf \
  --port 8080 \
  -ngl 99 \
  -c 768 \
  --chat-template none

Requirements: any GPU with β‰₯2 GB VRAM (full offload), or partial CPU offload as fallback.

Request

curl http://localhost:8080/completion \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "### Instruction:\nWrite a Python function that checks if a number is prime.\n\n### Response:\n",
    "n_predict": 256,
    "temperature": 0.0
  }'

Python

import requests

def generate(instruction, max_tokens=256):
    prompt = f"### Instruction:\n{instruction}\n\n### Response:\n"
    r = requests.post("http://localhost:8080/completion", json={
        "prompt": prompt,
        "n_predict": max_tokens,
        "temperature": 0.0,
        "stop": ["### Instruction:", "<|endoftext|>"],
    })
    return r.json()["content"]

print(generate("Write a Python function that reverses a string."))
# β†’ def reverse_string(s):
#       return s[::-1]

βš™οΈ Recommended Settings

Parameter Value Notes
--chat-template none Required. Quill Gen 1 is not a chat model
n_predict 128–256 Works best on short completions
temperature 0.0 Deterministic; use 0.2 for variation
repeat_penalty 1.1 Prevents repetition
-c 768 Matches training context length

⚠️ Limitations

  • Generation only β€” does not support fill-in-the-middle, tool calling, or chat
  • Python-only β€” trained exclusively on Python code
  • Small context β€” 768 tokens; long files are truncated
  • Weak on FIM β€” EditSim 0.16, Exact Match 0% (use Quill 1 instead)
  • Occasional over-explanation β€” may include comments when only code is requested
  • Number looping β€” may repeat large integers on some prompts (use repeat_penalty=1.1)

🧬 Training Details

Base model Qwen2.5-Coder-0.5B-Base
Method QLoRA (r=16, Ξ±=16)
Data FIM-converted Python instruction data
Epochs 1
Context 768 tokens
Optimizer paged_adamw_8bit

πŸ“ Files

File Size Description
quill-gen-1.Q4_K_M.gguf ~400 MB Ready to use with llama.cpp (recommended)
model.safetensors ~1 GB Full-precision merged weights
config.json β€” Model config
tokenizer.json β€” Tokenizer
tokenizer_config.json β€” Tokenizer config
generation_config.json β€” Generation defaults (optional)

πŸ—ΊοΈ Roadmap

  • Quill 2 β€” FIM autocomplete, fixed suffix handling, JS/TS/Rust support
  • Pepper 2 β€” improved MBPP and LiveCodeBench
  • Symphony β€” flagship agentic code model (3B MoE)

πŸ“œ License

Apache 2.0 β€” same as the base Qwen2.5-Coder-0.5B model.

πŸ™ Credits

πŸ“¬ Contact

Minsore β€” Ukrainian AI lab building open language models.

⭐ If Quill Gen 1 is useful, star the repo and share your results.

Downloads last month
-
Safetensors
Model size
0.5B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Evaluation results