--- license: apache-2.0 base_model: unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit tags: - text-generation-inference - transformers - unsloth - qwen2 - lora - fastapi - code-assistant language: - en library_name: transformers pipeline_tag: text-generation --- # Qwen2.5-Coder-7B-FastAPI-LoRA A LoRA fine-tune of **Qwen2.5-Coder-7B-Instruct** specialized as a **FastAPI documentation assistant**. The model is trained to answer questions, generate code, and explain concepts related to the FastAPI framework, covering everything from basic routing to advanced topics like security and testing. ## Model Details - **Developed by:** [LadiesMan69](https://huggingface.co/LadiesMan69) - **License:** apache-2.0 - **Base model:** [unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit](https://huggingface.co/unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit) - **Model type:** Causal decoder-only language model (Qwen2 architecture) - **Fine-tuning method:** LoRA (Low-Rank Adaptation) - **Language:** English - **Trained with:** [Unsloth](https://github.com/unslothai/unsloth) + Hugging Face TRL — 2x faster training ## Motivation General-purpose code models are often imprecise or outdated when it comes to framework-specific APIs. This model was fine-tuned on a curated dataset of FastAPI-focused instruction/response pairs to produce a lightweight, deployable assistant that gives accurate, idiomatic answers for building and debugging FastAPI applications. ## Training Data The fine-tuning dataset was built specifically for this task using the **ChatML** format and organized into topic categories, including: - **Tutorial** — core concepts: path/query parameters, request bodies, response models, dependency injection - **Advanced** — background tasks, middleware, WebSockets, custom exception handlers, lifespan events - **Security** — OAuth2/JWT authentication, password hashing, CORS, rate limiting - **Testing** — `TestClient` usage, pytest fixtures, mocking dependencies, async test patterns Examples were generated in batches per category to ensure balanced topic coverage and consistent formatting across the dataset. ## Intended Use - Answering questions about FastAPI concepts, patterns, and best practices - Generating FastAPI route handlers, Pydantic models, and dependency-injected services - Explaining and debugging FastAPI-related code snippets - Acting as an in-editor or chat-based documentation assistant for developers working with FastAPI ## How to Use ### With `transformers` ```python from transformers import AutoTokenizer, AutoModelForCausalLM model_id = "LadiesMan69/Qwen2.5-Coder-7B-FastAPI-LoRA" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto") messages = [ {"role": "system", "content": "You are a helpful FastAPI documentation assistant."}, {"role": "user", "content": "How do I add JWT-based authentication to a FastAPI route?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=512) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True)) ``` ### With `unsloth` ```python from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="LadiesMan69/Qwen2.5-Coder-7B-FastAPI-LoRA", max_seq_length=2048, ) ``` ### With `vLLM` ```bash pip install vllm vllm serve "LadiesMan69/Qwen2.5-Coder-7B-FastAPI-LoRA" ``` ```bash curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LadiesMan69/Qwen2.5-Coder-7B-FastAPI-LoRA", "messages": [ {"role": "user", "content": "Show me a minimal FastAPI app with a health check endpoint."} ] }' ``` ## Prompt Format This model uses the ChatML-style chat template built into the tokenizer (`apply_chat_template`). For best results, include a system message establishing the assistant's role as a FastAPI expert, followed by the user's question. ## Limitations - Focused specifically on FastAPI; general coding ability outside this domain is inherited from the base model and not specifically enhanced. - As with any LLM, generated code should be reviewed and tested before use in production. - May not reflect the very latest FastAPI releases if they postdate the training data. ## Training Procedure Fine-tuned using LoRA adapters on top of the 4-bit quantized `unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit` base model, leveraging Unsloth's optimized training kernels for faster, memory-efficient fine-tuning. ## Model Tree - Base: [Qwen/Qwen2.5-7B](https://huggingface.co/Qwen/Qwen2.5-7B) - → [Qwen/Qwen2.5-Coder-7B](https://huggingface.co/Qwen/Qwen2.5-Coder-7B) - → [Qwen/Qwen2.5-Coder-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct) - → [unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit](https://huggingface.co/unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit) (quantized) - → **LadiesMan69/Qwen2.5-Coder-7B-FastAPI-LoRA** (this model, LoRA fine-tune) ## Acknowledgements This qwen2 model was trained 2x faster with [Unsloth](https://github.com/unslothai/unsloth) and Hugging Face's TRL library. [![Made with Unsloth](https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png)](https://github.com/unslothai/unsloth)