--- license: mit language: - pt - en - es - fr base_model: - ornith-ai/Ornith-1.5-9B tags: - coder - code - programmer - edge - fast - smart ---
![image](https://cdn-uploads.huggingface.co/production/uploads/644afe169279988e0cbcd2d9/jCFCVTpFLOvaCUOF90b9o.png)
# LEVUM [👉 Visit the Velum AI Landing Page](https://guell11.github.io/velum-ai/) ### Local · Code · Software **A local model based on Qwen 3.5 9B, focused on software development.** Write. Understand. Fix. Refactor. Build. **Less ceremony. More working code.**
--- ## About **LEVUM** is a **9-billion-parameter model**, based on **Qwen 3.5 9B** and focused on software development tasks. The goal is simple: work close to the project and turn instructions into useful code. The model was designed for tasks such as: * code generation; * feature implementation; * debugging; * refactoring; * code explanation and understanding; * prototype creation; * project structuring; * complete application generation; * assistance during local development. LEVUM follows a **local-first** philosophy: the model can run close to the IDE, files, and the actual project context, depending on the chosen runtime and quantization. > **Software first. Local by default.** --- ## Specifications | | | | --------------- | ----------------------------------------------------- | | **Model** | LEVUM | | **Base** | Qwen 3.5 9B | | **Parameters** | 9B | | **Focus** | Code / Software | | **Primary use** | Software development | | **Execution** | Local | | **Language** | English + multilingual capabilities of the base model | | **Origin** | 🇧🇷 Brazil | --- # Quantizations LEVUM is available in different quantization levels to support a wider range of hardware. The ideal choice mainly depends on: **RAM / VRAM → speed → fidelity** Comparison of the model's relative quality after different quantization levels: ![image](https://cdn-uploads.huggingface.co/production/uploads/644afe169279988e0cbcd2d9/Kqx6ACbFDrb_UG0Zmah4z.png) Defined values: Q8 = 99%, Q4 = 50%, Q3 = 30%, Q2 = 25%, and Q1 = 10%. ### Q8 — Fidelity For machines with enough memory and users who want to preserve as much of the model's quality as possible. `Q8_0` --- ### Q4 / IQ4 — Balanced The recommended sweet spot for many local systems. A good balance between size, memory usage, and quality. `Q4_K_M` · `Q4_K_S` · `IQ4_XS` · `IQ4_NL` · `Q4_1` · `Q4_0` --- ### Q3 / IQ3 — Compact For more limited hardware or situations where reducing RAM/VRAM usage is a priority. `Q3_K_L` · `Q3_K_M` · `Q3_K_S` · `IQ3_M` · `IQ3_S` · `IQ3_XS` · `IQ3_XXS` --- ### Q2 / IQ2 — Ultra-Compact Aggressive compression for environments where larger versions simply do not fit. `Q2_K` · `Q2_K_S` · `IQ2_M` · `IQ2_S` · `IQ2_XS` · `IQ2_XXS` · `Q2_0` · `TQ2_0` --- ### IQ1 — Minimum The extreme option. `IQ1_M` Recommended only when saving memory is more important than preserving maximum model fidelity. --- > **Rule of thumb:** start with `Q4_K_M`. > If you have memory to spare, try `Q8_0`. > If memory is limited, move down to Q3, Q2, or IQ1. --- # Running Locally ## llama.cpp Download one of the GGUF versions of LEVUM and run it with a GGUF-compatible runtime. ```bash llama-cli \ -m ./LEVUM-Q4_K_M.gguf \ -p "Create a REST API in Python using FastAPI." ``` To start a local server: ```bash llama-server \ -m ./LEVUM-Q4_K_M.gguf \ -c 8192 ``` After that, the model can be integrated with local tools that support compatible endpoints. --- ## Ollama Create a `Modelfile` pointing to the GGUF: ```dockerfile FROM ./LEVUM-Q4_K_M.gguf PARAMETER temperature 0.6 PARAMETER top_p 0.9 ``` Then: ```bash ollama create levum -f Modelfile ollama run levum ``` Example: ```text >>> Create a FastAPI API for managing projects and tasks. ``` --- ## LM Studio 1. Download a GGUF quantization of LEVUM. 2. Import the file into LM Studio. 3. Load the model. 4. Adjust the context size according to the available memory. 5. Start a conversation or the local server. No remote infrastructure is required for inference when the model is running locally. --- # Prompts LEVUM works best when the task, context, and expected output format are explicit. ## Generate a Project ```text Create an interactive financial dashboard in a single HTML file. Requirements: - HTML, CSS, and JavaScript in the same file - interactive charts - responsive - sample data - no mandatory external dependencies Return only the complete HTML. ``` --- ## Implement a Feature ```text Analyze the code below and implement JWT authentication. Requirements: - preserve the current architecture - validate expired tokens - add authentication middleware - do not modify public endpoints - explain only important decisions Code: [paste the code here] ``` --- ## Debug ```text Find the cause of the bug in the code below. Expected behavior: [describe] Current behavior: [describe] Error: [paste the error] Code: [paste the code] Identify the cause and return the complete fix. ``` --- ## Refactoring ```text Refactor this code. Goals: - reduce duplication - improve readability - preserve current behavior - preserve the public API - avoid unnecessary abstractions Return the refactored code first, followed by a short summary of the changes. ``` --- # Prompt Format For larger tasks, a simple structure usually produces more predictable results: ```text OBJECTIVE What needs to be built. CONTEXT Stack, existing files, and architecture. REQUIREMENTS Mandatory behaviors. CONSTRAINTS What must not be changed. OUTPUT Exact expected format. ``` For example: ```text OBJECTIVE Create an analytics page. CONTEXT React + TypeScript + Tailwind project. REQUIREMENTS - revenue chart - period filters - metric cards - transactions table CONSTRAINTS - do not add new dependencies - reuse existing components OUTPUT Return the complete files that need to be created or modified. ``` --- # Code-first LEVUM was designed to work within the normal development cycle: ```text PROMPT / CODE ↓ LEVUM ↓ UNDERSTAND ↓ GENERATE ↓ DEBUG ↓ REFACTOR ↓ RUNNING SOFTWARE ``` ### Generation Projects, components, APIs, scripts, and features. ### Debug Analysis of errors, unexpected behavior, and fixes. ### Refactoring Structure, readability, and maintenance of existing code. ### Prototyping Turn an idea into something executable quickly. --- # Example **Prompt** ```text Create a Flappy Bird-style game in a single HTML file. Use only HTML, CSS, and JavaScript. Include: - physics - obstacles - collision detection - scoring - restart - keyboard and click controls Return only the complete HTML. ``` **Expected result** ```text prompt ↓ LEVUM ↓ HTML + CSS + JavaScript ↓ browser ↓ running software ``` --- # Hardware Actual resource usage depends on several factors, including: * quantization; * context size; * runtime; * KV cache; * CPU; * GPU; * number of layers offloaded to the GPU; * inference configuration. Because of this, memory requirements can vary significantly between systems. As a general rule: ```text more bits ↑ more fidelity ↑ more memory fewer bits ↓ less memory ↓ more compression ``` Choose the quantization based on the available hardware and the quality required for the task. --- # Limitations LEVUM is still a language model. This means it may: * generate incorrect code; * hallucinate APIs or libraries; * produce insecure solutions; * misinterpret requirements; * introduce regressions; * suggest nonexistent dependencies; * generate code that looks correct without actually working. For important software, review, test, and validate the code before putting it into production. AI-generated code does not gain magical powers just because it compiled once. --- # Responsible Use Before running code generated by the model: 1. review the changes; 2. verify dependencies; 3. run tests; 4. validate external inputs; 5. review filesystem, network, and database operations; 6. do not expose secrets or credentials unnecessarily; 7. use isolated environments when testing unknown code. For critical applications, the model should serve as an assistance tool, not as the only layer of review. --- # Base Model LEVUM is based on: **Qwen 3.5 9B** The base model provides the general capabilities upon which LEVUM is built. Also review the base model's model card and license before distributing or using derivatives. --- # License Use of LEVUM is subject to the license published in this repository and, where applicable, the terms and conditions associated with the base model. Review the `LICENSE` file before commercial use, redistribution, or creating derivatives. --- # Brazil
### 🇧🇷 MADE IN BRAZIL **Intelligence that stays close.** Local-first · Code-first · Software-first
`BUILD` · `DEBUG` · `REFACTOR` · `SHIP`
**LEVUM © 2026**