--- license: apache-2.0 base_model: CohereLabs/North-Mini-Code-1.0 tags: - gguf - llama.cpp - cohere - code - moe - quantized - north-mini-code pipeline_tag: text-generation language: - en --- # North-Mini-Code-1.0-GGUF GGUF conversions and quantizations of **CohereLabs/North-Mini-Code-1.0** for use with: - llama.cpp - LM Studio - Ollama - Jan - KoboldCpp - Text Generation WebUI - Open WebUI - Other GGUF-compatible runtimes --- # About the Model North-Mini-Code-1.0 is a code-focused Mixture-of-Experts (MoE) model released by CohereLabs. This repository provides ready-to-use GGUF conversions for local inference across a range of hardware configurations. --- # Available Files ### Full Precision - `North-Mini-Code-1.0-F16.gguf` ### Quantized Versions - `North-Mini-Code-1.0-Q4_K_M.gguf` - `North-Mini-Code-1.0-Q5_K_M.gguf` - `North-Mini-Code-1.0-Q6_K.gguf` - `North-Mini-Code-1.0-Q8_0.gguf` --- # Recommended Quantization For most users: ```text North-Mini-Code-1.0-Q4_K_M.gguf ``` It offers the best balance of: - Quality - Memory usage - Inference speed If you have more available RAM/VRAM, consider: ```text North-Mini-Code-1.0-Q5_K_M.gguf ``` or ```text North-Mini-Code-1.0-Q6_K.gguf ``` for slightly higher output quality. --- # Approximate File Sizes ```text F16 ~60+ GB Q4_K_M ~20 GB Q5_K_M ~23 GB Q6_K ~27 GB Q8_0 ~34 GB ``` Actual sizes may vary slightly depending on conversion tooling versions. --- # Usage ## llama.cpp Prompt mode: ```bash ./llama-cli \ -m North-Mini-Code-1.0-Q4_K_M.gguf \ -p "Write a Python function that reverses a linked list." ``` Chat mode: ```bash ./llama-cli \ -m North-Mini-Code-1.0-Q4_K_M.gguf \ -cnv ``` --- ## LM Studio 1. Download your preferred GGUF file. 2. Open LM Studio. 3. Import the model. 4. Start chatting. --- ## Ollama Create a `Modelfile`: ```text FROM North-Mini-Code-1.0-Q4_K_M.gguf ``` Create the model: ```bash ollama create north-mini-code -f Modelfile ``` Run it: ```bash ollama run north-mini-code ``` --- # Hardware Recommendations ### Q4_K_M Recommended minimum: ```text 24 GB RAM ``` ### Q5_K_M Recommended minimum: ```text 32 GB RAM ``` ### Q6_K Recommended minimum: ```text 32-40 GB RAM ``` ### Q8_0 Recommended minimum: ```text 48+ GB RAM ``` ### F16 Recommended minimum: ```text 80+ GB RAM ``` --- # Prompting Tips This model is optimized for programming-related tasks. Example prompts: ```text Implement a fast Rust HTTP server. ``` ```text Explain this C++ compiler error. ``` ```text Write comprehensive unit tests for the following Python code. ``` ```text Convert this JavaScript function to TypeScript. ``` ```text Optimize this SQL query. ``` --- # Base Model Base model: ```text CohereLabs/North-Mini-Code-1.0 ``` All training, architecture, benchmarks, licensing terms, and usage restrictions belong to the original model authors. Please refer to the original repository for official documentation and licensing information. --- # Conversion Details Converted using: ```text llama.cpp ``` Generated quantizations: ```text F16 Q4_K_M Q5_K_M Q6_K Q8_0 ``` A tokenizer compatibility workaround was applied during conversion to support current GGUF conversion tooling. --- # Credits - Base Model: CohereLabs - GGUF Conversion & Quantization: NANI-Nithin - Tooling: llama.cpp --- # Repository 👉 https://huggingface.co/NANI-Nithin/north-mini-code-gguf