NANI-Nithin's picture
Create README.md
5ffb616 verified
|
Raw
History Blame Contribute Delete
3.46 kB
---
license: apache-2.0
base_model: CohereLabs/North-Mini-Code-1.0
tags:
- gguf
- llama.cpp
- cohere
- code
- moe
- quantized
- north-mini-code
pipeline_tag: text-generation
language:
- en
---
# North-Mini-Code-1.0-GGUF
GGUF conversions and quantizations of **CohereLabs/North-Mini-Code-1.0** for use with:
- llama.cpp
- LM Studio
- Ollama
- Jan
- KoboldCpp
- Text Generation WebUI
- Open WebUI
- Other GGUF-compatible runtimes
---
# About the Model
North-Mini-Code-1.0 is a code-focused Mixture-of-Experts (MoE) model released by CohereLabs.
This repository provides ready-to-use GGUF conversions for local inference across a range of hardware configurations.
---
# Available Files
### Full Precision
- `North-Mini-Code-1.0-F16.gguf`
### Quantized Versions
- `North-Mini-Code-1.0-Q4_K_M.gguf`
- `North-Mini-Code-1.0-Q5_K_M.gguf`
- `North-Mini-Code-1.0-Q6_K.gguf`
- `North-Mini-Code-1.0-Q8_0.gguf`
---
# Recommended Quantization
For most users:
```text
North-Mini-Code-1.0-Q4_K_M.gguf
```
It offers the best balance of:
- Quality
- Memory usage
- Inference speed
If you have more available RAM/VRAM, consider:
```text
North-Mini-Code-1.0-Q5_K_M.gguf
```
or
```text
North-Mini-Code-1.0-Q6_K.gguf
```
for slightly higher output quality.
---
# Approximate File Sizes
```text
F16 ~60+ GB
Q4_K_M ~20 GB
Q5_K_M ~23 GB
Q6_K ~27 GB
Q8_0 ~34 GB
```
Actual sizes may vary slightly depending on conversion tooling versions.
---
# Usage
## llama.cpp
Prompt mode:
```bash
./llama-cli \
-m North-Mini-Code-1.0-Q4_K_M.gguf \
-p "Write a Python function that reverses a linked list."
```
Chat mode:
```bash
./llama-cli \
-m North-Mini-Code-1.0-Q4_K_M.gguf \
-cnv
```
---
## LM Studio
1. Download your preferred GGUF file.
2. Open LM Studio.
3. Import the model.
4. Start chatting.
---
## Ollama
Create a `Modelfile`:
```text
FROM North-Mini-Code-1.0-Q4_K_M.gguf
```
Create the model:
```bash
ollama create north-mini-code -f Modelfile
```
Run it:
```bash
ollama run north-mini-code
```
---
# Hardware Recommendations
### Q4_K_M
Recommended minimum:
```text
24 GB RAM
```
### Q5_K_M
Recommended minimum:
```text
32 GB RAM
```
### Q6_K
Recommended minimum:
```text
32-40 GB RAM
```
### Q8_0
Recommended minimum:
```text
48+ GB RAM
```
### F16
Recommended minimum:
```text
80+ GB RAM
```
---
# Prompting Tips
This model is optimized for programming-related tasks.
Example prompts:
```text
Implement a fast Rust HTTP server.
```
```text
Explain this C++ compiler error.
```
```text
Write comprehensive unit tests for the following Python code.
```
```text
Convert this JavaScript function to TypeScript.
```
```text
Optimize this SQL query.
```
---
# Base Model
Base model:
```text
CohereLabs/North-Mini-Code-1.0
```
All training, architecture, benchmarks, licensing terms, and usage restrictions belong to the original model authors.
Please refer to the original repository for official documentation and licensing information.
---
# Conversion Details
Converted using:
```text
llama.cpp
```
Generated quantizations:
```text
F16
Q4_K_M
Q5_K_M
Q6_K
Q8_0
```
A tokenizer compatibility workaround was applied during conversion to support current GGUF conversion tooling.
---
# Credits
- Base Model: CohereLabs
- GGUF Conversion & Quantization: NANI-Nithin
- Tooling: llama.cpp
---
# Repository
👉 https://huggingface.co/NANI-Nithin/north-mini-code-gguf