Instructions to use DKTechin/kanana with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use DKTechin/kanana with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf DKTechin/kanana:Q4_K_M # Run inference directly in the terminal: llama cli -hf DKTechin/kanana:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf DKTechin/kanana:Q4_K_M # Run inference directly in the terminal: llama cli -hf DKTechin/kanana:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf DKTechin/kanana:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf DKTechin/kanana:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf DKTechin/kanana:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf DKTechin/kanana:Q4_K_M
Use Docker
docker model run hf.co/DKTechin/kanana:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use DKTechin/kanana with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "DKTechin/kanana" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DKTechin/kanana", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/DKTechin/kanana:Q4_K_M
- Ollama
How to use DKTechin/kanana with Ollama:
ollama run hf.co/DKTechin/kanana:Q4_K_M
- Unsloth Studio
How to use DKTechin/kanana with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for DKTechin/kanana to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for DKTechin/kanana to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for DKTechin/kanana to start chatting
- Pi
How to use DKTechin/kanana with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf DKTechin/kanana:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "DKTechin/kanana:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use DKTechin/kanana with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf DKTechin/kanana:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "DKTechin/kanana:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use DKTechin/kanana with Docker Model Runner:
docker model run hf.co/DKTechin/kanana:Q4_K_M
- Lemonade
How to use DKTechin/kanana with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull DKTechin/kanana:Q4_K_M
Run and chat with the model
lemonade run user.kanana-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use DKTechin/kanana with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf DKTechin/kanana:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default DKTechin/kanana:Q4_K_M
Run Hermes
hermes
- Atomic Chat
| license: other | |
| license_name: kanana | |
| license_link: LICENSE | |
| base_model: kakaocorp/kanana-2-3b-instruct | |
| base_model_relation: quantized | |
| pipeline_tag: text-generation | |
| language: | |
| - ko | |
| - en | |
| tags: | |
| - gguf | |
| - llama.cpp | |
| - quantized | |
| - kanana | |
| # Kanana 2 3B Instruct โ GGUF (Q4_K_M) | |
| **Powered by Kanana** | |
| [`kakaocorp/kanana-2-3b-instruct`](https://huggingface.co/kakaocorp/kanana-2-3b-instruct) ๋ฅผ llama.cpp ์์ ๋ฐ๋ก ์ธ ์ ์๊ฒ GGUF ๋ก ๋ณํํ๊ณ 4bit(Q4_K_M) ์์ํํ ํ์ผ์ ๋๋ค. ์๋ณธ ๋ฐฐํฌ์๋ GGUF ๊ฐ ์์ด ์ฌ๋ด ์จ๋๋ฐ์ด์ค(๋ก์ปฌ LLM) ์ฉ๋๋ก ์ง์ ๋ณํํ์ต๋๋ค. | |
| | | | | |
| |---|---| | |
| | ํ์ผ | `kanana-2-3b-instruct-Q4_K_M.gguf` | | |
| | ํฌ๊ธฐ | 2,161,793,408 bytes (2.01 GiB) | | |
| | ์์ํ | Q4_K_M โ 4.92 BPW (์๋ณธ BF16 16.00 BPW) | | |
| | ์ํคํ ์ฒ | `qwen3` (์๋ณธ `config.json` ์ด `Qwen3ForCausalLM`) | | |
| | ํ ํฌ๋์ด์ | `tokenizer.ggml.pre = kanana2`, vocab 128,256 | | |
| | ์ปจํ ์คํธ | 32,768 (rope: yarn, factor 40, original 4,096) | | |
| | ๋ํ ์์ | ์๋ณธ `chat_template.jinja` ๋ฅผ GGUF ๋ฉํ๋ฐ์ดํฐ์ ํฌํจ | | |
| ## ๋ณ๊ฒฝ ์ฌํญ (Kanana Open License ยง3.1(iii)) | |
| ๊ฐ์ค์น์ ๊ฐ์ ๋ฐ๊พธ๋ ํ์ตยทํ์ธํ๋ยท๋ณํฉ์ ํ์ง ์์์ต๋๋ค. ์๋ณธ `safetensors` ๋ฅผ ์๋ ์ ์ฐจ๋ก **ํ์ ๋ณํ + 4bit ์์ํ**๋ง ํ์ต๋๋ค. | |
| ```bash | |
| # llama.cpp @ 7e1e28c (2026-07-28) | |
| python convert_hf_to_gguf.py kanana-2-3b-instruct --outtype bf16 \ | |
| --outfile kanana-2-3b-instruct-BF16.gguf | |
| ./build/bin/llama-quantize kanana-2-3b-instruct-BF16.gguf \ | |
| kanana-2-3b-instruct-Q4_K_M.gguf Q4_K_M | |
| ``` | |
| 4bit ์์ํ๋ ์๋ณธ ๋๋น ํ์ง ์์ค์ ์๋ฐํฉ๋๋ค. ์๋ณธ ํ์ง์ด ํ์ํ๋ฉด ์ ๋งํฌ์ ์๋ณธ ์ ์ฅ์๋ฅผ ์ฐ์ธ์. | |
| ## ์ฌ์ฉ๋ฒ | |
| ```bash | |
| # llama.cpp | |
| llama-cli -hf DKTechin/kanana:Q4_K_M -c 2048 --jinja \ | |
| -sys "๋น์ ์ ์ฌ๋ด ๋ฉ์ ์ ๋ํ๋ฅผ ๊ฐ๊ฒฐํ๊ฒ ์์ฝํ๋ ๋์ฐ๋ฏธ์ ๋๋ค." \ | |
| -p "์๋ ๋ํ๋ฅผ 3์ค๋ก ์์ฝํด์ค: ..." | |
| ``` | |
| LM StudioยทJanยทOllama ๋ฑ GGUF ๋ฅผ ์ฝ๋ ๋ฐํ์์์๋ ๊ทธ๋๋ก ์๋๋ค. ์์คํ ํ๋กฌํํธ ์์ด ์ฐ๋ฉด ์์ฝ ๋์ ์ ๋ ฅ์ ๋ํ์ดํ๋ ๊ฒฝํฅ์ด ์์ด, ์ญํ ์ ์ง์ ํ๋ ์์คํ ํ๋กฌํํธ๋ฅผ ํจ๊ป ์ฃผ๋ ํธ์ด ์์ ํฉ๋๋ค. | |
| ## ํ์ธํ ๊ฒ ยท ํ์ธํ์ง ๋ชปํ ๊ฒ | |
| - ํ์ธ: ๋ฉํ๋ฐ์ดํฐ(archยทtokenizer preยทropeยทchat template), ์ค์ ํ๊ตญ์ด ์์ฝ ์์ฑ, macOS/Metal ๊ธฐ์ค 2048 ํ ํฐ ์กฐ๊ฑด์์ ํ๋กฌํํธ 1,375 t/s ยท ์์ฑ 128 t/s. | |
| - ๋ฏธํ์ธ: 32k ์ฅ๋ฌธ ํ์ง. ์๋ณธ `config.json` ์ rope ๋ฐฐ์๋ฅผ 40 ์ผ๋ก ์ ์ด ๋์์ง๋ง ๊ธธ์ด ๋น์จ(32,768 / 4,096)๋ก ๊ณ์ฐํ๋ฉด 8 ์ด๋ผ ์๋ก ๋ง์ง ์์ต๋๋ค. ์๋ณธ ํ์ด์ฌ ๋ฐํ์๊ณผ ๋์์ ๋ง์ถ๊ธฐ ์ํด **๋ช ์๊ฐ 40 ์ ๊ทธ๋๋ก ๊ธฐ๋ก**ํ์ผ๋, ์์ฃผ ๊ธด ์ ๋ ฅ์์ ์ด์ํ๋ฉด ์ด ๊ฐ์ ๋จผ์ ์์ฌํ์ธ์. | |
| ## ๋ผ์ด์ ์ค | |
| ์ด ํ์ผ์ [Kanana Open License Agreement](LICENSE) ๋ฅผ ๋ฐ๋ฅด๋ ํ์๋ฌผ์ ๋๋ค. ์ฌ๋ณธ์ ์ด ์ ์ฅ์์ [`LICENSE`](LICENSE), ๊ณ ์ง ๋ฌธ๊ตฌ๋ [`NOTICE`](NOTICE) ์ ์์ต๋๋ค. | |
| - ์ฌ์ฉ์๋ ์๋ณธ๊ณผ ๋์ผํ๊ฒ **๊ธ์ง๋ ์ฌ์ฉ ์ ์ฑ (Agreement ยง2.2)** ๊ณผ KAKAO ์ Guidelines For Responsible AI ๋ฅผ ์ค์ํด์ผ ํ๋ฉฐ, ์ด ํ์ผ์ ์ฌ๋ฐฐํฌํ ๋์๋ ๊ฐ์ ์๋ฌด๋ฅผ ํ์ ์ฌ์ฉ์์๊ฒ ์๋ ค์ผ ํฉ๋๋ค. | |
| - APIยทํด๋ผ์ฐ๋ ๋ฑ์ผ๋ก ์ 3์์๊ฒ ์ ๊ทผ์ ์ ๊ณตํ๊ฑฐ๋ ์ฌํ๋งคํ๋ ค๋ฉด **KAKAO ์ ๋ณ๋ ์์ ๋ผ์ด์ ์ค**๊ฐ ํ์ํฉ๋๋ค(Agreement ยง4). | |
| - ์ด ํ์ผ์ ์ฌ์ฉํ๋ ์น์ฌ์ดํธยทUIยท๋ฌธ์์๋ **"Powered by Kanana"** ๋ฅผ ์์๋ณผ ์ ์๊ฒ ํ์ํด์ผ ํฉ๋๋ค(Agreement ยง3.1(v)). | |
| ``` | |
| Kanana is licensed in accordance with the Kanana Open License Agreement. | |
| Copyright ยฉ KAKAO Corp. All Rights Reserved. | |
| ``` | |