πŸ°πŸ‘‘ ChatBerry GGUF

https://i.pinimg.com/736x/bd/b9/be/bdb9bef25d336c9887e351cd1d7bfd57.jpg

GGUF quantizations of artindnr/chatberry, the flagship release of the ChatBerry family β€” a direct-answer chat fine-tune of artindnr/strawberry-1 (itself built on openai/gpt-oss-20b). ChatBerry's reasoning ("thinking") channel is disabled, so it responds directly instead of emitting a separate chain-of-thought trace.

These quants let you run ChatBerry locally with llama.cpp, Ollama, LM Studio, or any other GGUF-compatible runtime.

Files

Bits Quant Size Notes
2-bit chatberry.Q2_K.gguf 12.1 GB Smallest, most quality loss β€” only for tight memory budgets
3-bit chatberry.Q3_K_S.gguf 12.1 GB Small, low quality tier
3-bit chatberry.Q3_K_M.gguf 12.9 GB Balanced within the 3-bit tier
3-bit chatberry.Q3_K_L.gguf 13.3 GB Largest/highest quality of the 3-bit quants
4-bit chatberry.IQ4_XS.gguf 12.2 GB Compact 4-bit variant, good quality-per-GB
4-bit chatberry.Q4_K_S.gguf 14.7 GB Solid general-purpose 4-bit quant
4-bit chatberry.Q4_K_M.gguf 15.8 GB Recommended default β€” good balance of quality and size
5-bit chatberry.Q5_K_S.gguf 15.9 GB Higher fidelity, moderate size increase
5-bit chatberry.Q5_K_M.gguf 16.9 GB Best 5-bit option if you have the headroom
6-bit chatberry.Q6_K.gguf 22.2 GB Near-lossless, larger file
8-bit chatberry.Q8_0.gguf 22.3 GB Highest quality of the set, closest to full precision

If you're unsure which to pick: Q4_K_M is a solid default for most setups. Go up to Q5_K_M, Q6_K, or Q8_0 if you have the VRAM/RAM to spare and want maximum fidelity, and drop to Q3_K_M/Q2_K if you're tightly memory-constrained.

Usage

llama.cpp

./llama-cli -m chatberry.Q4_K_M.gguf -p "Ψͺو کی Ω‡Ψ³Ψͺی و Ψ§Ψ³Ω…Ψͺ Ϊ†ΫŒΩ‡ΨŸ" -n 512

Or serve it as an OpenAI-compatible endpoint:

./llama-server -m chatberry.Q4_K_M.gguf -c 4096

Ollama

Create a Modelfile:

FROM ./chatberry.Q4_K_M.gguf

Then:

ollama create chatberry -f Modelfile
ollama run chatberry

LM Studio

Download the .gguf file of your choice directly in LM Studio's model browser (search artindnr/chatberry-gguf), or drop the file into your local models folder.

About ChatBerry

ChatBerry is the culmination of the iterative ChatBerry line (1.0 β†’ 1.1 β†’ 1.2 β†’ ChatBerry), each release trained on more data than the last. It converts the reasoning behavior of strawberry-1 into a direct-answer chat model β€” no visible analysis/chain-of-thought channel, just a final response.

  • Base model: artindnr/strawberry-1 (fine-tuned from openai/gpt-oss-20b, 21B parameters)
  • Architecture: gpt_oss
  • Languages: Farsi (Persian), English, and multilingual support
  • Behavior: Reasoning disabled β€” responds directly without a separate analysis channel
  • Format: GGUF, for use with llama.cpp and compatible runtimes

See the full model card for training details and intended use.

Chat Template

ChatBerry uses the gpt-oss chat template (Harmony format) inherited from its base model. Most GGUF runtimes (llama.cpp, Ollama, LM Studio) apply this automatically when loading the model β€” no need to set a reasoning-language system message or parse separate channels; ChatBerry goes straight to its final answer.

Limitations

  • Quantization introduces some precision loss versus the original bf16 weights β€” expect small quality differences across the tiers, most noticeable at Q2_K/Q3_K.
  • ChatBerry trades away Strawberry-1's explicit chain-of-thought reasoning; for tasks that benefit from visible step-by-step reasoning, artindnr/strawberry-1 may be a better fit.
  • Inherits the general capabilities and limitations of the gpt-oss-20b base model and strawberry-1, including the possibility of hallucinated facts.
  • No formal safety fine-tuning beyond what is inherited from the base model and Strawberry-1 has been applied; use appropriate safeguards in production settings.

License

Released under the Apache 2.0 license, consistent with the base gpt-oss-20b model, strawberry-1, and chatberry.

Acknowledgements

Downloads last month
-
GGUF
Model size
21B params
Architecture
gpt-oss
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for artindnr/ChatBerry-GGUF

Quantized
(3)
this model

Collection including artindnr/ChatBerry-GGUF