Text Generation
GGUF
llama.cpp
qwen2.5
quantized
llama.rn
on-device
lora
companion
text-rewriting
conversational
Instructions to use Depthark/activegotchi-ai with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use Depthark/activegotchi-ai with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="Depthark/activegotchi-ai", filename="activegotchi-ai-v1.0.0-Q4_K_M.gguf", )
llm.create_chat_completion( messages = [ { "role": "user", "content": "What is the capital of France?" } ] ) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Depthark/activegotchi-ai with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Depthark/activegotchi-ai:Q4_K_M # Run inference directly in the terminal: llama cli -hf Depthark/activegotchi-ai:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Depthark/activegotchi-ai:Q4_K_M # Run inference directly in the terminal: llama cli -hf Depthark/activegotchi-ai:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Depthark/activegotchi-ai:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Depthark/activegotchi-ai:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Depthark/activegotchi-ai:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Depthark/activegotchi-ai:Q4_K_M
Use Docker
docker model run hf.co/Depthark/activegotchi-ai:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Depthark/activegotchi-ai with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Depthark/activegotchi-ai" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Depthark/activegotchi-ai", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Depthark/activegotchi-ai:Q4_K_M
- Ollama
How to use Depthark/activegotchi-ai with Ollama:
ollama run hf.co/Depthark/activegotchi-ai:Q4_K_M
- Unsloth Studio
How to use Depthark/activegotchi-ai with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Depthark/activegotchi-ai to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Depthark/activegotchi-ai to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Depthark/activegotchi-ai to start chatting
- Pi
How to use Depthark/activegotchi-ai with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Depthark/activegotchi-ai:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Depthark/activegotchi-ai:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use Depthark/activegotchi-ai with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Depthark/activegotchi-ai:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Depthark/activegotchi-ai:Q4_K_M
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use Depthark/activegotchi-ai with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Depthark/activegotchi-ai:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Depthark/activegotchi-ai:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use Depthark/activegotchi-ai with Docker Model Runner:
docker model run hf.co/Depthark/activegotchi-ai:Q4_K_M
- Lemonade
How to use Depthark/activegotchi-ai with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Depthark/activegotchi-ai:Q4_K_M
Run and chat with the model
lemonade run user.activegotchi-ai-Q4_K_M
List all available models
lemonade list
File size: 2,914 Bytes
a99c714 18ff929 a99c714 18ff929 ba88c00 a99c714 ba88c00 a99c714 18ff929 ba88c00 18ff929 a99c714 ba88c00 a99c714 ba88c00 a99c714 ba88c00 a99c714 ba88c00 a99c714 ba88c00 a99c714 ba88c00 a99c714 ba88c00 a99c714 ba88c00 a99c714 ba88c00 a99c714 ba88c00 a99c714 6bf588f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 | ---
license: apache-2.0
base_model: Qwen/Qwen2.5-0.5B-Instruct
base_model_relation: finetune
pipeline_tag: text-generation
library_name: llama.cpp
language: [en, cs, sk, de, fr, es, it, pt, ja, ko, zh]
tags:
- gguf
- qwen2.5
- quantized
- llama.cpp
- llama.rn
- on-device
- lora
- companion
- text-rewriting
---
# ActiveGotchi AI — v1.0.0
A tiny multilingual **companion-voice rewriting model**. Given a short,
pre-computed activity or sleep summary (steps, minutes, hours, streaks), it
rewrites the text as a warm, playful letter from a virtual pet — in the
requested language, keeping every number, name and emoji exactly as given.
It is a narrow specialist: not a chatbot, not an assistant, not a coach, and
not a medical tool.
## Lineage
| | |
|---|---|
| Parent model | [Qwen/Qwen2.5-0.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct) (494M params, Apache-2.0) |
| Fine-tune | LoRA (2.93M trainable params, ~0.6%), merged into the base |
| Training | 3-stage curriculum (voice → summary rewriting → long-term reviews) + mixed rehearsal pass, on ~1M fully synthetic multilingual samples — no real user data |
| Quantization | GGUF quants of the merged fine-tune, produced with [llama.cpp](https://github.com/ggml-org/llama.cpp) `llama-quantize` |
| Released | 2026-07-19 |
## Files
| File | Size |
|---|---|
| `activegotchi-ai-v1.0.0-Q4_K_M.gguf` | 398 MB |
| `activegotchi-ai-v1.0.0-Q5_K_M.gguf` | 420 MB |
**Recommended:** `Q4_K_M` — the best size/quality balance for phones and
other memory-constrained devices.
## Usage
ChatML prompt format, built into the GGUF chat template — llama.cpp,
llama.rn, LM Studio, Ollama etc. apply it automatically. Send ONE user
message that states the persona/task and ends with the text to rewrite:
```
You are a small cheerful companion. Rewrite the following daily summary for
your human in your own voice. Keep every number exactly as given, keep it to
3-4 short sentences. Write your entire reply in Czech. Do not use any other
language. SUMMARY TO REWRITE: <template letter with the real numbers>
```
Suggested inference settings: `temperature 0.6`, `max tokens 220`,
context 2048.
```sh
llama-cli -m activegotchi-ai-v1.0.0-Q4_K_M.gguf -st \
-p "<prompt as above>" -n 220 --temp 0.6
```
## Trained behavior
- Never emits a number that is not present in the prompt.
- Replies only in the language the prompt pins.
- Short letters (3–6 sentences), no greetings, no filler.
- Warm, non-judgmental tone; no medical advice, diagnosis or shaming.
## Limitations
The model only knows what the prompt contains — it has no memory, no health
knowledge, and no general-assistant abilities. Outside its rewrite task,
output quality is undefined. Languages beyond en/cs/de/es/fr received less
training weight. Synthetic-data style ceiling applies.
## Release notes
Automated release from run_full_flow_mac.sh
|