Instructions to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M # Run inference directly in the terminal: llama cli -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M # Run inference directly in the terminal: llama cli -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Use Docker
docker model run hf.co/aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "aogavrilov/diffusiongemma-agent-iq3-cuda13" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aogavrilov/diffusiongemma-agent-iq3-cuda13", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
- Ollama
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Ollama:
ollama run hf.co/aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
- Unsloth Studio
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for aogavrilov/diffusiongemma-agent-iq3-cuda13 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for aogavrilov/diffusiongemma-agent-iq3-cuda13 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for aogavrilov/diffusiongemma-agent-iq3-cuda13 to start chatting
- Pi
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Docker Model Runner:
docker model run hf.co/aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
- Lemonade
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Run and chat with the model
lemonade run user.diffusiongemma-agent-iq3-cuda13-Q4_K_M
List all available models
lemonade list
| # Agent-Side RAG | |
| The fast local model cannot hold a whole repository in prompt. The practical | |
| agent pattern is: | |
| 1. Search locally with `rg`. | |
| 2. Build a tiny file map. | |
| 3. Read only matching snippets. | |
| 4. Send the compact context to the OpenAI-compatible local endpoint. | |
| 5. Let an outer agent or human apply and test patches. | |
| This repo includes a minimal read-only wrapper: | |
| ```text | |
| scripts/rag_code_agent.py | |
| scripts/dg_agent.sh rag | |
| .dg-agent/bin/rag | |
| scripts/dg_agent.sh repo-pack | |
| .dg-agent/bin/repo-pack | |
| scripts/dg_agent.sh repo-map | |
| .dg-agent/bin/repo-map | |
| scripts/dg_agent.sh ast-grep | |
| .dg-agent/bin/ast-grep | |
| scripts/dg_agent.sh code-outline | |
| .dg-agent/bin/code-outline | |
| ``` | |
| It does not edit files. It retrieves context and asks the currently running | |
| local model for the next action or a patch. | |
| ## Start The Fast Backend | |
| ```powershell | |
| powershell -ExecutionPolicy Bypass -File \\wsl.localhost\Ubuntu-24.04\root\diffusiongemma-agent\scripts\start_agent_fast_service_windows.ps1 -StopExisting | |
| ``` | |
| The expected endpoint is: | |
| ```text | |
| http://127.0.0.1:4100/v1 | |
| ``` | |
| ## Ask About A Repo | |
| From WSL: | |
| ```bash | |
| cd /root/diffusiongemma-agent | |
| scripts/dg_agent.sh rag --repo /path/to/repo \ | |
| --task "Find where the server starts and explain how to change the port" | |
| ``` | |
| Preview only the retrieved context: | |
| ```bash | |
| scripts/dg_agent.sh rag --repo /path/to/repo \ | |
| --task "where is CUDA env configured?" \ | |
| --print-context | |
| ``` | |
| After `workspace-init`, use the repo-local launcher: | |
| ```bash | |
| .dg-agent/bin/rag --task "where is CUDA env configured?" --print-context | |
| ``` | |
| ## Client Settings | |
| The fast backend has `MAXTOK=768`, so keep retrieval compact: | |
| ```text | |
| --max-context-chars 500-900 | |
| --max-files 2-3 | |
| --max-tokens 128-256 | |
| ``` | |
| For file-level coding tasks, ask for one scoped change at a time: | |
| ```bash | |
| scripts/dg_agent.sh rag --repo /path/to/repo \ | |
| --max-context-chars 650 --max-files 2 --max-tokens 128 \ | |
| --task "In server.py, find the health endpoint and propose the smallest patch to add uptime_ms" | |
| ``` | |
| The same retrieval path is exposed to MCP clients as `dg_rag_context` and | |
| `dg_rag_answer`. Prefer `dg_rag_context` when an external agent should inspect | |
| the retrieved file map/snippets before deciding whether to call the model. | |
| For OSS repository packing, use the Repomix wrapper: | |
| ```bash | |
| scripts/dg_agent.sh repo-pack --repo /path/to/repo \ | |
| --include "src/**" \ | |
| --style markdown \ | |
| --compress \ | |
| --token-budget 20000 \ | |
| --stdout | |
| ``` | |
| The MCP tool name is `dg_repo_pack`. Prefer tight include filters so the packed | |
| artifact stays within the local model's small working context. | |
| For an Aider-style repository sketch, use: | |
| ```bash | |
| scripts/dg_agent.sh repo-map --repo /path/to/repo \ | |
| --map-tokens 512 \ | |
| --map-only | |
| ``` | |
| The MCP tool name is `dg_repo_map`. It uses upstream Aider's repo-map logic but | |
| keeps history files temporary and passes `--no-gitignore`, so it should not | |
| dirty the target repo. | |
| For structural search, use the upstream ast-grep wrapper: | |
| ```bash | |
| scripts/dg_agent.sh ast-grep --repo /path/to/repo \ | |
| --lang python \ | |
| --pattern 'return $X' \ | |
| --json | |
| ``` | |
| The MCP tool name is `dg_ast_grep`. Use it when an agent needs language-aware | |
| matches instead of raw text matches from `rg`. | |
| For symbol maps, use the upstream ast-grep outline wrapper: | |
| ```bash | |
| scripts/dg_agent.sh code-outline --repo /path/to/repo \ | |
| --lang python \ | |
| --view expanded \ | |
| --json | |
| ``` | |
| The MCP tool name is `dg_code_outline`. Use it before file reads when class, | |
| function, import, or member names are enough to choose the next file. | |
| ## Why This Works | |
| The model sees only a small, relevant working set instead of the whole repo. | |
| The agent wrapper owns tools such as `rg`, file reading, patch application, git | |
| status, and tests. The model only reasons over the selected snippets. | |