Instructions to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M # Run inference directly in the terminal: llama cli -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M # Run inference directly in the terminal: llama cli -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Use Docker
docker model run hf.co/aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "aogavrilov/diffusiongemma-agent-iq3-cuda13" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aogavrilov/diffusiongemma-agent-iq3-cuda13", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
- Ollama
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Ollama:
ollama run hf.co/aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
- Unsloth Studio
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for aogavrilov/diffusiongemma-agent-iq3-cuda13 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for aogavrilov/diffusiongemma-agent-iq3-cuda13 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for aogavrilov/diffusiongemma-agent-iq3-cuda13 to start chatting
- Pi
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Docker Model Runner:
docker model run hf.co/aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
- Lemonade
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Run and chat with the model
lemonade run user.diffusiongemma-agent-iq3-cuda13-Q4_K_M
List all available models
lemonade list
File size: 5,818 Bytes
ef2127a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 | # OpenCode Local Profile
OpenCode is the primary ready-made open-source terminal-agent shell wired into
this Windows checkout. It runs locally from `.tools/opencode`; it does not
install a global npm package.
```text
OpenCode on Windows -> safe agent gateway :8090 -> WSL GPU model service :4100
```
The gateway restricts this model to a `768`-token input and `256`-token output
budget. It performs tool delegation and repository operations outside the
model, which is necessary for the current small-context DiffusionGemma setup.
## Install and Start
From PowerShell in this repository:
```powershell
.\scripts\install_opencode_windows.ps1
.\scripts\start_agent_gateway.ps1
Invoke-RestMethod http://127.0.0.1:8090/healthz
```
`install_opencode_windows.ps1` installs the upstream `opencode-ai` package
locally and explicitly runs its post-install binary setup. The gateway forwards
to the existing GPU model at `http://127.0.0.1:4100/v1`; it does not restart or
move the model.
## Use in a Repository
Start OpenCode from the target Git repository so its file tools and MCP server
are scoped to that repository:
```powershell
Set-Location C:\path\to\target-repo
C:\Users\alexg\Downloads\diffusiongemma-agent\scripts\run_opencode_windows.ps1
```
For a bounded non-interactive request:
```powershell
C:\Users\alexg\Downloads\diffusiongemma-agent\scripts\run_opencode_windows.ps1 run `
--format json `
--model diffusiongemma-local/diffusiongemma-26b-a4b-it-iq4xs-aider-local `
'Read src/app.py and explain the request flow. Do not edit files.'
```
## Primary Compact Delegate
For the practical Codex-like local workflow on this machine, use the compact
OpenCode profile instead of the generic one:
```powershell
Set-Location C:\path\to\target-repo
C:\Users\alexg\Downloads\diffusiongemma-agent\scripts\run_opencode_agent_windows.ps1
```
For one non-interactive task:
```powershell
C:\Users\alexg\Downloads\diffusiongemma-agent\scripts\run_opencode_agent_windows.ps1 run `
--format json `
'Fix src\math_utils.py so add(a, b) returns the sum of its two arguments. Verify the change.'
```
This profile exposes only OpenCode's built-in `bash` tool. The safe gateway
immediately redirects that call to the local DG workflow: read-only requests
use compact repository retrieval; edit requests use the persistent supervisor,
checkpointed session runner, verification, and rollback-on-failure. DiffusionGemma does not need to perform
native tool selection, which is unreliable for this runtime.
The launcher sets `OPENCODE_EXPERIMENTAL_BASH_DEFAULT_TIMEOUT_MS=450000` for
this profile so OpenCode does not interrupt the bounded 420-second edit
session. It restores the previous environment value on exit. Narrow, verified
deterministic repairs such as explicit Python return expressions and an
explicit two-argument sum/difference/product/quotient complete without a
model generation round-trip; broader edits still use Aider and may reach their
own timeout.
The same launcher is used by native Windows `dg_agent.py opencode`,
`opencode-mcp`, and `opencode-acp` commands. Provider discovery can run without
MCP:
```powershell
.\scripts\run_opencode_windows.ps1 -NoMcp models diffusiongemma-local
```
## MCP and Safety
By default, the Windows launcher creates a temporary OpenCode config that
mounts exactly one MCP server: `dg_agent`. It starts that server through WSL,
passes the current Windows repository path as `DG_MCP_REPO`, and removes the
temporary config on exit.
```powershell
.\scripts\run_opencode_windows.ps1 mcp list
```
Serena is intentionally not mounted by this launcher. Its installed Windows
environment is separate from the working WSL Serena runtime. Keeping only
`dg_agent` in OpenCode's temporary profile bounds the tool schema for the
768-token model; IDE client profiles can mount Serena alongside DG MCP.
Read-only tasks delegate to the bounded read agent. Edit requests delegate to
the artifacted persistent supervisor, which selects files, verifies syntax and
optional tests, and can reverse only its own tracked diff when it starts from a
clean worktree. The runner uses the dedicated WSL Aider runtime for scoped
file edits and keeps Aider history in a temporary directory rather than the
target repository. Narrow deterministic repairs remain available as a fallback
for exact replacements and checked Python return-expression changes.
For non-interactive `opencode run`, the Windows runner propagates a nonzero
exit code when the delegated DG session reports failure. Automation should use
that exit code and the session report, not a textual model summary. File names
appearing after a `do not modify` constraint are excluded from bounded edit
selection.
## Validation
This host has verified all of the following against the live GPU gateway:
- OpenCode provider discovery and `dg_agent` MCP connection.
- A read-only file request through the OpenCode `bash` tool, PowerShell bridge,
and WSL read agent with no file mutation.
- A scoped Python edit through the same route, with a verified Git diff and
preserved session/task artifacts.
- Aider `0.86.2` through the WSL Python `3.12` runtime, including a verified
file-level edit with no `.aider*` or `__pycache__` artifacts in the target
repository.
The gateway itself continues to use WSL Python `3.14`; Aider runs separately
from `/root/diffusiongemma-agent/.venv-aider/bin/python` on Python `3.12`.
Use explicit file hints and small tasks, not broad repository-wide requests,
because the model budget is still `768` input tokens and `256` output tokens.
For semantic navigation before a wider task, use Serena from an IDE MCP bundle
or run `repo-map`/`code-outline`; Serena is intentionally excluded from the
compact OpenCode path because its startup time exceeds OpenCode's MCP connect
budget.
|