Instructions to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M # Run inference directly in the terminal: llama cli -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M # Run inference directly in the terminal: llama cli -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Use Docker
docker model run hf.co/aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "aogavrilov/diffusiongemma-agent-iq3-cuda13" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aogavrilov/diffusiongemma-agent-iq3-cuda13", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
- Ollama
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Ollama:
ollama run hf.co/aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
- Unsloth Studio
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for aogavrilov/diffusiongemma-agent-iq3-cuda13 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for aogavrilov/diffusiongemma-agent-iq3-cuda13 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for aogavrilov/diffusiongemma-agent-iq3-cuda13 to start chatting
- Pi
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Docker Model Runner:
docker model run hf.co/aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
- Lemonade
How to use aogavrilov/diffusiongemma-agent-iq3-cuda13 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull aogavrilov/diffusiongemma-agent-iq3-cuda13:Q4_K_M
Run and chat with the model
lemonade run user.diffusiongemma-agent-iq3-cuda13-Q4_K_M
List all available models
lemonade list
| set -euo pipefail | |
| DG_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" | |
| PYTHON="${DG_HAYSTACK_PYTHON:-$DG_ROOT/.venv-haystack/bin/python}" | |
| REPO="" | |
| CONFIG="" | |
| HELP_LOCAL=0 | |
| DRY_RUN=0 | |
| usage() { | |
| cat <<'EOF' | |
| Runs Haystack BM25 RAG against the local DiffusionGemma OpenAI-compatible profile. | |
| Usage: | |
| scripts/run_haystack_local.sh [--repo PATH] [--config PATH] [--dry-run|--smoke-import|--task TEXT] | |
| Default config: | |
| configs/client_profiles/haystack.dg.json | |
| Examples: | |
| scripts/run_haystack_local.sh --repo /repo --dry-run | |
| scripts/run_haystack_local.sh --repo /repo --smoke-import | |
| scripts/run_haystack_local.sh --repo /repo --task "Where is add(a, b) implemented?" | |
| EOF | |
| } | |
| while [[ $# -gt 0 ]]; do | |
| case "$1" in | |
| --repo) | |
| REPO="$2" | |
| shift 2 | |
| ;; | |
| --config) | |
| CONFIG="$2" | |
| shift 2 | |
| ;; | |
| --help-local) | |
| HELP_LOCAL=1 | |
| shift | |
| ;; | |
| --) | |
| shift | |
| break | |
| ;; | |
| *) | |
| break | |
| ;; | |
| esac | |
| done | |
| for arg in "$@"; do | |
| if [[ "$arg" == "--dry-run" ]]; then | |
| DRY_RUN=1 | |
| break | |
| fi | |
| done | |
| if [[ "$HELP_LOCAL" == "1" ]]; then | |
| usage | |
| exit 0 | |
| fi | |
| if [[ -z "$REPO" ]]; then | |
| REPO="$PWD" | |
| fi | |
| REPO="$(cd "$REPO" && pwd)" | |
| python_path() { | |
| if [[ "$PYTHON" == *.exe ]] && command -v wslpath >/dev/null 2>&1; then | |
| wslpath -w "$1" | |
| return | |
| fi | |
| if [[ "${OS:-}" == "Windows_NT" ]] && command -v cygpath >/dev/null 2>&1; then | |
| cygpath -w "$1" | |
| else | |
| printf '%s\n' "$1" | |
| fi | |
| } | |
| if [[ ! -x "$PYTHON" && "$DRY_RUN" != "1" ]]; then | |
| "$DG_ROOT/scripts/install_haystack_local.sh" >/tmp/dg-haystack-install.log | |
| fi | |
| if [[ -z "$CONFIG" ]]; then | |
| if [[ -s "$REPO/.dg-agent/haystack.dg.json" ]]; then | |
| CONFIG="$REPO/.dg-agent/haystack.dg.json" | |
| else | |
| CONFIG="$DG_ROOT/configs/client_profiles/haystack.dg.json" | |
| fi | |
| fi | |
| export OPENAI_BASE_URL="${OPENAI_BASE_URL:-http://127.0.0.1:8090/v1}" | |
| export OPENAI_API_KEY="${OPENAI_API_KEY:-dummy}" | |
| export HAYSTACK_MODEL="${HAYSTACK_MODEL:-diffusiongemma-local}" | |
| export DG_AGENT_CALLER_CWD="${DG_AGENT_CALLER_CWD:-$REPO}" | |
| REPO_FOR_PY="$(python_path "$REPO")" | |
| CONFIG_FOR_PY="$(python_path "$CONFIG")" | |
| RUNNER_FOR_PY="$(python_path "$DG_ROOT/scripts/dg_haystack_runner.py")" | |
| if [[ -x "$PYTHON" ]]; then | |
| exec "$PYTHON" "$RUNNER_FOR_PY" --repo "$REPO_FOR_PY" --config "$CONFIG_FOR_PY" "$@" | |
| fi | |
| exec python3 "$RUNNER_FOR_PY" --repo "$REPO_FOR_PY" --config "$CONFIG_FOR_PY" "$@" | |