Instructions to use hipinis/20260718 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use hipinis/20260718 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf hipinis/20260718:Q8_0 # Run inference directly in the terminal: llama cli -hf hipinis/20260718:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf hipinis/20260718:Q8_0 # Run inference directly in the terminal: llama cli -hf hipinis/20260718:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf hipinis/20260718:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf hipinis/20260718:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf hipinis/20260718:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf hipinis/20260718:Q8_0
Use Docker
docker model run hf.co/hipinis/20260718:Q8_0
- LM Studio
- Jan
- Ollama
How to use hipinis/20260718 with Ollama:
ollama run hf.co/hipinis/20260718:Q8_0
- Unsloth Studio
How to use hipinis/20260718 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for hipinis/20260718 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for hipinis/20260718 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for hipinis/20260718 to start chatting
- Pi
How to use hipinis/20260718 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf hipinis/20260718:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "hipinis/20260718:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use hipinis/20260718 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf hipinis/20260718:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "hipinis/20260718:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use hipinis/20260718 with Docker Model Runner:
docker model run hf.co/hipinis/20260718:Q8_0
- Lemonade
How to use hipinis/20260718 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull hipinis/20260718:Q8_0
Run and chat with the model
lemonade run user.20260718-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use hipinis/20260718 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf hipinis/20260718:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default hipinis/20260718:Q8_0
Run Hermes
hermes
- Atomic Chat
| from __future__ import annotations | |
| from concurrent.futures import FIRST_COMPLETED, ThreadPoolExecutor, wait | |
| from typing import Callable, List, Optional | |
| import comfy.model_management | |
| ResultType = dict | |
| TaskType = tuple | |
| class BatchGenerationRunner: | |
| """统一的批次任务调度器,负责线程池并发、进度条与日志回调。""" | |
| def __init__( | |
| self, | |
| logger, | |
| ensure_not_interrupted: Callable[[], None], | |
| progress_bar_factory: Callable[[int], object], | |
| ): | |
| self.logger = logger | |
| self.ensure_not_interrupted = ensure_not_interrupted | |
| self.progress_bar_factory = progress_bar_factory | |
| def run( | |
| self, | |
| tasks: List[TaskType], | |
| worker_fn: Callable[[TaskType], ResultType], | |
| batch_size: int, | |
| actual_workers: int, | |
| continue_on_error: bool, | |
| progress_callback: Callable[[ResultType, int, int, object], None], | |
| ) -> List[ResultType]: | |
| """通过线程池或串行方式执行任务,并在每个结果返回时调用 progress_callback。""" | |
| if batch_size <= 0: | |
| return [] | |
| progress_bar = self.progress_bar_factory(batch_size) | |
| self.ensure_not_interrupted() | |
| if actual_workers > 1 and batch_size > 1: | |
| return self._run_parallel( | |
| tasks, | |
| worker_fn, | |
| batch_size, | |
| actual_workers, | |
| continue_on_error, | |
| progress_callback, | |
| progress_bar, | |
| ) | |
| return self._run_sequential( | |
| tasks, | |
| worker_fn, | |
| batch_size, | |
| continue_on_error, | |
| progress_callback, | |
| progress_bar, | |
| ) | |
| def _run_parallel( | |
| self, | |
| tasks: List[TaskType], | |
| worker_fn: Callable[[TaskType], ResultType], | |
| batch_size: int, | |
| actual_workers: int, | |
| continue_on_error: bool, | |
| progress_callback: Callable[[ResultType, int, int, object], None], | |
| progress_bar: object, | |
| ) -> List[ResultType]: | |
| results: List[ResultType] = [] | |
| completed = 0 | |
| executor = ThreadPoolExecutor(max_workers=actual_workers) | |
| should_stop = False | |
| try: | |
| future_to_task = { | |
| executor.submit(worker_fn, task): task | |
| for task in tasks | |
| } | |
| pending = set(future_to_task.keys()) | |
| while pending: | |
| done, pending = wait( | |
| pending, | |
| timeout=0.1, | |
| return_when=FIRST_COMPLETED | |
| ) | |
| if not done: | |
| continue | |
| for future in done: | |
| task = future_to_task.pop(future, None) | |
| try: | |
| self.ensure_not_interrupted() | |
| result = future.result() | |
| except comfy.model_management.InterruptProcessingException: | |
| for future_ref in list(future_to_task.keys()): | |
| future_ref.cancel() | |
| raise | |
| except Exception as exc: # pragma: no cover - worker 应返回统一结构 | |
| self.logger.error(f"批次任务异常: {exc}") | |
| result = {"success": False, "index": -1, "error": str(exc)} | |
| results.append(result) | |
| completed += 1 | |
| progress_callback(result, completed, batch_size, progress_bar) | |
| if not continue_on_error and not result.get("success"): | |
| should_stop = True | |
| break | |
| if should_stop: | |
| for future_ref in pending: | |
| future_ref.cancel() | |
| break | |
| finally: | |
| executor.shutdown(wait=False, cancel_futures=True) | |
| return results | |
| def _run_sequential( | |
| self, | |
| tasks: List[TaskType], | |
| worker_fn: Callable[[TaskType], ResultType], | |
| batch_size: int, | |
| continue_on_error: bool, | |
| progress_callback: Callable[[ResultType, int, int, object], None], | |
| progress_bar: object, | |
| ) -> List[ResultType]: | |
| results: List[ResultType] = [] | |
| completed = 0 | |
| for task in tasks: | |
| self.ensure_not_interrupted() | |
| result = worker_fn(task) | |
| results.append(result) | |
| completed += 1 | |
| progress_callback(result, completed, batch_size, progress_bar) | |
| if not continue_on_error and not result.get("success"): | |
| break | |
| return results | |