Text Generation
GGUF
Japanese
English
llama.cpp
Mixture of Experts
expert-pruning
intel-mac
cpu
local-agent
imatrix
conversational
Instructions to use miutti/intel-mac-local-llm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use miutti/intel-mac-local-llm with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL # Run inference directly in the terminal: llama cli -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL # Run inference directly in the terminal: llama cli -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf miutti/intel-mac-local-llm:UD-Q2_K_XL # Run inference directly in the terminal: ./llama-cli -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf miutti/intel-mac-local-llm:UD-Q2_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Use Docker
docker model run hf.co/miutti/intel-mac-local-llm:UD-Q2_K_XL
- LM Studio
- Jan
- vLLM
How to use miutti/intel-mac-local-llm with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "miutti/intel-mac-local-llm" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "miutti/intel-mac-local-llm", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/miutti/intel-mac-local-llm:UD-Q2_K_XL
- Ollama
How to use miutti/intel-mac-local-llm with Ollama:
ollama run hf.co/miutti/intel-mac-local-llm:UD-Q2_K_XL
- Unsloth Desktop
- Pi
How to use miutti/intel-mac-local-llm with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "miutti/intel-mac-local-llm:UD-Q2_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use miutti/intel-mac-local-llm with Docker Model Runner:
docker model run hf.co/miutti/intel-mac-local-llm:UD-Q2_K_XL
- Lemonade
How to use miutti/intel-mac-local-llm with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull miutti/intel-mac-local-llm:UD-Q2_K_XL
Run and chat with the model
lemonade run user.intel-mac-local-llm-UD-Q2_K_XL
List all available models
lemonade list
- Hermes Agent
How to use miutti/intel-mac-local-llm with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default miutti/intel-mac-local-llm:UD-Q2_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use miutti/intel-mac-local-llm with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "miutti/intel-mac-local-llm:UD-Q2_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download source/kernel/bench_speed.py from miutti/intel-mac-local-llm: direct link, hf CLI and curl.
- Browser
- Download file 5.85 kB
-
https://huggingface.co/miutti/intel-mac-local-llm/resolve/main/source/kernel/bench_speed.py
- Command line
-
hf download hf://miutti/intel-mac-local-llm/source/kernel/bench_speed.py
-
curl -L -o bench_speed.py https://huggingface.co/miutti/intel-mac-local-llm/resolve/main/source/kernel/bench_speed.py
5.85 kB
| #!/usr/bin/env python3 | |
| # -*- coding: utf-8 -*- | |
| """同じ条件で llama-server の速度を測る小さな再利用可能ハーネス。""" | |
| from __future__ import annotations | |
| import argparse | |
| import datetime as dt | |
| import json | |
| import statistics | |
| import sys | |
| import time | |
| from pathlib import Path | |
| sys.path.insert(0, str(Path(__file__).resolve().parent)) | |
| from public_bench import ( # noqa: E402 | |
| HARNESS_VERSION, | |
| LocalClient, | |
| atomic_write, | |
| collect_runtime_metadata, | |
| collect_server_metadata, | |
| redact_url, | |
| ) | |
| PROMPTS = [ | |
| "Return only the next prime number after 100.", | |
| "Write one short sentence explaining why the sky appears blue.", | |
| "Compute 37 * 29. Return only the integer.", | |
| "Return a Python function named square(x) that returns x*x, with no explanation.", | |
| ] | |
| def percentile(values: list[float], q: float) -> float | None: | |
| if not values: | |
| return None | |
| xs = sorted(values) | |
| index = (len(xs) - 1) * q | |
| low = int(index) | |
| high = min(low + 1, len(xs) - 1) | |
| return xs[low] + (xs[high] - xs[low]) * (index - low) | |
| def write_json(path: Path, obj: dict) -> None: | |
| atomic_write(path, json.dumps(obj, ensure_ascii=False, indent=2) + "\n") | |
| def main() -> int: | |
| ap = argparse.ArgumentParser() | |
| ap.add_argument("--url", default="http://127.0.0.1:8080/v1/chat/completions") | |
| ap.add_argument("--out-dir", default="") | |
| ap.add_argument("--repeat", type=int, default=1) | |
| ap.add_argument("--warmup", type=int, default=1) | |
| ap.add_argument("--max-tokens", type=int, default=128) | |
| ap.add_argument("--timeout", type=int, default=900) | |
| ap.add_argument("--retries", type=int, default=2) | |
| ap.add_argument("--retry-backoff", type=float, default=0.25) | |
| ap.add_argument("--seed", type=int, default=20260909) | |
| ap.add_argument("--think", action="store_true") | |
| args = ap.parse_args() | |
| if args.repeat <= 0 or args.warmup < 0: | |
| ap.error("--repeat は正数、--warmup は0以上が必要です") | |
| stamp = dt.datetime.now().strftime("%Y%m%d-%H%M%S") | |
| out = Path(args.out_dir or Path(__file__).resolve().parent / "bench_results" / | |
| f"speed-{stamp}") | |
| out.mkdir(parents=True, exist_ok=True) | |
| client = LocalClient(args.url, args.think, args.max_tokens, args.timeout, | |
| retries=args.retries, retry_backoff=args.retry_backoff, | |
| seed=args.seed) | |
| for i in range(args.warmup): | |
| result = client.call(PROMPTS[i % len(PROMPTS)], f"warmup/{i}") | |
| if result.get("error"): | |
| print(f"warmup error: {result['error']}", file=sys.stderr) | |
| rows = [] | |
| started = time.monotonic() | |
| for repeat in range(args.repeat): | |
| for prompt_index, prompt in enumerate(PROMPTS): | |
| rid = f"p{prompt_index}/r{repeat}" | |
| result = client.call(prompt, rid) | |
| rows.append({ | |
| "id": rid, | |
| "prompt_index": prompt_index, | |
| "repeat": repeat, | |
| "elapsed_sec": result.get("elapsed_sec", 0), | |
| "timings": result.get("timings") or {}, | |
| "usage": result.get("usage"), | |
| "model": result.get("model"), | |
| "error": result.get("error"), | |
| "error_type": result.get("error_type"), | |
| "http_status": result.get("http_status"), | |
| "attempts": result.get("attempts") or [], | |
| "retry_count": result.get("retry_count", 0), | |
| }) | |
| print(f" {rid}: {rows[-1]['elapsed_sec']:.3f}s " | |
| f"{rows[-1]['error'] or 'ok'}", flush=True) | |
| successful = [r for r in rows if not r.get("error")] | |
| elapsed = [float(r.get("elapsed_sec", 0)) for r in successful] | |
| tps = [float((r.get("timings") or {}).get("predicted_per_second", 0) or 0) | |
| for r in successful if (r.get("timings") or {}).get("predicted_per_second")] | |
| prompt_tps = [float((r.get("timings") or {}).get("prompt_per_second", 0) or 0) | |
| for r in successful if (r.get("timings") or {}).get("prompt_per_second")] | |
| summary = { | |
| "harness_version": HARNESS_VERSION, | |
| "created_at": dt.datetime.now(dt.timezone.utc).isoformat(), | |
| "url": redact_url(args.url), | |
| "repeat": args.repeat, | |
| "warmup": args.warmup, | |
| "n": len(rows), | |
| "successful_n": len(successful), | |
| "failed_n": len(rows) - len(successful), | |
| "elapsed_total_sec": time.monotonic() - started, | |
| "mean_sec": statistics.mean(elapsed) if elapsed else None, | |
| "median_sec": statistics.median(elapsed) if elapsed else None, | |
| "p95_sec": percentile(elapsed, 0.95), | |
| "mean_predicted_tps": statistics.mean(tps) if tps else None, | |
| "median_predicted_tps": statistics.median(tps) if tps else None, | |
| "mean_prompt_tps": statistics.mean(prompt_tps) if prompt_tps else None, | |
| "error_types": { | |
| str(k): sum(1 for r in rows if (r.get("error_type") or "error") == k) | |
| for k in sorted({r.get("error_type") or "error" for r in rows | |
| if r.get("error")}) | |
| }, | |
| "runtime": collect_runtime_metadata(), | |
| "server": collect_server_metadata(args.url), | |
| } | |
| jsonl = out / "speed.jsonl" | |
| atomic_write(jsonl, "".join(json.dumps(r, ensure_ascii=False) + "\n" for r in rows)) | |
| write_json(out / "speed.summary.json", summary) | |
| write_json(out / "config.json", { | |
| "harness_version": HARNESS_VERSION, | |
| "argv": sys.argv, | |
| "url": redact_url(args.url), | |
| "max_tokens": args.max_tokens, | |
| "think": args.think, | |
| "seed": args.seed, | |
| "repeat": args.repeat, | |
| "warmup": args.warmup, | |
| }) | |
| print(json.dumps(summary, ensure_ascii=False, indent=2)) | |
| print(f"結果: {out}") | |
| return 0 | |
| if __name__ == "__main__": | |
| raise SystemExit(main()) | |