Text Generation
GGUF
Japanese
English
llama.cpp
Mixture of Experts
expert-pruning
intel-mac
cpu
local-agent
imatrix
conversational
Instructions to use miutti/intel-mac-local-llm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use miutti/intel-mac-local-llm with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL # Run inference directly in the terminal: llama cli -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL # Run inference directly in the terminal: llama cli -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf miutti/intel-mac-local-llm:UD-Q2_K_XL # Run inference directly in the terminal: ./llama-cli -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf miutti/intel-mac-local-llm:UD-Q2_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Use Docker
docker model run hf.co/miutti/intel-mac-local-llm:UD-Q2_K_XL
- LM Studio
- Jan
- vLLM
How to use miutti/intel-mac-local-llm with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "miutti/intel-mac-local-llm" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "miutti/intel-mac-local-llm", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/miutti/intel-mac-local-llm:UD-Q2_K_XL
- Ollama
How to use miutti/intel-mac-local-llm with Ollama:
ollama run hf.co/miutti/intel-mac-local-llm:UD-Q2_K_XL
- Unsloth Desktop
- Pi
How to use miutti/intel-mac-local-llm with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "miutti/intel-mac-local-llm:UD-Q2_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use miutti/intel-mac-local-llm with Docker Model Runner:
docker model run hf.co/miutti/intel-mac-local-llm:UD-Q2_K_XL
- Lemonade
How to use miutti/intel-mac-local-llm with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull miutti/intel-mac-local-llm:UD-Q2_K_XL
Run and chat with the model
lemonade run user.intel-mac-local-llm-UD-Q2_K_XL
List all available models
lemonade list
- Hermes Agent
How to use miutti/intel-mac-local-llm with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default miutti/intel-mac-local-llm:UD-Q2_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use miutti/intel-mac-local-llm with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "miutti/intel-mac-local-llm:UD-Q2_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download source/kernel/bench.py from miutti/intel-mac-local-llm: direct link, hf CLI and curl.
- Browser
- Download file 6.52 kB
-
https://huggingface.co/miutti/intel-mac-local-llm/resolve/main/source/kernel/bench.py
- Command line
-
hf download hf://miutti/intel-mac-local-llm/source/kernel/bench.py
-
curl -L -o bench.py https://huggingface.co/miutti/intel-mac-local-llm/resolve/main/source/kernel/bench.py
6.52 kB
| #!/usr/bin/env python3 | |
| # -*- coding: utf-8 -*- | |
| """ | |
| bench.py -- 性能を測る | |
| 「賢くなった気がする」では直すところを間違える。 | |
| ・ちゃんと答えられた割合 | |
| ・かかった時間(はじめて / 2回目) | |
| ・ノートが当たった割合 | |
| を、まとめて出す。 | |
| """ | |
| import os, sys, io, time, json, contextlib | |
| HERE = os.path.dirname(os.path.abspath(__file__)) | |
| sys.path.insert(0, HERE) | |
| import kernel | |
| # 入力 → 期待する答えの一部(含まれていれば正解とみなす) | |
| CASES = [ | |
| ("デスクトップの写真を数えて", "5 個"), | |
| ("机の上の画像はいくつ?", "5 個"), | |
| ("デスクトップの動画は何個?", "1 個"), | |
| ("デスクトップのPDFを数えて", "1 個"), | |
| ("デスクトップには何がある?", "旅行1.jpg"), | |
| ("ダウンロードのPDFを数えて", "3 個"), | |
| ("デスクトップの去年の写真を数えて", "4 個"), | |
| ("デスクトップの書類を数えて", "2 個"), | |
| ("デスクトップの一覧を見せて", "会議.pdf"), | |
| ("ダウンロードには何がある?", "請求書.pdf"), | |
| ("デスクトップのテキストを数えて", "1 個"), | |
| ("机の上の去年の画像はいくつ", "4 個"), | |
| ("デスクトップの音楽を数えて", "0 個"), | |
| ("ダウンロードの画像を数えて", "1 個"), | |
| # --- ここから、実際のやりとりで出てきた言い方 --- | |
| ("机の上の去年のスナップ、片付けといて", "個を"), | |
| ("デスクトップに散らかってる写真を集めて", "個を"), | |
| ("さっき撮ったやつを見せて", "jpg"), | |
| ("デスクトップの一番大きいファイルは?", "mp4"), | |
| ("ダウンロードで一番古いのは何?", "古い.pdf"), | |
| ("デスクトップに同じファイルある?", "組"), | |
| ("デスクトップぜんぶで何メガ?", "MB"), | |
| ("デスクトップの画像は何個?", "5 個"), | |
| ("先月のダウンロードを数えて", "個"), | |
| ("デスクトップのjpgだけ見せて", "旅行1.jpg"), | |
| ("デスクトップの画像以外を数えて", "3 個"), | |
| # --- ここから、文法(述語と助詞)で読むもの --- | |
| ("ダウンロードから書類をデスクトップに移して", "3 個を"), | |
| ("デスクトップの写真を寄せといてくれる?", "個を"), | |
| ("デスクトップのテキストをしまっとく", "個を"), | |
| ("デスクトップの画像を並べてもらえる?", "旅行1.jpg"), | |
| ("デスクトップの画像を数えといて", "5 個"), | |
| ("机の上の去年の写真をまとめてほしい", "個を"), | |
| ("デスクトップの動画を数えてください", "1 個"), | |
| ("デスクトップのPDFを見せてよ", "会議.pdf"), | |
| ] | |
| import chat | |
| def once(text): | |
| """1回動かして (答え, ミリ秒, 記録) を返す | |
| 聞かれているのか命令なのかは、本番と同じ振り分けで決める。 | |
| ここを全部「聞かれている」にしていたので、 | |
| 「片付けといて」まで一覧が返って、不正解に数えていた | |
| """ | |
| # 動かす命令は練習用フォルダを書き換えるので、毎回まっさらに戻す。 | |
| # 戻していなかったため、前の命令が動かした結果を次が見てしまい、 | |
| # 2回目以降の正解率が下がっていた(ノートのせいに見えていた) | |
| with contextlib.redirect_stdout(io.StringIO()): | |
| kernel.make_demo() | |
| slots = kernel.draw_cards(text) | |
| kind = chat.classify(text, slots) | |
| ro = (kind != "命令") | |
| buf = io.StringIO() | |
| t0 = time.time() | |
| with contextlib.redirect_stdout(buf): | |
| try: | |
| ans = kernel.handle(text, readonly=ro, quiet=True) | |
| except Exception as e: | |
| ans = f"失敗: {e}" | |
| ms = (time.time() - t0) * 1000 | |
| return (ans or ""), ms, buf.getvalue() | |
| def note_uses(): | |
| """ノートが当たった回数。 | |
| 前はノートの「回数」の合計を前後で引き算していた。 | |
| だが手順を組み立て直すたびに 回数 が 1 に戻る作りだったので、 | |
| 合計が下がることがあり、当たった回数がマイナスになっていた | |
| (実測で -23/33 が出た)。 | |
| いまは kernel がその場で数えているので、それを読む。 | |
| """ | |
| try: | |
| return kernel.note_hit_total() | |
| except AttributeError: | |
| return 0 | |
| def run(label, reset_note): | |
| if reset_note: | |
| for f in ("notebook.json", "policy.json"): | |
| p = os.path.join(HERE, f) | |
| if os.path.exists(p): | |
| os.remove(p) | |
| before = note_uses() | |
| ok = 0 | |
| total = 0.0 | |
| rows = [] | |
| for text, want in CASES: | |
| ans, ms, _log = once(text) | |
| good = want in (ans or "") | |
| ok += good; total += ms | |
| rows.append((text, want, ans, ms, good, False)) | |
| hit = note_uses() - before | |
| n = len(CASES) | |
| print(f"\n■ {label}") | |
| print(f" 正解 : {ok}/{n} ({ok/n*100:.0f}%)") | |
| print(f" ノート当たり: {hit}/{n} ({hit/n*100:.0f}%)") | |
| print(f" 合計時間 : {total:.0f} ミリ秒 (1件あたり {total/n:.1f})") | |
| for text, want, ans, ms, good, _h in rows: | |
| if not good: | |
| print(f" ✗ {text}") | |
| print(f" ほしい「{want}」/ 出た「{str(ans)[:60]}」") | |
| return ok, n, total, hit | |
| if __name__ == "__main__": | |
| # KERNEL_THINK=1 で「じっくり(何人かで考える)」を測る | |
| if os.environ.get("KERNEL_THINK") == "1": | |
| kernel.THINK_HARD = True | |
| print("※ じっくりモード(何人かで考えて多数決)で測ります") | |
| a = run("1回目(ノート空っぽ)", reset_note=True) | |
| b = run("2回目(ノートが溜まった状態)", reset_note=False) | |
| c = run("3回目", reset_note=False) | |
| print("\n■ まとめ") | |
| print(f" 正解 : {a[0]}/{a[1]} → {b[0]}/{b[1]} → {c[0]}/{c[1]}") | |
| print(f" 合計時間: {a[2]:.0f} → {b[2]:.0f} → {c[2]:.0f} ミリ秒") | |
| print(f" ノート : {a[3]} → {b[3]} → {c[3]} 件が当たった") | |