Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
Upload anthropic_proxy.py with huggingface_hub
Browse files- anthropic_proxy.py +10 -1
anthropic_proxy.py
CHANGED
|
@@ -29,13 +29,20 @@ TIMEOUT = float(os.environ.get("VL_TIMEOUT", "1800"))
|
|
| 29 |
# "claude-" : il valide le nom avant d'emettre la requete. On expose donc des
|
| 30 |
# alias conformes, et on ignore le nom recu pour router vers l'unique modele
|
| 31 |
# reellement charge -- le client choisit une etiquette, pas un moteur.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 32 |
ALIASES = [
|
|
|
|
| 33 |
"claude-kimi-k3",
|
| 34 |
"claude-kimi-k3-linear",
|
| 35 |
"claude-qwen3-coder",
|
| 36 |
"claude-sonnet-4-5", # alias de compatibilite : certains clients
|
| 37 |
"claude-3-5-haiku", # codent en dur un modele "rapide" et un "lent"
|
| 38 |
]
|
|
|
|
| 39 |
|
| 40 |
app = FastAPI(title="anthropic-bridge")
|
| 41 |
_client = httpx.AsyncClient(base_url=UPSTREAM, timeout=TIMEOUT)
|
|
@@ -350,7 +357,9 @@ async def models(request: Request):
|
|
| 350 |
data = [{
|
| 351 |
"type": "model",
|
| 352 |
"id": a,
|
| 353 |
-
|
|
|
|
|
|
|
| 354 |
"created_at": "2026-01-01T00:00:00Z",
|
| 355 |
**({"context_window": ctx} if ctx else {}),
|
| 356 |
} for a in ALIASES]
|
|
|
|
| 29 |
# "claude-" : il valide le nom avant d'emettre la requete. On expose donc des
|
| 30 |
# alias conformes, et on ignore le nom recu pour router vers l'unique modele
|
| 31 |
# reellement charge -- le client choisit une etiquette, pas un moteur.
|
| 32 |
+
# Le PREMIER alias nomme le modele reellement charge : sans cela, un client qui
|
| 33 |
+
# voit "claude-kimi-k3" croit legitimement executer du Kimi alors que le moteur
|
| 34 |
+
# sert du Qwen. Les suivants sont des etiquettes de compatibilite, et tous
|
| 35 |
+
# routent vers l'unique modele charge.
|
| 36 |
+
_REAL = {"qwen": "claude-qwen3-coder-30b", "kimi": "claude-kimi-linear-48b"}
|
| 37 |
ALIASES = [
|
| 38 |
+
_REAL.get(MODEL, f"claude-{MODEL}"),
|
| 39 |
"claude-kimi-k3",
|
| 40 |
"claude-kimi-k3-linear",
|
| 41 |
"claude-qwen3-coder",
|
| 42 |
"claude-sonnet-4-5", # alias de compatibilite : certains clients
|
| 43 |
"claude-3-5-haiku", # codent en dur un modele "rapide" et un "lent"
|
| 44 |
]
|
| 45 |
+
ALIASES = list(dict.fromkeys(ALIASES)) # dedoublonne en gardant l'ordre
|
| 46 |
|
| 47 |
app = FastAPI(title="anthropic-bridge")
|
| 48 |
_client = httpx.AsyncClient(base_url=UPSTREAM, timeout=TIMEOUT)
|
|
|
|
| 357 |
data = [{
|
| 358 |
"type": "model",
|
| 359 |
"id": a,
|
| 360 |
+
# Le nom affiche porte le depot exact : c'est la seule facon pour un
|
| 361 |
+
# utilisateur de savoir quel modele repond derriere une etiquette.
|
| 362 |
+
"display_name": f"{a} -> {os.environ.get('VL_REAL_MODEL', MODEL)}",
|
| 363 |
"created_at": "2026-01-01T00:00:00Z",
|
| 364 |
**({"context_window": ctx} if ctx else {}),
|
| 365 |
} for a in ALIASES]
|