Instructions to use webai-community/ai-models with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use webai-community/ai-models with llama-cpp-python:

# !pip install llama-cpp-python

from llama_cpp import Llama

llm = Llama.from_pretrained(
	repo_id="webai-community/ai-models",
	filename="CodeLlama-7b-Instruct-hf/gguf/codellama-7b-instruct.Q4_K_M.gguf",
)

output = llm(
	"Once upon a time,",
	max_tokens=512,
	echo=True
)
print(output)

Notebooks
Google Colab
Kaggle
Local Apps Settings

llama.cpp

How to use webai-community/ai-models with llama.cpp:

Install from brew

brew install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama-server -hf webai-community/ai-models:Q4_K_M
# Run inference directly in the terminal:
llama-cli -hf webai-community/ai-models:Q4_K_M

Install from WinGet (Windows)

winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama-server -hf webai-community/ai-models:Q4_K_M
# Run inference directly in the terminal:
llama-cli -hf webai-community/ai-models:Q4_K_M

Use pre-built binary

# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf webai-community/ai-models:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf webai-community/ai-models:Q4_K_M

Build from source code

git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf webai-community/ai-models:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf webai-community/ai-models:Q4_K_M

Use Docker

docker model run hf.co/webai-community/ai-models:Q4_K_M

LM Studio
Jan
Ollama
How to use webai-community/ai-models with Ollama:
```
ollama run hf.co/webai-community/ai-models:Q4_K_M
```

Unsloth Studio

How to use webai-community/ai-models with Unsloth Studio:

Install Unsloth Studio (macOS, Linux, WSL)

curl -fsSL https://unsloth.ai/install.sh | sh
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for webai-community/ai-models to start chatting

Install Unsloth Studio (Windows)

irm https://unsloth.ai/install.ps1 | iex
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for webai-community/ai-models to start chatting

Using HuggingFace Spaces for Unsloth

# No setup required
# Open https://huggingface.co/spaces/unsloth/studio in your browser
# Search for webai-community/ai-models to start chatting

Docker Model Runner
How to use webai-community/ai-models with Docker Model Runner:
```
docker model run hf.co/webai-community/ai-models:Q4_K_M
```

Lemonade

How to use webai-community/ai-models with Lemonade:

Pull the model

# Download Lemonade from https://lemonade-server.ai/
lemonade pull webai-community/ai-models:Q4_K_M

Run and chat with the model

lemonade run user.ai-models-Q4_K_M

List all available models

lemonade list

qjia7 commited on 23 days ago

Commit

a615462

verified ·

1 Parent(s): a118411

Add muffin ort-webgpu model

Browse files

Files changed (10) hide show

muffin/onnx-webgpu/adapter_cache.bin +0 -0
muffin/onnx-webgpu/chat_template.jinja +73 -0
muffin/onnx-webgpu/edge_on_device_model_execution_config.pb +3 -0
muffin/onnx-webgpu/encoder_cache.bin +0 -0
muffin/onnx-webgpu/genai_config.json +61 -0
muffin/onnx-webgpu/manifest.json +14 -0
muffin/onnx-webgpu/model.onnx +3 -0
muffin/onnx-webgpu/model.onnx.data +3 -0
muffin/onnx-webgpu/tokenizer.json +3 -0
muffin/onnx-webgpu/tokenizer_config.json +12 -0

muffin/onnx-webgpu/adapter_cache.bin ADDED Viewed

File without changes

muffin/onnx-webgpu/chat_template.jinja ADDED Viewed

	@@ -0,0 +1,73 @@

+{%- if tools %}
+{{- '<|system|>' }}
+{%- if messages and messages[0].role == 'system' %}
+{{- messages[0].content + '\n\n' }}
+{%- else %}
+{{- 'You are Muffin, a large language model trained by Microsoft.\n\n' }}
+{%- endif %}
+{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nThe available tools are provided as JSON between <|tool|> and <|/tool|> tokens:\n" }}
+{{- '<|tool|>' }}
+{{- tools | tojson }}
+{{- '<|/tool|>\n\nFor each function call, respond with a single <|tool_call|>{\"name\": \"<function-name>\", \"arguments\": <json-args>}<|/tool_call|> block.\n' }}
+<|end|>
+{%- set start_index = 1 if messages and messages[0].role == 'system' else 0 %}
+{%- else %}
+{%- if messages and messages[0].role == 'system' %}
+{{- '<|system|>' + messages[0].content + '<|end|>' }}
+{%- set start_index = 1 %}
+{%- else %}
+{{- '<|system|>You are Muffin, a large language model trained by Microsoft.<|end|>' }}
+{%- set start_index = 0 %}
+{%- endif %}
+{%- endif %}
+{%- for message in messages[start_index:] %}
+{%- if message.content is string %}
+{%- set content = message.content %}
+{%- else %}
+{%- set content = '' %}
+{%- endif %}
+{%- if message.role == 'user' %}
+{{- '<|user|>' + content + '<|end|>' }}
+{%- elif message.role == 'assistant' %}
+{{- '<|assistant|>' }}
+{{- content }}
+{%- if message.tool_calls %}
+{%- for tool_call in message.tool_calls %}
+{%- if tool_call.function %}
+{%- set tc = tool_call.function %}
+{%- else %}
+{%- set tc = tool_call %}
+{%- endif %}
+{{- '<|tool_call|>' }}
+{{- '{\"name\": \"' ~ tc.name ~ '\", \"arguments\": ' }}
+{%- if tc.arguments is string %}
+{{- tc.arguments }}
+{%- else %}
+{{- tc.arguments | tojson }}
+{%- endif %}
+{{- '}' }}
+{{- '<|/tool_call|>' }}
+{%- endfor %}
+{%- endif %}
+{%- if loop.last and not add_generation_prompt %}
+{{- '<|endofprompt|>' }}
+{%- else %}
+{{- '<|end|>' }}
+{%- endif %}
+{%- elif message.role == 'tool' %}
+{%- if loop.first or (messages[loop.index0 + start_index - 1].role != 'tool') %}
+{{- '<|user|>' }}
+{%- endif %}
+{{- '<|tool_response|>' }}
+{{- content }}
+{%- if loop.last or (messages[loop.index0 + start_index + 1].role != 'tool') %}
+<|end|>
+{%- endif %}
+{%- elif message.role == 'system' and (start_index + loop.index0) != 0 %}
+{{- '<|system|>' + content + '<|end|>' }}
+{%- endif %}
+{%- endfor %}
+{%- if add_generation_prompt %}
+{{- '<|assistant|>' }}
+{%- else %}
+{%- endif %}

muffin/onnx-webgpu/edge_on_device_model_execution_config.pb ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:d6174e6bddd62fc999e8578ef48a449b30976076088613f118091ce303ecf0b5
+size 8812

muffin/onnx-webgpu/encoder_cache.bin ADDED Viewed

File without changes

muffin/onnx-webgpu/genai_config.json ADDED Viewed

	@@ -0,0 +1,61 @@

+{
+  "model": {
+    "bos_token_id": 1,
+    "context_length": 8192,
+    "decoder": {
+      "session_options": {
+        "log_id": "onnxruntime-genai",
+        "provider_options": [
+          {
+            "webgpu": {
+              "enableGraphCapture": "1",
+              "validationMode": "disabled"
+            }
+          }
+        ]
+      },
+      "filename": "model.onnx",
+      "head_size": 128,
+      "hidden_size": 2048,
+      "inputs": {
+        "input_ids": "input_ids",
+        "attention_mask": "attention_mask",
+        "past_key_names": "past_key_values.%d.key",
+        "past_value_names": "past_key_values.%d.value"
+      },
+      "outputs": {
+        "logits": "logits",
+        "present_key_names": "present.%d.key",
+        "present_value_names": "present.%d.value"
+      },
+      "num_attention_heads": 24,
+      "num_hidden_layers": 28,
+      "num_key_value_heads": 8
+    },
+    "eos_token_id": [
+      200018,
+      200018,
+      200020,
+      200019
+    ],
+    "pad_token_id": 199999,
+    "type": "qwen3",
+    "vocab_size": 200029
+  },
+  "search": {
+    "diversity_penalty": 0.0,
+    "do_sample": true,
+    "early_stopping": true,
+    "length_penalty": 1.0,
+    "max_length": 8192,
+    "min_length": 0,
+    "no_repeat_ngram_size": 0,
+    "num_beams": 1,
+    "num_return_sequences": 1,
+    "past_present_share_buffer": true,
+    "repetition_penalty": 1.2,
+    "temperature": 0.6,
+    "top_k": 20,
+    "top_p": 0.95
+  }
+}

muffin/onnx-webgpu/manifest.json ADDED Viewed

	@@ -0,0 +1,14 @@

+{
+    "manifest_version":  2,
+    "name":  "Edge On Device Model",
+    "version":  "2026.5.8.1",
+    "BaseModelSpec":  {
+                          "supported_performance_hints":  [
+                                                              2,
+                                                              1
+                                                          ],
+                          "name":  "Muffin",
+                          "version":  "2026.5.8.1",
+                          "type":  "webgpu"
+                      }
+}

muffin/onnx-webgpu/model.onnx ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:9adf24cf528a984dd5d50d8496ad90054c6422da87f119135321c5c88502bfdf
+size 353943

muffin/onnx-webgpu/model.onnx.data ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:dca047f823e188347b2bedf11e5b6b1f13eef88c8a8f3c80f900df06f5ab8a88
+size 1494558336

muffin/onnx-webgpu/tokenizer.json ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:7ea8bdf68c3e7549a3fb4342523288ce628f6ab56a618f9a4dfb234a0b4d46a8
+size 15524476

muffin/onnx-webgpu/tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,12 @@

+{
+  "add_prefix_space": false,
+  "backend": "tokenizers",
+  "bos_token": null,
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "<|endofprompt|>",
+  "is_local": true,
+  "model_max_length": 8192,
+  "pad_token": "<|endoftext|>",
+  "tokenizer_class": "TokenizersBackend",
+  "unk_token": null
+}