On-device inference · GGUF + MLX
2.69×

Big-model answers, faster on the CPU you already own.

POCKET runs a 35B-class open model on a phone, a laptop, or a GPU-less server — and decodes 2.69× faster than the most-downloaded on-device model, at matched quality.

Same laptop. Same prompt. No accelerator. POCKET returns tokens while the baseline is still warming up.

measured · CPU decode throughput
vs. the most-downloaded on-device model
Apple M3, single request

2.22×
GPU decode
Throughput on a single consumer GPU, same quality target.
≈ 1.0×
Quality
Matched on standard evals — parity, not a downgrade.
0
GPUs required
Runs on CPU. A phone or an air-gapped box is enough.
2
Formats
GGUF for llama.cpp, MLX for Apple silicon. No custom fork.
Measured, not marketed

The full scorecard — including where we trail.

Four axes against the most-downloaded on-device model, same quality target. We publish the axis we lose, because a benchmark you can't check is just a claim.

CPU decode throughputtokens per second, single request
2.69×
GPU decode throughputsingle consumer GPU
2.22×
Answer qualitystandard reasoning + knowledge evals
matched
Prefill on long promptstime-to-first-token, 8k context
0.41×

Reading it straight: POCKET wins decode by a wide margin and holds quality, but the baseline ingests very long prompts faster — POCKET's prefill trails at 0.41×. For chat, agents and short-to-medium prompts, decode is what you feel; for 8k-token document dumps, weigh the trade. Full harness and configs are published.

On-device by design

One model. Three places it has no business running — and does.

01A phone in your hand

The model lives on the device. Prompts and answers never leave it — no round-trip, no cloud bill, no signal required.

02A laptop with no GPU

Apple silicon via MLX, everything else via GGUF. A MacBook Air is a workstation; POCKET decodes 2.69× faster than the on-device model most people already downloaded.

03An air-gapped server

Closed networks, no accelerators, no outbound. The model runs inside the perimeter, which is exactly where regulated work has to stay.

For the public sector

Sovereign by default.

No cloud dependency, no GPU procurement, no data leaving the building. POCKET ships as software you license or as a sealed appliance — built for closed networks, defense, healthcare, and administrative systems that cloud AI cannot enter. Built on Google’s open Gemma, tuned for Korean, distributed as GGUF and MLX.

Big-model answers. Pocket-sized footprint.

Put a frontier-class model where the work is — on the device, inside the network, off the cloud.