AI & ML interests
Hardware- and workload-adaptive LLM inference, efficient local AI, persistent inference state, and tools for long-running AI workflows. Building Imparo and Zeraix.
Recent Activity
Zeraix
LLM inference that fits your hardware. AI tools built around your work.
We build open-source software to make LLMs more useful on the hardware people already own — from the inference runtime to the everyday workspace.
Website · GitHub · X / Twitter
Imparo — the inference engine
Imparo is a hardware- and workload-adaptive LLM inference engine built in Rust.
- Less execution overhead: measured tuning and workload-aware execution, with Metal megakernel decode on supported paths.
- Reusable context: matching prefixes and shared paged state reduce repeated prefill work across conversations and agents.
- Persistent state: retained KV and recurrent state support resuming and branching long conversations and tool loops.
- Measured improvements: performance work is paired with output and state-transition checks.
Apple Silicon with Metal is the primary development and validation path. Other backend/model combinations remain qualification-specific. Context reuse does not imply concurrent decoding or continuous batching.
Source & Quick Start · Performance · Contribute
Zeraix — the AI workspace
Zeraix brings models, conversations, agents, files, and tools into an open-source desktop workspace, with local inference and optional cloud providers.
The desktop application and Imparo are separate projects. The current desktop release uses its documented llama.cpp-based runtimes; Imparo is not its default bundled engine.
Download Zeraix · Source & Docs
On Hugging Face
| Repository | Purpose |
|---|---|
| imparo-benchmarks | Published inference benchmark results and measurement notes. |
| llama-builds | Runtime and sandbox assets consumed by the Zeraix application. |
| llama-seeds | Version-matched prefix-cache seeds and auxiliary drafter assets for Zeraix. |
The two application asset repositories keep their existing download paths. They are not general-purpose standalone chat models or Imparo source mirrors.
Recent updates
- 2026-09-10: Imparo extended validation to Apple M4 Pro, reproducing optimization benefits demonstrated on M3 Pro. The published four-engine numerical table currently covers M3 Pro.
- 2026-09-09: Zeraix 2.0 released.
- 2026-09-07: Imparo published Metal megakernel decode and interleaved M3 Pro comparisons with llama.cpp, oMLX, and rapid-mlx.
Build with us
We welcome contributions to model support, kernels, hardware validation, state management, benchmarks, integrations, and documentation. For performance changes, include a reproducible baseline and correctness checks.
Imparo issues · Zeraix issues · Contact
Project source licenses and third-party artifact licenses are documented in their respective repositories. Model weights and upstream runtime components retain their own terms.