Reference implementations of LLM inference at the metal — Gemma 4 and Llama on CUDA, built from explicit C++23 components you can read and understand. Runs on consumer hardware.