AI & ML interests

Reference implementations of LLM inference at the metal — Gemma 4 and Llama on CUDA, built from explicit C++23 components you can read and understand. Runs on consumer hardware.

Recent Activity

toddt  published a model 1 day ago
mila-llm/gpt2-small
toddt  updated a model 1 day ago
mila-llm/gpt2-small
toddt  updated a Space 4 days ago
mila-llm/README
View all activity

toddt 
published a model 1 day ago
toddt 
updated a Space 4 days ago
toddt 
published a Space 4 days ago