AI & ML interests

Reference implementations of LLM inference at the metal โ€” Gemma 4 and Llama on CUDA, built from explicit C++23 components you can read and understand. Runs on consumer hardware.

Recent Activity

toddtย  published a model 1 day ago
mila-llm/gpt2-small
toddtย  updated a model 1 day ago
mila-llm/gpt2-small
toddtย  updated a Space 4 days ago
mila-llm/README
View all activity

mila-llm 's datasets

None public yet