LLM Foundations Applied: Optimizing Large Language Models for Apple M2 Ultra Consumer Hardware — Hayula Research
Hayula AI Lab
Abstract
This technical report bridges the gap between textbook large language model (LLM) theory and practical deployment on consumer-grade Apple Silicon hardware. We apply key results from Foundations of Large Language Models (Xiao & Zhu, 2025) to the specific constraints and capabilities of the Mac Studio M2 Ultra (192 GB unified memory, 76 GPU cores) running in the Hayula AI ecosystem. We present actionable guidance across five dimensions: (1) inference optimization — KV cache management, continuous
Files
| File | Description |
|---|---|
paper.md |
Full paper (Markdown) |
README.md |
This model card |
Citation
@techreport{hayulalab2026llmfoundationsapplied,
title={LLM Foundations Applied: Optimizing Large Language Models for Apple M2 Ultra Consumer Hardware — Hayula Research},
author={Hayula AI Lab},
year={2026},
url={https://huggingface.co/hayulalab/llm-foundations-applied-paper}
}
hayulalab — Open Source AI Research
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support