LLM Foundations Applied: Optimizing Large Language Models for Apple M2 Ultra Consumer Hardware — Hayula Research

Hayula AI Lab

Abstract

This technical report bridges the gap between textbook large language model (LLM) theory and practical deployment on consumer-grade Apple Silicon hardware. We apply key results from Foundations of Large Language Models (Xiao & Zhu, 2025) to the specific constraints and capabilities of the Mac Studio M2 Ultra (192 GB unified memory, 76 GPU cores) running in the Hayula AI ecosystem. We present actionable guidance across five dimensions: (1) inference optimization — KV cache management, continuous

Files

File Description
paper.md Full paper (Markdown)
README.md This model card

Citation

@techreport{hayulalab2026llmfoundationsapplied,
    title={LLM Foundations Applied: Optimizing Large Language Models for Apple M2 Ultra Consumer Hardware — Hayula Research},
    author={Hayula AI Lab},
    year={2026},
    url={https://huggingface.co/hayulalab/llm-foundations-applied-paper}
}

hayulalab — Open Source AI Research

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support