AI & ML interests

On-device AI, GGUF quantization, Apple Silicon, macOS automation

Recent Activity

hero775Β  updated a model 5 days ago
batiai/LFM2.5-8B-A1B-GGUF
hero775Β  updated a model 5 days ago
batiai/gemma-4-12B-it-GGUF
View all activity

batiai 's collections 8

πŸ”₯ Qwen 3.8 β€” Newest 27B, Korean-verified
Qwen's newest 27B: vision, 262K context, Korean-verified. Start with IQ4_XS β€” we measured every tier and Q3_K_M is smaller yet slower. Needs 24GB+.
🧠 NVIDIA Nemotron 3 β€” Hybrid Mamba+Attention
NVIDIA Nemotron 3 family β€” NemotronH architecture combining Mamba state-space + standard attention. Mac-runnable, BatiAI-quantized + signed.
πŸŽ™οΈ ν•œκ΅­μ–΄ μŒμ„± μŠ€μœ„νŠΈ β€” STT + ν™”μžλΆ„λ¦¬
batisay(STT, 무엇을 λ§ν–ˆλ‚˜) + batispeak(ν™”μžλΆ„λ¦¬, λˆ„κ°€ λ§ν–ˆλ‚˜) = ν†΅ν™”Β·νšŒμ˜ ν™”μžλ³„ 전사. 16GB Mac on-device.
🍎 Gemma 4 β€” Google's Latest
Gemma 4 quantizations from Google's official weights. Best entry for 16GB Mac mini M4 (E4B Q4 = 57 t/s).
BatiAI RAG Stack
Complete Mac-first on-device RAG stack β€” chat LLM + reranker + text/VL embedder, direct from BF16, BatiAI-signed. For BatiFlow.
πŸ”₯ Qwen 3.8 β€” Newest 27B, Korean-verified
Qwen's newest 27B: vision, 262K context, Korean-verified. Start with IQ4_XS β€” we measured every tier and Q3_K_M is smaller yet slower. Needs 24GB+.
πŸŽ™οΈ ν•œκ΅­μ–΄ μŒμ„± μŠ€μœ„νŠΈ β€” STT + ν™”μžλΆ„λ¦¬
batisay(STT, 무엇을 λ§ν–ˆλ‚˜) + batispeak(ν™”μžλΆ„λ¦¬, λˆ„κ°€ λ§ν–ˆλ‚˜) = ν†΅ν™”Β·νšŒμ˜ ν™”μžλ³„ 전사. 16GB Mac on-device.
🧠 NVIDIA Nemotron 3 β€” Hybrid Mamba+Attention
NVIDIA Nemotron 3 family β€” NemotronH architecture combining Mamba state-space + standard attention. Mac-runnable, BatiAI-quantized + signed.
🍎 Gemma 4 β€” Google's Latest
Gemma 4 quantizations from Google's official weights. Best entry for 16GB Mac mini M4 (E4B Q4 = 57 t/s).
BatiAI RAG Stack
Complete Mac-first on-device RAG stack β€” chat LLM + reranker + text/VL embedder, direct from BF16, BatiAI-signed. For BatiFlow.