build: use prebuilt llama.cpp CPU binary (b9895) instead of source compile — fixes mtmd OOM hang, ~20x faster build ea5d56a verified Leon4gr45 commited on Jul 7
build: pull model at container runtime instead of baking into image (faster/reliable build) 7cc60cb verified Leon4gr45 commited on Jul 7
v2.0: single-instance CPU proxy, native OpenAI tools+streaming, flash-attn+q4_0 KV, all-cores e24791f verified Leon4gr45 commited on Jul 7