feat: self-hosted Qwen2.5-1.5B-Instruct via transformers β no external API, no compilation deea70e imtrt004 commited on Feb 27
feat: replace llama-cpp-python/Groq with free HF InferenceClient (zero compilation) 98e3f05 imtrt004 commited on Feb 27
fix: restore build-essential+cmake, pin llama-cpp-python==0.3.16 for stable layer cache bfaa120 imtrt004 commited on Feb 27
fix: use pre-built llama-cpp-python CPU wheel β eliminates 8min C++ compile 6e6147b imtrt004 commited on Feb 27
fix: Dockerfile β pre-install CPU torch, upgrade llama-cpp-python to >=0.3.14 (qwen3 support) 5cfcd30 imtrt004 commited on Feb 27