fix(api): run 2 uvicorn workers — a blocking LLM call in an async endpoint was freezing the single event loop (HTTP 000 outages) 9318680 verified MHamdan commited on 21 days ago
feat(gpu): Ollama LLM server (qwen2.5:3b) for GPU inference + on-demand T4 with 15min auto-sleep 39e5349 verified MHamdan commited on 22 days ago
Deploy Auralynq RAG (Llama-3.3-70B via HF Inference Providers) 8c1b9fe verified MHamdan commited on 29 days ago