Commit History

Pull ollama model at runtime, switch to qwen3:8b
4638ba4

azettl Claude commited on

Run Qwen3-14B locally via ollama on T4 GPU
dd23ce5

azettl Claude commited on

Fix LLM timeout: return StreamingResponse before agentic loop
6ab878f

azettl Claude commited on

Cap max_tokens at 4096 to stay within HF Inference Provider limits
7ebe9fb

azettl Claude commited on

Fix STT audio + faster speech: dual AudioContext, bump LLM to Qwen3-32B
a390020

azettl Claude commited on

Initial Postman Voice Agent
adcabed

azettl commited on

Initial Postman Voice Agent implementation
9a62af9

azettl commited on