Switch to Qwen-1.5-1.8B-Chat - verified multilingual model with good Indic support 3862877 hardkpentium101 Qwen-Coder commited on Mar 10
Switch to AI4Bharat IndicLLM - better support for 11 Indic languages 057cc64 hardkpentium101 Qwen-Coder commited on Mar 10
Use bitsandbytes 4-bit quantization instead of AirLLM (more stable) 83eb81f hardkpentium101 Qwen-Coder commited on Mar 9
Fix cache directory permissions for non-root user 2e1f00e hardkpentium101 Qwen-Coder commited on Mar 9
Use AirLLM 4-bit quantization for Sarvam-1 (uses ~1.5GB RAM) c47fb58 hardkpentium101 Qwen-Coder commited on Mar 9
Switch to TinyLlama-1.1B with float16 for lower memory 916bdad hardkpentium101 Qwen-Coder commited on Mar 9
Pre-download models in Dockerfile, use cache at runtime d69e53e hardkpentium101 Qwen-Coder commited on Mar 8
Simplify: backend-only for HF Spaces, frontend on Netlify a47b955 hardkpentium101 Qwen-Coder commited on Mar 8
Add root Dockerfile combining backend and frontend 555caca hardkpentium101 Qwen-Coder commited on Mar 7