Optimize inference performance: add continuous-batching, NO_WARMUP support, improve thread/batch recommendations 392efee 0xarchit commited on May 19
Save models using repo filename, set MODEL_PATH from marker, ensure HF xet log dirs writable 885a74c 0xarchit commited on May 19
Auto-tune ctx-size, batch and cache types for low-memory (<=16GB) Spaces; export CACHE_TYPE_{K,V} 78e4a96 0xarchit commited on May 19
Add QUANT_PREFERENCE and PERF_PROFILE (low_latency/balanced/throughput) with server defaults f5bcacf 0xarchit commited on May 19