Optimize inference performance: add continuous-batching, NO_WARMUP support, improve thread/batch recommendations 392efee 0xarchit commited on May 19
Add LangSearch web search tool for small model knowledge augmentation c459436 0xarchit commited on May 19
Use temp download dir and robust rmtree handler to avoid permission errors on cleanup b3c9994 0xarchit commited on May 19
Save models using repo filename, set MODEL_PATH from marker, ensure HF xet log dirs writable 885a74c 0xarchit commited on May 19
Allow MODEL_FILE override; improve diagnostics when no .gguf files present f237530 0xarchit commited on May 19
Auto-tune ctx-size, batch and cache types for low-memory (<=16GB) Spaces; export CACHE_TYPE_{K,V} 78e4a96 0xarchit commited on May 19
Add QUANT_PREFERENCE and PERF_PROFILE (low_latency/balanced/throughput) with server defaults f5bcacf 0xarchit commited on May 19
Speed up build (ccache) and fix /data permissions: create appuser and run downloader/server as appuser c5437e1 0xarchit commited on May 19