--- title: Zero-GPU Side-by-Side LLM Comparison emoji: "\U0001F916" sdk: gradio app_file: app.py --- # Hugging Face Space: Zero-GPU Side-by-Side LLM Comparison This Space compares four instruction-tuned models side-by-side in a chat interface: - Gemma4 31B-IT - Gemma4 26B-IT - Qwen 3.6 35B-IT - Qwen 3.6 27B The UI is designed for **zero local GPU** usage. It calls Hugging Face hosted inference endpoints via `huggingface_hub.InferenceClient`, so the Space itself only needs CPU resources. ## What is included - `app.py`: Gradio app with four parallel chat columns. - `requirements.txt`: Runtime dependencies. ## Important setup 1. Add your Hugging Face token in Space Secrets: - `HF_API_TOKEN` (or `HUGGINGFACE_HUB_TOKEN` or `HF_TOKEN`) 2. Ensure model IDs in `MODEL_MATRIX` are available through the Hugging Face Inference API. 3. This project is configured to use the fixed best Unsloth checkpoints for each family: - `unsloth/gemma-4-31B-it-unsloth-bnb-4bit` - `unsloth/gemma-4-26B-A4B-it` - `unsloth/Qwen3.6-35B-A3B-GGUF` - `unsloth/Qwen3.6-27B-GGUF` ## Optional customization - Add more precision variants in each model card inside `MODEL_MATRIX`. - Change `MODEL_MATRIX` entries to your preferred checkpoints. - If your model requires a different prompt format, edit `_SYSTEM_PROMPT` and `_build_messages` in `app.py`. ## Notes - For private or gated models, confirm your token has access. - If a checkpoint does not support the chat endpoint, the app tries a text- - ZeroGPU is the `zero-a10g` hardware flavor. In Spaces config, set this at the Space level (the README `hardware:` field is not always respected for creation).