Spaces:
Running on Zero
Running on Zero
| title: Zero-GPU Side-by-Side LLM Comparison | |
| emoji: "\U0001F916" | |
| sdk: gradio | |
| app_file: app.py | |
| # Hugging Face Space: Zero-GPU Side-by-Side LLM Comparison | |
| This Space compares four instruction-tuned models side-by-side in a chat interface: | |
| - Gemma4 31B-IT | |
| - Gemma4 26B-IT | |
| - Qwen 3.6 35B-IT | |
| - Qwen 3.6 27B | |
| The UI is designed for **zero local GPU** usage. It calls Hugging Face hosted | |
| inference endpoints via `huggingface_hub.InferenceClient`, so the Space itself only | |
| needs CPU resources. | |
| ## What is included | |
| - `app.py`: Gradio app with four parallel chat columns. | |
| - `requirements.txt`: Runtime dependencies. | |
| ## Important setup | |
| 1. Add your Hugging Face token in Space Secrets: | |
| - `HF_API_TOKEN` (or `HUGGINGFACE_HUB_TOKEN` or `HF_TOKEN`) | |
| 2. Ensure model IDs in `MODEL_MATRIX` are available through the Hugging Face | |
| Inference API. | |
| 3. This project is configured to use the fixed best Unsloth checkpoints for each family: | |
| - `unsloth/gemma-4-31B-it-unsloth-bnb-4bit` | |
| - `unsloth/gemma-4-26B-A4B-it` | |
| - `unsloth/Qwen3.6-35B-A3B-GGUF` | |
| - `unsloth/Qwen3.6-27B-GGUF` | |
| ## Optional customization | |
| - Add more precision variants in each model card inside `MODEL_MATRIX`. | |
| - Change `MODEL_MATRIX` entries to your preferred checkpoints. | |
| - If your model requires a different prompt format, edit `_SYSTEM_PROMPT` and | |
| `_build_messages` in `app.py`. | |
| ## Notes | |
| - For private or gated models, confirm your token has access. | |
| - If a checkpoint does not support the chat endpoint, the app tries a text- | |
| - ZeroGPU is the `zero-a10g` hardware flavor. In Spaces config, set this at the | |
| Space level (the README `hardware:` field is not always respected for creation). | |