Spaces:
Running on Zero
Running on Zero
| title: Gemma 4 E2B QAT Mobile Demo | |
| emoji: "\U0001F48E" | |
| colorFrom: blue | |
| colorTo: indigo | |
| sdk: gradio | |
| sdk_version: 6.22.0 | |
| app_file: app.py | |
| short_description: Multimodal chat demo of Gemma 4 E2B QAT Mobile | |
| python_version: "3.12" | |
| startup_duration_timeout: 30m | |
| # Gemma 4 E2B (QAT Mobile) — Multimodal Demo | |
| A Gradio + ZeroGPU demo running **[`unsloth/gemma-4-E2B-it-qat-mobile`](https://huggingface.co/unsloth/gemma-4-E2B-it-qat-mobile)**, the mobile-optimized (wNa8o8) quantization-aware-trained checkpoint of Google's Gemma 4 E2B model. | |
| ## What it does | |
| - Multimodal chat: **text + image** input, streaming text output. | |
| - Optional **thinking mode** (the model reasons step-by-step before answering). | |
| - Adjustable sampling (temperature / top-p / top-k / max tokens). | |
| The model is a 2.3B-effective-parameter multimodal LLM (text, image, audio) quantized to ~2.5 GB via Unsloth's mobile-optimized wNa8o8 scheme, making it runnable on a single ZeroGPU. | |
| ## Stack | |
| - **SDK**: Gradio 6 on Hugging Face Spaces | |
| - **Hardware**: ZeroGPU (`zero-a10g`) — free-tier friendly | |
| - **Model loading**: `transformers` `AutoModelForMultimodalLM` + `AutoProcessor` | |
| > The first message after a cold start may take a few extra seconds while ZeroGPU streams the packed weights into VRAM. | |
| ## Notes | |
| - Image inputs are placed **before** the text, per the Gemma 4 best practices. | |
| - Default sampling follows the official recommendation: `temperature=1.0`, `top_p=0.95`, `top_k=64`. | |