diffusiongemma-26b-a4b-it / MODEL_RESEARCH.md
Mike0021's picture
Document verified deployment
77baae9 verified
|
Raw
History Blame Contribute Delete
4 kB
# DiffusionGemma Space Research
Verified on 2026-06-10 against the Hugging Face model card, model files, current Hub metadata, Hugging Face docs, and the supplied Spaces quickstart.
## Model
- Model ID: `google/diffusiongemma-26B-A4B-it`
- Task: `image-text-to-text`
- Library: `transformers`
- Architecture: `diffusion_gemma`
- Loading class in Hub metadata: `AutoModelForMultimodalLM`
- Concrete generation class in model files: `DiffusionGemmaForBlockDiffusion`
- Parameters: about 25.2B total, 3.8B active per token
- Weight storage: BF16 safetensors, about 51.7 GB
- Context length: up to 256K tokens
- Canvas length: 256
- Vision support: image input with soft-token budgets `70`, `140`, `280`, `560`, `1120`
- Chat template: supports `system`, `user`, and `assistant` roles, tool-call formatting, and `enable_thinking`
## Generation Defaults
The model ships a custom `DiffusionGemmaGenerationConfig`; ordinary autoregressive controls like top-p/top-k are not the primary controls for this model.
- `max_new_tokens`: 256 in `generation_config.json`; the Space defaults to 512 for better thinking-mode headroom
- `max_denoising_steps`: 48
- Sampler: `EntropyBoundSamplerConfig(entropy_bound=0.1)`
- Temperature schedule: `t_max=0.8`, `t_min=0.4`
- Adaptive stopping: `confidence_threshold=0.005`, `stability_threshold=1`
- EOS IDs: `[1, 106, 50]`
- PAD ID: `0`
## Space Design Decisions
- Use Gradio SDK because ZeroGPU is Gradio-only.
- Use `@spaces.GPU(size="xlarge")` because BF16 weights exceed the default 48 GB ZeroGPU allocation.
- Import `spaces` before `torch`.
- Set `HF_HOME`, `HF_MODULES_CACHE`, `MPLCONFIGDIR`, and `GRADIO_SSR_MODE` before heavy imports.
- Load the model once at module import and place it on `cuda`, following current ZeroGPU model-loading guidance.
- Use `AutoModelForMultimodalLM` from `transformers.models.auto.modeling_auto` because the pinned source commit maps `diffusion_gemma` there.
- Pin Transformers from source commit `513146c175a534a913b8cfd448b9a06f0caf9365`; the latest PyPI wheel visible locally did not include `transformers.models.diffusion_gemma`.
- Include unpinned `torchvision`; the Gemma 4 image processor used by DiffusionGemma imports `torchvision.transforms.v2`.
- Do not include bypass model-loading paths, local server checks, integer CUDA device IDs, or pinned `torch`/`torchaudio`.
- Preserve only final assistant content in model history; do not feed previous thinking traces back into future turns.
- Put image content before text content in each multimodal user turn.
- Gradio 6 `Chatbot` uses message dictionaries by default; do not pass the removed `type="messages"` parameter.
- DiffusionGemma emits fixed 256-token canvases. The Space trims EOS/PAD markers from displayed text, reports both displayed and canvas token counts, and exposes max-new-token values in 256-token increments.
- Warm live measurements on ZeroGPU were much shorter than the initial conservative reservation; `_estimate_gpu_seconds()` is tuned to the measured default text and image calls while preserving headroom for longer, image, and thinking-mode requests.
## Deployment Checks
Local checks should stop at:
```bash
python3 -m py_compile app.py
```
After upload, use Space logs and live API verification:
Use the Hub UI or `huggingface_hub.HfApi.fetch_space_logs()` for logs if the installed `hf` CLI does not include Space log commands.
Then call `Client("<namespace>/<space-name>").view_api()` before any `predict()` call.
## Final Deployment
- Space: <https://huggingface.co/spaces/Mike0021/diffusiongemma-26b-a4b-it>
- Final app commit: `a67bfb0f7f81b6357e35a9e8c29e4581baf2f132`
- Runtime: `RUNNING` on `zero-a10g`
- Text API smoke test: `What is 2 + 2?` returned `2 + 2 equals 4.` with no special-token leakage.
- Default-style text test: block-diffusion paragraph prompt returned a clean 103-token answer in 1.5s model time.
- Image API smoke test: generated PNG containing `BLUE 7` returned `BLUE7 is visible in the image.`