Text Generation
Transformers
Safetensors
English
diffusion_gemma
image-text-to-text
diffusion-language-model
generative-ui
openui
openui-lang
gemma
conversational
Instructions to use thesysdev/OUI-1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use thesysdev/OUI-1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="thesysdev/OUI-1") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("thesysdev/OUI-1") model = AutoModelForMultimodalLM.from_pretrained("thesysdev/OUI-1", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use thesysdev/OUI-1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "thesysdev/OUI-1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "thesysdev/OUI-1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/thesysdev/OUI-1
- SGLang
How to use thesysdev/OUI-1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "thesysdev/OUI-1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "thesysdev/OUI-1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "thesysdev/OUI-1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "thesysdev/OUI-1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use thesysdev/OUI-1 with Docker Model Runner:
docker model run hf.co/thesysdev/OUI-1
| license: gemma | |
| base_model: google/diffusiongemma-26B-A4B-it | |
| pipeline_tag: text-generation | |
| library_name: transformers | |
| tags: | |
| - diffusion-language-model | |
| - generative-ui | |
| - openui | |
| - openui-lang | |
| - gemma | |
| language: | |
| - en | |
| # OUI-1 | |
| **OUI-1** is the first diffusion model built for generative UI. It is a finetune of Google's | |
| DiffusionGemma 26B-A4B-it that writes user interface screens in | |
| [openui-lang](https://openui.com/docs), the declarative UI language behind | |
| [OpenUI](https://github.com/thesysdev/openui). It scores 71.7% on | |
| [Generative UI Benchmark](https://www.openui.com/benchmarks), 5.5x its base model, with 4B active parameters. | |
| Given a component library's signatures in the system prompt and a plain-language brief, it | |
| returns the screen as code, one component per line, wired into a root. Any component library | |
| with signatures can be the prompt, and any OpenUI app can point at it. | |
| Generative UI is unusually sensitive to latency, and this is a text diffusion model: it writes a | |
| 256-token block at once, starting from noise and committing each token the moment it is sure of | |
| it, so a screen arrives in about a second on one GPU. | |
| | | | | |
| |---|---| | |
| | Base model | google/diffusiongemma-26B-A4B-it (26B total, 4B active) | | |
| | Method | LoRA finetuning, merged into the base weights; bf16 safetensors | | |
| | Context | 16,384 tokens as served | | |
| | License | [Gemma Terms of Use](https://ai.google.dev/gemma/terms). This is a Gemma derivative; the base license and its use restrictions apply. | | |
| | Publisher | Thesys | | |
| ## Results on the Generative UI Benchmark | |
| The [generative-ui-bench](https://github.com/thesysdev/generative-ui-bench) protocol: 46 screen | |
| briefs in five size bands, 4 generations each, thinking off, one shared system prompt for every | |
| model, scored by the benchmark's own validator. Base DiffusionGemma 24/184 (13.0%); OUI-1 | |
| 132/184 (71.7%). Raw outputs for OUI-1 are committed to the benchmark repo so the number can be | |
| rescored offline. | |
|  | |
| Every number here was measured with the serving settings below (vLLM 0.24, FP8, the checkpoint's | |
| own sampler at 48 denoising steps). One request at a time on an A100 80GB, a light screen takes | |
| about a second and a dense one three to six seconds, prompt included. | |
| ## How to use | |
| The model reads the component library from the system prompt and answers with an openui-lang | |
| program. Build the system prompt for your own component library with `@openuidev/cli` | |
| (`npx @openuidev/cli generate <library.ts> --out system-prompt.txt`), or use the benchmark's | |
| reference prompt (`protocols/openui/prompt.ts` in generative-ui-bench) to try it as is. Render | |
| the output with `@openuidev/react-lang` (or the Vue and Svelte renderers) and validate it with | |
| `@openuidev/lang-core`. | |
| ### vLLM (recommended) | |
| vLLM 0.24 or newer. | |
| ```bash | |
| pip install "vllm>=0.24" | |
| vllm serve thesysdev/OUI-1 --trust-remote-code --max-model-len 16384 --quantization fp8 \ | |
| --served-model-name OUI-1 --max-num-seqs 4 \ | |
| --enable-auto-tool-choice --tool-call-parser gemma4 | |
| ``` | |
| The sampler settings come from the checkpoint: the 256-token canvas from `config.json`, the | |
| entropy-bound sampler (entropy bound 0.1) and 48 denoising steps from `generation_config.json`. | |
| No other flags or environment variables are needed; every number above was measured with this | |
| line. The weights take 25.8 GiB at FP8; vLLM then fills the rest of the card with KV cache by | |
| default, so pass `--gpu-memory-utilization` if the model has to share a GPU. On Ampere (A100) | |
| there is no native FP8, so vLLM uses weight-only FP8 through Marlin: the memory saving is real, | |
| the compute speedup is not. | |
| Then call it as an OpenAI-compatible chat model. `temperature` and `seed` are ignored: the | |
| sampler runs the checkpoint's own schedule and no per-request seed is plumbed through, so two | |
| identical requests can return differently worded screens. The benchmark runs at `max_tokens` | |
| 8192; 4096 covers every screen in the bands above. | |
| ```python | |
| from openai import OpenAI | |
| client = OpenAI(base_url="http://localhost:8000/v1", api_key="local") | |
| system = open("system-prompt.txt").read() # your component library, from @openuidev/cli | |
| brief = """Status page for the platform team. Build a single screen for this. It must show: | |
| 1. current uptime percentage for the API this month | |
| 2. a short note on the most recent incident and when it was resolved | |
| Cover every numbered item.""" | |
| r = client.chat.completions.create( | |
| model="OUI-1", | |
| messages=[{"role": "system", "content": system}, {"role": "user", "content": brief}], | |
| max_tokens=4096, | |
| stream=False, | |
| ) | |
| print(r.choices[0].message.content) # openui-lang: one component per line, wired into root | |
| ``` | |
| `stream=True` works, with one difference from an autoregressive model: text arrives one | |
| 256-token canvas at a time, so a screen that fits in one canvas arrives as a single chunk. | |
| To change the step cap use `--diffusion-config '{"canvas_length":256,"max_denoising_steps":32}'` | |
| (32 and 48 score the same on the benchmark, 16 breaks the reference graph). `--hf-overrides` keys | |
| such as `diffusion_max_denoising_steps` are not read; vLLM takes the step count from | |
| `--diffusion-config` or, failing that, from `generation_config.json`. | |
| ### Function calling | |
| The model keeps Gemma 4's native tool-call format. With `--tool-call-parser gemma4` it returns | |
| standard OpenAI `tool_calls` for tools you pass in the request, then writes the screen from the | |
| tool result on the next turn. It needs one instruction in the system prompt to prefer the tool | |
| over inventing data, e.g.: "When the user asks about the weather or a stock price you MUST call the | |
| matching tool first and output nothing else in that turn; write the screen from the returned | |
| values on the next turn." In our test (weather, stock price tools) it called the right tool with | |
| the right arguments on 9 of 9 asks, never called one on 3 unrelated asks, and used the returned | |
| numbers verbatim. | |
| Leave `tool_choice` at its default, `"auto"`. `"required"` and a named function are accepted but | |
| ignored (vLLM has no structured outputs for diffusion models yet): the model answers with an | |
| ordinary screen, and with `"required"` the response carries `finish_reason: "tool_calls"` while | |
| `tool_calls` is empty, so do not branch on `finish_reason` alone. | |
| ```python | |
| import json | |
| TOOL_RULE = ("\n\nWhen the user asks about the weather or a stock price you MUST call the " | |
| "matching tool first and output nothing else in that turn; write the screen " | |
| "from the returned values on the next turn.") | |
| tools = [{"type": "function", "function": {"name": "get_weather", | |
| "description": "Current weather for a city", | |
| "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}}}] | |
| msgs = [{"role": "system", "content": system + TOOL_RULE}, | |
| {"role": "user", "content": "weather in Goa this weekend"}] | |
| r = client.chat.completions.create(model="OUI-1", tools=tools, max_tokens=4096, messages=msgs) | |
| m = r.choices[0].message | |
| call = m.tool_calls[0] # get_weather({"city": "Goa"}) | |
| result = {"city": "Goa", "condition": "Partly cloudy", "high_c": 31, "low_c": 25, "rain_chance_pct": 20} | |
| msgs.append({"role": "assistant", "content": m.content or "", | |
| "tool_calls": [{"id": call.id, "type": "function", | |
| "function": {"name": call.function.name, | |
| "arguments": call.function.arguments}}]}) | |
| msgs.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)}) | |
| r2 = client.chat.completions.create(model="OUI-1", tools=tools, max_tokens=4096, messages=msgs) | |
| print(r2.choices[0].message.content) # the screen, built from the tool result | |
| ``` | |
| ### Transformers | |
| Transformers 5.11 or newer. `AutoModelForCausalLM` does not resolve this architecture, so import | |
| the model class directly. bf16 peaks at about 52 GiB of GPU memory (one A100 80GB or H100). | |
| `generate` chains 256-token canvases and stops at EOS; keep `max_new_tokens` above 256 or the | |
| screen is cut off after the first canvas. | |
| ```python | |
| import torch | |
| from transformers import AutoProcessor | |
| from transformers.models.diffusion_gemma.modeling_diffusion_gemma import DiffusionGemmaForBlockDiffusion | |
| model_id = "thesysdev/OUI-1" | |
| processor = AutoProcessor.from_pretrained(model_id) | |
| model = DiffusionGemmaForBlockDiffusion.from_pretrained(model_id, dtype=torch.bfloat16, device_map="cuda") | |
| messages = [{"role": "system", "content": open("system-prompt.txt").read()}, | |
| {"role": "user", "content": brief}] | |
| ids = processor.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, | |
| return_tensors="pt").to("cuda") | |
| gc = model.generation_config # entropy-bound sampler, 48 steps, from the checkpoint | |
| gc.max_new_tokens = 256 * 32 # up to 32 canvases; EOS stops earlier | |
| out = model.generate(input_ids=ids, generation_config=gc) | |
| print(processor.decode(out.sequences[0, ids.shape[1]:], skip_special_tokens=True)) | |
| ``` | |
| ### Adapter | |
| The LoRA adapter that produced these weights is in `adapter/` (PEFT format, r=64, alpha=128, | |
| base `unsloth/diffusiongemma-26B-A4B-it`, a mirror of `google/diffusiongemma-26B-A4B-it`). It is | |
| a tied LoRA: its 205 modules target the decoder projections, and DiffusionGemma's encoder shares | |
| the same weight storage, so the merge adapts both passes. Plain `PeftModel.from_pretrained(base, | |
| adapter)` wraps the decoder only and gives a different model; call `.merge_and_unload()` to | |
| reproduce these weights, or use the merged checkpoint above. | |
| ## Intended use | |
| Generating UI screens for OpenUI-based applications where latency matters and a small, | |
| self-hosted model is preferred. Not a general chat model. Outputs should be validated by the | |
| openui-lang parser and rendered through a component library that checks its props. | |