OUI-1 / README.md
andhakanoon's picture
model card: drop the training-data section
251b1f1 verified
|
Raw
History Blame Contribute Delete
9.97 kB
---
license: gemma
base_model: google/diffusiongemma-26B-A4B-it
pipeline_tag: text-generation
library_name: transformers
tags:
- diffusion-language-model
- generative-ui
- openui
- openui-lang
- gemma
language:
- en
---
# OUI-1
**OUI-1** is the first diffusion model built for generative UI. It is a finetune of Google's
DiffusionGemma 26B-A4B-it that writes user interface screens in
[openui-lang](https://openui.com/docs), the declarative UI language behind
[OpenUI](https://github.com/thesysdev/openui). It scores 71.7% on
[Generative UI Benchmark](https://www.openui.com/benchmarks), 5.5x its base model, with 4B active parameters.
Given a component library's signatures in the system prompt and a plain-language brief, it
returns the screen as code, one component per line, wired into a root. Any component library
with signatures can be the prompt, and any OpenUI app can point at it.
Generative UI is unusually sensitive to latency, and this is a text diffusion model: it writes a
256-token block at once, starting from noise and committing each token the moment it is sure of
it, so a screen arrives in about a second on one GPU.
| | |
|---|---|
| Base model | google/diffusiongemma-26B-A4B-it (26B total, 4B active) |
| Method | LoRA finetuning, merged into the base weights; bf16 safetensors |
| Context | 16,384 tokens as served |
| License | [Gemma Terms of Use](https://ai.google.dev/gemma/terms). This is a Gemma derivative; the base license and its use restrictions apply. |
| Publisher | Thesys |
## Results on the Generative UI Benchmark
The [generative-ui-bench](https://github.com/thesysdev/generative-ui-bench) protocol: 46 screen
briefs in five size bands, 4 generations each, thinking off, one shared system prompt for every
model, scored by the benchmark's own validator. Base DiffusionGemma 24/184 (13.0%); OUI-1
132/184 (71.7%). Raw outputs for OUI-1 are committed to the benchmark repo so the number can be
rescored offline.
![OpenUI score vs active parameters, open-weight models up to 31B active](scoreboard.png)
Every number here was measured with the serving settings below (vLLM 0.24, FP8, the checkpoint's
own sampler at 48 denoising steps). One request at a time on an A100 80GB, a light screen takes
about a second and a dense one three to six seconds, prompt included.
## How to use
The model reads the component library from the system prompt and answers with an openui-lang
program. Build the system prompt for your own component library with `@openuidev/cli`
(`npx @openuidev/cli generate <library.ts> --out system-prompt.txt`), or use the benchmark's
reference prompt (`protocols/openui/prompt.ts` in generative-ui-bench) to try it as is. Render
the output with `@openuidev/react-lang` (or the Vue and Svelte renderers) and validate it with
`@openuidev/lang-core`.
### vLLM (recommended)
vLLM 0.24 or newer.
```bash
pip install "vllm>=0.24"
vllm serve thesysdev/OUI-1 --trust-remote-code --max-model-len 16384 --quantization fp8 \
--served-model-name OUI-1 --max-num-seqs 4 \
--enable-auto-tool-choice --tool-call-parser gemma4
```
The sampler settings come from the checkpoint: the 256-token canvas from `config.json`, the
entropy-bound sampler (entropy bound 0.1) and 48 denoising steps from `generation_config.json`.
No other flags or environment variables are needed; every number above was measured with this
line. The weights take 25.8 GiB at FP8; vLLM then fills the rest of the card with KV cache by
default, so pass `--gpu-memory-utilization` if the model has to share a GPU. On Ampere (A100)
there is no native FP8, so vLLM uses weight-only FP8 through Marlin: the memory saving is real,
the compute speedup is not.
Then call it as an OpenAI-compatible chat model. `temperature` and `seed` are ignored: the
sampler runs the checkpoint's own schedule and no per-request seed is plumbed through, so two
identical requests can return differently worded screens. The benchmark runs at `max_tokens`
8192; 4096 covers every screen in the bands above.
```python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="local")
system = open("system-prompt.txt").read() # your component library, from @openuidev/cli
brief = """Status page for the platform team. Build a single screen for this. It must show:
1. current uptime percentage for the API this month
2. a short note on the most recent incident and when it was resolved
Cover every numbered item."""
r = client.chat.completions.create(
model="OUI-1",
messages=[{"role": "system", "content": system}, {"role": "user", "content": brief}],
max_tokens=4096,
stream=False,
)
print(r.choices[0].message.content) # openui-lang: one component per line, wired into root
```
`stream=True` works, with one difference from an autoregressive model: text arrives one
256-token canvas at a time, so a screen that fits in one canvas arrives as a single chunk.
To change the step cap use `--diffusion-config '{"canvas_length":256,"max_denoising_steps":32}'`
(32 and 48 score the same on the benchmark, 16 breaks the reference graph). `--hf-overrides` keys
such as `diffusion_max_denoising_steps` are not read; vLLM takes the step count from
`--diffusion-config` or, failing that, from `generation_config.json`.
### Function calling
The model keeps Gemma 4's native tool-call format. With `--tool-call-parser gemma4` it returns
standard OpenAI `tool_calls` for tools you pass in the request, then writes the screen from the
tool result on the next turn. It needs one instruction in the system prompt to prefer the tool
over inventing data, e.g.: "When the user asks about the weather or a stock price you MUST call the
matching tool first and output nothing else in that turn; write the screen from the returned
values on the next turn." In our test (weather, stock price tools) it called the right tool with
the right arguments on 9 of 9 asks, never called one on 3 unrelated asks, and used the returned
numbers verbatim.
Leave `tool_choice` at its default, `"auto"`. `"required"` and a named function are accepted but
ignored (vLLM has no structured outputs for diffusion models yet): the model answers with an
ordinary screen, and with `"required"` the response carries `finish_reason: "tool_calls"` while
`tool_calls` is empty, so do not branch on `finish_reason` alone.
```python
import json
TOOL_RULE = ("\n\nWhen the user asks about the weather or a stock price you MUST call the "
"matching tool first and output nothing else in that turn; write the screen "
"from the returned values on the next turn.")
tools = [{"type": "function", "function": {"name": "get_weather",
"description": "Current weather for a city",
"parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}}}]
msgs = [{"role": "system", "content": system + TOOL_RULE},
{"role": "user", "content": "weather in Goa this weekend"}]
r = client.chat.completions.create(model="OUI-1", tools=tools, max_tokens=4096, messages=msgs)
m = r.choices[0].message
call = m.tool_calls[0] # get_weather({"city": "Goa"})
result = {"city": "Goa", "condition": "Partly cloudy", "high_c": 31, "low_c": 25, "rain_chance_pct": 20}
msgs.append({"role": "assistant", "content": m.content or "",
"tool_calls": [{"id": call.id, "type": "function",
"function": {"name": call.function.name,
"arguments": call.function.arguments}}]})
msgs.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)})
r2 = client.chat.completions.create(model="OUI-1", tools=tools, max_tokens=4096, messages=msgs)
print(r2.choices[0].message.content) # the screen, built from the tool result
```
### Transformers
Transformers 5.11 or newer. `AutoModelForCausalLM` does not resolve this architecture, so import
the model class directly. bf16 peaks at about 52 GiB of GPU memory (one A100 80GB or H100).
`generate` chains 256-token canvases and stops at EOS; keep `max_new_tokens` above 256 or the
screen is cut off after the first canvas.
```python
import torch
from transformers import AutoProcessor
from transformers.models.diffusion_gemma.modeling_diffusion_gemma import DiffusionGemmaForBlockDiffusion
model_id = "thesysdev/OUI-1"
processor = AutoProcessor.from_pretrained(model_id)
model = DiffusionGemmaForBlockDiffusion.from_pretrained(model_id, dtype=torch.bfloat16, device_map="cuda")
messages = [{"role": "system", "content": open("system-prompt.txt").read()},
{"role": "user", "content": brief}]
ids = processor.apply_chat_template(messages, tokenize=True, add_generation_prompt=True,
return_tensors="pt").to("cuda")
gc = model.generation_config # entropy-bound sampler, 48 steps, from the checkpoint
gc.max_new_tokens = 256 * 32 # up to 32 canvases; EOS stops earlier
out = model.generate(input_ids=ids, generation_config=gc)
print(processor.decode(out.sequences[0, ids.shape[1]:], skip_special_tokens=True))
```
### Adapter
The LoRA adapter that produced these weights is in `adapter/` (PEFT format, r=64, alpha=128,
base `unsloth/diffusiongemma-26B-A4B-it`, a mirror of `google/diffusiongemma-26B-A4B-it`). It is
a tied LoRA: its 205 modules target the decoder projections, and DiffusionGemma's encoder shares
the same weight storage, so the merge adapts both passes. Plain `PeftModel.from_pretrained(base,
adapter)` wraps the decoder only and gives a different model; call `.merge_and_unload()` to
reproduce these weights, or use the merged checkpoint above.
## Intended use
Generating UI screens for OpenUI-based applications where latency matters and a small,
self-hosted model is preferred. Not a general chat model. Outputs should be validated by the
openui-lang parser and rendered through a component library that checks its props.