File size: 9,971 Bytes
4d103df
 
97c9e59
 
 
 
 
 
 
 
 
 
 
4d103df
97c9e59
fa9d881
97c9e59
fa9d881
87a7756
 
 
96e620a
97c9e59
87a7756
 
 
 
 
 
 
97c9e59
 
 
 
 
 
 
 
 
96e620a
97c9e59
 
 
fa9d881
 
2eedc6a
 
 
 
 
 
 
97c9e59
 
 
d380b91
 
 
 
 
 
 
 
 
87a7756
97c9e59
 
d380b91
97c9e59
fa9d881
97c9e59
 
 
87a7756
 
 
 
 
 
 
 
 
 
 
 
d380b91
 
 
 
 
 
87a7756
 
 
d380b91
 
fa9d881
d380b91
 
 
 
 
 
 
87a7756
 
d380b91
87a7756
d380b91
87a7756
 
d380b91
 
 
 
 
 
 
 
 
 
87a7756
 
 
 
 
 
d380b91
 
87a7756
 
 
 
 
 
d380b91
 
 
87a7756
 
fa9d881
87a7756
 
 
 
 
 
 
 
 
fa9d881
87a7756
d380b91
 
 
 
87a7756
 
 
 
d380b91
 
 
87a7756
 
d380b91
 
 
87a7756
d380b91
 
 
87a7756
 
d380b91
87a7756
d380b91
 
 
 
87a7756
 
 
 
 
 
 
 
d380b91
97c9e59
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
---
license: gemma
base_model: google/diffusiongemma-26B-A4B-it
pipeline_tag: text-generation
library_name: transformers
tags:
  - diffusion-language-model
  - generative-ui
  - openui
  - openui-lang
  - gemma
language:
  - en
---

# OUI-1

**OUI-1** is the first diffusion model built for generative UI. It is a finetune of Google's
DiffusionGemma 26B-A4B-it that writes user interface screens in
[openui-lang](https://openui.com/docs), the declarative UI language behind
[OpenUI](https://github.com/thesysdev/openui). It scores 71.7% on
[Generative UI Benchmark](https://www.openui.com/benchmarks), 5.5x its base model, with 4B active parameters.

Given a component library's signatures in the system prompt and a plain-language brief, it
returns the screen as code, one component per line, wired into a root. Any component library
with signatures can be the prompt, and any OpenUI app can point at it.

Generative UI is unusually sensitive to latency, and this is a text diffusion model: it writes a
256-token block at once, starting from noise and committing each token the moment it is sure of
it, so a screen arrives in about a second on one GPU.

| | |
|---|---|
| Base model | google/diffusiongemma-26B-A4B-it (26B total, 4B active) |
| Method | LoRA finetuning, merged into the base weights; bf16 safetensors |
| Context | 16,384 tokens as served |
| License | [Gemma Terms of Use](https://ai.google.dev/gemma/terms). This is a Gemma derivative; the base license and its use restrictions apply. |
| Publisher | Thesys |

## Results on the Generative UI Benchmark

The [generative-ui-bench](https://github.com/thesysdev/generative-ui-bench) protocol: 46 screen
briefs in five size bands, 4 generations each, thinking off, one shared system prompt for every
model, scored by the benchmark's own validator. Base DiffusionGemma 24/184 (13.0%); OUI-1
132/184 (71.7%). Raw outputs for OUI-1 are committed to the benchmark repo so the number can be
rescored offline.

![OpenUI score vs active parameters, open-weight models up to 31B active](scoreboard.png)

Every number here was measured with the serving settings below (vLLM 0.24, FP8, the checkpoint's
own sampler at 48 denoising steps). One request at a time on an A100 80GB, a light screen takes
about a second and a dense one three to six seconds, prompt included.

## How to use

The model reads the component library from the system prompt and answers with an openui-lang
program. Build the system prompt for your own component library with `@openuidev/cli`
(`npx @openuidev/cli generate <library.ts> --out system-prompt.txt`), or use the benchmark's
reference prompt (`protocols/openui/prompt.ts` in generative-ui-bench) to try it as is. Render
the output with `@openuidev/react-lang` (or the Vue and Svelte renderers) and validate it with
`@openuidev/lang-core`.

### vLLM (recommended)

vLLM 0.24 or newer.

```bash
pip install "vllm>=0.24"
vllm serve thesysdev/OUI-1 --trust-remote-code --max-model-len 16384 --quantization fp8 \
  --served-model-name OUI-1 --max-num-seqs 4 \
  --enable-auto-tool-choice --tool-call-parser gemma4
```

The sampler settings come from the checkpoint: the 256-token canvas from `config.json`, the
entropy-bound sampler (entropy bound 0.1) and 48 denoising steps from `generation_config.json`.
No other flags or environment variables are needed; every number above was measured with this
line. The weights take 25.8 GiB at FP8; vLLM then fills the rest of the card with KV cache by
default, so pass `--gpu-memory-utilization` if the model has to share a GPU. On Ampere (A100)
there is no native FP8, so vLLM uses weight-only FP8 through Marlin: the memory saving is real,
the compute speedup is not.

Then call it as an OpenAI-compatible chat model. `temperature` and `seed` are ignored: the
sampler runs the checkpoint's own schedule and no per-request seed is plumbed through, so two
identical requests can return differently worded screens. The benchmark runs at `max_tokens`
8192; 4096 covers every screen in the bands above.

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="local")
system = open("system-prompt.txt").read()  # your component library, from @openuidev/cli
brief = """Status page for the platform team. Build a single screen for this. It must show:
1. current uptime percentage for the API this month
2. a short note on the most recent incident and when it was resolved
Cover every numbered item."""
r = client.chat.completions.create(
    model="OUI-1",
    messages=[{"role": "system", "content": system}, {"role": "user", "content": brief}],
    max_tokens=4096,
    stream=False,
)
print(r.choices[0].message.content)  # openui-lang: one component per line, wired into root
```

`stream=True` works, with one difference from an autoregressive model: text arrives one
256-token canvas at a time, so a screen that fits in one canvas arrives as a single chunk.

To change the step cap use `--diffusion-config '{"canvas_length":256,"max_denoising_steps":32}'`
(32 and 48 score the same on the benchmark, 16 breaks the reference graph). `--hf-overrides` keys
such as `diffusion_max_denoising_steps` are not read; vLLM takes the step count from
`--diffusion-config` or, failing that, from `generation_config.json`.

### Function calling

The model keeps Gemma 4's native tool-call format. With `--tool-call-parser gemma4` it returns
standard OpenAI `tool_calls` for tools you pass in the request, then writes the screen from the
tool result on the next turn. It needs one instruction in the system prompt to prefer the tool
over inventing data, e.g.: "When the user asks about the weather or a stock price you MUST call the
matching tool first and output nothing else in that turn; write the screen from the returned
values on the next turn." In our test (weather, stock price tools) it called the right tool with
the right arguments on 9 of 9 asks, never called one on 3 unrelated asks, and used the returned
numbers verbatim.

Leave `tool_choice` at its default, `"auto"`. `"required"` and a named function are accepted but
ignored (vLLM has no structured outputs for diffusion models yet): the model answers with an
ordinary screen, and with `"required"` the response carries `finish_reason: "tool_calls"` while
`tool_calls` is empty, so do not branch on `finish_reason` alone.

```python
import json

TOOL_RULE = ("\n\nWhen the user asks about the weather or a stock price you MUST call the "
             "matching tool first and output nothing else in that turn; write the screen "
             "from the returned values on the next turn.")

tools = [{"type": "function", "function": {"name": "get_weather",
          "description": "Current weather for a city",
          "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}}}]
msgs = [{"role": "system", "content": system + TOOL_RULE},
        {"role": "user", "content": "weather in Goa this weekend"}]
r = client.chat.completions.create(model="OUI-1", tools=tools, max_tokens=4096, messages=msgs)
m = r.choices[0].message
call = m.tool_calls[0]                              # get_weather({"city": "Goa"})

result = {"city": "Goa", "condition": "Partly cloudy", "high_c": 31, "low_c": 25, "rain_chance_pct": 20}
msgs.append({"role": "assistant", "content": m.content or "",
             "tool_calls": [{"id": call.id, "type": "function",
                             "function": {"name": call.function.name,
                                          "arguments": call.function.arguments}}]})
msgs.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)})
r2 = client.chat.completions.create(model="OUI-1", tools=tools, max_tokens=4096, messages=msgs)
print(r2.choices[0].message.content)                # the screen, built from the tool result
```

### Transformers

Transformers 5.11 or newer. `AutoModelForCausalLM` does not resolve this architecture, so import
the model class directly. bf16 peaks at about 52 GiB of GPU memory (one A100 80GB or H100).
`generate` chains 256-token canvases and stops at EOS; keep `max_new_tokens` above 256 or the
screen is cut off after the first canvas.

```python
import torch
from transformers import AutoProcessor
from transformers.models.diffusion_gemma.modeling_diffusion_gemma import DiffusionGemmaForBlockDiffusion

model_id = "thesysdev/OUI-1"
processor = AutoProcessor.from_pretrained(model_id)
model = DiffusionGemmaForBlockDiffusion.from_pretrained(model_id, dtype=torch.bfloat16, device_map="cuda")

messages = [{"role": "system", "content": open("system-prompt.txt").read()},
            {"role": "user", "content": brief}]
ids = processor.apply_chat_template(messages, tokenize=True, add_generation_prompt=True,
                                    return_tensors="pt").to("cuda")
gc = model.generation_config          # entropy-bound sampler, 48 steps, from the checkpoint
gc.max_new_tokens = 256 * 32          # up to 32 canvases; EOS stops earlier
out = model.generate(input_ids=ids, generation_config=gc)
print(processor.decode(out.sequences[0, ids.shape[1]:], skip_special_tokens=True))
```

### Adapter

The LoRA adapter that produced these weights is in `adapter/` (PEFT format, r=64, alpha=128,
base `unsloth/diffusiongemma-26B-A4B-it`, a mirror of `google/diffusiongemma-26B-A4B-it`). It is
a tied LoRA: its 205 modules target the decoder projections, and DiffusionGemma's encoder shares
the same weight storage, so the merge adapts both passes. Plain `PeftModel.from_pretrained(base,
adapter)` wraps the decoder only and gives a different model; call `.merge_and_unload()` to
reproduce these weights, or use the merged checkpoint above.

## Intended use

Generating UI screens for OpenUI-based applications where latency matters and a small,
self-hosted model is preferred. Not a general chat model. Outputs should be validated by the
openui-lang parser and rendered through a component library that checks its props.