Instructions to use Shelter/UI-testjev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Shelter/UI-testjev with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
UI-TestJev
Comparable local UI-check performance to GPT-6 Luna Instant, with lower observed latency and estimated cost.
UI-TestJev uses DiffusionGemma to compare approved and current screenshots, identify UI defects, and return structured decisions for testing pipelines: pass, unknown, or a failure category.
Dataset · Technical report · Use · Input format
What it does
- Requirement checks: detect clipped text, missing content, low contrast, occlusion, distorted images and undersized controls.
- Visual regression: compare full screenshots for a newly introduced defect.
- Keyboard checks: judge reachability from an observed Tab-navigation trace.
The release is a LoRA adapter for DiffusionGemma-26B-A4B-it, using one-step decision inference. Vision weights remain BF16; frozen expert weights use INT8.
Selected results
Fine-tuning improves local-check macro-F1 by 41.70 percentage points and full-screen regression by 19.98 points.
| Development-test measure | Base DiffusionGemma | UI-TestJev | GPT-6 Luna Instant* |
|---|---|---|---|
| Local UI checks — macro-F1 | 31.32% | 73.02% | 71.18% |
| Full-screen regression — macro-F1 | 20.55% | 40.53% | 50.80% |
| Weighted task score | 31.71% | 59.47% | 63.50% |
| Recorded median latency | — | 1.36 s (local reference) | 1.97 s (HTTPS) |
| Estimated cost / 1,941 decisions | — | $0.26 | $0.77–$1.61 |
Base DiffusionGemma is evaluated without the adapter, using the same INT8-expert/BF16-vision setup, screenshots and decision protocol. Fine-tuning increases false alarms: 1.77% → 16.31% on local checks and 0.76% → 3.05% on full-screen checks.
Luna is gpt-6-luna with reasoning off. These are development-test results. The weighted score combines local, full-screen and keyboard macro-F1 (40/50/10%). Jev latency is a local reference time on an RTX PRO 6000 Blackwell Server Edition; Luna includes HTTPS overhead. Jev cost uses djev's announced tariff as an estimate, not a quote for hosting this adapter. Measurement and pricing details.
Technical report
Data
The dataset has 1,590 training, 846 validation and 1,941 development-test decisions, plus 311 external Scry pairs. Training uses 50 pages from four template families: Start Bootstrap, Flat UI, Material Dashboard and HTML5 UP. Validation uses Bulma Templates; the development test uses AdminLTE.
Synthetic examples are built as defect/harmless-change pairs. Browser measurements check the requirement; screenshots provide the model evidence. Related variants stay in one split. Local questions include a reference target and detail crops; full-screen questions receive no target-location hint.
The audit checked 6,536 images, removed 120 duplicate rows and excluded 21 ambiguous clipping pairs. No exact image or public-input hashes cross splits. The AdminLTE family informed earlier repairs, so the development test is not an untouched final holdout. Dataset construction and coverage.
Evaluation setup
The main comparison uses 1,941 development-test decisions: 1,127 local checks, 787 full-screen checks and 27 keyboard checks. Base DiffusionGemma and UI-TestJev use the same pinned model, INT8 experts, BF16 vision, processor, prompts, option order and one-step decoding. The base arm has no LoRA adapter. This measures adaptation gains, not the effect of INT8 versus BF16 weights.
GPT-6 Luna uses reasoning off and the same public requirements, ordered images and answer options. It uses its own image processing and autoregressive decoding through Codex HTTPS. All 3,098 Luna cases across validation, development test and Scry returned valid answers with zero reported reasoning tokens.
Training source families are separate from validation and test. The AdminLTE test family was inspected during earlier dataset repair, however, so these are development results. Neither these test scores nor Scry selected the adapter checkpoint.
Detection and false alarms
Defect recall counts a defect as detected when the model predicts any failure category. Precision is the fraction of failure predictions made on actual defects. False-alarm rate is the fraction of passing examples assigned a failure category. A wrong defect category can therefore count as detection while still being an incorrect answer.
Local requirement checks
| Metric | Base | UI-TestJev | Luna |
|---|---|---|---|
| Defect precision | 92.31% | 83.21% | 84.90% |
| Defect recall | 21.31% | 80.99% | 76.91% |
| False-alarm rate | 1.77% | 16.31% | 13.65% |
| Exact-answer accuracy | 53.06% | 82.34% | 79.86% |
| Answers marked unknown | 17.30% | 0.00% | 2.84% |
Full-screen regression
| Metric | Base | UI-TestJev | Luna |
|---|---|---|---|
| Defect precision | 95.08% | 92.59% | 95.90% |
| Defect recall | 14.76% | 38.17% | 59.54% |
| False-alarm rate | 0.76% | 3.05% | 2.54% |
| Exact-answer accuracy | 53.37% | 63.15% | 70.01% |
| Answers marked unknown | 0.00% | 0.00% | 0.38% |
On local checks, Jev detects 456/563 defects, compared with 120/563 for the base. False alarms increase from 10/564 to 92/564 passing cases. On full-screen checks, detection rises from 58/393 to 150/393, while false alarms rise from 3/394 to 12/394.
The benchmark has roughly equal numbers of defective and passing visual examples. Precision in a production pipeline will depend on how often defects actually occur.
Macro-F1 in the model card averages F1 over the labels present in each task's gold answers. Visual tasks have no gold unknown labels: an unknown prediction counts as a miss for the true class, without adding an extra class to that average. The weighted score uses 40% local, 50% full-screen and 10% keyboard macro-F1. Missing or invalid predictions count as errors; none occurred in these runs.
Which errors remain?
The chart requires the correct defect category, not merely any failure prediction. A dash means the category was not evaluated for that task. Chart values.
- Control size: local recall is 142/153, but control-size warnings account for 56 of 92 local false alarms. In full-screen checks, recall falls to 61/153; another 22 cases are called missing content.
- Contrast: local recall is 88/103. Full-screen recall is 6/103, with 91 called pass. The model is much less reliable when it must find the affected region itself.
- Image aspect: local recall is 2/12 and full-screen recall is 0/12. This remains a weak category, and the sample is small.
- Category confusion: Jev flags 150 full-screen defects, but only 115 receive the right category. Reporting detection recall alone would hide 35 wrong-category answers.
Local and full-screen tasks differ in their prompts, crops and category mix. Their score gap does not isolate the causal effect of cropping.
Abstention and keyboard checks
Jev gives no unknown answers on either visual test task. The training set has no visual unknown targets, so this does not demonstrate good handling of ambiguous screenshots. Its answer scores are not calibrated probabilities.
Keyboard test accuracy is 24/27 for the base, 27/27 for Jev and 26/27 for Luna. Jev and Luna both recover all nine unknown cases; the base recovers eight. These are trace-based questions across nine scenarios, not evidence of broad keyboard-testing reliability.
Validation and uncertainty
Epoch 3 is the exported checkpoint. Its weighted validation score is 67.26%, versus 39.88% for the base, but no checkpoint met every acceptance rule. In particular, local false alarms rose from 2.07% to 12.03%; the permitted maximum was 4.07%. Full-screen clipping recall also fell from 11/26 to 10/26. A 48-case noise probe changed decisions on 4 cases, versus 1 for the base.
The export is therefore not recommended over the base under the original acceptance rules. It offers higher defect recall with a false-alarm tradeoff that requires review.
For the Luna comparison, 2,000 paired bootstrap resamples keep each page or scenario together. On local checks, Luna minus Jev macro-F1 is -1.83 percentage points, with a 95% interval of [-6.17, +3.03]. On full-screen checks it is +10.27 points, [+4.46, +16.15]. Local results are close; the full-screen difference favors Luna. These intervals describe this corpus, not generalization to new repository families.
External Scry evaluation
Scry contributes 311 reference/implementation pairs from 52 app groups. We ask 12 fixed category-presence questions per pair and score 505 annotated category instances.
| Diagnostic | Base | UI-TestJev | Luna |
|---|---|---|---|
| Annotated categories recovered | 266/505 | 337/505 | 410/505 |
| Annotated-category recall | 52.67% | 66.73% | 81.19% |
| Additional unannotated predictions | 994 | 1,214 | 1,856 |
Jev recovers more annotated categories than the base, but also flags more unannotated categories. Scry does not contain exhaustive negatives, so those extra predictions cannot all be called false positives. This evaluation measures category presence, not localization or the original Scry benchmark score. Scry has been observed during development and is not a fresh blind test.
Training method
| Setting | Completed run |
|---|---|
| Base | google/diffusiongemma-26B-A4B-it, revision f7f5b7f5fa82ffc52addd066915886d497f5517b |
| Adapter | LoRA rank 32, alpha 64, dropout 0.05; decoder attention and dense MLP projections |
| Frozen weights | Encoder, vision, experts, router, embeddings and output head |
| Learning rate / epochs | 5e-5 / 3; one configuration, seed 3407 |
| Optimizer | AdamW, betas (0.9, 0.95), weight decay 0.01, epsilon 1e-8 |
| Batch / schedule | Microbatch 1, accumulation 16; 5% warmup, cosine decay to 10% of peak; gradient norm clipped to 1.0 |
| Visual input | Original aspect-preserving processor; at most 1,120 visual tokens per image; 8,192 total input tokens; overflow rejected |
The objective is cross-entropy over each question's allowed answer tokens. Supervised answer slots are always replaced with random vocabulary tokens before the decoder sees them. Option order changes during training and stays fixed during evaluation. This is a task-specific decision objective, not ordinary next-token chat training.
All 1,590 training rows are visited once per epoch. Source families and pages are interleaved; related examples stay in sampling blocks, reshuffled each epoch. Passing visual examples receive 1.5x loss weight. No minority examples are duplicated to inflate the dataset.
Latency and cost
Jev's median development-test latency is 1.36 seconds, with p95 1.63 seconds, on an RTX PRO 6000 Blackwell Server Edition. This is a local reference time, including preprocessing and model work with a vision-feature cache. Luna's HTTPS client median is 1.97 seconds, p95 3.00 seconds, including request setup and network. The hardware and timing boundaries differ; these numbers do not establish a controlled hosted speedup.
For the same 1,941 decisions, the estimated Jev cost is $0.2589: 7,396,753 local input tokens priced once at djev's announced $0.035/M input tariff. Its billing contract counts physical reads, so repeated context or provider scaffolding can change that estimate. The provider reported free preview on October 1, 2026. This adapter has not been deployed there.
Luna's estimate is $0.7667 with the recorded cache mix, or $1.6112 without caching, using Standard API rates checked on October 1: $0.10/M input, $0.01/M cached input and $0.50/M output. The run used Codex authentication, not a paid API invoice. Both cost estimates exclude retries, training and infrastructure.
Data and reuse
See the dataset card for source families, processing, checks and the public-input schema. The released archives preserve the original screenshots, labels and grouping records. They are dominated by controlled web-template mutations; native-app interactions, natural bugs and visual ambiguity need separate evaluation.
Diagnostic metrics · Category-recall values. These tables contain aggregate results, not per-case predictions or development logs.
Use
Use Linux, Python 3.12/3.13 and a CUDA GPU with at least 40 GB VRAM. The loader downloads the pinned base model automatically.
python -m pip install huggingface_hub
hf download Shelter/UI-testjev --local-dir ui-testjev
cd ui-testjev
python -m pip install torch==2.10.0 torchvision==0.25.0 --index-url https://download.pytorch.org/whl/cu128
python -m pip install -r requirements.lock.txt
python modeling.py --adapter . --input example-regression.json --image-root /path/to/screenshots
Provide reference.png and current.png, then set their viewport in example-regression.json. See Input format for local and keyboard checks. Use the supplied helper and keep the adapter unmerged.
Scope
Full-screen checks assume one introduced defect category. Multi-bug detection, localization and generated explanations are not supported. Natural-bug and native-app coverage remain limited; visual abstention was not trained.
- Downloads last month
- 65
Model tree for Shelter/UI-testjev
Base model
google/diffusiongemma-26B-A4B-it