Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

UI-TestJev

Comparable local UI-check performance to GPT-6 Luna Instant, with lower observed latency and estimated cost.

UI-TestJev uses DiffusionGemma to compare approved and current screenshots, identify UI defects, and return structured decisions for testing pipelines: pass, unknown, or a failure category.

Dataset · Technical report · Use · Input format

What it does

  • Requirement checks: detect clipped text, missing content, low contrast, occlusion, distorted images and undersized controls.
  • Visual regression: compare full screenshots for a newly introduced defect.
  • Keyboard checks: judge reachability from an observed Tab-navigation trace.

The release is a LoRA adapter for DiffusionGemma-26B-A4B-it, using one-step decision inference. Vision weights remain BF16; frozen expert weights use INT8.

Selected results

Fine-tuning improves local-check macro-F1 by 41.70 percentage points and full-screen regression by 19.98 points.

Development-test measure Base DiffusionGemma UI-TestJev GPT-6 Luna Instant*
Local UI checks — macro-F1 31.32% 73.02% 71.18%
Full-screen regression — macro-F1 20.55% 40.53% 50.80%
Weighted task score 31.71% 59.47% 63.50%
Recorded median latency — 1.36 s (local reference) 1.97 s (HTTPS)
Estimated cost / 1,941 decisions — $0.26 $0.77–$1.61

Base DiffusionGemma is evaluated without the adapter, using the same INT8-expert/BF16-vision setup, screenshots and decision protocol. Fine-tuning increases false alarms: 1.77% → 16.31% on local checks and 0.76% → 3.05% on full-screen checks.

Luna is gpt-6-luna with reasoning off. These are development-test results. The weighted score combines local, full-screen and keyboard macro-F1 (40/50/10%). Jev latency is a local reference time on an RTX PRO 6000 Blackwell Server Edition; Luna includes HTTPS overhead. Jev cost uses djev's announced tariff as an estimate, not a quote for hosting this adapter. Measurement and pricing details.

Technical report

Data

The dataset has 1,590 training, 846 validation and 1,941 development-test decisions, plus 311 external Scry pairs. Training uses 50 pages from four template families: Start Bootstrap, Flat UI, Material Dashboard and HTML5 UP. Validation uses Bulma Templates; the development test uses AdminLTE.

Synthetic examples are built as defect/harmless-change pairs. Browser measurements check the requirement; screenshots provide the model evidence. Related variants stay in one split. Local questions include a reference target and detail crops; full-screen questions receive no target-location hint.

The audit checked 6,536 images, removed 120 duplicate rows and excluded 21 ambiguous clipping pairs. No exact image or public-input hashes cross splits. The AdminLTE family informed earlier repairs, so the development test is not an untouched final holdout. Dataset construction and coverage.

Evaluation setup

The main comparison uses 1,941 development-test decisions: 1,127 local checks, 787 full-screen checks and 27 keyboard checks. Base DiffusionGemma and UI-TestJev use the same pinned model, INT8 experts, BF16 vision, processor, prompts, option order and one-step decoding. The base arm has no LoRA adapter. This measures adaptation gains, not the effect of INT8 versus BF16 weights.

GPT-6 Luna uses reasoning off and the same public requirements, ordered images and answer options. It uses its own image processing and autoregressive decoding through Codex HTTPS. All 3,098 Luna cases across validation, development test and Scry returned valid answers with zero reported reasoning tokens.

Training source families are separate from validation and test. The AdminLTE test family was inspected during earlier dataset repair, however, so these are development results. Neither these test scores nor Scry selected the adapter checkpoint.

Detection and false alarms

Defect recall counts a defect as detected when the model predicts any failure category. Precision is the fraction of failure predictions made on actual defects. False-alarm rate is the fraction of passing examples assigned a failure category. A wrong defect category can therefore count as detection while still being an incorrect answer.

Local requirement checks

Metric Base UI-TestJev Luna
Defect precision 92.31% 83.21% 84.90%
Defect recall 21.31% 80.99% 76.91%
False-alarm rate 1.77% 16.31% 13.65%
Exact-answer accuracy 53.06% 82.34% 79.86%
Answers marked unknown 17.30% 0.00% 2.84%

Full-screen regression

Metric Base UI-TestJev Luna
Defect precision 95.08% 92.59% 95.90%
Defect recall 14.76% 38.17% 59.54%
False-alarm rate 0.76% 3.05% 2.54%
Exact-answer accuracy 53.37% 63.15% 70.01%
Answers marked unknown 0.00% 0.00% 0.38%

On local checks, Jev detects 456/563 defects, compared with 120/563 for the base. False alarms increase from 10/564 to 92/564 passing cases. On full-screen checks, detection rises from 58/393 to 150/393, while false alarms rise from 3/394 to 12/394.

The benchmark has roughly equal numbers of defective and passing visual examples. Precision in a production pipeline will depend on how often defects actually occur.

Macro-F1 in the model card averages F1 over the labels present in each task's gold answers. Visual tasks have no gold unknown labels: an unknown prediction counts as a miss for the true class, without adding an extra class to that average. The weighted score uses 40% local, 50% full-screen and 10% keyboard macro-F1. Missing or invalid predictions count as errors; none occurred in these runs.

Which errors remain?

Correct-category recall for base DiffusionGemma, UI-TestJev and Luna

The chart requires the correct defect category, not merely any failure prediction. A dash means the category was not evaluated for that task. Chart values.

  • Control size: local recall is 142/153, but control-size warnings account for 56 of 92 local false alarms. In full-screen checks, recall falls to 61/153; another 22 cases are called missing content.
  • Contrast: local recall is 88/103. Full-screen recall is 6/103, with 91 called pass. The model is much less reliable when it must find the affected region itself.
  • Image aspect: local recall is 2/12 and full-screen recall is 0/12. This remains a weak category, and the sample is small.
  • Category confusion: Jev flags 150 full-screen defects, but only 115 receive the right category. Reporting detection recall alone would hide 35 wrong-category answers.

Local and full-screen tasks differ in their prompts, crops and category mix. Their score gap does not isolate the causal effect of cropping.

Abstention and keyboard checks

Jev gives no unknown answers on either visual test task. The training set has no visual unknown targets, so this does not demonstrate good handling of ambiguous screenshots. Its answer scores are not calibrated probabilities.

Keyboard test accuracy is 24/27 for the base, 27/27 for Jev and 26/27 for Luna. Jev and Luna both recover all nine unknown cases; the base recovers eight. These are trace-based questions across nine scenarios, not evidence of broad keyboard-testing reliability.

Validation and uncertainty

Epoch 3 is the exported checkpoint. Its weighted validation score is 67.26%, versus 39.88% for the base, but no checkpoint met every acceptance rule. In particular, local false alarms rose from 2.07% to 12.03%; the permitted maximum was 4.07%. Full-screen clipping recall also fell from 11/26 to 10/26. A 48-case noise probe changed decisions on 4 cases, versus 1 for the base.

The export is therefore not recommended over the base under the original acceptance rules. It offers higher defect recall with a false-alarm tradeoff that requires review.

For the Luna comparison, 2,000 paired bootstrap resamples keep each page or scenario together. On local checks, Luna minus Jev macro-F1 is -1.83 percentage points, with a 95% interval of [-6.17, +3.03]. On full-screen checks it is +10.27 points, [+4.46, +16.15]. Local results are close; the full-screen difference favors Luna. These intervals describe this corpus, not generalization to new repository families.

External Scry evaluation

Scry contributes 311 reference/implementation pairs from 52 app groups. We ask 12 fixed category-presence questions per pair and score 505 annotated category instances.

Diagnostic Base UI-TestJev Luna
Annotated categories recovered 266/505 337/505 410/505
Annotated-category recall 52.67% 66.73% 81.19%
Additional unannotated predictions 994 1,214 1,856

Jev recovers more annotated categories than the base, but also flags more unannotated categories. Scry does not contain exhaustive negatives, so those extra predictions cannot all be called false positives. This evaluation measures category presence, not localization or the original Scry benchmark score. Scry has been observed during development and is not a fresh blind test.

Training method

Setting Completed run
Base google/diffusiongemma-26B-A4B-it, revision f7f5b7f5fa82ffc52addd066915886d497f5517b
Adapter LoRA rank 32, alpha 64, dropout 0.05; decoder attention and dense MLP projections
Frozen weights Encoder, vision, experts, router, embeddings and output head
Learning rate / epochs 5e-5 / 3; one configuration, seed 3407
Optimizer AdamW, betas (0.9, 0.95), weight decay 0.01, epsilon 1e-8
Batch / schedule Microbatch 1, accumulation 16; 5% warmup, cosine decay to 10% of peak; gradient norm clipped to 1.0
Visual input Original aspect-preserving processor; at most 1,120 visual tokens per image; 8,192 total input tokens; overflow rejected

The objective is cross-entropy over each question's allowed answer tokens. Supervised answer slots are always replaced with random vocabulary tokens before the decoder sees them. Option order changes during training and stays fixed during evaluation. This is a task-specific decision objective, not ordinary next-token chat training.

All 1,590 training rows are visited once per epoch. Source families and pages are interleaved; related examples stay in sampling blocks, reshuffled each epoch. Passing visual examples receive 1.5x loss weight. No minority examples are duplicated to inflate the dataset.

Latency and cost

Jev's median development-test latency is 1.36 seconds, with p95 1.63 seconds, on an RTX PRO 6000 Blackwell Server Edition. This is a local reference time, including preprocessing and model work with a vision-feature cache. Luna's HTTPS client median is 1.97 seconds, p95 3.00 seconds, including request setup and network. The hardware and timing boundaries differ; these numbers do not establish a controlled hosted speedup.

For the same 1,941 decisions, the estimated Jev cost is $0.2589: 7,396,753 local input tokens priced once at djev's announced $0.035/M input tariff. Its billing contract counts physical reads, so repeated context or provider scaffolding can change that estimate. The provider reported free preview on October 1, 2026. This adapter has not been deployed there.

Luna's estimate is $0.7667 with the recorded cache mix, or $1.6112 without caching, using Standard API rates checked on October 1: $0.10/M input, $0.01/M cached input and $0.50/M output. The run used Codex authentication, not a paid API invoice. Both cost estimates exclude retries, training and infrastructure.

Data and reuse

See the dataset card for source families, processing, checks and the public-input schema. The released archives preserve the original screenshots, labels and grouping records. They are dominated by controlled web-template mutations; native-app interactions, natural bugs and visual ambiguity need separate evaluation.

Diagnostic metrics · Category-recall values. These tables contain aggregate results, not per-case predictions or development logs.

Use

Use Linux, Python 3.12/3.13 and a CUDA GPU with at least 40 GB VRAM. The loader downloads the pinned base model automatically.

python -m pip install huggingface_hub
hf download Shelter/UI-testjev --local-dir ui-testjev
cd ui-testjev
python -m pip install torch==2.10.0 torchvision==0.25.0 --index-url https://download.pytorch.org/whl/cu128
python -m pip install -r requirements.lock.txt
python modeling.py --adapter . --input example-regression.json --image-root /path/to/screenshots

Provide reference.png and current.png, then set their viewport in example-regression.json. See Input format for local and keyboard checks. Use the supplied helper and keep the adapter unmerged.

Scope

Full-screen checks assume one introduced defect category. Multi-bug detection, localization and generated explanations are not supported. Natural-bug and native-app coverage remain limited; visual abstention was not trained.

Apache 2.0 · Attribution

Downloads last month
65
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shelter/UI-testjev

Adapter
(8)
this model

Dataset used to train Shelter/UI-testjev