Cross-format generative UI benchmark
A benchmark that runs four generative-UI formats against their own published prompts and their own validators, with auto-repair switched off for every format including MDMA. 1107 generations across three model tiers, every raw generation committed to the repo.
- Feature
MDMA holds its rate from flagship down to open weights
On the 'renders every time' measure MDMA scores 94.4% on both Opus 5 and the open-weights Gemma-4-26B — a 0.0pp drop between tiers. OpenUI Lang drops 27.8pp over the same two models and json-render drops 44.4pp.
- Feature
Each format judged by its own validator
Every format was given its own published system prompt and scored with its own validator, on a byte-identical user message with no format hints. The report discloses where a format's own validator disagrees with the harness.
- Feature
Prompt cost published alongside reliability
System-prompt size is part of the bill: OpenUI Lang 5172 tokens, MDMA 5910, json-render 8466, A2UI 19689. Mean output tokens and a renderable-per-1k-tokens efficiency figure are reported per model.