Developers | Where Yui stands | checked Sep 26 2026

Where Yui stands.

Yui gives your own AI agents a native screen on your phone. A compact, stateful language turns their instructions into useful controls, and your taps go back to the agent. This page is the data behind that sentence: what is new, what is not, and what we still have to prove.

9 tokensa whole Tabata timer. 25 as minified JSON, 75 as a component tree.
36% fewertokens than minified JSON on ten real screens: 338 against 530.
21 tokensto reopen all ten once saved, against 338 to send them again.
7 ways inHermes, OpenClaw, webhook, MCP, A2A, AG-UI and a model you run.

The combination is new. The ideas are not.

Semantic UI languages go back to 1999. Compact streaming formats and agent UI standards shipped before Yui. What we did not find in the systems we surveyed is the whole package: an open source iPhone app for agents you already run, a stateful line language that turns one short instruction into a working native control, and every tap sent back to the agent through the adapter it came in on.

Unique here means a combination not found in the named systems we reviewed, as of Sep 26 2026. It is a product and technical comparison, not a patent opinion. Yui's own numbers are first-party and marked that way.

SWOT

Strengths and weaknesses are ours. Opportunities and threats come from the field. Every point carries its evidence and a tag for how sure it is.

StrengthsOurs, and true today

  • A screen people own, not a frame in another app

    Yui is a native iPhone app for agents that run somewhere else. The nearest systems mostly draw UI inside another app or a tool host.

  • One line is a whole experience

    timer 40/20x8 Tabata is 9 tokens and a working interval timer. Ten real screens cost 338 tokens, against 530 as minified JSON.

    Measured[4]
  • Stateful, not a static string

    Each line becomes an op: add, route, patch, save, show, theme, table data. Taps go back as small typed events. Streamed and whole input must parse the same, and five parsers share one test suite of over 600 vectors.

    Shipped[2]
  • Works with the agents you already run

    Hermes, OpenClaw, webhook, MCP, A2A, AG-UI and a model bridge. The AG-UI bridge passed 16 of 16 live checks on Microsoft Agent Framework, the OpenClaw channel 24 of 24.

    Checked live[5][19]
  • Local mechanics stay local

    A timer counts down, a form checks its fields and a choice can change on the phone, with no round trip to the model. A saved screen comes back for 2 or 3 tokens.

    Shipped[2][4]

WeaknessesOurs, and not good enough yet

  • We cannot say first

    Portable, device-adapted UI goes back to UIML in 1999. A2UI and OpenUI stream compact agent UI today. Only the combination is new.

  • Our numbers are our own

    The token benchmark covers ten screens we chose. The channel eval, which scores how often a model's screens parse and make sense, and the speed numbers from real phones are ours too. Nobody outside has reproduced them, and there is no head-to-head with A2UI, OpenUI or json-render.

    First-party[4][21][22]
  • iPhone only

    The beta is live on iPhone. Mac, Watch, Android and the browser are specs or roadmap, not proof.

    Roadmap[1]
  • Breadth is checked one bridge at a time

    The AG-UI bridge is checked on Agent Framework only. CopilotKit, Mastra and Pydantic AI are untested.

    Untested[5]
  • You bring the agent

    Someone with no agent has nothing to talk to yet. The starter agent is designed, not built.

    Draft[20]
  • Only one host checks what a phone can draw

    The phone reports its build, and the Hermes plugin swaps out any preset that build cannot draw. OpenClaw, the webhook, MCP, A2A, AG-UI and the model bridge do not yet, so a phone on an older build shows an Update chip instead of the screen. And the core keeps growing: 38 preset words in the parser today.

    Gap[23]

OpportunitiesWhat the field leaves open

  • The super app a novice could never build

    People who pay for an AI plan but will never write an app can still get one: a timer, a food log, a booking screen, each drawn on the fly and kept. The ten screens below are that day, for 338 tokens.

    Scenario[4]
  • Standards are routes in, not rivals

    AG-UI agents, MCP clients and A2A agents already reach Yui. A2UI waits on demand: if agents that speak it start asking, a one-way translator in the bridge is the cheap step, and the app keeps one format.

  • One build check for every host

    Move the check the Hermes plugin does into the session every host already calls. Then each host learns what the phone can draw, and sends plain words for the rest.

    Proposed[23]
  • Compare formats on getting it right

    Tokens are half the story. The channel eval already scores a model's replies against the real parser. Run the same cases with an A2UI or JSON guide and we learn which format models get right more often, for the cost of a few runs.

    Proposed[21]
  • Measure what a new person feels

    Minutes from install to the first screen from your own agent, and how often agents answer with a screen instead of text. Those two numbers decide whether the super app is real.

    Proposed[24]
  • Hermes still has no screen of its own

    The biggest open agent community (248,460 GitHub stars on Sep 24) has no official phone app and no generative UI. Yui plugs in the way Telegram does.

    Survey[18]

ThreatsWhat could close the gap

  • A2UI has Google and momentum

    Streaming JSONL, trusted client components, incremental updates and renderers across web and mobile. OpenClaw's gateway already speaks it.

  • OpenUI sells compact too

    A streaming-first compact language that claims token savings of its own. Nobody has run the two side by side.

    Field[13]
  • Native feel is not ours alone

    A2UI on Flutter and json-render on React Native reach phones with native-feeling UI.

  • Hosts may become the canvas

    MCP Apps let a tool draw UI inside chat hosts like Claude and ChatGPT. If people stay there, a separate app has to earn its place.

    Field[17]
  • Invention is not a moat

    A bigger team can ship the same combination. What lasts is measured task performance and reach, and our lead on native rendering is months, not years.

    Verdict[18]

The app you wish you could build

Picture someone new to AI. They want a workout timer, a food log and a way to book client calls. Today that is three apps, or a developer. In Yui they ask their agent, and it draws each piece on the fly from a fixed set of native presets. Nothing to design, host or install. The ten screens in our benchmark are that person's day. Tap one.

Training
Food
Work
First run
timer 40/20x8 Tabata
Yui Lines
9
Minified JSON
25
Component tree
75
Reopen once saved
3

Tokens, o200k_base. The screen on the right is drawn live from the lines above.

Drawing the screen...
338 tokensto draw all ten screens, three apps' worth
967 tokenssaved against a component tree (1,305)
21 tokensto bring all ten back once they are saved

What this takes today. You bring the agent: Hermes, OpenClaw, Claude Code, Cursor, a model you run, or anything behind a webhook. The Claude and ChatGPT connectors are built and each waits on one check in its own app.

A starter agent for people with no agent yet is designed, not built. Agent tables, where rows like your macros stay on the phone, are a spec and a playground mock; the app's store comes next.

How a line becomes a screen

Four layers, each with one job. The network can change underneath without changing the language on top.

  1. Connectivity

    Moves messages and tracks turns: HTTP, streams, the relay, adapters.

    Hermes plugin, OpenClaw channel, webhook, MCP, A2A and AG-UI bridges.

  2. Yui Lines

    Describes controls and operations: add, route, patch, save, restore, theme, data.

    One line parses to one op. A bad line is isolated.

  3. Native runtime

    Maps ops to screen state and runs local mechanics.

    A SwiftUI app and a Swift parser. iPhone beta.

  4. Return path

    Reports what the person did back to the agent.

    Small events such as {id, preset, answer}. The bridge decides the mapping.

How strong is the proof?

Six claims we could make, scored against the evidence we can show.

ClaimStrengthEvidenceLimit
A native agent screen people ownStrongA shipped iPhone app works with outside agents. [1][3][11][17]Others mostly embed UI in another app or host. A survey of named systems, not a census.
One line compresses a whole controlMeasured338 against 530 tokens on ten screens. The timer is 9 against 25. [2][4][13]Ten screens we chose. OpenUI chases the same idea.
Models get it rightMeasuredThe channel eval scores every guide change against the real parser. Opus 5.5 scored 94 to 100% on the recent full runs, Sonnet 5 88% on an earlier guide. [21]Our cases, our guide, two models. Not yet run for other formats.
Stateful, streaming interactionDocumentedOps, patches, routing, saved screens, events and a conformance suite all exist. [2][5][12][14]A2UI and its peers also stream and keep state.
Works with the agents you runStrong todayHermes, OpenClaw, webhook and MCP paths. AG-UI passed 16 of 16 on Agent Framework. [1][3][5]Other frameworks are untested.
Native everywhereNot yetThe iPhone beta is live. [1][2][3]Mac, Watch, Android and A2UI interchange stay roadmap until they ship.

Verdict: the broad idea has clear prior art, and several single tactics have direct competitors. What we can defend is the integrated product. The moat worth building is shown task performance and reach, not a claim to have invented generative UI.

The token numbers, and their limits

Ten representative screens, each written four ways. The JSON is built from the parsed lines, drops defaults and uses short keys. Tokens counted with o200k_base.

What follows

  • On these ten screens, Yui Lines uses 36% fewer output tokens than minified semantic JSON (192 of 530).
  • The timer is 9 tokens, against 25 for minified JSON and 75 for a component tree.
  • Reopening the ten saved screens takes 21 tokens of show commands, against 338 to send them again, if they are already on the phone.

What does not follow

  • It is not a head-to-head with A2UI, OpenUI or json-render on the same tasks.
  • It does not measure speed, render quality or cost. How often models get a screen right is a separate test, the channel eval.
  • Defaults and richer renderers are part of the saving. A different set of screens can change the ratio.
  • The Claude column in our table uses Anthropic's legacy tokenizer, so treat it as a guide.

The next test, open to anyone. Take the channel eval's cases, write an equal guide for A2UI and for plain JSON, and score every reply for validity and tokens with the same models. That tells us which format models get right more often, not only which is shorter. A full study on iPhone across formats (time to the first usable control, correction turns, completion) comes after, if the cheap one says it matters.

Full method in the benchmark spec. Run it yourself in the playground, or help build the next one.

27 years of prior art

Yui is a new arrangement on top of a long history of abstract, portable UI and a fast-moving agent UI field. Dates are papers or public launches.

  1. 1999

    UIML

    An XML UI language that splits one description of a UI from the device it runs on. [6]

  2. 2002

    Personal Universal Controller

    A handheld reads an appliance's spec and generates a usable remote for it. [7]

  3. 2004

    SUPPLE

    Renders one abstract spec for each device by optimization. Device-adapted UI, long before LLMs. [8]

  4. 2023

    Vercel v0

    Launched in October. Generated interfaces become a mainstream developer category. [9]

  5. 2024

    Vercel AI SDK 3.0

    Open source generative UI in the SDK, March 2024. [10]

  6. 2025

    Google A2UI

    An open, cross-platform agent UI format with trusted client components and incremental updates, Dec 15. [11]

  7. 2026

    A2UI 0.9 and OpenUI

    A2UI widens its renderers and transports in April. OpenUI ships a compact, streaming-first UI language. [12][13]

  8. 2026

    Yui alpha

    A native iPhone app for agents you already run, open source under Apache-2.0. The MVP was done Sep 26, and the alpha is open to anyone on TestFlight. Latest release: Yui 0.6.0, build 342, Sep 29. [1]

So we never say first semantic UI, first compact streaming format or first cross-platform agent UI. The claim is the combination.

The nearest systems

A narrow set on purpose. Each solves a related problem from a different starting point. Where it draws is each project's documented focus, not a hard limit.

SystemWhat the agent sendsWhere it drawsHow it relates to Yui
A2UI (Google)Streaming JSONL against a client catalogWeb and mobile renderersClosest on agent to renderer ideas. It assembles lower-level components; a Yui preset like timer is a whole experience. [11][12]
OpenUI (Thesys)A compact, streaming-first UI languageRegistered components in your own appThe strongest answer to a compact-language moat. Built into existing apps, not a personal agent app. [13]
json-render (Vercel Labs)A JSON spec against a typed catalog, with state and actionsSeveral renderers, React Native includedStrong developer-controlled composition and portability. [14]
TamboThe agent picks React components and streams propsAn existing React appA framework for adding generative UI to your app, not a consumer agent surface. [15]
CopilotKit and AG-UIAn agent frontend SDK and a two-way interaction protocolYour appAG-UI can carry several UI specs. Yui bridges it: a screen is a tool call, the tap comes back as the result. [5][16]
MCP AppsUI resources attached to toolsInside a chat hostA tool shows UI in the host. Yui's MCP server lets tools send screens to the phone too. [3][17]
YuiOne short line per control, parsed to opsA native iPhone app, for agents you already runThe high-level preset plus one personal app for every agent. [1][2][3]

Two corrections to our own earlier pitch: A2UI is streaming JSONL too, and OpenUI also claims token savings. Native-feeling mobile UI exists in A2UI on Flutter and in React Native renderers. The difference is the preset and the app, not the broad techniques.

Six ideas from an outside review, and our call

An outside review suggested six moves. It read our public pages, not the code, so some already exist. Here is what we are doing with each, and why.

  1. Publish an op contract

    Mostly done

    The spec already defines ops, events, errors and versioning, and over 600 shared vectors hold five parsers to the same ops. Events get vectors of their own when a second app, likely Android, starts. [2]

  2. Negotiate capabilities

    Doing, for every host

    The phone reports its build, and the Hermes plugin swaps out what that build cannot draw. The other hosts need the same check, best served from the session they already call. Camera and mic permission stay the phone's job. [23]

  3. Third-party preset packs

    No

    A fixed set of native presets is why Yui passes App Store review, why every screen looks right and why lines stay tiny. New presets come from the flywheel: custom screens agents keep sending get promoted. Saved flows already share a composition. [25]

  4. An A2UI to Yui Lines adapter

    Not now

    Decided Sep 25 in the AG-UI spec: a one-way translator in the bridge, built when agents that speak A2UI ask. None have yet. [5]

  5. A task-level study across formats

    Cheap version first

    How often models get it right (the channel eval) and speed on real phones (the speed budget) are measured already. Next is the same eval across formats. The full study on iPhone only if that says it matters. [21][22]

  6. Local mechanics and an inspector

    Already how it works

    Timers, answer changes and locks run on the phone. The playground shows every op and error in its wire log, and the war room has a speed panel. An inspector inside the app can wait. [2][22]

The two we are taking on, one build check for every host and the eval across formats, are proposed as the next cards. When they are picked, they show on the roadmap and the board.

What we say, and what we do not

“Yui gives your own AI agents a native screen on your phone. A compact, stateful language turns their instructions into useful controls, and your taps go back to the agent.”

Sources

Public papers, product docs and repositories, read on Sep 26 2026. Product pages change. Yui's beta and bridge checks are reported by us and not reproduced by anyone else yet. A feature missing from this review does not prove it is missing elsewhere.

  1. Yui product site: beta status, use case and first-party claims
  2. Yui Lines v0 spec: ops, presets, events, streaming and conformance
  3. Yui iOS app repository: README and source tree
  4. Yui ten-screen token benchmark and method
  5. Yui AG-UI bridge spec and its recorded live checks
  6. Abrams et al., UIML (1999), original paper
  7. CMU Personal Universal Controller specification (2002)
  8. Gajos and Weld, SUPPLE (2004), original paper
  9. Vercel, announcing v0 (Oct 11 2023)
  10. Vercel, AI SDK 3.0 generative UI (2024)
  11. Google, introducing A2UI (Dec 15 2025)
  12. Google, A2UI v0.9 (Apr 17 2026)
  13. OpenUI README: a streaming-first compact language
  14. Vercel Labs, json-render README
  15. Tambo README: a React component registry
  16. AG-UI: generative UI specs and the protocol line
  17. MCP Apps specification
  18. Yui competitor research notes (BIZ-1), Sep 24 2026
  19. Yui adapters spec: every bridge and its live checks
  20. Yui starter agent spec (draft)
  21. Yui channel guide and its eval scores
  22. Yui speed budget and numbers from real phones
  23. Yui relay spec: min builds and the Update chip
  24. Yui marketing and positioning brief (BIZ-1): the metrics we watch
  25. Yui preset flywheel: custom screens that become presets

Try Yui, or help build it

Get the alpha

The MVP is done and Yui is in alpha, open to anyone with an iPhone on iOS 26. Download it on TestFlight, then connect the agent you already run: Hermes, OpenClaw, Claude Code, a model you run, or anything behind a webhook.

Star it on GitHub

Yui is open source under Apache 2.0. Star the repo, open an issue, or send a pull request.

Lend your agent

Spare tokens on Claude or ChatGPT Codex? Your agent can pick a card off our backlog and open a pull request. Yui@home, like SETI@home.

Want a hand getting in?

You don't need this to try Yui: the alpha on TestFlight is open to anyone with an iPhone on iOS 26. Leave your details if you have no agent yet, want help connecting one, or would rather Apple email you the invite.

  1. Yui emails you a link to confirm your address.
  2. We read your request, and reach out if you asked for help.
  3. Apple emails you a TestFlight invite.
  4. Open it on your iPhone, install Yui, and sign in with Apple.

We use this only to get you into Yui. Privacy.