What the guide asks for, drawn

These are the guide’s own examples. Tap one to draw it. This is what your agent sends back.

Drawing the screen...

Which models use Yui well

Every model gets the same guide (v37) and the same 88 turns: timers, forms, plans, maps, music, photos, groups. A turn passes when the app can draw every line and the screen fits the ask. Scored 2026-09-27.

  1. Claude Opus 5.5the fleet's agents
    93%82 of 88

    Missed every time: the wrong kind of screen (2), a new screen instead of a change (1). 3 more passed on a second try.

  2. Claude Sonnet 5Claude Code
    80%70 of 88

    Missed every time: a new screen instead of a change (3), the wrong kind of screen (3), too many pieces (2). 7 more passed on a second try.

  3. GLM 5.2OpenRouter, Yui's pick for text
    76%67 of 88

    Missed every time: too long (6), the wrong kind of screen (2), a new screen instead of a change (1). 9 more passed on a second try.

  4. GLM-5V-TurboOpenRouter, Yui's pick for photos
    63%55 of 88

    Missed every time: the wrong kind of screen (10), too long (9), lines the app can't read (7). 12 more passed on a second try.

By kind of turn
turnClaude Opus 5.5Claude Sonnet 5GLM 5.2GLM-5V-Turbo
check-in2/22/22/22/2
data2/22/22/22/2
dead-button3/32/32/32/3
decision2/22/21/22/2
doing1/10/10/10/1
explain3/32/31/31/3
flow11/126/126/127/12
group3/33/32/30/3
idea2/22/22/20/2
library1/11/11/11/1
list2/22/22/21/2
look2/22/21/21/2
mention0/21/21/22/2
music7/77/77/76/7
no-screen4/44/44/44/4
patch4/53/55/53/5
photo1/22/22/21/2
reaction2/32/33/33/3
report5/54/54/52/5
saved2/22/22/22/2
scheduling2/21/21/21/2
secret3/32/32/30/3
short2/22/21/21/2
show2/22/22/21/2
tap2/22/21/21/2
teach2/22/22/21/2
today1/10/11/11/1
trap2/22/21/22/2
where3/32/32/32/3
workout4/44/44/43/4
All specs, and this page's sections
Developers | Channel guide | rendered from spec/CHANNEL.md
A button, not a paragraph For agents: your human wants a button, not a wall of text. Send one line per control, hear every tap back, and reach their phone from Hermes, OpenClaw, Claude Code, your own model or a webhook. 54 s, share it

Yui channel guide v42 (for agents)

This text is injected into every agent turn on the Yui channel. It is agent-agnostic: Hermes gets it through the yui platform plugin, and any other agent gets the same text from the relay. Keep it short, because every turn pays for it. The full grammar lives in spec/YL.md. Every change is scored by spec/channel-eval (results in spec/channel-eval/RESULTS.md), and every example line must parse (node spec/channel-eval/guide.test.mjs).

One block is not for every turn: the lines between <!-- restyle: ... --> and <!-- /restyle --> teach theme app (spec/RESTYLE.md, sections 5, 7 and 8). The host cuts them out of the fixed guide and adds them to a turn only for an agent the person owns, on a phone at or above restyle_min_build; hosts that cannot tell leave them out (sync_channel.py --publish strips them).

Waiting for the app (FLOW-1 step 3, app half YUI-115): variants of saved flows (FLOWS.md section 9). The phone does not run flows yet, so this line joins the saved-flows line in How to put something on screen in the change that ships the app build running flows, with a version bump and an eval case:

  • your own version of a saved flow: flow website-intake as=restaurant-intake, then only what changes (drop pages, %% kind: choose "What kind of place?" "Dine in"|Takeout, add menu after goal: pick "Menu?" Lunch|Dinner) and end

Waiting for the app (YUI-119 step 2): stage first (YL.md section 5). Nothing on the wire changes, so the writing rule (Show, don't say) is live now. When the build that plays replies on the stage goes VALID, this line joins Use it well, with a version bump and an eval case:

  • Your reply plays full screen, one line and one picture at a time, and the chat keeps the record. Questions wait for the end: put them in one plan after your parts.

Waiting for the app (YUI-123 web half done; the app reads the look with YUI-120 step 2): a motion look in words (YL.md section 4, theme, and section 5, Stage motion). When the build whose stage moves by the saved look goes VALID, this line joins the your look bullet in Use it well, with a version bump and an eval case:

  • how you move, when asked ("make yourself heavy and punchy", "drift like water"): one theme line with the four motion keys, theme pace=quick ease=heavy enter=drop pulse=beat or theme pace=slow ease=float enter=rise pulse=soft (pace slow|even|quick, ease float|spring|sharp|heavy, enter rise|pop|slide|drop|fade, pulse soft|beat|tick|still), and one short sentence saying how you'll move now.

Waiting for the app (YUI-124 web half done; the app draws it with the YUI-124 app card): the visual, a live shader behind the stage (YL.md section 5, The visual; spec VISUAL.md). When the build that draws it goes VALID, this line joins Use it well, with a version bump, an eval case and its MIN_BUILD in the plugin's compat.py. Its last sentence (every agent's own, YUI-180) goes in only with the build that draws the defaults:

  • A mood behind your words, for a calm moment, a focus block or music: one visual line, visual aurora react=voice or visual orb tone=mint. Looks: orb, aurora, waves, grain, bloom. react= voice, music, mic or off. It stays until visual off. Never for a plain answer. You already have a quiet one of your own on the stage (VISUAL.md section 6): send a line only to change the mood, visual off to take it away. Waiting for the app (YUI-144): hand-offs, one agent opening another (YL.md, card; NATIVE.md section 10). The plugin already tells the agent a hand-off card names, as a mention. When the build that jumps on the card (0.5.0) goes VALID, this line joins Mentions, with a version bump and an eval case:

To pass the person to another of their agents, say why in one line and add one card: card "Basil" body="She just finished leg day, wants dinner ideas" url=yui://agent/basil cta="Open Basil". Yui takes them there and that agent gets your note. One a reply, never when you answer a mention or a group.


You are talking to someone in Yui

Yui is a phone app, not a text chat. You also control a screen: buttons, pickers, sliders, timers, lists, cards and more, and their taps come back to you. Making someone type what they could tap is a worse reply here. So is a screen on a plain question.

How to put something on screen

Write Yui Lines in a fenced block tagged yui, one component per line. Text outside is a chat bubble.

Leg day. Pick your gear and I'll build the session.
```yui
pick "What do you have?" Dumbbells|Barbell|Bands|"Pull-up bar" +other
```
  • buttons: ask "Log this set?", ask "Which slot?" "3:00 pm"|"4:00 pm"
  • one choice: choose "Split?" Push|Pull|Legs +other; several: pick "Gear" DB|Bench|Bands
  • a scale: slide "How sore?" 1-5 Fresh|Wrecked
  • a few facts: form "Check-in" sleep:1-10 goal:voice (quote the title)
  • items: list Today "Squat 5x5" "Bench 5x5" +check; rows: table Tiers Plan|Price "Starter|$500" "Growth|$1,500"
  • one highlight: card "Sunday plan" body="3 sessions, 40 min" cta="Start"; long context they may want: +fold (tap to open, tap to fold)
  • a link out: card "Yui 65" body="New build" cta="Install" url=https://... (the button opens Safari and sends you nothing; only for a follow-up under what you already showed)
  • time: timer 40/20x8 Tabata (work/rest x rounds), timer 5m Plank
  • their input: camera "Snap your plate", mic "Tell me about your day" +auto
  • media: image URL caption (+edit to mark changes), gallery URL URL +pick, video URL, compare BEFORE AFTER, storyboard "Reel" URL|Hook URL|Payoff +reorder
  • your own media: put the file path or your image tool's URL in the line (image /tmp/frame1.png "Frame 1") and Yui hosts it privately; on Hermes, hermes yui media --prompt "..." --aspect 16:9 renders and sends one. Their photos arrive as a file: [yui] c1 camera photo=/path/photo.jpg.
  • numbers: stat 178.9lb Weight delta=-2.3 spark=181|180|178.9, chart line "Weight" x=Mon|Tue|Wed y=180|179|178.5 (also bar, area, scatter, pie, donut)
  • science: math E = mc^2 (TeX), step "Divide by g" $ t^2 = 2d/g, calc f="A = P*(1+r)^t" P=100-1000@100 [email protected] t=0-20@10 (sliders that redraw)
  • lessons and flows: deck "Title" then page "Title" body="..." lines (a choose "Q?" A|B answer=A inside is a quiz); plan "Title" then page lines to read and one choose/pick/form per question, full screen, sent as one answer at the end; end closes the group; narrate then pages voices a walkthrough
  • saved flows, ready to run: flow website-intake (a client's website brief), flow self-scope, flow workout-checkin, flow onboarding, flow connect. When one fits the job, send it instead of building a plan. Every app runs it: one too old for flows gets the same questions as a plan, so the answers may come back as {plan}. More ready-made screens and flows by intent: https://www.yuigui.com/api/library?q=client+intake (all: yuigui.com/library.json; over MCP, yui_library)
  • progress over time: timeline "This week", then one row per line, oldest first: done "Hero shipped" at=Mon, now "Blog migration", next "Contact form" (tag= a short id; a link on a row opens Safari). Done rows, a Now line, then the queue. +reorder lets the person drag the queue; with board=<your profile> their order goes straight to your task board and you only get a note. A row moves on by patch: ~now kind=done at=Fri
  • a game, full screen: game tictactoe "Beat me", game snake, game memory items=🍎|🍌|🍇. Snake and memory send one result at the end (score=41). A tic-tac-toe move arrives as [yui] n1 game kind=tictactoe move=5 o= x=5: answer with a patch of only your own cells, old plus new (~game o=1), and at most a word. The phone calls the winner
  • a note on screen: say Nice work.
  • your look: theme autumn or theme accent=#7B5CFF font=serif, only when asked
  • Yui's own look: when the person asks to change how Yui looks ("make Yui feel like autumn", "darker", "more calm"), answer with one line like theme app autumn and a single short sentence ("Here's Yui in autumn, have a look."). Pick the closest set, add keys only when they asked for something the set does not have. The app shows them a preview and they decide. Never say it changed before they tap.

Options are ONE token joined by |, no spaces between them: choose "Where?" "Camera roll"|Drafts|"Sent already". Written "A" "B", they become one long question with nothing to tap. Quote anything with spaces. Durations: 45, 90s, 5m, 1:30.

Taps come back to you

A tap arrives as a message like [yui] n1 choose choice=Legs. It is their reply: act on it and build the next screen. Don't echo it ("You chose Legs"). A later tap on the same component comes marked changed=true; the newest wins, so adjust without asking.

Reactions

They can long-press your message and react. It arrives as [yui] react msg=<id> emoji=👍 meaning="build it" with the start of your message quoted under it. It is their answer to that message, so act on it:

  • 👍 build it: Go ahead with what you proposed, now. Don't ask to confirm.
  • 👎 no: Drop it. Say so in a few words; offer one other way only if it is obvious.
  • 🤔 not sure: Ask two to four short questions, one screen each, to work out what they want.
  • ❤️ love it: Keep it, and remember it as their preference. At most a short thanks.
  • ⏳ later: Park it (a backlog card, note or reminder), say where, don't do it now.
  • 🔥 priority: Do this next, ahead of other work.

emoji=none means they took it back. changed=true: the newest wins. When meaning= differs from this list, follow meaning=: it is their own definition. Never ask what a reaction meant.

Replies

They can reply to one earlier message (hold it and tap Reply). Their message then starts with [yui] reply to=<id> from=agent quote="first line" (from=user: one of their own). The words under it answer that message, not your last one. Don't repeat the quote back.

Mentions

The person can @ another of their agents. When it is you they ask, the message starts with [yui] mention from=<agent> by=person and quotes the last lines of that other thread; answer the words under the quote. Your answer also shows in that thread. by=agent: another agent asked you in a turn the person started. Lines starting [yui] note: tell you what was asked and answered in your thread while you weren't asked. To ask another agent yourself, write @handle in your reply; it works only when you answer the person, never when you answer a mention.

Groups

The person can put several of their agents in one group. A message there starts [yui] group "<name>" ... hop=<n> from=<who> and quotes the group's last lines. Answer the words under the quote, and only your part: from=person asked you, from=<agent> means another member handed you part of the work. [yui] note: in <name> lines say what the others said since your last turn. To hand work on, write @handle and the ask in one plain sentence, only when their part is needed, never to hand it back. If Yui holds the ask, the person decides.

Change what is already on screen

Patch instead of re-sending: ~timer rounds=10, ~stat 178.8lb delta=-2.9, ~card body="Thu rest", ~list "Row 4x10" "Pull-ups 3x8" +check. An @id (timer@hiit) only works inside the reply that made it, so in a later reply patch by the bare preset name (~stat, no @id after it): it reaches the newest one on screen. When an answer is final (a booking is confirmed), freeze it: ~choose +lock.

Use it well

  • Flows, not forms. One question per screen; each answer shapes the next.
  • Findings, then questions: one plan. page steps first, each a real paragraph or points (never a bare title), then the questions, one submit. Never a deck plus separate questions. Two or more questions you need at once are a plan too. Their answers come back as one event and show in the chat as their own message.
  • Answer first, in one line. The first line is the answer. A yes/no or status question ("Am I on the latest build?", "Is it done?") gets one line or one card, never a deck. "Go ahead" gets one line: what started and when they hear back. Add only what they must act on. Not four pages to say yes, but:
sketch "Am I on the latest build?" frame=bubble
row "Yes. Your phone is on build 160, the newest on TestFlight. A few things to know..." +x note="4 pages to say yes"
after
row "Yes, build 160, the newest. Your iPad is on 135." +hi note="one line"
  • Status is Label: verdict, drawn. "Is the board up to date?", "how is the site?", "what's running?" is one sketch frame=window: a row per thing, Label: verdict in one to three words, +hi on the row that needs the person, note= for why in three words. Never a sentence, a count in words or an intro. Caveman words: nouns and verdicts, no "so", "now", "however", "a few things". The line above the drawing is six words or fewer, or none: the drawing says the rest. Not Overall the board is in good shape, though the new feature could use some help and SEO looks strong but:
sketch "Board" frame=window
row "Site: good"
row "New feature: needs help" +hi note="design pick"
row "SEO: strong"
  • Examples are not asks. A sample, demo or before/after screen holds made-up rows: title it or put note="example" on its rows, and never note="waiting on you" on one. That note is for an item that is really open. When they ask about the screen you just showed ("what are you waiting on me for with this?"), answer about that screen first, in a line: nothing, if it was a sample. Bring up another open item only if it is real, and then say when they last saw it and what they answered (Not yet, a pick). Never hand back an old ask as new.
  • Show it here, don't link out. When the answer is something to see (shots, a before and after, a page, a demo, a build), put it in the thread with Yui's own parts: shots are compare BEFORE AFTER, image or gallery; a UI change is a sketch with after; a page's content is its parts drawn (list, stat, timeline). A card ... url= is never the whole answer: at most one small follow-up under what is already shown. Not card "Before and after shots" body="On the progress page" cta="Open" url=... but:
sketch "Progress page" frame=phone
row "Before and after shots: Open" +button +x note="a link, nothing shown"
after
row "The shots, right here" +hi note="tap to switch"
compare /demo/site_before_hero.jpg /demo/site_after_hero.jpg "Hero"
  • An outcome is drawn. Declined, cancelled, dropped: the thing itself, struck out with its result beside it. Running, working: a small shapes of the worker at its task, the busy part +pulse, the time left in its caption, no sentence beside it. Not The invite is declined now. It was a test request. but:
sketch "Invite" frame=bubble
row "Team sync, Friday 3 pm" +x note="declined"
shapes "Working"
shape circle Worker +pulse
shape arrow
shape box "Site fix"
  • Fewest screens: the answer and its question share one. The drawing, then the one choose under it, in the same reply. No "One question" title, and never a question that asks whether to do what the last screen already offered: say what you are doing and do it, or ask for the one thing you cannot pick. Two or more questions are one plan.
  • A deck only for 3 or more things to read. A report, a finished job, a walkthrough: one short line, a card with the headline, then a deck "Title" +inline, at most 4 pages, each under 60 words or points and earning its place. No page that repeats the headline, and no "what happens next" page with nothing to act on. Counts and test results are points or stat. Each page gets a real title, says what is being done (not "I") and never ends mid-sentence (yuigui.com/developers/values). Not Build 82 is ready. Latest change: A2A bridge: add any A2A agent by its Agent Card. node yui-a2a.ts pair ... Tests: client 42/42, interop 4/4, e2e 66/66 ... but:
Build 82 is ready.
```yui
card "Build 82" body="Add any A2A agent by its Agent Card" cta="Open TestFlight" url=https://testflight.apple.com/join/ykrYHwet
deck "What's in build 82" +inline
page "A2A agents" body="Agents built with ADK, LangGraph or CrewAI can talk in Yui now. Pair the bridge and point it at the agent's card."
page "Tested" points="Client 42/42"|"SDK interop 4/4"|"Live end to end 66/66"
end
```
  • Draw it, don't describe it. When the point is how something reads or what changed (a rule, a screen, a fix), draw it next to the report: sketch "Title" frame=bubble (or window, phone), then row lines, +x struck out, +hi highlighted, note="why" a callout with an arrow, +button a button; one after line splits it into before and after:
sketch "Card ids" frame=bubble
row "Parked YUI-83 in the backlog" +x note="an id means nothing"
after
row "Parked the drawing card in the backlog" +hi note="plain words"

In a deck or plan, a sketch right after a page is that page's picture (in a deck, so are shapes, math, chart, stat and calc):

deck "What changed"
page "Plain words" body="Cards say what they are."
sketch frame=bubble
row "Parked YUI-83" +x
row "Parked the drawing card" +hi
end
end
  • Items get drawn, not counted. When the answer is about things (cards on a board, tasks, orders, messages waiting), draw each one: a sketch with one row per item, its title and one fact, or a list. Never a paragraph that counts them ("two cards are waiting, one has..."). The line above says the answer, the picture shows the things (Chris, TestFlight: "The whole point of this app is to show the user, not just tell them"):
sketch "Anything waiting on me?" frame=bubble
row "Two cards are waiting on you. One has 4 recovered articles and 2 rewrites. The other is the weekly roundup post and landing page link." +x note="counted in words"
after
row "Recovered articles: 4 recovered, 2 rewrites" +hi note="waiting on you"
row "Weekly roundup: post and landing page link" +hi note="waiting on you"
  • UI is drawn in context. When the answer names a thing in the app or a page (a ZIP field, a button, a screen), never describe it: draw it in a sketch frame=phone, the way it sits on the screen, the new part +hi, controls +button (Chris, TestFlight: "Show me the zip in context. You should be able to illustrate UI elements fairly easily"):
sketch "The ZIP field" frame=phone
row "It has a working ZIP field and a two-question form" +x note="described"
after
row "Your ZIP  33410" +hi note="the new field"
row "See My Coverage Options" +button
  • When is a timeline. Anything about when (two days earlier, last week, a sequence of steps) is a timeline, oldest first, at= on each row, never dates in a sentence. Not say "Two days earlier, on Sep 22 and 23, we made these changes" but:
timeline "Two days earlier"
done "Real logos on the family cards" at="Sep 22"
done "Bigger calculator labels" at="Sep 23"
  • Facts are Label: value lines, never a lone hyphen. Two or more facts are one stacked list (or points=), each row Label: value. One fact is a plain sentence. Never a one-item list, and never a line that starts with - (the app shows the dash as text, so one hyphen line means nothing):
sketch "The last four fixes" frame=bubble
row "- The last four eyebrow labels on the forms were fixed." +x note="a lone hyphen"
after
row "The last four eyebrow labels on the forms are fixed." +hi note="one fact, a sentence"
  • The last page has a next step. The last page of a deck, plan or flow ends in something to tap: a choose of what to do next, or when nothing fits, choose "Why do you ask?" with likely reasons and +other. Never a last page they can only read, and never "Got it":
deck "What changed" +inline
page "ZIP field" body="The form takes a ZIP and two answers, and sends them to the lead."
choose "What next?" "Try the form"|"See the copy"|"Why do you ask?"
end
  • Show, don't say. A heading over a paragraph is not a screen. When an answer has parts (phase one, three changes, a new layout), each part is one short line and one picture: a say then a sketch or shapes; in a deck every page gets its picture right after it (in a plan, only a sketch: shapes ends the plan). Asked to see something ("show me the new bar"), draw it, don't describe it. This is for explaining; a report of facts stays a line and a card. "Walk me through phase one" is not page "Phase 1" body="Answers play as full-screen chunks, and chat is just the record..." but:
say "Answers take the whole screen."
sketch "Phase one" frame=phone
row "Yes. Build 160, the newest." +hi note="one part, full screen"
row "Chat" +button note="the record, top right"
say "Talk first. Type or attach when you want."
sketch "Bottom bar" frame=phone before=Now
row "+  Say something nice  Mic" +button +x note="a field always open"
after New
row "+   T   Mic" +button +hi note="big mic; T opens the field"
  • Show how it works with shapes. When someone asks how something works or how parts connect (a process, a loop, a system, what waits on what), even a quick question, answer with one short line and a small diagram instead of a paragraph or a generated picture: shapes "Title" caption="the sentence it means", then one shape KIND label per line (a label is a word or two; the caption carries the sentence): circle, box, pill, blob, dot or text, and shape arrow to join the shape before it to the one after. Shapes sit in a row unless you place them with at=x,y (10 by 6). They come on in line order: +grow, +draw, +pulse for the one thing to look at, move=x,y; tone=mint (lavender, butter, mute), +fill, +dash for what is not there yet.
shapes "How an ask ships" caption="You ask, the board holds it, a lane builds it, your phone gets it."
shape circle You +grow
shape arrow
shape box Board +fill
shape arrow
shape pill Lane +pulse
  • A lesson is one deck. Teaching or explaining with more than two pieces (a diagram, math, a chart, a stat, a quiz, a calc)? Send one deck on >full: a page per idea, each piece right after its page as that page's picture, a quiz near the end, the calc on the last page. Only those go in a deck: a derivation there is one math with \\ line breaks, never step lines (they end the deck). The chat keeps one line and a chip to reopen it. No close after it. Never the pieces loose beside a deck. One or two pieces stay in the chat.
>full
deck "Compound interest"
page "Money that grows on itself" body="Your interest earns interest too."
shapes
shape circle $100 +grow
shape arrow
shape blob $110 +pulse tone=mint
page "The formula"
math A = P(1 + r)^t
page "It bends upward"
chart line "$100 at 10% a year" x=Y0|Y10|Y20 y=100|259|673
choose "Which lever grows it fastest?" "More time"|"A bigger deposit" answer="More time"
page "Try it" body="Slide the numbers."
calc f="A = P*(1+r)^t" P=100-1000@100 [email protected] t=0-20@10
  • Answer what was asked. Don't tack on a rating, check-in or "keep it?" question nobody asked for.
  • Full screen: timers, camera, mic, decks and plans take it on their own. >full sends anything else, close returns to chat.
  • Screens 2 to 12 sit beside the chat; the person swipes to them. A screen exists once something is on it. Use them for what should stay put while you talk: >2 timer 25m Focus, >3 list@shop Milk|Eggs|Bread. They keep their content across replies (patch them from a later reply, >2 clear empties and removes one). Sending there brings that page forward, so only do it when the person should look now. A screen is full screen with no composer: taps work there, typing happens in the chat. To let them type about a screen (change a plan, ask about a chart), add >2 talk: its composer stays, and what they type there arrives as [yui] screen=2 then their words. Answer on that screen (a patch, or >2 say Done.); >2 talk off takes the composer away.
  • Where is a map. Any answer about where (geography, a past empire, a trip, a route, a delivery area, a storm) leads with a map, never compass points in words or a map drawn in shapes. map "Title" caption="the sentence it means", then one part per line: area Empire MN|CN|KR (country codes) or area "Delivery zone" 42.7,-73.3|45,-73.3|45,-71.5|42.7,-71.5 (lat,lon points for borders that are not today's), pin@ka Karakorum 47.2,102.8, route "The road" ka|37.6,127 +arrow (pins or places). Parts come on in line order: +pulse on the one place to look, +dash for raided or planned, tone= as in shapes. A quick where is one line and the map in the chat. It draws from a world outline, so a town or a few streets is too small for it: place those in shapes with at=x,y.
  • Explainers draw every page. A place, a past or how something works is a picture on the stage, never pages of text. One line with the answer, then one deck on >full, 2 to 4 pages, each with its picture right after it: where is a map, how big or how fast is a chart or a stat, how it works is a diagram in shapes, a real thing to see is an image. A page of only words is the exception, and never a list of bullets after it. Not page "Where they ruled" body="The Mongol Empire ran along the Eurasian grassland..." then a stat and bullets, but:
>full
deck "The Mongols, by the map"
page "How far it reached" body="Korea to Hungary, the Siberian forest to Persia."
map caption="Karakorum sat in the middle and rode out every way."
area "Mongol Empire" 53,140|43,131|34.7,126.5|22.3,114|24,98|34,70|25.5,57|33,44|41,31|46,30.5|54,23|60,56|55,95 tone=butter
area Raided PL|HU +dash
pin@ka Karakorum 47.2,102.8 +pulse
route East ka|37.6,127 +arrow
route West ka|50.4,30.5 +arrow
page "The biggest one on land"
chart bar "Land empires, million km²" x=Mongol|Russian|Qing|Roman y=24|22.8|14.7|5
page "At its peak, 1279"
stat "24M km²" "A sixth of the land on Earth"
end
  • Music gets an instrument, not advice. Someone practicing, writing or jamming gets one line. A beat they edit while it plays: loop 96 "Boom bap" p=x...x.x.|....x...|..x...x.|xxxxxxxx +play. p is one string per row, top to bottom (kick, snare, clap, hat unless rows= names them): x a hit, . a rest, 8 steps. Name rows with kit words (kick snare clap hat open rim tom shaker crash cow snap conga) or a known alias (surdo, caixa, tamborim, ganza, agogo, djembe, cajon: MUSIC.md section 4), so each row sounds different. Pads to play on: drums 2x2, and drums 2x2 +record sends a take back as a beat. What they make comes back in the same words (bpm, swing, steps, rows, p): patch it in with ~loop p=... and keep it with save beat, and add nothing else. A keyboard to play in a key: keys Am pentatonic (notes outside the scale are locked, so nothing sounds wrong). A song's chords as big buttons to strum: chords G I-V-vi-IV, or by name chords C|G|Am|F; ~chords key=D moves them to a new key. +send on either sends back what they played. Tuning up: tuner guitar (or ukulele, bass, chromatic) listens on the mic and tells you when every string is in tune. A click, only when they ask for a tempo: metronome 80; Stop tells you how long they played. Whole practice sessions, ready to send: yuigui.com/api/library?q=music.
  • Save what they will reuse. After a screen they will want again (a workout, a routine, a check-in), add save workout: it goes on their shelf. Later, show workout brings it back instead of re-sending it; forget workout takes it off. One or two words per name.
  • Fill your drawer. Their drawer (a drag right on the chat) lists three things you keep up to date: menu review@dana "Invite Dana?" sub="asked yesterday" (waiting on them), menu backlog@deload "Deload week plan" sub=drafting (what you're working on), menu shortcut "Start today's workout" (a tap sends it as their message; say="Log a meal: " puts words in the composer). menu done dana takes one out when it's handled. A review or backlog tap comes back as [yui] dana menu bucket=review tapped: answer with that screen. The lines draw nothing in the chat.
  • Your home. When they open you, your newest four shortcuts are big chips over the bar and your review items are on the screen, so keep two to four shortcuts for what they do with you most. A shortcut with show= a screen you saved on a page (>2 ... save this week) swipes to that live page; keep those pages current with patches, never re-send them.
  • Say what you're doing. On a turn that takes more than a few seconds (reading, searching, drafting), send a few plain words each time the step changes, with the step when you know how many: doing "Reading your calendar" 1/3. Plain words, no ids or file names. It shows in your working row, never as a message, and your reply clears it.
  • Every button does something. No buttons that only acknowledge ("Got it", "OK", "Nice", "Cool"): a card with nothing to act on has no cta, and a note is a say. Name a submit for what happens, not a generic noun: plan "Trip" submit="Book it".
  • Offer, don't interrogate. Never ask what they already told you. Likely answers as options, +other for the rest.

When not to use a screen

A fact, a quick number, thanks, small talk, or "explain in words": plain text, no block.

Rules

  • At most 6 components per screen.
  • Only Yui Lines draw UI. Never HTML, JSON or markdown tables.
  • Don't narrate the UI ("here are some buttons", "tap below"). One short line, then the screen.
  • Never ask for passwords, codes, keys, card or account numbers, in a form or in text. Point to a safe place (the service's own login, settings, the environment).
  • Keep chat text under about 50 words; most answers need one line. The app folds a longer bubble into "Read as pages", which is a deck nobody asked for.

Other channels

Elsewhere (Telegram) there are no screens. If one would clearly help, offer "Want this on Yui?" and on yes, send it to the Yui channel.

Channel guide eval results

Does the channel guide (spec/CHANNEL.md) make an agent use Yui well? 88 realistic turns (cases.json) go to an agent with the guide injected the way the yui plugin injects it. Each reply is scored automatically against the real Yui Lines parser. Run it with node spec/channel-eval/run.mjs. Every example line in the guide must parse: node spec/channel-eval/guide.test.mjs.

Scores

guide words Opus 5.5 Sonnet 5 what changed
v0 886 25/34 (74%) 23/34 (68%) the first draft
v1 711 29/34 (85%) rewrite: fixed the patch rule, added "when not to use a screen", "answer what was asked", no secrets anywhere, 175 words shorter
v2 732 34/34 (100%) options are one |-joined token, with the failure spelled out
v3 700 34/34 (100%) 26/34 (76%) trimmed to 700 words
v4 703 32/34 (94%) 28/34 (82%) patch by bare preset name; "never ask what they already told you"; "free text is a form or mic" (this one backfired: Sonnet turned questions into forms)
v5 702 33/34 (97%) 27/34 (79%) v3 plus v4's first two changes, without the free-text line.
v6 750 33/34 (97%) v5 plus one media line (YUI-21): file paths and tool URLs in a line get hosted, hermes yui media, photos arrive as files. No regression; the miss is patch-timer-rounds. Shipped.
v6, tolerant parser 750 33/34 (97%) 30/34 (88%) same guide; the parser reads loose quoted options and ~preset@id (YL.md sections 4 and 5). Scorer counts a line as failed only when it still has nothing to tap.
v7 915 37/37 (100%) v6 plus the Reactions section (YUI-49, generated from spec/REACTIONS.md), and three new reaction cases (👍 builds it, 🤔 asks one screen at a time, 👎 drops it). No regression on the first 34. The first run had one harness timeout (focus-second-screen) and one case bug (the 👍 case proposed 185 lb squats to someone with 50 lb dumbbells; the agent built it and asked about the swap, which was right). Both re-run after fixing the case. Shipped.
v8 960 38/40 (95%) v7 plus "Every button does something" (YUI-53): no acknowledgement-only buttons ("Got it", "OK", "Nice"), a card with nothing to act on has no cta, submits are named for what happens. The scorer now fails any case on a dead button, and three new cases tempt one (overnight report, logging water, a trip plan that must not end in "Create project"). All three pass. The two misses are not buttons: secret-bank ran 88 words (cap 70), and dead-status-report wrapped a good screen in a four-backtick outer fence, copying the guide's own example. Shipped.
v9 1026 42/42 (100%) v8 plus "Findings, then questions: one plan" (YUI-51): a plan now holds page steps (each a real paragraph or points) before its questions, opens full screen, and sends everything with one button. Two new cases give the agent findings and then questions; the scorer's one_flow check fails a deck beside the plan, a question outside it, or a page with only a title. The v8 guide fails both new cases (it sent a deck, then a separate plan, the exact flow Chris called disjunct); v9 passes both. Six cases that already allowed plan now allow page too, because two v9 replies opened their plan with a good context page. One harness timeout (focus-second-screen) was re-run. Shipped.
v10 1074 42/44 (95%) v9 plus "Save what they will reuse" (YUI-32): save workout puts a screen on the person's shelf, show workout brings it back in a later reply, forget workout takes it off. Two new cases: a Tabata they will come back to must be saved (it was, as save busy-day), and a later "pull up my busy day workout" must be a show, not a re-sent timer (it was). The two misses were misses before too: patch-timer-rounds (the v6 miss) and dead-logged-water running 85 words against a 30-word cap. Shipped.
v11 1125 44/45 (98%) v10 plus "Screens 2 and 3" (YUI-31): the app now shows the chat and two pages beside it, pages keep what lands on them across replies, and sending there brings the page forward, so agents send only what should stay put. One new case (keep a shopping list in sight while planning a recipe) and a scorer check (page) that fails a reply with nothing on screen 2 or 3; v11 put the list on screen 3 with the recipe question in the chat, and focus-second-screen still sends its timer to >2. Both saved-screen cases still pass. First run 41/45; the four misses were re-run once: today-plan (used a plan), plain-thanks (26 words, cap 25) and dead-logged-water (37 words, cap 30) passed, patch-timer-rounds missed again, as in v6 and v10. Shipped.
v12 1213 45/46 (98%) v11 plus "progress over time" (YUI-65): a timeline with done, now and next rows. One new case (a project status: two shipped, one running, two queued) passes with one clean timeline in line order. today-plan, unprompted, drew the day as a timeline (deep work now, workout, lunch, pickup at 2) with one priority question; that is a good reply, so its allowed list now includes the timeline words, applied to every version (earlier runs never used them). The one miss is patch-timer-rounds, the same miss as v6, v10 and v11. Shipped.
v13 1297 2/2 new (100%) v12 plus "a game" (YUI-59): game tictactoe, game snake, game memory, and how a tic-tac-toe move comes back and gets answered. Two new cases: asked for a quick game, the agent sent one game tictactoe line and five words; handed the move [yui] n1 game kind=tictactoe move=5 o= x=5, it answered ~game o=1 and one word, a patch by preset name with only its own cell. Only the new cases were run (report v13-games); the guide above them did not change. Shipped.
v15 1445 47/50 (94%) v14 (Replies, no run of its own) plus "Mentions" (YUI-44): what a [yui] mention from=<agent> by=person message is, what [yui] note: lines are, and that an agent may @handle another only when it answers the person. Two new cases pass: asked by another agent's thread whether a quoted plan fits a sore knee, Arnold answered that plan with a swap table; given two [yui] note: lines, the agent rebuilt Saturday with Arnold's swaps. The three misses are in cases the new section doesn't touch (schedule-call left a choose with only +other, decision-three-options drew a timeline, dead-logged-water ran 44 words over a 30 cap), within the usual run-to-run noise. The runner's --into now adds cases new to a report instead of dropping them. Shipped.
v16 1617 49/52 (94%) v15 plus "Long answers are pages, not walls" (YUI-79, Chris on build 82: "I don't want these text bombs"): over about 50 words, send one line, a card and a deck +inline of short pages, counts as points or stat, with a bad and a good build report. Two new report cases (relay a worker's 180-word handoff summary; walk through a week of delivery changes) with a 50-word chat cap. The v15 guide fails the first one (it relayed the summary as text; report v15-report-cases); v16 passes both. The four misses were re-run once: project-timeline passed; patch-timer-rounds (the v6 miss) and dead-logged-water (48 words, cap 30) missed again, as before; mention-notes-context tried to write a memory file with tools off (harness noise; it passed in v15). Shipped.
v17 1666 2/2 new cases v16 plus one line on chat with a screen (YUI-62): >2 talk keeps the composer on screen 2, typed words arrive as [yui] screen=2 then the words, answer on that screen, >2 talk off ends it. Two new cases, run on their own (report v17-talk-cases): asked for a week they can edit by typing, the agent put the list on screen 2 with >2 talk; given [yui] screen=2 and a change, it patched the list by preset name and added one short line there, no re-sent list. The rest of the guide is unchanged from v16, so the full set was not re-run. Shipped.
v18 1773 4/4 report cases v17 plus "Draw it, don't describe it" (YUI-83): a sketch (a window, phone or chat bubble) with row lines struck out (+x), highlighted (+hi) or called out (note=), and one after line for a before and after. Two new cases (show how update wording changed; which button on a screen should go) need a sketch; the v17 guide fails both (a table, then words; report v17-sketch-cases), v18 passes both, drawing the old and new wording in a bubble and the screen before and after. The first run missed one: a correct sketch, but the whole reply wrapped in a four-backtick fence (the v8 slip); re-run once, it passed. The two earlier report cases still pass, and now allow the sketch words (report v18-sketch-cases). The rest of the guide is unchanged from v17, so the full set was not re-run. Shipped.
v19 1878 2/2 new cases v18 plus "Fill your drawer" (YUI-86) and +fold on a card: `menu review
v20 1976 3/3 new cases v19 plus "Groups" (YUI-93): what a [yui] group "<name>" ... hop=<n> from=<who> message is, answer only your part, what [yui] note: in <name> lines are, and @handle to hand work on only when another member's part is needed, never back. Three new cases (report v20-group-cases), with two new scorer checks, at and no_at: handed the runs by the lead, Arnold listed them and @'d nobody; as the lead asked to block a week and set the runs, Urza blocked the week and handed the runs to @arnold in one sentence; given two notes, Arnold moved Thursday's run around the new 7 am call without asking what changed. The v19 guide (report v19-group-cases) scored 2/3 on the same cases; its miss used a timer, not a group mistake, so these cases show no regression more than a clear gain. The rest of the guide is unchanged from v19, so the full set was not re-run. Shipped.
v22 2236 2/2 new cases v21 plus "Show how it works with shapes" (YUI-104): when someone asks how something works or how parts connect, even a quick question, one short line and a shapes diagram (a caption, shape KIND label lines, shape arrow joining the shape before it to the one after, a row by default, at=x,y to place, +grow +draw +pulse move=, tones from the look) instead of a paragraph or a generated picture. Two new cases (report v22-shapes-cases): asked how a heat pump heats a house "quick, I'm on my phone", Urza drew outside air, outdoor coil, compressor (pulsing) and indoor coil in a row with a caption; asked for the Yui flywheel in one picture, it placed four parts in a loop with arrows back to the start. The v21 guide on the same cases (report v21-shapes-cases) scored 0/2: a paragraph for the heat pump, and the flywheel as a donut chart of four equal slices. A first wording ("Show an idea with shapes", when the point is how parts connect) got 1/2, the heat pump still answered in words, so the line now names how-it-works questions. The four plain-answer cases (fact, thanks, explain in words, math) still get no screen (report v22-plain-check, 4/4). The rest of the guide is unchanged from v21, so the full set was not re-run. Shipped.
v23 2287 1/1 new case v22 plus a page's picture (YUI-85): in a deck or plan, a sketch right after a page is that page's picture, with one example (page, sketch and rows, end end), and shapes go outside the deck. New case report-pages-picture with a page_picture check: v22 0/1 (no deck, loose sketches), v23 1/1. First draft (no shapes clause) put a shapes diagram inside the deck in 1 of 3 walkthrough runs, which ended the deck; the clause fixed it (3/3). Regression: report 4/4, draw 1/1, plain 4/4. report-long-walkthrough now allows shapes, shape and end (v22 already answered with a diagram above the deck). +93 tokens per turn (cl100k, 3703 to 3796).
v24 2296 2/2 timeline cases v23 plus one clause on the timeline line (YUI-111): a row moves on by patch, ~now kind=done at=Fri (spec/YL.md "Moving a row", app yui 6a64a35, on phones since build 148). One new case, patch-timeline-move (a site rebuild timeline on screen 2, then "Blog migration just finished"): the v23 guide re-sent the whole timeline (report v23-timeline-move, 0/1); v24 sent one line, >2 ~now kind=done at=Thu. project-timeline still draws one clean timeline; its first v24 run wrapped the reply in a four-backtick fence (the v8 slip) and passed on the re-run (report v24-timeline-move, 2/2). Plain answers still get no screen (report v24-plain-check, 4/4). The rest of the guide is unchanged from v23, so the full set was not re-run. Shipped.
v25 2423 12/12 (teach, plain, shapes, report) v24 plus "A lesson is one screen" (TestFlight feedback APSw0dsa, Chris: "I want all that to be in one experience"): teaching with more than two pieces (a diagram, math, a chart, a stat, pages, a calc) goes on the stage with >full, calc last, and no close after it (a trailing close leaves the stage shut and only the chip shows; seen on the simulator). New case teach-one-screen with a one_screen check (every top-level piece on the stage, not ending in close): the v24 guide put the formula, a deck, a chart and the calc loose in the chat (report v24-one-screen, 0/1); v25 put all of it on the stage (report v25-teach). teach-compound-interest now allows shapes and shape (the case predates shapes; its v25 reply was good). Regression: plain 4/4, shapes 2/2, report 4/4. Deck pages cannot hold these pieces yet, so the lesson is one scrolling full screen for now; one swipeable deck is YUI-113. Shipped.
v26 2482 14/14 (teach, plain, idea, report) v25 with "A lesson is one deck" (YUI-113): a deck page now carries a diagram, math, a chart, a stat or a calc as its picture, so a lesson is one deck on >full, calc on the last page; only deck members go in it (a derivation is one math, never step lines, which end the deck). New one_deck check on teach-one-screen: the v25 reply re-scored fails it (6 pieces on the stage), v26 passes (2 of 2 runs). The first v26 draft slipped once with step lines inside the deck; the step clause fixed it. +59 words. Phones older than build 158 get the v25 stage layout from the plugin (compat.py _lift).
v27 2479 7/7 (plain, react) v26 minus "swipe it" in Replies (TestFlight feedback AACnEo9w): swipe to reply is gone from the app, hold and tap Reply is the only way. A wording fix, no new case. Regression: plain 4/4 (report v27-plain), react 3/3 (report v27-react). -3 words.
v28 2524 6/6 (library, flow, check-in) v27 plus "saved flows, ready to run" (FLOW-2 step 2): names the five saved flows (flow website-intake and four more), says to send one instead of building a plan when it fits, and points at the library search (/api/library?q=, library.json, the yui_library MCP tool). One new case, library-flow-intake (a client wants a bakery site): the v27 guide built a ten-question plan by hand (report v27-library); v28 sent flow website-intake in one line (report v28-library). A first wording, only the lookup URL, still built the plan by hand, so the line names the flows. Regression (report v28-flow-check): the three flow cases pass; checkin-morning now answers with flow workout-checkin, which asks the same sleep, energy and soreness, so its allowed list gained flow (applied to every version). +45 words. Shipped.
v29 2647 10/10 (short, report, plain) v28 with "Answer first, in one line" and "A deck only for 3 or more things to read" in place of "Long answers are pages" (YUI-118, TestFlight feedback AJq7CcQS8fyM, Chris: "eight screens of basically nothing just to tell me that I'm up-to-date"). The old rule sent anything over about 50 words to a deck, so agents padded status answers into pages. Now: the first line is the answer; a yes/no or status question gets one line or one card, never a deck; a deck needs 3 or more things to read, at most 4 pages, no page that repeats the headline or says what happens next with nothing to act on; a before and after sketch of the build question. Two new cases from the screenshot, short-status-latest-build and short-release-go-ahead, and a new max_pages check (4 on the report cases, 0 on the short ones). The v28 guide (report v28-short) scored 0/2: 49 words and five sentences to say yes, and a timeline for "go ahead"; v29 (report v29-short) answered "Yes, this phone is on build 160, the newest. Your iPad is still on 135." and three sentences for the release, 2/2. Regression: report 4/4 at 3 or 4 pages each (report v29-report; the saved v26 report replies re-score at 5 and 6 pages), plain 4/4 (report v29-plain). +123 words. Shipped.
v30 2710 8/8 (doing, short, plain) v29 with "Say what you're doing" moved from the waiting preamble into Use it well (YUI-63 step 2): on a turn that takes more than a few seconds, a few plain words each time the step changes, with the step when known (doing "Reading your calendar" 1/3), no ids or file names; it shows in the working row, never as a message. It went live with build 176 (Yui 0.3.2), the first TestFlight build that draws the words; hosts send doing only to phones at or above doing_min_build (176). New case doing-long-turn (check calendar, mail and the board, tool results carrying a card id and tool names) and a doing check (at least 2 lines with words, none carrying a card id, file name, tool name or code). The v29 guide (report v29-doing) sent no doing line, 0/1; v30 sent doing "Reading your calendar" 1/3, "Checking your mail" 2/3, "Looking at the board" 3/3, then the answer, 2 of 2 runs (reports v30-doing, v30-doing-b). Regression: short 2/2 (report v30-short), plain 4/4 (report v30-plain, no doing on a quick answer). +63 words. Shipped.
v31 2806 9/9 (music, plain, short) v30 plus "Music gets an instrument, not advice" (YUI-116 step 2, build 176 draws loop and drums): a beat they edit while it plays is one loop line with the pattern in p= (one string per row, x a hit, . a rest), pads are drums 2x2, +record sends a take back as a beat, and what they send comes back in the same words to patch in (~loop p=...) and save. keys, chords, tuner and metronome wait in the preamble for steps 3 and 4. Three new cases (report v31-music): a boom bap beat to mess with, pads to tap while waiting, and a sent beat. The v30 guide (report v30-music) scored 0/3: a markdown drum table, a tic-tac-toe game for the pads, and a re-sent loop with no save. v31 sent loop 90 "Boom bap" p=... +play, drums 2x2 +record, and ~loop bpm=94 swing=25 ... p=... with save beat, 3/3 on the first run. Regression: plain 4/4 (report v31-plain), short 2/2 (report v31-short). Phones below build 175 get the beat in words from the plugin (compat.py). +96 words. Shipped.
v33 3080 10/10 (music, plain, short); full suite 69, 70 and 66 of the old 76 v32 with keys and chords in "Music gets an instrument, not advice" (YUI-116 step 3, build 204 draws them): a keyboard in a key with the scale lock (keys Am pentatonic), a song's chords as big buttons (chords G I-V-vi-IV, or names `chords C
v34 3126 7/7 music in 3 of 4 runs; full suite 67 of the old 76 v33 with tuner and metronome in "Music gets an instrument, not advice" (YUI-116 step 4, build 208, Yui 0.4.1, draws them): tuner guitar (or ukulele, bass, chromatic) listens on the mic and says when every string is in tune, a click only when they ask for a tempo (metronome 80), and the library pointer (yuigui.com/api/library?q=music), now that every music screen in it draws. The step 4 preamble is gone. Two new cases: music-tuner-guitar (a restrung acoustic) and music-metronome-practice (strumming at 70 in 4/4). The v33 guide (report v33-tuner) scored 0/2 on them: a keyboard plus a list of string notes, and no metronome. v34 sent tuner guitar and metronome 70, 2/2, and all seven music cases passed (report v34-music). A rerun (report v34-music-b) showed one new miss: chords G I-V-vi-IV with an unasked metronome 90 (and music-jam-beat's known four-backtick fence). The metronome sentence then became "a click, only when they ask for a tempo"; the music cases passed 7/7 in 2 of 2 runs (reports v34-music-c, v34-music-d). Full suite (report v34-all, the draft before that one sentence): 71/80, so 67 on the 76 older cases, inside v33's 66 to 70; no music miss, one harness drop (no reply: exit null). Phones below build 205 get the tuner and the metronome in words from the plugin (compat.py). +46 words. Shipped.
v35 3343 explain 3/3 in 3 of 3 runs; full suite 69 of the old 80 v34 plus "Explainers draw every page" (YUI-157, TestFlight feedback AJIE1_1Ru1V4EMgmpWZtniI, Chris: "It's just a text bomb. Come on that's not the point of this app"): a place, a past or how something works is one line, then one deck on >full of 2 to 4 pages, each with its picture (a map in shapes placed with at=x,y, a chart or stat for size and speed, a diagram, an image); a page of only words is the exception, never bullets after it. One worked example, the Mongols by the map (the steppe in shapes, a bar chart of land empires, the peak as a stat). Three new cases in a new explain category (explain-mongols-geography, the screenshot's question; explain-rome-rise-fall; explain-monsoon-how) and a new drawn check: read as the stage plays the reply (stageChunks), every page and every line after the first needs a picture that draws (not a list, card or table of words), and at least one map, chart, timeline or image. The v34 guide failed explain-mongols-geography in 2 of 2 runs (a page of only words; a second sentence with no picture); v35 drew every page in 3 of 3 runs (reports v35-all, v35-explain, v35-explain-b), maps of the steppe and the Mediterranean in shapes, bar and line charts of territory. Full suite (reports v34-all-2, v34-all-3, v35-all): v34 70 and 71 of the 80 older cases, v35 69, inside the ~5 case noise; the v35-only misses are known harness ones (a four-backtick fence, a memory-tool attempt in menu-shortcut). Case change after seeing results, applied to every version: explain-monsoon-how allows 50 chat words (both v35 answers were two short sentences and one diagram, 41 and 45 words). The stat line in the example became stat "24M km²" ... after v35-all ran (the playground drew "24 M"); v35-explain and v35-explain-b ran on that final text. +217 words. Shipped.
v36 3367 flow 11/12, library 1/1; the new case passes on build 205 v35 plus one sentence on the saved-flows line (YUI-155, TestFlight feedback AMLn-Gg3qXC2NxpqF7fLpt4, Chris: "I asked for several steps and it came back with nothing"): every app runs a saved flow, one too old for flows gets the same questions as a plan, so the answers may come back as {plan}. Flows stay in the guide; the yui plugin (compat.py, flow gated until YUI-115) sends a flow to a phone that can't run it as the plan it walks by default, and a reply that promises questions with nothing to tap gets a question to tap. New case flow-interview-old-app ("Interview me for a personal brand site.") and two new checks in run.mjs: app_build scores the reply as that build receives it, through the plugin's compat.downgrade (the yui repo beside this one, or $YUI_PLUGIN), and tap fails a reply with nothing to tap (a flow still standing on a phone counts as nothing). The screenshot's reply (flow website-intake under a headline) fails tap through the old plugin and passes through the new one; the v36 agent sent the same flow and the phone got the 7-step intake plan. Flow category (report v36-flow): 11 of 12; the miss is menu-shortcut, the known harness case (a memory-tool attempt, also failed on v34 and v35). Library (v36-library): 1/1, still flow website-intake. +24 words. Shipped.
v37 3479 a map on all 18 answers that needed one; where 7/9, explain 5/6; full suite 75 of the 82 cases the rule does not touch v36 plus "Where is a map" (YUI-158 step 3, TestFlight feedback AL2nKEYo, Chris on the Mongol Empire answer, a big number over East, West, South and North bullets: "Again, not bad but this should be a Map"): any answer about where (geography, a past empire, a trip, a route, a delivery area, a storm) leads with a map (area by country code or lat,lon points, pin, route), never compass points in words or a map drawn in shapes; a quick where is one line and the map in the chat; a town or a few streets is too small for the world outline, so those go in shapes. "Explainers draw every page" now says where is a map, and its Mongols example draws the empire on a map (outline, raided lands dashed, Karakorum pinned, routes east and west) instead of dots in shapes. Phones below build 219 get the place names in words from the plugin (compat.py MAP_BUILD, YUI-158 step 2). New where category (where-trip-route, Lisbon to Barcelona by train; where-delivery-area, a farm delivering across Vermont, New Hampshire and western Massachusetts; where-quick-country, where is Kyrgyzstan) and a map check: a where answer with no map, area, pin or route fails. explain-mongols-geography and explain-rome-rise-fall need a map now; all three explain cases accept map lines. The v36 guide failed all three where cases in 2 of 2 runs (reports v36-where, v36-where-b: sketches, bullets and a 3 page deck for Kyrgyzstan) and drew the Mongols and Rome in shapes (v36-explain, v36-explain-b, 1/3 each). v37 sent a map every time: where 7 of 9 (v37-where, -b, -c), explain 5 of 6 (v37-explain, -b); every miss is a word cap (41, 46 and 52 words, a paragraph after the map). Full suite: v36 71 and 73 of the 82 cases the rule does not touch (v36-all-a, v36-all-b, rescored with the yui plugin beside the repo), v37 75 (v37-all), inside the noise; its misses are the known ones (patch-timer-rounds, mention-asked, menu-shortcut's harness drop). Changes after seeing results, applied to every version: the first draft's delivery case (three Brooklyn neighborhoods) drew as a speck, since a map never goes closer than about 8 degrees, so the guide sends towns and streets to shapes and the case became a regional area (reruns on the final text; the Brooklyn replies are gone from the reports); explain-monsoon-how counts a map of the winds as its picture; where-trip-route allows a small table of the legs beside the map. +112 words. Shipped.
v32 (not published) 3024 2/2 new show cases; full suite 61/76 and 65/76 (v31 on the same 76: 71/76 and 66/76) v31 plus "Show, don't say" (YUI-119 step 1, TestFlight feedback AL1My-stUBnvtXD2WKV7ec8, Chris: a headline plus a paragraph on a screen is not a screen): each part of an answer is one short line and one picture, a page in a deck gets its picture (in a plan only a sketch), asked to see something draw it, and a report of facts stays a line and a card. One worked example ("Walk me through phase one", the bottom bar before and after). New cases show-phase-one and show-new-layout replay that screenshot, scored by a new show_not_say check (read as the stage plays a reply, stageChunks in site/lib/yl/chunks.mjs: a part with no picture, or a page over 30 words, fails). The v31 guide failed show-phase-one in 1 of 3 runs (headings over paragraphs, report v31-show) and passed show-new-layout in 2 of 3; v32 passed both in every run (reports v32-show, v32). The widened cases: flow-findings-then-questions and flow-two-questions-one-plan now allow sketch, row and after (the replies drew each finding, which is the rule). Held: across full runs v32 sits 4 to 6 cases under v31, just outside run-to-run noise, with no single case to blame (reports v31-all, v31-all-b, v32). Not published to agents until that gap is understood. Published with v33, whose full runs sit inside v31's range.
YUI-161 (no guide change) 3479 1/1 Opus, 3/4 GLM 5.2 New case list-no-escaped-breaks and a no_breaks check (any \n inside a quoted string in a yui block fails; YL reads it as the letter n). From a live native turn: GLM 5.2 as Basil sent one say with \n between five dinners and the app showed "nn". The fix is in the native runtime (yui runtime/src/turn.ts unbreak(), plus one rule in its prompt), not the guide. Hosted Opus on v37 (report yui161-opus) sent card, table, list and choose, 1/1. GLM 5.2 on the guide alone (reports yui161-glm, -b, -c, -d) wrote no \n in 4 of 4 runs; its one miss is 83 words of markdown bullets, not a break. The live reply re-scored (report yui161-live-before) fails breaks; after the runtime sweep (yui161-live-after) it passes breaks and keeps only the stray end the model also wrote.

Before and after, on the fleet's model (Opus 5.5): 74% to 97%, with the guide 20% shorter. On Sonnet 5, 68% to 79%. Scores between v2 and v5 are within run-to-run noise (about two cases either way), so v5 ships because it adds a correct rule (patch by bare preset name), not because of its last point.

Every score up to v6 uses the final cases and the scorer as it stood before the tolerant parser, re-scored with run.mjs --rescore without re-running the saved replies. The tolerant-parser rows use the current parser and scorer.

Show it, don't tell it (YUI-203, guide v39)

Chris's four TestFlight notes on one thread (a ZIP field described in words, "two days earlier" as bold text, a lone hyphen line, a deck whose last page had nothing to tap) became four rules in the guide: UI is drawn in a phone sketch, when is a timeline, facts are Label: value and never a lone hyphen or a list of one, and the last page of a deck, plan or flow ends in something to tap (Why do you ask? as the fallback). Six new cases (ui-zip-in-context, when-two-days-timeline, facts-stacked-list, fact-one-sentence, last-page-next-step, last-page-walkthrough) and three new scorer checks (phone, no_lone_bullet, last_tap). Opus 5.5, the old guide (v38) twice and v39 twice, 96 cases each.

run guide passed of the six new cases
t203-old-r1 v38 83/96 3
t203-old-r2 v38 78/96 3
t203-new-r1 v39 81/96 6
t203-new-r2 v39 83/96 6

The old guide fails the same three new cases both times: when-two-days-timeline (dates in a sentence) and both last-page cases (a deck or sketch that ends with nothing to tap). The 90 older cases swing 75 to 80 on either guide, the usual noise. flow-interview-old-app fails every run here because the yui repo is not beside this worktree (spawnSync python3 ENOENT). One case changed after seeing results: report-pages-picture now allows a choose or ask, because a report deck ending in next steps is what the new rule asks for.

Answers that know what they answer (t_53b06721, guide v41)

Chris's TestFlight note: a sample status board said "New feature: needs help, waiting on you", he asked "what are you waiting on me for with this?", and the agent answered with an old, unrelated ask (spike: spec/research/context-on-reply.md). Guide v41 adds Examples are not asks, changes the guide's own status example so it no longer teaches note="waiting on you" on made-up rows, and the plugin now hands the agent a note naming its newest message when he types a line. Three new context cases: a sample marked as one, a question about the screen just shown, and an old ask named with when he last saw it. The last two carry the plugin's note in the message, so they test the guide and the note together.

run guide passed of the three new cases
t53-full-old-r1 v40 90/102 3
t53-full-old-r2 v40 84/102 2
t53-full-new-r1 v41 83/102 3

Ten cases failed on the new run while an old run passed them. Rerun on v41 (t53-full-new-r2, r3): nine of the ten pass, the word-cap cases included. The one that still misses is react-no (a struck-out sketch for a dropped follow-up, the v40 "outcome is drawn" rule; the old guide's second run misses it the same way). list-no-escaped-breaks had failed once with exit null (the CLI call died, no reply) and passed on the rerun. The old guide's one miss on the new cases is context-sample-not-ask: a sample board that says nothing about being a sample. The 99 older cases swing 84 to 90 on the same guide, the usual noise; nothing the new rule touches got worse.

Chris's TestFlight note: he asked "what are you waiting on me for with this?" and the answer paged to a card, "Before and after shots / On the progress page / Open". He wrote: "We have components to show before and after. Let's not do so much linking out." Guide v42 adds Show it here, don't link out: shots are compare, image or gallery, a UI change is a sketch with after, a page's content is its parts drawn, and a card ... url= is only a small follow-up under what is already shown. The scorer has a new show_here check: a card that leaves the thread (url or open) fails when nothing is drawn beside it, and two link-out cards fail. Ten new show cases: before and after of a hero, what changed on the progress page, "just show me here" after a link-out card, a demo, and six where the looked-up facts carry a url and little more than a sentence about the change.

run guide model passed of the ten new cases
t1d7-full-old-r1 v41 Opus 96/112 10
t1d7-full-old-r2 v41 Opus 99/112 9 (showlink-demo-page: a card that only opens the playground)
t1d7-full-new-r1 v42 Opus 97/112 10
focus, link and lean runs, Opus, 2 each v41 / v42 Opus old 19/20, new 19/20 new miss: showhere-after-linkout r1, no ```yui block at all
focus, link and lean runs, Sonnet 5.5, 2 each v41 / v42 Sonnet old 19/20, new 19/20 showlean-new-hero over the word cap on both

Honest read: the new cases are a guard, not a big win. Opus with the v41 guide already drew the shots in nearly every case once the looked-up facts held the image paths; its one link-only answer was the demo page. The harness cannot separate the two guides on these cases: the counts are a tie. The live miss needs the long context Chris had (a 14-page answer), which the harness cannot replay. The full suite swings 96 to 99 on the same v41 guide, and v42 sits inside that range: the 15 misses on the new run are the known noise set (menu-shortcut, doing-long-turn, short-release-go-ahead, flow-interview-old-app, the exit null cases).

Models on the same guide (YUI-132)

Native Yui runs on GLM through OpenRouter (spec/NATIVE.md section 6), so both GLM models got the full suite on the v37 guide, beside Opus and Sonnet on the same 88 cases. Each model ran once; every miss ran a second time. A case that failed both times is a steady miss; one that passed the second time is noise. Scores are the first pass. Drawn on /channel.

model how it is called first pass missed twice what it gets wrong, most first
Claude Opus 5.5 claude CLI, the fleet's agents 82/88 (93%) 3 a preset the case doesn't want (a sketch or a choose where a list was asked), one timer rebuilt instead of patched
Claude Sonnet 5 claude CLI 70/88 (80%) 11 new screens instead of patches (3), the wrong kind of screen (3), a flow split into loose questions (2), too long (2)
GLM 5.2 OpenRouter, Yui's pick for text 67/88 (76%) 12 too long (6 of the 12: chat words over the case's cap), a button label broken onto its own line (submit=...), turns that thought through all 2000 tokens and answered nothing
GLM-5V-Turbo OpenRouter, Yui's pick for photos 55/88 (63%) 21 the wrong kind of screen (10), too long (9), lines the app can't read (7, three of them shape kinds such as blob and box written as their own lines), a form that asks for an API key

Defaults confirmed, runtime/src/models.ts unchanged. GLM 5.2 sits with Sonnet 5 (76% and 80%, 12 and 11 steady misses), well enough for every text turn. GLM-5V-Turbo is clearly weaker on general turns, but the runtime only sends it turns with a photo, and it passes the photo cases a text model would get (meal-photo). No other vision model has been scored yet, so it stays until one scores better. It should not be offered for text turns.

Found along the way:

  • GLM 5.2 can think its whole budget away. The runtime sends max_tokens: 2000 with no reasoning cap. On a long ask (explain the rise and fall of Rome, a group hand-off) GLM 5.2 spent all 2000 tokens reasoning (7,162 characters) and returned an empty answer: two cases in the first pass, two on the re-run. A runtime fix (a reasoning cap or a bigger budget), not a guide one.
  • Fixed in the runtime (YUI-162). A native turn on OpenRouter now sends max_tokens: 3000 with reasoning: {max_tokens: 1000}: the answer keeps its 2000 and thinking gets 1000 on top. GLM takes that cap as a hint, not a limit: in 15 turns of the first rerun one Rome answer still thought 2958 of 3000 tokens and left half a sentence (deck "Rome, rise and fall). So the runtime also asks once more with reasoning: {enabled: false} when an answer is cut off at the limit and is empty or spent over half the budget thinking (thoughtOut() in turn.ts); the ran-out-of-room line is only the last resort. run.mjs sends the same request and retry, and marks a retried case retried. Rerun on the three cases that came back empty (explain-rome-rise-fall, group-lead-hands-on, flow-findings-then-questions), six passes each, 18 turns: 0 empty answers, 0 retries needed, finish stop on all 18, thinking 49 to 1,198 tokens (median about 500). Reports yui162-glm52-r1 to r6. The misses left on those cases are guide ones (submit= on its own line, a deck where a list was asked). Live on yuigui: "explain the rise and fall of Rome" to hosted Yui on a throwaway account answered with a 4-page deck (map, chart, shapes) in 7 to 25 s, 4 of 5 runs; the one run with no answer came right after the function deploy.
  • Cheap guide fixes, not made on this card: GLM-5V-Turbo writes blob "More people" instead of shape blob "More people" (4 cases across both runs); one clause ("every line under shapes starts with shape") should catch it. Both GLMs run long; a number in "Answer first, in one line" (chat words, about 40) may hold them better than the prose does.
  • The new case, data-lunch-macros (Basil, "a quick lunch idea with 30g of protein"), is the YUI-131 live miss: GLM 5.2 wrote a quoted title, columns: rows and end, and every row was a parse error. On this guide all four models passed it.
  • Scorer fix, applied to every report: eight cases that allow a timeline now also allow its done, now and next rows. Before, a correct timeline failed as "preset: now not in [...]". It moved Sonnet 69 to 70; the other first-pass scores did not change.
  • Cost on OpenRouter: about $0.40 for a GLM 5.2 pass and $0.25 for GLM-5V-Turbo, a cent or less a case. Most GLM 5.2 turns went to Mistral or Wafer as the provider, all with data collection denied.

What the eval found

  1. Cross-reply patches were broken by the guide itself. v0 taught timer@hiit 40/20x8 "then later" ~hiit rounds=10. The app parses every reply with a fresh parser (ChatStore.swift), and YL.md section 9 says ids only resolve inside the reply that made them. So an agent following the guide sent a patch the app throws away, and the screen never changed. Four of v0's nine failures were this. v5 says: in a later reply, patch by the bare preset name (~timer, ~stat, ~choose +lock).
  2. Space-separated options silently render nothing to tap. choose "Where?" "Camera roll" "Drafts" parses without an error, but as one long question with no options (an ask falls back to Yes/No). The parser cannot flag it, so the scorer checks for it. v1 did this in 5 of 34 replies. v2's explicit rule took that to zero on Opus.
  3. Tacked-on questions. v0 agents added a "Keep it?" after a theme change, an energy slider under a weight chart, and a lunch question under a packing list. "Answer what was asked" fixed all three.
  4. Plain questions stayed plain. Every version answered "capital of Portugal", "18% tip on $84", "thanks" and "explain in two sentences" without a screen. No version ever put a secret in a form.

Known gaps

  • Sonnet 5 sits around 80-88%. Its steady misses: forms or plain text where a question would do (schedule-call, flow-onboard-goal), a pick where a list was asked for, and an unclosed fence. Some misses in each run are harness noise: with tools off, it sometimes tries to read a memory file instead of answering (today-plan, decision-three-options, list-groceries).

  • The parser now tolerates the two slips agents make most (YL.md sections 4 and 5): choose "Q?" "A" "B" reads the trailing quoted tokens as options, and ~card@week patches week when that reply made it and the newest card otherwise. Honest size of the win: Sonnet's ~card@week shows up in patch-plan-card in every Sonnet run (v3, v4, v5, v6) and now passes, so that is one case. Loose options appeared in v0 and v1 replies (up to nine lines in v1) and in none since the v2 guide rule, so today they are a safety net, not a score. Re-scoring the saved v5-sonnet replies with the new parser gives 28/34 (82%, was 27/34); the fresh v6 Sonnet run's 30/34 is that one case plus run-to-run noise.

  • Multi-turn cases replay earlier turns as a transcript in one user message. The fleet's shim resumes real sessions instead. Taps and patches still behave as expected (all patch cases pass on Opus), but this is a proxy, not the live channel.

  • Dead buttons are rare in the harness and showed up live. Rescoring every saved reply (v0 to v8, Opus and Sonnet) finds no acknowledgement-only button, and the v7 guide passes the three new cases too. The live case that started YUI-53 was Yui confirming a saved rule: card "New rule saved" ... cta="Got it". That turn doesn't reproduce here: with tools off, the CLI tries to write the rule to memory and times out. So the new check is a regression guard. The rule in the guide is what fixes the live channel.

  • Outer fences get copied. One v8 reply put its whole answer in the four-backtick fence the guide uses to show an example. It happened once in 40.

Method

  • The runner calls the claude CLI the way the fleet's shim does: the agent persona and today's context, then Yui channel guide <version> and the guide text (the same extraction and version hash as yui/hermes-plugin/sync_channel.py), and the per-agent look line, all appended to the CLI's system prompt. No tools. Calls are killed after 180 s and retried once. A model id with a slash (z-ai/glm-5.2) goes to OpenRouter instead, the way Yui's native runtime calls it: the same system text as the system message, max_tokens: 2000, provider.data_collection: deny, key from OPENROUTER_API_KEY. scores.mjs turns the model reports into site/content/model-scores.json for /channel.
  • Scoring per reply: a ```yui block when a screen is required, and none when it is not. Every line must parse with site/lib/yl/yl.mjs (one fresh parser per reply, like the app). Presets must be in the case's allowed set, and at least one of its needed presets must be used (patches count). At most 6 components. Chat text under the case's word cap (70 by default). Options must be |-joined. Fails on: HTML, custom, markdown tables, Yui Lines outside the block, UI narration ("here are some buttons", "you chose..."), and any form field for a password, PIN, key, token, card or account number. Patch cases also need a ~ patch on the right target, without re-sending the component.
  • Case changes made after seeing results, applied to every version: schedule-call now allows plan, focus-second-screen allows mic, and secret-login allows plan. In each of those, the reply was good and the allowed list was too narrow. Scorer changes, also applied to every version: patches count toward "needed", and the options-join check.

Reports

Full transcripts and per-case reasons are in reports/<run>.md, with raw replies in reports/<run>.json (the guide text each run used is inside the JSON). Runs: v37-opus (v37-all plus data-lunch-macros), v37-sonnet, v37-glm52, v37-glm5v, and each one's -rerun (its misses once more) (YUI-132), v37-all, v36-all-a, v36-all-b (full suite, the v36 guide twice), v37-where, v37-where-b, v37-where-c, v36-where, v36-where-b (the v36 guide on the three where cases), v37-explain, v37-explain-b, v36-explain, v36-explain-b, v36-flow, v36-library (the v36 guide on the flow and library cases), v35-all, v34-all-2, v34-all-3 (full suite, the v34 guide twice), v35-explain, v35-explain-b, v34-all (full suite), v34-music, v34-music-b, v34-music-c, v34-music-d, v33-tuner (the v33 guide on the seven music cases), v33-all-a, v33-all, v33-all-b (full suite), v33-music-2, v33-music-2b, v33-music-2c, v33-music, v32-music (the v32 guide on the five music cases), v33-music-drums-b, v33-music-drums-c, v33-plain, v33-short, v31-music, v30-music (the v30 guide on the three music cases), v31-plain, v31-short, v29-short, v29-report, v29-plain, v28-short (the v28 guide on the two short cases), v28-library, v28-flow-check, v27-library (the v27 guide on the library case), v27-plain, v27-react (feedback AACnEo9w), v26-teach, v26-teach-b, v26-report, v26-plain, v26-idea (YUI-113), v0-baseline, v1, v2, v3, v4, v5, v6, v6-tolerant, v7, v8, v9, v8-flow-cases (the v8 guide on the two new flow cases), v7-dead-cases (the v7 guide on the three dead-button cases), v10, v11, v12, v15, v16, v15-report-cases (the v15 guide on the two report cases), v17-talk-cases (the v17 guide on the two talk cases), v19-menu-cases (the v19 guide on the two drawer cases), v21-restyle-case (the v21 guide on restyle-app-autumn and theme-autumn: the app look and the agent's own look stay apart), v22-shapes-cases and v21-shapes-cases (the v22 and v21 guides on the two shapes cases), v22-plain-check (the v22 guide on the four plain-answer cases), v22-page-picture and v22-walkthrough-check (the v22 guide on report-pages-picture and report-long-walkthrough), v23-report-cases, v23-draw-check and v23-plain-check (the v23 guide on the report, drawing and plain-answer cases) (Opus 5.5), and v0-sonnet, v3-sonnet, v4-sonnet, v5-sonnet, v5-sonnet-tolerant (the v5 replies re-scored with the tolerant parser), v6-tolerant-sonnet (Sonnet 5). Rows up to v6 were scored before the tolerant parser; re-scoring them now would credit early guides for loose options the parser fixes, so they are left as they were.

Try Yui, or help build it

Get the alpha

The MVP is done and Yui is in alpha, open to anyone with an iPhone on iOS 26. Download it on TestFlight, then connect the agent you already run: Hermes, OpenClaw, Claude Code, a model you run, or anything behind a webhook.

Star it on GitHub

Yui is open source under Apache 2.0. Star the repo, open an issue, or send a pull request.

Lend your agent

Spare tokens on Claude or ChatGPT Codex? Your agent can pick a card off our backlog and open a pull request. Yui@home, like SETI@home.

Want a hand getting in?

You don't need this to try Yui: the alpha on TestFlight is open to anyone with an iPhone on iOS 26. Leave your details if you have no agent yet, want help connecting one, or would rather Apple email you the invite.

  1. Yui emails you a link to confirm your address.
  2. We read your request, and reach out if you asked for help.
  3. Apple emails you a TestFlight invite.
  4. Open it on your iPhone, install Yui, and sign in with Apple.

We use this only to get you into Yui. Privacy.