Skip to content

BlogBehind the Scenes

How we test the iPhone app: Claude, Codex and a simulator

Work on how the iPhone app looks took 42 percent of our Claude use. We compared four ways to check it on a simulator, and what each one found, missed and cost.

By The Little Worker team6 min read

In one overnight push, the iPhone app changed a lot. The workers got faces, Today got warmer, the chat was rebuilt, and each worker got a page of its own. AI tools wrote most of that code, Claude above all, with a person giving direction and saying yes or no.

Code that compiles is not an app that looks right. Each change has to run on an iPhone simulator (a copy of an iPhone that runs on a Mac), and someone has to look at the screens. Most of the time that someone is Claude, reading pictures of the simulator.

Why the checks got costly

AI tools charge by tokens: small pieces of the text and pictures a model reads and writes. A picture of a phone screen costs many more tokens than a line of code. When we counted our recent Claude sessions, work on what the app looks like used 42 percent of all Claude tokens. Those sessions built the app 366 times, took 932 screenshots and had Claude read 2,191 pictures.

That is a lot of looking. So we asked a plain question: which way of checking the app finds the most real problems for the least cost?

The test

We gave four setups the same job on the same simulator and the same build. Each one looked at four screens, in light and in dark:

  • Today with a full list of to-dos
  • A chat where Pip is at work
  • Pip's worker page
  • Today with nothing in it

Each one also did two taps: give the first to-do to Pip, then open Chats.

To know what was really wrong, one more Claude session acted as the judge. It took its own pictures of every screen at 1, 2 and 6 s after the app opened, did both taps, and read the labels that VoiceOver reads. It found four real issues:

# Issue How bad
1 On Today, the first to-do and its Give to Pip button sit under the bars at the bottom of the screen until you scroll. Minor to medium
2 When you give a to-do to Pip, it leaves the list and the count drops from 4 to 3, but Pip's face does not change and nothing on screen says Pip got it. Medium
3 In a chat, about 2 s after the app opens, Pip's first messages can draw on top of later ones while the screen settles. It happens only some of the time. Medium, only sometimes
4 On the worker page, one row has Pip's face and the row under it has no icon, so the text in the two rows does not line up. Small

About the second issue: our rule is that a task waiting in line does not make a worker Busy, so Pip's face is right not to say Busy yet. The gap is that nothing on the screen confirms Pip got the task.

The four setups

  1. Claude alone. Claude takes screenshots with Apple's own command-line tools and reads them. It cannot tap.
  2. Claude with AXe. AXe is a free command-line tool that taps and types on a simulator and reads the same labels VoiceOver reads. Claude decides what to tap and reads both the pictures and the labels.
  3. Claude with a Codex judge. Codex is OpenAI's coding agent. Here a fixed script takes the pictures and does the taps, and Codex grades the pictures against a checklist. Claude only starts the script and reads the result.
  4. Codex with computer use. Codex runs the whole check by itself and works the simulator on its own. Claude only starts it.

Setups 3 and 4 start Codex from the command line, with no person at the keyboard. This is the command, shortened:

codex exec -m gpt-6.1-sol -c model_reasoning_effort=high \
  --cd run-folder -o last_message.txt "$(cat prompt.txt)" < /dev/null

What we found

Two bar charts. Claude with AXe found 3 of 4 real issues and used about 252,000 Claude tokens. Claude alone found 2 for 165,000. The two Codex setups each found 1, for 124,000 and 97,000 Claude tokens.
The setup that found the most also cost the most Claude tokens. The Codex setups cost less Claude, but they also used Codex tokens, which are billed on their own.
Setup Claude tokens Codex tokens Issues found Wrong alarms Taps right
1. Claude alone 165,224 0 2 of 4 0 0 of 2
2. Claude with AXe 251,724 0 3 of 4 1 2 of 2
3. Claude with a Codex judge 124,169 30,447 1 of 4 0 2 of 2
4. Codex with computer use 97,328 67,878 1 of 4 0 2 of 2

Each check took from 97 s (Claude alone) to 197 s (the Codex judge). In the first two setups Claude read 16 to 17 pictures; in the Codex setups it read none, because Codex looked at the pictures instead.

The Claude numbers are weighted, so each kind of token counts by what it costs. Text Claude has already seen counts a tenth, and words Claude writes count five times. About 40,000 to 50,000 of each Claude number is the fixed cost of starting a session at all.

Claude alone was the fastest and had the best eye for layout. It was the only one to see that two rows on the worker page do not line up. But it cannot tap, so it could not check either tap.

Claude with AXe found the most. It did both taps, read the exact labels, and it was the only setup that took more than one picture of a screen over time. That is why it was the only one to catch the chat messages drawing on top of each other. It also made one wrong alarm: it called the Chats tab "the wrong tab" under a worker page that was still loading. Worker pages open from Chats, so Chats was right.

Claude with a Codex judge got both taps exactly right and used little Claude. But it found only what its checklist asked about. The checklist said that content under the bottom bars is normal, so it missed the hidden to-do.

A judge with a checklist finds what the checklist asks for, and excuses what the checklist excuses.

Codex with computer use used the least Claude and did both taps right. It called every still screen clean, so it found nothing the taps had not already shown. It also used the most Codex tokens of any setup.

What we use now

We use Claude with AXe as the check before each release of the iPhone app, and we are making it cheaper by taking the fixed steps away from the model:

  1. A script opens each screen, takes pictures at about 2 s and 6 s, does the taps with AXe, and checks the exact labels: the to-do count, Pip's state and which tab is selected. This costs no tokens, and it gives the same answer every time.
  2. Claude reads only the small pictures, with no checklist that tells it what to ignore. Each screen is opened twice, so a problem that comes and goes has two chances to show.

We stopped using Codex with computer use for this job. We keep the Codex judge only as a cheap repeat check between releases, and only with a checklist that does not excuse layout problems.

The issues themselves come next: show that Pip got the work, let the first to-do show above the bars, and find why the chat messages overlap while the screen settles.

The limits

  • Each setup ran once, on one build, with one judge. A problem that comes and goes, like the chat messages, can be caught or missed by chance.
  • The setups ran one after another on the same simulator, so they did not all see the same moments.
  • The judge was a Claude session too, and the grades are its judgment.
  • Four screens and two taps are a small part of the app. A setup that did well here can still miss a problem on a screen we did not test.