Blog ·
Giving phones thumbs
For the last year or so I've been watching AI agents get properly good at using computers. Claude can open a browser, find its way around a website and fill in a form, and most of the time it gets there. Codex can sit in a terminal for an hour and come back with a pull request that mostly works. What kept bothering me is that almost none of this reaches the device I actually spend my day on. My bank is an app. The taxi is an app. Parking, the gym booking, the parcel locker, and the two-factor code that every one of those things sends me all happen on my phone, and none of the agents I use could touch any of it.
There are good reasons for that. Phones are built for one person using them with their hands, and both iOS and Android go to a lot of trouble to stop apps from reaching into each other. That's a big part of why phones don't get the kind of malware that Windows machines used to get, so I'm not complaining about it. But it does mean there's no obvious place for an agent to plug in. On a laptop you can take a screenshot and move the mouse, and that's more or less what computer use is. On a phone, an app that wants to do the same thing has to ask for permissions that almost no app asks for, and on an iPhone there's no way to do it from inside the phone at all.
Starting with Android
So we started with Android, because Android lets you do this properly. Opposable is an app you install on an Android phone. It uses the accessibility service, which is the same part of the system that screen readers for blind people use, to read what's on the screen and to tap, swipe and type. Then it offers all of that as tools that Claude Code, Codex, Cursor or anything else that speaks MCP can call. You install the app, paste one line into your terminal, and your agent has a phone.
One decision that made a big difference early on was to give the agent the screen as a list of things it can press, with their labels and positions, instead of just a picture. Models are surprisingly bad at guessing exactly where a small button is in a screenshot, and they're very good at reading a list and saying "tap number 14". It's faster too, because a list of a few dozen items is much cheaper to send than an image. The agent can still ask for a screenshot when it needs to see something, like a photo or a map, but most of the time it doesn't have to.
The first time the whole thing worked end to end, I asked Claude Code to turn on dark mode and set a five minute timer. It took seventeen tool calls and about fifty seconds, and watching it scroll up and down the Settings app looking for the Display menu felt a lot like watching my parents use their first smartphone. It's a good deal quicker now. Most of that speed came from fairly boring changes, like keeping one connection to the phone open instead of opening a new one for every tap, and waiting until the screen has actually stopped changing instead of waiting a fixed half second and hoping for the best.
What surprised me
The first thing I didn't expect was how much of using a phone is dealing with interruptions. There are permission dialogs, "rate this app" pop-ups, cookie banners, login screens you didn't know were coming, and a six digit code that arrives by text message right in the middle of everything. You handle all of these without really noticing. An agent that doesn't know about them will stare at the same screen forever. A lot of the work has gone into making the phone point these out clearly, dismissing the harmless ones on its own, and reading one-time codes from messages and notifications so the agent can finish what it started.
The second was that a lot of what people want a phone agent to do doesn't need a big model. Setting an alarm, adding a contact, sending a text or switching on Wi-Fi are the same few steps every time, and the app can do them reliably with no model at all, just by knowing how those screens are laid out and checking afterwards that the thing actually happened. For everything else there's a small screen model that runs on the phone itself. It looks at the screen and decides where to tap. It isn't as clever as Claude and it's slower than I'd like, but it works offline, your screen never leaves the phone, and it's been getting faster pretty much every week. I suspect this is where all of this ends up: something small on the device for the everyday stuff, and a big model in the cloud for the hard parts.
The third thing I should have seen coming. People get nervous when you put "AI" and "my banking app" in the same sentence, and they're right to. So anything that looks like a payment stops and asks you on the phone before it goes through, with an Approve and a Deny button, and everything the agent does goes into a log on the phone that you can read afterwards. Nobody should have to just trust that the agent did what it said it did.
And the iPhone
Apple doesn't let any app control other apps, and I don't expect that to change. What an iPhone will happily do is accept a Bluetooth keyboard and mouse, and if you turn on AssistiveTouch, a mouse click is just a tap. So the iPhone version is a small USB stick that plugs into your computer and introduces itself to your phone as a keyboard and a mouse, while the computer watches the screen through mirroring. The parts cost about twelve dollars, the firmware is open, and you can install it on the stick from your browser in about a minute. I want to be upfront that this part is still in beta. We haven't put it through the same testing as the Android app yet, and I'd rather say that than have you find out.
Who it's for
Right now it's mostly for people who already use coding agents and have wished they could hand them their phone. Developers use it to test their own apps on a real device instead of a simulator. Some people use an old phone from a drawer as a dedicated agent phone with its own apps and accounts. Others want their agent to finish jobs that dead-end in an app, like checking whether a parcel arrived, booking a table, or getting past the verification code it couldn't read before.
The Android app is free and the code that runs on your phone is open. We charge for the parts that are annoying to run yourself, like reaching your phone from anywhere, payment approvals and syncing, and later for phones we run in the cloud for you.
Where this goes
Computer use got a lot of attention because it was the first time an AI could use software that was made for people rather than for other programs. Most of the software made for people lives on phones now, and for a lot of the world the phone is the only computer they have. That seemed like a big enough gap to spend some time on. It's early, there's a lot that doesn't work yet, and we're fixing it in the open. If you try it and it gets stuck somewhere, I'd really like to hear where.