Claude Code Controls My iPhone: A Skill Built on iPhone Mirroring
What this post is. How a small Claude Code skill works, taken from its source: about 45 lines of Bash, a 17 line Swift helper, and a 49 line SKILL.md. The skill is private for now. The snippets are fragments, not the full scripts.
Short version
- macOS 15 can show a live iPhone in a window, through the iPhone Mirroring app. A window can be captured and clicked, so an agent can use the phone.
- The skill gives Claude Code one command,
iphone, with subcommands to capture, tap, swipe, type, and press keys. - The loop is screenshot, read the image, pick a point, tap, screenshot again. The agent works from pixels only. There is no element tree.
- The main costs: every action steals focus on the Mac, every step needs a new screenshot, and the session drops when someone picks up the phone.
Why an agent needs a phone
Some apps have no public API and no web version. To automate a task, an agent usually needs one of three things: an API, a web page that a browser agent can use, or developer access to the app, as in XCUITest or Appium UI tests. For an app that you do not own, you often have none of them. A bank or delivery app where you need to read a status is a typical example.
iPhone Mirroring removes that wall. Apple shipped it in macOS Sequoia 15. The phone must run iOS 18 or later, be locked, and be near the Mac. The Mac shows the phone screen in a normal window, and mouse and keyboard input in that window goes to the phone. For an agent, that is a screen it can see and a pointer it can move.
The interface the agent sees
A Claude Code skill is a folder with a SKILL.md file. Claude loads its body only when the task needs it. This one holds a list of commands and the loop to follow. Its front matter sets allowed-tools: [Bash, Read]. Bash runs the script. Read opens the screenshot as an image.
iphone info # window id, origin, size
iphone shot out.png # capture; prints "path WIDTH HEIGHT"
iphone tap X Y # X Y are pixels in the LAST captured image
iphone dtap X Y # double tap
iphone swipe X1 Y1 X2 Y2 # drag
iphone type "text" # keyboard text
iphone key return # return, esc, arrow-down, space, delete
iphone home # go to home screen
iphone switcher # app switcher
iphone spotlight # phone Spotlight search
The key design choice: coordinates are pixels in the last screenshot. The agent points at what it saw, and the script does the math.
How it works
1. Find the window by id
A small Swift helper, winid, calls CGWindowListCopyWindowInfo and walks the window list. It keeps the first window whose owner is iPhone Mirroring, whose title is the main or welcome window, and which is wider than 100 points and taller than 300 points. It prints five numbers: window id, x, y, width, height.
The script runs winid again on every call. So the origin is always fresh, even if someone moved the window.
2. Capture the window, not the screen
screencapture -x -o -l "$WID" -t png out.png
-l captures one window by its id. -x turns off the shutter sound and -o drops the window shadow. A capture by window id works when other windows cover the phone, and it does not bring the window to the front. Reading the phone does not disturb the person at the Mac.
The script also saves a copy as ref.png, the reference for the next tap. sips reads the pixel size, and the command prints it next to the path.
3. The loop
The skill tells the agent to follow four steps:
- Take a screenshot with
shot. - Open the PNG with the Read tool, so the model sees the image.
- Pick the target coordinates from that image.
- Tap, wait 1 to 3 seconds, and take a new screenshot to confirm.
Step 4 matters most. The agent does not assume that a tap worked. It checks.
4. Map retina pixels to screen points
On a retina display, a window capture has 2 pixels per point. A window 300 points wide gives an image 600 pixels wide. The mouse tool works in screen points. So the script converts:
screen_x = win_x + img_x * win_w / img_w
screen_y = win_y + img_y * win_h / img_h
It takes img_w and img_h from ref.png, the last capture, not from a fixed factor of 2. If there is no capture yet, it falls back to 2x. The result is cut to whole points. At 2x the error is under 1 point.
5. Send input with cliclick and System Events
cliclick is a command line tool for mouse and keyboard events on macOS. The skill uses it like this:
- tap moves the pointer to the point, then clicks (
m:thenc:). - dtap sends a double click (
dc:). - swipe is a drag: press at the start (
dd:), move to the midpoint, move to the end, release (du:). - type sends text with a 20 ms wait after each event (
-w 20 t:). - key presses one named key (
kp:).
home, switcher, and spotlight send keyboard shortcuts to iPhone Mirroring through osascript and System Events. Home, for example, is Command Shift H.
Every input command first runs activate on iPhone Mirroring and waits 0.35 seconds, because the input needs the window at the front.
Setup and permissions
- macOS 15 or later, with iPhone Mirroring connected and the iPhone locked. The skill targets Apple silicon Macs.
- cliclick, from Homebrew:
brew install cliclick. - Screen Recording permission for the terminal app, so
screencapturecan read the window. - Accessibility permission for the terminal app, so cliclick and System Events can send input.
Both permissions go to the terminal app, so every process in that terminal gets them too.
Limits and failure modes
- Focus stealing. Each tap, key, or text command brings iPhone Mirroring to the front and moves the real pointer. If you type at the same moment, your keys can go to the phone. Screenshots do not steal focus.
- The session drops. When someone picks up the iPhone, mirroring stops and the window shows "iPhone in Use". The skill tells the agent to stop and ask the user to lock the phone and click Connect. Long runs with no person nearby are fragile.
- Stale screenshots. The origin is fresh on every call, so a moved window still maps correctly. A resized window, a scroll, or an animation after the last shot does not. The skill rule is to take a screenshot before every tap.
- Pixels only. The agent sees no buttons, labels, or element ids. It must find targets by eye. A misread means a tap in the wrong place, which the next screenshot must catch.
- One step per screenshot. Each action needs a capture, an image read by the model, and the 1 to 3 second wait that the skill asks for. Long flows add up.
- One pointer. cliclick drives a single mouse pointer. The skill has no long press, no pinch, and no multi-finger gesture. A swipe is a drag with no speed control.
What it fits, and what it does not
Good fits share three traits: the task lives only in an app, it takes a few screens, and the result is visible on the screen, so the agent can check it.
- Read a value or a status from an app with no API and no web version, for example the state of an order in a delivery app.
- Change a setting that exists only in an app.
Bad fits:
- Payments and messages without a person who confirms each step. A wrong tap has a real cost.
- Anything fast or timed, such as games. The skill waits 1 to 3 seconds after each tap.
- Bulk work. One screenshot per step makes 100 items expensive.
- Unattended overnight runs. The session drops easily, and the Mac must stay free of other input.
Two security notes
Screen text is input to the model. A notification or a web page on the phone goes into the agent's context. If that text contains instructions, the model can follow them. This is the usual prompt injection risk. Treat the screen as data.
Screenshots stay on disk. By default they go to ~/.cache/iphone-mirror, and IPHONE_SHOT_DIR changes the folder. They hold whatever was on the phone screen, so clear the folder. Also, allowed-tools pre-approves Bash and Read for the turn that invokes the skill. In that turn, taps run with no permission prompt.
Sources
- Apple Support: iPhone Mirroring, use your iPhone from your Mac (requirements)
- cliclick on GitHub (command reference)
- Claude Code docs: Extend Claude with skills (
SKILL.mdandallowed-tools)