education
Supervise AI Coding Agents From Your Phone With Moshi
September 29, 2026
You start an agent on your computer. It needs ten minutes. You want to leave the desk.
Then one of two things happens. You come back early and the agent has been sitting there for twenty minutes, waiting on a permission prompt you didn’t see. Or you check your phone every few minutes to find out, which means you didn’t really leave.
Moshi fixes both. It’s an iPhone terminal app built for exactly this job, and it does three things that matter:
- The terminal doesn’t drop. It uses Mosh, so walking out of Wi-Fi onto 5G, or closing the phone for an hour, doesn’t kill your session.
- The phone tells you when an agent needs you. A push when a turn finishes, a push when it wants permission. You stop polling.
- You answer from the phone, and the agent keeps going. Approve, deny or reply without opening a laptop.
All of it rides on the private network from the previous article . Nothing is exposed to the internet.
Why plain SSH isn’t enough
I started with Termius. It works. But SSH is a single TCP connection between two IP addresses. Change either one and the connection is dead.
Phone terminals hit this all day. Lock the screen, switch from Wi-Fi to mobile data, go into an elevator. herdr (or tmux) saves your work on the computer, so nothing is lost, but you still reconnect, retype, find your window again. Ten times a day, that’s the friction that makes you stop bothering.
Mosh is built for this. It runs over UDP and doesn’t care when your address changes or the connection goes quiet. It picks up where it left off.
flowchart TB
subgraph ssh["Plain SSH"]
direction TB
a1["Phone on Wi-Fi<br/>session open"]
a2["Switch to 5G<br/>connection dies"]
a3["Reconnect, find<br/>your window again"]
a1 --> a2 --> a3
end
subgraph mosh["Mosh"]
direction TB
b1["Phone on Wi-Fi<br/>session open"]
b2["Switch to 5G<br/>brief pause"]
b3["Same screen,<br/>keep typing"]
b1 --> b2 --> b3
end
classDef good stroke:#4ade80,color:#f5f5f5
classDef bad stroke:#ef4726,color:#f5f5f5
class a2,a3 bad
class b3 good
style ssh fill:#141414,stroke:#ef4726,stroke-width:2px,color:#f5f5f5
style mosh fill:#141414,stroke:#4ade80,stroke-width:2px,color:#f5f5f5Moshi supports SSH, Mosh and Eternal Terminal, and its site says sessions survive network switches and even the app being killed. Mosh works over Tailscale for both the login step and the session itself. Set the connection type to Auto or Mosh.
One gotcha: if Tailscale SSH is on, it grabs port 22 and key login hangs for about a minute, then fails. Moshi’s docs say to turn it off and use normal SSH keys.
On the computer I run herdr . Mosh keeps the connection alive. herdr keeps the agents alive. You want both. Moshi knows about herdr and shows a picker when you connect, so you jump straight to the tab where an agent is running. (It does the same for plain tmux, if that’s what you use.)
The real upgrade: the phone taps you on the shoulder
A better terminal is nice. This is the part that changes how you work.
Moshi has a small program for your computer called moshi-hook. It installs hooks into your coding agent (Claude Code, Codex and several others are supported). When something happens, the hook tells your phone.
The events I care about:
- The agent finished its turn. Go look.
- The agent needs permission or an answer. It’s stuck until you respond.
Moshi’s docs list a few more, like session started and tool activity, but those are throttled and I mostly ignore them. Two events run my day.
Here’s what a permission request looks like:
sequenceDiagram participant A as Agent participant H as moshi-hook participant P as Your phone participant Y as You A->>H: needs permission to run a command H->>P: push with the command text P->>Y: buzz, lock screen Y->>P: approve P->>H: answer H->>A: continue
The push includes the command or question itself (Moshi caps it at 256 characters), so you can decide from the notification. You don’t open a terminal to find out what it’s asking. The daemon holds a live connection back to Moshi, so your answer goes straight to the agent. It also shows up on Live Activities and Apple Watch.
The inbox keeps one row per agent session. New events update the row instead of stacking twenty notifications. If you run several agents, that matters more than it sounds. It’s a list of who’s waiting on you, not a feed.
The practical effect: I start an agent, walk away, and my phone is the only thing I watch. If it’s quiet, the agent is working. If it buzzes, it needs me, and usually it needs one tap.
Setup, short version
On the computer:
- Install
mosh(and herdr or tmux). Make sure both are found in a non-interactive login shell. This is the most common reason Mosh fails to start. - Run
moshi-hook host setupand scan the QR code from the app. This is Easy Pair: it creates the saved connection and installs the phone’s key on your machine. - Pair the hooks with a token from the app, then run
moshi-hook install. It writes the hook entries into your agent’s config and leaves your own hooks alone. - Start a herdr (or tmux) session, reconnect from the phone, and confirm the picker appears.
Two warnings, because I’d rather you hear them from me.
Don’t share your screen while the QR code is up. Anyone who scans it before it expires gets SSH access to your machine. Treat it like a password.
Test with a real agent task before you trust it. Run something short, watch for the push, approve it from the phone, and see the agent move. Don’t assume.
Custom summaries
The automatic pushes say that something finished. Sometimes I want to know what. Moshi also has a webhook, so after a long job or a deploy I can have the agent send a one-line result: what shipped, what failed. I keep this for meaningful completions only. Routine turns are already covered by the hook, and a phone that buzzes for everything trains you to ignore it.
What else is in there
Moshi has more than I use daily. From their site: on-device voice input (handy for dictating a prompt while walking), image paste, a diff viewer, an in-app browser to preview a local dev server, and Face ID protection for the SSH keys stored on the phone. I haven’t leaned on all of them, so try the ones that fit how you work.
The stack, in one line
Tailscale gives your devices a private network. Mosh keeps the connection alive across it. herdr keeps your agents running. Moshi turns the phone into the place where you supervise them. Each layer does one job.
Try it
If you already have the private network from the first article, this is an afternoon. If you don’t, start there.
In my workshops we set this up on your own devices: private network, Mosh and herdr, hooks, pairing, and one real agent task supervised from your phone before you leave. You go home with it working, not a list of things to try later.
That’s the difference between “I can code from my phone” and “I stopped babysitting my computer.”
Interested in this?
Learn about AI Education →