Codex CLI remote: running the agent on the Mac you left behind
The setup is common enough. Codex CLI is installed on the Mac at the desk, the repository is checked out there, the environment variables and the local database are all on that machine. The work is somewhere else: a train, a client office, a laptop that has none of it. The question is not whether the agent can run somewhere else. It is which of several unrelated mechanisms called "remote" actually matches the situation, because they fail in different ways.
Three routes exist today, and the phrase "codex cli remote" gets used for all three. One is a terminal session over SSH into a machine that belongs to you. One is a pairing between the ChatGPT mobile app and a desktop host, where the phone sends prompts and the desktop does the work. One is a container on OpenAI infrastructure that never touches the Mac at all. Choosing correctly means knowing which one holds your uncommitted changes and which one goes dark when the lid closes.
Three things go by the name remote
The differences that matter are not about features. They are about where the files live and what happens when the connection drops.
| Route | Where the code lives | What stops it | Best for |
|---|---|---|---|
| SSH into your own machine | The machine you left behind | The host sleeping, or the session closing without a multiplexer | Work that depends on local state, credentials, or a running service |
| Phone paired to a desktop host | The paired desktop | The host sleeping or going offline | Approving and steering a run while away from the desk |
| Cloud container | A fresh checkout on remote infrastructure | Nothing local, but nothing local is available either | Self-contained tasks on a pushed branch |
The official documentation is direct about the first constraint on the pairing route. The host has to be a machine that stays on.
Repository files and local documents come from the connected host. Source: learn.chatgpt.com
That single sentence rules out a large share of the setups people try first. Pairing a phone to a laptop that sleeps in a bag is pairing to nothing. The documentation names the alternative plainly: an existing laptop or desktop is convenient but stops when asleep, a dedicated always-on computer stays available, and an SSH host suits a remote development environment. A Mac mini left running is the usual answer for anyone who wants this to work more than once.
The SSH route, and the two places it breaks
For a machine you own and can reach, SSH is the shortest path. Remote Login has to be on first, under Sharing in System Settings, and the account has to be listed among the users allowed to log in. Apple documents the switch and the exact address to connect to in the macOS User Guide. After that, codex behaves the way it does in any terminal, because it is a terminal program.
The first break is the login. Codex CLI prompts on first launch to sign in with ChatGPT, and that flow wants a browser. On a headless session there is no browser to open, and the callback lands on a port on the remote host that nothing local is listening to. The usual fix is to forward the port when opening the session, so the callback reaches the browser on the machine in front of you:
ssh -L 1455:localhost:1455 mac-mini.local
The port number is whatever the sign-in prompt names. Forwarding it turns a dead end into an ordinary login. An API key is the other option, but it moves billing onto usage-based API pricing rather than a ChatGPT plan, which is a different decision than a connectivity one.
The second break is the session itself. An SSH session is a process tree, and closing the laptop lid ends it. A twenty minute refactor ends with it. This is the oldest problem in remote work and it has the oldest fix: start the agent inside a terminal multiplexer so the process survives the disconnect.
ssh mac-mini.local
tmux new -s codex
codex
Detaching with the multiplexer prefix and then d leaves the run going. Reconnecting later with tmux attach -t codex picks up mid-sentence, scrollback included. screen ships with macOS at /usr/bin/screen if installing tmux is not an option. Skipping this step is the single most common reason a long run appears to have failed when it was simply killed.
Keeping the host awake is a different job from keeping it reachable
These two get conflated and they have separate fixes.
Reachable means the network can find the machine. On a LAN, the .local hostname works. From outside, exposing SSH to the open internet is the option to avoid, and the Codex documentation says so about its own server component: avoid exposing the app server to public networks, and use a VPN or mesh networking instead. A mesh VPN gives the Mac a stable private address that follows it between networks, which also removes the need to think about the router.
Awake means the machine is not asleep when the connection arrives. An incoming SSH connection does not wake a sleeping Mac by itself, and a long agent run can be cut short by idle sleep partway through. macOS ships a tool for exactly this. Wrapping the session in caffeinate holds the assertion only for as long as the command runs:
caffeinate -dimsu tmux new -s codex
The flags block display sleep, idle sleep, disk sleep, and system sleep, and declare the user active. The -s assertion is only honoured while the machine is on AC power, which is worth knowing before trusting it on battery. Whether it took effect is visible in pmset -g assertions, where PreventSystemSleep flips to 1. Setting the whole machine to never sleep in System Settings works too, and costs more power on a machine that mostly idles.
Approvals are what the phone is actually for
The pairing route is not a small terminal on a small screen. It is built around the part of an agent run that needs a human: the moment the agent asks to run something. Setup runs from the desktop app's settings, where remote access is enabled and a QR code is shown, and the phone scans it. From there the phone lists the host and its projects.
What the phone contributes is narrow and deliberate. The documentation describes the host as supplying the shell, the MCP servers, browser access and the security controls, while the phone supplies prompts and approvals. In practice that means reviewing requested commands before the agent continues, sending a correction mid-run, and inspecting changed files, diffs and test results afterwards.
On the CLI side the same decision is made once rather than per prompt. The /permissions command chooses when Codex may edit files or run commands without asking, and shows the active sandbox and the writable roots. An AGENTS.md file in the repository carries the standing instructions so they do not have to be retyped from a phone keyboard. The combination that works for unattended runs is a narrow writable root plus explicit permissions, not a blanket approval, because the cost of a wrong command is paid on the machine holding the uncommitted work.
Reading the result is a file problem, not a chat problem
The part that stays awkward is not starting the run. It is the twenty minutes afterwards, when the agent has touched thirty files and the question is what actually changed. A chat transcript is a poor shape for that. A diff is better. A directory listing sorted by modification time is often better still, because it shows the files nobody mentioned.
This is where the window count starts to hurt. Checking a path means a file browser, reading a diff means a terminal, and asking a follow up question means the agent again. On a desktop that is three windows and a lot of switching. A file manager with a built-in terminal collapses two of them, and the reason it matters for remote work specifically is that the same layout has to work on a small screen. The trade-offs between the available options are laid out in Compared with other file managers, and the terminal and agent panes themselves are described in Features.
The phone side has its own version of the problem. Approving a command is a two second action, but deciding whether to approve it usually means looking at a file first. Being able to read the tree and open a file from a phone, rather than only reading what the agent chose to quote, is the difference between approving and guessing. That specific workflow is what From iPhone and iPad covers.
A fourth consideration sits underneath all of this and rarely gets written down: which route leaves an audit trail. An SSH session in a multiplexer keeps its scrollback, so the commands the agent ran are still readable an hour later. A phone-driven session keeps the approvals. A container keeps the diff. Deciding in advance which record you will want is easier than reconstructing it afterwards from a repository that has already moved on.
Where the cloud route stops being the same question
The third route is worth separating out because it answers a different problem. A cloud container starts from a fresh checkout of a branch that has been pushed, runs in an environment defined by configuration rather than by whatever happened to be installed on the Mac, and returns a diff. Nothing on your machine has to be awake, reachable, or even switched on.
That is genuinely useful for a certain shape of task. A dependency bump, a test suite that needs to run against a clean tree, a mechanical refactor across a package with no local services involved. The container is reproducible in a way a personal Mac never is, and two of them can run at once without fighting over the same port.
It is the wrong route the moment the task depends on something that only exists locally. A staging database reachable only from the office network, credentials in a keychain, a half-finished branch that was never pushed, a file sitting in a downloads folder. None of that is present in a fresh container, and no amount of prompting conjures it. The practical test is simple: if the task would fail on a colleague's freshly cloned copy, it will fail in a container for the same reason.
There is also a working-style difference that gets underrated. The SSH and pairing routes keep the run in front of you while it happens, which means a wrong turn can be stopped after ten seconds. The cloud route is closer to submitting a job and reading the result, so a misunderstood instruction costs the whole run rather than a sentence. For work where the specification is clear, that is fine. For exploratory work where the next instruction depends on what the last one found, it is the slower arrangement despite looking like the faster one.
What to change first
Pick the host before touching anything else, because every other decision follows from it: a machine that stays awake and reachable, not the laptop that travels. Then put the agent inside tmux under caffeinate, which costs one line and removes the most common failure. If the plan involves approving work from a phone, make sure you can read the files from there too, not only the diff the agent chose to show, which is what Atriens was built around.
Frequently asked questions
Does Codex CLI keep running if the SSH connection drops?
No. The agent is a child of the SSH session, so closing the laptop or losing the network ends it mid-task. Start it inside tmux or screen and detach instead, then reattach later to find the run still going with its scrollback intact.
How do you sign in to Codex CLI on a headless machine?
The ChatGPT sign-in flow opens a browser and waits for a callback on a local port. Forward that port when opening the SSH session, with ssh -L <port>:localhost:<port> host, so the callback reaches the browser on the machine in front of you. An API key avoids the browser entirely but bills as API usage rather than under a ChatGPT plan.
Can a phone drive the agent without leaving a computer switched on?
Not on the pairing route. Repository files and local documents come from the connected host, so the host has to be awake, online and signed into the same account. A cloud container is the route that needs no machine of yours running, at the cost of having none of your local state.
Is it safe to open SSH to the internet for this?
Exposing a port directly is the option to avoid, and the Codex documentation recommends a VPN or mesh networking rather than public exposure. A mesh VPN gives the Mac a stable private address across networks, which keeps the host reachable without opening anything on the router.
Why does a long run stop on its own with no error?
Idle sleep is the usual cause. An agent working through a large refactor produces no user input, so the machine sleeps partway through and the session dies quietly. Wrapping the session in caffeinate -dimsu holds it awake for the duration, and pmset -g assertions confirms the assertion is active.