Claude Code vs Codex: the differences you notice after a week
For the first day the two look interchangeable. Install a binary, sign in with a subscription, open a project directory, describe a change, watch a diff appear. The claude code vs codex question only becomes answerable around day four, when the differences that matter show up: where configuration lives, what runs without asking, what happens when the terminal is closed, and which one can be dropped into a script without a rewrite. Those are the four places worth looking, and none of them are visible in a first-day demo.
Getting each one onto a Mac
Both install without a package manager if that is the preference, and both have a package manager route.
Codex CLI installs through a standalone script for macOS and Linux, through npm install -g @openai/codex, or through brew install --cask codex. The same script is used to update. Signing in happens on first run inside a project directory, with Sign in with ChatGPT as the default path and other methods available.
Claude Code installs its own native binary, with claude install accepting a specific version, stable, or latest, which is genuinely useful when a release introduces a regression mid-project. Signing in uses claude auth login, with flags for pre-filling an email address, forcing SSO, or signing in through the Console for API billing rather than a subscription. There is also claude doctor, which prints installation and settings diagnostics without starting a session, and it is the first thing to run when something is behaving oddly.
The install step is not where either one wins. It is worth mentioning only because pinning a version and diagnosing an install are the two things people reach for in week two and then have to search for separately.
What each costs, including the tier most comparisons skip
Figures below come from each company's own pricing page.
| Tier | Codex | Claude Code |
|---|---|---|
| Free | $0, limited capability | Not included on the free plan |
| Light paid | Go, $8 per month | No equivalent tier |
| Standard | Plus, $20 per month | Pro, $20 per month, or $17 billed annually at $200 up front |
| Heavy | Pro, from $100 per month, at 5x or 20x Plus limits | Max, from $100 per month, at 5x or 20x Pro limits |
| Key-based | API key, suited to shared environments like CI | Console sign-in for API billing |
The eight dollar tier is the one that changes the calculation. For someone who wants an agent in the terminal a few times a week rather than all day, Go has no counterpart on the Claude side, where Claude Code starts at the twenty dollar plan. In the other direction, Claude's Pro plan can be paid annually at a lower monthly equivalent, which the Codex individual tiers do not currently offer.
One detail that matters for anyone doing capacity planning: ChatGPT Work usage inside ChatGPT draws on the same pricing, credits, and usage limits as Codex, so a team already using one is spending from a shared pool rather than two.
There is also a dated change worth knowing before pinning a workflow to a specific model: GPT-5.5 retires from ChatGPT, ChatGPT Work, and Codex on all plans on October 14, 2026, with the OpenAI API unaffected. Any script that names a model explicitly needs a review before then.
Configuration, which is the first real divergence
This is where day four arrives.
Codex keeps settings in a TOML file, at ~/.codex/config.toml for personal defaults, with project overrides in a .codex/config.toml inside the repository, loaded only for projects marked as trusted. Profiles selected with --profile add another layer, and a system-level file can exist at /etc/codex/config.toml. The resolution order is published and strict: CLI flags first, then project files from the repository root down to the current directory with the closest winning, then profiles, then user config, then cloud-managed defaults, then system config, then built-in defaults. On managed machines an organization can enforce constraints through a requirements.toml, for example forbidding an approval policy of never.
Claude Code splits the same job in two. Behaviour settings live in settings files, and instructions live in markdown: CLAUDE.md for project context, with AGENTS.md read either on its own or alongside it, plus a .claude/rules/ directory for rules scoped to particular file types. On top of that there is an auto memory that writes its own notes based on corrections given during a session.
The trade-off is real and it is about determinism. A TOML file with a documented precedence chain is easier to reason about, easier to diff, and easier to enforce across a fleet of machines. A markdown instructions file is easier to write, carries reasoning a config key cannot express, and is read by more than one tool, since a well-written AGENTS.md is portable.
The caveat applies to the markdown side and is stated plainly in the documentation: instructions in these files are treated as context, not as enforced configuration. Something that must be blocked regardless of what the model decides belongs in a hook, not in a sentence.
Sandboxing and approvals on macOS
Both tools land in the same place conceptually, with two layers: what is technically possible, and when the tool stops to ask. Both lean on the operating system rather than on good intentions.
Codex runs with network access off by default, and locally uses an OS-enforced sandbox that typically limits writes to the current workspace. The documented Auto preset combines a workspace-write sandbox with on-request approvals, so reads, edits, and commands inside the working directory proceed automatically while leaving the workspace or reaching the network requires approval. Switching to read-only for planning is a /permissions command. The older untrusted approval policy has been retired, and the stricter behaviour is now expressed by marking a project's trust level in user config, which also disables project-local configuration for that path.
Claude Code's sandboxed Bash tool uses the built-in Seatbelt framework on macOS, with nothing to install, and covers Bash commands and their child processes. Sandboxed commands can write to the working directory, a per-user temporary directory, and any directories added explicitly. The first time a command needs a new network domain, approval is requested for that domain. A /sandbox panel exposes the mode, whether commands that fail under the sandbox may fall back to running unsandboxed, and the resolved configuration. Commands that cannot be sandboxed fall back to the ordinary permission flow and are labelled as unsandboxed in the prompt.
| Behaviour | Codex CLI | Claude Code |
|---|---|---|
| Network default | Off | Per-domain approval on first use |
| Write scope by default | The active workspace | Working directory, temp directory, added directories |
| Mechanism on macOS | OS-enforced sandbox | Seatbelt |
| Switch to read-only | /permissions |
Permission modes and the /sandbox panel |
| Organization enforcement | requirements.toml on managed machines |
Managed settings files |
Read that table as evidence that the safety argument between the two is close to a tie. The practical difference is in wording and in which panel a setting hides behind, not in whether an agent can quietly reach the network.
Working when the terminal is not in front of you
This is the widest gap in day-to-day use, and it runs in opposite directions.
Codex moves the work off the machine. Codex cloud runs in isolated containers managed by OpenAI, using a two-phase model: a setup phase that can reach the network to install dependencies, then an agent phase that is offline by default unless internet access is enabled for that environment. Secrets configured for a cloud environment are available during setup only and are removed before the agent phase begins. Codex is also available on the web, in an IDE extension, and on iOS.
Claude Code keeps the work on the machine and moves the controls. Remote Control connects claude.ai/code or the mobile app to a session running locally, so the filesystem and local MCP servers stay available while the steering happens elsewhere. It is available on Pro, Max, Team, and Enterprise plans, off by default on Team and Enterprise until an owner enables it, and unavailable with API keys or a custom base URL. Claude Code also has a background session model, with claude agents for monitoring and dispatching parallel sessions, claude attach for attaching to one in a terminal, and a supervisor process with its own status and stop commands.
Which of those is better depends entirely on whether the work needs the local machine. A task touching a half-finished branch, a running dev server, or a folder of files that exist nowhere else has to be local. A task on a repository that is not cloned locally is better off in a container.
Scripting, and repeatable runs
Both are usable non-interactively, which is the requirement for putting an agent in CI or in a nightly job.
Codex exposes codex exec for non-interactive runs, codex resume for reopening a recent chat from the current repository or searching older ones, codex --image for passing a screenshot or diagram with the first prompt, and codex --search for switching a run to live web search when the task depends on current documentation. Subagents split larger investigations, and a dedicated review mode reports findings against uncommitted changes, a commit, or a base branch without modifying the working tree.
Claude Code uses claude -p for a single query that exits, accepts piped input on the same flag, claude -c to continue the most recent conversation in the current directory, and claude -r to resume a session by ID or name. Agent view can print active sessions as JSON for scripting.
For anyone building automation, the deciding question is narrow: whether the job needs one clean answer and an exit code, which both handle, or whether it needs to rejoin a long conversation later, where the resume semantics differ enough to be worth testing with a throwaway repository before committing.
What a week reveals that neither one fixes
After a week with either tool, the remaining friction is usually not the agent. It is the traffic around it.
Work that is about files rather than code keeps breaking the loop. Open the folder to see what is actually there. Copy the path. Paste it into the shell. Get a listing. Read it in the terminal. Decide. Go back to the folder to confirm. Each of those steps is quick, and none of them get faster when the model improves, because the cost is in the window switching rather than the thinking. On a few hundred files, the switching outlasts the inference by a wide margin.
Neither Codex nor Claude Code is built to remove that, and neither claims to be. A file manager with a built-in terminal is the piece that closes it, by making the folder already on screen the context for the command. What that arrangement looks like is set out on the Features page, and the limits of what to expect from it are stated on the FAQ page.
What to change first
Spend one week with each on the same real repository and keep a tally of interruptions rather than judging the diffs, because after day three the diffs look similar and the interruptions do not. Then write the rules down in a file the tool reads, AGENTS.md or a TOML config, so the second week is not a rerun of the first. If most of the interruptions turn out to be moving between a folder and a shell, that is a desk problem rather than a model problem, and Atriens is aimed at that half.
Frequently asked questions
Is there a cheaper way in than twenty dollars a month?
On the Codex side there is a Go plan at eight dollars a month and a free tier with limited capability. Claude Code is not included on the free Claude plan, so the entry point there is the twenty dollar Pro plan, which can be paid annually at a lower monthly equivalent.
Can both be configured per project rather than globally?
Yes, in different formats. Codex reads a .codex/config.toml inside the repository for trusted projects, resolved against a documented precedence chain. Claude Code reads CLAUDE.md and AGENTS.md from the repository plus a .claude/rules/ directory, and treats those as context rather than enforced settings.
Will either of them reach the network without permission?
Not by default. Codex runs with network access turned off and asks before a command needs it. Claude Code's sandboxed Bash tool asks for approval the first time a command needs a new network domain. Both also allow an organization to enforce limits from managed configuration rather than leaving it to each developer.
Which one is better for a long job left running overnight?
If the job needs the local machine, a locally running session steered remotely fits better, and Claude Code's Remote Control plus its background session commands cover that. If the job only needs a repository and a container, Codex cloud avoids depending on a laptop staying awake at all.
Is it worth running both?
It can be, because the two behave differently under the same prompt and a second opinion on an awkward refactor has real value. The cost is two configuration systems to maintain, so the reasonable version is keeping the rules in AGENTS.md, which both can read, and accepting that the per-tool extras will drift.