APIs, integration & security — in depth

Claude Code vs Codex for Autonomous Pull Requests

Claude Code keeps work local and supervised; Codex delegates to the cloud and waits for results.

Contributing Writer · · 11 min read
Cover illustration for “Claude Code vs Codex for Autonomous Pull Requests”
Cloud Coding Agents · October 7, 2026 · 11 min read · 2,477 words

Claude Code and Codex converge on the same deliverable: a merged pull request. The path each tool takes to get there exposes a real architectural difference, not a matter of preference or polish.

Most comparisons rank coding agents by feature lists or benchmark scores, but that misses what a team actually needs to decide. Watching how each agent produces a PR raises the real question: how much control a team retains, where the execution actually happens, and what obligations appear once the PR lands in a reviewer's queue. That framing turns "which tool is better" into a more useful question: better at which class of task, for which kind of team.

The architectural fork: local-supervised versus cloud-autonomous

Claude Code and Codex did not diverge because one is more capable than the other. They diverged on philosophy. One was built to be supervised closely and configured in detail; the other was built to be handed a task and left alone until the work is done.

Claude Code runs locally, so proprietary code and sensitive work never leave the developer's machine when it executes. That constraint shapes everything downstream of it, long before you get to comparing output quality. Codex takes the opposite bet: it operates inside a network-isolated cloud sandbox that clones the GitHub repository at the moment a task is submitted, then executes changes, runs the test suite, and opens a PR without the engineer present for any of it. Submitting a task to Codex and waiting for a notification is closer to filing a ticket with a contractor than pairing with a colleague.

The direction of travel reinforces the split. In July 2026, the standalone Codex desktop app was folded into the ChatGPT desktop app, consolidating both products under one roof. That move signals where OpenAI sees Codex's primary surface: cloud-first, not terminal-first. Claude Code, by contrast, remains anchored to the local environment, with memory and configuration files like CLAUDE.md and AGENTS.md doing the work of carrying context across sessions. How each of those files actually works needs its own treatment later, but the names are worth planting here, because they say the same local-versus-cloud split in miniature.

A third path lets teams use coding agents without committing to one architecture. A third-party tool runs coding agents, including Claude Code and Codex, inside isolated virtual machines pre-loaded with a team's codebase and tooling, letting engineers trigger runs from Slack, Linear, or GitHub without committing to a single harness.

How Claude Code produces a pull request

Claude Code's route to a pull request is interactive, and it triggers on events. An engineer, or an automated hook acting on the engineer's behalf, points the agent at an issue. From there it reads the relevant code, drafts a plan, implements changes, runs tests, and opens the PR, with the engineer free to redirect the work at any point along the way.

Teams wire that event-triggered design into their workflows in several concrete patterns. A nightly test run can auto-file GitHub issues for failures that Claude Code then picks up. A weekly dependency-update PR can be scheduled without a human initiating it each time. A CI step can write release notes straight from a changelog diff, and a GitHub Issue label can start a full agent run through GitHub Actions. Each of these patterns treats Claude Code as a responsive participant in the development pipeline.

The clearest expression of Claude Code's supervised design is Plan mode. Before writing any code, Claude Code can explore a codebase and propose a structured implementation plan, and the engineer reviews and approves it before execution proceeds. That approval step is a mid-run checkpoint that has no equivalent in Codex's asynchronous model, where the task is already running in isolation by the time anyone sees the first line of generated code.

How Codex produces a pull request end to end

Codex's path to a pull request is fully asynchronous from the first step. An engineer delegates a task description, the cloud sandbox clones the repository, executes the requested changes, runs the test suite, and opens a PR, all without a human present or reachable during the run.

Once that PR is open, the task Codex was given is technically complete, but the interaction doesn't have to stop there. Codex can respond to review comments through @codex mentions placed directly on the pull request, such as "@codex fix the P1 issue," which starts a cloud chat session that can push fixes back to the same branch without requiring a brand-new task submission. A related but distinct capability lets a reviewer comment "@codex review" on a pull request, prompting Codex to post a standard GitHub review that flags only P0 and P1 issues, while a separate "@codex security review" command runs a dedicated security pass. Both are review-assist functions layered on top of the generation workflow, not substitutes for it.

Codex connects natively to VS Code, JetBrains, GitHub, GitLab (Beta), Slack, Linear, and any external or internal tool reachable through MCP. Tagging "@codex review" in a GitHub PR comment triggers automated reviews or patches, subject to the account's subscription limits. Getting this running requires a one-time repository connection and some configuration in Codex settings, but no separate CI pipeline work.

OpenAI's own internal usage gives a sense of what this throughput looks like at scale. Starting from an empty git repository, roughly 1,500 pull requests were opened and merged over five months by a team that began with three engineers, averaging 3.5 PRs per engineer per day, with throughput increasing as the team grew to seven. That figure shows what a fully async pipeline can sustain when conditions are favorable. It says less about what any given team will see on its own codebase, with its own review standards.

For teams that want delegation without giving up a local view of their codebase, platforms like Replicas let Claude Code run inside a cloud-isolated environment instead of on a developer's laptop, extending supervised-mode agents into fully async workflows while keeping the engineer in control of what the agent sees.

Models, context windows, and benchmark results

On benchmark leaderboards, Claude Code and Codex sit close enough that picking the right model and task class affects outcomes more than the small score gap between them. The cost structure is where the two tools actually pull apart, especially on large-context work.

As of September 22, 2026, Claude Code defaults to Claude Opus 5.5, while Codex defaults to GPT-6.1 Sol. The recommended Codex model costs half of Claude Code's price per token, and OpenAI has described the new rates as promotional, set to run at least through November 21, 2026. Pricing moves fast in this market, so the landscape could look different again by the time the promotional window closes.

The frontier models behind each tool, Claude Fable 5.1 and GPT-6 Astra, share the same headline rate per million tokens on both input and output, but they diverge in two places that matter for real engineering workloads. Fable 5.1 charges less than Astra for cache reads, and the two handle long-context surcharges differently: Astra doubles its input rate above a certain token threshold, while Fable 5.1 holds to a flat rate regardless of context length. Both tools carry roughly the same large context window, but because Claude Code bills flat, it removes a pricing cliff that would otherwise make the cost of large-codebase runs unpredictable. For a team running agents against sprawling, interconnected repositories, that predictability is a planning advantage independent of raw model quality.

Memory and context persistence: CLAUDE.md versus AGENTS.md

The single-run comparison only tells part of the story. Whether autonomous PRs get better over time, rather than requiring the same corrections on the tenth run as on the first, depends on the memory architecture each tool ships with.

Claude Code's memory system centers on CLAUDE.md, with MEMORY.md serving as a companion file that holds evolving state: recent migrations, active decisions, what changed in the codebase last week. The two files work together. CLAUDE.md content gets delivered as a system-reminder inside the conversation rather than embedded in the system prompt itself, which gives it high priority but not deterministic enforcement, and a model's compliance with CLAUDE.md instructions is probabilistic. Permission rules give you the deterministic enforcement layer, and they work no matter what the memory file says.

Codex relies on AGENTS.md to carry equivalent context forward into async runs, and the investment required to set that file up properly is not optional if PR quality matters. When teams put real time into an initial AGENTS.md setup, they get measurably better first-draft PR quality across subsequent Codex tasks. That upfront configuration cost pays forward directly into async throughput, because AGENTS.md is the only channel you have for shaping a Codex run before the PR appears. There is no mid-session conversation to redirect it, the way there is with Claude Code's Plan mode or its Routines and scheduled sessions. Where Claude Code offers an engineer the chance to interject while the agent works, Codex offers only the chance to write better instructions in advance.

How review burden shapes each tool's fit

Autonomous PR generation runs into a scaling problem that has nothing to do with model quality: PR volume increases faster than review capacity does, and Claude Code and Codex each produce that imbalance differently.

H1 2026 data captured the shape of the problem directly: PR volume rose industry-wide over that period, and incidents per PR rose alongside it. Faster output arrived paired with declining reliability per PR. Autonomous PR generation needs the tool's output rate matched to the task's actual risk tolerance, since the review queue, not the generation step, becomes the constraint once agents are producing PRs faster than a team can read them.

Codex's async, high-throughput design is built for volume, parallel tasks, and bulk PRs generated from a single product spec, and that same design amplifies the review burden if it runs without governance controls in place. Codex builds in enterprise audit logs and an approval-policy layer so it can manage that risk at scale. Claude Code's interactive model keeps review burden lower per individual PR because the engineer is there, intervening during execution. That advantage holds only as long as a human stays in the loop; once Claude Code runs inside a fully automated pipeline with no one watching, the same review-burden dynamics that apply to Codex apply to it as well.

Vendor-reported productivity figures, including OpenAI's internal PR counts and customer-reported reductions in review turnaround, show what you can achieve if you configure a deployment well. The METR methodological caution is relevant here: those numbers show what is possible under favorable conditions, not what a given team should expect by default. The review-burden data from H1 2026 is a useful counterweight to keep in view alongside any throughput claim.

One pattern in the usage data shows how teams can resolve this tension. Staff-plus engineers reach for coding agents more often than regular engineers or engineering managers do. Autonomous PR generation appears to work best where the reviewing engineer already has the judgment to catch what the agent misses, which turns the question of fit away from the tool alone and toward who is on the other end of the review.

Fit by task class, team size, and infrastructure posture

None of this amounts to a winner-takes-all verdict. The choice between Claude Code and Codex is a task-routing decision, and the two architectures turn out to be complementary more often than they compete directly for the same job.

Claude Code is the better fit when a task demands deep reasoning across a large, interconnected codebase, when a team needs to supervise or redirect the agent mid-run, when proprietary code cannot leave the local machine for contractual or security reasons, or when the goal is a surgical, high-quality edit with context costs that stay predictable even on large files. Small and mid-size teams tilt heavily toward this profile: Claude Code is used by a large majority of the smallest companies and teams, and it tends to enter most organizations over roughly nine months through individual engineer advocacy.

Codex fits better when the task is well-specified and lower-risk, when the work benefits from running many instances in parallel, when a team wants fully async throughput, submitting a task description and receiving a finished PR in return, or when enterprise governance requirements, audit logs, approval gates, and repository restrictions among them, are non-negotiable. Cisco reported roughly a 20% reduction in project build times and a 10 to 15 times increase in defect resolution throughput after rolling out Codex. Duolingo reported a significant reduction in median review turnaround alongside a large increase in PR volume. Both outcomes are consistent with what an async, high-throughput architecture is designed to produce.

Senior teams most often run both tools side by side rather than picking one: Claude Code as the daily driver for design work and surgical edits, Codex for bulk parallel PRs generated from a single product spec. It's a direct match between architecture and task class, applied twice within the same organization.

Running both agents at scale with Replicas

Running Claude Code and Codex side by side raises an operational question neither tool answers on its own: how a team delegates work to either agent, or both, without building separate infrastructure for each one. Replicas addresses that by letting engineering teams delegate from Linear, Slack, GitHub, GitLab, or a dashboard, without requiring a new workflow layered on top of the ones already in place.

Each delegated task runs inside its own Linux VM, with the agent doing the work chosen for the task. The output matches what either tool produces natively: a pull request, a reply on an existing thread, or a recording of the session, ready to merge, comment on, or send back for another pass. The dashboard offers a mode close to Claude Code's Plan mode, letting an engineer watch the diff, steer the plan, and pair with the agent directly when the task calls for that kind of attention, while Slack and Linear integrations let a team start a run wherever the work already surfaces, whether that's a ticket turning into an active agent run or a ping dropped into a channel.

Codex's network-isolated sandbox is a real constraint on which tasks are viable: no outbound internet access during execution means the agent cannot install packages from the network mid-run. Teams whose work depends on offline agent execution, or on coordinating multiple agents across a single workflow, often run Codex, or another agent entirely, through a platform like Replicas, which handles sandbox isolation and the integration layer separately from whichever agent's native surface is doing the generation. A mid-sized engineering group using Replicas, for instance, is reported to ship over 30% of its pull requests through the platform, a sign that the delegate-sandbox-PR loop can carry real production load.

More in Cloud Coding Agents