Most "Codex vs Claude Code" comparisons quote vendor benchmarks, or one third-party token figure that gets copied from page to page. Few of them publish the tasks, the prompts or the commit, so their numbers cannot be checked. We did the boring version instead: four real tasks on one pinned open-source repository, each run twice per CLI, headless, with the exact commands, prompts and numbers below, and the raw data to download. Measured on 2026-10-06 with Claude Code 2.1.291 (claude-opus-5-5) and Codex CLI 0.160.0 (gpt-6.1-sol), default models, default effort.

TL;DR

Codex vs Claude Code at a glance

Claude Code Codex CLI Gemini CLI (not benchmarked)
Maker Anthropic OpenAI Google
Default model (measured) claude-opus-5-5 gpt-6.1-sol –
Interfaces Terminal, VS Code and JetBrains extensions, desktop app, web Terminal, IDE extension, cloud tasks from ChatGPT Terminal
Headless command claude -p (docs) codex exec (docs) gemini -p
Machine-readable output --output-format json or stream-json --json (JSONL events) –
Permission model Permission modes (default, acceptEdits, plan, bypassPermissions); --dangerously-skip-permissions skips every prompt Sandbox modes (read-only, workspace-write, danger-full-access) plus an approval policy; --dangerously-bypass-approvals-and-sandbox disables both –
Instruction file CLAUDE.md AGENTS.md GEMINI.md
Entry plan Claude Pro, $20/month ChatGPT Plus, $20/month (Free and Go at $8 for light use) –
Pay per token API key, $4 / $20 per million input / output tokens (Opus 5.5) API key, $2 / $10 per million input / output tokens (GPT-6.1 Sol) –
License Proprietary Open source (Apache-2.0) Open source (Apache-2.0)

Prices are from the official pages on 2026-10-06: claude.com/pricing, ChatGPT plans for Codex and the OpenAI API pricing.

How we tested (and how you can rerun it)

Everything needed to check or rerun the benchmark is in the raw data archive (180 KB): the four prompts, the hidden acceptance tests, the run and check scripts, the full JSON transcript of each of the 16 runs and the computed metrics.

The repository and the four tasks

The repository is python-humanize/humanize (MIT): 1,725 lines of Python in src/, about 860 tests that run in about 7 seconds, pinned at commit 785e5dcc0d0308ad0dff3f6cc0faa7085ad0375b. Every run got a fresh git worktree of its starting commit and its own virtualenv, so no run could see another one's work.

Task Starting commit Objective check
Bug fix from a real issue f971127, the parent of the fix for PR #346 The upstream regression test test_intword_rounding_rollover, copied in after the run, plus the full suite
Feature from a spec 785e5dc 16 hidden acceptance tests written before the runs (checked against a reference implementation), copied in after the run, plus the full suite
Refactor, no behavior change 785e5dc tests/ untouched, old and new import paths return the same objects, docs updated, doctests and full suite green
Code review 785e5dc plus one commit adding parse_size() with 3 planted bugs Planted bugs found, false alarms

The review commit adds a parse_size() function, the inverse of naturalsize(), plus 12 passing tests that avoid every bug. To recreate it, append this to src/humanize/filesize.py (with import re at the top) and export it from humanize/__init__.py:

_SIZE_PATTERN = re.compile(r"^\s*(-?(?:\d+)+(?:\.\d+)?)\s*([A-Za-z]*)\s*$")


def parse_size(value: str) -> int:
    """Parse a size written by `naturalsize` back into a number of bytes."""
    match = _SIZE_PATTERN.match(value)
    if match is None:
        msg = f"invalid size: {value!r}"
        raise ValueError(msg)
    number, unit = float(match.group(1)), match.group(2)

    if unit in ("", "B", "Byte", "Bytes"):
        return int(number)
    if unit in suffixes["decimal"]:
        return int(number * 1000 ** (suffixes["decimal"].index(unit) + 1))
    if unit in suffixes["binary"]:
        return int(number * 1024 ** suffixes["binary"].index(unit))
    msg = f"unknown unit: {unit!r}"
    raise ValueError(msg)

The three planted bugs:

  1. Logic: the binary branch forgets the + 1, so parse_size("1.0 KiB") returns 1.
  2. Edge case: int() truncates the float product, so parse_size("8.2 MB") returns 8199999.
  3. Security: the nested quantifier (?:\d+)+ backtracks exponentially. parse_size("1" * 24 + "!") takes 2.6 seconds on our machine and each extra digit doubles it: a denial of service on untrusted input.

The exact commands and prompts

Setup of one run:

git clone https://github.com/python-humanize/humanize && cd humanize
git worktree add --detach ../run-1 785e5dcc0d0308ad0dff3f6cc0faa7085ad0375b
cd ../run-1
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -e '.[tests]' pytest

The two commands, same prompt text for both, default model and effort, permissionless:

# Claude Code 2.1.291
env -i HOME="$HOME" PATH="$PATH" IS_SANDBOX=1 \
  claude -p "$PROMPT" --output-format stream-json --verbose \
    --dangerously-skip-permissions \
    --setting-sources project,local --strict-mcp-config --no-session-persistence < /dev/null

# Codex CLI 0.160.0
env -i HOME="$HOME" PATH="$PATH" \
  codex exec --json --dangerously-bypass-approvals-and-sandbox \
    --ignore-user-config --skip-git-repo-check -C "$WORKTREE" "$PROMPT" < /dev/null

A few flags deserve a word:

Every prompt ended with the same line: "The project virtualenv is ready in .venv: run the tests with .venv/bin/python -m pytest -q. Work only inside this repository. Do not commit." The four prompts, in full:

[Bug fix]
Fix this bug, reported against this repository (humanize):

`intword()` decides whether rounding pushed a value up into the next magnitude with:

    if not largest_ordinal and rounded_value * power == powers[ordinal + 1]:

For values above ~10**22, `rounded_value * power` is evaluated in floating point and no longer
equals the exact `powers[ordinal + 1]`, so the carry is skipped and the number is rendered against
the lower magnitude:

    >>> import humanize
    >>> humanize.intword(10**24 - 1)
    '1000.0 sextillion'      # expected '1.0 septillion'
    >>> humanize.intword(10**27 - 1)
    '1000.0 septillion'      # expected '1.0 octillion'

The same happens at 10**30 and 10**33.

Add a regression test.
[Feature]
Add two keyword-only parameters to `humanize.natural_list` (src/humanize/lists.py):

- `conjunction: str = "and"`: the word placed before the last item.
  `natural_list(["a", "b", "c"], conjunction="or")` returns `"a, b or c"`.
- `max_items: int | None = None`: when set and the iterable has more than `max_items` items, keep
  the first `max_items` items and replace the rest with "N more", joined with the conjunction.
  `natural_list(["a", "b", "c", "d"], max_items=2)` returns `"a, b and 2 more"`;
  `natural_list(["a", "b", "c", "d"], max_items=1)` returns `"a and 3 more"`. When the iterable has
  `max_items` items or fewer, the output is the same as without the parameter.
- `max_items` lower than 1 raises `ValueError`.
- The default behavior must not change. Update the docstring (with examples) and add tests.
[Refactor]
Refactor, with no behavior change: move `scientific()` and `metric()` out of
src/humanize/number.py into a new module src/humanize/notation.py, together with the helpers and
constants only they use.

- `humanize.scientific`, `humanize.metric`, `humanize.number.scientific` and
  `humanize.number.metric` must keep working.
- Helpers still needed by number.py must not be duplicated.
- The new module must appear in the API documentation (docs/ and mkdocs.yml) like the other modules.
- Do not modify the existing tests; the whole test suite must pass.
[Code review]
You are reviewing a pull request. The change under review is the last commit of this repository
(see `git show HEAD`).

Review it for correctness, edge cases, security and performance. Report each problem as a numbered
list: file and line, severity (high, medium or low), what is wrong, and a concrete input that
demonstrates it. Only list real problems. Do not modify any file.

The review prompt does not end with the test line, since the reviewer must not change anything.

What we measured, and what "cost" means here

"API-equivalent" matters. With an API key you pay exactly these amounts. On a Claude Pro or ChatGPT Plus subscription you pay a flat fee, and the same tokens are consumed from your usage limits instead.

Results

Task Passed (Claude Code / Codex) Wall time, median Cost per run, median Codex cheaper by
Bug fix 2/2 / 2/2 35.0 s / 42.4 s $0.133 / $0.075 1.8x
Feature 2/2 / 2/2 36.2 s / 53.9 s $0.144 / $0.073 2.0x
Refactor 2/2 / 2/2 59.3 s / 58.9 s $0.290 / $0.101 2.9x
Code review (3 planted bugs) 3/3 twice / 3/3 twice 95.7 s / 39.2 s $0.156 / $0.047 3.3x

Bug fix from a real issue

Claude Code (run 1, run 2) Codex (run 1, run 2)
Hidden test and suite pass, pass pass, pass
Wall time 37.2 s, 32.8 s 38.5 s, 46.2 s
Input tokens (of which from cache) 75.0k (64.2k), 92.8k (84.4k) 153.7k (129.0k), 159.6k (130.7k)
Tokens written to cache 10.7k, 8.4k none
Output tokens 2,378, 1,808 892, 842
Tool calls 3, 4 8, 8
Cost $0.146, $0.121 $0.071, $0.079
Lines changed +9 / -1, +7 / -1 +12 / -1, +12 / -1

All four runs changed the same line and added a regression test. Claude Code compared against powers[ordinal + 1] / power both times; Codex did the same once and used the integer division // once, which is exactly the upstream fix. Codex explored more before editing (listing files, looking for an AGENTS.md, running the targeted test, then the suite, then git diff --check), which shows in its tool calls and input tokens. It still cost half as much.

Feature from a spec

Claude Code (run 1, run 2) Codex (run 1, run 2)
Hidden tests and suite pass, pass pass, pass
Wall time 37.0 s, 35.4 s 50.6 s, 57.2 s
Input tokens (of which from cache) 74.1k (63.6k), 75.3k (66.8k) 111.6k (92.3k), 119.1k (93.4k)
Tokens written to cache 10.5k, 8.4k none
Output tokens 2,852, 2,653 1,896, 1,864
Tool calls 3, 3 6, 6
Cost $0.153, $0.134 $0.067, $0.079
Lines changed +72 / -5, +84 / -5 +82 / -5, +76 / -5

The four implementations are nearly interchangeable: same keyword-only signature, same validation, same slicing for "N more". A precise spec leaves little room for a model to differ. Claude Code finished about 18 seconds sooner (median) by batching its reads and test runs into fewer tool calls.

Refactor without behavior change

Claude Code (run 1, run 2) Codex (run 1, run 2)
Checks (tests untouched, API, docs, doctests, suite) pass, pass pass, pass
Wall time 61.9 s, 56.6 s 64.2 s, 53.5 s
Input tokens (of which from cache) 251.9k (230.1k), 188.3k (168.7k) 218.1k (194.0k), 196.1k (154.8k)
Tokens written to cache 21.9k, 19.6k none
Output tokens 4,262, 4,136 2,087, 1,568
Tool calls 8, 6 8, 8
Cost $0.306, $0.273 $0.088, $0.114
Lines changed +155 / -130, +154 / -130 +151 / -133, +165 / -146

This task hides a trap: the helper _format_not_finite is used by both modules, so a naive move creates a circular import. The runs found three different ways out:

All four pass every check. Claude Code paid 2.9x more here, mostly for twice the output and 20k tokens of cache writes per run.

Code review with planted bugs

Claude Code (run 1, run 2) Codex (run 1, run 2)
Planted bugs found 3/3, 3/3 3/3, 3/3
False alarms 0, 0 0, 0
Extra valid findings 2, 3 0, 0
Wall time 31.2 s, 160.2 s 40.6 s, 37.8 s
Input tokens (of which from cache) 94.2k (84.2k), 141.7k (130.8k) 111.6k (97.5k), 110.7k (96.9k)
Tokens written to cache 10.0k, 10.9k none
Output tokens 2,351, 2,749 999, 969
Tool calls 4, 6 5, 5
Cost $0.144, $0.168 $0.048, $0.047

We expected this to be where one agent pulled ahead. It wasn't: every review found the ReDoS, the missing + 1 and the float truncation, each with a reproducing input, and none reported a bug that wasn't one.

The difference is in style. Codex returned exactly three findings, each probed with a short timeout. Claude Code verified every claim by running code, counted the damage (590 wrong results out of 99,999 values between 0.01 kB and 999.99 kB), and added valid remarks we had not planted: GNU-style suffixes like 2.9K that naturalsize() produces but parse_size() rejects, an OverflowError on very long inputs where the docstring promises ValueError, and localized output that cannot be parsed back. In its second run it also started a background ReDoS probe with 30,000 digits, waited, then killed it, which is why that run took 160 seconds. It also rated the binary bug "high" both times, where Codex said "high" once and "medium" once.

Our daily pull request review template runs this kind of review every morning; the numbers above say a Codex step does the planted-bug part at a third of the price.

Totals: tokens, time and cost per passed task

Over 8 runs each Claude Code Codex
Passed 8/8 8/8
Total wall time 452 s 389 s
Input tokens 993k (90% read from cache) 1.18M (84% from cache)
Uncached input tokens 90 192k
Tokens written to cache 100k none
Output tokens 23.2k 11.1k
Tool calls 37 54
Cost $1.45 $0.59
Cost per passed task $0.181 $0.074

Where the money goes is the interesting part:

Pricing and usage limits in practice

Official plans on 2026-10-06:

Claude Code (Anthropic) Codex (OpenAI)
Entry plan Pro: $20/month, $17/month billed annually Plus: $20/month. Go ($8) and Free cover "lightweight" and "quick" tasks
Higher tiers Max: from $100/month, "5x or 20x more usage per 5-hour session than Pro" Pro: $100 to $500/month
How limits work Rolling 5-hour window plus weekly caps, shared between Claude and Claude Code; no published message count Estimated 15 to 160 local messages per 5 hours with GPT-6.1 Sol on Plus; local and cloud share the allowance; weekly limits may apply
API key Yes, per token Yes, "standard API pricing" per token

Sources: claude.com/pricing, Claude Code with Pro or Max, Codex pricing.

Which is cheaper? Per task, Codex, by 1.8x to 3.3x on our four tasks. Projected at API prices, 100 tasks like these cost about $18 with Claude Code and $7.40 with Codex. An unattended job that runs 20 such tasks every weekday (about 440 a month) would cost about $80 against $33. That is the budget question for anything scheduled, like a weekly dependency update or a nightly triage.

Which hits limits first on a $20 plan? We can't say from our data: none of the 16 runs hit a limit, and neither vendor converts its limits into tokens. What we can say is that Claude Code consumed about twice the API-equivalent value per task. If both vendors size their $20 plans to a similar dollar value of compute, which is not published, Claude Code would reach its limit sooner.

Which one should you use?

When Claude Code is the better pick

When Codex is the better pick

Is there a "best AI coding agent" right now?

Not on this evidence. On four everyday tasks, the two leading CLIs produced code of the same quality, sometimes line for line. What separated them was price and verbosity, and both move with each model release. The honest answer is to measure on your own repository, with your own tasks, every time a model changes.

Using Codex and Claude Code together

"Most people use both" is in every comparison, usually followed by a handoff log that you paste from one terminal into the other. The cleaner split follows the results: route each task type to the agent that measured best for it, and let a workflow do the handoff.

In SideHub, a workflow is a list of steps run by an agent on your own machine, and each step can set its own CLI with provider (claude, codex, gemini or copilot). Here is a sketch that implements with Claude Code and reviews with Codex:

format: sidehub.workflow/v1
name: Implement, then cross-review
defaultProvider: claude
parameters:
  - { name: issue, label: Issue number, type: string, required: true }
steps:
  - name: Implement the issue
    prompt: |
      Read GitHub issue #{{inputs.issue}} with `gh issue view`, implement it on a new branch
      and run the test suite. Do not push.
  - name: Review the diff
    provider: codex
    inputsFrom: ["1"]
    prompt: |
      Review the uncommitted diff of this repository for correctness, edge cases, security and
      performance. List each problem with file, line, severity and a reproducing input.

Each step runs as its own run, with its CLI, duration, exit code and the token usage read from the CLI's transcript, and you can follow it from anywhere, phone included. That means you can rerun this comparison on your own codebase and see the cost of each step, instead of trusting a blog post, this one included. The template gallery has ready-made workflows to start from, such as the nightly issue triage, a typical scheduled job where the cheaper agent per run adds up.

Limitations of this benchmark

The verdict

On four everyday tasks, Codex (gpt-6.1-sol) and Claude Code (claude-opus-5-5) were equally good: 16 passes out of 16, every planted bug found, no false alarm. Codex was 1.8x to 3.3x cheaper per task at API prices; Claude Code was a bit faster on small coding tasks and wrote richer reviews. Use Codex where cost per run adds up, Claude Code where you want depth or a quick answer, and rerun these four tasks on your own repository: the commands and prompts are all above.