My split is simple. By default, Claude Code does the work itself. Codex only takes two kinds of job: several unrelated coding tasks that have to run in parallel, or one huge, one-off mechanical job. When I’m not sure, I don’t hand it to Codex — Claude Code does it itself.

This post doesn’t compare features, prices or versions. It covers how I split the work, why I split it that way, and the pitfalls I hit. It picks up from my June post, A Few Months with Claude — I’ve Put Codex and ChatGPT Aside, where my split was: Claude hands out the tasks, Codex does the mechanical work. For how to use Codex itself, the site also has a complete beginner’s guide for non-programmers.

The split

Here is the split I now have written into my rules, in the form “what kind of work → who does it → why”.

  • Small fixes, feature work, debugging, iterations that depend on a lot of context, and anything that takes real judgement or can’t afford mistakes → Claude Code does it itself. Why: faster and more accurate, no information lost, no handoff or review cost.
  • Several unrelated coding tasks that need to run at the same time → Codex. Why: what I need is throughput, with several things getting done at once.
  • One-off, very large mechanical jobs, such as bulk-editing hundreds of files → Codex. Why: I want to keep the current conversation lean.
  • Not sure → Claude Code does it itself. Why: handing work to Codex means writing a complete, self-contained brief and then reviewing the result strictly; doing it itself has no such handoff and review cost.
  • Going live, deleting, paying, sending anything to the outside world → I approve it, whoever does the work. Why: once done, these can’t be undone.

“Claude Code does it itself” also covers the Claude sub-agents it sends out: helpers the main AI dispatches to work on their own and bring back only the conclusion.

What each one is for

Claude Code is the lead: it breaks tasks down, hands them out, checks the results and reports back. Codex is the executor: it takes a task, breaks it down itself, carries it through, checks its own work and then reports. I wrote that into its rules: it is there to get the task finished, not just to plan.

In the June post I described how the two behave. Codex is like a programmer: you have to spell out the requirements very clearly, it is very strong from 0 to 90%, and the last 10% takes a lot of polishing. Claude is more like a boss, closer to how a person thinks, with a somewhat lower cost of communication. That matches my rules today: work sent to Codex has to be written up as a complete, self-contained brief, and the result gets a strict review.

Work that goes to Codex follows a fixed process. First, settle “what to change and how”, and do a dry run (look at what would change, without changing anything). Once I approve, Codex does the work on a dedicated branch (a separate line of changes, kept apart from the main version). When it’s done, Claude Code reviews the diff (the list of changes) and runs the tests, and only then is it merged. Anything that touches the live site, a deployment or deleting data waits for my go-ahead. When I set this process in late June, my words were: “The point isn’t to go and change everything in one go. Settle it first, then make the change.”

Why Claude Code is the default

First, there is the cost of handoff and review: every time I give a job to the other tool, I add one more round of “explain the background in full” and one more round of “review it afterwards”. That is especially true for iterations that depend on a lot of context: the background is already in the current conversation, so when Claude Code does it itself, no information is lost.

The other cost is the size of the conversation. My rule is that the process goes to sub-agents and the main conversation keeps only the conclusion. In my words: “I don’t need to know the detailed analysis in the conversation. I don’t have time to read it. Give that to sub-agents; all I need is the summarized result.” In a measurement at the end of September, the 1,110 past sub-agents on my computer used a median of about 42,000 tokens (the unit AI uses to count text) internally and brought back only about 900 to the main conversation. The fatter the main conversation gets, the more often it is compressed, and the more the AI forgets. I wrote this rule into the rule files of both tools, and it is one reason huge mechanical jobs go to Codex.

Four pitfalls

1. Freezing. This is the problem I complain about most. When I send work to Codex or run a long command, and it runs in the foreground, or the AI writes “I’ll wait for Codex to finish before continuing” and ends its turn, it won’t wake up on its own, and the conversation just stays stuck there. My rule now: this kind of work always runs in the background; when it finishes, the system wakes the AI automatically, and the AI checks the result and carries on, without me prodding it. Tests, crawlers and large downloads follow the same rule.

In Claude Code there is a related trap: a sub-agent running in the foreground is interrupted by any new message I send and gets logged as “stopped by user”. If I wait a while and type “is it done yet?”, I have just killed it. Background tasks aren’t affected.

Another way to freeze is to grind away at one small item, so both tools’ rules contain the same clause: save after every finished item; if an item fails twice, skip it and note “needs a human”; always deliver partial results, never zero output. For example, if 3 items out of 30 fail, still hand over the summary table for the other 27 plus the list of the 3.

2. Overwriting each other. On one project, Claude and Codex worked back and forth, and the old version overwrote the new one. It turned out that project wasn’t a git repository at the time (a project folder that records every change): whatever was written last simply overwrote whatever was written first, with no way back.

Now there are three layers. Projects become git repositories, so an accidental overwrite can be recovered. One task has one writer at a time: before changing a project, check whether anyone else is changing it; only if nobody is can you take the lock; if you can’t get it, go and do something else instead of pushing through; and a lock expires by itself after two hours. And before starting, read the project’s progress file. For Codex there is one more rule: before sending it work, check whether the target project has uncommitted changes; it works only on a clean branch and never deletes or overwrites anything.

3. Sub-agents in the same conversation collide too. One Claude conversation sent out two sub-agents: one fixing the website’s hamburger menu, the other working on performance at the same time, on the same set of files. The second one started from the live version, and after it deployed, the live site had “the performance work, without the menu fix”. The root cause was the wrong starting point, not who typed faster: starting from the live site naturally swallows every local change that hasn’t gone live yet. In the end a three-way merge put both back together, but that was luck, because the two changes happened not to touch the same line. My rule now: the local mirror (the copy kept locally) is the only baseline, and sub-agents that would change the same files either run one after the other or split the files between them.

4. Rules written into the wrong file. Claude Code reads CLAUDE.md; Codex reads the AGENTS.md in its own global folder. Each reads its own. At the end of September an AI went to edit Codex’s rules and edited a mirror file that Codex doesn’t normally read (the copy of CLAUDE.md with “Claude” replaced by “Codex”). It took a test the next day to find out. One more thing: Codex’s tool instructions say it must not spawn sub-agents unless the user or a rules file explicitly asks for it, so I wrote a paragraph of explicit authorization into its rules file.

If you don’t write code

A few things you can use as they are:

  • Let one tool carry the work through. Bring in a second one only when one of two signals appears: several unrelated things have to be done at the same time, or there is a large one-off mechanical job. When in doubt, stay with the first.
  • Keep progress in a file inside the project, not in the chat. I use COWORK.md, with four fixed lines at the top: current status, next step, blocker, update. When you switch tools or conversations, the newcomer reads it first and picks up from there, and you don’t have to be the messenger. The conversation is just the hands that do the work; the memory lives in files.
  • Turn the project into a git repository, and let only one AI change a given set of files at a time. The first guards against “it broke and I can’t get it back”, the second against “two changes at once”.
  • Ask for evidence when work is handed in, and don’t accept a bare “done”. Require the AI to verify end to end before it says “fixed”; when a sub-agent reports back, spot-check the key evidence. For outgoing email and messages, the AI only drafts, and it sends only when I explicitly say “send”; “OK” and “fine” don’t count.