For most teams in October 2026, Codex CLI is the best AI coding agent harness by default. Claude Code is an equal choice if you use Anthropic models and want the deepest extensibility. Choose OpenCode when model freedom matters most, especially if you need local models.
What is an agent harness?
An agent is a model plus a harness. The model reasons and produces text. The harness is the software around it: the loop, tools, context, permissions, sandbox, session state and extensions. LangChain puts it in one line: Agent = Model + Harness.
The VS Code documentation gives a similar definition. It describes the harness as the layer that runs the session, connects the model to context and tools, applies approval rules and keeps state. In plain terms, the model decides what it wants to do. The harness decides what it can see, what it can run and when a person must approve it.
A July 2026 source-code study of eleven coding harnesses found that they run on hand-written agent loops. None imported a general agent framework or retrieved code with vector embeddings. They relied on direct methods such as ripgrep, tree-sitter, glob patterns and Markdown context files. Skills appeared in nine projects and MCP in eight, but the study did not rank them.
Which is the best AI coding agent harness right now?
Our recommendation is Codex CLI for most engineering teams. Its secure default is the deciding factor. It limits writes to the workspace, blocks network access and uses an operating system enforced sandbox without asking teams to switch those controls on first. It is open source under Apache-2.0 and supports profiles, hooks, subagents, MCP and custom model providers.
Claude Code is not a lesser option. We think it is the better fit for teams centred on Anthropic models, or teams that want the broadest set of surfaces and extension points. The same engine, CLAUDE.md, settings and MCP servers work across its terminal, editors, desktop and web tools, with further routes into CI and Slack.
OpenCode is our choice when provider freedom comes first. It supports more than 75 providers and has a clear path to local models. Its permissions start from a more permissive position, however, and its documentation describes tool rules rather than an operating system sandbox. That changes the setup we would accept for work repositories.
How we judged them
- Model choice: whether the tool supports one model family, several hosted providers, or local inference.
- Tool use and extensibility: support for skills, hooks, MCP, subagents, context files and other ways to shape the agent loop.
- Sandboxing and permissions: what is isolated, what is allowed by default and where approval is required.
- Context handling: how the agent finds repository facts, keeps session state and follows project instructions.
- Cost: the cost of model use, while keeping benchmark run costs separate from normal billing.
- Team fit: editor or terminal use, policy control, shared configuration and the effort needed to set safe defaults.
Claude Code vs Codex: which is better?
Benchmark
The official Terminal-Bench 4.0 leaderboard tests a harness and model as a pair. Its leading Codex plus GPT-6 Astra result is 58.2%, with a 2.8 percentage point confidence interval half width. Claude Code plus Fable 5.1 reaches 57.9%, with a 3.8 point half width. Those intervals overlap, so the results are statistically tied.
At xhigh effort, both pairs score 57.9%. The listed benchmark run cost is $2,351 for Codex and $4,872 for Claude Code. That makes the Codex run about half the cost, but it is not a monthly bill. Terminal-Bench mostly measures the model inside the harness, so it cannot settle every product choice. OpenCode, Cursor, Gemini CLI and Aider have no rows on the 4.0 board, so we cannot compare them on it.
Sandbox defaults
Codex has the safer starting point. Its security documentation says the automatic preset uses workspace-write with on-request approval. Network access is off, workspace writes are constrained, and the sandbox uses Seatbelt on macOS, bwrap with seccomp on Linux, or a native Windows sandbox.
Claude Code's Bash sandbox is off by default. Its sandbox documentation explains that it covers shell commands only. File tools, MCP servers and hooks sit outside it. Sandboxed commands may also read much of the machine unless paths such as SSH and AWS credentials are denied.
Extensibility
Both products read repository instructions and support skills, hooks, subagents and MCP. Codex adds profiles and optional automated review of actions that already require approval. Claude Code adds parallel agents, an Agent SDK and scheduled routines. Our view is that Claude Code has the deeper extension surface, while Codex offers the stronger default boundary.
Model choice
Claude Code is designed around Claude, though its terminal and editor clients can use third-party providers. Codex supports compatible custom provider URLs, Amazon Bedrock and local use through --oss. Neither matches OpenCode for provider breadth.
Cost
Cost depends on the selected model, effort setting, workload and provider. The only sound direct comparison here is the Terminal-Bench run above. Teams should measure their own tasks, including retries and review time, rather than treating one benchmark invoice as a general price.
How do the other harnesses compare?
OpenCode
OpenCode is open source and works in a terminal, desktop app and editor extension. It has Plan and Build modes, plus undo and redo. Its provider range is its main advantage. The trade-off is that most permissions allow actions by default, while loop detection and access outside the working directory ask for approval.
Cursor
Cursor is editor-first and supports models from several vendors alongside its own models. Its run modes cover allowlists, Auto-review and Run Everything. Auto-review can sandbox some shell commands and passes the rest to a classifier that runs on Cursor's backend. Cursor states plainly that it is not a security boundary.
Gemini CLI
Gemini CLI offers several isolation choices: Seatbelt, Docker, Podman, native Windows, LXC and gVisor. Its documentation calls gVisor the strongest available option. It also supports tool-level isolation and requests to expand the sandbox. The documentation also warns that isolation reduces risk but does not remove it.
Aider
Aider is the clear Git-first option. Its Git integration commits each edit with a descriptive message and first commits dirty files so existing work remains separate. It provides /undo and /diff. By default, commits skip pre-commit hooks, so use --git-commit-verify when those checks form part of your policy.
Hermes Agent and OpenClaw
Hermes Agent and OpenClaw are general personal agents, not coding-first tools. They suit a different job and should not be forced into this comparison. Our earlier Hermes Agent and OpenClaw guide covers that choice without repeating its setup here.
What is the best agent harness for local LLMs?
OpenCode is the best agent harness for local LLMs. Its provider documentation covers more than 75 providers, with local routes through Ollama, LM Studio and llama.cpp. Codex is the next choice: --oss connects to Ollama or LM Studio. Aider also remains a sensible, simple option for a Git-led workflow.
Local availability does not mean equal agent quality. Tool calling varies sharply between models. OpenCode's own documentation recommends models with strong tool use and advises raising Ollama's num_ctx value. Test file edits, command selection, retries and long-context behaviour before setting a team standard.
Which settings should you change on day one?
- Claude Code: turn on /sandbox, or set sandbox.enabled to true. Set allowUnsandboxedCommands to false and failIfUnavailable to true. Add credential read denials for ~/.ssh and ~/.aws, because the default sandbox can read them.
- Codex: keep workspace-write with on-request approval. Leave network access off unless the task needs it. Avoid --yolo, which removes both approvals and the sandbox. In advanced configuration, set shell_environment_policy.ignore_default_excludes to false. Codex then strips variables whose names contain KEY, SECRET or TOKEN before it runs a command. Out of the box, it does not.
- OpenCode: set the bash permission to ask and run the agent inside a container. Tool permission rules help control intent, but they are not a substitute for operating system isolation. Do not use --auto on an untrusted repository.
- Cursor: use an allowlist for repeatable commands and treat Auto-review as a convenience, not a security boundary. Do not use Run Everything for ordinary repository work.
These controls matter more than a long feature list. Agent-generated commands can touch credentials, dependency scripts and external services. The right policy starts with the smallest useful access, then adds permissions for a defined task.
Where Alongside fits
Choosing and configuring a coding harness, including permissions and sandbox policy for a team, is the kind of decision covered by our technical advisory work. If the concern is code already written by agents, a code audit provides a focused review of that code and its risks.