IA al Día
Back to archive
Tools news Sep 10, 2026 5 min read

OpenAI rents out the Codex harness: the Agents API and who controls it

OpenAI opens the Agents API in public beta, hosting the Codex harness so developers keep the model and infra choice but hand over the orchestration layer.

primary source · openai.com — Introducing the Agents API

Building an agent has always meant writing the harness yourself: the loop that decides what stays in context, which tool gets called next, when to compress history and how to split work across subagents. OpenAI’s Agents API, opened in public beta to all developers on 10 September 2026, replaces that code with a call to infrastructure OpenAI owns. It is the same harness that runs Codex, now hosted, versioned and improved by OpenAI itself rather than tuned by hand inside each team.

What the developer still picks

The API is built around four primitives: the agent (model, instructions, tools and MCP servers), the environment, the session (“a durable instance of an agent”) and the events and items that make up its run, according to the Agents API overview. Inside that frame, three choices stay with the developer. The model is one, although every published example uses GPT-6 Astra and OpenAI has not published a list of supported models. The tools are another: MCP servers, custom functions and built-ins such as web search. The third is where the agent’s code actually executes: an OpenAI-managed sandbox, the developer’s own infrastructure, or one of nine sandbox partners OpenAI named at launch: Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop and Vercel.

Compaction, tool search and subagents come built in

What the developer no longer writes is the orchestration logic itself. The harness compacts context automatically as a session nears the model’s limit, loads tool definitions on demand through tool search instead of stuffing them all into the prompt, supports programmatic tool calling, and runs multi-agent setups where each subagent keeps its own context. “We maintain and continuously improve the harness alongside our models,” OpenAI says, and it versions that access with each model launch, so teams get harness improvements without touching orchestration code. One launch customer, Ciridae, reported an evaluation score rising from 0.71 to 0.85 and a fourfold latency drop after adopting subagents — a vendor-selected quote in OpenAI’s own announcement, without a published method behind it.

No separate fee, but containers are metered

OpenAI is not charging for the harness on top of the model. “There are no additional fees for using the Agents API – you simply pay for the tokens and tools your agents use,” the announcement states. What is metered is the OpenAI-hosted sandbox itself: Shell and Code Interpreter containers cost $0.03, $0.12, $0.48 and $1.92 per 20-minute session, for 1, 4, 16 and 64 GB of memory respectively, per the published pricing. The launch examples run on GPT-6 Astra, billed at $10 per million input tokens and $50 per million output tokens at short context.

An open-source harness as a counterweight

The harness is not a black box OpenAI is asking developers to trust: it is the same codebase published as open source for Codex, so a team can read exactly what it is renting and, in principle, run an equivalent version itself instead of depending on OpenAI’s hosted copy. Pluggable execution works the same way. Letting code run on a company’s own infrastructure or at a third-party sandbox provider is a concession to teams that will not move their data into OpenAI’s environment, and it keeps the API from being an all-or-nothing commitment.

What the beta does not offer yet

Two limits define what regulated buyers can do with this today. Data residency in the beta covers only the United States, and the API does not support Zero Data Retention; running the sandbox on a company’s own infrastructure does not change that, according to the Agents API overview. Several other questions have no public answer at all: OpenAI says the infrastructure keeps agents “running reliably for days” without an uptime figure, session limit or SLA attached; subagent concurrency defaults to six with no documented maximum; every code example uses gpt-6-astra and no model list is published; and the announcement says nothing about human-approval gates, audit logs or tracing for what an agent does.

The launch retires nothing. The Assistants API had already shut down on 26 August, and the docs place the new API beside the other two routes rather than above them: the Agents API for “long-running tasks where OpenAI manages the agent and saves its progress”, the Agents SDK for agents built inside the developer’s own application, and the Responses API for calling models directly or building an agent from scratch.

The trade is plain from OpenAI’s own description: developers stop maintaining orchestration code and get every harness improvement for free, and OpenAI holds the layer that most shapes how an agent behaves, with the agent’s state sitting in its sessions inside US data centers for now. Whether that trade is acceptable depends on a detail the beta has already made concrete — data residency and retention — more than on any feature in the launch demo.