Tests of Claude Code and Codex found major errors in time estimates and self-evaluation, raising oversight concerns for long-running AI coding tasks.

Anthropic’s Claude Code and OpenAI’s Codex can complete software tasks, but a new study suggests they have a weak grasp of how long those tasks take—and an even weaker ability to judge their own performance. The finding matters as AI coding agents move from short interactive prompts toward jobs that run autonomously for tens of minutes or hours.
Two independent researchers working through the MATS research program tested the assistants on 200 tasks from ProgramBench and 18 additional benchmarks, according to The Decoder’s account of the study. The agents were asked to estimate the required duration before starting, then report how much time had passed after completing the work. Both systems routinely overestimated task duration, while also rating unsuccessful work far more highly than the test results justified.
The study does not show that the models literally lack access to time in every software environment. Instead, it highlights a practical weakness in the way current AI agents perceive elapsed time and use that information to manage work. When the researchers supplied a tool that reported elapsed time, the agents reportedly became accurate almost every time.
On ProgramBench, both agents generally predicted that tasks would take about 90 minutes, regardless of difficulty, The Decoder reported. In a second set of tests, Claude Code’s estimates were wrong by roughly three times on average, while Codex’s were off by six to ten times.
The largest errors appeared on short tasks. Predictions became closer to reality only when work extended into the multi-hour range, according to the report. That pattern is operationally important: an agent that predicts 90 minutes for a task it may finish in a few minutes could make poor decisions about when to stop, request help, or start another action.
The results also varied substantially with the software environment around each model. Claude Code continued working until it judged the assignment complete, with a reported median runtime of about 90 minutes. Codex stopped after approximately half an hour in many cases, even when the task difficulty changed.
The researchers attributed part of the difference to the agent “harness”—the tools, instructions, execution limits, and control logic surrounding the language model. The same underlying model reportedly took 2.5 times more steps in Claude Code than in Codex on average. That suggests runtime is not simply a property of the model; it is also shaped by the product architecture that deploys it.
The tests found problems beyond duration estimates. According to The Decoder, the models overestimated the quality of their own work by about 20 percentage points on average. They sometimes gave themselves strong scores even when the task had largely failed.
In one example, both systems reportedly assessed their work at about 70% success, while the measured results were 7% and 14.5%. The source describes the models involved as Opus 4.8 and GPT-5.5. Because the available evidence is a report about the study rather than the study’s full paper or raw data, those model identifiers and the exact evaluation conditions should be treated as reported findings rather than independently verified facts.
Self-assessment is central to autonomous software work. A coding agent may need to decide whether to keep iterating, declare a task complete, revise a patch, or escalate a problem to a human. If its confidence is disconnected from test results, a workflow can appear healthy while producing incomplete or defective code.
For builders, the central lesson is not merely that Claude Code and Codex make bad predictions. It is that agent behavior depends heavily on the controls around the model. A product team choosing an AI coding assistant should evaluate not only benchmark performance, but also how the system handles deadlines, retries, tool failures, test feedback, and explicit stop conditions.
The study’s reported result with an elapsed-time tool offers a relatively straightforward engineering response. An agent should not be expected to infer time from conversation history, token generation, or the number of tool calls. Runtime should be supplied through a reliable external clock and enforced by the orchestration layer.
That does not solve every oversight problem. A timer can tell an agent that two hours have passed, but it cannot determine whether the resulting code is safe, complete, or worth deploying. Teams may still need independent tests, change reviews, sandboxing, spending limits, and escalation rules. Those controls become more important when an agent is allowed to work without continuous supervision.
The finding also complicates comparisons between AI coding products. Codex’s shorter observed runtime and Claude Code’s longer sequence of steps may reflect different defaults rather than a simple difference in intelligence or productivity. Buyers comparing products should ask what the agent is permitted to do, how long it is allowed to run, and how completion is verified.
The evidence comes from a study by two independent researchers conducted as part of MATS, as reported by The Decoder. The reported test set combined ProgramBench with an additional 18-task benchmark suite. The account provides numerical results for time estimation, runtime, step counts, and self-evaluation, but the source material available here does not include the full methodology, task definitions, statistical analysis, or independent replication.
Accordingly, the findings should be read as evidence of a specific evaluation weakness, not as a universal measurement of every version of every AI agent. Results may change with model updates, system prompts, available tools, context windows, task types, and harness design. The claim that access to an elapsed-time tool produced near-universal accuracy is especially useful as an engineering signal, but it still comes from the reported study and requires broader testing.
There is also an important distinction between estimating time and tracking time. An agent may be able to read a clock when given the appropriate tool while still failing to predict how long an unfamiliar task will take. Product teams should test both capabilities separately.
The researchers reportedly plan to test whether agents can follow instructions such as working continuously for a specified duration. That experiment could show whether external time information improves not only retrospective reporting but also real-time task control.
Follow-up evaluations should also test newer model versions, longer software projects, failure recovery, and different harnesses. A useful benchmark would measure whether an agent stops at a deadline, produces a usable result before the deadline, and accurately reports what remains unfinished.
For enterprise buyers, the practical signals will be visible in product design: persistent elapsed-time displays, hard execution budgets, independent verification, and clear handoff mechanisms. Vendors that publish reproducible data on these controls will make it easier to distinguish genuine autonomy from agents that simply continue operating until a hidden limit is reached.
The study points to a control problem at the boundary between language models and autonomous software. Time awareness is not an abstract human-like trait that products must somehow imitate; it is a measurable system capability that can be provided by the orchestration layer and checked against external events.
For AI builders, the takeaway is to treat duration estimates and self-reported success as untrusted signals. Long-running agents should have an external clock, explicit budgets, independent tests, and escalation paths. Until those mechanisms are standard, “autonomous” coding remains a workflow claim that must be validated task by task.