
London-based AI startup Inherent says its new agent, Faraday, has outperformed larger systems from Anthropic and OpenAI at independently reproducing the findings of published scientific papers. The claim marks the company’s first public demonstration since emerging from stealth with a $50 million seed round and offers an early test of its ambition to build AI systems that contribute to scientific discovery.
The comparison is notable because Inherent says Faraday runs on Qwen 3.6, a model with 27 billion parameters. The startup measured it against Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5, which Inherent describes as much larger frontier models. However, the performance results are company-reported through TechCrunch, and the available evidence does not include a public benchmark methodology or independent validation.
Faraday was designed to reproduce results from scientific papers without being given the expected answer in advance. Inherent cofounder and chief scientist Edward Hughes told TechCrunch that paper replication is a standard early exercise for human scientists, including PhD students.
The startup says it evaluated more than whether an agent could reach a correct result. It also wanted Faraday to show what Hughes called “research taste”: the ability to identify worthwhile experiments and make sensible choices about how to run them.
That distinction matters for research-oriented AI. A system that follows a fixed procedure may verify an existing result, but a useful scientific collaborator must decide which questions deserve time, which variables to test, and when evidence is strong enough to justify a conclusion. Inherent’s longer-term objective is an AI scientist agent that can operate across scientific fields rather than a tool limited to one benchmark or discipline.
Inherent attributes Faraday’s reported result partly to its use of reinforcement learning. Instead of relying primarily on training data about how scientists conduct research, the company says it rewards agents for achieving useful outcomes. The approach is intended to help the system acquire broader decision-making instincts, including the judgment the startup associates with research taste.
The choice also reflects a deliberate product strategy. Inherent is not developing a dedicated coding tool for Faraday. According to the company, the agent uses OpenAI’s GPT-5.5 Codex for coding tasks, similar to how human researchers rely on existing software rather than building every tool themselves.
That architecture could be important for teams building scientific and technical agents. The most capable system may not be a single monolithic model. It could instead be a smaller reasoning model connected to specialized tools, coding systems, experiment environments, and evaluation loops. Faraday’s reported comparison does not prove that approach is broadly superior, but it provides a concrete example of the tradeoff Inherent is exploring.
The central performance claim comes from Inherent and was reported by TechCrunch. The available source material does not specify the number of papers tested, the scientific fields represented, the scoring criteria, the degree of human supervision, or whether Anthropic and OpenAI systems were run under identical tool and compute conditions.
Those details are essential when comparing research agents. Paper replication can vary considerably in difficulty depending on whether source code, datasets, laboratory protocols, and compute resources are available. Results may also depend on how an agent is prompted, how failed experiments are handled, and whether evaluators judge only the final answer or the quality of the experimental process.
The use of a 27-billion-parameter model is nevertheless an important vendor-reported signal. If independently reproduced, the result would suggest that model size alone is not a reliable proxy for performance on long-horizon scientific workflows. It could also indicate that reinforcement learning and tool orchestration are contributing as much as base-model scale for this particular task.
Inherent’s framing is deliberately broader than a leaderboard result. Hughes told TechCrunch that beating frontier systems was less important to the company than the way Faraday was built. He described the intended agent as a collaborator that can return with unexpected experiments and results, rather than one that simply produces answers designed to satisfy its user.
For AI builders, Faraday highlights a growing design question: should agents be optimized for conversational quality, benchmark accuracy, or the ability to make and learn from decisions over many steps? Scientific research exposes weaknesses that are easier to hide in shorter workflows, including poor experiment selection, brittle planning, unverified assumptions, and a tendency to report confident but unsupported conclusions.
A smaller model could make research agents cheaper to deploy, particularly when the workflow requires many experiments or repeated calls to a model. But lower model size does not automatically mean lower total cost. Tool use, code execution, failed runs, data access, human review, and infrastructure can dominate the economics of an agent that operates for hours or days.
Enterprise buyers should therefore treat Faraday as an early signal rather than a ready-made replacement for researchers. The practical questions will be whether the system produces reproducible work, exposes its assumptions, preserves an audit trail, and knows when it lacks enough evidence. Those requirements apply to drug discovery, materials research, engineering, and internal R&D teams, where an attractive result is less valuable than a result that can be checked.
Inherent’s hiring plans also place the product in a competitive market for advanced AI talent. The company has about a dozen employees working in person in London and plans to expand to roughly 20 to 25 people by the end of the year, according to the TechCrunch report. Hughes also criticized the U.K.’s garden-leave restrictions, saying they had affected his own ability to move from a previous role.
The most important follow-up will be an independently reproducible evaluation of Faraday against Claude Opus 4.8 and GPT-5.5. Inherent would strengthen its claim by publishing the paper set, prompts, tool configurations, compute budgets, failure rates, and scoring process.
Researchers and product teams should also watch for evidence beyond replication. Inherent says its north star is an AI scientist agent, but the current public demonstration concerns reproducing existing research. The next meaningful test would be whether Faraday can generate novel hypotheses, design experiments that experts consider worthwhile, and produce findings that survive external review.
Other signals include the system’s error-handling behavior, its reliance on GPT-5.5 Codex, and whether Inherent releases access for outside users or evaluators. Hiring growth in London may indicate confidence and momentum, but headcount is not evidence that the agent’s scientific capabilities generalize.
Inherent’s announcement is interesting because it shifts attention from raw model scale to the construction of long-running research workflows. A smaller base model paired with reinforcement learning, coding tools, and experiment feedback could be a more practical path to specialized AI agents than waiting for every capability to arrive inside one frontier model.
The claim should still be read as an opening result, not a settled lead over Anthropic or OpenAI. For builders, the real question is whether Faraday’s reported advantage survives transparent evaluation and transfers from paper replication to reliable scientific work. That evidence will determine whether Inherent has demonstrated a durable approach or simply a strong result on a narrow task.
Inherent says its Faraday AI agent beat larger Anthropic and OpenAI systems at paper replication, highlighting smaller models and research-focused agents.