
London-based AI research lab Inherent has released Faraday, a 27-billion-parameter artificial intelligence agent designed to reproduce results from scientific research papers. The company says Faraday outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 on a benchmark created to test how effectively AI systems can replicate published research.
Inherent, which was founded by former Google DeepMind researchers, published its research paper on Faraday and the Replica benchmark on August 13, 2026. The company published a detailed explanation of the work the following day. TechCrunch independently reported the results on August 22.
Faraday is based on Qwen3.6-27B and was post-trained with long-horizon reinforcement learning. Rather than operating only as a standalone language model, Faraday uses coding agents as tools to carry out experiments. Inherent says the system was trained to develop the kind of judgment required to investigate research problems where important experimental details are not explicitly stated.
The company developed Replica specifically for this purpose. The initial benchmark contains 310 tasks drawn from 100 machine-learning and AI-for-science papers, covering areas including natural language processing, materials science, structural biology and weather forecasting. The benchmark is divided into 242 machine-learning tasks used for training and 68 AI-for-science tasks held out for evaluation.
For each task, the AI system receives a research paper with a selected results figure removed, along with the figure’s caption. The original figure remains hidden from the system and is used as the reference for evaluation. The agent must then determine how to reproduce the result within a limited time and computing budget.
Inherent says this setup is intended to test more than a model’s ability to copy existing code. Research papers generally document successful methods and results rather than every unsuccessful experiment or adjustment made along the way. An agent attempting replication therefore has to make decisions about implementation, experiment design and resource use while working from incomplete information.
The benchmark normally gives agents 60 minutes and access to a one-seventh MIG slice of an H200 GPU. When a paper’s original experiment requires significantly more resources, the agent can create a scaled-down version while attempting to preserve the experiment’s underlying scientific purpose.
Inherent evaluates the resulting experiments across several dimensions, including visual similarity to the original figure, whether the experiment supports the scientific claim, implementation and experimental fidelity, use of available computing resources and scientific integrity.
The company uses an automated rubric-based judging system to score the experiments. Claude Opus 4.7 is used to generate task-specific rubrics, while GPT-5.5 Codex evaluates agent outputs against those rubrics. The evaluation can consider an agent’s code, generated figure and interaction history, and can also rerun code produced during an experiment.
Inherent also conducted a human evaluation involving 20 experts and 117 rankings to compare the automated judging system with human assessments. The company reported that its rubric-based judge showed greater consistency than its baseline judge and somewhat stronger agreement with human judgments, although the agreement between the automated judge and human evaluators was not perfect.
On the 242 machine-learning tasks in the training distribution, Faraday achieved a mean replication score of 0.856. Claude Opus 4.8 scored 0.828, while GPT-5.5 Codex scored 0.796. The Qwen3.6-27B base model scored 0.678.
On the 68 held-out AI-for-science tasks, Faraday scored 0.791. Claude Opus 4.8 scored 0.748 and GPT-5.5 Codex scored 0.729. The unmodified Qwen3.6-27B model scored 0.554.
Inherent reports that Faraday performed better than the two frontier systems on 73% of the machine-learning tasks and 60% of the held-out AI-for-science tasks. The results indicate that the post-training process produced a substantial improvement over the underlying Qwen model, particularly on the held-out tasks.
The comparison is not a simple contest between a 27-billion-parameter model and much larger models. Faraday itself uses GPT-5.5 Codex as a coding tool. Inherent describes Faraday as a scientific layer that directs a more capable coding system, rather than as a replacement for the underlying coding model.
According to the company, Faraday was originally trained using GPT-5.4-mini as its coding agent and was later tested with GPT-5.5 Codex. It was able to work with the more capable coding agent without being retrained specifically for that model.
Inherent says this result supports its approach of separating scientific decision-making from the execution of technical tasks. In this setup, Faraday determines what should be investigated and how an experiment should proceed, while the coding agent handles much of the implementation.
The company also examined individual cases in which Faraday outperformed the other systems. In one experiment involving the Darwin-Gödel Machine, Faraday implemented the evolutionary search procedure described in the research rather than simply reproducing a discovered result. In another experiment involving an LSTM model, Faraday changed the training approach when the initial setup failed to converge instead of manipulating the experiment to obtain a desired result.
Inherent says these examples illustrate what it calls a more scientifically principled approach to research replication. The company argues that reproducing a result requires decisions that are not always written explicitly in a paper, including which approaches to abandon and which experimental changes preserve the original scientific claim.
The research also tests whether the approach can move beyond replication. Inherent created 20 modified research tasks by altering aspects of existing papers, such as changing the dataset or research environment while retaining the original claim, or changing the claim while keeping the experimental setting. The company says Faraday was preferred by its rubric-based judge on 19 of those 20 tasks.
However, Inherent does not present that result as definitive evidence that Faraday can independently make scientific discoveries. The company notes that its judging system was not validated on those imagined tasks in the same way as the main Replica evaluation.
The lab’s broader goal extends beyond reproducing existing experiments. Inherent says it wants to build AI systems capable of scientific innovation and has described Faraday as an early step toward a research system in which AI tools contribute to the development of new knowledge while humans remain involved in the process.
Inherent recently emerged from stealth with a $50 million seed financing round led by Index Ventures and Radical Ventures. The company has positioned itself as a London-based AI research lab focused on scientific discovery and the development of new forms of human-machine collaboration.
The Index Ventures investment announcement identifies Inherent’s founders as Tantum Collins, Edward Hughes, Louis Kirsch and Kaloyan Aleksiev. Collins, Hughes and Kirsch have backgrounds at Google DeepMind, while Aleksiev has infrastructure experience from Reka AI and Microsoft. Edward Hughes serves as co-founder and chief scientist.
Faraday’s reported results do not establish that a 27-billion-parameter model is broadly superior to frontier systems for scientific work. Replica is a benchmark developed by Inherent, the main evaluation uses an automated judge, and the held-out test set contains 68 tasks. Independent testing of the system on other benchmarks and research environments would provide additional evidence about how broadly the results generalize.
For now, Inherent’s release provides a concrete demonstration of a different approach to AI research systems: rather than relying on one model to perform every part of an experiment, Faraday combines a relatively small model trained for scientific decision-making with a more powerful coding agent that carries out the technical work.
The company is now positioning that architecture as a possible path from AI systems that reproduce known results to systems that can conduct increasingly open-ended research. Whether that transition produces reliable scientific discoveries remains an open question, but the Replica results give Inherent a measurable starting point for testing the approach.
References
- Training AI Scientists to Replicate Research — arXiv research paper
- Inherent: Training AI Scientists to Replicate Research
- Index Ventures: Inherent and its AI research approach
- TechCrunch: Inherent’s Faraday research agent
Discover more from Aree Blog
Subscribe now to keep reading and get access to the full archive.

