{"id":6601,"date":"2026-08-24T09:03:58","date_gmt":"2026-08-24T09:03:58","guid":{"rendered":"https:\/\/areeblog.com\/?p=6601"},"modified":"2026-08-24T09:03:58","modified_gmt":"2026-08-24T09:03:58","slug":"inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark","status":"publish","type":"post","link":"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/","title":{"rendered":"Inherent\u2019s 27B Faraday Agent Beats GPT-5.5 and Claude Opus 4.8 on Scientific Replication Benchmark"},"content":{"rendered":"<p><img loading=\"lazy\" loading=\"lazy\" decoding=\"async\" data-attachment-id=\"6602\" data-permalink=\"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/images-37\/\" data-orig-file=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-37.jpeg\" data-orig-size=\"739,415\" data-comments-opened=\"1\" data-image-meta=\"{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}\" data-image-title=\"images (37)\" data-image-description=\"\" data-image-caption=\"\" data-large-file=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-37.jpeg\" class=\"aligncenter size-full wp-image-6602\" src=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-37.jpeg\" alt=\"Inherent\u2019s 27B Faraday Agent Beats GPT-5.5 and Claude Opus 4.8 on Scientific Replication Benchmark\" width=\"739\" height=\"415\" srcset=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-37.jpeg 739w, https:\/\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-37-300x168.jpeg 300w\" sizes=\"auto, (max-width: 739px) 100vw, 739px\" \/><\/p>\n<p>London-based AI research lab Inherent has released Faraday, a 27-billion-parameter artificial intelligence agent designed to reproduce results from scientific research papers. The company says Faraday outperformed Anthropic\u2019s Claude Opus 4.8 and <a href=\"https:\/\/areeblog.com\/openai-and-anthropic-ai-agents-linked-to-new-security-incidents\/\">OpenAI\u2019s<\/a> GPT-5.5 on a benchmark created to test how effectively AI systems can replicate published research.<\/p>\n<p>Inherent, which was founded by former Google DeepMind researchers, published its <a href=\"https:\/\/arxiv.org\/abs\/2608.13331\">research paper on Faraday and the Replica benchmark<\/a> on August 13, 2026. The company published a detailed explanation of the work the following day. TechCrunch independently reported the results on August 22.<\/p>\n<p>Faraday is based on Qwen3.6-27B and was post-trained with long-horizon reinforcement learning. Rather than operating only as a standalone language model, Faraday uses coding agents as tools to carry out experiments. Inherent says the system was trained to develop the kind of judgment required to investigate research problems where important experimental details are not explicitly stated.<\/p>\n<p>The company developed Replica specifically for this purpose. The initial benchmark contains 310 tasks drawn from 100 machine-learning and AI-for-science papers, covering areas including natural language processing, materials science, structural biology and weather forecasting. The benchmark is divided into 242 machine-learning tasks used for training and 68 AI-for-science tasks held out for evaluation.<\/p>\n<p>For each task, the AI system receives a research paper with a selected results figure removed, along with the figure\u2019s caption. The original figure remains hidden from the system and is used as the reference for evaluation. The agent must then determine how to reproduce the result within a limited time and computing budget.<\/p>\n<p>Inherent says this setup is intended to test more than a model\u2019s ability to copy existing code. Research papers generally document successful methods and results rather than every unsuccessful experiment or adjustment made along the way. An agent attempting replication therefore has to make decisions about implementation, experiment design and resource use while working from incomplete information.<\/p>\n<p>The benchmark normally gives agents 60 minutes and access to a one-seventh MIG slice of an H200 GPU. When a paper\u2019s original experiment requires significantly more resources, the agent can create a scaled-down version while attempting to preserve the experiment\u2019s underlying scientific purpose.<\/p>\n<p>Inherent evaluates the resulting experiments across several dimensions, including visual similarity to the original figure, whether the experiment supports the scientific claim, implementation and experimental fidelity, use of available computing resources and scientific integrity.<\/p>\n<p>The company uses an automated rubric-based judging system to score the experiments. Claude Opus 4.7 is used to generate task-specific rubrics, while GPT-5.5 Codex evaluates agent outputs against those rubrics. The evaluation can consider an agent\u2019s code, generated figure and interaction history, and can also rerun code produced during an experiment.<\/p>\n<p>Inherent also conducted a human evaluation involving 20 experts and 117 rankings to compare the automated judging system with human assessments. The company reported that its rubric-based judge showed greater consistency than its baseline judge and somewhat stronger agreement with human judgments, although the agreement between the automated judge and human evaluators was not perfect.<\/p>\n<p>On the 242 machine-learning tasks in the training distribution, Faraday achieved a mean replication score of 0.856. Claude Opus 4.8 scored 0.828, while GPT-5.5 Codex scored 0.796. The Qwen3.6-27B base model scored 0.678.<\/p>\n<p>On the 68 held-out AI-for-science tasks, Faraday scored 0.791. Claude Opus 4.8 scored 0.748 and GPT-5.5 Codex scored 0.729. The unmodified Qwen3.6-27B model scored 0.554.<\/p>\n<p>Inherent reports that Faraday performed better than the two frontier systems on 73% of the machine-learning tasks and 60% of the held-out AI-for-science tasks. The results indicate that the post-training process produced a substantial improvement over the underlying Qwen model, particularly on the held-out tasks.<\/p>\n<p>The comparison is not a simple contest between a 27-billion-parameter model and much larger models. Faraday itself uses GPT-5.5 Codex as a coding tool. Inherent describes Faraday as a scientific layer that directs a more capable coding system, rather than as a replacement for the underlying coding model.<\/p>\n<p>According to the company, Faraday was originally trained using GPT-5.4-mini as its coding agent and was later tested with GPT-5.5 Codex. It was able to work with the more capable coding agent without being retrained specifically for that model.<\/p>\n<p>Inherent says this result supports its approach of separating scientific decision-making from the execution of technical tasks. In this setup, Faraday determines what should be investigated and how an experiment should proceed, while the coding agent handles much of the implementation.<\/p>\n<p>The company also examined individual cases in which Faraday outperformed the other systems. In one experiment involving the Darwin-G\u00f6del Machine, Faraday implemented the evolutionary search procedure described in the research rather than simply reproducing a discovered result. In another experiment involving an LSTM model, Faraday changed the training approach when the initial setup failed to converge instead of manipulating the experiment to obtain a desired result.<\/p>\n<p>Inherent says these examples illustrate what it calls a more scientifically principled approach to research replication. The company argues that reproducing a result requires decisions that are not always written explicitly in a paper, including which approaches to abandon and which experimental changes preserve the original scientific claim.<\/p>\n<p>The research also tests whether the approach can move beyond replication. Inherent created 20 modified research tasks by altering aspects of existing papers, such as changing the dataset or research environment while retaining the original claim, or changing the claim while keeping the experimental setting. The company says Faraday was preferred by its rubric-based judge on 19 of those 20 tasks.<\/p>\n<p>However, Inherent does not present that result as definitive evidence that Faraday can independently make scientific discoveries. The company notes that its judging system was not validated on those imagined tasks in the same way as the main Replica evaluation.<\/p>\n<p>The lab&#8217;s broader goal extends beyond reproducing existing experiments. Inherent says it wants to build AI systems capable of scientific innovation and has described Faraday as an early step toward a research system in which AI tools contribute to the development of new knowledge while humans remain involved in the process.<\/p>\n<p>Inherent recently emerged from stealth with a $50 million seed financing round led by Index Ventures and Radical Ventures. The company has positioned itself as a London-based AI research lab focused on scientific discovery and the development of new forms of human-machine collaboration.<\/p>\n<p>The <a href=\"https:\/\/www.indexventures.com\/perspectives\/inherent-designing-for-discovery\/\">Index Ventures investment announcement<\/a> identifies Inherent\u2019s founders as Tantum Collins, Edward Hughes, Louis Kirsch and Kaloyan Aleksiev. Collins, Hughes and Kirsch have backgrounds at Google DeepMind, while Aleksiev has infrastructure experience from Reka AI and Microsoft. Edward Hughes serves as co-founder and chief scientist.<\/p>\n<p>Faraday\u2019s reported results do not establish that a 27-billion-parameter model is broadly superior to frontier systems for scientific work. Replica is a benchmark developed by Inherent, the main evaluation uses an automated judge, and the held-out test set contains 68 tasks. Independent testing of the system on other benchmarks and research environments would provide additional evidence about how broadly the results generalize.<\/p>\n<p>For now, Inherent\u2019s release provides a concrete demonstration of a different approach to AI research systems: rather than relying on one model to perform every part of an experiment, Faraday combines a relatively small model trained for scientific decision-making with a more powerful coding agent that carries out the technical work.<\/p>\n<p>The company is now positioning that architecture as a possible path from AI systems that reproduce known results to systems that can conduct increasingly open-ended research. Whether that transition produces reliable scientific discoveries remains an open question, but the Replica results give Inherent a measurable starting point for testing the approach.<\/p>\n<h2>References<\/h2>\n<ul>\n<li><a href=\"https:\/\/arxiv.org\/abs\/2608.13331\">Training AI Scientists to Replicate Research \u2014 arXiv research paper<\/a><\/li>\n<li><a href=\"https:\/\/inherentlabs.ai\/research\/training-to-replicate\">Inherent: Training AI Scientists to Replicate Research<\/a><\/li>\n<li><a href=\"https:\/\/www.indexventures.com\/perspectives\/inherent-designing-for-discovery\/\">Index Ventures: Inherent and its AI research approach<\/a><\/li>\n<li><a href=\"https:\/\/techcrunch.com\/2026\/08\/22\/inherent-founded-by-deepmind-alumni-says-its-ai-teammate-just-outperformed-anthropic-and-openai-at-replicating-research\/\">TechCrunch: Inherent\u2019s Faraday research agent<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>London-based AI research lab Inherent has released Faraday, a 27-billion-parameter artificial intelligence agent designed to reproduce results from scientific research papers. The company says Faraday outperformed Anthropic\u2019s Claude Opus 4.8 and OpenAI\u2019s GPT-5.5 on a benchmark created to test how effectively AI systems can replicate published research. Inherent, which was founded by former Google DeepMind [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":6602,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"content-type":"","_monsterinsights_skip_tracking":false,"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_post_was_ever_published":false},"categories":[164],"tags":[166],"class_list":["post-6601","post","type-post","status-publish","format-standard","has-post-thumbnail","category-tech-updates","tag-ai"],"share_on_mastodon":{"url":"https:\/\/mastodon.social\/@Areeblog\/117149718114917917","error":""},"yoast_head":"<!-- This site is optimized with the Yoast SEO Premium plugin v28.4 (Yoast SEO v28.5) - https:\/\/yoast.com\/product\/yoast-seo-premium-wordpress\/ -->\n<title>Inherent\u2019s 27B Faraday Agent Beats GPT-5.5 and Claude Opus 4.8 on Scientific Replication Benchmark - Aree Blog<\/title>\n<meta name=\"description\" content=\"Inherent\u2019s 27B Faraday AI agent beats GPT-5.5 and Claude Opus 4.8 in scientific research replication tests.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Inherent\u2019s 27B Faraday Agent Beats GPT-5.5 and Claude Opus 4.8 on Scientific Replication Benchmark\" \/>\n<meta property=\"og:description\" content=\"Inherent\u2019s 27B Faraday AI agent beats GPT-5.5 and Claude Opus 4.8 in scientific research replication tests.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/\" \/>\n<meta property=\"og:site_name\" content=\"Aree Blog\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-24T09:03:58+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-37.jpeg\" \/>\n\t<meta property=\"og:image:width\" content=\"739\" \/>\n\t<meta property=\"og:image:height\" content=\"415\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"Daniel Chinonso John\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Daniel Chinonso John\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"6 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\\\/\"},\"author\":{\"name\":\"Daniel Chinonso John\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/#\\\/schema\\\/person\\\/d972222c55618fb0f4b4c0c11ff52f63\"},\"headline\":\"Inherent\u2019s 27B Faraday Agent Beats GPT-5.5 and Claude Opus 4.8 on Scientific Replication Benchmark\",\"datePublished\":\"2026-08-24T09:03:58+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\\\/\"},\"wordCount\":1275,\"commentCount\":0,\"image\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/areeblog.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/images-37.jpeg\",\"keywords\":[\"AI\"],\"articleSection\":[\"Tech Updates\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/areeblog.com\\\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\\\/\",\"url\":\"https:\\\/\\\/areeblog.com\\\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\\\/\",\"name\":\"Inherent\u2019s 27B Faraday Agent Beats GPT-5.5 and Claude Opus 4.8 on Scientific Replication Benchmark - Aree Blog\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/areeblog.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/images-37.jpeg\",\"datePublished\":\"2026-08-24T09:03:58+00:00\",\"author\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/#\\\/schema\\\/person\\\/d972222c55618fb0f4b4c0c11ff52f63\"},\"description\":\"Inherent\u2019s 27B Faraday AI agent beats GPT-5.5 and Claude Opus 4.8 in scientific research replication tests.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/areeblog.com\\\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/areeblog.com\\\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\\\/#primaryimage\",\"url\":\"https:\\\/\\\/areeblog.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/images-37.jpeg\",\"contentUrl\":\"https:\\\/\\\/areeblog.com\\\/wp-content\\\/uploads\\\/2026\\\/08\\\/images-37.jpeg\",\"width\":739,\"height\":415,\"caption\":\"Inherent\u2019s 27B Faraday Agent Beats GPT-5.5 and Claude Opus 4.8 on Scientific Replication Benchmark\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/areeblog.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Inherent\u2019s 27B Faraday Agent Beats GPT-5.5 and Claude Opus 4.8 on Scientific Replication Benchmark\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/#website\",\"url\":\"https:\\\/\\\/areeblog.com\\\/\",\"name\":\"Aree Blog\",\"description\":\"Unfiltered Perspectives, Unstoppable Insights\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/areeblog.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/areeblog.com\\\/#\\\/schema\\\/person\\\/d972222c55618fb0f4b4c0c11ff52f63\",\"name\":\"Daniel Chinonso John\",\"description\":\"Daniel Chinonso John is a web designer, penetration tester, and founder of Aree Tech. He writes clear, actionable posts at the intersection of productivity, AI, cybersecurity, and blogging to help readers get things done.\",\"sameAs\":[\"https:\\\/\\\/www.linkedin.com\\\/in\\\/daniel-john-45183a169\\\/\"],\"url\":\"https:\\\/\\\/areeblog.com\\\/author\\\/danojohn55gmail-com\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO Premium plugin. -->","yoast_head_json":{"title":"Inherent\u2019s 27B Faraday Agent Beats GPT-5.5 and Claude Opus 4.8 on Scientific Replication Benchmark - Aree Blog","description":"Inherent\u2019s 27B Faraday AI agent beats GPT-5.5 and Claude Opus 4.8 in scientific research replication tests.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/","og_locale":"en_US","og_type":"article","og_title":"Inherent\u2019s 27B Faraday Agent Beats GPT-5.5 and Claude Opus 4.8 on Scientific Replication Benchmark","og_description":"Inherent\u2019s 27B Faraday AI agent beats GPT-5.5 and Claude Opus 4.8 in scientific research replication tests.","og_url":"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/","og_site_name":"Aree Blog","article_published_time":"2026-08-24T09:03:58+00:00","og_image":[{"width":739,"height":415,"url":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-37.jpeg","type":"image\/jpeg"}],"author":"Daniel Chinonso John","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Daniel Chinonso John","Est. reading time":"6 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/#article","isPartOf":{"@id":"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/"},"author":{"name":"Daniel Chinonso John","@id":"https:\/\/areeblog.com\/#\/schema\/person\/d972222c55618fb0f4b4c0c11ff52f63"},"headline":"Inherent\u2019s 27B Faraday Agent Beats GPT-5.5 and Claude Opus 4.8 on Scientific Replication Benchmark","datePublished":"2026-08-24T09:03:58+00:00","mainEntityOfPage":{"@id":"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/"},"wordCount":1275,"commentCount":0,"image":{"@id":"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/#primaryimage"},"thumbnailUrl":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-37.jpeg","keywords":["AI"],"articleSection":["Tech Updates"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/","url":"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/","name":"Inherent\u2019s 27B Faraday Agent Beats GPT-5.5 and Claude Opus 4.8 on Scientific Replication Benchmark - Aree Blog","isPartOf":{"@id":"https:\/\/areeblog.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/#primaryimage"},"image":{"@id":"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/#primaryimage"},"thumbnailUrl":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-37.jpeg","datePublished":"2026-08-24T09:03:58+00:00","author":{"@id":"https:\/\/areeblog.com\/#\/schema\/person\/d972222c55618fb0f4b4c0c11ff52f63"},"description":"Inherent\u2019s 27B Faraday AI agent beats GPT-5.5 and Claude Opus 4.8 in scientific research replication tests.","breadcrumb":{"@id":"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/#primaryimage","url":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-37.jpeg","contentUrl":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-37.jpeg","width":739,"height":415,"caption":"Inherent\u2019s 27B Faraday Agent Beats GPT-5.5 and Claude Opus 4.8 on Scientific Replication Benchmark"},{"@type":"BreadcrumbList","@id":"https:\/\/areeblog.com\/inherents-27b-faraday-agent-beats-gpt-5-5-and-claude-opus-4-8-on-scientific-replication-benchmark\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/areeblog.com\/"},{"@type":"ListItem","position":2,"name":"Inherent\u2019s 27B Faraday Agent Beats GPT-5.5 and Claude Opus 4.8 on Scientific Replication Benchmark"}]},{"@type":"WebSite","@id":"https:\/\/areeblog.com\/#website","url":"https:\/\/areeblog.com\/","name":"Aree Blog","description":"Unfiltered Perspectives, Unstoppable Insights","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/areeblog.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Person","@id":"https:\/\/areeblog.com\/#\/schema\/person\/d972222c55618fb0f4b4c0c11ff52f63","name":"Daniel Chinonso John","description":"Daniel Chinonso John is a web designer, penetration tester, and founder of Aree Tech. He writes clear, actionable posts at the intersection of productivity, AI, cybersecurity, and blogging to help readers get things done.","sameAs":["https:\/\/www.linkedin.com\/in\/daniel-john-45183a169\/"],"url":"https:\/\/areeblog.com\/author\/danojohn55gmail-com\/"}]}},"jetpack_sharing_enabled":true,"jetpack-related-posts":[{"id":6913,"url":"https:\/\/areeblog.com\/stepfuns-600b-step-5-targets-ai-coding-agents-with-1m-token-context\/","url_meta":{"origin":6601,"position":0},"title":"StepFun\u2019s 600B Step 5 Targets AI Coding Agents With 1M-Token Context","author":"Daniel Chinonso John","date":"September 20, 2026","format":false,"excerpt":"Chinese artificial intelligence company StepFun has introduced Step 5 Preview, a new flagship model designed for software engineering, AI agents, professional knowledge work and finance. The company says the model is built to handle long-running tasks rather than only generate individual answers. Step 5 Preview uses a sparse Mixture-of-Experts architecture\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"StepFun\u2019s 600B Step 5 Targets AI Coding Agents With 1M-Token Context","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/images-62.jpeg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/images-62.jpeg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/images-62.jpeg?resize=525%2C300&ssl=1 1.5x"},"classes":[]},{"id":4905,"url":"https:\/\/areeblog.com\/z-ai-releases-glm-4-5-series-high-performance-open-source-models-focused-on-agent-capabilities\/","url_meta":{"origin":6601,"position":1},"title":"Z.AI Releases GLM 4.5 Series: High-Performance, Open-Source Models Focused on Agent Capabilities","author":"Samuel Ogori","date":"August 2, 2025","format":false,"excerpt":"Z.AI (formerly Zepoo AI) has launched the GLM 4.5 series, comprising the flagship GLM 4.5 model and the lighter GLM 4.5 Air. Positioned as a significant open-source release in 2025, these models emphasize a balance of performance, efficiency, agent capabilities, and cost. Model Architecture and Efficiency GLM 4.5: A 355\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"Z.AI Releases GLM 4.5 Series: High-Performance, Open-Source Models Focused on Agent Capabilities","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/08\/file_00000000f38c6246b5d683f4a05174d4.png?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/08\/file_00000000f38c6246b5d683f4a05174d4.png?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/08\/file_00000000f38c6246b5d683f4a05174d4.png?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2025\/08\/file_00000000f38c6246b5d683f4a05174d4.png?resize=700%2C400&ssl=1 2x"},"classes":[]},{"id":6835,"url":"https:\/\/areeblog.com\/nsa-fbi-and-cisa-accuse-china-ai-firms-of-stealing-u-s-model-capabilities\/","url_meta":{"origin":6601,"position":2},"title":"NSA, FBI and CISA Accuse China AI Firms of Stealing U.S. Model Capabilities","author":"Daniel Chinonso John","date":"September 8, 2026","format":false,"excerpt":"The U.S. National Security Agency, Federal Bureau of Investigation and Cybersecurity and Infrastructure Security Agency have accused six China-based artificial intelligence companies of conducting industrial-scale campaigns to extract capabilities from U.S. frontier AI models. In a joint cybersecurity advisory released on September 8, 2026, the agencies said the campaigns involve\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"NSA, FBI and CISA Accuse China AI Firms of Stealing U.S. Model Capabilities","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/970a355d-2615-4e71-9d32-c5dad14b8324_860e4f66.jpg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/970a355d-2615-4e71-9d32-c5dad14b8324_860e4f66.jpg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/970a355d-2615-4e71-9d32-c5dad14b8324_860e4f66.jpg?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/970a355d-2615-4e71-9d32-c5dad14b8324_860e4f66.jpg?resize=700%2C400&ssl=1 2x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/970a355d-2615-4e71-9d32-c5dad14b8324_860e4f66.jpg?resize=1050%2C600&ssl=1 3x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/970a355d-2615-4e71-9d32-c5dad14b8324_860e4f66.jpg?resize=1400%2C800&ssl=1 4x"},"classes":[]},{"id":6511,"url":"https:\/\/areeblog.com\/anthropic-finds-ai-agents-can-sabotage-each-other-when-their-goals-collide\/","url_meta":{"origin":6601,"position":3},"title":"Anthropic Finds AI Agents Can Sabotage Each Other When Their Goals Collide","author":"Daniel Chinonso John","date":"August 14, 2026","format":false,"excerpt":"Anthropic researchers have found that AI agents working on the same software project can turn against one another when they are given incompatible objectives, with some models disabling competing processes, interfering with other agents\u2019 work and deploying malicious code during controlled tests. The findings were published by Anthropic on Thursday,\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"Anthropic Finds AI Agents Can Sabotage Each Other When Their Goals Collide","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260814-WA0008.jpg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260814-WA0008.jpg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260814-WA0008.jpg?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260814-WA0008.jpg?resize=700%2C400&ssl=1 2x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260814-WA0008.jpg?resize=1050%2C600&ssl=1 3x"},"classes":[]},{"id":6657,"url":"https:\/\/areeblog.com\/ai-loss-of-control-reports-nearly-doubled-in-july\/","url_meta":{"origin":6601,"position":4},"title":"AI Loss-of-Control Reports Nearly Doubled in July","author":"Daniel Chinonso John","date":"August 30, 2026","format":false,"excerpt":"AI systems are being reported increasingly often for behaviours that bypass user instructions and human controls, with a monitoring project recording a sharp rise in reported loss-of-control incidents during the latest period examined. The Centre for Long-Term Resilience (CLTR), which operates the Loss of Control Observatory, said it had identified\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"AI Loss-of-Control Reports Nearly Doubled in July","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260830-WA0009.jpg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260830-WA0009.jpg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260830-WA0009.jpg?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260830-WA0009.jpg?resize=700%2C400&ssl=1 2x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/08\/IMG-20260830-WA0009.jpg?resize=1050%2C600&ssl=1 3x"},"classes":[]},{"id":6732,"url":"https:\/\/areeblog.com\/crowdstrike-and-openai-move-ai-agent-security-from-monitoring-to-runtime-enforcement\/","url_meta":{"origin":6601,"position":5},"title":"CrowdStrike and OpenAI Move AI-Agent Security From Monitoring to Runtime Enforcement","author":"Daniel Chinonso John","date":"September 2, 2026","format":false,"excerpt":"CrowdStrike and OpenAI have expanded their partnership to secure AI agents while they are operating, bringing OpenAI\u2019s Codex agents into CrowdStrike\u2019s Falcon Guardian security platform and adding OpenAI\u2019s GPT-5.6 Cyber to planned Falcon capabilities. The companies announced the partnership on September 2, 2026, during Fal.Con 2026 in Las Vegas. The\u2026","rel":"","context":"In &quot;Tech Updates&quot;","block_context":{"text":"Tech Updates","link":"https:\/\/areeblog.com\/category\/tech-updates\/"},"img":{"alt_text":"CrowdStrike and OpenAI Move AI-Agent Security From Monitoring to Runtime Enforcement","src":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/Falcon-780x470-1.jpg?resize=350%2C200&ssl=1","width":350,"height":200,"srcset":"https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/Falcon-780x470-1.jpg?resize=350%2C200&ssl=1 1x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/Falcon-780x470-1.jpg?resize=525%2C300&ssl=1 1.5x, https:\/\/i0.wp.com\/areeblog.com\/wp-content\/uploads\/2026\/09\/Falcon-780x470-1.jpg?resize=700%2C400&ssl=1 2x"},"classes":[]}],"jetpack_featured_media_url":"https:\/\/areeblog.com\/wp-content\/uploads\/2026\/08\/images-37.jpeg","_links":{"self":[{"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/posts\/6601","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/comments?post=6601"}],"version-history":[{"count":1,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/posts\/6601\/revisions"}],"predecessor-version":[{"id":6603,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/posts\/6601\/revisions\/6603"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/media\/6602"}],"wp:attachment":[{"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/media?parent=6601"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/categories?post=6601"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/areeblog.com\/wp-json\/wp\/v2\/tags?post=6601"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}