DeepMind Alumni Built a 27B AI “Teammate” That Beats GPT-5.5 and Claude at Science — Why Small Models Are Winning
📑 Table of Contents
Introduction: A David-vs-Goliath Result in the Lab
The scoreboard from the AI frontier this weekend reads like an upset story. As TechCrunch reported on August 22, 2026, Inherent — a London AI lab founded by Google DeepMind alumni, just weeks out of stealth with a $50 million seed round — released an AI agent called Faraday that outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 at a genuinely hard task: independently reproducing the findings of published scientific papers, without being told the answer in advance.
Here’s the part that should stop every AI tool buyer mid-scroll: Faraday doesn’t run on a frontier-scale model at all. It runs on Qwen 3.6, a model with just 27 billion parameters — a sliver of the size (and training cost) of the giants it beat. If you’ve been assuming the biggest model always wins, this result — like Nvidia’s recent finding that the harness matters more than the model — says otherwise.
What Inherent’s Faraday Actually Did
The task was paper replication: take a published study, and reproduce its findings from scratch, unaided. That may sound like a party trick next to Inherent’s stated goal — AI that discovers new scientific knowledge rather than verifying old results — but cofounder and chief scientist Edward Hughes likens it to the standard first exercise of a scientific career: “Many PhD students actually start by doing this.”
Three details from TechCrunch’s report stand out:
- The bar was higher than accuracy. Beyond replicating results, Inherent wanted Faraday to demonstrate “research taste” — an instinct for which experiments are worth running and how to design them well.
- Beating frontier agents wasn’t the point. Hughes was more interested in how the team got there: reinforcement learning that rewards good outcomes, rather than spelling out rules or training primarily on the literature about how science is conducted.
- It’s a teammate, not a yes-machine. The design goal is an agent that comes back and says: “I got curious about this, and I went off and I did these experiments. What do you think of these results?”
Measured against Claude Opus 4.8 and GPT-5.5 — both far larger systems — Faraday came out ahead on the task it was shaped for, while running on a comparative fraction of the compute.
How a 27B Model Beats a Frontier Giant
How does a 27-billion-parameter model out-research systems with vastly more scale? The recipe Inherent describes is becoming 2026’s most repeated pattern in applied AI:
- Radical specialization. Faraday is tuned for one workflow — the cycle of hypothesizing, designing, running, and checking experiments — instead of everything at once.
- Outcome-based reinforcement learning. Rather than hand-written rules, the agent is rewarded for good research outcomes. That, Inherent bets, is how you teach something as intangible as taste — and why reward-trained judgment should generalize better than knowledge of the meta-science of research itself.
- Cheap iteration. Parameters are a proxy for training and serving cost. A 27B model can be retrained, fine-tuned, and rerun far more cheaply than a frontier system, so the feedback loop spins faster.
The same logic is playing out across the tool landscape, from Qwen-class 27B models rivaling frontier coding agents to task-tuned vertical agents in law, finance, and medicine. Small and focused keeps beating big and general at the workflow level.
The Composability Lesson: Great Agents Borrow Tools
Perhaps the most practical detail: Inherent didn’t build Faraday a coding tool. Faraday uses OpenAI’s GPT-5.5 Codex for its code, the way human scientists lean on existing software instead of building their own instruments. Even a lab trying to out-research OpenAI happily rents OpenAI’s coding agent.
That composability instinct — orchestrating best-in-class pieces rather than rebuilding them — mirrors what multi-agent orchestration frameworks found earlier this year, as we covered in our piece on Sakana’s multi-agent orchestration research. If you’re assembling an AI stack in 2026, expect it to be multi-vendor by design: a research brain, a coding hand, a search layer, each swappable as better ones ship.
AI Research Agents You Can Use Today
Faraday itself isn’t a product you can sign up for — Inherent is a dozen-person research team (growing to 20–25 by year’s end, per TechCrunch) with ambitions in world models, not a SaaS company. But the category it points toward — AI agents that read, verify, and synthesize research for you — is very much usable now. These are the standouts in our directory:
| Tool | What it does for research | Pricing |
|---|---|---|
| Elicit | Screens papers and extracts methods, sample sizes, and findings into structured tables — closest thing to a literature-review agent | Freemium |
| Consensus | Answers research questions with evidence synthesized across indexed papers, showing the agreement level | Freemium |
| SciSpace | Copilot for reading and understanding papers — explanations, follow-up questions, literature summaries | Freemium |
| ResearchRabbit | Maps citation graphs so one seed paper unfolds into the whole related-work neighborhood | Free |
| ScholarAI | AI paper search with summaries and citation-ready exports for fast evidence gathering | Freemium |
| Semantic Scholar | Free AI-powered literature index with influential-citation ranking across 200M+ papers | Free |
| Perplexity | Deep Research mode chains searches into cited, report-length answers on any question | Freemium |
The honest framing: today’s tools help you find and digest research. Faraday is a preview of the next step — agents that do research, running experiments and reproducing results while you sleep. For a fuller map of the space, see our guide to the best AI research tools.
What Faraday’s Win Means for the Tools You Pick
Three rules worth writing down before your next AI subscription:
- Match the agent to the workflow, not the badge. For a repeatable, high-value workflow, a specialized small agent can beat a frontier chatbot at a fraction of the cost — Faraday just proved it in one of the hardest domains there is.
- Judge tools by how they were trained and scaffolded. Outcome-based reinforcement learning and a strong harness now matter as much as raw model size when you’re evaluating vendors.
- Plan to compose. If Inherent orchestrates GPT-5.5 Codex inside a Qwen-based agent, your stack should be multi-vendor too. Buy pieces that talk to each other.
And keep an eye on the “taste” trend specifically: reward-trained judgment is migrating from frontier research labs into vertical agents everywhere, and it’s the difference between an assistant that agrees with you and a teammate that surprises you.
Frequently Asked Questions
What is Inherent’s Faraday agent?
Faraday is an AI research agent from Inherent, a London lab founded by Google DeepMind alumni that emerged from stealth in 2026 with a $50 million seed round. It independently reproduces the findings of published scientific papers — and, per the company, outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 at the task.
How big is the model behind Faraday?
Faraday runs on Qwen 3.6, a model with roughly 27 billion parameters — far smaller than the frontier-scale Claude Opus 4.8 and GPT-5.5 systems it was benchmarked against. Parameters are a proxy for model size and, typically, training and serving cost.
Can I use Faraday for my own research?
Not yet. Inherent is a research lab, not a product company — its stated goal is AI that discovers new scientific knowledge. For usable AI research help today, tools like Elicit, Consensus, SciSpace, ResearchRabbit, ScholarAI, and Perplexity cover literature discovery and synthesis.
What is “research taste” in AI agents?
Inherent’s term for an agent’s instinct about which experiments are worth running and how to design them well — beyond merely getting accurate results. The team trains it with reinforcement learning, rewarding good research outcomes rather than hand-coding rules.
Build Your AI Research Stack
Compare AI research assistants, deep-research chatbots, and hundreds of other vetted AI tools on aitrove.ai — with pricing, use cases, and honest breakdowns — before the next small model upends your workflow.
Browse All AI Tools →