DeepMind Alumni Built a 27B AI “Teammate” That Beats GPT-5.5 and Claude at Science — Why Small Models Are Winning

Introduction: A David-vs-Goliath Result in the Lab

The scoreboard from the AI frontier this weekend reads like an upset story. As TechCrunch reported on August 22, 2026, Inherent — a London AI lab founded by Google DeepMind alumni, just weeks out of stealth with a $50 million seed round — released an AI agent called Faraday that outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 at a genuinely hard task: independently reproducing the findings of published scientific papers, without being told the answer in advance.

Here’s the part that should stop every AI tool buyer mid-scroll: Faraday doesn’t run on a frontier-scale model at all. It runs on Qwen 3.6, a model with just 27 billion parameters — a sliver of the size (and training cost) of the giants it beat. If you’ve been assuming the biggest model always wins, this result — like Nvidia’s recent finding that the harness matters more than the model — says otherwise.

What Inherent’s Faraday Actually Did

The task was paper replication: take a published study, and reproduce its findings from scratch, unaided. That may sound like a party trick next to Inherent’s stated goal — AI that discovers new scientific knowledge rather than verifying old results — but cofounder and chief scientist Edward Hughes likens it to the standard first exercise of a scientific career: “Many PhD students actually start by doing this.”

Three details from TechCrunch’s report stand out:

Measured against Claude Opus 4.8 and GPT-5.5 — both far larger systems — Faraday came out ahead on the task it was shaped for, while running on a comparative fraction of the compute.

How a 27B Model Beats a Frontier Giant

How does a 27-billion-parameter model out-research systems with vastly more scale? The recipe Inherent describes is becoming 2026’s most repeated pattern in applied AI:

  1. Radical specialization. Faraday is tuned for one workflow — the cycle of hypothesizing, designing, running, and checking experiments — instead of everything at once.
  2. Outcome-based reinforcement learning. Rather than hand-written rules, the agent is rewarded for good research outcomes. That, Inherent bets, is how you teach something as intangible as taste — and why reward-trained judgment should generalize better than knowledge of the meta-science of research itself.
  3. Cheap iteration. Parameters are a proxy for training and serving cost. A 27B model can be retrained, fine-tuned, and rerun far more cheaply than a frontier system, so the feedback loop spins faster.

The same logic is playing out across the tool landscape, from Qwen-class 27B models rivaling frontier coding agents to task-tuned vertical agents in law, finance, and medicine. Small and focused keeps beating big and general at the workflow level.

The Composability Lesson: Great Agents Borrow Tools

Perhaps the most practical detail: Inherent didn’t build Faraday a coding tool. Faraday uses OpenAI’s GPT-5.5 Codex for its code, the way human scientists lean on existing software instead of building their own instruments. Even a lab trying to out-research OpenAI happily rents OpenAI’s coding agent.

That composability instinct — orchestrating best-in-class pieces rather than rebuilding them — mirrors what multi-agent orchestration frameworks found earlier this year, as we covered in our piece on Sakana’s multi-agent orchestration research. If you’re assembling an AI stack in 2026, expect it to be multi-vendor by design: a research brain, a coding hand, a search layer, each swappable as better ones ship.

AI Research Agents You Can Use Today

Faraday itself isn’t a product you can sign up for — Inherent is a dozen-person research team (growing to 20–25 by year’s end, per TechCrunch) with ambitions in world models, not a SaaS company. But the category it points toward — AI agents that read, verify, and synthesize research for you — is very much usable now. These are the standouts in our directory:

ToolWhat it does for researchPricing
ElicitScreens papers and extracts methods, sample sizes, and findings into structured tables — closest thing to a literature-review agentFreemium
ConsensusAnswers research questions with evidence synthesized across indexed papers, showing the agreement levelFreemium
SciSpaceCopilot for reading and understanding papers — explanations, follow-up questions, literature summariesFreemium
ResearchRabbitMaps citation graphs so one seed paper unfolds into the whole related-work neighborhoodFree
ScholarAIAI paper search with summaries and citation-ready exports for fast evidence gatheringFreemium
Semantic ScholarFree AI-powered literature index with influential-citation ranking across 200M+ papersFree
PerplexityDeep Research mode chains searches into cited, report-length answers on any questionFreemium

The honest framing: today’s tools help you find and digest research. Faraday is a preview of the next step — agents that do research, running experiments and reproducing results while you sleep. For a fuller map of the space, see our guide to the best AI research tools.

What Faraday’s Win Means for the Tools You Pick

Three rules worth writing down before your next AI subscription:

And keep an eye on the “taste” trend specifically: reward-trained judgment is migrating from frontier research labs into vertical agents everywhere, and it’s the difference between an assistant that agrees with you and a teammate that surprises you.

Frequently Asked Questions

What is Inherent’s Faraday agent?

Faraday is an AI research agent from Inherent, a London lab founded by Google DeepMind alumni that emerged from stealth in 2026 with a $50 million seed round. It independently reproduces the findings of published scientific papers — and, per the company, outperformed Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 at the task.

How big is the model behind Faraday?

Faraday runs on Qwen 3.6, a model with roughly 27 billion parameters — far smaller than the frontier-scale Claude Opus 4.8 and GPT-5.5 systems it was benchmarked against. Parameters are a proxy for model size and, typically, training and serving cost.

Can I use Faraday for my own research?

Not yet. Inherent is a research lab, not a product company — its stated goal is AI that discovers new scientific knowledge. For usable AI research help today, tools like Elicit, Consensus, SciSpace, ResearchRabbit, ScholarAI, and Perplexity cover literature discovery and synthesis.

What is “research taste” in AI agents?

Inherent’s term for an agent’s instinct about which experiments are worth running and how to design them well — beyond merely getting accurate results. The team trains it with reinforcement learning, rewarding good research outcomes rather than hand-coding rules.

Build Your AI Research Stack

Compare AI research assistants, deep-research chatbots, and hundreds of other vetted AI tools on aitrove.ai — with pricing, use cases, and honest breakdowns — before the next small model upends your workflow.

Browse All AI Tools →