The Mythical Fable 5.1 Is Here!

AIClaude

Anthropic has rewritten the coordinates of machine intelligence once again.

Today, Claude Fable 5.1 is live across all platforms, available to everyone.

The fully uncaged Mythos 5.1 remains locked behind a whitelist, exclusively for vetted cybersecurity and life sciences teams.

Claude Fable 5.1

Claude Fable 5.1

From重构复杂的工程代码 to tackling long-horizon high-dimensional scientific research, this twin pair is redefining what a top-tier「cyber brain worker」looks like.

Their specialty: taking on the kinds of complex, multi-step tasks that used to stump even human experts.

The new model's benchmark scores pull out a断层式代差 gap over the previous generation.

Benchmark results

  • Scientific Research (Terminal-Bench-Science 0.1)

Fable 5.1 scores 52.6%. The previous generation Fable 5 and OpenAI's GPT-5.6 Sol scored just 24.7% and 22.4% respectively.

  • Code Engineering (Terminal-Bench 4.0)

Fable 5.1 surges from 42.0% to 55.8%; with Mythos 5.1's guardrails removed, the number climbs to a striking 60.9%.

  • Complex Cognition & Knowledge Work

On the knowledge-work benchmark GDPval-AA, Fable 5.1 scores 1853; on the「last exam for humanity」difficult benchmark, tool-equipped real-world testing crosses the 65.0% threshold.

Even with「thinking effort」dialed to minimum, it still outperforms the prior generation at lower token consumption.

Performance charts

Performance charts

Performance charts

Performance charts

But the real breakthrough isn't capability — it's cost.

Base pricing for Fable 5.1 is identical to the previous generation: $10 per million tokens input, $50 output.

However, on the most resource-intensive「cached read」, it drops from $1.00 to $0.25 — a 75% reduction.

In real tasks, typical scenarios see a 25% cost drop; for heavy long-horizon agent workloads, savings reach up to 45%.

By comparison, GPT-5.6 Sol's current short-context promotional pricing: $2 input, $10 output, $0.20 cache.

Pricing comparison

Pricing comparison

Benchmarks and price cuts prove its strength — but that's just the surface.

The real dividing line: starting with this generation, AI has fully shed its assistant-tool skin.

It is now doing science itself.

The Claude Twin Kings Enter Scientific Research

Remapping Venus, Hand-Crafting GPU Kernels

To prove this twin pair's real-world devastation, Anthropic rolled out three frontier cases far more直观 than any leaderboard.

  • Remapping One-Third of Venus

Over 30 years ago, NASA's Magellan probe took a radar sweep of Venus.

But its elevation data resolution was only 10–20 km. Translated to Earth, an entire city would occupy just a few pixels.

Previously, humans had mapped only one-fifth of that data into elevation maps.

Fable 5.1 took this dataset, trained a neural network, and regenerated a high-resolution terrain map covering roughly one-third of Venus's surface.

New details pushed to 2–3 km, with height estimation accuracy improving up to 25%.

Venus map comparison

Venus terrain detail

  • Molecular & Protein Design, 50% Hit Rate

Handling static surface data is one thing; designing biological proteins is another order of magnitude harder.

Many modern drugs work by designing a small protein that precisely locks onto a target in the body, blocking or activating a specific reaction.

How tightly they bind is called binding affinity — the higher the affinity, the more likely the drug works at low doses.

After calling open-source design tools, Mythos 5.1's designs were sent directly to external labs for real-world wet lab experiments.

The results: strongest tier measured so far. On three targets including EGFR and Nipah G, binding affinity reached 10x the best Adaptyv Bio competition solutions.

Across 12 targets, the overall hit rate approaches 50%. The current industry norm is just 10–15%.

Protein design results

  • Hand-Written GPU Kernels, 2.5x Speedup for Bioinformatics Models

The third case may have rewritten the compute bills of research institutions.

Genomics and protein research often runs specialized deep learning models repeatedly on GPUs.

A full genome analysis tests大量的基因附近突变, with inference running thousands of times. Each millisecond of slowness compounds quickly.

Optimizing and hand-writing GPU kernels is typically the job of top performance engineers. A team would spend weeks; academic labs often can't afford to hire such people.

Mythos 5.1, armed only with publicly available source code, finished in days.

Not only did it hand-write GPU kernels for 7 open-source bioinformatics models, it achieved up to 2.5x speedup.

Critically, outputs match the originals exactly.

Translating to dollar terms: a 3-million-variant analysis that cost $18,000 on Evo 2 now costs $8,000.

GPU optimization comparison

GPU cost reduction

World's Strongest Coding AI, Jumping 14%

Fable 5.1's ability to do science directly comes from its extremely hardcore code logic.

Today's agents aren't afraid to write code — they're afraid of losing track of the original goal after hours of tool calls, or of never finding the real root cause after an error.

What Fable 5.1 reinforces is this complete action loop:

Read repo docs → break down tasks → make changes → run tests → check results → locate failures → continue to next round.

In Anthropic's words, when Fable 5.1 hits an obstacle it explicitly points to where it's stuck, rarely taking the「appears complete but leaves hidden traps」shortcut anymore.

Coding agent workflow

On CursorBench 3.2, measuring real IDE coding水平, it scored 73.4%,刷新内部纪录; on the commercial workflow AutomationBench, it跃升至31.4%, nearly doubling.

With thislong-horizon stability, it even helped hedge fund Millennium solve a years-long buried bug.

Facing a bug that crashed once per million runs, it traced to an external dependency, disassembled the third-party library, and根拔起 the root cause in someone else's territory.

This time, Fable 5.1 truly delivered a「complete long-chain task in one shot」.

Coding results

After Fable 5.1's full launch, a wave of real-world tests followed.

First wave benchmarks

From AI/ML API's public demo comparisons, the gap between Fable 5.1 and GPT-5.6 Sol is clear.

Same prompt, single-shot output for both.

Fable 5.1 cost $5.69; GPT-5.6 Sol used just $0.88.

Fable goes for realism — waves, depth of field, sand rings around islands are more逼真. Sol is minimalist, and maintains its风格 consistently across all 5 scenes, never breaking.

In one test, Anthropic researcher Alex Albert fed Fable 5.1 a photo of a plot of land.

It autonomously completed architectural design, high-quality rendering, and even generated a film-grade guided tour video.

In the following hardcore demo, Fable 5.1 hand-built a fully interactive human brain model.

Not only did it precisely replicate every fold of the cerebral cortex in code, it let you watch the complete neural pathway for「pass the salt」run in real time.

Interactive brain model

Generating an ARC survival game, Fable 5.1 passed in one shot.

ARC game generation

Finding Vulnerabilities, Rejecting Distillation

Anthropic's security moves this round follow a「one loosening, one tightening」principle.

Loosening: guardrails are no longer so「neurotic」.

Previously, programmers using Claude to review their own code were constantly flagged as hackers. Now Fable 5.1 finally allows vulnerability identification, cutting cybersecurity false positives by 60%.

The bottom line stays firm: only find vulnerabilities, never write exploitation code. High-risk operations like penetration testing still go to the stricter Opus.

Tightening: the moat becomes an「anti-distillation mechanism」.

To completely close the previous loophole of modifying context in multi-turn conversations to extract Claude's「thinking chain」, Fable 5.1 locked the API at the底层.

From now on, the API will scrutinize context strictly.

If your submitted prompt doesn't match what generated the thinking block, the API either errors and blocks, or strips the entire thinking block — no exceptions.

Currently, this mechanism targets new API accounts registered after August 31st.

But Anthropic has stated: for all future new models, this lock will be mandatory for everyone.

Security features

Security features

Silicon Valley's top researchers have repeatedly articulated a vision: AI's greatest value will compress decades of biomedical and basic science progress into a few years.

From remapping Venus's terrain, to multiplying protein hit rates in test tubes, this twin pair is making that预言 a reality.

The threshold for scientific discovery is visibly collapsing.

Just don't forget: the hand that passed the torch is simultaneously, quietly pulling up the ladder behind it with「anti-distillation」.

Comments

Sign in to leave a comment Sign in

Loading…

Back to AI News