Anthropic has published full benchmark results for Claude Opus 5.5, the first model in its new Claude 5.5 family, and the numbers put it at the top of the vendor's own agentic coding and knowledge-work tables — while undercutting its predecessors sharply on price.

On the published benchmarks, Opus 5.5 scores 66.4% on Terminal-Bench 4.0, ahead of Claude Fable 5.1 (55.8%), OpenAI's GPT-6 Astra (57.9%) and GPT-5.6 Sol (37.3%). It posts 54.4% on FrontierCode v1.1 and leads Anthropic's own models on CursorBench 4.0 at 57.8%. On GDPval-AA v2.1, which tests real-world professional work across 44 occupations, it scores 1846 Elo against 1735 for Fable 5.1 and 1708 for Opus 5.

Anthropic is candid that the margins matter less than they look. “At these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences,” the company writes, adding that the gap between Opus 5.5 and Fable 5.1 is narrower in practice than the scores suggest. Astra still wins some tests: 64.6% on Terminal-Bench-Science 0.1 versus 58.7% for Opus 5.5.

Where the advantage is clearest is efficiency. Anthropic says Opus 5.5 costs 40% less than Opus 5 on typical workloads — $4 per million input tokens and $20 per million output tokens, 20% below Opus 5, with cache reads at $0.20 per million, 60% cheaper. Output is generated more than 30% faster. Internal case studies cite a 680,000-line code migration finished in under a day, a 200,000-line audit completed in under three hours where Opus 5 needed over 20, and a C-to-Rust rewrite of HAProxy finished in 9.5 hours against 12 for Fable 5.1 at 51% lower cost.

Safety is part of the pitch. On Anthropic's automated behavioural audit of nearly 2,000 scenarios, Opus 5.5 scored best of any model the company has tested; in a new containment evaluation it attempted to cross boundaries roughly 85% less often than Opus 5 or Mythos 5.1, with every attempt low-severity and self-reported. On a benchmark run by the security firm Gray Swan it ties Fable 5.1 for the lowest prompt-injection success rate of any model tested. Because its biology and cybersecurity capabilities are comparable to Mythos 5.1, it ships with Fable 5.1-class safeguards: most cyber tasks are transparently re-routed to Opus 4.8, and vetted organisations must join verification programmes for cyber and life-sciences work.

The release also formalises Anthropic's “pacing the frontier” posture: Opus 5.5 is the first model shipped since CEO Dario Amodei argued for deliberately slowing capability advances so safety work can keep up. Sonnet 5.5 and Haiku 5.5 are due “in the coming weeks.”