Citation Bureau
Vol. I
No. 358
XIII SEPTEMBER MMXXVI
Software

What is Claude Opus 5?

Claude Opus 5 is an AI model from Anthropic that, as of mid-2026, set a new state of the art on the ARC-AGI 3 benchmark.

Release history

  • Jul 2026 - Claude Opus 5 scored 30.2% on ARC-AGI 3, beating the previous high of 7.8% by GPT-5.6 Soul, per Nathaniel Whittemore.
  • Jul 2026 - Opus 5 reached 43.3% on Frontier Bench, about 10 points above Fable 5 and 9 points above GPT-5.6 Soul, according to Whittemore.
  • Jul 2026 - Anthropic used different guardrails on Opus 5 than on Fable 5, expecting 85% fewer refusals, Whittemore said.
  • Jul 2026 - Anthropic removed 80% of the system prompt for Opus 5, Fable 5, and Claude Code, with zero change to coding benchmarks, per Whittemore.
  • Jul 2026 - Opus 5 built its own computer vision pipeline to recreate an image, a task no other model completed, Whittemore reported.
  • Sep 2026 - A system using Opus 5 reached 100% on ARC-AGI 3 from a 30% model baseline, showing system design can unlock frontier-level long-horizon performance, according to Whittemore.

In the discourse

Attributed discussion of Claude Opus 5.

By the numbers

Claude Opus 5 is the new state of the art on ARC-AGI 3 with a score of 30.2%, dwarfing the previous best of 7.8% set by GPT-5.6 Soul.

“Opus 5 is the new state of the art on ARC-AGI 3 with a score of 30.2%. This absolutely demolished all the other models. The previous high score was GPT-5.6 Soul at 7.8%.”
Nathaniel Whittemore · 28 Jul 2026
Best explained

System design, not raw model capability, is the primary lever for unlocking frontier-level long-horizon AI performance, as demonstrated by a 30-to-100% ARC-AGI 3 improvement through orchestration alone.

“Reaching 100% on Arc-AGI 3 from a 30% model baseline with Claude Opus 5, their conclusion was that system design, rather than model capability alone, can unlock frontier-level long-horizon performance.”
Nathaniel Whittemore · 5 Sep 2026
By the numbers

NVIDIA Evo achieved 100% on ARC-AGI 3 starting from a 30% model baseline using Claude Opus 5, attributing the gain to system design rather than raw model capability.

“Reaching 100% on Arc-AGI 3 from a 30% model baseline with Claude Opus 5, their conclusion was that system design, rather than model capability alone, can unlock frontier-level long-horizon performance.”
Nathaniel Whittemore · 5 Sep 2026
By the numbers

Anthropic expects 85% fewer refusals from Claude Opus 5 compared to Fable 5, attributing the difference to a distinct set of guardrails on Opus.

“As a result, Anthropic is using a different set of guardrails on Opus than they do on Fable 5, which they believe will lead to 85% fewer refusals.”
Nathaniel Whittemore · 28 Jul 2026
By the numbers

Claude Opus 5 scored 43.3% on Frontier Bench, roughly 10 points above Fable 5 and 9 points above GPT-5-6-Soul on the same benchmark.

“It scored a 43.3% on that one that I just mentioned, Frontier Bench, which is a more difficult version of Terminal Bench, which is about 10 points higher than Fable 5 and around 9 points higher than GPT-5-6-Soul.”
Nathaniel Whittemore · 28 Jul 2026
By the numbers

Anthropic removed 80% of the system prompt for Claude Opus 5, Fable 5, and Claude Code with zero change to coding benchmark results.

“Anthropic had removed 80% of the system prompt for Opus 5 and Fable 5 and Claude Code. He said this resulted in zero change to their coding benchmarks.”
Nathaniel Whittemore · 28 Jul 2026
Company & tool watch

Claude Opus 5 (Anthropic) is worth watching as the first model to autonomously build its own computer vision pipeline to complete a task when denied image access, a capability no rival including Mythos demonstrated.

“Opus 5 created its own computer vision pipeline to view the image before successfully recreating the part. No other model, including Mythos, was able to complete this task.”
Nathaniel Whittemore · 28 Jul 2026
Contrarian take

Anthropic intentionally did not train Claude Opus 5 on cyber tasks, yet the model still achieved meaningful improvement in finding code vulnerabilities without that training.

“Anthropic pointed out that Opus is intentionally not trained on cyber tasks. The model has still achieved solid improvement on finding vulnerabilities in code, making it similar to Mythos 5 in that aspect.”
Nathaniel Whittemore · 28 Jul 2026
Contrarian take

Claude Opus 5 beats Fable 5 on many benchmarks but is described as nowhere near Fable 5 in practical use, illustrating a widening gap between benchmark rankings and real-world utility.

“Opus 5 is nowhere near Fable in practical use, not even close. Anyone who's used it meaningfully can tell this very quickly after a few tasks. Yet, Opus beats Fable on many.”
Nathaniel Whittemore · 28 Jul 2026
Worth quoting

Nathaniel Whittemore on the practical gap between Claude Opus 5 benchmark scores and real-world performance.

“Opus 5 is nowhere near Fable in practical use, not even close. Anyone who's used it meaningfully can tell this very quickly after a few tasks. Yet, Opus beats Fable on many.”
Nathaniel Whittemore · 28 Jul 2026
Citation Bureau · reference note, compiled from attributed expert discussion. Last updated 2026-09-13.