ObviousBench

PUBLIC MODEL RELIABILITY · LAB VIEW

Anthropic

59 public configurations across 13 model families in the 2026-07-25 release. The global cost frontier stays available as context.

GLOBAL REFERENCE RETAINED

Where Anthropic sits in the complete public field

The lab is highlighted without changing the frontier calculation or zoom. This preserves the comparison frame that a provider-only chart would otherwise hide.

Anthropic in the global fieldAll 432 public configurations on the same global log-cost and pass3 scale as the homepage frontier. Public field AnthropicGlobal Pareto
Anthropic configurations highlighted in the global ObviousBench cost and reliability field 100% 95% 90% 80% 60% 40% 20% 10% More expensive Cheaper

This is a fixed global reference, not a recomputed Anthropic frontier.

CURATED DETAIL

Decision pages for Anthropic

These highlighted cards have an indexable evidence page. The rest remain in the full explorer rather than being turned into thin pages.

FULL LAB VIEW

All observed Anthropic model families

Sorted by highest observed answer pass3, then the lowest cost at that score. Release dates are not inferred because this release data does not encode a consistent vendor release-date field.

Claude Opus 5Best observed 100% · $0.5998
Claude Fable 5Best observed 99.3% · $0.8536
Claude Opus 4.8Best observed 99.3% · $1.004
Claude Opus 4.5Best observed 99.3% · $1.602
Claude Opus 4.6Best observed 97.9% · $0.6526
Claude Sonnet 4.5Best observed 97.2% · $1.074
Claude Sonnet 4.6Best observed 96.5% · $0.5734
Claude Sonnet 5Best observed 96.5% · $0.7646
Claude Sonnet 4Best observed 95.8% · $1.215
Claude Haiku 4.5Best observed 93.1% · $0.4391
Claude Opus 4.7Best observed 91.7% · $0.2859
Claude 3.5 HaikuBest observed 59% · $0.02929
Claude 3 HaikuBest observed 44.4% · $0.008526