ObviousBench

PUBLIC MODEL RELIABILITY · LAB VIEW

Meta

14 public configurations across 10 model families in the 2026-07-25 release. The global cost frontier stays available as context.

GLOBAL REFERENCE RETAINED

Where Meta sits in the complete public field

The lab is highlighted without changing the frontier calculation or zoom. This preserves the comparison frame that a provider-only chart would otherwise hide.

Meta in the global fieldAll 432 public configurations on the same global log-cost and pass3 scale as the homepage frontier. Public field MetaGlobal Pareto
Meta configurations highlighted in the global ObviousBench cost and reliability field 100% 95% 90% 80% 60% 40% 20% 10% More expensive Cheaper

This is a fixed global reference, not a recomputed Meta frontier.

CURATED DETAIL

Decision pages for Meta

These highlighted cards have an indexable evidence page. The rest remain in the full explorer rather than being turned into thin pages.

FULL LAB VIEW

All observed Meta model families

Sorted by highest observed answer pass3, then the lowest cost at that score. Release dates are not inferred because this release data does not encode a consistent vendor release-date field.

Muse Spark 1.1Best observed 98.6% · $0.3613
Llama 4 MaverickBest observed 63.9% · $0.003954
Llama 3.3 70B InstructBest observed 56.9% · $0.002586
Llama 3 70B InstructBest observed 56.9% · $0.01207
Llama 4 ScoutBest observed 50.7% · $0.002185
Llama 3.1 70B InstructBest observed 45.1% · $0.009085
Llama 3 8B InstructBest observed 39.6% · $0.0008989
Llama 3.2 1B InstructBest observed 24.3% · $0.0009379
Llama 3.2 11B Vision InstructBest observed 23.6% · $0.005567
Llama 3.1 8B InstructBest observed 20.8% · $0.0005088