ObviousBench

PUBLIC MODEL RELIABILITY · LAB VIEW

Meta

28 public configurations across 13 model families in the 2026-08-13 release. The global cost frontier stays available as context.

GLOBAL REFERENCE RETAINED

Where Meta sits in the complete public field

The lab is highlighted without changing the frontier calculation or zoom. This preserves the comparison frame that a provider-only chart would otherwise hide.

Meta in the global fieldAll 555 public configurations on the same global log-cost and pass3 scale as the homepage frontier. Public field MetaGlobal Pareto
Meta configurations highlighted in the global ObviousBench cost and reliability field 100% 95% 90% 80% 60% 40% 20% 10% Llama 3 70B Instruct · None / disabled · 56.9% · $0.01207 Llama 3 8B Instruct · None / disabled · 39.6% · $0.0008989 Llama 3.1 70B Instruct · None / disabled · 45.1% · $0.009085 Llama 3.1 8B Instruct · None / disabled · 20.8% · $0.0005088 Llama 3.2 11B Vision Instruct · None / disabled · 23.6% · $0.005567 Llama 3.2 1B Instruct · None / disabled · 24.3% · $0.0009379 Llama 3.3 70B Instruct · None / disabled · 56.9% · $0.002586 Llama 4 Maverick · None / disabled · 63.9% · $0.003954 Llama 4 Scout · None / disabled · 50.7% · $0.002185 Muse Spark 1.3 · Minimal · 97.2% · $0.2146 Muse Glimmer 30B · Low · 95.8% · $0.0828 Muse Spark 1.3 · Low · 99.3% · $0.4899 Muse Glimmer 30B · Medium · 97.2% · $0.1063 Muse Spark 1.3 · Medium · 98.6% · $0.6989 Muse Glimmer 30B · High · 98.6% · $0.1658 Muse Spark 1.3 · High · 99.3% · $0.7571 Muse Glimmer 30B · Extra high · 97.9% · $0.1758 Muse Spark 1.3 · Extra high · 99.3% · $0.774 Muse Spark 1.1 · Minimal · 97.9% · $0.3706 Muse Spark 1.2 · Minimal · 98.6% · $0.2798 Muse Spark 1.1 · Low · 98.6% · $0.3613 Muse Spark 1.2 · Low · 97.2% · $0.5529 Muse Spark 1.1 · Medium · 97.9% · $0.5266 Muse Spark 1.2 · Medium · 97.9% · $0.7482 Muse Spark 1.1 · High · 98.6% · $0.7417 Muse Spark 1.2 · High · 97.9% · $0.7419 Muse Spark 1.1 · Extra high · 98.6% · $0.8084 Muse Spark 1.2 · Extra high · 97.9% · $0.6709 More expensive Cheaper

This is a fixed global reference, not a recomputed Meta frontier. Hover a highlighted point for its observed setting, then select it to continue in the interactive cost frontier.

CURATED DETAIL

Decision pages for Meta

These highlighted cards have an indexable evidence page. The rest remain in the full explorer rather than being turned into thin pages.

FULL LAB VIEW

All observed Meta model families

Sorted by highest observed answer pass3, then the lowest cost at that score. Release dates are not inferred because this release data does not encode a consistent vendor release-date field.

Muse Spark 1.3Best observed 99.3% · $0.4899
Muse Glimmer 30BBest observed 98.6% · $0.1658
Muse Spark 1.2Best observed 98.6% · $0.2798 Muse Spark 1.1Best observed 98.6% · $0.3613
Llama 4 MaverickBest observed 63.9% · $0.003954
Llama 3.3 70B InstructBest observed 56.9% · $0.002586
Llama 3 70B InstructBest observed 56.9% · $0.01207
Llama 4 ScoutBest observed 50.7% · $0.002185
Llama 3.1 70B InstructBest observed 45.1% · $0.009085
Llama 3 8B InstructBest observed 39.6% · $0.0008989
Llama 3.2 1B InstructBest observed 24.3% · $0.0009379
Llama 3.2 11B Vision InstructBest observed 23.6% · $0.005567
Llama 3.1 8B InstructBest observed 20.8% · $0.0005088