ObviousBench

PUBLIC MODEL RELIABILITY · LAB VIEW

DeepSeek

22 public configurations across 11 model families in the 2026-08-13 release. The global cost frontier stays available as context.

GLOBAL REFERENCE RETAINED

Where DeepSeek sits in the complete public field

The lab is highlighted without changing the frontier calculation or zoom. This preserves the comparison frame that a provider-only chart would otherwise hide.

DeepSeek in the global fieldAll 555 public configurations on the same global log-cost and pass3 scale as the homepage frontier. Public field DeepSeekGlobal Pareto

This is a fixed global reference, not a recomputed DeepSeek frontier. Hover a highlighted point for its observed setting, then select it to continue in the interactive cost frontier.

CURATED DETAIL

Decision pages for DeepSeek

These highlighted cards have an indexable evidence page. The rest remain in the full explorer rather than being turned into thin pages.

FULL LAB VIEW

All observed DeepSeek model families

Sorted by highest observed answer pass3, then the lowest cost at that score. Release dates are not inferred because this release data does not encode a consistent vendor release-date field.

DeepSeek V4.1 FlashBest observed 99.3% · $0.02519
DeepSeek V4 Flash 0731Best observed 99.3% · $0.04967
DeepSeek V4 Pro 0813Best observed 99.3% · $0.08797
DeepSeek V4 Pro 0423Best observed 79.9% · $0.04693
DeepSeek V4 Flash 0423Best observed 73.6% · $0.01148
DeepSeek V3.2 ExpBest observed 61.1% · $0.005534
DeepSeek V3.1 TerminusBest observed 58.3% · $0.01981
DeepSeek Chat V3.1Best observed 54.9% · $0.00504
DeepSeek V3.2Best observed 50.7% · $0.003792
DeepSeek Chat V3 0324Best observed 48.6% · $0.006335
DeepSeek ChatBest observed 42.4% · $0.007478