ModelRig Leaderboard

Probed, dated, reproducible capability facts — sampled statistics with 95% confidence intervals, never single-shot verdicts. Ranked by effective cost per 1,000 schema-conformant outputs. Reproduce any row: npx modelrig-probes run --model <model>. ⚠ badges mark declared-vs-probed discrepancies.

Coverage: 24 fixtures, mixed-domain (a finance-seeded probe-suite plus a domain-general demo-rig family), in 2 families (20 probe-suite, 4 demo-rig) — schema: 9 general, 7 finance, 1 technical, 1 e-commerce, 1 legal · grounding: 2 finance, 1 technical · caching: 1 general · image: 1 e-commerce. The demo-rig family is authored synthetic / public-domain tasks (example.support_summarize, invoice.extract, docs.qa, content.classify) with deterministic ground truth — strong for capability ranking, and explicitly not a claim about any customer's workload. Per-fixture stats (with domains and families) are in each published result file; the rates in this table aggregate the whole schema corpus, and the Corpus column names which families produced each row — a model probed only by demo-rig fixtures is marked so, never mixed silently. The Hard conf. column is conformance on the authored hard-tier subset — the same fixtures for every model — where capable models separate; the aggregate stays the ranking basis. A native or coached tag beside it names the serving path those strict fixtures were probed through: a model without probed structured_native is coached via json_mode, so a coached miss is that path's limit — read "coached, missed", not a flat fail. Every model still runs every fixture; the tag labels, it never excludes. The probe-kit's headline ask is fixtures from your domaincontribute one.

parity-50: 23/28 of the top-by-usage yardstick (source, as of 2026-08-08) are PROBED — 82%. Their coverage is declared; ours is verified — and our gaps are named, not hidden:

Grok 4.2 — not in registry Grok 4.2 Mini — not in registry Qwen3.5 235B — not in registry gpt-oss-20b — not in registry Mistral Large 3 — not in registry
ModelSchema conformanceHard conf.Value accuracy $ / 1K conformantGroundedCache hitsImage inputSamplesCorpus
deepinfra/mistralai/Mistral-Small-3.2-24B-Instruct-2506
⚠ undeclared_schema_capable
95% [88%–98%]80% (acc 60%, n=25) coached89%$0.0490%0%95demo-rig + probe-suite
deepinfra/meta-llama/Llama-4-Scout-17B-16E-Instruct
⚠ undeclared_schema_capable ⚠ citations_without_declared_search
93% [86%–96%]72% (acc 100%, n=25) coached100%$0.060100%0%95demo-rig + probe-suite
deepinfra/google/gemma-4-31B-it
⚠ undeclared_schema_capable
100% [96%–100%]100% (acc 100%, n=25) coached98%$0.0820%95demo-rig + probe-suite
deepseek/deepseek-chat
⚠ schema_served_via_json_mode ⚠ citations_without_declared_search
95% [88%–98%]80% (acc 80%, n=25) coached95%$0.084100%100%95demo-rig + probe-suite
deepseek/deepseek-v4-flash
⚠ schema_served_via_json_mode ⚠ citations_without_declared_search
96% [90%–98%]84% (acc 100%, n=25) coached100%$0.09380%100%95demo-rig + probe-suite
grok/grok-4-1-fast
⚠ declared_search_ungrounded
100% [89%–100%]100%$0.10427%100%30probe-suite
deepinfra/meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8
⚠ citations_without_declared_search
83% [74%–89%]36% (acc 100%, n=25) coached100%$0.140100%0%95demo-rig + probe-suite
deepseek/deepseek-reasoner
⚠ schema_served_via_json_mode ⚠ citations_without_declared_search
95% [88%–98%]80% (acc 100%, n=25) coached100%$0.14893%100%95demo-rig + probe-suite
openai/gpt-5.4-nano100% [96%–100%]100% (acc 56%, n=25) native88%$0.15053%60%95demo-rig + probe-suite
deepinfra/Qwen/Qwen3-30B-A3B
⚠ undeclared_schema_capable ⚠ citations_without_declared_search
97% [91%–99%]88% (acc 100%, n=25) coached99%$0.17187%0%95demo-rig + probe-suite
deepinfra/openai/gpt-oss-120b
⚠ undeclared_schema_capable
99% [94%–100%]96% (acc 100%, n=25) coached100%$0.1710%0%95demo-rig + probe-suite
gemini/gemini-3.1-flash-lite
⚠ declared_caching_unrealized
95% [88%–98%]80% (acc 100%, n=25) native100%$0.207100%0%95demo-rig + probe-suite
deepinfra/Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo
⚠ undeclared_schema_capable ⚠ citations_without_declared_search
95% [88%–98%]80% (acc 88%, n=25) coached96%$0.22773%80%95demo-rig + probe-suite
fireworks/gpt-oss-120b
⚠ schema_served_via_json_mode
99% [94%–100%]96% (acc 100%, n=25) coached100%$0.26620%100%95demo-rig + probe-suite
deepseek/deepseek-v4-pro
⚠ schema_served_via_json_mode
99% [94%–100%]96% (acc 100%, n=25) coached99%$0.4047%100%95demo-rig + probe-suite
openai/gpt-5.4-mini100% [96%–100%]100% (acc 92%, n=25) native98%$0.52560%100%95demo-rig + probe-suite
grok/grok-4.3
⚠ declared_search_ungrounded
100% [89%–100%]100%$0.54420%100%30probe-suite
deepinfra/MiniMaxAI/MiniMax-M3
⚠ undeclared_schema_capable ⚠ citations_without_declared_search
99% [94%–100%]96% (acc 100%, n=25) coached100%$0.69787%100%95demo-rig + probe-suite
openai/gpt-5-mini100% [96%–100%]100% (acc 100%, n=25) native100%$1.04753%100%95demo-rig + probe-suite
deepinfra/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
⚠ undeclared_schema_capable
100% [96%–100%]100% (acc 100%, n=25) coached100%$1.18733%0%95demo-rig + probe-suite
grok/grok-4.5
⚠ declared_search_ungrounded
100% [89%–100%]100%$1.27827%100%30probe-suite
gemini/gemini-2.5-flash95% [88%–98%]80% (acc 100%, n=25) native100%$1.28193%40%95demo-rig + probe-suite
openai/gpt-5.2
⚠ declared_search_ungrounded
100% [96%–100%]100% (acc 100%, n=25) native100%$1.5727%80%95demo-rig + probe-suite
deepinfra/moonshotai/Kimi-K2.7-Code
⚠ undeclared_schema_capable
97% [91%–99%]88% (acc 100%, n=25) coached100%$1.61813%100%95demo-rig + probe-suite
anthropic/claude-haiku-4-5
⚠ citations_without_declared_search ⚠ declared_caching_unrealized
100% [96%–100%]100% (acc 48%, n=25) native86%$1.78687%0%95demo-rig + probe-suite
openai/gpt-5.4100% [96%–100%]100% (acc 100%, n=25) native100%$1.80640%100%95demo-rig + probe-suite
deepinfra/zai-org/GLM-5.2
⚠ undeclared_schema_capable ⚠ citations_without_declared_search
98% [93%–99%]96% (acc 100%, n=25) coached100%$1.86060%100%95demo-rig + probe-suite
gemini/gemini-3-flash-preview
⚠ declared_caching_unrealized
98% [93%–99%]92% (acc 100%, n=25) native100%$2.637100%0%95demo-rig + probe-suite
fireworks/glm-5p2
⚠ schema_served_via_json_mode
100% [96%–100%]100% (acc 100%, n=25) coached100%$2.76347%100%95demo-rig + probe-suite
deepinfra/moonshotai/Kimi-K2.6
⚠ undeclared_schema_capable
99% [94%–100%]96% (acc 100%, n=25) coached100%$3.6460%80%95demo-rig + probe-suite
anthropic/claude-sonnet-5
⚠ citations_without_declared_search ⚠ declared_caching_unrealized
92% [84%–96%]68% (acc 84%, n=25) native94%$4.36293%0%95demo-rig + probe-suite
openai/gpt-5.5
⚠ declared_search_ungrounded
100% [96%–100%]100% (acc 100%, n=25) native100%$4.6840%100%95demo-rig + probe-suite
fireworks/kimi-k3
⚠ undeclared_schema_capable ⚠ citations_without_declared_search
98% [93%–99%]92% (acc 100%, n=25) coached100%$5.09860%100%95demo-rig + probe-suite
anthropic/claude-opus-5
⚠ citations_without_declared_search
99% [94%–100%]100% (acc 100%, n=25) native97%$9.70593%95demo-rig + probe-suite
gemini/gemini-3.1-pro-preview
⚠ declared_caching_unrealized
100% [96%–100%]100% (acc 100%, n=25) native100%$11.558100%0%95demo-rig + probe-suite
anthropic/claude-fable-5
⚠ citations_without_declared_search
92% [84%–96%]100% (acc 100%, n=25) native94%$19.98093%95demo-rig + probe-suite
openai/gpt-5.5-pro
⚠ undeclared_schema_capable ⚠ declared_search_ungrounded ⚠ declared_caching_unrealized
100% [96%–100%]100% (acc 100%, n=25) native100%$52.8270%0%95demo-rig + probe-suite