Probed, dated, reproducible capability facts — sampled statistics with 95% confidence
intervals, never single-shot verdicts. Ranked by effective cost per 1,000 schema-conformant outputs.
Reproduce any row: npx modelrig-probes run --model <model>. ⚠ badges mark
declared-vs-probed discrepancies.
Coverage: 24 fixtures, mixed-domain (a finance-seeded probe-suite plus a domain-general demo-rig family), in 2 families (20 probe-suite, 4 demo-rig) — schema: 9 general, 7 finance, 1 technical, 1 e-commerce, 1 legal · grounding: 2 finance, 1 technical · caching: 1 general · image: 1 e-commerce. The demo-rig family is authored synthetic / public-domain tasks (example.support_summarize, invoice.extract, docs.qa, content.classify) with deterministic ground truth — strong for capability ranking, and explicitly not a claim about any customer's workload. Per-fixture stats (with
domains and families) are in each published result file; the rates in this table aggregate the whole
schema corpus, and the Corpus column names which families produced each row — a model
probed only by demo-rig fixtures is marked so, never mixed silently. The Hard conf.
column is conformance on the authored hard-tier subset — the same fixtures for every model — where
capable models separate; the aggregate stays the ranking basis. A native
or coached tag beside it names the serving path those strict
fixtures were probed through: a model without probed structured_native is coached via
json_mode, so a coached miss is that path's limit — read "coached, missed", not a flat fail. Every
model still runs every fixture; the tag labels, it never excludes. The probe-kit's headline
ask is fixtures from your domain —
contribute one.
parity-50: 23/28 of the top-by-usage yardstick (source, as of 2026-08-08) are PROBED — 82%. Their coverage is declared; ours is verified — and our gaps are named, not hidden:
| Model | Schema conformance | Hard conf. | Value accuracy | $ / 1K conformant | Grounded | Cache hits | Image input | Samples | Corpus |
|---|---|---|---|---|---|---|---|---|---|
| deepinfra/mistralai/Mistral-Small-3.2-24B-Instruct-2506 ⚠ undeclared_schema_capable | 95% [88%–98%] | 80% (acc 60%, n=25) coached | 89% | $0.049 | 0% | 0% | — | 95 | demo-rig + probe-suite |
| deepinfra/meta-llama/Llama-4-Scout-17B-16E-Instruct ⚠ undeclared_schema_capable ⚠ citations_without_declared_search | 93% [86%–96%] | 72% (acc 100%, n=25) coached | 100% | $0.060 | 100% | 0% | — | 95 | demo-rig + probe-suite |
| deepinfra/google/gemma-4-31B-it ⚠ undeclared_schema_capable | 100% [96%–100%] | 100% (acc 100%, n=25) coached | 98% | $0.082 | — | 0% | — | 95 | demo-rig + probe-suite |
| deepseek/deepseek-chat ⚠ schema_served_via_json_mode ⚠ citations_without_declared_search | 95% [88%–98%] | 80% (acc 80%, n=25) coached | 95% | $0.084 | 100% | 100% | — | 95 | demo-rig + probe-suite |
| deepseek/deepseek-v4-flash ⚠ schema_served_via_json_mode ⚠ citations_without_declared_search | 96% [90%–98%] | 84% (acc 100%, n=25) coached | 100% | $0.093 | 80% | 100% | — | 95 | demo-rig + probe-suite |
| grok/grok-4-1-fast ⚠ declared_search_ungrounded | 100% [89%–100%] | — | 100% | $0.104 | 27% | 100% | — | 30 | probe-suite |
| deepinfra/meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 ⚠ citations_without_declared_search | 83% [74%–89%] | 36% (acc 100%, n=25) coached | 100% | $0.140 | 100% | 0% | — | 95 | demo-rig + probe-suite |
| deepseek/deepseek-reasoner ⚠ schema_served_via_json_mode ⚠ citations_without_declared_search | 95% [88%–98%] | 80% (acc 100%, n=25) coached | 100% | $0.148 | 93% | 100% | — | 95 | demo-rig + probe-suite |
| openai/gpt-5.4-nano | 100% [96%–100%] | 100% (acc 56%, n=25) native | 88% | $0.150 | 53% | 60% | — | 95 | demo-rig + probe-suite |
| deepinfra/Qwen/Qwen3-30B-A3B ⚠ undeclared_schema_capable ⚠ citations_without_declared_search | 97% [91%–99%] | 88% (acc 100%, n=25) coached | 99% | $0.171 | 87% | 0% | — | 95 | demo-rig + probe-suite |
| deepinfra/openai/gpt-oss-120b ⚠ undeclared_schema_capable | 99% [94%–100%] | 96% (acc 100%, n=25) coached | 100% | $0.171 | 0% | 0% | — | 95 | demo-rig + probe-suite |
| gemini/gemini-3.1-flash-lite ⚠ declared_caching_unrealized | 95% [88%–98%] | 80% (acc 100%, n=25) native | 100% | $0.207 | 100% | 0% | — | 95 | demo-rig + probe-suite |
| deepinfra/Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo ⚠ undeclared_schema_capable ⚠ citations_without_declared_search | 95% [88%–98%] | 80% (acc 88%, n=25) coached | 96% | $0.227 | 73% | 80% | — | 95 | demo-rig + probe-suite |
| fireworks/gpt-oss-120b ⚠ schema_served_via_json_mode | 99% [94%–100%] | 96% (acc 100%, n=25) coached | 100% | $0.266 | 20% | 100% | — | 95 | demo-rig + probe-suite |
| deepseek/deepseek-v4-pro ⚠ schema_served_via_json_mode | 99% [94%–100%] | 96% (acc 100%, n=25) coached | 99% | $0.404 | 7% | 100% | — | 95 | demo-rig + probe-suite |
| openai/gpt-5.4-mini | 100% [96%–100%] | 100% (acc 92%, n=25) native | 98% | $0.525 | 60% | 100% | — | 95 | demo-rig + probe-suite |
| grok/grok-4.3 ⚠ declared_search_ungrounded | 100% [89%–100%] | — | 100% | $0.544 | 20% | 100% | — | 30 | probe-suite |
| deepinfra/MiniMaxAI/MiniMax-M3 ⚠ undeclared_schema_capable ⚠ citations_without_declared_search | 99% [94%–100%] | 96% (acc 100%, n=25) coached | 100% | $0.697 | 87% | 100% | — | 95 | demo-rig + probe-suite |
| openai/gpt-5-mini | 100% [96%–100%] | 100% (acc 100%, n=25) native | 100% | $1.047 | 53% | 100% | — | 95 | demo-rig + probe-suite |
| deepinfra/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B ⚠ undeclared_schema_capable | 100% [96%–100%] | 100% (acc 100%, n=25) coached | 100% | $1.187 | 33% | 0% | — | 95 | demo-rig + probe-suite |
| grok/grok-4.5 ⚠ declared_search_ungrounded | 100% [89%–100%] | — | 100% | $1.278 | 27% | 100% | — | 30 | probe-suite |
| gemini/gemini-2.5-flash | 95% [88%–98%] | 80% (acc 100%, n=25) native | 100% | $1.281 | 93% | 40% | — | 95 | demo-rig + probe-suite |
| openai/gpt-5.2 ⚠ declared_search_ungrounded | 100% [96%–100%] | 100% (acc 100%, n=25) native | 100% | $1.572 | 7% | 80% | — | 95 | demo-rig + probe-suite |
| deepinfra/moonshotai/Kimi-K2.7-Code ⚠ undeclared_schema_capable | 97% [91%–99%] | 88% (acc 100%, n=25) coached | 100% | $1.618 | 13% | 100% | — | 95 | demo-rig + probe-suite |
| anthropic/claude-haiku-4-5 ⚠ citations_without_declared_search ⚠ declared_caching_unrealized | 100% [96%–100%] | 100% (acc 48%, n=25) native | 86% | $1.786 | 87% | 0% | — | 95 | demo-rig + probe-suite |
| openai/gpt-5.4 | 100% [96%–100%] | 100% (acc 100%, n=25) native | 100% | $1.806 | 40% | 100% | — | 95 | demo-rig + probe-suite |
| deepinfra/zai-org/GLM-5.2 ⚠ undeclared_schema_capable ⚠ citations_without_declared_search | 98% [93%–99%] | 96% (acc 100%, n=25) coached | 100% | $1.860 | 60% | 100% | — | 95 | demo-rig + probe-suite |
| gemini/gemini-3-flash-preview ⚠ declared_caching_unrealized | 98% [93%–99%] | 92% (acc 100%, n=25) native | 100% | $2.637 | 100% | 0% | — | 95 | demo-rig + probe-suite |
| fireworks/glm-5p2 ⚠ schema_served_via_json_mode | 100% [96%–100%] | 100% (acc 100%, n=25) coached | 100% | $2.763 | 47% | 100% | — | 95 | demo-rig + probe-suite |
| deepinfra/moonshotai/Kimi-K2.6 ⚠ undeclared_schema_capable | 99% [94%–100%] | 96% (acc 100%, n=25) coached | 100% | $3.646 | 0% | 80% | — | 95 | demo-rig + probe-suite |
| anthropic/claude-sonnet-5 ⚠ citations_without_declared_search ⚠ declared_caching_unrealized | 92% [84%–96%] | 68% (acc 84%, n=25) native | 94% | $4.362 | 93% | 0% | — | 95 | demo-rig + probe-suite |
| openai/gpt-5.5 ⚠ declared_search_ungrounded | 100% [96%–100%] | 100% (acc 100%, n=25) native | 100% | $4.684 | 0% | 100% | — | 95 | demo-rig + probe-suite |
| fireworks/kimi-k3 ⚠ undeclared_schema_capable ⚠ citations_without_declared_search | 98% [93%–99%] | 92% (acc 100%, n=25) coached | 100% | $5.098 | 60% | 100% | — | 95 | demo-rig + probe-suite |
| anthropic/claude-opus-5 ⚠ citations_without_declared_search | 99% [94%–100%] | 100% (acc 100%, n=25) native | 97% | $9.705 | 93% | — | — | 95 | demo-rig + probe-suite |
| gemini/gemini-3.1-pro-preview ⚠ declared_caching_unrealized | 100% [96%–100%] | 100% (acc 100%, n=25) native | 100% | $11.558 | 100% | 0% | — | 95 | demo-rig + probe-suite |
| anthropic/claude-fable-5 ⚠ citations_without_declared_search | 92% [84%–96%] | 100% (acc 100%, n=25) native | 94% | $19.980 | 93% | — | — | 95 | demo-rig + probe-suite |
| openai/gpt-5.5-pro ⚠ undeclared_schema_capable ⚠ declared_search_ungrounded ⚠ declared_caching_unrealized | 100% [96%–100%] | 100% (acc 100%, n=25) native | 100% | $52.827 | 0% | 0% | — | 95 | demo-rig + probe-suite |