Workout-note model matrix

OpenRouter bench envelope (temperature 0, seed, json_schema name=prescription_rows, strict:false, usage.include). Production interpretPrescription is Base44 InvokeLLM({prompt, response_json_schema}) with model omitted, then sane() drops declined rows. Same RULES string; different envelope.

extract SHA 73860e2c00f5dbf3d74270fe5d0264b7f77b6d0f · ledger docs/benchmarks/2026-09-11-matrix-runs.ndjson · scored 1822 · errors 2 · ledger spend $1.7727 · generated 2026-09-11T17:41:03.446Z

Read this first

Recommendation (rescored after dropping distance_m + parse_error)

openai/gpt-4o-mini

accuracy
80.7%
n
288
invented
10
missed
170
weight_0
1
p50 / p95
1425 / 2827
cost
$0.0417

openai/gpt-5.6-luna

accuracy
92%
n
288
invented
0
missed
75
weight_0
0
p50 / p95
2557 / 5486
cost
$0.0978

x-ai/grok-4.5

accuracy
93%
n
288
invented
0
missed
65
weight_0
0
p50 / p95
6452 / 36697
cost
$1.3254

Per-model × band

modelABCD
deepseek/deepseek-chat-v3.14.425.310071.6
google/gemini-2.5-flash99.191.610077.3
google/gemini-3-flash-preview10092.610093.3
mistralai/mistral-small-3.2-24b-instruct7.723.810010.7
openai/gpt-4o-mini99.779.310046.7
openai/gpt-5.6-luna99.190.910079.6
x-ai/grok-4.599.487.410088

llms.txt for a Base44 builder · every context page