deepseek/deepseek-chat-v3.1
- accuracy
- 35.6%
- n
- 288
- invented
- 1
- missed
- 600
- weight_0
- 1
- p50 / p95
- 1710 / 4911
- cost
- $0.0525
OpenRouter bench envelope (temperature 0, seed, json_schema name=prescription_rows, strict:false, usage.include). Production interpretPrescription is Base44 InvokeLLM({prompt, response_json_schema}) with model omitted, then sane() drops declined rows. Same RULES string; different envelope.
extract SHA 73860e2c00f5dbf3d74270fe5d0264b7f77b6d0f · ledger docs/benchmarks/2026-09-11-matrix-runs.ndjson · scored 1822 · errors 2 · ledger spend $1.7727 · generated 2026-09-11T17:41:03.446Z
terminal: true success per (model, corpus_id, seed). HTTP 429s and parse_error are error rows, not hits or misses. After that filter n=1822 (not 1824).distance_m or duration_seconds. Cardio lines are still graded as sets/reps (often 1×1).rows[] envelope. google/gemini-2.5-flash ran 1 seed; others 3.google/gemini-2.5-flashgoogle/gemini-3-flash-previewopenai/gpt-5.6-luna · slower peer x-ai/grok-4.5| model | A | B | C | D |
|---|---|---|---|---|
| deepseek/deepseek-chat-v3.1 | 4.4 | 25.3 | 100 | 71.6 |
| google/gemini-2.5-flash | 99.1 | 91.6 | 100 | 77.3 |
| google/gemini-3-flash-preview | 100 | 92.6 | 100 | 93.3 |
| mistralai/mistral-small-3.2-24b-instruct | 7.7 | 23.8 | 100 | 10.7 |
| openai/gpt-4o-mini | 99.7 | 79.3 | 100 | 46.7 |
| openai/gpt-5.6-luna | 99.1 | 90.9 | 100 | 79.6 |
| x-ai/grok-4.5 | 99.4 | 87.4 | 100 | 88 |
llms.txt for a Base44 builder · every context page