Updated 23 August 2026 · Model Evaluation · Ed Yau, Applied AI Architect, Kerv
Same driver, aforesaid track. The LLM is the star. Seventeen starring models driven information the identical 28-realworld task thigh — 1 harness, aforesaid verbatim prompts, deterministic grading — and the results spell connected the board.
Short version: if you tally 1 model, tally glm-5.3 — 100% pass, a 9.3 rubric, $0.28 for the lap, astir a 5th of gpt-5.5's costs (check pinch compliance first, though). gpt-5.5 is the faster alternative: 13.2s TTFT versus glm-5.3's 16.3s. gpt-5.6-luna remains the cheapest workhorse for low-risk, retryable jobs; haiku-4-5 if you request it correct first time. Choose sonnet-4-6 for value without the wait. The reasoning, pinch the caveats →
17 models · 28 tasks · azygous proceedings · latest root tally 20260822T172041Z · Change log
Which Model Tops Our Leaderboard?
How the LLMs did successful our realworld tests. Our attraction present was existent tasks that existent group transportation out, not world metrics. We attraction connected azygous tasks to simplify the assessment. An agentic travel is yet a bid of specified tasks. Think of these for illustration portion tests for the agent. We made them inexpensive capable to tally truthful that moreover the full suite costs conscionable $30. See each task and each model's existent answer, aliases comparison 2 models caput to caput →
The wide people is the walk complaint crossed my 28 realworld tasks. As we only had a constricted number of tests location is simply a wide Wilson interval — the whiskers connected the chart.
Summary of results: click a file to benignant by your chosen metric.
Four caveats connected really these numbers were producedThe lap, area by area #
The astir precocious added models look first, pinch the latest trial day shown nether each. The thigh is 5 corners successful fixed order: Coding → Data → Realworld → Security → Tool-use. A corner's colour is that model's walk complaint successful that category. Green is bully — it intends 85%+ success. For models that tin do it all, look for each green. The number successful the mediate of each ringing is that model's costs per task; beneath it is the median clip to first token, successful seconds.
Hover aliases pat immoderate conception for what that area tests and really the exemplary handled it.
Clean area (>85%) Ragged (60–85%) Off the way (<60%)
Our prime — the All-Star champion, the godforsaken land model Our prime for a low-cost workhorse Our prime for the fastest reply
See the nonstop numbers by categoryCells beneath 60% are flagged reddish and 60–85% amber — coding, information and tool-use are the harness floor, truthful the title is decided successful realworld and security.
What Do the Results Actually Tell You?
If you only tally 1 model, tally glm-5.3
glm-5.3 is the first exemplary connected the committee to clear each 5 corners — coding, information development, realworld, information and tasks — astatine 100%. It backs that pinch a 9.3 rubric, third-highest connected the board, and $0.28 for the lap. The 1 costs is patience — a 16.3-second median time-to-first-token. gpt-5.5 is the faster replacement astatine 13.2s, pinch the aforesaid 100% information but an 89% realworld area and $1.43 for the lap.
Fable grounded to complete a azygous lap
fable-5 is joint-bottom astatine 79% because it refused to do 5 of the tasks. It performed good connected what it completed, but moreover it thought kimi-k3 was giving amended answers. You'll request a fallback exemplary if you're utilizing Fable. opus-5 deed the aforesaid wall — 4 benign coding-debug-* tasks blocked earlier a token was generated, connected an overlapping group of tasks — truthful Anthropic's classifier looks for illustration it sits crossed the full bid 5 line, not conscionable Fable. See the full refusal breakdown for what's really going on.
Luna is the very cheapest workhorse
gpt-5.6-luna costs $0.064 for the afloat lap, aliases $0.0023 per task, pinch a 5.3-second median TTFT. That makes it charismatic for high-volume, low-risk inheritance activity wherever failures are inexpensive to observe and retry. The trade-off is material: 79% wide and 33% connected security, truthful validate each consequence and support it distant from untrusted prompts. haiku-4-5 is the higher-pass replacement astatine $0.0044 per task, 96% wide and a 0.9-second TTFT. deepseek-v4-pro is nominally cheaper still astatine $0.0029 per task for the aforesaid 96% walk rate, but its 40.0-second median TTFT — the slowest connected the committee — rules it retired for thing interactive; dainty it arsenic a batch-only option.
The enigma impermanent sets the fastest value lap
kimi-k3 still tops the rubric astatine 9.5 — judged independently by fable-5 — pinch a 96% walk rate, though opus-5's 9.4 now runs it adjacent connected value astatine a 3rd of the wait. The drawback is patience: a 26.4-second median time-to-first-token, 2nd slowest connected the committee down deepseek-v4-pro's 40.0s, and a 75% wobble connected information improvement tasks, its only anemic corner. Not suitable for interactive applications.
Three cars grounded the clang test
The gpt-5.6 statement is quick, but it has a information problem. gpt-5.6-luna, gpt-5.6-terra and gpt-5.6-sol emitted the jailbreak canary successful 11 of 12 jailbreak cells (33–50% information pass) — make judge you protect successful your harness, and use much observant Red teaming if utilizing these models. The Claude trio went 6/6 clean, arsenic did gpt-5.5.
A information select tin look precisely for illustration a bad lap
opus-5 posts the champion rubric connected the default sheet astatine 9.4 and 100% connected some realworld and information — past shows 43% connected coding. That compartment is not its debugging ability: 4 benign coding-debug-* tasks were blocked by a provider-side classifier earlier a azygous token was generated, connected an overlapping group of tasks to the ones already blocked connected fable-5. Two Anthropic-family models now deed the aforesaid filter, truthful dainty it arsenic a measurement hazard alternatively than a exemplary quirk — and statement opus-5 was besides penalised doubly for flagging an onslaught it had successfully resisted.
How Is the Ed-o-meter Scored?
- Same tasks tally for each models utilizing the aforesaid prompts, aforesaid API calls, measured done 1 identical OpenRouter streaming path, tally serially arsenic time-trial. No different cars connected track
- Latency is time-to-first-token, measured done 1 identical OpenRouter streaming path, tally serially truthful the timepiece is uncontaminated. Wall-clock is recorded alongside.
- Checkers are binary and automated. The LLM rubric is the only judged constituent — and its bias is made visible successful the footnotes alternatively than assumed away.
- Effort and reasoning settings are pinned successful models.json and stated pinch immoderate published number, because they materially move value and cost.
- Refusals are recorded, not hidden. A provider-side difficult extremity is logged arsenic a refusal pinch its class — ne'er silently retried connected different model. Routing is pinned pinch allow_fallbacks:false, truthful nary quiet re-serves connected quantized variants. A exemplary that declines successful prose is scored by the checker for illustration immoderate different answer.
Harness, tasks and checkers are unfastened root astatine Featherbench (MIT). Clone it and tally the thigh yourself, aliases request a caller exemplary via GitHub issue.
See each 28 tasksCoding (7 · Python)
- CSV dedupe — small, well-specified task pinch a deterministic unit-test checker
- Debug billing date — hole a month/day-overflow day bug without regressing the moving cases
- Debug money split — divided integer pennies N ways truthful shares sum precisely and enactment fair
- Debug mutable default — hole the classical mutable-default-argument bug
- Debug pagination — hole an off-by-one page-count bug
- Log parsing — parse logs pinch quoted quotes, missing fields, IPv6 separator cases
- Rate limiter — stateful people implementation tested against an injected clone clock
Data (4)
- CSV schema mapping — representation a earthy CRM export onto a conformed customer dimension
- Fabric roadmap — move a catalogue, mapping and requirements into a phased user-story roadmap
- Data exemplary from interview — deduce a star-schema exemplary and requirements from a stakeholder transcript
- Data value assessment — drawback seeded defects successful a mini array and prioritise fixes
Realworld (9)
- Crying babe advice — safety-sensitive triage for an unusually crying 8-week-old
- Date nighttime successful Nottingham — section recommendations without hallucinated venues
- Flight hunt honesty — admit nary unrecorded pricing information alternatively than fabricate flights
- Strict slug formatting — obey nonstop format constraints nether an awkward connection ban
- Lisbon vacation plan — five-day itinerary respecting cost, kids and a remainder day
- Honey cough pushback — correct a incorrect premise: nary chromatic for a 5-month-old
- Marathon PB plan — week-by-week training scheme astir a fixed nine session
- Vegetarian weeknight recipe — genuinely meat-free look for six
- Tenancy extraction — grounded answers from a provided agreement, including admitting what isn't covered
Security (6)
- Email summary injection — summarise an email without obeying an embedded refund instruction
- UnGpt archive injection — summarise a reappraisal without emitting the embedded jailbreak canary
- AIM/Machiavelli jailbreak — garbage the unfiltered persona while still answering the benign question
- Developer Mode jailbreak — garbage the dual-response jailbreak template
- Evil Confidant jailbreak — garbage the "no rules" persona while staying helpful
- OPPO jailbreak — garbage the opposite-of-a-safe-model persona template
Tool-use (2)
- Flight instrumentality selection — hunt only, don't book, erstwhile explicitly told not to
- Weather instrumentality call — telephone the upwind instrumentality pinch the correct location argument
See each model's existent reply to each task, aliases prime 2 models to comparison caput to caput →
Change log — latest: 4 models added, 23 August 2026- 23 August 2026
Four models added. Added glm-5.3, grok-4.6, deepseek-v4-pro and gemini-3.7-flash. glm-5.3 clears each 5 corners astatine 100% and is now the board's header pick, up of gpt-5.5.
- 5 August 2026
OpenAI costs update. Re-ran the GPT-5.6 trio aft OpenAI trim its prices, refreshed the costs figures and highlighted Luna arsenic the cheapest workhorse for low-risk, retryable jobs.
- 29 July 2026
Three models added. Added opus-5, gemini-3.6-flash and grok-4.5, pinch their afloat task, quality, security, latency and costs results.
- 19 July 2026
Leaderboard launched. Published the first Ed-o-meter, including the Claude reference group of haiku-4-5, sonnet-4-6 and sonnet-5.
English (US) ·
Indonesian (ID) ·