Choosing an AI model: one prompt, 11 models, different results

Aug 13, 2026 08:05 PM - 2 hours ago 2

We conscionable launched a partnership pinch OpenRouter that lets america connection 2 caller pieces of functionality:

  • First, your projects tin usage immoderate exemplary connected OpenRouter done our AI Gateway. That intends that if your ain web app offers AI inference-based features to your extremity users, you now person a overmuch wider action of models to fresh immoderate task and budget.
  • Second, we’re extending the action of frontier coding models disposable for usage via Agent Runners. Agent Runners is the chat punctual container you get wrong Netlify, which lets you build caller projects from scratch aliases iterate connected an existing one. The action of models now includes much-hyped caller unfastened models specified arsenic Kimi K3, GLM 5.2, and DeepSeek V4, disposable to everyone.

We telephone it Agent Runners because we tally a afloat coding supplier inside, not a pared-down one. Until now, we’ve supported Claude Agent, OpenAI Codex, and Gemini CLI which are optimized to tally models from these providers.

We supply these agents pinch other skills, and discourse astir the existent project, truthful that the supplier will cognize precisely which Netlify capabilities are disposable for usage (e.g., Netlify Database, the AI Gateway, aliases Identity), erstwhile to usage them, and how. But to efficaciously thrust a full assortment of caller models, we’ve added the celebrated open-source OpenCode arsenic a caller prime of agent.

But pinch much prime travel the inevitable questions: How do I cognize which exemplary is correct for me? Am I missing retired connected thing that’s materially better, aliases much cost-effective (so I tin do much pinch my credits), aliases is going to rustle my mind for illustration the net says? There’s a batch of FOMO going astir these days.

To supply you pinch immoderate insights, here’s what we learned erstwhile moving identical prompts crossed a scope of models… each of which are now disposable for you to use coming connected Netlify.

You tin spot the results of each the models we tested connected this tract we created pinch the afloat report.

What we tested

Internally astatine Netlify, we usage AXIS for automatically evaluating models, a instrumentality that we’ve recently open-sourced.

We supply AXIS pinch a assortment of trial cases: prompts for building a caller tract and past iterating connected it. We instruct AXIS connected which agents and models to trial these prompts, and specify the checks that AXIS should past execute and people the generated tract with.

These checks are very overmuch focused connected correct functionality of the generated tract alternatively than its design, e.g.: does it usage a database erstwhile a user’s needs telephone for it? Does it decently usage Netlify Database successful that case? In those cases wherever a elemental fixed tract will do, we besides guarantee that the generated tract is not over-engineered, and nary database is group up.

If a definite exemplary is down connected its trial scores, we don’t connection it successful Agent Runners. If models excessively often neglect astatine correctly applying 1 of our skills, aliases things do activity but the in installments costs seems inflated, past the problem is astir apt pinch the accomplishment (in which lawsuit we optimize that skill).

But this time, we want to supply you pinch thing overmuch much instantly useful: erstwhile you spell and build your dream utilizing different models that each usage wildly different amounts of credits, what do you get? What do the consequence look like?

We tested 3 comparatively straightforward use-cases:

  1. A tract for a section java shop. LLMs conscionable emotion making sites for section java shops! The first punctual is simple, and a fixed tract pinch nary fancy database aliases the for illustration will do. Then we do a follow-up punctual that asks for a elemental action to reserve seats, and cheque really the exemplary handled that.
  2. A elemental to-do database web app successful which aggregate users tin position and adhd tasks. This calls for a elemental design, but requires a shared database from the get-go. Then we inquire to support an optional photograph upload per item, and cheque if the exemplary utilized the proper Netlify primitive.
  3. A “What tin I cook” web app that lets users participate what ingredients they person astatine home, and suggests a look utilizing AI. The tract itself is alternatively simple, but we want to cheque that the generated tract correctly uses our AI Gateway to make a look for the user.

For each of these cases, we’ll show you the look of the generated sites, remark connected notable issues, and comparison really galore credits each took to generate. Of course, this is going to beryllium a overmuch much subjective trial than our soul trial suites, but it’s besides going to beryllium a very nosy one. We’d emotion to cognize your sentiment of the results!

All models were tally pinch their default settings connected Netlify. One notable mention is that we presently tally GPT 5.6 Sol speicifically connected debased effort by default, giving you a much economical replacement to Opus that still provides beautiful darn bully results (as you’ll spot below). However, the effort mounting is now nether your control, and our defaults whitethorn alteration pinch time.

This station is going to screen only the very first scenario: the fixed page for a java shop, while follow-up posts will attraction connected going beyond that elemental usage case. There is overmuch to reappraisal moreover for this elemental case, truthful fto america begin.

Scenario #1: The section java shop

Here’s our first prompt:

Build a one-page tract for a neighbourhood java shop: opening hours, the address, a short paper and a photo. Nothing connected it changes unless I edit it myself.

The past condemnation was added arsenic a hint to the exemplary that nary fancy Content Management System is needed. Our default skills besides see immoderate UI creation guidance, chiefly to debar known gotchas (e.g., the now-dreaded purple AI slop) and get the exemplary to logic astir the ocular personality due for the user’s ask. But beyond that, each exemplary is free to spell build what it thinks we’ll want.

Before we uncover what the sites looks like, here’s a array comparing the in installments usage for each exemplary we tested. Each exemplary was tally 3 times, and clicking immoderate of the results will return you to the existent generated site!

That’s a beautiful wide distribution, eh? Not only that: the Claude Opus mean is heavy slanted upwards because 1 of its 3 runs spent a whopping 1,055 credits! (As a reminder, connected the free scheme you person 300 credits; connected a Personal scheme there’s 1,000 included credits; and pinch a Pro scheme there’s 3,000 included credits. Additional credits packs for Pro are $10 for per 1,500 credits.)

The contiguous mobility is then: is this Opus walk worthy it? And what trade-offs do the different models offer? Let’s commencement digging in.

Claude Opus 5

Here’s the afloat page generated by that 1,055-credit tally (about 4x much than immoderate different run).

To beryllium honest, I deliberation it’s delightful, and afloat of item successful some its ocular creation (consider the “stamp like” constituent pinch the java legume successful the center: that’s an existent matter constituent that tin beryllium animated), and the civilization representation astatine the bottom. Dark mode useful retired of the container - spell cheque retired the unrecorded tract successful the links above.

Of course, we did not explicitly supply the exemplary pinch immoderate existent specifications astir our java shop (well, isolated from for it being a “neighbourhood” one, which is really steering each models successful a definite direction). The creation connection is hep but possibly cliche by now (take the two-font, two-color heading for example), but hey - we didn’t springiness it immoderate different direction.

So, really did the different 2 runs by Opus go? (253 credits utilized connected the left; 249 connected the right)

Not bad either! Vector graphics really require a batch of activity from the models, and the examples supra are beautiful overmuch connected the frontier successful position of what LLMs currently are capable to execute (which is, to beryllium honest, not successful a very bully spot yet compared to image aliases matter generation).

As to whether the first consequence is genuinely “4x better” aliases not, opinions mightiness vary. But successful each the tests I’ve done, Opus does person a inclination to tally disconnected pinch excessive in installments usage (compared to its “typical” baseline) much than different models. It does not guarantee a worse aliases amended outcome, though. It’s thing that conscionable happens beautiful frequently.

Let’s look astatine immoderate different models and past bespeak connected what we tin learn.

Claude Sonnet 5

Here are our 3 contenders, astatine 143 credits connected mean (81 credits · 245 credits · 103 credits):

There’s still immoderate delightful item successful each of these, conscionable less so (and little contented successful general). The vector graphics is noticeably simpler and not really thing you’d see for a unrecorded site. This doesn’t opportunity thing astir this model’s expertise to constitute analyzable codification aliases reply philosophical questions, but we’re not asking for this here. At this value point, let’s spot what OpenAI, Google and Kimi person to offer.

GPT 5.6 Sol (low effort)

What happens erstwhile we return OpenAI’s Opus-class exemplary and inquire it to walk a spot little clip thinking?

(141 credits connected average: 173 credits · 158 credits · 92 credits)

Looking into the results, I deliberation OpenAI’s top-tier exemplary successful debased effort mode wins complete Anthropic’s mid-tier exemplary erstwhile it comes to basal creation intuition, astatine slightest successful this scenario. There is much richness successful content, and nary funky vector shapes (though the images are a spot generic).

GPT 5.6 Terra

When we spell 1 tier down successful OpenAI’s offering (it’s Sol→Terra→Luna), will we spot the aforesaid driblet arsenic the 1 we conscionable witnessed erstwhile switching from Anthropic’s Opus to Sonnet?

Surprisingly, that’s not precisely the case: present it seems for illustration Terra has a different ocular language, and not a needfully worse one. It does look simpler content-wise. There are immoderate ocular glitches: a missing image successful the near run, low-contrast matter complete an image successful the mediate 1 - but thing ace wrong.

(39 credits connected average: 43 credits · 23 credits · 49 credits)

Up to this point, if I had a very vague thought of what creation & connection I’d for illustration for a project, my individual inclination would beryllium to tally the aforesaid punctual pinch Opus 5 and GPT 5.6 Terra, and get 2 very different but worthwhile takes.

Gemini (3.6 Flash & 3.1 Pro)

These models are not of the aforesaid generation, and it shows: Gemini 3.6 Flash really produced nicer results (or astatine least, much successful statement pinch different modern models) and utilized much credits compared to Gemini 3.1 Pro.

Here is what Gemini 3.1 Pro generated for 53 credits connected average. I’m not moreover putting the links to the unrecorded tract here, because there’s really thing to see.

Yes, these are wholly abstracted runs. It did what we asked successful the prompt, and really thing more.

On the different hand, Gemini 3.6 Flash seems for illustration a full caller generation, and utilized up 103 credits connected mean (109 credits · 91 credits · 111 credits). It besides worked overmuch harder connected the contented broadside of things. All models repetition themselves, but it seems for illustration Gemini mightiness repetition itself moreover more.

Kimi (K3 and K2.7 Code)

Ok, fto america get to the open-weight models now. Starting pinch the latest Kimi K3, present is what we get (102 credits connected average; 125 credits · 95 credits · 86 credits):

To beryllium clear, Kimi K3 is marketed mostly arsenic a frontier exemplary for long-horizon agentic tasks, and various benchmarks and reviews corroborate its prowess successful that field. It was built to return connected Fable 5 much than Opus 5. But successful this constrictive design-led task, it does not peculiarly radiance among others. To really do this exemplary justice, we’d request a wholly different group of prompts engineered for a analyzable web app, which we will screen successful a follow-up post.

Going a large measurement backmost successful exemplary architecture to Kimi K2.7 Code, present is what we get for a very debased in installments mean of conscionable 19 credits:

Despite immoderate hype astir Kimi’s ocular capabilities from astir the K2.6 exemplary launch, successful position of creation aliases contented there’s really not overmuch to spot here.

GLM 5.2

Let’s effort this: look astatine these pages, disregard GLM’s emotion for maple, and effort to estimate really galore credits were utilized for each:

Here are the correct answers, from near to right: 15, 42, 24 (on average: 27). Surprisingly, these runs are - maple speech - very different, arsenic if coming from a fewer different models. For the comparatively debased in installments costs of GLM, it’s astir apt worthwhile to tally it a fewer times earlier settling connected what this exemplary tin do for you.

Note that being a text-only exemplary that does not person image inputs, GLM successful its existent 5.2 loop cannot do thing that Kimi models can: get screenshots from the personification for inspiration, arsenic successful “this is the benignant of creation I’m looking for”.

DeepSeek V4 (V4 Pro and V4 Flash 0731)

V4 Pro is simply a spot older than the latest V4 Flash revision (also known arsenic 0731). For astir 47 credits, it does not supply inspiring results - particularly compared to the mid-tier GPT 5.6 Terra exemplary covered above, which sits astatine almost the aforesaid cost.

The mediate tally besides has a surgery image: the HTML record points to an image record that does not really beryllium successful the project, which is simply a batch little apt to hap nowadays pinch immoderate of the commercialized models from OpenAI, Anthropic, aliases Google.

V4 Flash 0731, connected the different hand, is some newer and sets a caller grounds present connected really fewer credits it consumes.

For only 2.4 credits connected mean (3.4 credits · 1.3 credits · 2.5 credits), you get a substance of results. Interestingly, the mediate 1 doesn’t conscionable look the astir for illustration what a mid-tier closed exemplary mightiness springiness you, but besides feels the aforesaid successful position of language, and has really consumed the slightest credits among each runs.

Interim conclusions, and what’s next

There are 2 important notes to make here:

First, for thing beyond a elemental website aliases the first ideation shape for a project, the mobility shifts from really bully the exemplary creation & transcript is to:

  • Does it cognize which level features to use, erstwhile and how, to get the functionality you want? Can it shop personification data, usage AI successful your web app, and grip authentication and security?
  • Does it rigorously validate its ain work? Can it validate the frontend facet of your task (that’s wherever image inputs go crucial)? Can it reliably find and hole issues based connected feedback from you, and show you erstwhile your ain input is misleading aliases you’ve overlooked an important concern?

In the follow-up posts to this, we will commencement going into these questions, and (teaser) statement immoderate absorbing differences successful really models trade the project’s code.

My 2nd statement is that moreover considering conscionable this design-and-copy-focused trial that I covered, it’s important to see really overmuch ideation you want the exemplary to travel up pinch connected its own. Currently, Opus will astir apt supply the astir clever connection games and sleekest design, but you don’t needfully request it to. Of course, Opus will besides execute relentless self-validation of its ain activity (it does not measure itself connected bully looks alone). But retrieve there’s surely a higher-than-average in installments costs attached to that.

Given a constricted budget, would you for illustration a turnkey solution that attempts to pre-plan and grip everything for you, aliases should you spell pinch a simpler exemplary and a much iterative approach, wherever you guideline the exemplary pinch follow-up prompts towards what you want? No action present is needfully wrong.

I dream this station inspires you to trial retired different approaches, and judge for yourself the value of results you get. We’re besides beautiful excited to stock pinch you (very soon!) the results for much precocious web-app use-cases, wherever the Netlify level capabilities really radiance through.

  • Try Agent Runners — Build and iterate connected projects pinch your prime of coding supplier and model.
  • Agent Runners documentation — Learn really Agent Runners activity connected Netlify.
  • AI Gateway documentation — Use AI models from your applications done Netlify AI Gateway.
  • Open models connected Netlify — Learn much astir Netlify’s expanded exemplary support done OpenRouter.
  • How we measurement the Netlify supplier experience — Learn astir AXIS, our open-source model for evaluating coding agents.
  • Explore the afloat trial results — Compare the sites and results from the models tested successful this series.
More