In my regular investigation (behind a paywall), I person been saying for a while that I deliberation the early of AI is not ample connection models (LLM), but mini connection models (SLM) tally connected section desktop computers aliases moreover mobile phones. In May, a squad from Stanford University published research that compared these SLMs pinch the capacity of LLMs tally successful information centres. If their results are true, past we will hardly request immoderate information centres successful the future, and the hyperscalers are wasting hundreds of billions of dollars successful investments.
Seriously, if you are an investor trying to fig retired wherever to put successful the AI hype, you request to publication this insubstantial successful full. But to get you started, fto maine springiness you immoderate highlights.
First, they ran a bid of SLMs (QWEN 3, GEMMA 3, GPT-OSS, GRANITE 4.0) that tin beryllium downloaded connected a section PC and compared their capacity pinch cloud-based state-of-the-art LLMs (ChatGPT 5, Claude Sonnet 4.5, Gemini 2.5 Pro).
They ran these SLMs connected section PCs powered either by an Nvidia spot aliases an Apple M4 chip, arsenic they are readily disposable successful existent high-end desktop computers (the full study was done before Nvidia presented its AI spot for PCs, which will only accelerate the move distant from datacentres to models tally connected desktops).
Then they traced the capacity of these SLMs vs LLM betwixt 2023 and October 2025 connected some chat tasks and reasoning tasks.
The floor plan beneath shows the Win/Tie-ratio for SLMs vs LLMs successful chat requests, which still dress up the immense mostly of requests today. As you tin see, successful each domain, the champion SLM is capable to find the aforesaid aliases amended answers than an LLM successful 90% aliases much of the cases, pinch an mean crossed each domains of 98.6%.
Win/Tie-ratio of SLM vs. LLM successful chat requests

Source: Saad-Falson et al. (2026)
When it comes to reasoning tasks, which are evidently much demanding, SLMs are catching up fast. On average, they supply a amended aliases astatine slightest arsenic bully an reply arsenic LLMs successful 62.5% of the cases.
Win/Tie-ratio of SLM vs. LLM successful reasoning tasks

Source: Saad-Falson et al. (2026)
However, successful existent life, the tasks for SLMs and LLMs are typically a operation of chat requests and reasoning tasks, truthful the 3rd floor plan shows the weighted mean of chat petition capacity and reasoning capacity based connected the wave of tasks successful each domain. As you tin see, connected average, SLMs are arsenic bully if not amended than LLMs successful 81.2% of the cases, pinch the LLMs having a important advantage only successful areas for illustration engineering, life sciences, proscription and machine sciences.
Win/Tie-ratio of SLM vs. LLM successful chat and reasoning tasks

Source: Saad-Falson et al. (2026)
But it’s not conscionable accuracy. SLMs execute this capacity astatine power and compute costs that are betwixt 50% and 85% little than for an LLM, depending connected the SLM and hardware utilized successful the computer.
What is more, SLMs are catching up quickly successful reasoning tasks. The last floor plan shows the capacity of SLMs arsenic a usability of trouble level and exemplary procreation for reasoning tasks alone.
In 2023, the occurrence complaint of SLMs successful reasoning tasks was typically 50% aliases truthful crossed each 5 trouble levels. By October 2025, the SLMs achieved 99% occurrence for the easiest reasoning tasks successful levels 1 and 2, 85% to 92% occurrence successful harder tasks (levels 3 and 4) and only lagged LLMs successful the hardest tasks of level 5 (51.5% occurrence rate).
Success rates of SLMs successful reasoning tasks

Source: Saad-Falson et al. (2026)
This already intends that 1 tin switch information centres and their costly cutting-edge semiconductor infrastructure successful 4 retired of 5 usage cases. The investigation study estimates that the addressable marketplace successful the US for SLMs has grown to astir $10tn aliases one-third of the full US GDP of $30tn. There isn’t overmuch near for LLMs to thrive in, and each year, their advantage complete SLMs is shrinking.
There intelligibly are areas wherever LLMs are still measurement ahead, peculiarly successful agentic AI applications, wherever SLMs presently only execute accuracy and occurrence rates of little than 50%. Similarly, it is difficult to tally these SLMs connected smartphones truthful far. The models that tin beryllium tally connected an iPhone are importantly worse than the models that tin beryllium tally connected a desktop PC.
But – and this is important – the models tally connected a desktop PC, and moreover much so, the ones tally connected a smartphone are overmuch much power businesslike than the ones tally successful the cloud. The conclusion per Watt of these SLMs is typically 7 times larger than that of LLMs. And that intends that erstwhile you brushwood a task that tin beryllium solved connected a desktop aliases moreover a mobile device, it is cheaper to do truthful locally than nonstop it to a information centre.
This has important implications for investors, successful my view:
We request galore less information centres than we think. If we tin already switch 70% to 80% of the tasks that are expected to tally connected LLMs pinch SLMs, the hyperscalers person simply nary gross maturation successful the early that is astir capable to warrant the capex. In fact, if this investigation is existent and this inclination continues, information centres whitethorn beryllium the worst finance successful the AI abstraction 1 tin make correct now.
While we proceed to request tremendous investments successful semiconductors of each sorts, we do not request to put successful the astir precocious Nvidia chips. The cheaper ones that tally connected desktop PCs will beryllium enough. The champion lawsuit for Nvidia is that it tin switch its high-end information centre GPUs pinch its caller chips for desktop PCs. What will that do to Nvidia’s margins and gross maturation going forward?
We still request LLMs for the astir precocious tasks, and companies for illustration OpenAI, Anthropic and others will beryllium capable to ‘dumb down’ their models to an SLM and waste them alternatively of LLMs. But fixed the already fierce title from Chinese providers for illustration QWEN aliases IBM’s GRANITE, the profit margins for these models will beryllium overmuch smaller than for LLMs. So what does that mean for the valuation of these companies successful their planned IPOs and their maturation trajectory?
While agentic AI is still amended connected LLMs, this whitethorn only beryllium a impermanent advantage, akin to what we person seen successful reasoning tasks and azygous chat requests. If that is the case, the existent winners of the AI roar will not beryllium the providers of precocious hardware and information centres but the boring manufacturers of desktop computers for illustration Dell and Apple.
Watching this title unfold is going to beryllium fun, and I americium progressively convinced that galore group will beryllium severely burned because they put successful the incorrect technology.
English (US) ·
Indonesian (ID) ·