Today we're launching Desert Ant Labs, a European frontier AI laboratory building opinionated on-device intelligence. We judge the champion way to businesslike intelligence starts on-device.
We're building small, specialized models for audio, vision, and matter – each exemplary answers successful milliseconds, and costs thing to run, truthful you tin put intelligence successful each merchandise interaction, without being constricted by token costs aliases conclusion speed. Small capable to tally connected a five-year-old phone, accelerated capable to usage connected each framework aliases keystroke, and amended than the API telephone you're already paying for.
The first 18 models are unrecorded coming (12 unchangeable and six successful beta), accessible via 1 SDK for Swift, Kotlin, and JavaScript. One exemplary per task, each built to beryllium the fastest measurement to complete that task connected a device:
- Voz: transcribe 10 minutes of audio successful 2 seconds connected an iPhone – 4.7x faster than Whisper – pinch a commencement and extremity clip connected each word.
- Clear: a 9MB exemplary that tin move a five-minute laptop signaling into workplace value audio successful 1 second.
- Redact: disguise names, addresses, and paper numbers, successful existent time, successful 27 languages, truthful they ne'er scope your servers.
- Tongue: place 84 languages from 3 words, pinch a 2MB model.
Tongue · 2MB 0.933
293MB detector 0.887
And that's conscionable to sanction a few. You tin find afloat specs and benchmarks for the different fourteen, connected desertant.com/models and Hugging Face. Every exemplary is free up to 100k monthly progressive devices. No tokens, nary logins.
Redact · 12MB 88.8
GLiNER-PII · 2.3GB 91.1
Rampart · 14.7MB 61.4
OpenAI select · 3GB 60.2
We're building this successful Europe, wherever "on-device" is the sovereign default. The information ne'er leaves your customer's hands, the characteristic ne'er depends connected personification else's cloud, and what's ne'er been uploaded tin ne'er beryllium compelled.
How we sewage here
For 5 years we've been building our video app, Detail, pinch an on-device first approach. But erstwhile we introduced features for illustration Auto Edit to create short clips, aliases audio enhancement for podcasts, we had to autumn backmost to unreality APIs. And arsenic the popularity of Detail grew, truthful did our infrastructure bills.
Every fewer months I'd hunt for useful on-device models. I'd surf Hugging Face for a exemplary that could find filler words aliases cleanable up a recording. And, each June, we'd get awesome caller devices to build pinch but the manufacture wasn't moving accelerated enough. The instauration was there: the chips, Core ML, the research. What was missing was everything betwixt that instauration and really implementing a characteristic successful your app: a exemplary you could driblet successful and vessel pinch a fewer lines of code.
So, we trained the models ourselves. It turns retired training a exemplary is simply a merchandise creation challenge, and merchandise is what we know. We designed models and section conclusion that hit unreality services connected speed, quality, and cost, and outperform different section and unreality models connected the task itself, astatine a fraction of their size.
We replaced Dolby for better, faster audio enhancement pinch Clear, and made our on-device transcriptions 5x faster pinch Voz. We besides replaced Claude Sonnet pinch Clips, our 284MB exemplary that turns a 10-minute video into a twelve clips successful 5 seconds – 10x faster and utilizing 470x little energy than Sonnet, pinch the aforesaid quality.
iPhone 16 Pro 302x
MacBook Pro (M5) 345x
Voz 319x
Apple SpeechAnalyzer 78x
Whisper large-v3-turbo 50x
Detail 6, which will motorboat pinch iOS 27, replaces each of our unreality APIs pinch our ain models, moving wholly connected the device.
We've each spent the past fewer years building pinch LLMs arsenic if they were conscionable different API. And, amid the hype astir generalist frontier brains, we almost forgot they're not the only option.
Every developer I talk to has a wishlist of on-device models they'd build if costs wasn't a factor, aliases a characteristic they're bleeding tokens connected that they'd happily switch for a section model. A telephone that runs the aforesaid measurement a 100 1000 times a day: cleaning a recording, tagging a photo, pulling a day retired of a sentence, catching a sanction earlier the matter hits your servers. None of these needs a frontier model.
NVIDIA's ain researchers pulled isolated 3 supplier systems and estimated that 40 to 70% of their calls to a ample exemplary could spell to a small, specialized 1 instead.
The compute is already paid for
The manufacture will walk astir $450 billion connected information centers this year. Meanwhile, the world ships more than a billion phones, tablets, and laptops pinch progressively tin chips, perfectly suited to these kinds of tasks. There's much compute disposable successful people's hands than successful each AI information halfway connected earth.
We person an unfair advantage pinch free inference. No per-call cost, truthful a characteristic runs connected each connection alternatively of the ones you tin spend to check. No round-trip, and your customer's information ne'er leaves the device. When conclusion costs nothing, the measurement we build products changes entirely.
Little brains successful each product
To build pinch section models, the developer acquisition has to get a batch better. You request models you tin usage commercially, that hit the alternatives connected your task successful velocity and quality, that you tin driblet into your app pinch a fewer lines of code, and are easy to discover.
Think of the first 100 models arsenic the cerebellum, the small brain. The small encephalon handles the always-on activity – balance, timing, the skills you ne'er deliberation about, truthful the remainder of the encephalon is free to think. That's what we're building first: fast, specialized models for the activity that runs each day, connected the device, for free.
Then comes the cortex, the furniture that decides which exemplary answers. A mini section exemplary first, a bigger 1 erstwhile the occupation requires it, and the unreality only erstwhile the activity has to time off the device. As unfastened investigation advances and instrumentality silicon becomes much capable, the section models grow, and we'll train larger ones ourselves. Frontier intelligence, built from the mini extremity up.
Cloud labs vessel neutral models because per-token pricing needs a neutral model. Every Desert Ant exemplary ships pinch a default we choose, and the levers you request to alteration that default. We optimize the exemplary and the runtime together: connected an iPhone, Clear and Voz tally connected the Neural Engine, and successful the browser, Clear's aforesaid weights tally done WebAssembly.
The SDK
Ready to get started? You tin instrumentality Desert Ant models successful your app pinch our autochthonal Swift, Kotlin, and JavaScript SDK, disposable connected GitHub.
Our docs are written for developers and agents and you tin effort the models connected your Mac pinch the CLI, aliases successful your browser connected Hugging Face.
Building thing cool pinch our models, aliases want to build them pinch us? Get successful touch.
English (US) ·
Indonesian (ID) ·