Toast 1, our first specialised hunt agent, is disposable today. It provides frontier hunt quality, matching aliases outperforming Claude Opus 5 and GPT-5.6 Sol while being up to 10× cheaper and 12× faster. It performs champion pinch Mixedbread Search, but it tin activity pinch immoderate hunt backend.
Today, frontier models are now capable to execute existent knowledge work. They tin reason, analyse, and find accusation successful analyzable archive collections. But they are besides the astir costly models successful the stack. As intelligence is progressively metered, the request for specialised agents capable to lucifer their capabilities astatine a fraction of the costs is greater than ever.
Toast 1 tin tally arsenic a standalone specialized retrieval agent, aliases arsenic 1 of galore subagents your frontier exemplary already knows really to trust on. It afloat takes complete the hunt loop: fixed an first query, it decomposes it into subqueries, gathers evidence, inspects sources, and curates the applicable discourse earlier returning it. This lets your supplier walk its discourse and compute connected the task that requires a generalist, frontier-level model: reasoning, acting, and producing the last answers.
This specialisation of agentic labour results successful considerably cheaper search, but besides successful amended end-to-end results connected galore realistic tasks. We recovered that Toast 1 establishes a caller Pareto frontier crossed agentic workloads crossed costs per task and velocity per task.
Financial Analysis: OfficeQA Pro V2Link to section
OfficeQA Pro V2, released by Databricks, evaluates reply correctness crossed 90 questions successful realistic, analyzable endeavor financial situations.
GPT‑5.6 Sol pinch Toast 1 made disposable arsenic a sub-agent wrong Codex reaches 70% reply correctness astatine astir $1.15 per task: that is the highest people among the systems evaluated by Databricks successful the OfficeQA v2 release, establishing caller state-of-the-art capacity successful some value and efficiency.
By comparison, the erstwhile champion performer, Claude Fable 5 connected Databricks Genie, reaches 60% correctness astatine astir $4 per task, while GPT-5.6 Sol wrong Codex without Toast 1 only reaches 33% correctness.
This betterment stems from reformulating the economics of grounds gathering. Toast 1's specialization allows it to nutrient high-quality, token-efficient grounds packages, leaving ample resources for the reasoning process to scope the last answer.
Legal Agentic Benchmark - Firm KnowledgeLink to section
Harvey LAB's Law Firm Knowledge benchmark seeks to measure really good an supplier tin hunt and usage organization ineligible knowledge astatine large, realistic scales.
Legal work, by nature, is context-heavy. You cannot outargue personification pinch entree to better, much applicable precedents and details. But it is besides noisy: galore situations are akin but alteration by elemental details, making it tricky to cod precocious value grounds packages without galore mendacious positives.
On a randomly selected subset of 33 tasks,1 we recovered that GPT-5.6 Sol's reply value remained changeless crossed hunt methods.
However, expanding hunt value drastically accrued token efficiency: replacing the vanilla agent's filesystem hunt pinch Mixedbread Search trim token usage from 80.6M to 47M astatine an identical task score. Subsequently adding Toast 1 arsenic its dedicated hunt subagent reduced it further to 23M, and allowed it to decorativeness successful half the turns required by vanilla agent.
The preamble of a Mixedbread Search-powered Toast 1 preserved reply quality, while consuming 3.5× less tokens, starring to a costs simplification of complete 60%. Toast 1 frees up the discourse model of frontier models to fto them walk their tokens connected reaching the correct answer.
Demo: Dig Deep Into Dwarkesh's PodcastLink to section
Benchmarks and numbers tin only show 1 portion of the story. To genuinely understand really Toast 1 works, location is nary amended measurement than watching it hunt successful action. At Mixedbread, we really bask Dwarkesh's podcast, and thought being capable to hunt heavy into its transcripts would beryllium fun.
You tin effort it yourself here.
Although it is simply a tin subagent for analyzable tasks, Toast 1 is besides a tin standalone model, trained specifically for heavy search. It represents the adjacent measurement of our co-design attack down our embedding models and Silo: the model, supplier harness, and retrieval primitives are designed to activity together.2
On a assortment of heavy hunt benchmarks, it reaches frontier exemplary performance, opinionated successful the aforesaid convention arsenic GPT-5.6 Sol and comfortably outperforming models specified arsenic Kimi K3 aliases GLM-5.2.
It remains lightweight successful doing so. A modular Toast 1 tally costs astir 0.016−0.016 - 0.023 per query and has an eight-second median latency. Our highest-quality fusion configuration costs astir 0.05−0.05 - 0.07 per query and has an eleven-second median latency. In practice, among the systems successful our information that reached akin performance, Toast 1 was 7–11× cheaper and considerably faster: Frontier-model retrieval agents took betwixt 20 seconds and 4 minutes connected the aforesaid evaluation.
Availability and PricingLink to section
Toast 1 is disposable instantly done the Mixedbread API astatine the discounted motorboat pricing:
- $0.30 per cardinal input tokens
- $0.036 per cardinal cached input tokens (cache writes are free)
- $0.72 per cardinal output tokens
Mixedbread hunt invoked by Toast 1 is priced astatine a typical rate.
With Your Existing Retrieval StackLink to section
Toast 1 was co-designed pinch Mixedbread Search's primitives and will beryllium astatine its strongest capacity pinch it. But we put typical attraction successful ensuring that it remains backend agnostic: it tin tally complete your existing retrieval indexes, and does not require migrating your existing backend. We conducted thorough testing to guarantee that Toast 1 remains competitory pinch the capacity of frontier models successful akin conditions astatine a fraction of the costs and latency, nary matter the provided index.
You tin usage Toast 1 pinch our Chat Completions API and adhd it arsenic a retrieval instrumentality to your existing agentic workflows successful conscionable a fewer minutes. Here is simply a golden harness you tin usage directly.
With Coding AgentsLink to section
Let your coding agents grip the integration pinch npx skills adhd mixedbread-ai/skills. Or usage Toast 1 straight arsenic a subagent pinch our OpenCode integration.
With Your Mixedbread StoresLink to section
Get an API key pinch $5 successful included credits to effort it out.