Introducing Toast 1

Hacker News by 7 min read 503x views
Introducing Toast 1

Share Post

Toast 1, our first specialised hunt agent, is disposable today. It provides frontier hunt quality, matching aliases outperforming Claude Opus 5 and GPT-5.6 Sol while being up to 10× cheaper and 12× faster. It performs champion pinch Mixedbread Search, but it tin activity pinch immoderate hunt backend.

Today, frontier models are now capable to execute existent knowledge work. They tin reason, analyse, and find accusation successful analyzable archive collections. But they are besides the astir costly models successful the stack. As intelligence is progressively metered, the request for specialised agents capable to lucifer their capabilities astatine a fraction of the costs is greater than ever.

Toast 1 tin tally arsenic a standalone specialized retrieval agent, aliases arsenic 1 of galore subagents your frontier exemplary already knows really to trust on. It afloat takes complete the hunt loop: fixed an first query, it decomposes it into subqueries, gathers evidence, inspects sources, and curates the applicable discourse earlier returning it. This lets your supplier walk its discourse and compute connected the task that requires a generalist, frontier-level model: reasoning, acting, and producing the last answers.

Waterfall trace of a Toast 1 agentic search: 16 instrumentality calls crossed 3 rounds answering an employment-rate comparison query successful conscionable complete 5 seconds. Expand the trace, past prime a measurement to spot the sub-query, grep pattern, aliases scheme the supplier produced astatine that point.

This specialisation of agentic labour results successful considerably cheaper search, but besides successful amended end-to-end results connected galore realistic tasks. We recovered that Toast 1 establishes a caller Pareto frontier crossed agentic workloads crossed costs per task and velocity per task.

OfficeQA Pro V2, released by Databricks, evaluates reply correctness crossed 90 questions successful realistic, analyzable endeavor financial situations.

GPT‑5.6 Sol pinch Toast 1 made disposable arsenic a sub-agent wrong Codex reaches 70% reply correctness astatine astir $1.15 per task: that is the highest people among the systems evaluated by Databricks successful the OfficeQA v2 release, establishing caller state-of-the-art capacity successful some value and efficiency.

Scatter crippled of reply correctness versus costs per rollout connected OfficeQA Pro V2, log-scale cost. GPT-5.6 Sol moving successful Codex pinch Toast 1 arsenic a sub-agent reaches 70 percent correctness astatine astir $1.20 per task, supra the erstwhile Pareto frontier from the Databricks evaluation, wherever Claude Fable 5 connected Databricks Genie reaches 60 percent astatine astir $4.Answer correctness vs. costs per rollout connected OfficeQA Pro V2. Genie and harness numbers arsenic reported by Databricks; Codex + Toast 1 runs are ours. Shaded region sits nether the erstwhile Pareto frontier.

By comparison, the erstwhile champion performer, Claude Fable 5 connected Databricks Genie, reaches 60% correctness astatine astir $4 per task, while GPT-5.6 Sol wrong Codex without Toast 1 only reaches 33% correctness.

This betterment stems from reformulating the economics of grounds gathering. Toast 1's specialization allows it to nutrient high-quality, token-efficient grounds packages, leaving ample resources for the reasoning process to scope the last answer.

Harvey LAB's Law Firm Knowledge benchmark seeks to measure really good an supplier tin hunt and usage organization ineligible knowledge astatine large, realistic scales.

Legal work, by nature, is context-heavy. You cannot outargue personification pinch entree to better, much applicable precedents and details. But it is besides noisy: galore situations are akin but alteration by elemental details, making it tricky to cod precocious value grounds packages without galore mendacious positives.

On a randomly selected subset of 33 tasks,1 we recovered that GPT-5.6 Sol's reply value remained changeless crossed hunt methods.

Bar floor plan of full tokens utilized connected the Harvey LAB firm-knowledge benchmark. A vanilla supplier uses 80.6 cardinal tokens astatine 21.7 turns per task. Adding Mixedbread Search cuts that by 42 percent to 47 cardinal tokens astatine 14.6 turns per task. Adding Toast 1 arsenic a subagent cuts it by different 51 percent to 23 cardinal tokens astatine 11.2 turns per task. All 3 configurations scope the identical task people of 55, truthful the extremity consequence is the aforesaid capacity pinch 3.5 times less tokens.Tokens are totals crossed the 33-task benchmark; turns are supplier loop iterations per task. All 3 configurations scope the identical task people of 55.

However, expanding hunt value drastically accrued token efficiency: replacing the vanilla agent's filesystem hunt pinch Mixedbread Search trim token usage from 80.6M to 47M astatine an identical task score. Subsequently adding Toast 1 arsenic its dedicated hunt subagent reduced it further to 23M, and allowed it to decorativeness successful half the turns required by vanilla agent.

The preamble of a Mixedbread Search-powered Toast 1 preserved reply quality, while consuming 3.5× less tokens, starring to a costs simplification of complete 60%. Toast 1 frees up the discourse model of frontier models to fto them walk their tokens connected reaching the correct answer.

Benchmarks and numbers tin only show 1 portion of the story. To genuinely understand really Toast 1 works, location is nary amended measurement than watching it hunt successful action. At Mixedbread, we really bask Dwarkesh's podcast, and thought being capable to hunt heavy into its transcripts would beryllium fun.

You tin effort it yourself here.

Although it is simply a tin subagent for analyzable tasks, Toast 1 is besides a tin standalone model, trained specifically for heavy search. It represents the adjacent measurement of our co-design attack down our embedding models and Silo: the model, supplier harness, and retrieval primitives are designed to activity together.2

Retrieval value versus costs and latency per query connected BrowseComp Plus, OfficeQA Pro, and LongSeal. Toast 1 matches aliases approaches the champion frontier-model sweeps connected each benchmark while costing a fraction per query and answering successful astir 8 to 10 seconds, acold faster than the frontier sweeps.Cost per query astatine database prices pinch punctual caching; latency is p50 per query. Lines show each model's Pareto-efficient reasoning sweep.

On a assortment of heavy hunt benchmarks, it reaches frontier exemplary performance, opinionated successful the aforesaid convention arsenic GPT-5.6 Sol and comfortably outperforming models specified arsenic Kimi K3 aliases GLM-5.2.

It remains lightweight successful doing so. A modular Toast 1 tally costs astir 0.016−0.016 - 0.023 per query and has an eight-second median latency. Our highest-quality fusion configuration costs astir 0.05−0.05 - 0.07 per query and has an eleven-second median latency. In practice, among the systems successful our information that reached akin performance, Toast 1 was 7–11× cheaper and considerably faster: Frontier-model retrieval agents took betwixt 20 seconds and 4 minutes connected the aforesaid evaluation.

Toast 1 is disposable instantly done the Mixedbread API astatine the discounted motorboat pricing:

  • $0.30 per cardinal input tokens
  • $0.036 per cardinal cached input tokens (cache writes are free)
  • $0.72 per cardinal output tokens

Mixedbread hunt invoked by Toast 1 is priced astatine a typical rate.

Toast 1 was co-designed pinch Mixedbread Search's primitives and will beryllium astatine its strongest capacity pinch it. But we put typical attraction successful ensuring that it remains backend agnostic: it tin tally complete your existing retrieval indexes, and does not require migrating your existing backend. We conducted thorough testing to guarantee that Toast 1 remains competitory pinch the capacity of frontier models successful akin conditions astatine a fraction of the costs and latency, nary matter the provided index.

You tin usage Toast 1 pinch our Chat Completions API and adhd it arsenic a retrieval instrumentality to your existing agentic workflows successful conscionable a fewer minutes. Here is simply a golden harness you tin usage directly.

Let your coding agents grip the integration pinch npx skills adhd mixedbread-ai/skills. Or usage Toast 1 straight arsenic a subagent pinch our OpenCode integration.

from mixedbread import Mixedbread client = Mixedbread() results = client.stores.search( store_identifiers=["legal-documents"], query="does the MSA let duty connected a alteration of control?", search_options={ "agentic": True, # alteration Toast 1 }, )

Get an API key pinch $5 successful included credits to effort it out.

Other Article Hacker News
↑
Close Right Ads
Close Left Ads