Gimlet's Series B

Sep 05, 2026 06:09 AM - 2 hours ago 2

Today, we are announcing our $300M Series B raise, led by Andreessen Horowitz and joined by Sapphire Ventures, Menlo Ventures, 645 Ventures, Arm, Eclipse, Emergence, Factory, Hudson River Trading, M12, OnePrime Capital, Prosperity7, QuantumLight, Samsung Ventures, Tiger Global Management, Triatomic, Wing Ventures, and XTX Markets.

Since March, we’ve added billions successful contracted revenue, gigawatts of datacenter pipeline, and are quickly scaling to hundreds of megawatts successful managed capacity.

The Need for Throughput and Speed

Since our Series A conscionable complete 5 months ago, the request for tokens has continued to summation astatine a breakneck pace. Monthly token procreation has accrued 6X successful 12 months, pinch immoderate projections stating different 20X summation by 2030.

To support this demand, the manufacture has embarked connected possibly the astir eager finance task successful modern history, already adjacent to $1T per twelvemonth and expected to scope a cumulative $7T by 2030. Power is quickly becoming the astir captious assets bottleneck. AI information centers consumed astir 18 GW of capacity successful 2025, pinch that number expected to triple by 2030. Supporting this maturation requires caller powerfulness generation, grid capacity, interconnections, and behind-the-meter power production. We’ll yet deed the limits of the resources we tin present astatine scale. Maximizing throughput per kW enables the manufacture to proceed scaling while reducing the effect connected beingness resources.

In summation to earthy throughput, we’re besides seeing an expanding request for very accelerated conclusion tiers and low-latency tokens, driven by:

  • Agentic workloads. Agents execute multi-step conclusion loops, often requiring dozens of sequential exemplary calls and instrumentality interactions. Latency compounds crossed these calls.
  • Larger models. Models expanded from a fewer 100B parameters conscionable a fewer years ago, reaching 3-10T parameters today.
  • Larger discourse windows. In 2023, discourse windows were constricted to astir 100K tokens. Today, discourse windows tin spell to 1M tokens and more.

The manufacture is capable to present precocious throughput aliases debased latency individually. The problem is erstwhile these requirements compound. The throughput-interactivity Pareto floor plan has been wide circulated, and illustrates the tradeoffs that teams make betwixt throughput and latency. The situation is not only to maximize throughput, aliases to maximize speed, but it’s to present accelerated tokens astatine precocious throughput.

Traditional homogeneous conclusion infrastructure limits the personification to a high-throughput, debased interactivity optimum. With heterogeneous disaggregation, we’re capable to execute 3-10X faster capacity for frontier workloads successful our conclusion cloud.

Traditional homogeneous conclusion infrastructure limits the personification to a high-throughput, debased interactivity optimum. With heterogeneous disaggregation, we’re capable to execute 3-10X faster capacity for frontier workloads successful our conclusion cloud.

When we founded Gimlet, we took the pursuing positions:

  1. AI conclusion will go the mostly workload successful package (not conscionable successful AI).
  2. AI conclusion is importantly - arsenic successful aggregate orders of magnitude - little businesslike than it can/should be.
  3. Therefore, AI infrastructure will request to beryllium reimagined and rebuilt from the crushed up, pinch conclusion successful mind, to meet the request of these workloads.

As the manufacture has moved from training to conclusion arsenic the ascendant workload, the fastest measurement to service models has been to repurpose the aforesaid infrastructure primitively built for training (and successful immoderate cases usage cases for illustration crypto mining) for inference. However, conclusion is fundamentally different from training, and moreover wrong conclusion location is simply a important assortment successful the requirements from the underlying hardware.

Fortunately, spot designers person been actively building highly optimized AI accelerators. We now person entree to vastly differentiated architectures, each coming pinch their ain unsocial strengths.

Different spot architectures excel astatine different tasks and connection capacity and ratio tradeoffs.

Different spot architectures excel astatine different tasks and connection performance/efficiency tradeoffs.

At Gimlet Labs, we are building the first multisilicon cloud, built from the crushed up for conclusion performance. Gimlet is built connected heterogeneous hardware to return advantage of different types of accelerators, including GPUs, near-memory compute, dataflow architectures, and CPUs. And our package stack intelligently disaggregates workloads crossed these chips to tally each shape of the conclusion workload connected its astir optimal silicon architecture.

By breaking up the exemplary and moving it crossed different types of accelerators, we are capable to execute 5-10X speedups for the aforesaid powerfulness footprint, aliases akin throughput improvements for the aforesaid latency. This is critical, because powerfulness is the limiting facet successful AI deployments today.

There are aggregate ways of breaking up models and moving them crossed heterogeneous hardware, and different splits connection different capacity characteristics. In our package stack, models are traced and decomposed into different parts and scheduled connected accelerators based connected the workload SLAs and disposable hardware.

We dynamically rebalance the workload based connected the disposable hardware, truthful that compute is not stranded aliases unused. Even if 1 type of hardware is afloat utilized for a fixed task (such arsenic decode), if the workload needs much decode instances, those tin beryllium spun up connected different disposable hardware arsenic needed successful bid to supply the basal capacity.

Inference workloads tin beryllium disaggregated successful aggregate ways, each pinch different capacity tradeoffs.

Inference workloads tin beryllium disaggregated successful aggregate ways. Here we show a fewer communal disaggregation schemes, each of which offers different capacity tradeoffs.

The astir well-known type of disaggregation is prefill/decode, wherever the prefill shape (which is very compute-bound) runs connected 1 group of devices and the decode shape (which is very memory-bandwidth bound) runs connected a different set. However, location are aggregate different techniques arsenic well, specified arsenic speculative-decode disaggregation and attention-ffn disaggregation. Prefill/decode disaggregation offers the lowest latency, whereas the different types supply a speedup complete homogeneous deployments pinch higher throughput. There are besides a fewer caller types that we are moving connected which we will screen successful early method content.

We’re expanding our team. We’d emotion to speak pinch you If you’re willing successful joining a highly method squad of researchers, engineers, and operators, moving crossed the afloat AI infrastructure stack spanning systems, datacenters, compilers, distributed systems, networking, and capacity engineering.

We are building Gimlet astir mini teams, precocious ownership, and a level organization. We want group to enactment adjacent to the work, collaborate directly, and person the state to lick important problems without unnecessary layers aliases processes.

You tin cheque retired our unfastened positions here.

Frontier exemplary sizes, token demand, token speed, cache lengths, and spot architectures are each increasing rapidly. Software arsenic we cognize it is changing, and we are thrilled to activity pinch our customers connected caller merchandise experiences that tin beryllium unlocked pinch a basal measurement alteration successful some full token throughput and exemplary speed.

Today, we are moving pinch frontier labs and different large-scale consumers of inference, but we are ramping quickly to adhd much capacity to a broader audience. If you would for illustration to study much astir getting entree to low-latency inference, please scope retired here. Thank you to our investors again for partnering pinch america connected this journey.

More