About a milliard Edge Functions run on Netlify all day — Sunweb personalizing pages, Loto-Québec routing traffic on a biscuit check, and hundreds of thousands of another sites doing everything from personalization to routing to auth. All of it runs on a complete JavaScript runtime that scales alongside our customers’ traffic.
This poses a important specialized challenge, as we strive to create the latency as low as possible. To run tens, sometimes hundreds of thousands, of border functions per second, we need to procedure all request, path it correctly, allocate compute capacity, and footwear the two our phase code and the customer’s code. All of that has to happen inside milliseconds.
Over the former multiple months, our squad has rebuilt the infrastructure rearward Edge Functions, operating closely alongside the squad at Unikraft, who wrote concerning the cognition from their side. In the past, requests went out to a hosted implementation service. Today, they run on MicroVMs inner our own border network — approximately 5x faster at the median. That change additionally improves safety and reliability, and opens up additional possibilities for operating complex compute at the edge.
This doesn’t alter how Edge Functions are written or used — URL imports, npm packages, Node built-ins, netlify.toml declarations, local betterment — all of it plant exactly as it did before. It’s now faster and additional resilient. In this part we’d akin to portion additional concerning the new architecture and our learnings on construction a new compute phase that’s capable to assist elevated quantity at a low achievement overhead.
The numbers, first
An border function runs in forefront of a site, on all petition that matches it. The period it takes is period a client spends waiting, so milliseconds current figure for additional than they do nearly anyplace else.
A heated invocation — routing to a compute node, entering a MicroVM, operating the function, producing reply headers — now costs:
- ~5–6ms at median (p50), downward from 25–40ms on our former infrastructure
- 47.4% faster p99 invocations
- 99.998% availability
- 5x faster border function log delivery

A chilly invocation is value stating too. When a petition arrives in a area that no compute node has seen before, it needs to fetch the applicable images before it’s capable to run anything. This happens on concerning 1.2% of invocations and takes concerning 9ms on average.
What happens during petition time
What follows is the way a sole petition takes, in order: it arrives at the border node, it’s turned into a specification, is routed to a compute node, and afterward handed off to a MicroVM that may or may not already be according to whether it’s a chilly or heated invocation.
Request arrives at the border node
Every petition lands on the Netlify border node closest to the client. The node terminates the TLS association and checks the petition way against the Edge Functions’ routes for that deploy.
If nothing matches, the petition carries on to the cache and onto the base as usual. If a path does match, this is the item anywhere the petition used to depart our network. With our old infrastructure it went out complete the internet, ran the border function, and came rear to us to continue on. With the new compute platform, the petition is forwarded to a compute node inside our network.

Creating an Edge Function service
When the compute node receives the petition alongside the device specification and the assistance ID, it archetypal checks to see if a assistance alongside that ID already exists. If it does, it forwards the petition into the assistance for it to dispatch into the MicroVM. A assistance lets us have multiple MicroVMs connected alongside the identical site’s Edge Functions, and allows us to configure parameters for whenever to measure MicroVMs in and out. For example, we configure all assistance to lone authorize a fixed figure of requests to be handled by a MicroVM before we close it down, to evade MicroVMs operating indefinitely. We use the identical parameters to cognize whenever to eagerly footwear up another MicroVM in anticipation of one shutting down.
If a assistance for the site’s border functions doesn’t already be on the compute node, one is created, and we inspect to see if we have all the images in the device specification on disk. If any are missing, they’re fetched from the border node and written to disk. This method method we lone fetch the border function images that are receiving traffic in that region.
The border node writes a spec
Before the petition goes anywhere, the border node writes a specification for the device that volition run a function. The spec names three images: the runtime, our phase image, and the border function image. It additionally sets the CPU, memory, and association limits.
The spec travels alongside the request, on all request. Its hash and site-specific data is computed to rotate into a assistance ID. This allows for isolation, since two deploys alongside distinct code or distinct surroundings variables are distinct services, and they never portion a MicroVM.
This isolation matters most for failures we don’t desire to be possible. A possibly compromised deploy runs in a distinct MicroVM, and equal if it escapes the runtime, it cannot poison another customers or the compute tier itself. V8 isolates, no matter their name, do not provision this flat of isolation.
Choosing a compute node to run an border function
Each area has a collection of compute nodes. The border node picks one for the assistance using rendezvous hashing: the identical assistance lands on the identical node all time, which is what keeps a MicroVM heated and the code already on disk and in cache formerly it’s been read. This stickiness gives us a caching strategy. If we dispersed requests evenly throughout the swarm, we’d end up alongside a higher flat of chilly starts.
It’s crucial to recall that during sending all petition for a function to the identical compute node is the accelerated path, it’s additionally how a hot place forms — anywhere one occupied function competes for resources alongside everything alternatively on that box. A assistance taking a ample portion of a region’s traffic pinned to a sole node volition saturate the node at the disbursal of another services.
We balance this by relaxing the stickiness. Over a certain threshold, we dispersed the assistance throughout a piece of nodes. This lets us assimilate abrupt spikes in traffic from a sole client without affecting another services that hashed to the identical node.
Finally, formerly a node is chosen, it pulls in the function’s code. A compute node that has served the function before already has it. A node seeing it for the archetypal period fetches it formerly and caches it, so lone the archetypal petition pays that cost.

Starting the MicroVM
Each function runs in its own Firecracker MicroVM. These are created in under a millisecond and commencement in concerning 2ms at p99, since the VM starts a stripped-down Linux surroundings fairly than a complete functioning system. The border function’s records are mounted as an uncompressed EROFS depiction and afterward memory-mapped, so the VM says lone the parts of the bundle it really uses alternatively of loading all of it.
When the MicroVM boots up and the JavaScript server starts to hear on a port, we obtain a snapshot of the MicroVM. When the border function isn’t being invoked, the MicroVMs operating it measure to zero alternatively of sitting idle. The next period it’s invoked, we commencement a new MicroVM from that snapshot. The snapshot is memory-mapped, so the VM can commencement executing without waiting for the complete snapshot to be peruse rear into memory.
The lifecycle of the VM — boot, snapshot, restore, and measure to zero — is the activity of Unikraft’s product. We worked closely alongside them throughout the immigration to create certain it holds up under our petition quantity and traffic patterns.
Running and reply handling
After functioning this project for multiple years, we already had learnings we unified to maximize achievement and the capability to debug. At scale, we’ve run into all kinds of issues, from operating out of ports on virtual switches to DNS (we were surprised, but it wasn’t continually DNS).
In this iteration, we made certain compute nodes run local DNS resolvers. We’ve additionally expanded the metrics we collect, record things akin footwear time, period to archetypal harbor open, and period to commencement person code. There are additionally multiple circuit breakers in location to justify immediate rerouting and decommissioning of compute nodes.

That’s the entire path, and on a heated case it adds concerning 6ms. None of it leaves our network, and we’re in authority of the entire petition cycle. Everything complete happens between the petition arriving and the reply going rear out.
Designing resilient compute infrastructure
When you build a scheme akin the one described above, you’re optimizing for two things at once: the end-user cognition and rollout resiliency. We need to be capable to rotate out changes quickly but balance that alongside the capability to rotate rear fair as quickly.
The compute nodes are built from a basis depiction published by Unikraft and instal a set of packages. These nodes are built separately from our border nodes for a brace of reasons: it keeps our border nodes lightweight and fast, it lets us use distinct case types for our compute nodes, and it lets us measure these nodes independently.
A authority aircraft keeps track of which compute nodes be and which are healthy, and the border nodes study it for that list. It’s additionally what drives a deploy — a new fleet comes up alongside the operating one, scales to equivalent it, and takes complete traffic lone formerly it’s healthy.
Building the compute infrastructure has required near collaboration alongside the Unikraft team. Throughout the migration, we’ve worked alongside them on evaluation correctness, handling ample volumes of requests, and construction capabilities particular to our platform.
It’s live
The activity to rebuild our border compute architecture is additional than fair a speed boost. It’s a faster basis we can keep construction on and have additional authority over. The finest part is that it’s already serving your manufacturing traffic today, at the identical pricing, alongside no immigration stage and nothing to alter in any project.
Running the compute ourselves method the ceiling on Edge Functions is ours to raise. Three things this makes tractable that weren’t before:
- npm bundle support, out of beta. npm packages activity in border functions today, in beta, alongside caveats about native binaries and importing records at runtime. A genuine VM alongside a genuine filesystem removes most of the reasons those caveats exist.
- Room to revisit the procedure limits. The documented limits of 50ms of CPU per request, 512MB of memory, and 20MB of compressed code came from the isolate-based implementation model.
- Compute inner our own network. Anything that depends on controlling the network path, alternatively of reaching a third gathering throughout the internet, is now item we can build.
We’re not done here. The limits and coarse edges we couldn’t contact before are the ones we’re operating on now, so remain tuned.