I spent concerning a day porting 37signals’ Rust Campfire to Kotlin alongside Claude Code. I wanted to see anywhere the JVM would district in DHH’s benchmark table. On the x86 machines I tested, it kept up alongside Rust and frequently pulled ahead.
David Heinemeier Hansson has been rewriting Campfire, 37signals’ open-source conversation app, in one tongue following another, alongside AI agents doing most of the typing. The scoreboard is a sole table: the identical app and the identical pages in Rails, Django, Laravel, Express, Elixir, Go, Rust and C, in requests per second. Today’s version, for the area page:
Rails Django Laravel Express Elixir Go Rust C ----- ------ ------- ------- ------ ------ ------- ------- 4,101 1,507 3,872 42,636 5,350 53,060 106,494 137,524
Source: basecamp/once-campfire-rust README. 16 clients on a Ryzen AI MAX+ 395, four cores per app.
It’s uncomplicated to peruse that array as a ranking of languages: compiled strikes interpreted, and Rust and C attack everything else. I don’t think that’s what it measures.
On October 7 the Rails area leaf in this difference served 230 requests per second. A day afterward it serves 4,101: an 18x jump alongside the language, example and hardware unchanged. What changed was the architecture. The improvements from the fastest ports were carried into all implementation. In DHH’s words:
I let the clankers run for a day ... and afterward ask it to bring all architectural improvements to all implementation.
After another day of delegate work, C passed Rust too. DHH announced: “Good old C is now in the lead!”
These numbers measure how much notice all harbor got at smallest as much as they measure the language. The Rust harbor has had additional of it than any of the others: hundreds of commits, profiling write-ups, a replacement for Thruster, a vendored HTML parser, and a gzip pipeline tuned until compression stopped dominating the profile. The JVM isn’t in the array at all.
So I asked: what happens if a garbage-collected tongue gets the identical attention, ported from the Rust type fairly than from Rails, held to the identical output, and measured on the identical hardware?
once-campfire-kotlin is a harbor of the Rust harbor of Campfire to Kotlin. It runs on the JVM (Java 25) or as a GraalVM native image. It says the identical SQLite repository alongside the identical signed and encrypted cookies, so it could sit in forefront of an existing install.
It isn’t a toy endpoint. Every petition goes through the identical layers the Rust harbor runs:
a forefront server doing the jobs of Thruster (which the Rust harbor replaced) and Rack::Deflater;
the Rails middleware, public/ records and the complete path table;
the petition environment alongside its biscuit jar, meeting and flash;
the controller’s before-actions, including the hourly meeting refresh;
the bounded part cache alongside Rails’ keys;
the layout;
and, in circular two, the finished-response cache from upstream.
It serves one page, the area leaf (GET /rooms/:id), since that’s the header row of DHH’s table. Rich content is included. That meant porting html5ever, the HTML parser the Rust harbor vendors, to Kotlin. The harbor matches the Rust crate’s 658-case corpus byte for byte, and 180,000 additional fuzzed cases against the Rust implementation established no differences.
Under the hood it’s Netty for HTTP, Java’s FFM API for SQLite (compiled from the identical amalgamation and flags as the Rust build, so the two apps run the identical repository engine), and concerning 22,000 lines of Kotlin. In a native image, SQLite is connected statically and called through GraalVM’s C interface.
A benchmark between two ports method nothing if the ports do distinct work. So the Kotlin harbor has one difficult rule: every reply must be byte-identical to the Rust port’s, headers included, in the identical order.
Bodies and ETags are compared for all area in the seed, for the two user-agent branches and the two encodings.
Twenty-odd another requests are checked too, including 304s, HEAD, Turbo Frames, redirects, 404s, fixed records and byte ranges. So are cache misses, cache hits and invalidation.
Under load, the burden generator checks all reply byte for byte against the Rust page, using the shared harness’s exact-body contract.
Across the final four benchmark runs that’s 306 closed-loop measurements and millions of responses, alongside zero invalid responses.
Matching Rust byte for byte meant porting several things you’d never compose by hand. Kotlin’s header map reproduces how Rust’s http::HeaderMap orders headers, including how removing one moves another into its place. A zero Content-Length goes following all another header, since hyper writes it last.
Here’s the area page’s approval inspect in the two ports. Rust:
Kotlin:
The Rust type takes a Scope since another controllers reuse it, and the Kotlin harbor lone serves show, so this isn’t a strictly equal comparison. But most of the difference is the language: no async/await threading, no move closures, no match arms for None. Nullable types and ?: do the identical job in a line. The act itself says nearly the identical in both:
Kotlin:
I favor study the Kotlin version. In these throughput tests, that didn’t appear to arrive alongside a achievement penalty.
The Rust harbor is the activity of DHH and “GPT on behalf of DHH”: additional than 400 commits since delayed September, concerning 64,000 lines of Rust affirmative 161 templates, held to pixel-for-pixel parity alongside Rails by a Playwright harness throughout browsers.
The Kotlin harbor was built alongside Claude Code (Claude Opus 5.5) complete approximately 24 hours of wall-clock time, overnight included:
A spike in the morning: one route, concerning 2,100 lines, to see whether the JVM could get near at all. It was suspiciously fast, but it skipped the example layers, so most of its guide was small work, not a faster runtime. I threw it away.
The example port in the evening: seven agents operating in parallel, component by module, checked against the aureate vectors the Rust harbor generates from the Rails app.
Parity, afterward profiling: decision all byte of difference against Rust, afterward two rounds of profiling (cheaper SQLite content reads, allocation-free header lookups, caching the stylesheet Link header) before any figure counted.
A assessment round. An autonomous assessment established three genuine bugs: pipelined HTTP requests could be answered out of order; a SQLite call that can obstacle for seconds was marked critical in the FFM API, which Java says must never block; and the JDBC fallback couldn’t run writes. All three were fixed alongside regression tests, afterward remeasured.
The reply cache: whenever upstream Rust gained its finished-response cache mid-project, an delegate ported it (about 300 lines affirmative tests) and matched upstream’s output on hits, misses and invalidation.
The token count: concerning 5.3 myriad caller tokens (0.6 myriad generated and 4.7 myriad written to the immediate cache), affirmative concerning 410 myriad cache says (the environment re-read on all turn). The haze invoice for all benchmark run on Hetzner, including a night whenever I forgot to rotate the servers off: concerning €4.65.
I spent a lot of period checking whether a faster outcome meant we’d made item faster or accidentally stopped doing several of the work.
Here’s how I ran the comparison:
Hardware. Hetzner Cloud dedicated-vCPU servers in Hillsboro, AMD EPYC Milan at 2.0 GHz. The app server was either a ccx13, which turns out to be one bodily core alongside its two SMT threads, or a ccx23, two bodily cores. The burden generator was a distinct ccx23 on a personal network measured at 12–17 Gbit/s.
One app at a time, all in Docker alongside presenter networking, on a caller copy of the kernel database.
Rust as shipped: its own manufacturing image, built from its Dockerfile.
Kotlin on eclipse-temurin:25-jre alongside ParallelGC, and as an Oracle GraalVM 25 native depiction alongside -O3 and a PGO overview recorded on the identical code, targeted at x86-64-v3 so it runs on any latest x86 box.
Closed loop: 1 to 4,096 connections, a 20-second warmup, afterward 15 seconds per stage and three reps in alternating order. The tables display medians. CPU per petition comes from all container’s cgroup.
Logging off in the header numbers. Docker’s log controller charges the petition log to CPU exterior the container, which disbursal all app concerning 30% on these boxes. With logging on, the ratios are similar.
Rust gets the advantage of the doubt. The Rust port’s listener asks the kernel for a queue of lone 128 pending connections (more below). That makes it stall intermittently at elevated association counts. I patched it to use the kernel’s somaxconn, as Go does, and the tables use the patched Rust. Below 4,096 connections the place changes throughput by 2% or less; it lone removes stalls.
A speed inspect before all run. One midday run in Ashburn came out 45% slower for all app, on two caller servers in a row. I threw it distant and moved regions fairly than average it in.
This is the like-for-like comparison: Kotlin against the Rust harbor at the revision it was ported from (5eb6d2d), the two rendering the leaf on all petition through the identical layers, alongside the part cache warm.
One core (ccx13). Requests per second, alongside portion of Rust and CPU per request:
Conns Rust Kotlin JVM Kotlin native ----- -------------- -------------------- -------------------- 1 2,379 (268 µs) 2,810 (118%; 183 µs) 2,716 (114%; 189 µs) 16 7,920 (245 µs) 8,850 (112%; 217 µs) 7,647 (97%; 245 µs) 64 8,530 (228 µs) 8,961 (105%; 212 µs) 7,747 (91%; 242 µs) 256 8,788 (222 µs) 8,945 (102%; 213 µs) 7,408 (84%; 249 µs) 1,024 7,196 (271 µs) 8,536 (119%; 225 µs) 6,887 (96%; 265 µs) 4,096 5,906 (329 µs) 7,193 (122%; 270 µs) 6,002 (102%; 307 µs)
Two cores (ccx23). Requests per second, alongside portion of Rust:
Conns Rust Kotlin JVM Kotlin native ----- ------ ------------- ------------- 1 2,484 3,081 (124%) 2,844 (115%) 16 14,464 18,340 (127%) 16,095 (111%) 64 15,122 20,664 (137%) 17,562 (116%) 256 16,407 20,564 (125%) 17,703 (108%) 1,024 13,448 19,613 (146%) 16,759 (125%) 4,096 11,654 17,300 (148%) 12,605 (108%)
[Chart: circular one, requests/s by connections, 1 center and 2 cores, three lines each]
On one center the JVM is onward of Rust at all association count, by 2% to 22%, and the native depiction is inside 84–114%. On two cores the gap widens: the JVM serves 25–48% additional than Rust, and equal the native depiction is 8–25% ahead. Every app kept the box occupied (about 3.8 of 4 threads) during the burden generator stayed below fractional its capacity, so this isn’t a measure artifact. On this hardware, Rust spends additional CPU per request: 250 µs against the JVM’s 184 µs at 64 connections.
Tail latency follows the identical pattern. At 1,024 connections on one core, p99 was 284 ms for Rust, 147 ms for the JVM and 169 ms for the native image.
Then upstream changed underneath me. The Rust harbor gained a cache of finished, authenticated responses, valid until the repository changes (SQLite’s data_version) or 15 seconds pass. That’s anywhere most of the jump to 106,494 came from. A benchmark against yesterday’s Rust would be uncomplicated to dismiss, so the reply cache was ported to Kotlin row for row and benchmarked against upstream HEAD (ac6d655), the identical code rearward DHH’s latest table. Byte-for-byte parity held on misses, hits and invalidation.
One core (ccx13). Requests per second, alongside portion of Rust and CPU per request:
Conns Rust Kotlin JVM Kotlin native ----- --------------- -------------------- --------------------- 1 3,787 (123 µs) 4,411 (116%; 86 µs) 4,264 (113%; 90 µs) 16 18,614 (101 µs) 19,175 (103%; 97 µs) 16,201 (87%; 115 µs) 64 18,634 (102 µs) 21,643 (116%; 88 µs) 16,993 (91%; 110 µs) 256 19,155 (99 µs) 24,049 (126%; 81 µs) 18,431 (96%; 102 µs) 1,024 15,686 (122 µs) 21,377 (136%; 92 µs) 16,419 (105%; 113 µs)
Two cores (ccx23). Requests per second, alongside portion of Rust:
Conns Rust Kotlin JVM Kotlin native ----- ------ ------------- ------------- 1 4,337 4,318 (100%) 3,876 (89%) 16 25,505 28,701 (113%) 24,557 (96%) 64 30,045 38,057 (127%) 36,106 (120%) 256 30,154 35,581 (118%) 34,746 (115%)
[Chart: circular two, requests/s by connections, 1 center and 2 cores]
The cache approximately doubles all implementation’s throughput (1.8x to 2.4x), and the ordering doesn’t change. With rendering cached away, what’s remaining is the HTTP stack, biscuit verification and copying bytes. The JVM does that in small CPU than Rust: 77 µs against 109 µs per petition on two cores.
The two-core array stops at 256 connections since former that the test stopped measuring the apps. At 1,024 cached connections the app server pushes 4–5 Gbit/s of pages, and Hetzner’s virtual network starts dropping packets. In 15 seconds Rust retransmitted 252,000 TCP segments and Kotlin 452,000. Both apps sat fractional idle, waiting out retransmission timeouts. I confirmed that alongside thread dumps, a escape recording, additional employee threads and a second burden generator, and all of it pointed at the network. Kotlin suffers additional from the packet defeat (fewer requests per second), plausibly since Netty flushes bigger bursts, but I haven’t proven that.
Here’s anywhere Rust is better:
Memory. Rust idles at 15–17 MB and stayed under concerning 80 MB until 4,096 connections. The JVM idles at 117–170 MB and runs at 170–500 MB under load. If recollection is your constraint, the JVM is the incorrect choice, but peruse on.
Startup. Rust answers its archetypal petition in concerning 175 ms from docker run, Docker’s own receptacle commencement included. The JVM takes concerning 1.5 s.
Warmup. The JVM needs its JIT to heated up on the exact petition mix. At a fixed 4,000 requests/s, its p99 was 315–389 ms alongside the benchmark’s customary warmup, against 5 ms for Rust. Warming up alongside the identical requests the measure sends brought it downward to 31–80 ms, but that’s motionless additional than Rust’s 2–3 ms. The remainder is most apt refuse collection, although I haven’t measured it.
Apple silicon. In Docker on an M4 Max (two CPUs per app, before the assessment fixes), the JVM was flat alongside Rust (93–105%) and the native depiction trailed at 75–87%. Kotlin’s guide on x86 doesn’t transport complete to all CPU.
This is the part I discover most interesting. The identical Kotlin code compiled onward of period alongside GraalVM:
starts faster than Rust: 165 ms to the archetypal leaf against 178 ms, Docker’s commencement included;
idles at 28–29 MB, twice Rust’s 15–17 MB but a fifth of the JVM’s;
stays about 30–50 MB under load, below Rust’s 70–225 MB at elevated association counts;
runs at 84–125% of Rust’s throughput uncached, and 87–120% cached.
There’s one recollection catch: on two cores at 1,024 uncached connections its heap grew to 555 MB. Like any garbage-collected runtime, it needs a heap cap (-Xmx) in production.
For this endpoint, GraalVM gave me another helpful deployment option: much lesser recollection use than the JVM alongside throughput fairly near to Rust. One codebase can be deployed for highest throughput on the JVM or for a small footprint as a native image.
Rust’s 128-connection hear queue. In 4 of 6 runs at 4,096 connections, inventory Rust stalled. Half the CPU sat idle, throughput cut to 2,100–3,100 requests/s, and the worst petition took 27.5 seconds. The kernel counted 13,465 dropped association attempts during one run, alongside clients retrying on Linux’s 1, 3, 7 and 15 second schedule. The logic is one number: the Rust harbor listens through Tokio and mio, and mio hard-codes a backlog of 128, the norm library’s value. Thruster, the Go server it replaced, uses the kernel’s somaxconn (65,535 on these boxes). It’s a one-line fix, and the benchmark repo carries it as a patch.
Cloud noise is bigger than the differences you’re measuring. A 45% slowdown throughout all app in one area one midday is larger than most of the gaps in this article. Without a speed inspect against a known-good run, I’d have published it.
Logging is costly in Docker. The petition log expenses concerning 30% of throughput on these boxes, additional than any tongue difference here, since Docker’s log controller does the activity exterior the container.
I started this since the JVM was missing from the table. These results create me think it deserves a location there.
They additionally made me additional careful concerning study these tables as tongue rankings. Rails improved 18x without changing languages, and adding a reply cache approximately doubled throughput in my tests. The architecture plainly matters a lot.
One page. The Kotlin harbor serves the area leaf and nothing else. Posting messages (where the repository author dominates), WebSockets, hunt and attachments aren’t measured.
Two small boxes. One and two bodily cores on AMD EPYC Milan. DHH’s array uses four cores of a Ryzen AI MAX+ 395, so the complete numbers aren’t comparable, lone the ratios inside all table.
One burden shape. Closed-loop clients hammering one cached page, affirmative an open-loop run at fixed rates in circular one. In that open-loop run the JVM sustained concerning 7,500 requests/s on one center against Rust’s 6,400 and the native image’s 6,050.
A lab, not production. No TLS, and logging off in the header numbers.
Everything is public:
The Kotlin port, parity scripts and benchmark harness: mysticcoders/once-campfire-kotlin. It includes bench/cloud/cloud.sh (up, provision, run, fetch, down) for Hetzner, and the Rust backlog patch.
The Rust port: basecamp/once-campfire-rust
The shared burden generator and its byte-for-byte contract: basecamp/once-campfire-verification
A complete run on Hetzner takes concerning an hr and expenses concerning a euro. If you discover a error in the method, open an issue.
