Deser: Rethinking Rust Serialization

Hacker News by 16 min read 39x views
Deser: Rethinking Rust Serialization

Share Post

written on September 29, 2026

Serde is an amazing serialization archive for Rust and it has been a huge logic why I felt productive alongside it for years. However already while at Sentry I got fairly disappointed alongside several of the restrictions alongside it but actually replacing Serde is tricky since of the power that it has in the ecosystem. Also since it’s fairly difficult to really do improved without also making several possibly achy compromises.

Here are three examples of Serde border cases that display mediocre interactions of Serde features or unexpected limitations:

A figure that is a map

An internally tagged enum, alongside serde_json‘s arbitrary_precision feature turned on:

#[derive(Deserialize)] #[serde(tag = "type")] enum Shape {  Circle { radius: f64 }, } serde_json::from_str::<Shape>(r#"{"type": "Circle", "radius": 1.5}"#) // error: invalid type: map, expected f64 

Serde’s data example has no location for arbitrary precision numbers, so serde_json uses in-band signalling alongside a map alongside a magic key. The enum has to buffer the fields until it has seen the tag, and the buffer does not cognize concerning the magic key. Because Cargo features are unified, it’s adequate for any crate in your dependency chart to rotate the characteristic on.

Flattening breaks entire figure keys
#[derive(Deserialize)] struct Stats {  scores: HashMap<u32, u32>, } #[derive(Deserialize)] struct Report {  name: String,  #[serde(flatten)]  stats: Stats, } serde_json::from_str::<Report>(r#"{"name": "x", "scores": {"42": 23}}"#) // error: invalid type: cord "42", expected u32 at row 1 pillar 35 

Stats on its own parses {"scores": {"42": 23}} fair fine. JSON keys are always strings, and serde_json lone turns them into integers if the category asks for one. However formerly flatten buffers the value, "42" is fair a string. The error additionally points at the end of the document fairly than at the key.

Adapters do not compose
fn from_hex<'de, D: Deserializer<'de>>(d: D) -> Result<u32, D::Error> { ... } #[derive(Deserialize)] struct Theme {  #[serde(deserialize_with = "from_hex")]  primary: u32,  #[serde(deserialize_with = "from_hex")]  accent: Option<u32>, } //error[E0308]: `?` controller has incompatible types // | // | #[serde(deserialize_with = "from_hex")] // | ^^^^^^^^^^ expected `Option<u32>`, established `u32` // | //help: try wrapping the expression in `Some` // | // | #[serde(deserialize_with = Some("from_hex"))] // | +++++ + 

A function cannot be passed as a category parameter, so there is no way to apply from_hex to the inner of an Option, a Vec or a map. You compose another function for all wrapper, and formerly you have from_opt_hex the site is no longer optional unless you additionally recall to add #[serde(default)].

None of these are bugs that are uncomplicated to fix in Serde. They autumn out of its design, and that scheme is protected by Serde’s stability guarantees.

Back in 2022 I started an test called Deser. It’s a serialization archive for Rust that takes the person cognition of Serde and puts it on top of a entirely distinct architecture inspired by miniserde. I never really completed it and it sat about for a few years. I picked it rear up, and it has now reached a item anywhere I think it’s value looking at. Even fair to motivate others to see if they desire to examine the space.

The Name And Idea

The name is Serde alongside its two halves swapped. Deser is Serde but the another way around. In Serde, a category drives the deserialization process: a Deserialize impl asks the deserializer for the benevolent of value it expects, the format calls back into a visitor. Every nested value is handled by recursion which makes Serde deserialization inherently develop the stack alongside all flat of nesting.

Deser on the another hand turns this about and the format tells the category of the next value and pushes events into a sink. When a descend hits the commencement of a nested value, it doesn’t call into it but hands rear a new descend to a driver, which keeps all province on the heap (in fact, in an arena). On the way out, emitters come back their nested values alternatively of recursing into them.

That additionally method that Deser cannot assistance formats akin protobuf that are not self describing. They are in fact fairly intentionally remaining out of the design entirely. Which is one way to say: if you desire to “fix” Serde, you need to make several another compromises.

Most of the reasons for Deser’s ideas go rear to Sentry Relay, which processes enormous amounts of untrusted JSON. Over the years whenever I was at Sentry we ran into the identical set of problems again and again, and many of them are not really bugs in Serde but consequences of its design. Serde’s stability guarantees average that a lot of them cannot be fixed without breaking all format and all hand written implementation. Most of these problems arrive from three decisions:

  1. One set of traits for all formats. Serde serves the two oneself describing formats (JSON, YAML, TOML, …) and formats anywhere the audience has to cognize the type upfront (postcard, bincode, protobuf, …). That is incredibly useful, but it method that several features lone activity alongside several formats, and you discover out at runtime. In case of Serde it additionally has several odd wrinkles anywhere a derived struct quietly accepts an gathering in location of an entity in JSON for instance.

  2. A fixed data example that loses data whenever buffering. Internally tagged enums, untagged enums and flatten need to buffer values before they cognize what to do alongside them. The buffer can’t clasp everything the format knew, errors endure their location and extensions to the ecosystem depend on in-band signalling to province things specified as arbitrary precision numbers.

  3. Recursion on the call stack. Every flat of nesting uses stack space. Formats defend against this alongside a recursion limit, but the instant you go through a code way that doesn’t have one (writing, energetic values), deeply nested data can obtain downward your process. It additionally method that a deserialization cannot be paused during you delay for additional input.

Many of the corresponding Serde issues have been open for years, and I wrote about abusing Serde before. People have tried different angles on this complete the years. Some went minimal and dropped most features to get accelerated compiles and no recursion. dtolnay’s own miniserde is the finest example of that, and deser’s trait scheme was initially modelled following it. Other recent attempts went for runtime reflection, or for a new data example alongside a concentration on binary formats.

If you desire to peruse up on all of the collected challenges alongside Serde’s design, I keep a lengthy catalog here.

Dethroning Serde

First of all I don’t think it’s apt that one can substitute Serde. The orphan rule entrenches Serde incredibly fine in the ecosystem. But several things are within the attain of a crate author’s control. In case of Deser it’s completeness.

Deser today implements all crucial oneself describing formats from YAML, JSON, TOML, CBOR, JSON5 and the likes, but additionally XML and plist to really near the gap. XML in particular is item Serde has declined to support, and it shows (more on that below). At the extremely smallest format assistance should not be the reason not to use Deser.

The second issue normally is that really solving Serde’s issues comes at a significant disbursal in compile period and/or runtime performance. Deser is no different. While Deser’s compile times are a bit improved than Serde’s, the binary bloat is fairly a bit worse and the runtime achievement is mixed. It’s roughly comparable if you appearance at the numbers but depending on the format structure you are losing considerably from several of the tradeoffs.

That said, it’s now in a province anywhere it’s at smallest in regulation a drop-in replacement anywhere the tradeoffs power activity fine for users.

Deser’s Design

Deser does not try to be considerably distinct than Serde on the surface level. For most uses you get Serialize and Deserialize and afterward start using it alongside your format implementing crate of choice. Most attributes are very similar, although they are taking Rust expressions alternatively of strings.

use deser::{Serialize, Deserialize}; #[derive(Debug, Serialize, Deserialize)] #[deser(rename_all = "camelCase")] pub struct Account {  id: u64,  account_holder: String,  #[deser(default)]  is_deactivated: bool, } let account: Account = deser_json::from_str(json)?; 

The difference in the scheme would rotate into additional apparent if you execute a serializer or deserializer yourself. Instead of visitors that call into each other recursively, deserializing a category creates a sink which receives events that are immediately emitted by the parser, and serializing produces emitters that hand out values. Nested sinks and emitters are handed rear to a driver, which keeps them on the heap. This design, which is entirely taken from miniserde, gives several engaging consequences:

  • No stack overflows. You can arbitrarily nest structures without issues. For untrusted input you set limits alongside a layer, and you choice the number that you are comfortable with, which is autonomous of your stack space.
  • Suspendable. Because the province lives in the driver, a deserialization can be fed input as it arrives. It’s additionally Send, so it can move between threads during you delay on IO which makes it much nicer to use alongside tokio. Formats akin JSON, CBOR and MessagePack can be parsed as a stream if you so desire.
  • An extensible data model. The center data example is small and made of atoms, maps and sequences. For all else, there are expansion values (DateTime, Uuid, etc.) that additionally all transport a fallback for formats that don’t understand them. Unlike Serde this method it does not depend on in-band signalling of objects alongside magic keys to smuggle values through.
  • Lossless buffering. When a value have to be buffered (for case because the tag of an internally tagged enum comes last), Deser records the events together alongside everything the format knew concerning them. Protocol specific extension types or error locations all survive.
  • Layers are a middleware scheme that sit between the format and your types and can track things akin paths, enforce safety limits, rename keys or redact values without having to contact particular code paths.
  • Native flattening that doesn’t buffer at all.

On top of that are a lot of things that I fair wanted to have:

  • Allow enum tags to be of any type, not fair strings
  • Adapters that create (as = Option<Vec<DisplayFromStr>>)
  • Enabling validation as an adapter
  • derive attributes that are genuine Rust expressions alternatively of strings
  • bytes as a center functionality in the data model
  • duplicate keys rejected by default and errors that item at the problem

Here is a small configuration category that shows a few of these together:

use deser::adapters::DisplayFromStr; use deser::de::Recording; use deser::{Deserialize, Serialize}; use deser_encoding::Hex; use deser_validate::{Check, NonEmpty, Range}; use ipnet::IpNet; #[derive(Debug, Serialize, Deserialize)] pub struct Config {  // at smallest one 256-bit key, all written as hex  #[deser(as = Check<NonEmpty, Vec<Hex>>)]  secret_keys: Vec<[u8; 32]>,  // `IpNet` knows nothing concerning deser, but has `FromStr` and `Display`  #[deser(as = Option<Vec<DisplayFromStr>>)]  allowed_networks: Option<Vec<IpNet>>,  listeners: Vec<Listener>, } #[derive(Debug, Serialize, Deserialize)] #[deser(tag = "type", rename_all = "snake_case")] pub enum Listener {  Unix { path: PathBuf },  Tcp {  host: IpAddr,  #[deser(as = Check<Range<1, 65535>>)]  port: u16,  },  // types this type does not cognize are kept and written back  #[deser(other)]  Other(#[deser(tag)] String, Recording), } 

Adapters are types, so Hex can go inner a Vec, and DisplayFromStr inside a Vec inner an Option. Validators are adapters too, so Check<NonEmpty, Vec<Hex>> decodes the keys and afterward checks that there is at least one. The catch-all type keeps the tag and a record of everything else in case person wants to procedure it later.

Errors are item I attention a lot about, so current is what happens whenever a value is wrong:

secret_keys = ["9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08"] allowed_networks = ["10.0.0.0/8", "fd00::/8"] [[listeners]] type = "unix" path = "/run/app.sock" [[listeners]] host = "127.0.0.1" port = 0 type = "tcp" [[listeners]] type = "quic" host = "::1" alpn = ["h3"] 
let config: Config = deser_toml::Deserializer::from_str(input)  .deserialize_with(|driver| driver.push_layer(PathLayer::new()))?; 

Note that current the tag of the internally tagged enum comes final which method that the values have to be buffered until the tag is known. In Serde this is tricky and we would endure the location if we used several tricks to add it. With Deser however, alongside the way tier enabled Deser you anywhere in the construction the problem is:

Unexpected: invalid value: must be between 1 and 65535 at row 10 pillar 8 (path: listeners[1].port) 

Deser Meta Data

Deser really wants to be extensible, and XML is a additional extreme example of the differences between Deser and Serde. Here is an Atom admission that mixes in Dublin Core for the authors:

use chrono::{DateTime, Utc}; use deser::Deserialize; use deser_value::Value; use deser_xml::DeserializerConfig; deser_xml::namespace!(  atom = "http://www.w3.org/2005/Atom",  dc = "http://purl.org/dc/elements/1.1/", ); #[derive(Debug, Deserialize)] struct Entry {  #[deser(rename = atom!("title"))]  title: String,  #[deser(rename = dc!("creator"))]  creators: Vec<String>,  #[deser(rename = atom!("updated"))]  updated: DateTime<Utc>, } // entries we understand, and everything alternatively is kept as it is #[derive(Debug, Deserialize)] #[deser(untagged)] enum Item {  Entry(Entry),  Other(Value), } let item: Item = DeserializerConfig::new()  .resolve_namespaces(true)  .from_str(r#"  <entry xmlns="http://www.w3.org/2005/Atom"  xmlns:d="http://purl.org/dc/elements/1.1/">  <title>Deser</title>  <d:creator>John</d:creator>  <updated>2026-09-29T21:00:00Z</updated>  <d:creator>Jane</d:creator>  </entry>  "#)?; 

XML uses namespaces which method that names need to be matched by their namespace, not by the prefix the document happens to use. Here the document says d: and the category says dc!. atom!("title") is fair the string {http://www.w3.org/2005/Atom}title, which plant since attributes are expressions. The two creators are collected into one Vec equal although there is another component between them, and the content of updated goes direct into a chrono datetime. Because the enum is untagged, the admission have to be buffered before a type is picked, and deser’s buffer keeps the two creators. So the result is an Entry alongside John and Jane.

quick-xml, the most famous XML crate for Serde, drops the prefixes and ignores namespaces entirely, so a <x:title> from several another namespace is happily accepted as the heading of the entry. The divided catalog part although is considerably worse. A plain Entry fails alongside a copy site error for creator, unless you rotate on the overlapped-lists characteristic (which, remember, is a global additive emblem that any crate could set). That characteristic makes quick-xml read ahead to the end of the component and buffer everything in between, without a limit unless you set one.

But the characteristic lone helps whenever quick-xml is hooked up to the struct directly and no buffering is taking place. Wrap the struct in the untagged enum and Serde buffers the admission itself. Read from that buffer, Entry sees creator twice and fails again. The fallback is a map, which keeps lone the last creator, and there is no error. With or without the characteristic you get this:

Other({"creator": {"$text": "Jane"}, "title": {"$text": "Deser"}, ...}) 

Notice how John is gone.

Format particular expansion types specified as TOML datetimes are another case. TOML has them natively, Serde’s data example does not, so the toml crate passes them on as a map alongside a magic key. In Deser a datetime is an expansion value, which formats that cognize it keep and all others compose as a string:

let value: Value = deser_toml::from_str("released = 2026-09-29T21:00:00+02:00")?; deser_json::to_string(&value)?; // {"released":"2026-09-29T21:00:00+02:00"} deser_toml::to_string(&value)?; // released = 2026-09-29T21:00:00+02:00 

The identical alongside serde_json::Value gives you {"released":{"$__toml_private_datetime":"2026-09-29T21:00:00+02:00"}}, and reading the value into a chrono::DateTime fails outright alongside invalid type: map, expected an RFC 3339 formatted date and period string.

The Cost

So now that you cognize Deser is at smallest in theory cool, at what cost?

It is not free. The scheme relies on energetic dispatch and on sinks and emitters that live on the heap, and that has significant runtime overhead. In my own measurements for JSON, Deser says location between 33% faster and 60% slower than serde_json depending on the data. On average it’s concerning 10% slower for reading. Writes are between three times as accelerated and 70% slower and a cleanse on average. For YAML and TOML it’s noticeably faster than the Serde based crates, but that is additional concerning the format implementations than the architecture.

Compile times slightly are better, but not dramatically so. Because it doesn’t monomorphize everything, publish builds of derived code are concerning 2.3 times as fast as alongside Serde and that get a small bit improved in custom for your own code as small recompilation is necessary.

To create Deser’s scheme activity at all, it additionally uses unsafe internally. Most of this is to keep the sequence of borrowed sinks on the heap. I awareness akin this is fine in the days of Miri and agents, but I cognize it makes several folks uneasy.

And well, the biggest disbursal is that it’s fair not Serde.

How Much Is There?

Quite a lot really which power be surprising. In supplement to the core there is assistance for derive.

It supports all flavorts of JSON you can think of: JSON, JSONC, JSON5 and HJSON. (Fun fact here: they are all generated out of one shared parser template) For binary handling it supports CBOR and MessagePack. Additionally it does YAML 1.1 and 1.2, TOML, XML and all three flavors of Apple’s plist as fine as CSV/TSV, urlencoded data and environment variables. For additional insane contraptions you can attach way info or capture location data as fine as assistance for debug printing. You can execute validation as you parse, opt into different binary encodings in supplement to base64, you can bridge to serde or capture dynamic values, transcode between formats or hook it up with tokio.

For records see docs.rs/deser and the code itself is on GitHub alongside many examples.

This admission was tagged rust

copy as / view markdown

Other Article Hacker News
↑
Close Right Ads
Close Left Ads