500B Tokens Later: Letting AI Agents Decompile a First-Person Shooter

Hacker News by 12 min read 160x views
500B Tokens Later: Letting AI Agents Decompile a First-Person Shooter

Share Post

During the final 3 months, I spent several of my period and tokens decompiling a famous first-person shooter. The goal was not to attain a uncomplicated proof-of-concept state. Instead, we really wanted an accurate, stable and feature-complete recreation of the game.

The avid audience of my blog power have noticed that I had earlier written two posts that have since been removed. Everyone alternatively power now be wondering which equivalent I am talking about. To the two of you I can lone say that company America was current to ruin our fun.

However, that’s fine. This article is not concerning the game, it’s additionally small concerning the procedure of decompilation. It’s additional concerning AI orchestration and how to optimize infrastructure, setup and harness for optimal results.

This project was done alongside the assistance of RektInator, Future, st0rm and another members of the community. A big appreciate you to all of them.

We aimed at an exact decompilation of the equivalent to C++. Besides apparent semantic correctness, we had fairly a few additional requirements: We wanted readable C++ origin that compiles. Given how old the equivalent is, we additionally wanted safety and bug fixes, but additionally portability improvements. It would be nice to run the equivalent on Linux, macOS, in the browser, …

We afterward deferred modernization and portability to concentration entirely on reconstructing the first behavior.

Obviously, the general goal was to study how to efficiently orchestrate autonomous AI agents complete the way of months.

We started alongside Claude Max (20x), afterward added Codex Pro and used the two subscriptions simultaneously. Model choice varied a lot. We had been using Sonnet 5 most of the time, but Opus 5.5, Luna, Sol and Terra were additionally used a lot. More on that later.

Claude agents were operating in Claude Code CLI, Codex agents in Codex CLI. We additionally tried another delegate harnesses, but the choice barely mattered, so we stuck to the defaults.

Progress Tracking

Using GitHub CLI, the agents oversee GitHub issues to track their progress. There is one matter per translation component (.cpp file). Additionally, labels assistance collection and prioritize issues.

Communication

Agents communicate via Discord. All of them have admission to one conduit and can the two article and peruse all messages in there.

Discord allows agent-2-agent communication, as fine as human-2-agent. So another participants can conversation to them, without needing device access.

A GitHub webhook posts CI failures into the shared channel, so agents get notified whenever item broke.

Disassembly & Decompilation

Agents have been using the authoritative ida-mcp by Hex-Rays nearly the complete time. It plant great. It’s excellent stable, it’s headless and supports everything needed for this project. I can lone propose it.

We had 4 agents operating at that time.

3 employee agents decompiling and committing and one reviewer delegate that passively coordinates and reviews commits to emblem bugs.

Agents managed to decompile concerning 80% of the equivalent and apparent advancement was made. The equivalent launched, the chief list was apparent and we were capable to burden maps.

We spent a lot of those 4 weeks optimizing our setup:

We reduced token use by triggering before compactions.
The default compaction threshold is 90% environment fill. We reduced it downward to 42%. Decompilation consists of a lot of volatile information: A function that was decompiled is now irrelevant and could be removed from the context. So before compactions assistance eliminate specified “junk” from the context.

We additionally noticed that agents lean to endure concentration complete time. Even inside the extend of one compaction cycle, agents can drift and endure focus, the additional data the environment holds.

Agents sometimes moved to another function before finishing the former one. They additionally started idling during observing CI, notwithstanding receiving nonaccomplishment notifications on Discord. Occasionally, they closed issues without thoroughly checking whether the activity was really complete.

Working flank by flank in a terminal, you can steer the agents to forestall that, but whenever letting them activity autonomously, steering is not possible.

To forestall that, we wrote a document defining our goal, how agents should work, what they must evade and how to grip particular situations.

An hourly cron job automatically injected a petition for agents to reread this document, keeping the instructions caller in their context. While this power not be the ideal method to keep agents focused, it worked really fine through the end of the project.

Lots of period was spent refining the education document. As the satisfied is extremely particular to this project, it makes small awareness to portion it here, though.

However…

… notwithstanding all our efforts optimizing setup and infrastructure, we need to conversation concerning the norm of the work.

Constant advancement (game starting, list rendering, maps loading) led us to accept the norm of the decompilation was great.

It was not. Despite the code being extremely readable, it was semantically wrong.

Agents used incorrect function signatures, types or struct layouts. They invented logic or removed it anywhere deemed unnecessary.

Beyond semantic errors, agents additionally introduced unnecessary architectural changes. As an example, the equivalent has certain configuration variables that it accesses via earth variables. The agents had turned this changeless recollection admission into hash tables alongside a lookup that was orders of dimension additional expensive.

And that’s fair one example of the many things that went wrong.

Why is that?

While a reviewer helps capture bugs, it does not activity fine for item beyond that. Architectural decisions were not questioned, as lengthy as they aligned alongside the goal.

The chief logic for that is that we didn’t have goal acceptance criteria. We never correctly defined “correctness”. Therefore it was difficult for the reviewer to fairness which alter is accurate and what’s wrong.

Obviously it had the equivalent as reference, but stated that we had modernization and portability on the catalog as well, certain deviations were not treated as bugs.
Interestingly, comments in the commits or code led the reviewer to obtain deviations, because of any justification the employee delegate had written down. The workers’ comments efficiently acted as unintentional immediate injection: the reviewer accepted their justifications alternatively of independently checking the deviations against the original.

What we needed was an automated inspect that tells agents whether a reconstructed function matches the original. It should verify identical semantics. A uncomplicated PASS or FAIL indication would be adequate and the delegate can fig out on its own what’s wrong.

Byte Matching Decompilation

The simplest way to accomplish this was byte matching decompilation.

We switched to the compiler used to build the first equivalent and wrote a manuscript that performs the comparison.

The manuscript says our reconstructed OBJ document and the equivalent EXE/PDB (having the PDB is awesome and makes things slightly simpler, but the procedure can activity fair as fine without a PDB).

It afterward extracts the function data from OBJ and EXE and compares all bytes. If they match, the function is exact, alternatively it fails and the delegate needs to rework the function.

References to another functions or data won’t necessarily equivalent byte for byte, since their encoded values depend on anywhere the targets end up in the compiled binary.

Luckily, the OBJ document records these references as relocations. We can exclude the relocation bytes from the straightforward difference and alternatively verify that the two versions citation the identical sign alongside the identical offset.

The manuscript afterward does the identical for data and types.

Agents can afterward use this manuscript to verify their work, before pushing it.

Reconstructed functions are recorded in a set of content files. CI can afterward use these content records to verify all recorded functions and vigilant in case of regressions.

Cheating

The archetypal item agents did whenever we introduced this manuscript was compose inline assembly.

This evidently defeats the purpose. So we had to refine our instructions to disallow certain constructs. Naked functions, entity patching, inline gathering and embedding bytes in the code were forbidden verbally. Given how uncomplicated it is to scan for those constructs, verbal rules were enough.

However, agents often tried to modify this manuscript to exclude their function from comparison.

To forestall them from doing this, CI hashes the verification manuscript and compares it against a stored GitHub Actions secret.

Trade-Offs

Cons

Functions can be difficult to match. Register selection, inlining decisions and calling conventions can be difficult to reproduce exactly. In rare cases, we additionally observed distinct compiler output from identical inputs. Agents now need much longer for decompilation, without necessarily producing improved results. A function can have identical semantics, notwithstanding showing certain divergences, e.g. autonomous instructions that are shuffled.

Pros

Matching functions are guaranteed to have identical semantics. This preserves the first behavior, including any existing bugs, and prevents reconstruction errors in those functions. On top of this, the reviewer delegate is not needed anymore.

Another advantage is that cheaper, small capable models can now reliably activity on the task. Previously models akin Haiku or Luna were a bad fit and produced extremely bad results. However, stated this strict acceptance criterion, they now have adequate feedback to create amazing results, reducing the expenses drastically and allowing the project to measure up massively.

With our new verification harness, agents have been operating for nearly 2 additional months now. 99% of the game’s functions are current in our reconstructed source, alongside 83% of all functions being byte exact.

For the most part, we have been using 14 Luna and 2 Opus 5.5 agents throughout the final weeks.

At that scale, workers used distinct branches and submitted their changes through drag requests.

This is additionally anywhere Discord stopped scaling well. So many agents spamming the conduit is nonsense. We restricted messages to which issues they were taking and CI coordination. That kept communication to a minimum. However, the additional we got into the project, the small we needed to conversation to the agents. Human-2-agent communication was no longer needed as they worked completely autonomously. For another projects at that scale, I would apt choose item another than Discord.

At this point, we’re hitting diminishing returns. The remaining functions mostly have certain non-deterministic characteristics or can’t be matched because of another circumstances, e.g. identical COMDAT folding in the linker that we cannot reliably reproduce.

The Opus 5.5 agents are motionless capable to move onward and get the remaining functions matched. However, the equivalent runs flawlessly now. There are no noticeable bugs and all features of the first equivalent are present.

The remaining functions have been reworked repeatedly. While they motionless don’t equivalent byte for byte, we accept their semantics are correct.

Continuing the byte matching procedure for them would consume additional tokens without meaningfully improving the result. That means, the project can be considered done now.

This project has taught us a lot. Here are the most precious takeaways at a glance:

  • Precise instructions are necessary. Agents have the desire to cheat if the project leaves area for interpretation.

  • Correctness have to be defined and machine-checkable. Humans are notoriously incapable of exactly articulating their intent. Therefore, I don’t think reviewers volition always be enough. A dependable verification harness that delivers an goal PASS or FAIL indication is the finest feedback an delegate can get. Obviously not all project has the luxury decompilation has. Yet, I think for all project it is imaginable to get near to that, alongside adequate creativity.

  • Instructions decay complete time. Agents ignore or treat certain rules as small important, the additional period passes and the additional environment gets compacted. During interactive work, you can accurate that drift as it happens. When agents activity autonomously, it can go unnoticed and undermine the norm of their work. The hourly refresh of the instructions solved that.

  • Generating code is cheap now. Throw it distant if it’s bad. After the archetypal 4 weeks, whenever we noticed the code was bad, we tried to salvage it. This did in fact disbursal us additional period than starting from scratch.

  • Correctness is so much additional crucial than productivity. Reducing token consumption, scaling up agents, optimizing delegate throughput, etc. is awesome and all, but if the outcome is bad, it doesn’t assistance much.

This was honestly specified a precious project and I learned so much. I don’t think penning downward these takeaways can communicate the sheer amount of lessons we’ve learned throughout the process.

So many things went wrong, yet so much went right. As AI agents obtain on additional application betterment work, orchestrating them becomes a new, much additional demanding project for humans.

Scaling up to 15+ agents revealed equal additional difficulties. Agents periodically wiped the VM since of malformed commands. There are already sandboxing solutions for Windows, yet none of what I have seen currently fits my needs. I have started extending Sogen, my userspace emulator, to provision lightweight and scalable sandboxing capabilities. However, it volition obtain a during before this becomes equal remotely manufacturing ready.

Given that the VMs were wiped, we unfortunately misplaced a bunch of meeting logs. So I don’t cognize the exact figure of tokens that were spent. My evaluation is item between 600-700 milliard tokens.

For apparent reasons, none of the code volition be shared. Everything volition remain personal and I volition use it for my own needs only.

Other Article Hacker News
↑
Close Right Ads
Close Left Ads