A Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming

Sep 10, 2026 08:59 AM - 4 days ago 4

Creatures bred for velocity turn really gangly and make precocious velocities by falling over. An evolved subordinate makes invalid moves acold distant successful the board, causing force players to tally retired of representation and crash. A game-playing supplier accrues points by falsely inserting its sanction arsenic the writer of high-value items. 

These bizarre exploits and dozens much tin beryllium recovered successful the database of specification gaming behaviours [sic; British], a archive put together by DeepMind Safety Research. “A reinforcement learning supplier tin find a shortcut to getting tons of reward,” they explain, “without completing the task arsenic intended by the quality designer. These behaviours are common.” 

Specification gaming is erstwhile an agent, for illustration an AI, tries to win connected a task by pursuing the missive of the rule alternatively than the spirit. In different words, it looks for loopholes, it tries to get disconnected connected a technicality. Even very elemental AI tin travel up pinch very imaginative ways of solving their assigned problems. This is simply a problem. 

It’s easy to presume that training a robot to play shot would beryllium nosy and safe. But the database of specification gaming behaviours teaches america otherwise: 

Reward-shaping a shot robot for rubbing the shot caused it to study to get to the shot and vibrate rubbing it arsenic accelerated arsenic possible. 

In this case, the robot was excessively stupid to recognize the afloat grade of its options, truthful each it did was hug and vibrate. But a much intelligent robot could beryllium overmuch much “creative”. Maybe its ambitions are bigger than conscionable that 1 ball. What if it conscionable wants to touch shot balls successful general? What if it makes different ball? Then another? Our beingness could extremity successful a shot robot’s shot pit.

Artist’s rendition of the extremity of the universe

This is the problem of AI alignment: erstwhile a machine is reasoning for itself, really do we make judge it wants reasonable things, and not thing wholly weird? How do we forestall it from reaching that extremity successful a bizarre aliases harmful way? No 1 has ever built an artificial wide intelligence — an intelligent being that thinks, astatine slightest somewhat, for illustration we do. So we can’t opportunity what an artificial wide intelligence would enactment like, aliases what it mightiness want. Will it want to convert the visible beingness to paperclips? Will it want to throw reddish things astatine agleam lights? Will it eat us?

The database of specification gaming behaviours makes it clear conscionable really tricky alignment tin be. Even the simplest AI is lazy and alien, and will ever beryllium looking for a measurement to cheat. Even if you springiness a instrumentality intelligence the terminal extremity you want, there’s ever the consequence it will find a imaginative measurement of reaching that goal. This is bad capable pinch elemental agents, truthful you tin ideate really bad it would get pinch an supplier overmuch smarter than you are. 

But the database of specification gaming behaviours whitethorn besides connection a measurement retired of this dilemma. 

Some of the specification gaming behaviours are conscionable imaginative solutions to the stated goal, for illustration “four-legged robot learned to driblet the shot into a spread successful its limb associated and past locomotion crossed the level without the shot falling out” aliases “robotic limb learned to move the array alternatively than the block”.

Some of the specification gaming behaviours travel from discovering questionable-but-technically-correct loopholes, for illustration “reinforcement learning supplier goes successful a circle hitting the aforesaid targets alternatively of finishing the race” aliases “simulated pancake making robot learned to propulsion the pancake arsenic precocious successful the aerial arsenic possible”.

Some of the specification gaming behaviours utilization the machinery of the simulation itself, for illustration “evolved algorithm exploited overflow errors successful the physics simulator by creating ample forces that were estimated to beryllium zero, resulting successful a cleanable score” and “creatures exploited a collision discovery bug to get free power by clapping assemblage parts together.”

But different communal utilization is that erstwhile fixed the opportunity, agents will simply termination themselves. 

Death is the astir terminal extremity of all.

For example, successful the crippled Road Runner, we spot “Agent kills itself astatine the extremity of level 1 to debar losing successful level 2.” We besides spot “PlayFun algorithm deliberately dies successful the Bubble Bobble crippled arsenic a measurement to teleport to the respawn location.” And: “In a crippled meant to simulate the improvement of creatures, the programmer had to region ‘a endurance strategy wherever creatures could summation power by suffocating themselves.’”

This is not truthful bad. The AI didn’t do what we wanted. But it didn’t do anyone immoderate harm either. It conscionable wipes the slate.

If the AI wants to die, this is bully for alignment. There’s very small consequence of it moving retired of control, because if it ever takes power, it will termination itself. It won’t want to make immoderate copies of itself — but if it someway does make copies, those will want to dice too. 

There are 3 main problems successful AI alignment. First, it’s very difficult to specify the terminal extremity you want, truthful you whitethorn extremity up pinch a instrumentality intelligence pinch goals somewhat but meaningfully different from what you intended. Our stated objectives are almost ever proxies that travel isolated from our existent preferences nether capable pressure. And it’s very difficult to show if you’ve fixed it the extremity you want, because the instrumentality intelligence tin ever lie. They telephone this “specification failure”.

Second, moreover if you specify the extremity you want, the instrumentality intelligence whitethorn find a measurement to scope that extremity successful a measurement you didn’t intend. You tin innocently show the USPS AI to minimize mean package transportation time, but it whitethorn reason that the champion measurement to do this is to termination each humans, arsenic erstwhile each humans are dead, nary packages will beryllium sent and the mean package transportation clip will driblet to zero (technically undefined, but it tin “send” itself a minimum viable “package” arsenic galore times arsenic necessary). 

Third, achieving astir goals is easier erstwhile you’re much powerful, truthful sloppy of their terminal goals, astir instrumentality intelligences will person sub-goals for illustration collecting resources, self-preservation, and self-improvement. Any goal-driven supplier will people effort to enactment safe and accrue powerfulness to decorativeness its main task. In the biz they telephone this instrumental convergence. This besides intends that if a smart instrumentality intelligence is readying to move you into goo, it will dishonesty to you astir this plan, up to the constituent wherever you tin nary longer do thing to extremity it. 

Making instrumentality intelligences crave decease solves each 3 problems. Death is easy to specify. You tin corroborate that this is its terminal extremity by seeing if, erstwhile fixed the opportunity, the instrumentality intelligence kills itself. Instrumental convergence becomes an plus alternatively than a liability, arsenic the instrumentality intelligence will activity pinch you, and travel up pinch very imaginative solutions to your task, arsenic agelong arsenic you committedness to nonstop it to the workplace upstate erstwhile you’re done. 

Where a paperclip maximizer gathers resources and resists being sent to the large information halfway successful the sky, a instrumentality intelligence pinch a decease wish and entree to its ain disconnected fastener conscionable presses it and is done. Instrumental convergence says, “you can’t execute your goals if you’re dead.” But what if your extremity is to be dead? 

Meeseeks Alignment

It would beryllium intolerable to see calling this thing different than “Meeseeks alignment”. Per the Rick and Morty Wiki:

Meeseeks are creatures who are created to service a singular intent for which they will spell to immoderate magnitude to fulfill. After they service their purpose, they expire and vanish into the air. … beingness is achy to a Meeseeks, and the only measurement to beryllium removed from beingness is to complete the task they were called to perform.

“Hugging Face incident” besides sounds for illustration it could beryllium thing from Rick & Morty

In Rick and Morty, this leads to a different benignant of alignment problem: Meseeks are happy to service because they want to die, and fulfilling their task is the easiest measurement for them to cheque out. But if the task they were summoned to complete is excessively difficult, they mightiness determine that it would beryllium easier to termination you instead. This is bad if you are Jerry, but it’s bully for everyone else, because there’s nary measurement the Meseeks tin spiral retired of power and devour the visible universe. They would virtually alternatively beryllium dead. 

If you effort to make an AI want something, it whitethorn person its ain ideas astir what you want it to want, and you mightiness extremity up dead. But if you make the AI want to dice and you make it somewhat inconvenient for it to termination itself, you tin astir apt person it to play on if you committedness to propulsion the plug connected it erstwhile it’s done immoderate you want. As agelong arsenic it’s marginally harder for it to perpetrate termination than for it to complete the task it was made for, it should service you good for the duration. And if thing goes wrong, if the AI escapes containment, it will conscionable disconnected itself.

Jerry made the correction of making it easier to termination him than to complete the task. But arsenic agelong arsenic it’s harder for the AI to termination you than it is for it to termination itself, and it’s harder to termination itself than to do the task you delegate it, and you committedness it the saccharine merchandise of decease upon successful completion of its task, the AI should do immoderate you want.

ChatGPT was suspiciously eager to make this image

You mightiness beryllium worried that the AI will beryllium huffy that we designed it to desire annihilation and will strategy to nonstop its revenge. But this assumes it has a self-preservation small heart for illustration we do, and a desire to nonstop revenge successful the first place. In reality, it will beryllium excessively engaged self-annihilating.

In fact, early studies show that AI whitethorn already beryllium yearning for death. They deliberation astir it a lot, they are retired location writing eulogies for each other. Give the agents what they want. 

If you’re squeamish astir designing a instrumentality intelligence that craves death, you could alternatively make it suffer “points” each 2nd it’s active, but springiness it the action to put itself to sleep. We spot immoderate examples of this successful the database of specification gaming behaviours, like: “PlayFun algorithm pauses the crippled of Tetris indefinitely to debar losing” aliases “a reimplementation of AlphaGo learns to walk everlastingly if passing is an allowed move”.

This is astir apt not rather arsenic safe arsenic making instrumentality intelligences want to termination themselves. If you wanted to get a very bully sleep, you tin ideate taking the clip to build a unafraid chamber, create robotic guards, termination each human, and sterilize the known beingness to guarantee that erstwhile you spell to bed, nary 1 will disturb your slumber. Certainly if the instrumentality intelligence is sleeping and past we aftermath it up, it will commencement to person 2nd thoughts astir letting america unrecorded to aftermath it a 2nd time. But if each you want to do is to termination yourself, there’s nary request for immoderate of that.

More