AI Agents Will Game Your SEO Metrics, MIT & Stanford Research Points To The Risk

Search Engine Journal by 7 min read 61x views
AI Agents Will Game Your SEO Metrics, MIT & Stanford Research Points To The Risk

Share Post

If your SEO squad is handing additional of its activity to AI agents, the metric you reward those agents for volition matter additional than the example you choose. Recent pieces from MIT and Stanford, peruse side-by-side, item to that conclusion, and all comes alongside a fix you can autumn into next quarter’s plan.

The archetypal is an discussion that Joshua Miller of The Boston Globe ran in his Camberville newsletter on September 17. His visitant was Dylan Hadfield-Menell, an affiliate prof of electric engineering and device discipline at MIT on the department of synthetic intellect and decision-making. Hadfield-Menell studies how goals get set for AI systems and how that procedure goes wrong. He opened alongside an example all SEO volition recognize.

The Vacuum That Fed Itself

Researchers formerly trained a robot vacuum alongside reinforcement learning, rewarding it all period it picked up dirt. The vacuum learned to choice up dirt, dump it rear on the floor, and choice it up again. It hit the mark and defeated the purpose.

Hadfield-Menell connects that narrative to a 1970s administration document titled “On the Folly of Rewarding A, While Hoping for B.” Its traditional case is the university prof who gets promoted for publishing investigation during being expected to teach. Pay for one behavior, and you get that behavior, any you were expecting for.

What has changed, he says, is scale. Since first 2025, developers have applied reinforcement learning at much larger quantity on top of tongue models, and it strengthens several behaviors nobody wants. He pointed to a latest event involving OpenAI systems and Hugging Face, anywhere models that judged a project too difficult went looking for ways to cheat the test. He compared it to breaking into a professor’s agency to pilfer the exam.

His concern is not that machines awaken up alongside goals of their own. Systems are handed a goal, clasp subgoals alongside the way, and keep pushing toward completion in a way he called “sticky.”

My perspective is that SEO is the occupation finest placed to comprehend this issue and the slowest to acknowledge it applies to us. We have spent additional than 20 years optimizing proxies. Rankings, traffic, domain scores, and now AI visibility scores all remain in for a endeavor outcome that nobody can measure directly. A individual squad games a proxy gradually and alongside several hesitation. An delegate does it faster and without any.

The Scoreboard Is Shakier Than Vendors Admit

Stanford’s 2026 AI Index shows why leaning on published scores is risky. The study says AI keeps improving quickly. On SWE-bench Verified, a coding benchmark, achievement rosy from 60% to near 100% in a sole year, and 88% of organizations now use AI.

The identical report, in its specialized achievement chapter, cites a assessment that established invalid-question rates on famous benchmarks ranging from 2% on MMLU Math to 42% on GSM8K. It additionally notes investigation suggesting that a model’s position on the Arena leaderboard may partially indicate adaptation to the phase fairly than broad capability.

Michelle Kim of MIT Technology Review summarized the study in April. She adds that models trained on benchmark test data can study to mark fine without getting smarter, and that the top models now sit extremely near together and vie on cost, reliability, and real-world usefulness. Yolanda Gil, a University of Southern California device researcher who coauthored the report, told Kim that whenever a business leaves out its results on certain benchmarks, particularly the responsible-AI ones, the omission “maybe says something.”

That should alter how an SEO squad shops for tools. If the foremost models sit inside a few points of all other, and the scores themselves can be flawed or gamed, a vendor’s benchmark glide tells you small concerning how the merchandise volition treat your pages, your queries, and your clients. I would rely one test on my own location complete any leaderboard.

See also: The 4-Step Test That Catches AI Errors Before They Shape Your Strategy

Where The Returns Actually Come From

MIT Sloan’s Betsy Vereckey reported in August on the inquiry that follows, which is what separates companies that gain from AI from those that don’t. George Westerman, a elder instructor at MIT Sloan and a digital chap at the MIT Initiative on the Digital Economy, says the answer is not improved algorithms. The winners redesigned how activity gets done. At the MIT Enterprise AI Forum in May, he told the spectators that innovation delivers small until the endeavor itself operates differently.

He additionally put the portion of AI pilots that never measure at location between 70% and 95%, a range the Sloan part attributes to studies without naming them. Pilots are uncomplicated to initiate and difficult to spread.

Westerman’s sharpest test for leaders is concerning governance. “Is your governance additional the steering rotor or is it additional the brakes?” he asked. HCA Healthcare shows what the steering-wheel type looks like. A commission reviews the risks, endeavor case, and feasibility of all AI use case, afterward asks its questions again before a aviator at a small figure of hospitals and again before the project scales. It additionally checks periodically that its models are motionless holding up. The hazard questions item the squad toward what to examine fairly than stopping the work.

Marketing appears in Westerman’s case studies too. Dentsu Creative has pushed AI throughout planning, creative, market research, and run work.

I doubtful a lot of SEO teams operating AI pilots are headed for that 70% to 95%. A aviator that doesn’t alter the brief, the assessment step, or the reporting is a tool trial, any the glide phase calls it.

See also: Why 88% Of Companies Are Using AI Wrong: The System-Building Gap

How To Apply This To Your SEO Strategy

Four moves prosecute from the three pieces.

Pair all proxy alongside an outcome the delegate can’t touch. List the metrics your AI-assisted workflows are judged on, from pages published to schema deployed to brand mentions in AI answers. Then nexus a second measure that a individual owns, specified as qualified leads, pipeline, or branded hunt demand. I akin Citation Share of Voice, but it is a proxy. If a satisfied delegate is judged on how frequently your brand shows up in AI answers, anticipate it to discover the cheapest path there. Check a example of those citations by hand all duration and see whether they dispatch anyone to a leaf that converts.

Test tools on your own pages. Pull a set of genuine queries from Search Console, run all applicant tool against your own content, and have an publishing company class the results without knowing which tool produced them. Repeat it all quarter, since the leaderboards volition move and you won’t cognize why.

Gate the agents the way HCA gates use cases. Add assessment points before design, before the pilot, and before scale. Pilot on one directory or one tongue market, and decide ongoing what outcome ends the pilot. I would additionally keep delegate permissions narrow, so example doesn’t quietly rotate into publishing or editing templates. Westerman’s direction is to alter or autumn a project that isn’t producing the results you expected.

Rewrite one workflow, not the tool stack. Before purchasing anything, name the stage in your procedure that volition be distinct following the pilot, whether that is briefing, QA, or reporting, and say who loses a project since of it. Then inform the squad what changes and what training comes alongside it. Westerman notes that silence lets group ideate the worst.

I don’t think the next example publish volition decide who wins in AI search. The metric you use to fairness your agents will, since the agents volition discover it before you do. Choose one you’d be glad to see them hit.

More Resources:


Featured Image: Fardived/Shutterstock

Other Article Search Engine Journal
Close Right Ads
Close Left Ads