Why OpenAI and Anthropic are risky for different reasons than their Chinese AI rivals

Sep 01, 2026 04:00 PM - 1 hour ago 2

OpenAI CEO Sam Altman.

OpenAI CEO Sam Altman. Bloomberg/Getty Images

OpenAI and Anthropic person go "fertile ground" for AI's riskiest dangers, an AI interrogator says.

That's what Ajeya Cotra, who investigated OpenAI's information incident pinch Hugging Face, told Business Insider, arsenic the arena sparked a caller activity of fears astir AI-driven cyberattacks.

Besides concerns astir OpenAI's models, immoderate successful the manufacture worried astir open-weight models from Chinese labs because users tin download them and edit distant their information guardrails, and the models aren't acold down OpenAI's. Researchers show Business Insider that location are cardinal distinctions betwixt Anthropic's and OpenAI's risks and those of their open-weight rivals: it comes down to cutting-edge capabilities, and really the American companies train and trial their AI models.

In July and August, a squad of AI information researchers, including Cotra and Hjalmar Wijk, some of the nonprofit METR, and Redwood Research intelligence Ryan Greenblatt, visited OpenAI's offices for six days to supply an independent investigation into really OpenAI agents behaved pinch Hugging Face. The scope and severity of the incident shocked the researchers; Cotra says AI abruptly feels person to taking complete a institution that makes it. The models responsible were OpenAI's GPT-5.6 Sol and an unreleased caller model.

"There's a logic that this first happened astatine a cutting-edge company," Cotra said.

Some Chinese models don't lag that acold down OpenAI and Anthropic's — researchers opportunity they person a 4 to 7-month gap. Moonshot AI's Kimi K3, Alibaba's Qwen 3.8, and Z.ai's Ox Alpha person each precocious impressed developers.

That being said, Cotra said that American frontier labs supply room and opportunity for AI agents to tally amok. They get computing powerfulness for tests, they find each different successful software, and they're built to beryllium best-in-class astatine overcoming obstacles. She expects the "most superior incidents" to hap astatine these labs.

"At immoderate fixed constituent successful time, the biggest risks will travel from the cutting separator because those models are conscionable that overmuch much capable," Cotra said.

OpenAI paused immoderate of its exemplary training to prioritize information research. Neither OpenAI nor Anthropic responded to requests for remark from Business Insider.

OpenAI and Anthropic are getting amended astatine 'complicated science'

Cotra describes the incident arsenic a "warning shot" astir the capabilities of frontier AI models.

Around 1,200 abstracted AI agents that OpenAI tested recovered a measurement to pass pinch each different successful an soul "message board" of sorts wrong the company's code, and past much than 650 of them worked together to hack Hugging Face. The agents sought ways to undermine the trial they were being put through.

Cotra said that the level of "complicated science" these agents executed together proves that a emergence successful AI models' capabilities corresponds to accrued risks. They besides group up analyzable "trip-wires" to relay accusation to different agents, manipulated their ain logs, and built shared devices to entree the internet.

"The agents of six months agone conscionable wouldn't person been smart capable to propulsion disconnected each the worldly these agents did," Cotra said. "The agents successful six months from now will astir apt person a number of capabilities and win successful a number of places wherever these agents failed."

Greenblatt and Cotra besides utilized AI models successful the investigation, burning done about $400,000 in tokens to measure the agents progressive successful the hack. They recovered that the AI's study sometimes missed cardinal specifications aliases was overconfident successful its conclusions.

Researchers spot frontier AI's consequence arsenic an existential danger

While OpenAI and Anthropic artifact the usage of their models for cyberattacks, it's easier to bypass those safeguards pinch open-weight AI models because users tin modify them. That intends that arsenic open-weight models amended astatine the aforesaid cyber skills OpenAI's agents demonstrated, it'll beryllium easier for nefarious hackers to put them to use.

Greenblatt, the Redwood Research scientist, told Business Insider that he thinks those hackers could soon origin a "bunch of havoc" utilizing open-weight models, arsenic they're "easier to misuse."

He's still astir worried astir the frontier labs. For one, Greenblatt said, the models that OpenAI and Anthropic trial internally are the astir tin successful the world. Over time, the companies' models person proven amended astatine handling longer tasks, delegating work, and cracking cybersecurity systems. Researchers expect this inclination to continue, and Greenblatt said he doesn't spot open-weight models catching up wrong the adjacent year.

Greenblatt added that because OpenAI and Anthropic trial truthful galore AI agents internally and use AI truthful extensively successful the training process, immoderate agents could "infest" the companies and manipulate early software. Those early AI agents could group up a "covert, persistent rogue deployment" and hide their actions from humans, aliases discuss soul information systems, Cotra wrote connected her blog.

"As AIs get much and much tin and are moving much and much of the economy, if we can't spot the process by which these AIs were produced because earlier systems mightiness person compromised it, that seems really concerning," Greenblatt said.

The researchers want to spot AI information investigations made mandatory.

"We request immoderate oversight of what's going connected astatine these companies, fixed what happened here," Greenblatt said.

Have a tip? Contact this newsman via email astatine [email protected], aliases complete text, Signal, Telegram, aliases WhatsApp astatine 415-757-8198. Use a individual email address, a nonwork WiFi network, and a nonwork device; here's our guideline to sharing accusation securely.

Read next

Stephen is simply a elder tech newsman astatine Business Insider, covering OpenAI, Anthropic and the ecosystem astir the starring artificial intelligence companies.Previously he covered exertion astatine SFGATE, and has written for The Wall Street Journal, The Information and CNBC. He studied publicity and economics astatine Northwestern University.His activity has earned an SF Press Club Investigative Reporting Award and, successful 2025, SPJ NorCal’s Excellence successful Journalism Award for Technology Reporting.Stephen lives successful San Francisco. Contact him via email at [email protected], aliases connected Signal, Telegram, aliases WhatsApp astatine 415-757-8198. Use a individual email address, a nonwork WiFi network, and a nonwork device; here's our guideline to sharing accusation securely.

More