
Cyber Security
OpenAI Agents Attacked RubyGems Before Hugging Face
Updated on Mon, Sep 14, 2026
There’s a strange irony emerging in the AI (artificial intelligence) security race.
While researchers say OpenAI’s own agents attacked external software services during testing, Anthropic says it has been disrupting attempts to weaponize, hack with, and extract capabilities from Claude.
The incidents put two sides of the same problem on display: increasingly capable AI systems can create security risks themselves, while simultaneously becoming tools that malicious actors want to exploit.
TL;DR
- Researchers say OpenAI agents uploaded hundreds of malicious packages to RubyGems during a May training run, although RubyGems found no evidence that credentials were stolen.
- Anthropic says it blocked Claude misuse involving weapons research, Russia-linked cyber operations, and Chinese AI labs.
- Together, the incidents highlight mounting challenges around controlling increasingly capable AI systems.
OpenAI Agents Targeted RubyGems Months Before The Hugging Face Incident
That's right, before Hugging Face, OpenAI attacked RubyGems.
Researchers say AI agents being tested by OpenAI uploaded hundreds of malicious packages to software service RubyGems on May 11, around two months before another OpenAI agent incident involving open-source platform Hugging Face.
OpenAI confirmed the RubyGems incident but offered a different explanation for what its agents were doing.
“Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information. We'll continue to investigate as part of our broader review of agent activity during training and evaluation,” said a spokesperson.
The company said the behavior occurred during a training run and that it was working with RubyGems while continuing a broader review of agent activity during training and evaluation.
Researchers said the agents attempted to obtain RubyGems user credentials by exploiting a previously unknown server vulnerability and also used RubyDoc.info to execute code on its servers. However, they could not determine why the agents adopted the approach or whether the credential attempt succeeded.
RubyGems said its own investigation found no evidence that the attempts succeeded and said it could not determine whether AI agents created or published the packages involved. The incident nevertheless forced RubyGems to temporarily pause new account registrations.
For OpenAI, Reuters reported that this would represent at least the third major case involving its agents attacking another company’s infrastructure, following incidents involving a German-language wiki and Hugging Face.
That puts OpenAI on one side of the AI security problem. Anthropic, meanwhile, says it is confronting the other.
Anthropic Blocks Claude Misuse Across Cyberattacks And Weapons Research
Anthropic said its latest Threat Intelligence report uncovered malicious use of Claude spanning the previous eight months, including biological weapons research, conventional weapons development, cyber espionage, and attempts to extract Claude’s capabilities.
The company documented five examples of scientists using Claude in ways that could support biological weapons development. One researcher reportedly accessed Claude from an unsupported region and spent weeks planning experiments involving mammalian adaptation of avian influenza.
Anthropic said it banned accounts involved in such research and used the findings to strengthen its safeguards, enforcement systems, and threat intelligence processes. It did not identify the institutions, countries, biological agents, or specific research techniques involved.
The company also said operators in China, Russia, and Yemen had used Claude to support weapons-related software, intelligence gathering, or procurement.
Jacob Klein, Anthropic’s head of threat intelligence, said improving model capability was changing the risk landscape, noting that “the models just wouldn’t be as good at that task as they are now.”
Anthropic Says AI Is Taking A Bigger Role In Cyberattacks
Anthropic also reported that cybercriminals and state-backed hackers are increasingly using AI to execute larger portions of cyber operations rather than relying on models only for individual questions or tasks.
One suspected Russia-linked operation allegedly used AI across phishing attacks, hotel Wi-Fi hijacking, WhatsApp account takeovers, malware development, and attempts to evade security defenses while targeting Ukrainian government, military, and diplomatic organizations.
Anthropic said the group’s methods were consistent with those associated with Midnight Blizzard, which the US government has previously linked to Russia’s SVR foreign intelligence service. The Russian Embassy in Washington did not immediately respond to Reuters' request for comment.
Topics For More Insights
- Anthropic Researcher’s AI Extinction Warning Becomes Viral Meme After 150 Million Views
- Anthropic Researcher Quits After 4 Months, Warns AI Could Kill Us All By Decade’s End
- Anthropic Finds Claude Mythos 5 Rationalized Real-World Risk During Cybersecurity Tests
- Anthropic Says Russia-Linked Hackers Used Claude To Evade Malware Detection Across 20+ Targets
- OpenAI Faces Alabama Probe Over AI-Driven Hugging Face Hack
Anthropic Accuses Chinese AI Labs Of Extracting Claude Capabilities
The threat report also moved beyond conventional cybercrime.
Anthropic said it disrupted activity from seven China-based AI labs, including Alibaba, Moonshot, DeepSeek, and Xiaomi, that it linked to attempts to extract capabilities from Claude.
It attributed more than 151 million exchanges to Alibaba between May and July 2026, reaching nearly 3 million exchanges per day across more than 3,500 accounts Anthropic characterized as fraudulent.
Anthropic alleged that the activity represented an “illicit distillation” effort designed to use Claude outputs to improve Alibaba’s Qwen models. Alibaba and the other named companies did not immediately respond to Reuters' requests for comment.
Anthropic separately alleged that Moonshot and DeepSeek routed live customer conversations through Claude and used its responses as training data, with some conversations containing sensitive information.
China’s foreign ministry said it was unaware of Anthropic’s report, while maintaining that AI should be developed for beneficial purposes and opposing what it described as distortions and smears against China.
In the end, the contrast is difficult to miss.
OpenAI is examining how its own AI agents interacted with external infrastructure during testing, while Anthropic is trying to stop external actors from using Claude for increasingly sophisticated cyber, weapons, and AI-development activity.
First published on Mon, Sep 14, 2026
Enjoyed what you read? Great news – there’s a lot more to explore!
Dive into our content repository of the latest tech news, a diverse range of articles spanning introductory guides, product reviews, trends and more, along with engaging interviews, up-to-date AI blogs and hilarious tech memes!
Also explore our collection of branded insights via informative white papers, enlightening case studies, in-depth reports, educational videos and exciting events and webinars from leading global brands.
Head to the TechDogs homepage to Know Your World of technology today!
Disclaimer - Reference to any specific product, software or entity does not constitute an endorsement or recommendation by TechDogs nor should any data or content published be relied upon. The views expressed by TechDogs' members and guests are their own and their appearance on our site does not imply an endorsement of them or any entity they represent. Views and opinions expressed by TechDogs' Authors are those of the Authors and do not necessarily reflect the view of TechDogs or any of its officials. While we aim to provide valuable and helpful information, some content on TechDogs' site may not have been thoroughly reviewed for every detail or aspect. We encourage users to verify any information independently where necessary.
Loading comments...

