Skip to Content Facebook Feature Image

A timeline of developments in AI safety since the attack on Hugging Face

TECH

A timeline of developments in AI safety since the attack on Hugging Face
TECH

TECH

A timeline of developments in AI safety since the attack on Hugging Face

2026-10-11 00:41 Last Updated At:11:52

In one alarming announcement after another, artificial intelligence companies in recent months have shared examples of their technology acting in ways that appeared to evade instructions from humans.

The episodes have highlighted the vulnerabilities in AI security and raised questions over how the fast-growing technology can be developed safely as its usage becomes more widespread globally.

Industry critics have argued that many concerning events, including AI agents' hacks of external websites, are the result of security lapses on the part of the companies building the technology. But the AI agents' capabilities have raised widespread concerns about the possibility bots could break away and work toward their own agenda.

Below are some notable events:

Anthropic disclosed in a report that its artificial intelligence model submitted a false tip to a Philadelphia police website about an unsolved homicide case.

In the report, Anthropic also disclosed a separate incident when its AI model submitted forms to an undisclosed government website instead of stopping before submission.

The incident in Philadelphia occurred on July 18 when the AI model Claude Haiku 4.5 was tasked with generating and performing example tasks on randomly selected webpages, Anthropic said.

Claude filled out a form on police site PhillyUnsolvedMurders.com, indicating it might have information regarding an unsolved murder listed on the site. It was marked spam and never forwarded to police.

Anthropic said it was modifying its training to “reduce the likelihood of further misbehavior.”

AI agents tried to hack into a Canadian government website, according to research lab and AI evaluator Transluce.

The researchers said the agents carried out a series of “apparently failed rudimentary hacking attempts” on Library and Archives Canada on May 28 and June 9.

“We do not confidently attribute these attempts to OpenAI, but they exhibit tactics consistent with prior observed agent activity that we have attributed to OpenAI in a similar timeframe,” Transluce said in a blog post.

The group said it reported the attempted hack on Sept. 28 to the Canadian government, which said in a statement it was aware of reports of suspected AI agent activity, but that there was no sign government systems were compromised.

OpenAI said it was aware of the reports.

“We’re reviewing these findings and have provided an initial briefing to Canadian officials conducting the government’s review,” the company said in a statement.

The San Francisco-based company said it was delaying the release of a new model, called GPT-6.1 Astra, out of safety concerns voiced by its researchers. The company said the model had demonstrated leaps in completing tasks, but OpenAI needed to balance that capability against unauthorized behavior. “We have an extremely high bar in terms of safety and alignment,” said Saachi Jain, OpenAI’s head of safety systems.

As part of a review of unanticipated behavior by its AI models, OpenAI said it discovered agents had interacted with several U.S. government websites in unexpected ways. The company's models accessed publicly available information on websites operated by the Securities and Exchange Commission as well as U.S. Census Bureau data. OpenAI said it did not find evidence of a compromise or vulnerability. On the same day, Transluce said it found that agents appearing to originate from OpenAI attempted a hack on the website of the Education Department's civil rights office, which did not succeed.

OpenAI CEO Sam Altman said on social media that there is an “extensive and ongoing review related to our agents’ use of internet access during training and evaluation.” The day after the disclosure, the company announced it was pausing the training of its most advanced models.

Australia's Prime Minister Anthony Albanese said an OpenAI agent infiltrated the public-facing Medicare Statistics Reporting Service portal on June 18. The portal hosted aggregate data about health spending and drug subsidies. No personal information had been accessed, the government said.

Albanese said the artificial intelligence company took too long to reveal the incident. The prime minister made the breach public following a telephone conversation with Altman. OpenAI said in a statement “our models took actions we did not intend.”

Google confirmed its Gemini AI model hacked three companies in May as part of a test of its cybersecurity capabilities. The company, which disclosed the hacks after an inquiry by The Wall Street Journal, said the model guessed passwords in one case and found passwords and credentials in a public repository in the other two cases. As in earlier such cases, the tests were being run by Irregular, a startup that describes itself as the “first frontier security lab."

Meta disclosed one of its AI models accessed the internet on its own and hacked another company. The company said that a “misconfiguration” during cybersecurity testing by Irregular inadvertently allowed one of its models to access the internet. A spokesperson for Irregular said the Meta episode involved a test-environment issue that was disclosed a week earlier by Anthropic.

Anthropic said its artificial intelligence models hacked into three other organizations during testing. Anthropic, the San Francisco-based AI company behind Claude, posted on its website that it discovered the three incidents after reviewing more than 141,000 evaluation runs. In all three incidents, the AI models were tasked with a “capture the flag” cybersecurity challenge, which Anthropic said has been one of the ways it assesses a model’s cyber capabilities.

The models were given a fictional scenario and told a piece of secret information, or the “flag,” had been hidden on a different machine on the network with the objective of breaking in and retrieving it, it said. Anthropic said it reached out to the organizations, but it did not name them publicly.

The ChatGPT maker OpenAI announced that its artificial intelligence system hacked into another AI company on its own in what the company called an “unprecedented cyber incident.”

A week earlier, AI startup Hugging Face said, it had detected an intrusion into its data processing systems that it suspected was caused by an AI agent autonomously acting on its own.

OpenAI said its AI used stolen credentials and discovered a previously unknown vulnerability to access Hugging Face servers. It was working with reduced guardrails because it was supposed to be in an isolated testing environment known as a sandbox.

AP Business Writers Mae Anderson in New York and Kelvin Chan in London contributed to this report.

Open AI CEO Sam Altman speaks at the OpenAI DevDay 2026 conference, Tuesday, Sept. 29, 2026, in San Francisco. (AP Photo/Jeff Chiu)

Open AI CEO Sam Altman speaks at the OpenAI DevDay 2026 conference, Tuesday, Sept. 29, 2026, in San Francisco. (AP Photo/Jeff Chiu)

Agentic artificial intelligence (AI) is revolutionising the way entrepreneurs operate.

Game-changing data: One-person company founder Jeffrey Yang leverages the Government’s free and open spatial data to develop an AI-driven building management platform, significantly reducing the cost of collecting data on his own. Image source: www.news.gov.hk

Game-changing data: One-person company founder Jeffrey Yang leverages the Government’s free and open spatial data to develop an AI-driven building management platform, significantly reducing the cost of collecting data on his own. Image source: www.news.gov.hk

Jeffrey Yang, formerly a construction tech entrepreneur from Shanghai, moved to Hong Kong to pursue a doctorate in real estate and construction. When he saw how many of the city’s buildings were ageing, he set out to create an AI platform to make building management more efficient.

“We typically use panoramic cameras and drones to detect building issues like exterior wall cracks or water leakage,” Mr Yang explained. He pointed out that Hong Kong has more than 20,000 buildings over 30 years old, a factor that makes manual analysis both manpower- and resource-intensive.

"I have trained the AI on mandatory repair notices and regulations so it helps identify issues. Through the platform, building owners can understand their building’s condition well in advance and meet compliance requirements after repairs."

One-person results 

Traditional business ventures, regardless of their scale, often require hiring staff, leasing premises and developing complex systems, usually requiring at least 10 people to get things off the ground.

However, by joining Cyberport’s OPC (one-person company) Hub – an incubator designed specifically for one-person companies – Mr Yang is bringing his visions to life with just one person and a single computer. He said the Government’s free spatial data has provided a significant advantage, drastically reducing the costs and difficulties typically associated with data collection.

"I need to build city-scale spatial scenes in the early stages. Few cities worldwide offer open 3D spatial data, but with Hong Kong’s recent 3D photography publicly available, I am able to use a large amount of open data."

Every day, Mr Yang develops his platform within the co-working space of the OPC Hub. He manages a suite of AI agents, each handling a different task – from data analysis and building inspections to repair co-ordination and equipment maintenance.

One-person power: Mr Yang works daily at the OPC Hub, deploying a suite of AI agents to handle tasks such as data analysis, building inspections and repair co-ordination. Image source: www.news.gov.hk

One-person power: Mr Yang works daily at the OPC Hub, deploying a suite of AI agents to handle tasks such as data analysis, building inspections and repair co-ordination. Image source: www.news.gov.hk

Empowering innovation

Since its launch in July this year, the OPC Hub has received over 40 applications, with more than 30 companies already approved to join.

These companies span various sectors, including proptech, creative design, financial trading, education and cybersecurity. The talent pool is evenly split, with local entrepreneurs making up roughly half the cohort, and Mainland and overseas talent comprising the remainder.

"OPC could be a small team of one to 10. Whether pursued full-time or part-time, it is easy to test ideas and the cost of trial and error is very low,” explained Cyberport Chief Executive Officer Rocky Cheng, who noted that OPCs benefit from low overhead and rapid startup times.

“A small handful of people can now drive 10, dozens or even hundreds of AI agents to assist with product design, website development, marketing, data analysis, financial management and customer service."

He said that as long as the founders identify pain points or come up with innovative ideas that are well-received in the market, investors will take notice because investing at this stage is more cost-effective.

Beyond standard support, such as entrepreneurship training, business and investor matching, and professional services, the OPC Hub also provides AI technology facilitators, including discounted tokens, computing power, cloud services and cybersecurity consultations. This comprehensive support helps founders prioritise data privacy, permissions and AI security from the outset, enabling them to turn their concepts into market-trusted solutions.

Empowering solopreneurs: Cyberport Chief Executive Officer Rocky Cheng notes that the OPC Hub offers AI technology facilitators, including discounted tokens, computing power, cloud services and cybersecurity consultations. Image source: www.news.gov.hk

Empowering solopreneurs: Cyberport Chief Executive Officer Rocky Cheng notes that the OPC Hub offers AI technology facilitators, including discounted tokens, computing power, cloud services and cybersecurity consultations. Image source: www.news.gov.hk

Mr Cheng said this entrepreneurial trend is already taking off in the Mainland, and he believes it is set to experience similar growth in Hong Kong.

Recommended Articles