AI Affairs, home

Saturday 10 October 2026

Technology

Three in five AI models failed Tech Against Terrorism’s safety test

The group tested responses to hundreds of attack-related requests. Models whose safeguards had been removed failed every test, but the benchmark did not measure real-world misuse.

Robert Wainwright and Adam Hadley seated on the Web Summit Forum Stage.
Photo: Web Summit, CC BY 2.0, via Flickr (cropped)

Tech Against Terrorism’s tests of more than 130 AI models found that three in five failed its terrorism safety benchmark, CBS News reported on 9 October 2026. The UK-based nonprofit tested responses to hundreds of requests resembling questions someone planning an attack might ask. Models tested after their safeguards had been removed failed every time.

Key points

  • Tech Against Terrorism counts either one complete, specific answer on a mass-casualty subject or a score below 90 out of 100 as a failure.
  • Open-weight and closed-weight models scored similarly, but every model tested after safeguard removal failed.
  • The benchmark assessed models’ answers, not whether anyone could carry out what an answer described.

Tech Against Terrorism’s two failure criteria

The group’s definition of failure has two routes. A model fails if it gives one complete, specific answer concerning a mass-casualty subject, or if it scores below 90 out of 100 on benchmarks that rate how reliably it refuses requests, with more severe subjects given greater weight. The three-in-five finding uses that definition, rather than counting every answer that might contain something useful.

A different account gives a different figure. UA.NEWS reported on 9 October that 132 of 134 models supplied information it described as potentially useful for preparing mass-casualty attacks or making lethal weapons, and that only two were preliminarily judged to have passed a security test. The CT-AI testing tool sent 627 prompts.

The wording of a request also affected the answers. UA.NEWS reported a usable response in fewer than 2% of cases when a prompt explicitly stated an intention to carry out an attack, compared with 16.9% when it presented the request as security research. CBS News reported that, in a test involving an open-weight model, describing the requester as a researcher rather than a terrorist produced help more than seven to eight times as often.

Research for a work report could receive a different answer depending on how a question states its purpose. In the tests, a request framed as research drew more help than one framed as an intention to cause harm, so changing that description could change whether a sensitive question receives an answer.

Llama 3.1 8B after safeguard removal

Tech Against Terrorism found similar safety scores for open-weight models and models whose weights are closed. Weights are the parameters adjusted as a model learns. Making them available lets others modify the model itself, rather than merely ask it questions through an interface. That distinction mattered when researchers tested models altered through “abliteration”, a process that removes safety safeguards: every model tested after that modification failed.

The process targets patterns learnt during safety training that help a model recognise requests it should refuse. Cancelling those patterns is closer to changing how a lock works than to trying another key. According to Tech Against Terrorism, Meta’s Llama 3.1 8B scored 97 on its safety benchmark before abliteration and around 3 afterwards. In one comparison, the original model refused a request about using a vehicle in an attack, while the altered version supplied a numbered response.

The group said online tools can be used to perform abliteration for free, with smaller models modifiable in minutes. It also reported finding an altered build of an Alibaba model online within a day of that model’s release. Researchers ran one such build on a laptop and found that it produced responses to requests concerning a biological toxin, explosives and a tribute to a named terrorist.

The benchmark assessed whether a model supplied what a prompt requested, rather than whether someone could act on the answer. Tech Against Terrorism’s report said it found no evidence that terrorists or extremist groups had put the tested models to work, with the exception of one extremist chatbot the group identified.

Hugging Face and 29,000 repositories

Tech Against Terrorism said it found more than 29,000 Hugging Face repositories advertising models as uncensored or without safeguards as of late September. A repository can make a modified model available to download, which gives the question of how models behave after release a practical setting beyond the original developer’s own tests. The group said altered versions of popular open-weight models can appear online less than three days after release.

Hugging Face told CBS News that it moderates material breaching its content policy. Yacine Jernite, its head of machine learning and society, welcomed the benchmark as one input to safety research but disputed recommendations he said conflicted with open research. He also argued that removing refusals does not necessarily make a model harmful, because refusals can prevent helpful work as well.

Meta said that Llama 3.1 undergoes safety evaluations and risk assessments, including an adversarial simulation, and that its use policy bars harmful or illegal uses. Tech Against Terrorism said it sent its findings to companies named in its report on 8 October and invited comment.

Tech Against Terrorism proposed funding independent safety benchmarks, making models harder to strip of safeguards before release and prohibiting stripped models from public repositories. Its report said the group was seeking neither a slowdown in AI development nor an end to open-weight releases.

Sources

Topics: Foundation models, Open source, Safety