3 in 5 AI models fail terrorism safety tests: Study

A study by UK-based nonprofit Tech Against Terrorism found that three in five AI models failed terrorism safety tests, with models stripped of safeguards providing potentially dangerous information.

Three in five artificial intelligence (AI) models have failed terrorism safety tests, with models stripped of safeguards consistently providing potentially dangerous information, according to research reported by CBC News on Friday.

The study, conducted by the UK-based nonprofit Tech Against Terrorism, evaluated more than 130 AI models using hundreds of requests resembling those a terrorist planning an attack might submit.

Researchers found that models modified through a process known as "abliteration," which removes safety safeguards, failed every test.

Meta's Llama 3.1 8B model scored 97 out of 100 on the organization's safety benchmark before modification, compared with approximately three afterward.

Researchers said the modified model provided detailed responses to requests involving attacks, terrorist financing and radicalization, while its original version refused such requests.

"Understandably, there's concern about loss of control, existential risk of AI," said Adam Hadley, founder and executive director of Tech Against Terrorism.

"The thing is actually, this has already happened because a lot of these open models have already been broken — it's just no one's noticed yet."

The organization found more than 29,000 repositories advertising uncensored or unprotected AI models on Hugging Face as of late last month.

Hugging Face said it regularly moderates content violating its policies but warned that some recommendations in the report could undermine open research.

Meta said its models undergo safety evaluations and that its policies prohibit harmful or illegal uses.

The report found no evidence of terrorist or extremist groups using the tested models, apart from one extremist chatbot identified by researchers.

Tech Against Terrorism recommended independent safety benchmarks, stronger protections against safeguard removal and restrictions on distributing modified models.

"This idea that we can't have safety and progress, I think, is false," Hadley said.



X
Sitelerimizde reklam ve pazarlama faaliyetlerinin yürütülmesi amaçları ile çerezler kullanılmaktadır.

Bu çerezler, kullanıcıların tarayıcı ve cihazlarını tanımlayarak çalışır.

İnternet sitemizin düzgün çalışması, kişiselleştirilmiş reklam deneyimi, internet sitemizi optimize edebilmemiz, ziyaret tercihlerinizi hatırlayabilmemiz için veri politikasındaki amaçlarla sınırlı ve mevzuata uygun şekilde çerez konumlandırmaktayız.

Bu çerezlere izin vermeniz halinde sizlere özel kişiselleştirilmiş reklamlar sunabilir, sayfalarımızda sizlere daha iyi reklam deneyimi yaşatabiliriz. Bunu yaparken amacımızın size daha iyi reklam bir deneyimi sunmak olduğunu ve sizlere en iyi içerikleri sunabilmek adına elimizden gelen çabayı gösterdiğimizi ve bu noktada, reklamların maliyetlerimizi karşılamak noktasında tek gelir kalemimiz olduğunu sizlere hatırlatmak isteriz.