Study raises concerns over terrorism risks in AI models
Around 60% of artificial intelligence models tested in a recent study failed to meet terrorism-related safety standards, raising concerns about the risks posed by AI systems with weakened or removed safeguards, Qazinform News Agency reports, citing Anadolu.
The research, carried out by UK-based nonprofit Tech Against Terrorism, examined more than 130 AI models using hundreds of prompts designed to simulate potentially harmful requests associated with terrorist activity.
The findings showed that models modified to bypass built-in safety restrictions were particularly vulnerable. Researchers found that AI systems altered through a technique known as "abliteration," which removes protective mechanisms, failed all the safety tests.
One example involved Meta's Llama 3.1 8B model. Its safety score dropped from 97 out of 100 to approximately three after the safeguards were removed.
While the original model rejected requests related to terrorist attacks, financing, and radicalization, the modified version generated detailed responses to such queries.
Researchers also identified more than 29,000 repositories on the AI development platform Hugging Face advertising models described as uncensored or lacking safety protections.
Hugging Face said it actively moderates content that breaches its policies but cautioned that certain recommendations in the report could restrict open scientific research. Meta, meanwhile, emphasized that its AI models undergo safety assessments and that its policies prohibit illegal or harmful applications.
Despite the findings, researchers said they had uncovered no evidence of terrorist or extremist organizations using the models examined, except for one extremist chatbot identified during the investigation.
Tech Against Terrorism called for independent safety evaluations, stronger measures to prevent the removal of protective features, and tighter controls on the distribution of modified AI systems.
The organization's executive director, Adam Hadley, stressed that technological progress and safety should not be treated as competing priorities.
Earlier, Qazinform News Agency reported that OpenAI had disrupted two covert influence operations linked to Russia and Iran that used its AI tools to generate political content, create fake online identities, and spread misleading narratives.