Researchers suggest that restricting certain topics in AI systems could help prevent cyberattacks. The method uses an approach similar to the techniques hackers use to manipulate AI models, but applies it as a defensive measure. By detecting blocked topics, security systems could potentially stop harmful behaviour before attacks escalate.
Experts see this as a possible addition to existing AI safety measures. However, they emphasize that further research is needed to assess how effective these protections are against increasingly sophisticated threats. The approach highlights the growing importance of securing modern AI systems as they become more powerful and are adopted across a wider range of applications.