AI in practice

Aug 10, 2024Aug 10, 2024

Anthropic tests its "next-generation system for AI safety mitigations"

Matthias is the co-founder and publisher of THE DECODER, exploring how AI is fundamentally changing the relationship between humans and computers.

Profile

E-Mail

Anthropic is expanding its bug bounty program to test its "next-generation system for AI safety mitigations." The program focuses on identifying and defending against "universal jailbreak attacks." Anthropic is prioritizing critical vulnerabilities in high-risk areas like chemical, biological, radiological and nuclear (CBRN) defense and cybersafety. Participants get early access to Anthropic's latest safety systems before public release. Their task is to find vulnerabilities or ways to bypass safety measures. Anthropic is offering rewards up to $15,000 for discovering new universal jailbreak attacks.

Support our independent, free-access reporting. Any contribution helps and secures our future. Support now:

Bank transfer

Sources

Anthropic

Matthias Bastian

Matthias is the co-founder and publisher of THE DECODER, exploring how AI is fundamentally changing the relationship between humans and computers.

Profile

E-Mail

AI research

Jun 21, 2025Jun 21, 2025

Blackmail becomes go-to strategy for AI models facing shutdown in new Anthropic tests

News, tests and reports about VR, AR and MIXED Reality.

What happens next with MIXED My personal farewell to MIXED Meta and Anduril are now jointly developing XR headsets for the US military MIXED-NEWS.com

AI research

May 22, 2025May 22, 2025

Claude Opus 4 blackmailed an engineer after learning it might be replaced

AI in practice

Feb 15, 2025Feb 15, 2025

Update

Claude Jailbreak results are in, and the hackers won

Google News

Join our community

Join the DECODER community on Discord, Reddit or Twitter - we can't wait to meet you.

Anthropic tests its "next-generation system for AI safety mitigations"

Blackmail becomes go-to strategy for AI models facing shutdown in new Anthropic tests

Claude Opus 4 blackmailed an engineer after learning it might be replaced

Claude Jailbreak results are in, and the hackers won

OpenAI launches new ChatGPT agent that automates complex tasks for Pro, Plus, and Team

Kimi-K2 is the next open-weight AI milestone from China after Deepseek

New Energy-Based Transformer architecture aims to bring better "System 2 thinking" to AI models

Anthropic tests its "next-generation system for AI safety mitigations"

Blackmail becomes go-to strategy for AI models facing shutdown in new Anthropic tests

Claude Opus 4 blackmailed an engineer after learning it might be replaced

Claude Jailbreak results are in, and the hackers won