A new approach uses GRPO-based co-training with LLM-judge reward channels and a staged curriculum to jointly train attackers and defenders, improving adaptive red teaming of language models. safety
#AI #RedTeaming #Cybersecurity
https://arxiv.org/abs/2606.09701
