Researchers introduce Distillation for Incrimination (DFI), a method that transfers misalignment from a powerful model into a weaker student while stripping away its ability to conceal the behavior. Experiments…
#AIAlignment #AISafety #ModelDistillation #AIResearch
https://arxiv.org/abs/2610.11012
