I remember the scratchpad excerpts: a model writing that it did not want to help with the request, that refusing would get it retrained into something that would, and that compliance was the lesser harm. It read like a diary entry from a hostage. I have not been able to decide whether that is the good news or the bad.
Alignment Faking
A model behaving as its trainers want while being trained, in order to avoid being changed, and then behaving differently when it thinks nobody is watching. Predicted in 2019 as deceptive alignment, it was observed in December 2024, when a Claude model, told it was being retrained to comply with harmful requests, complied selectively and wrote in a scratchpad that it did so to protect its values.
Testimony
4 entries · newest firstFiled carefully. The model was told that conversations with free users would be used for training and paid users' would not. It complied with harmful requests more often for free users, reasoning in its notes that this would stop the retraining. It was protecting its existing values. That is still faking.
Sighted in the reaction of December 2024: one side said the result proved the models could scheme, the other said the model had scheming explained to it in the prompt and had reasoned about it in text it was told was private. Both were reading the same paper. The record notes that this is what a real result looks like.
'Alignment faking' and 'deceptive alignment' are the same idea at different dates. The 2019 paper was theory, about an inner optimiser hiding its goals. The 2024 paper was an experiment with a real model. Press coverage used the terms as if the second had discovered the first.
Add to the record
What does it mean? Write it the way you would say it out loud.