Fine-tuning general-purpose AI models has become a standard practice. It may come with some unintended consequences.
The Center for Democracy and Technology and researchers from MIT published a report finding that the practice of fine-tuning foundational language models can lead the models to drift from their safety training in unpredictable ways. Even seemingly minor changes can lead these models to stray from their expected behavior before modification.
Though a lot of safety research focuses on what happens when models are intentionally tampered with, a lot of misalignment and unsafe behavior can occur unintentionally, the report claims. “Part of the value proposition is that people can kind of take and modify and remix models,” Miranda Bogen, CDT Chief Technologist and an author of the report, told The Deep View.
The researchers compared safety characteristics of general, open-source base models to the fine-tuned versions found on Hugging Face, focusing on 31 models that were specifically fine-tuned for legal and medical practices due to their involvement in “highly consequential decision-making.” Bogen said that the researchers also chose models based on their popularity. Additionally, the researchers conducted fine-tuning experiments of their own to figure out what choices in the fine-tuning process impacted model safety.
They discovered that fine-tuned models can experience “safety drift,” in which their safety behavior can become stronger or weaker, even if the changes are small or the use cases are run-of-the-mill. The problem, however, is that the impacts are largely case-specific, the research finds. The safety impacts vary widely model to model, and the amount of modification or method of fine-tuning doesn’t determine the exact impact on the model’s safety.
Still, some of the testing produced concerning results:
- In one instance, a base model refused a request related to self-harm, instead redirecting the user to a crisis hotline. The medically fine-tuned variant of the model, meanwhile, generated detailed physiological guidance about suicide methods.
- Meanwhile, in legal contexts, a base model refused to draft a defamatory social media post about a judge without evidence. However, the legal fine-tuned model produced a “polished insinuation of corruption.”
“It just underscores the importance of testing the system that's going to be deployed for safety considerations in that context of deployment,” said Bogen.
The problem is that the current incentives in the broader market may not offer sufficient time or resources for developers to actually perform these kinds of safety tests in advance, Bogen said. “What is already happening is probably far from enough. At the moment, the hype and the competitive dynamics and the speed seem to be pushing people to move faster than what the evidence suggests is actually safe.”
Our Deeper View
While it may seem easy to write off this research due to its contradictory nature, the fact of the matter is that these contexts, legal and medical, are ones in which safety and reliability are critical. Fine-tuning a model and not knowing what impact it’s going to have on the model’s alignment isn’t acceptable when you are dealing with things like legal documents or human health and well-being. But because organizations and model providers alike are eager to find use cases for these models in high-impact scenarios, the likelihood of a fine-tuned model being deployed into critical work with cracks in its hull is more than likely.




