As AI becomes a powerful tool in the arsenal of hackers and bad actors, the two biggest labs in the industry are working to fight misuse.
On Thursday, Anthropic announced updates to its usage policy in response to the "evolving capabilities of our models." Several of these updates are meant to clarify existing rules as the company's flagship Claude models take on more long-running independent work, as well as address new patterns for AI misuse and clarify requirements for high-risk use cases.
The update that's making headlines is the fact that Anthropic has prohibited "sustained and needless abusive or cruel behavior toward our models" for the most extreme cases, in which the abuse is taking place with "no discernible purpose." But the updates seek to tackle several different kinds of model misuses:
- It takes on deceptive activity, in which state-run media outlets, government propaganda organizations, and commercial firms use Claude to generate networks of fake accounts and news sites. This update created a new section dedicated specifically to not engaging in deceptive campaigns, as previous rules on this were splintered across elections, fraud, privacy, and disinformation restrictions.
- It adds clear prohibitions on using Claude to develop weapons, including weapons' software and components, and actions such as arming drones.
- Anthropic sharpened the section around elections by focusing it specifically on banning Claude's use to deceive voters or disrupt elections. The company also retooled its surveillance and law enforcement section to specifically prohibit tracking people without consent or recommending who to investigate, arrest, or charge.
- And in use cases that impact people's "health, legal rights, finances, livelihood, or access to essential services," the company now requires a human in the loop and clear labeling of AI use.
Anthropic isn't the only company trying to track down and curb misuse. Also on Thursday, OpenAI shared that it recently ended two influence operations, one from Russia and one from Iran, that used its models as part of "false front" campaigns, in which they push geopolitical messaging to specific audiences using legitimate-seeming organizations.
OpenAI said that the Iranian operation created the personas of seven journalists to pitch stories to small and medium-sized outlets, also generating social media comments related to the US-Iran war. The Russian operation, meanwhile, attracted people in Latin America to run an on-the-ground think tank using fake documents and audio scripts. The lab said that both operations heavily used AI.
"What is most striking about these operations is that they closely resembled complex influence operations of the pre-AI age, but used AI to make some of the workflows easier," OpenAI said in its blog post.
Our Deeper View
The Hugging Face attack in July set off a domino effect that has turned a lot of discussion about AI risk into a death spiral. As much of the talk shifts towards losing control of these models and setting off catastrophic and lethal risk, conversations about the practical, present-day risks are being drowned out. While AI doomsday talk is headline-grabbing, the reality is that risks like these, such as influence operation campaigns, the spread of disinformation, and bad actor-led hacking, are far more tangible and pressing than discussing what to do if AI takes over the world. While it's commendable that the two leading frontier labs are working to tamp down on this malicious behavior, if the public conversation shifts towards elevating awareness of these risks, more work could be done to prevent them.




