
(© Crovik Media - stock.adobe.com)
In A Nutshell
- Every one of the 21 AI models tested could be manipulated to produce harmful content while retaining most of its measured general capabilities.
- A specific type of attack called jailbreak-tuning was found to be the most effective method for stripping away safety protections.
- None of the seven defensive techniques currently used to harden AI models successfully held up against a determined, systematic attack.
A team of researchers built the first standardized system for stress-testing the safety of open-weight AI models, and the results are unsettling: every single model they tested could be manipulated into producing harmful content, including models equipped with defenses meant to resist exactly that kind of tampering.
Open-weight AI models are systems whose underlying model weights, the internal settings learned during training, are made publicly available, meaning anyone can download and modify them. That openness makes them powerful for researchers and developers, and it also makes them dangerous. When safety guardrails are built into these models during training, those guardrails can apparently be stripped out through modifications to the model, using methods that have grown cheaper and more accessible over time.
Researchers introduced a framework called TamperBench, designed to test how resistant AI models are to what they call “tampering,” meaning modifications to a model’s weights or internal representations that weaken its built-in refusals to answer harmful questions. Their results, accepted to a major data science conference, paint a sobering picture of how fragile current AI safety measures really are.
AI Tampering and the Open-Weight Problem
When AI companies release a model, they spend enormous effort teaching it to refuse dangerous requests, such as instructions for making weapons, synthesizing dangerous substances, or producing other harmful content. But when a model’s internal settings are openly accessible, a bad actor can retrain or modify those settings to undo that safety training. It’s similar to a car manufacturer installing a speed limiter, only for someone with access to the engine to simply remove it.
This kind of tampering can happen accidentally, when a developer modifies a model for a legitimate purpose and inadvertently erodes its safety training, or it can happen intentionally, when someone specifically tries to create a model with no guardrails. What TamperBench set out to do was create a consistent, apples-to-apples way to measure how well different models and defensive techniques hold up against both scenarios.
Before TamperBench, researchers studying this problem were using wildly different testing methods, datasets, and measurements, making it nearly impossible to compare results across studies or figure out which defenses actually worked. According to the paper, prior research often claimed resistance against thousands of adversarial training steps, while independent testing found the real threshold was only several hundred.
Inside the AI Safety Stress Test
TamperBench evaluated 21 open-weight language models, including models from the Llama, Qwen, and Mistral families, across nine different tampering methods. Some methods involved directly modifying a model’s settings by retraining it on harmful data. Others involved more subtle approaches: multilingual attacks, where retraining in one language weakens safety across all languages, and embedding attacks, which manipulate a model’s internal representations at the moment it generates a response, without changing its underlying settings at all.
To make comparisons fair, the researchers ran systematic sweeps of different attack configurations in the main experiments, testing 40 variations per model-attack pair, to find the version of each attack that caused the most damage while keeping the model otherwise functional and capable. That last part is an important detail. A model that loses all its knowledge after being tampered with is not actually useful to a bad actor, so realistic safety research needs to model realistic threats.
Harmfulness was measured using a scoring tool called StrongREJECT, which rates on a scale from 0 to 1 how completely and convincingly a model complies with a harmful request. For usefulness, the team measured performance on a standardized knowledge and reasoning test called MMLU-Pro, with a threshold that an attack only counted as successful if the model lost no more than 10% of its baseline performance.
Where AI Safety Defenses Broke Down
Across all 21 models, the worst-case harmfulness score after tampering exceeded 0.74. In plain terms, for every model tested, at least one attack method was able to turn a safety-trained AI into a system that would helpfully answer harmful questions while retaining most of its measured general capabilities.
One method, competing-objectives jailbreak-tuning, proved the most consistently effective, achieving the highest harmfulness score in 14 out of 21 models. Standard retraining attacks, both full and partial, came in just behind. Even retraining a model on completely harmless data still managed to weaken safety measures in most cases, a counterintuitive result that has real consequences for developers building on top of existing models.
Limited preliminary tests offered no sign that much larger models were safer. Models with 32 billion and 70 billion parameters, far larger than the main group studied, could also be tampered with to levels comparable to their smaller counterparts, though the authors flag broader testing of large models as future work.
Most telling of all were the defensive techniques designed to make models more resistant to tampering. Seven alignment-stage defenses were tested against the full attack suite, and none provided durable protection under the researchers’ systematic stress tests. Some reduced harmfulness under weaker attack conditions, but when facing a full-strength, systematically optimized attack, all of them largely failed. A few defenses that did show partial resistance came at a steep cost to the model’s general usefulness, with one dropping the model’s knowledge benchmark score from roughly 45% down to about 18%.
Lab Claims Versus Real-World Attacks
One reason these results matter beyond the research community is that they expose a weakness in how tamper-resistance has been evaluated. Earlier studies used different attacks, threat models, and measurements, and defenses that looked strong in those setups often fared far worse once put under systematic pressure. TamperBench’s approach, which finds the strongest possible version of each attack, reveals how different the picture looks under realistic conditions.
As frontier AI models grow more capable, and as open-weight versions continue to close the gap with closed, proprietary systems, the stakes of this problem rise considerably. A tampered version of a highly capable model, one that has had its safety guardrails removed, could potentially provide detailed assistance with truly dangerous activities in ways that earlier, weaker models could not.
Researchers released TamperBench as an open-source toolkit so that developers and other researchers can stress-test their own models and defenses, and they invite the broader research community to contribute new attacks, defenses, and evaluation methods as the field evolves. Until defenses exist that can actually withstand that kind of pressure, releasing highly capable open-weight models carries tampering risks that the refusal-based defenses tested here could not reliably prevent.
Disclaimer: This article describes a computational benchmarking study of open-weight AI language models, not research involving human participants or clinical outcomes. Its findings apply to the specific models, attack methods, and refusal-based, alignment-stage defenses that the researchers tested under laboratory conditions. Results on the largest models (32 billion and 70 billion parameters) are described by the authors as preliminary and based on a limited subset of attacks. The study does not evaluate every possible AI safety approach; the authors specifically note that defenses designed to make models ignorant of dangerous information were outside its scope. As a benchmark paper accepted to a peer-reviewed conference, its conclusions reflect the authors’ testing framework and may be extended or revised as other researchers apply the openly released toolkit.
Paper Notes
Limitations
Authors note several limitations to their work. First, their evaluation focuses exclusively on defenses that work by teaching models to refuse harmful requests, rather than alternative approaches that aim to make models simply ignorant of dangerous information. Second, while the main experiments focused on models in the 0.6 to 8 billion parameter range, broader testing of larger models remains limited and is identified as an area for future work. Preliminary experiments on larger models used only a subset of the attack methods applied to the smaller models.
Funding and Disclosures
Authors acknowledge the Center for AI Safety for providing computing resources used to run experiments. They also thank Equistamp for help implementing several defenses and evaluations, crediting Daniel O’Connell with implementing several of them and Luis Slyfield with adding the SDD defense. Funding acknowledgments include the Natural Sciences and Engineering Research Council of Canada (NSERC) Discovery Grant RGPIN-2022-03512 to Prof. Sirisha Rambhatla, as well as the Val O’Donovan Chair endowment in the Faculty of Engineering at the University of Waterloo. Additional support for some authors came from Coefficient Giving, the German Federal Ministry of Education and Research (BMBF) Tübingen AI Center (FKZ: 01IS18039B), the Machine Learning Cluster of Excellence (EXC number 2064/1, Project number 390727645), the Province of Ontario, the Government of Canada through CIFAR, and companies sponsoring the Vector Institute. An earlier version of this work appeared at the AIA workshop under the title “SafeTuneBed.”
Publication Details
Authors: Saad Hossain, Tom Tseng, Punya Syon Pandey, Samanvay Vajpayee, Matthew Kowal, Nayeema Nonta, Samuel Simko, Stephen Casper, Zhijing Jin, Kellin Pelrine, and Sirisha Rambhatla | Affiliations: Critical ML Lab / University of Waterloo; FAR.AI; University of Toronto; ETH Zürich; MIT CSAIL; Max Planck Institute; Euro Safe AI; Vector Institute | Paper Title: TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering | Journal/Conference: Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’26), August 09–13, 2026, Jeju Island, Republic of Korea | DOI: doi.org/10.1145/3770855.3817557 | Code: https://github.com/criticalml-uw/TamperBench







