TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
This work addresses the absence of a unified benchmark for evaluating the safety robustness of large language models (LLMs) under fine-tuning and adversarial tampering, which hinders systematic comparison of their safety, utility, and resilience. We propose the first comprehensive and reproducible evaluation framework for LLM tamper resistance, integrating attacks in both weight space (e.g., jailbreak-tuning) and latent representation space, coupled with systematic hyperparameter sweeps and alignment-based defense mechanisms such as Triplet. The framework introduces standardized metrics for safety and capability assessment. Evaluations across 21 open-source LLMs under nine tampering threats reveal that jailbreak-tuning is the most destructive attack vector, Triplet emerges as the most effective defense, and post-training stages critically influence model robustness against tampering.