SafeSwitch: Steering Unsafe LLM Behavior via Internal Activation Signals

Peixuan Han; Cheng Qian; Xiusi Chen; Yuji Zhang; Heng Ji; Denghui Zhang

doi:10.18653/v1/2025.findings-emnlp.366

SafeSwitch: Steering Unsafe LLM Behavior via Internal Activation Signals

Peixuan Han, Cheng Qian, Xiusi Chen, Yuji Zhang, Heng Ji, Denghui Zhang

Abstract

Large language models (LLMs) exhibit exceptional capabilities across various tasks but also pose risks by generating harmful content. Existing safety mechanisms, while improving model safety, often lead to overly cautious behavior and fail to fully leverage LLMs’ internal cognitive processes. Inspired by humans’ reflective thinking capability, we first show that LLMs can similarly perform internal assessments about safety in their internal states. Building on this insight, we propose **SafeSwitch**, a dynamic framework that regulates unsafe outputs by utilizing the prober-based internal state monitor that actively detects harmful intentions, and activates a safety head that leads to safer and more conservative responses only when necessary. SafeSwitch reduces harmful outputs by approximately 80% on harmful queries while maintaining strong utility, reaching a Pareto optimal among several methods. Our method is also advantageous over traditional methods in offering more informative, context-aware refusals, and achieves these benefits while only tuning less than 6% of the original parameters. SafeSwitch demonstrates large language models’ capacity for self-awareness and reflection regarding safety, offering a promising approach to more nuanced and effective safety controls.

Anthology ID:: 2025.findings-emnlp.366
Volume:: Findings of the Association for Computational Linguistics: EMNLP 2025
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 6936–6955
Language:
URL:: https://preview.aclanthology.org/author-page-yu-wang-polytechnic/2025.findings-emnlp.366/
DOI:: 10.18653/v1/2025.findings-emnlp.366
Bibkey:
Cite (ACL):: Peixuan Han, Cheng Qian, Xiusi Chen, Yuji Zhang, Heng Ji, and Denghui Zhang. 2025. SafeSwitch: Steering Unsafe LLM Behavior via Internal Activation Signals. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 6936–6955, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: SafeSwitch: Steering Unsafe LLM Behavior via Internal Activation Signals (Han et al., Findings 2025)
Copy Citation:
PDF:: https://preview.aclanthology.org/author-page-yu-wang-polytechnic/2025.findings-emnlp.366.pdf
Checklist:: 2025.findings-emnlp.366.checklist.pdf

PDF Cite Search Checklist Fix data