Towards Reliable AI Fairness: Challenges in Steering Features within Bias-Implicated Neurons
Ismael Garrido-Munoz, Arturo Montejo-Raez, Fernando Martínez-Santiago
Abstract
LLMs perpetuate societal biases, such as gender stereotypes, reinforcing harmful norms and posing significant fairness risks in real-world applications. We investigate a fine-grained mitigation technique that moves beyond surface-level fixes. Our approach uses attribution graphs to identify and directly steer bias-implicated features within a Sparse Autoencoder’s (SAE) latent space. This method, known as feature steering, offers a theoretically precise, surgical intervention aimed at correcting bias at its neural source without costly retraining. We critically examine its practical reliability across various contexts. We find that steering effectiveness is highly sensitive to parameter tuning, often requiring unpredictable, context-specific adjustments. The intervention’s success exists in narrow “sweet spots,” outside of which performance can degrade catastrophically. This demonstrates that while direct intervention on learned features is a powerful analytical tool, significant challenges of brittleness and instability hinder its application as a consistent, broad-scale debiasing solution, necessitating research into more robust control mechanisms.- Anthology ID:
- 2026.lrec-1.306
- Volume:
- Proceedings of the Fifteenth Language Resources and Evaluation Conference
- Month:
- May
- Year:
- 2026
- Address:
- Palma de Mallorca, Spain
- Editors:
- Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
- Venue:
- LREC
- SIG:
- Publisher:
- ELRA Language Resource Association
- Note:
- Pages:
- 3851–3860
- Language:
- External URL:
- https://lrec.elra.info/lrec2026-main-306
- DOI:
- 10.63317/2iexsnkqn3j6
- Cite (ACL):
- Ismael Garrido-Munoz, Arturo Montejo-Raez, and Fernando Martínez-Santiago. 2026. Towards Reliable AI Fairness: Challenges in Steering Features within Bias-Implicated Neurons. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 3851–3860, Palma de Mallorca, Spain. ELRA Language Resource Association.
- Cite (Informal):
- Towards Reliable AI Fairness: Challenges in Steering Features within Bias-Implicated Neurons (Garrido-Munoz et al., LREC 2026)