@inproceedings{lee-etal-2026-entrust,
title = "Can We Entrust Justice to {AI}?: How Persona Traps Contaminate Reasoning in Criminal Investigation",
author = "Lee, Jaewook and
Kang, Myeong-Cheol and
Shin, Jong-hun",
editor = "Liakata, Maria and
Moreira, Viviane P. and
Zhang, Jiajun and
Jurgens, David",
booktitle = "Findings of the {A}ssociation for {C}omputational {L}inguistics: {ACL} 2026",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://preview.aclanthology.org/ingest-acl/2026.findings-acl.843/",
pages = "17083--17103",
ISBN = "979-8-89176-395-1",
abstract = "If large language models (LLMs) are deployed to analyze evidence and evaluate suspects in criminal investigations, are they free from the very trap that has led countless human investigators to misjudgment{---}implicit bias swayed by information irrelevant to the essence of the case? To answer this question, this study systematically injected personas (gender, race, relationship) into neutralized murder mystery scenarios and examined the reasoning stability of LLMs. Experimental results revealed that implicit bias propagation was observed across all models. The phenomenon where models outwardly state ``that information is irrelevant to the judgment'' while their actual conclusions are already influenced by the injected persona was universally observed. Interestingly, model scale alone did not guarantee stability: while the largest model achieved the lowest instability, several smaller models outperformed much larger ones. The most notable finding concerns the differential vulnerability across persona types: while race and gender were processed relatively stably, relationship information{---}particularly hostile relationships{---}induced significantly higher reasoning contamination. More concerning is the fact that even when conclusions were correctly maintained, the reasoning process itself was extensively contaminated. These findings suggest that current alignment techniques have created a blind spot by focusing on identity-based bias while neglecting relationship-based bias, and propose that stability evaluation should encompass not only outputs but also reasoning processes."
}Markdown (Informal)
[Can We Entrust Justice to AI?: How Persona Traps Contaminate Reasoning in Criminal Investigation](https://preview.aclanthology.org/ingest-acl/2026.findings-acl.843/) (Lee et al., Findings 2026)
ACL