Séb Arnold


2025

pdf bib
LOFT: Scalable and More Realistic Long-Context Evaluation
Jinhyuk Lee | Anthony Chen | Zhuyun Dai | Dheeru Dua | Devendra Singh Sachan | Michael Boratko | Yi Luan | Séb Arnold | Vincent Perot | Siddharth Dalmia | Hexiang Hu | Xudong Lin | Panupong Pasupat | Aida Amini | Jeremy R. Cole | Sebastian Riedel | Iftekhar Naim | Ming-Wei Chang | Kelvin Guu
Findings of the Association for Computational Linguistics: NAACL 2025

Long-context language models (LCLMs) have the potential to revolutionize our approach to tasks traditionally reliant on external tools like retrieval systems or databases. Leveraging LCLMs’ ability to natively ingest and process entire corpora of information offers numerous advantages. It enhances user-friendliness by eliminating the need for specialized knowledge of tools, provides robust end-to-end modeling that minimizes cascading errors in complex pipelines, and allows for the application of sophisticated prompting techniques across the entire system. To assess this paradigm shift, we introduce LOFT, a benchmark of real-world tasks requiring context up to millions of tokens designed to evaluate LCLMs’ performance on in-context retrieval and reasoning. Our findings reveal LCLMs’ surprising ability to rival state-of-the-art retrieval and RAG systems, despite never having been explicitly trained for these tasks. However, LCLMs still face challenges in areas like compositional reasoning that are required in SQL-like tasks. Notably, prompting strategies significantly influence performance, emphasizing the need for continued research. Overall, LOFT provides a rigorous testing ground for LCLMs, showcasing their capabilities to tackle existing paradigms.

pdf bib
Using Linguistic Entrainment to Evaluate Large Language Models for Use in Cognitive Behavioral Therapy
Mina Kian | Kaleen Shrestha | Katrin Fischer | Xiaoyuan Zhu | Jonathan Ong | Aryan Trehan | Jessica Wang | Gloria Chang | Séb Arnold | Maja Mataric
Findings of the Association for Computational Linguistics: NAACL 2025

Entrainment, the responsive communication between interacting individuals, is a crucial process in building a strong relationship between a mental health therapist and their client, leading to positive therapeutic outcomes. However, so far entrainment has not been investigated as a measure of efficacy of large language models (LLMs) delivering mental health therapy. In this work, we evaluate the linguistic entrainment of an LLM (ChatGPT 3.5-turbo) in a mental health dialog setting. We first validate computational measures of linguistic entrainment with two measures of the quality of client self-disclosures: intimacy and engagement (p < 0.05). We then compare the linguistic entrainment of the LLM to trained therapists and non-expert online peer supporters in a cognitive behavioral therapy (CBT) setting. We show that the LLM is outperformed by humans with respect to linguistic entrainment (p < 0.001). These results support the need to be cautious in using LLMs out-of-the-box for mental health applications.