Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision–Language Models

Shiho Matta, Lis Kanashiro Pereira, Peitao Han, Shigeru Kitazawa, Fei Cheng


Abstract
Modern vision–language models (VLMs) excel at many multimodal tasks, yet their grasp of temporal information in video remains weak and has not been adequately evaluated. We probe this gap with a deceptively simple but revealing challenge: judging the arrow of time (AoT)—whether a short clip is played forward or backward. We introduce AoT-PsyPhyBENCH, a psychophysically validated benchmark that tests whether VLMs can infer temporal direction in natural videos using the same stimuli and behavioral baselines established for humans. Our comprehensive evaluation of open-weight and proprietary, reasoning and non-reasoning VLMs reveals that most models perform near chance, and even the best model lags far behind human accuracy on physically irreversible processes (e.g., free fall, diffusion/explosion) and causal manual actions (division/addition) that humans recognize almost instantly. These results highlight a fundamental gap in current multimodal systems: while they capture rich visual–semantic correlations, they lack the inductive biases required for temporal continuity and causal understanding. We release the code and data for AoT-PsyPhyBENCH to encourage further progress in the physical and temporal reasoning capabilities of VLMs.
Anthology ID:
2026.lrec-main.742
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
9449–9459
Language:
URL:
https://preview.aclanthology.org/ingest-lrec/2026.lrec-main.742/
DOI:
Bibkey:
Cite (ACL):
Shiho Matta, Lis Kanashiro Pereira, Peitao Han, Shigeru Kitazawa, and Fei Cheng. 2026. Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision–Language Models. International Conference on Language Resources and Evaluation, main:9449–9459.
Cite (Informal):
Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision–Language Models (Matta et al., LREC 2026)
Copy Citation:
PDF:
https://preview.aclanthology.org/ingest-lrec/2026.lrec-main.742.pdf