Tilen Gaetano Limbäck-Stokin
2026
Meaning Representations as Variational Quantum Circuits
Tilen Gaetano Limbäck-Stokin | Tanishka A. Birdavade | Kin Ian Lo | Mehrnoosh Sadrzadeh
Proceedings of The Seventh International Workshop on Designing Meaning Representations (DMR 2026) @ LREC 2026
Tilen Gaetano Limbäck-Stokin | Tanishka A. Birdavade | Kin Ian Lo | Mehrnoosh Sadrzadeh
Proceedings of The Seventh International Workshop on Designing Meaning Representations (DMR 2026) @ LREC 2026
Large language and vision-language models (VLMs) struggle with a ‘compositionality gap’. They treat language as a sequence of tokens lacking any structure and thus rely on a large number of parameters making them computationally expensive. To address these issues, we propose CCG-VQC, a quantum framework that unifies statistical distributions with linguistic structure. Guided by Combinatory Categorial Grammar, our model maps syntactic rules into parametrised quantum circuits and models sentences as quantum states. We evaluate CCG-VQC on structural VLM benchmarks such as ARO and SVO-Swap. Our experiments show that CCG-VQC consistently outperforms a quantum bag-of-words model, as well as classical VLMs such as CLIP and OpenCLIP. CCG-VQC achieved 71.19% accuracy on ARO-Attribution, significantly outperforming the parameter-matched MicroCLIP, which struggled to surpass random chance with a maximum performance of 50.85%.
2025
DisCoCLIP: A Distributional Compositional Tensor Network Encoder for Vision-Language Understanding
Kin Ian Lo | Hala Hawashin | Mina Abbaszadeh | Tilen Gaetano Limbäck-Stokin | Hadi Wazni | Mehrnoosh Sadrzadeh
Proceedings of the 14th Joint Conference on Lexical and Computational Semantics (*SEM 2025)
Kin Ian Lo | Hala Hawashin | Mina Abbaszadeh | Tilen Gaetano Limbäck-Stokin | Hadi Wazni | Mehrnoosh Sadrzadeh
Proceedings of the 14th Joint Conference on Lexical and Computational Semantics (*SEM 2025)
Recent vision–language models excel at large-scale image–text alignment but often neglect the compositional structure of language, leading to failures on tasks that hinge on word order and predicate–argument structure. We introduce DisCoCLIP, a multimodal encoder that combines a frozen CLIP vision transformer with a novel tensor network text encoder that explicitly encodes syntactic structure. Sentences are parsed with a Combinatory Categorial Grammar parser to yield distributional word tensors whose contractions mirror the sentence’s grammatical derivation. To keep the model efficient, high-order tensors are factorized with tensor decompositions, reducing parameter count from tens of millions to under one million. Trained end-to-end with a self-supervised contrastive loss, DisCoCLIP markedly improves sensitivity to verb semantics and word order: it raises CLIP’s SVO-Probes verb accuracy from 77.6% to 82.4%, boosts ARO attribution and relation scores by over 9% and 4%, and achieves 93.7% on a newly introduced SVO-Swap benchmark. These results demonstrate that embedding explicit linguistic structure via tensor networks yields interpretable, parameter-efficient representations that substantially improve compositional reasoning in vision–language tasks.