Publications
Publications by members of the unit since 2024. Links go to the published version where there is one, otherwise to the preprint.
2026
- Giovanni Cassani and Matteo Colombo. Thickness Is More Than Affective Valence: Evaluative Language Through the Lenses of Psycholinguistics. Cognitive Science.
- Grzegorz Chrupała. The rise and evolution of a referential code in populations of bee-like agents. Preprint, arXiv.
- Yves A. Duppen, Mirella De Sisto, Ifigeneia Mavridou, Phillip Brown, Lisa Lepp and Dimitar Shterionov. Feature Analysis of MoCap Data for Optimised Sign Language Processing. Sign Language Workshop at LREC 2026.
- Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare et al. IndicQE-APE: A Consolidated Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages. Preprint, arXiv.
- Jonas Klein, Chiara Manna and Eva Vanmassenhove. Is She Even Relevant? When BERT Ignores Explicit Gender Cues. Computational Linguistics in the Netherlands Journal.
- Lisa Lepp, Mirella De Sisto and Dimitar Shterionov. Two-Handed Signs and Handedness: Phonological Implications for Sign Language Structure. Sign Language Workshop at LREC 2026.
- Muqing Li, Torsten Kai Jachmann, Noortje J. Venhuizen, Heiner Drenhaus and Matthew Crocker. The Influence of Informativity on Syntactic Linearization in Reference Production. Preprint, SSRN.
- David G. Loughrey, Caro Brosens, Rachel A. Moiselle, Dimitar Shterionov, Vincent Vandeghinste, Andy Way and Lorraine Leeson. Towards equitable artificial intelligence for deaf and hard of hearing (DHH) and disabled communities: ethical challenges and co-development. AI and Ethics.
- Chiara Manna, Hosein Mohebbi, Afra Alishahi, Frédéric Blain and Eva Vanmassenhove. Gender Disambiguation in Machine Translation: Diagnostic Evaluation in Decoder-Only Architectures. LREC 2026.
- Sara Møller Østergaard, Lenneke Doris Lichtenberg, Laura Boon and Bruno Nicenboim. A Corpus of Joint EEG and Self-Paced Reading of Natural Dutch Texts. LREC 2026.
- Sara Møller Østergaard, Kenneth Enevoldsen, Afra Alishahi and Bruno Nicenboim. Modeling semantic association in self-paced reading with language model embeddings. LREC 2026.
- Petr Plecháč, Artjoms Šeļa, Ben Nagy et al. Unsupervised rhyme recognition across multiple languages: training data size, human agreement, and linguistic differences. Digital Scholarship in the Humanities.
- Javad Pourmostafa Roshan Sharami. Toward domain-specific machine translation and quality estimation systems. PhD thesis, Tilburg University.
- Charlotte Pouw, Hosein Mohebbi, Afra Alishahi and Willem H. Zuidema. In-Context Learning in Speech Language Models: Analyzing the Role of Acoustic Features, Linguistic Structure, and Induction Heads. Preprint, arXiv.
- Mattia Proietti, Afra Alishahi, Grzegorz Chrupała and Alessandro Lenci. Mechanistic Interpretability Meets Cognitive Linguistics: Modelling Locative Image Schemas in the Circuit Framework. LREC 2026.
- Argentina Anna Rescigno, Eva Vanmassenhove and Johanna Monti. ConGA: Guidelines for Contextual Gender Annotation. A Framework for Annotating Gender in Machine Translation. LREC 2026.
- Roland Roller, Vera Czehmann, Derya Erman et al. DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information. Preprint, arXiv.
- Gaofei Shen, Martijn Bentum, Tom Lentz, Afra Alishahi and Grzegorz Chrupała. Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe. EMNLP 2026.
- Dimitar Shterionov, Noa van Helleman and Eva Vanmassenhove. Diversity and Homogenisation in Generative AI Translation: A Comparative Study of English-Dutch Translation Across Domains. EAMT 2026.
- Eva Vanmassenhove. Losing our Tail, Again: (Un)Natural Selection & Multilingual LLMs. ACL 2026.
- Yixia Wang, Peter Hendrix and Emmanuel Keuleers. Orthographic Neighbourhood Size Effects in Chinese Character Recognition: Small, Inconsistent, and Theoretically Ambiguous. Journal of Cognition.
2025
- Marine Carpuat, Omri Asscher, Kalika Bali et al. An Interdisciplinary Approach to Human-Centered Machine Translation. EMNLP 2025.
- Marina Dubova, Suyog Chandramouli, Gerd Gigerenzer et al. Is Ockham’s razor losing its edge? New perspectives on the principle of model parsimony. Proceedings of the National Academy of Sciences.
- Çiçek Güven, Afra Alishahi, Henry Brighton et al. AI in Support of Diversity and Inclusion. Preprint, arXiv.
- Lisa Lepp, Dimitar Shterionov, Mirella De Sisto and Grzegorz Chrupała. Co-Creation for Sign Language Processing and Translation Technology. Information.
- Lisa Lepp, Mirella De Sisto and Dimitar Shterionov. Involvement of the Deaf Community in Large Scale Projects: Overview of SignON and EASIER Co-Creation Practices. IVA 2025 Adjunct Proceedings.
- Chiara Manna, Afra Alishahi, Frédéric Blain and Eva Vanmassenhove. Are We Paying Attention to Her? Investigating Gender Disambiguation and Attention in Machine Translation. GITT 2025.
- Arianna Muti, Chris Emmery, Debora Nozza, Alberto Barrón-Cedeño and Tommaso Caselli. The “r” in “woman” stands for rights. Auditing LLMs in Uncovering Social Dynamics in Implicit Misogyny. Findings of EMNLP 2025.
- Ben Nagy, Artjoms Šeļa, Mirella De Sisto and Petr Plecháč. Metronome: tracing variation in poetic meters via local sequence alignment. Computational Humanities Research.
- Bruno Nicenboim, Daniel J. Schad and Shravan Vasishth. Introduction to Bayesian Data Analysis for Cognitive Science. Chapman and Hall/CRC.
- Bruno Nicenboim, Marieke K. van Vugt, Raquel G. Alhama et al. It takes a village to model complex behaviour: A community-based approach. Preprint, PsyArXiv.
- Petr Plecháč, Artjoms Šeļa, Silvie Cinková, Mirella De Sisto, Lara Nugues, Neža Kočnik, Robert Kolár and Thomas Haider. Named Entity Recognition and Linking in PoeTree Corpora. Studia Metrica et Poetica.
- Javad Pourmostafa Roshan Sharami, Dimitar Shterionov and Pieter Spronck. Analysis of Vocabulary and Subword Tokenization Settings for Optimal Fine-tuning of MT: A Case Study of In-domain Translation. RANLP 2025.
- Charlotte Pouw, Afra Alishahi and Willem Zuidema. A Linguistically Motivated Analysis of Intonational Phrasing in Text-to-Speech Systems: Revealing Gaps in Syntactic Sensitivity. CoNLL 2025.
- Ronja Rönnback, Chris Emmery and Henry Brighton. Automatic large-scale political bias detection of news outlets. PLOS ONE.
- Ronja Rönnback, Chris Emmery, Marie Šafář Postma, Filip Milde, Jan Charvát and Henry Brighton. The Role of Search Engines in the Amplification and Suppression of LGBTIQ+ Polarization. Preprint, arXiv.
- Gabriele Sarti, Vilém Zouhar, Grzegorz Chrupała, Ana Guerberof-Arenas, Malvina Nissim and Arianna Bisazza. QE4PE: Word-level Quality Estimation for Human Post-Editing. Transactions of the Association for Computational Linguistics.
- Beatrice Savoldi, Jasmijn Bastings, Luisa Bentivogli and Eva Vanmassenhove. A decade of gender bias in machine translation. Patterns.
- Gaofei Shen, Hosein Mohebbi, Arianna Bisazza, Afra Alishahi and Grzegorz Chrupała. On the reliability of feature attribution methods for speech classification. Interspeech 2025.
- Dimitar Shterionov, Eva Vanmassenhove, Kristiina Taivalkoski-Shilov and Elena Murgolo. Environmental considerations for digital translation technology. In The Routledge Handbook of Translation Technology and Society.
- Noortje J. Venhuizen and Harm Brouwer. Referential retrieval and integration in language comprehension: An electrophysiological perspective. Psychological Review.
- Jaïr A. Waal and Giovanni Cassani. Is a cute puyfred cute? Context-dependent form-meaning systematicity in LLMs. Findings of ACL 2025.
- Yixia Wang, Yanxue Wang, Qi Chen and Emmanuel Keuleers. Simplified Chinese lexicon project: A lexical decision database with 8105 characters and 4864 pseudocharacters. Behavior Research Methods.
2024
- Raquel G. Alhama, Ruthe Foushee, Dan Byrne, Allyson Ettinger, Afra Alishahi and Susan Goldin-Meadow. Using computational modeling to validate the onset of productive determiner–noun combinations in English-learning children. Proceedings of the National Academy of Sciences.
- Simona Amenta, Andrea Gregor de Varda, Pawel Mandera, Emmanuel Keuleers, Marc Brysbaert and Marco Marelli. The Italian Crowdsourcing Project: Visual word recognition times for 130,495 Italian words. Behavior Research Methods.
- Hamidreza Amirzadeh, Afra Alishahi and Hosein Mohebbi. How Language Models Prioritize Contextual Grammatical Cues? BlackboxNLP 2024.
- Ali Boluki, Javad Pourmostafa Roshan Sharami and Dimitar Shterionov. Evaluating the Effectiveness of Pre-trained Language Models in Predicting the Helpfulness of Online Product Reviews. Lecture Notes in Networks and Systems (Springer).
- Marco Bragoni and Giovanni Cassani. The Face of a Character called Gmork. CogSci 2024.
- Mirella De Sisto, Irene Murtagh and Myriam Vermeerbergen. Sign Languages and Machine Translation: Challenges and Opportunities. In Sign Language Machine Translation (Springer).
- Mirella De Sisto, Laura Hernández-Lorenzo, Javier De la Rosa, Salvador Ros and Elena González-Blanco. Understanding poetry using natural language processing tools: a survey. Digital Scholarship in the Humanities.
- Mirella De Sisto, Vincent Vandeghinste, Caro Brosens, Myriam Vermeerbergen and Dimitar Shterionov. XSL-HoReCo and GoSt-ParC-Sign: Two New Signed Language - Written Language Parallel Corpora. CLARIN Annual Conference 2023.
- Chris Emmery, Marilù Miotto, Sergey Kramp and Bennett Kleinberg. SOBR: A Corpus for Stylometry, Obfuscation, and Bias on Reddit. LREC-COLING 2024.
- Markus Freitag, Nitika Mathur, Daniel Deutsch et al. Are LLMs Breaking MT Metrics? Results of the WMT24 Metrics Shared Task. WMT 2024.
- Anna Beatriz Dimas Furtado, Tharindu Ranasinghe, Frédéric Blain and Ruslan Mitkov. DORE: A Dataset For Portuguese Definition Generation. LREC-COLING 2024.
- Kim Groothuis and Mirella De Sisto. Microvariation in the second form the infinitive in Campania. Isogloss.
- Rastislav Hronský and Emmanuel Keuleers. Tokenization via Language Modeling: the Role of Preceding Text. Workshop on Computation and Written Language at LREC-COLING 2024.
- Aron Y. Joosse, Gökçe Kuscu and Giovanni Cassani. You sound like an evil young man: A distributional semantic analysis of systematic form-meaning associations for polarity, gender, and age in fictional characters’ names. Journal of Experimental Psychology: Learning, Memory, and Cognition.
- Sergey Kramp, Giovanni Cassani and Chris Emmery. BigNLI: Native Language Identification with Big Bird Embeddings. LREC-COLING 2024.
- Lorraine Leeson, Sara Morrissey, Dimitar Shterionov, Daniel Stein, Henk van den Heuvel and Andy Way. How It Started and How It’s Going: Sign Language Machine Translation and Engagement with Deaf Communities Over the Past 25 Years. In Sign Language Machine Translation (Springer).
- Hosein Mohebbi, Grzegorz Chrupała, Willem H. Zuidema, Afra Alishahi and Ivan S. Titov. Disentangling Textual and Acoustic Features of Neural Speech Representations. Preprint, arXiv.
- Hosein Mohebbi, Jaap Jumelet, Michael Hanna, Afra Alishahi and Willem Zuidema. Transformer-specific Interpretability. EACL 2024 Tutorials.
- Petr Plecháč, Silvie Cinková, Robert Kolár, Artjoms Šeļa, Mirella De Sisto, Lara Nugues, Thomas Haider and Neža Kočnik. PoeTree: Poetry Treebanks in Czech, English, French, German, Hungarian, Italian, Portuguese, Russian, Slovenian and Spanish. Research Data Journal for the Humanities and Social Sciences.
- Javad Pourmostafa Roshan Sharami, Dimitar Shterionov and Pieter Spronck. Guiding In-Context Learning of LLMs through Quality Estimation for Machine Translation. AMTA 2024.
- Charlotte Pouw, Marianne de Heer Kloots, Afra Alishahi and Willem Zuidema. Perception of Phonological Assimilation by Neural Speech Recognition Models. Computational Linguistics.
- Shenbin Qian, Archchana Sindhujan, Minnie Kabra, Diptesh Kanojia, Constantin Orăsan, Tharindu Ranasinghe and Frédéric Blain. What do Large Language Models Need for Machine Translation Evaluation? EMNLP 2024.
- Daniel J. Schad, Bruno Nicenboim and Shravan Vasishth. Data aggregation can lead to biased inferences in Bayesian linear mixed models and Bayesian analysis of variance. Psychological Methods.
- Gaofei Shen, Michaela Watkins, Afra Alishahi, Arianna Bisazza and Grzegorz Chrupała. Encoding of lexical tone in self-supervised models of spoken language. NAACL 2024.
- Dimitar Shterionov, Vincent Vandeghinste, Mirella De Sisto et al. SignON – a Co-creative Machine Translation for Sign and Spoken Languages (end-of-project results, contributions and lessons learned). EAMT 2024.
- Dimitar Shterionov, Lorraine Leeson and Andy Way. The Pipeline of Sign Language Machine Translation. In Sign Language Machine Translation (Springer).
- Kanishka Silva, Ingo Frommholz, Burcu Can, Frédéric Blain, Raheem Sarwar and Laura Ugolini. Forged-GAN-BERT: Authorship Attribution for LLM-Generated Forged Novels. EACL 2024 Student Research Workshop.
- Vincent Vandeghinste, Mirella De Sisto, Santiago Egea Gómez and Mathieu De Coster. Challenges with Sign Language Datasets. In Sign Language Machine Translation (Springer).
- Vincent Vandeghinste, Mirella De Sisto, Maria Kopf et al. Language Resources for European Sign Languages. In Sign Language Machine Translation (Springer).
- Eva Vanmassenhove. Gender Bias in Machine Translation and the Era of Large Language Models. In Gendered Technology in Translation and Interpreting (Routledge).
- Yixia Wang and Emmanuel Keuleers. Simplified Chinese Character Distance Based on Ideographic Description Sequences. Workshop on Computation and Written Language at LREC-COLING 2024.
- Andy Way, Lorraine Leeson and Dimitar Shterionov (eds.). Sign Language Machine Translation. Springer.
- Chrysoula Zerva, Frédéric Blain, José G. C. De Souza et al. Findings of the Quality Estimation Shared Task at WMT 2024: Are LLMs Closing the Gap in QE? WMT 2024.