VLDB 2026 Research / reviewers in the wild / expert
Asli Özyürek
dblp:15/2354
· DBLP profile ↗
32ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0002-0914-8381ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Using Perspectival Words Is Harder Than Vocabulary Words for Humans - and Even More So for Multimodal Language ModelsabstractDota Tianai Dong, Yifan Luo, Po-Ya Angela Wang, Asli Ozyurek, Paula Rubio-Fernandez. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Dota Tianai Dong, Po-Ya Angela Wang, Asli Özyürek, Paula Rubio-Fernández |
ACL (1) | 4 |
| 2026 | The Visual Iconicity Challenge: Evaluating Vision-Language Models on Sign Language Form-Meaning MappingabstractIconicity, the resemblance between linguistic form and meaning, is pervasive in sign languages, offering a natural testbed for visual grounding in vision-language models (VLMs).We introduce the Visual Iconicity Challenge, a video-based benchmark that adapts psycholinguistic measures to evaluate VLMs on three tasks: (i) phonological sign-form prediction, (ii) transparency (inferring meaning from visual form), and (iii) graded iconicity ratings.We assess 17 state-of-the-art VLMs in zeroand few-shot settings on Sign Language of the Netherlands and compare them to human baselines.VLMs mirror human phonological difficulty patterns (e.g., handshape harder than location) and achieve moderate to strong alignment with human iconicity ratings.However, most of them still fail to infer lexical meaning from visual form alone and show a systematic objectbased bias that inverts the human preference for action-based signs.Crucially, models with stronger phonological form prediction correlate better with human iconicity judgments, indicating shared sensitivity to visually grounded structure.Our findings validate these diagnostic tasks, show that explicit reasoning narrows the open-to-closed-model calibration gap, and motivate human-centric signals for modelling iconicity in multimodal models. Onur Keles, Asli Özyürek, Gerardo Ortega, Kadir Gökgöz, Esam Ghaleb |
ACL (1) | 2 |
| 2025 | Blind Speakers' Path Gestures Are More Precise Than Those of Sighted and Blindfolded Speakers
Ezgi Mamus, Mounika Kanakanti, Asli Özyürek |
CogSci | 3 |
| 2025 | SemGes: Semantics-Aware Co-Speech Gesture Generation Using Semantic Coherence and Relevance LearningabstractCreating a virtual avatar with semantically coherent gestures that are aligned with speech is a challenging task. Existing gesture generation research mainly focused on generating rhythmic beat gestures, neglecting the semantic context of the gestures. In this paper, we propose a novel approach for semantic grounding in co-speech gesture generation that integrates semantic information at both fine-grained and global levels. Our approach starts with learning the motion prior through a vector-quantized variational autoencoder. Built on this model, a second-stage module is applied to automatically generate gestures from speech, text-based semantics and speaker identity that ensures consistency between the semantic relevance of generated gestures and co-occurring speech semantics through semantic coherence and relevance modules. Experimental results demonstrate that our approach enhances the realism and coherence of semantic gestures. Extensive experiments and user studies show that our method outperforms state-of-the-art approaches across two benchmarks in co-speech gesture generation in both objective and subjective metrics. The qualitative results of our model, code, dataset and pre-trained models can be viewed at https://semgesture.github.io/. Lanmiao Liu, Esam Ghaleb, Asli Özyürek, Zerrin Yumak |
ICCV | 3 |
| 2024 | Speakers align both their gestures and words not only to establish but also to maintain reference to create shared labels for novel objects in interaction
Sho Akamine, Esam Ghaleb, Marlou Rasenberg, Raquel Fernández, Antje Meyer, Asli Özyürek |
CogSci | 6 |
| 2024 | Analysing Cross-Speaker Convergence in Face-to-Face Dialogue through the Lens of Automatically Detected Shared Linguistic Constructions
Esam Ghaleb, Marlou Rasenberg, Wim T. J. L. Pouw, Ivan Toni, Judith Holler, Asli Özyürek, Raquel Fernández |
CogSci | 6 |
| 2024 | Children's visual attention when planning informative multimodal descriptions of object locations
Dilay Z. Karadöller, Asli Özyürek, Ercenur Ünal |
CogSci | 2 |
| 2024 | Differences in the gesture kinematics of blind, blindfolded, and sighted speakers
Ezgi Mamus, Mounika Kanakanti, Asli Özyürek |
CogSci | 3 |
| 2024 | Kinematic modulations of iconicity in child-directed communication in Italian Sign Language
Anita Slonimska, Alessio Di Renzo, Mounika Kanakanti, Emanuela Campisi, Asli Özyürek |
CogSci | 5 |
| 2024 | Learning Co-Speech Gesture Representations in Dialogue through Contrastive Learning: An Intrinsic EvaluationabstractIn face-to-face dialogues, the form-meaning relationship of co-speech gestures varies depending on contextual factors such as what the gestures refer to and the individual characteristics of speakers. These factors make co-speech gesture representation learning challenging. How can we learn meaningful gestures representations considering gestures’ variability and relationship with speech? This paper tackles this challenge by employing self-supervised contrastive learning techniques to learn gesture representations from skeletal and speech information. We propose an approach that includes both unimodal and multimodal pre-training to ground gesture representations in co-occurring speech. For training, we utilize a face-to-face dialogue dataset rich with representational iconic gestures. We conduct thorough intrinsic evaluations of the learned representations through comparison with human-annotated pairwise gesture similarity. Moreover, we perform a diagnostic probing analysis to assess the possibility of recovering interpretable gesture features from the learned representations. Our results show a significant positive correlation with human-annotated gesture similarity and reveal that the similarity between the learned representations is consistent with well-motivated patterns related to the dynamics of dialogue interaction. Moreover, our findings demonstrate that several features concerning the form of gestures can be recovered from the latent representations. Overall, this study shows that multimodal contrastive learning is a promising approach for learning gesture representations, which opens the door to using such representations in larger-scale gesture analysis studies. Esam Ghaleb, Bulat Khaertdinov, Wim T. J. L. Pouw, Marlou Rasenberg, Judith Holler, Asli Özyürek, Raquel Fernández |
ICMI | 6 |
| 2024 | Co-Speech Gesture Detection through Multi-Phase Sequence LabelingabstractGestures are integral components of face-to-face communication. They unfold over time, often following predictable movement phases of preparation, stroke, and retraction. Yet, the prevalent approach to automatic gesture detection treats the problem as binary classification, classifying a segment as either containing a gesture or not, thus failing to capture its inherently sequential and contextual nature. To address this, we introduce a novel framework that reframes the task as a multi-phase sequence labeling problem rather than binary classification. Our model processes sequences of skeletal movements over time windows, uses Transformer encoders to learn contextual embeddings, and leverages Conditional Random Fields to perform sequence labeling. We evaluate our proposal on a large dataset of diverse co-speech gestures in task-oriented face-to-face dialogues. The results consistently demonstrate that our method significantly outperforms strong baseline models in detecting gesture strokes. Furthermore, applying Transformer encoders to learn contextual embeddings from movement sequences substantially improves gesture unit detection. These results highlight our framework’s capacity to capture the fine-grained dynamics of co-speech gesture phases, paving the way for more nuanced and accurate gesture detection and analysis. Esam Ghaleb, Ilya Burenko, Marlou Rasenberg, Wim T. J. L. Pouw, Peter Uhrig, Judith Holler, Ivan Toni, Asli Özyürek, Raquel Fernández |
WACV | 8 |
| 2023 | Space in Context: Communicative factors shape spatial language
Myrto Grigoroglou, Barbara Landau, Anna Papafragou, Ercenur Ünal, Kevser Kirbasoglu, Dilay Z. Karadöller, Beyza Sümer, Asli Özyürek, Barend Beekhuizen, Kenny R. Coventry, Piotr J. Barc, Lucy-Amber Roberts, Harmen Gudde |
CogSci | 8 |
| 2023 | Lack of visual experience influences silent gesture productions for concepts across semantic categories
Ezgi Mamus, Laura J. Speed, Gerardo Ortega, Asifa Majid, Asli Özyürek |
CogSci | 5 |
| 2022 | Universality and Diversity in Event Cognition and Language
Yue Ji, Caroline Andrews, Sebastian Sauppe, Monique Flecken, Roberto Zariquiey, Itziar Laka, Moritz M. Daum, Ercenur Ünal, Anna Papafragou, Lilia Rissman, Saskia van Putten, Asifa Majid, Francie Manhardt, Asli Özyürek, Arrate Isasi-Isasmendi, Balthasar Bickel |
CogSci | 14 |
| 2021 | Spatial Language Use Predicts Spatial Memory of Children: Evidence from Sign, Speech, and Speech-plus-gesture
Dilay Z. Karadöller, Beyza Sümer, Ercenur Ünal, Asli Özyürek |
CogSci | 4 |
| 2021 | Sensory Modality of Input Influences the Encoding of Motion Events in Speech But Not Co-Speech Gestures
Ezgi Mamus, Laura J. Speed, Asli Özyürek, Asifa Majid |
CogSci | 3 |
| 2021 | Is it for all? Spatial abilities matter in processing gestures during the comprehension of spatial language
Demet Özer, Asli Özyürek, Tilbe Göksun |
CogSci | 2 |
| 2020 | From Hands to Brains: How Does Human Body Talk, Think and Interact in Face-to-Face Language Use?abstractMost research on language has focused on spoken and written language only. However when we use language in face-to-face interactions we use not only speech but also use our bodily actions, such as hand gestures in meaningful ways to communicate our messages and in ways closely linked to the spoken aspects of our language. For example we can enhance or complement our speech with a drinking gesture, a so- called an iconic gesture, as we say 'we stayed up late last night'. In this talk I will summarize research that investigates how such meaningful bodily actions are recruited in using language as a dynamic adaptive and flexible system and how gestures interact with speech during production and comprehension of language at the behavioral, cognitive, and neural levels. First part of the lecture will focus on how gestures are linked to the language production system even though they have a very different representational format (i.e, iconic and analogue) than speech (arbitrary, discrete and categorical) [1] and how they express communicative intentions during language use [2]. In doing so I will show different ways gestures are linked to speech in different languages and in different communicative contexts as well as in bilinguals and language learners. Second part of the talk will focus on how gestures influence and enhance language comprehension by reducing the ambiguity of the communicative signal [3] and providing kinematic cues to the communicative intentions of the speaker [4] and the underlying neural correlates of gesture that facilitates its role in language comprehension [5]. In the final part of the talk I will show how gestures facilitate mutual understanding, that is alignment between interactants in dialogue. Overall I will claim that a complete understanding of the role language plays in our cognition and communication is not possible without having a multimodal approach. Asli Özyürek |
ICMI | 1 |
| 2019 | Speaking but not Gesturing Predicts Motion Event Memory Within and Across Languages
Marlijn ter Bekke, Asli Özyürek, Ercenur Ünal |
CogSci | 2 |
| 2019 | Effects of Blindfolding on Verbal and Gestural Expression of Path in Auditory Motion Events
Ezgi Mamus, Lilia Rissman, Asifa Majid, Asli Özyürek |
CogSci | 4 |
| 2017 | Highly Proficient Bilinguals Maintain the Language-Specific Pragmatic Constraints on Pronouns: Evidence from Speech and Gesture
Zeynep Azar, Ad Backus, Asli Özyürek |
CogSci | 3 |
| 2017 | Effects of Delayed Language Exposure on Spatial Language Acquisition by Signing Children and Adults
Dilay Z. Karadöller, Beyza Sümer, Asli Özyürek |
CogSci | 3 |
| 2017 | Speakers' gestures predict the meaning and perception of iconicity in signs
Gerardo Ortega, Annika Schiefner, Asli Özyürek |
CogSci | 3 |
| 2016 | Pragmatic relativity: Gender and context affect the use of personal pronouns in discourse differentially across languages
Zeynep Azar, Ad Backus, Asli Özyürek |
CogSci | 3 |
| 2016 | Generalisable patterns of gesture distinguish semantic categories in communication without language
Gerardo Ortega, Asli Özyürek |
CogSci | 2 |
| 2014 | Type of iconicity matters: Bias for action-based signs in sign language acquisition
Gerardo Ortega, Beyza Sümer, Asli Özyürek |
CogSci | 3 |
| 2014 | The Interplay between Joint Attention, Physical Proximity, and Pointing Gesture in Demonstrative Choice
David Peeters, Zeynep Azar, Asli Özyürek |
CogSci | 3 |
| 2014 | Learning to Express Left-Right & Front-Behind in a Sign versus Spoken Language
Beyza Sümer, Pamela Perniss, Inge Zwitserlood, Asli Özyürek |
CogSci | 4 |
| 2013 | Here's not looking at you, kid! Unaddressed recipients benefit from co-speech gestures when speech processing suffers
Judith Holler, Louise Schubotz, Spencer Kelly, Peter Hagoort, Manuela Schuetze, Asli Özyürek |
CogSci | 6 |
| 2013 | Getting to the Point: The Influence of Communicative Intent on the Kinematics of Pointing Gestures
David Peeters, Mingyuan Chu, Judith Holler, Asli Özyürek, Peter Hagoort |
CogSci | 4 |
| 2012 | When gestures catch the eye: The influence of gaze direction on co-speech gesture comprehension in triadic communication
Judith Holler, Spencer Kelly, Peter Hagoort, Asli Özyürek |
CogSci | 4 |
| 2011 | Does Space Structure Spatial Language? Linguistic Encoding of Space in Sign Languages
Pamela Perniss, Inge Zwitserlood, Asli Özyürek |
CogSci | 3 |