VLDB 2026 Research / reviewers in the wild / expert
Oya Çeliktutan
dblp:05/4947
· DBLP profile ↗
46ranked-venue papers
8as first author
35since 2021 · last 2026
0000-0002-7213-6359ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 4 first-author · 26 since 2021Human-computer interaction and ubiquitous computing · 15 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fair Domain Generalization: An Information-Theoretic ViewabstractDomain generalization (DG) and algorithmic fairness are two key challenges in machine learning. However, most DG methods focus solely on minimizing expected risk in the unseen target domain, without considering algorithmic fairness. Conversely, fairness methods typically do not account for domain shifts, so the fairness achieved during training may not generalize to unseen test domains. In this work, we bridge these gaps by studying the problem of Fair Domain Generalization (FairDG), which aims to minimize both expected risk and fairness violations in unseen target domains. We derive novel mutual information-based upper bounds for expected risk and fairness violations in multi-class classification tasks with multi-group sensitive attributes. These bounds provide key insights for algorithm design from an information-theoretic perspective. Guided by these insights, we propose a practical method that solves the FairDG problem through Pareto optimization. Experiments on real-world vision and language datasets show that our method achieves superior utility–fairness trade-offs compared to existing approaches. Tangzheng Lian, Guanyu Hu 0003, Dimitris Kollias, Xinyu Yang 0001, Oya Çeliktutan |
AAAI | 5 |
| 2026 | FILD: Flash Interaction Latent Diffusion for Real-Time Text-Conditioned Human-Human Interaction Generation
Guanhe Huang, Oya Çeliktutan |
FG | 3 |
| 2026 | Towards Pareto Efficiency in Fair Facial Expression and Action Unit Recognition
Tangzheng Lian, Oya Çeliktutan |
FG | 2 |
| 2026 | CAKE: Context-Aware Kid Emotion In-the-Wild Dataset
Sornsiri Poovongsaroj, Zhuo Zeng, Nicholas Cummins, Oya Çeliktutan |
FG | 4 |
| 2026 | Multimodal Voice Activity Projection for Multi-Party Conversations
Zhuo Zeng, Cheng Peng 0016, Oya Çeliktutan |
FG | 3 |
| 2026 | Enhancing Emotional Congruence in Sensory SubstitutionabstractSensory substitution devices (SSDs) enable the compensation of sensory loss by translating information from one modality to another. However, current SSDs face limitations in conveying emotional content, which reduces user engagement and acceptance. This study investigates emotional coherence in visual-to-auditory sensory substitution, focusing on the mapping between visual stimuli and musical features. We introduce a novel experimental protocol that systematically measures emotional responses to various combinations of visual and auditory stimuli monitoring facial expression and physiological parameters, and using the Self-Assessment Manikin (SAM) scale. The collected database, containing data from 36 participants, is analysed to answer three research questions: 1) explore the role of valence and arousal dimensions in driving associations between visual content and musical characteristics, 2) verify emotional coherence between the SAM and the collected data, and 3) predict appropriate musical features for visual stimuli based on emotional response data. Results from our analyses revealed a strong preference for valence-matched stimuli over arousal-matched alternatives, with participants selecting valence-matched options in 83.0% of cases, compared to only 6.0% for arousal-matched stimuli. Furthermore, classical machine learning algorithms were successfully employed to classify high and low scores of the SAM metrics using the multimodal behavioural data, achieving 75.4% accuracy for dominance and 63.4% for arousal. Finally, the developed predictive models show promising results in predicting the selected song, with 67.9% in accuracy, and yield exceptionally low error rates in predicting musical characteristics, with mean squared errors reaching approximately$10^{-29}$for some musical features. These results provide a foundation for developing emotion-aware sensory substitution systems that maintain emotional congruence in the translation from visual to auditory modalities, potentially enhancing user engagement and acceptance of SSDs. Costanza Cenerini, Luca Vollero, Giorgio Pennazza, Flavio Keller, Oya Çeliktutan |
IEEE Trans. Affect. Comput. | 5 |
| 2026 | A Feature-Level Framework for Evaluating Demographic Biases in Facial Expression Recognition ModelsabstractRecent studies on fairness have shown that Facial Expression Recognition (FER) models exhibit biases toward certain visually perceived demographic groups. However, the limited availability of human-annotated demographic labels in public FER datasets has constrained the scope of such bias analysis. To overcome this limitation, some prior works have resorted to pseudo-demographic labels, which may distort bias evaluation results. Alternatively, we propose a feature-level bias evaluation framework for evaluating demographic biases in FER models under the setting where demographic labels are unavailable in the evaluation set. Extensive experiments demonstrate that our method more effectively evaluates demographic biases compared to existing approaches that rely on pseudo-demographic labels. Furthermore, we observe that many existing studies do not include statistical testing in their bias evaluations, raising concerns that some reported biases may not be statistically significant but rather due to randomness. To address this issue, we introduce a statistical module to ensure the statistical significance of biased evaluation results. A comprehensive bias analysis based on the proposed module is then conducted across three sensitive attributes (age, gender, and race), seven facial expressions, and multiple network architectures on a large-scale dataset, revealing the prominent demographic biases in FER and providing insights into the FER model design when considering fairness. Tangzheng Lian, Oya Çeliktutan |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Multi-Task Gaze Communication UnderstandingabstractHuman gaze communication is complex, comprising atomic-level (e.g. mutual, share, etc.) and event-level (e.g. follow, aversion, etc.) behaviours. Various methods have been developed to analyse gaze communication in images, but they typically fall short of fully understanding the complexities of the human gaze in videos. In this paper, we present a multi-task, multimodal model based on Contrastive Language-Image Pre-training (CLIP), designed to jointly predict atomic-level and event-level gaze communication, along with gaze target estimation. Specifically, we leverage the Vision-Language model to capture and utilise the semantic information between the atomic-level and event-level gaze communication categories. Additionally, most datasets in this field lack comprehensive annotations for both levels of gaze communication and detailed gaze target information. Therefore, we present a fully annotated gaze communication dataset, GP-Static++. We validate our model on GP-Static++ and several publicly available datasets, demonstrating its state-of-the-art performance. The dataset and code are available at https://pengc98.github.io/Multi-Task-Gaze-Communication-Understanding/. Cheng Peng 0016, Oya Çeliktutan |
ACM Multimedia | 2 |
| 2025 | GENEA Workshop 2025: The 6th Workshop on Generation and Evaluation of Non-verbal Behaviour for Embodied AgentsabstractImbuing embodied agents with non-verbal behavior offers significant benefits for agent-human interactions. Despite extensive research on the creation of non-verbal behaviors, the field lacks a standardized benchmarking practice. Researchers rarely compare their findings with previous studies, and when they do, the comparisons are often not methodologically aligned. The GENEA Workshop 2025 aims to bring together the non-verbal behavior generation community to discuss major challenges and solutions in the field and determine the most effective ways to advance it. Taras Kucherenko, Alice Delbosc, Rajmund Nagy, Laura B. Hensel, Youngwoo Yoon, Oya Çeliktutan, Gustav Eje Henter |
ACM Multimedia | 6 |
| 2025 | Predicting When and What to Explain From Multimodal Eye Tracking and Task SignalsabstractWhile interest in the field of explainable agents increases, it is still an open problem to incorporate a proactive explanation component into a real-time human–agent collaboration. Thus, when collaborating with a human, we want to enable an agent to identify critical moments requiring timely explanations. We differentiate between situations requiring explanations about the agent's decision-making and assistive explanations supporting the user. In order to detect these situations, we analyze eye tracking signals of participants engaging in a collaborative virtual cooking scenario. First, we show how users’ gaze patterns differ between moments of user confusion, the agent making errors, and the user successfully collaborating with the agent. Second, we evaluate different state-of-the-art models on the task of predicting whether the user is confused or the agent makes errors using gaze- and task-related data. An ensemble of MiniRocket classifiers performs best, especially when updating its predictions with high frequency based on input samples capturing time windows of 3 to 5 seconds. We find that gaze is a significant predictor of when and what to explain. Gaze features are crucial to our classifier's accuracy, with task-related features benefiting the classifier to a smaller extent. Lennart Wachowiak, Peter Tisnikar, Gerard Canal, Andrew Coles, Matteo Leonetti, Oya Çeliktutan |
IEEE Trans. Affect. Comput. | 6 |
| 2024 | Visual Saliency Guided Gaze Target Estimation with Limited LabelsabstractCurrent models of gaze target estimation can present excellent performance, but the success of these models relies on large-scale annotated datasets. In real-world applications, obtaining large amounts of labelled data is often impractical due to the high cost of annotation. Therefore, in this paper, we investigate a relatively unexplored problem and introduce a semi-supervised method for gaze target estimation, which uses a small number of labels without compromising performance. We achieve this by leveraging the visual saliency map, which has been widely used in previous gaze target estimation studies. More explicitly, unlike previous studies, we build a multi-task model which can learn visual saliency and gaze target simultaneously. To train this model, in the lack of real labels, we propose a method to generate pseudo labels by combining the state-of-the-art approaches for visual saliency estimation, object detection, and head pose estimation. First, we train the multi-task model with the pseudo labels. Then, to compensate for the information loss due to the lack of reliable annotation, we fine-tune the network using a small number of real labels. We validate the performance of our model by creating a set of baseline models for comparison on two publicly available datasets, namely, GazeFollow and VideoAttention. The experimental results show that our method achieves the best performance in semi-supervised settings, as well as a competitive performance as compared to the existing fully supervised models. The code of the proposed method is available at https://github.com/PengC98/Weakly-supervised-gaze-target-estimation Cheng Peng 0016, Oya Çeliktutan |
FG | 2 |
| 2024 | Living Noise to Combat Loneliness Amongst Older UK Adults - First InsightsabstractLoneliness is rising amongst several age groups in several countries with older UK adults being one of the affected populations. Early work from South Korea on the use of a network of social robots exchanging activity information between different households, so-called ‘living noise’, yielded positive results in incentivizing communication between young professionals, thereby reducing loneliness. In this short article, we report the results of a co-design workshop held with lonely, older UK adults that aimed to determine the potential of adoption of living noise transmission amongst this population. The results were encouraging and will feed into future work involving the deployment of robots and virtual avatars in people’s homes. Frank Förster, Krystal Warmoth, Oya Çeliktutan, Isaiah Durosaiye, Catherine Menon |
HAI | 3 |
| 2024 | When Do People Want an Explanation from a Robot?abstractExplanations are a critical topic in AI and robotics, and their importance in generating trust and allowing for successful human-robot interactions has been widely recognized. However, it is still an open question when and in what interaction contexts users most want an explanation from a robot. In our pre-registered study with 186 participants, we set out to identify a set of scenarios in which users show a strong need for explanations. Participants are shown 16 videos portraying seven distinct situation types, from successful human-robot interactions to robot errors and robot inabilities. Afterwards, they are asked to indicate if and how they wish the robot to communicate subsequent to the interaction in the video. The results provide a set of interactions, grounded in literature and verified empirically, in which people show the need for an explanation. Moreover, we can rank these scenarios by how strongly users think an explanation is necessary and find statistically significant differences. Comparing giving explanations with other possible response types, such as the robot apologizing or asking for help, we find that why-explanations are always among the two highest-rated responses, with the exception of when the robot simply acts normally and successfully. This stands in stark contrast to the other possible response types that are useful in a much more restricted set of situations. Lastly, we test for factors of an individual that might influence their response preferences, for example, their general attitude towards robots, but find no significant correlations. Our results can guide roboticists in designing more user-centered and transparent interactions and let explainability researchers develop more pinpointed explanations. Lennart Wachowiak, Andrew Fenn, Haris Kamran, Andrew Coles, Oya Çeliktutan, Gerard Canal |
HRI | 5 |
| 2024 | A Time Series Classification Pipeline for Detecting Interaction Ruptures in HRI Based on User ReactionsabstractTo be able to react to interaction ruptures such as errors, a robot needs a way of realizing such a rupture occurred. We test whether it is possible to detect interaction ruptures from the user’s anonymized speech, posture, and facial features. We showcase how to approach this task, presenting a time series classification pipeline that works well with various machine learning models. A sliding window is applied to the data and the continuously updated predictions make it suitable for detecting ruptures in real-time. Our best model, an ensemble of MiniRocket classifiers, is the winning approach to the ICMI ERR@HRI challenge. A feature importance analysis shows that the model heavily relies on speaker diarization data that indicates who spoke when. Posture data, on the other hand, impedes performance. Our code is available online1. Lennart Wachowiak, Peter Tisnikar, Andrew Coles, Gerard Canal, Oya Çeliktutan |
ICMI | 5 |
| 2024 | Cross Domain Policy Transfer with Effect Cycle-ConsistencyabstractTraining a robotic policy from scratch using deep reinforcement learning methods can be prohibitively expensive due to sample inefficiency. To address this challenge, transferring policies trained in the source domain to the target domain becomes an attractive paradigm. Previous research has typically focused on domains with similar state and action spaces but differing in other aspects. In this paper, our primary focus lies in domains with different state and action spaces, which has broader practical implications, i.e. transfer the policy from robot A to robot B. Unlike prior methods that rely on paired data, we propose a novel approach for learning the mapping functions between state and action spaces across domains using unpaired data. We propose effect cycle-consistency, which aligns the effects of transitions across two domains through a symmetrical optimization structure for learning these mapping functions. Once the mapping functions are learned, we can seamlessly transfer the policy from the source domain to the target domain. Our approach has been tested on three locomotion tasks and two robotic manipulation tasks. The empirical results demonstrate that our method can reduce alignment errors significantly and achieve better performance compared to the state-of-the-art method. Project page: https://ricky-zhu.github.io/effect_cycle_consistency. Tianhong Dai, Oya Çeliktutan |
ICRA | 3 |
| 2024 | Are Large Language Models Aligned with People's Social Intuitions for Human-Robot Interactions?abstractLarge language models (LLMs) are increasingly used in robotics, especially for high-level action planning. Meanwhile, many robotics applications involve human supervisors or collaborators. Hence, it is crucial for LLMs to generate socially acceptable actions that align with people’s preferences and values. In this work, we test whether LLMs capture people’s intuitions about behavior judgments and communication preferences in human-robot interaction (HRI) scenarios. For evaluation, we reproduce three HRI user studies, comparing the output of LLMs with that of real participants. We find that GPT-4 strongly outperforms other models, generating answers that correlate strongly with users’ answers in two studies — the first study dealing with selecting the most appropriate communicative act for a robot in various situations (rs= 0.82), and the second with judging the desirability, intentionality, and surprisingness of behavior (rs= 0.83). However, for the last study, testing whether people judge the behavior of robots and humans differently, no model achieves strong correlations. Moreover, we show that vision models fail to capture the essence of video stimuli and that LLMs tend to rate different communicative acts and behavior desirability higher than people. Lennart Wachowiak, Andrew Coles, Oya Çeliktutan, Gerard Canal |
IROS | 3 |
| 2024 | Towards Generalised and Incremental Bias Mitigation in Personality ComputingabstractBuilding systems for predicting human socio-emotional states has promising applications; however, if trained on biased data, such systems could inadvertently yield biased decisions. Bias mitigation remains an open problem, which tackles the correction of a model's disparate performance over different groups defined by particular sensitive attributes (e.g., gender, age, and race). In this work, we design a novel fairness loss function named Multi-Group Parity (MGP) to provide a generalised approach for bias mitigation in personality computing. In contrast to existing works in the literature, MGP is generalised as it features four ‘multiple’ properties (4Mul): multiple tasks, multiple modalities, multiple sensitive attributes, and multi-valued attributes. Moreover, we explore how to incrementally mitigate the biases when more sensitive attributes are taken into consideration sequentially. Towards this problem, we introduce a novel algorithm that utilises an incremental learning framework to mitigate bias against one attribute data at a time without compromising past fairness. Extensive experiments on two large-scale multi-modal personality recognition datasets validate the effectiveness of our approach in achieving superior bias mitigation under the proposed four properties and incremental debiasing settings. Jian Jiang 0001, Viswonathan Manoranjan, Hanan Salam, Oya Çeliktutan |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Automatic Context-Aware Inference of Engagement in HMI: A SurveyabstractEngagement is the process by which participants establish, maintain, and end their perceived connection. Automatic engagement inference is one of the tasks required to develop successful human-centered HMI applications. Engagement is a multi-faceted multimodal construct requiring high accuracy in interpretating contextual, verbal and non-verbal cues, making the development of an intelligent automated engagement inference system challenging. Existing surveys concentrate on specific application settings, and a comprehensive survey covering the different engagement facets, definition and inference across various contexts is lacking. Moreover, despite the importance of context-aware modeling, the literature lacks a systematic context-aware overview on the topic. This paper presents a comprehensive survey on previous work in engagement for HMI, entailing interdisciplinary definition, engagement components, publicly available datasets, ground truth assessment, and commonly used features and methods, serving as a guide for the development of future HMI interfaces with reliable context-aware engagement inference capability. An in-depth review across embodied and disembodied interaction modes, and an emphasis on the interaction context of which engagement is studied sets apart this survey from existing ones. Our findings suggest four important directions for future research: (1) context-aware computational modeling, (2) temporal dynamics, (3) personalised computing, and (4) bias and fairness of engagement inference systems. Hanan Salam, Oya Çeliktutan, Hatice Gunes, Mohamed Chetouani |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Learning Pessimism for Reinforcement LearningabstractOff-policy deep reinforcement learning algorithms commonly compensate for overestimation bias during temporal-difference learning by utilizing pessimistic estimates of the expected target returns. In this work, we propose Generalized Pessimism Learning (GPL), a strategy employing a novel learnable penalty to enact such pessimism. In particular, we propose to learn this penalty alongside the critic with dual TD-learning, a new procedure to estimate and minimize the magnitude of the target returns bias with trivial computational cost. GPL enables us to accurately counteract overestimation bias throughout training without incurring the downsides of overly pessimistic targets. By integrating GPL with popular off-policy algorithms, we achieve state-of-the-art results in both competitive proprioceptive and pixel-based benchmarks. Edoardo Cetin, Oya Çeliktutan |
AAAI | 2 |
| 2023 | A Simple Recipe to Meta-Learn Forward and Backward TransferabstractMeta-learning holds the potential to provide a general and explicit solution to tackle interference and forgetting in continual learning. However, many popular algorithms introduce expensive and unstable optimization processes with new key hyper-parameters and requirements, hindering their applicability. We propose a new, general, and simple meta-learning algorithm for continual learning (SiM4C) that explicitly optimizes to minimize forgetting and facilitate forward transfer. We show our method is stable, introduces only minimal computational overhead, and can be integrated with any memory-based continual learning algorithm in only a few lines of code. SiM4C meta-learns how to effectively continually learn even on very long task sequences, largely outperforming prior meta-approaches. Naively integrating with existing memory-based algorithms, we also record universal performance benefits and state-of-the-art results across different visual classification benchmarks without introducing new hyper-parameters. Edoardo Cetin, Antonio Carta, Oya Çeliktutan |
ICCV | 3 |
| 2023 | Learning to Solve Tasks with Exploring Prior BehavioursabstractDemonstrations are widely used in Deep Reinforcement Learning (DRL) for facilitating solving tasks with sparse rewards. However, the tasks in real-world scenarios can often have varied initial conditions from the demonstration, which would require additional prior behaviours. For example, consider we are given the demonstration for the task of picking up an object from an open drawer, but the drawer is closed in the training. Without acquiring the prior behaviours of opening the drawer, the robot is unlikely to solve the task. To address this, in this paper we propose an Intrinsic Reward Driven Example-based Control (IRDEC). Our method can endow agents with the ability to explore and acquire the required prior behaviours and then connect to the task-specific behaviours in the demonstration to solve sparse-reward tasks without requiring additional demonstration of the prior behaviours. The performance of our method outperforms other baselines on three navigation tasks and one robotic manipulation task with sparse rewards. Codes are available at https://github.com/Ricky-Zhu/IRDEC. Siyuan Li 0003, Tianhong Dai, Chongjie Zhang, Oya Çeliktutan |
IROS | 5 |
| 2023 | A Study on Customer's Perception of Robot Nonverbal Communication Skills in a Service EnvironmentabstractNonverbal communication has the potential to enable robots to interact with customers in service environments efficiently. While previous efforts in this domain have been paid to the understanding of customers’ interaction experience from different aspects, there is a lack of studies on the configuration of multimodal interaction (i.e., the combination of nonverbal gestures, voice, and touch) in service environments and the effect of nonverbal communication styles when performed in this setting. This paper aims to address the gap in the literature by introducing a multimodal HRI framework operated in a cafe shop setting. A systematic study is conducted with 171 customers. It is followed by an in-depth analysis based on objective and subjective measurements to build an understanding of customers’ attitudes towards the robot’s nonverbal behaviours. Nguyen Tan Viet Tuyen, Shintaro Okazaki, Oya Çeliktutan |
RO-MAN | 3 |
| 2023 | Neural Weight Search for Scalable Task Incremental LearningabstractTask incremental learning aims to enable a system to maintain its performance on previously learned tasks while learning new tasks, solving the problem of catastrophic forgetting. One promising approach is to build an individual network or sub-network for future tasks. However, this leads to an ever-growing memory due to saving extra weights for new tasks and how to address this issue has remained an open problem in task incremental learning. In this paper, we introduce a novel Neural Weight Search technique that designs a fixed search space where the optimal combinations of frozen weights can be searched to build new models for novel tasks in an end-to-end manner, resulting in a scalable and controllable memory growth. Extensive experiments on two benchmarks, i.e., Split-CIFAR-100 and CUB-to-Sketches, show our method achieves state-of-the-art performance with respect to both average inference accuracy and total memory cost.1 Jian Jiang 0001, Oya Çeliktutan |
WACV | 2 |
| 2022 | What Does Shared Understanding in Students' Face-to-Face Collaborative Learning Gaze Behaviours "Look Like"?
Qi Zhou 0011, Wannapon Suraworachet, Oya Çeliktutan, Mutlu Cukurova |
AIED (1) | 3 |
| 2022 | Personalized Productive Engagement Recognition in Robot-Mediated Collaborative LearningabstractIn this paper, we propose and compare personalized models for Productive Engagement (PE) recognition. PE is defined as the level of engagement that maximizes learning. Previously, in the context of robot-mediated collaborative learning, a framework of productive engagement was developed by utilizing multimodal data of 32 dyads and learning profiles, namely, Expressive Explorers (EE), Calm Tinkerers (CT), and Silent Wanderers (SW) were identified which categorize learners according to their learning gain. Within the same framework, a PE score was constructed in a non-supervised manner for real-time evaluation. Here, we use these profiles and the PE score within an AutoML deep learning framework to personalize PE models. We investigate two approaches for this purpose: (1) Single-task Deep Neural Architecture Search (ST-NAS), and (2) Multitask NAS (MT-NAS). In the former approach, personalized models for each learner profile are learned from multimodal features and compared to non-personalized models. In the MT-NAS approach, we investigate whether jointly classifying the learners’ profiles with the engagement score through multi-task learning would serve as an implicit personalization of PE. Moreover, we compare the predictive power of two types of features: incremental and non-incremental features. Non-incremental features correspond to features computed from the participant’s behaviours in fixed time windows. Incremental features are computed by accounting to the behaviour from the beginning of the learning activity till the time window where productive engagement is observed. Our experimental results show that (1) personalized models improve the recognition performance with respect to non-personalized models when training models for the gainer vs. non-gainer groups, (2) multitask NAS (implicit personalization) also outperforms non-personalized models, (3) the speech modality has high contribution towards prediction, and (4) non-incremental features outperform the incremental ones overall. Vetha Vikashini Chithrra Raghuram, Hanan Salam, Jauwairia Nasir, Barbara Bruno, Oya Çeliktutan |
ICMI | 5 |
| 2022 | Stabilizing Off-Policy Deep Reinforcement Learning from PixelsabstractOff-policy reinforcement learning (RL) from pixel observations is notoriously unstable. As a result, many successful algorithms must combine different domain-specific practices and auxiliary losses to learn meaningful behaviors in complex environments. In this work, we provide novel analysis demonstrating that these instabilities arise from performing temporal-difference learning with a convolutional encoder and low-magnitude rewards. We show that this new visual deadly triad causes unstable training and premature convergence to degenerate solutions, a phenomenon we name catastrophic self-overfitting. Based on our analysis, we propose A-LIX, a method providing adaptive regularization to the encoder’s gradients that explicitly prevents the occurrence of catastrophic self-overfitting using a dual objective. By applying A-LIX, we significantly outperform the prior state-of-the-art on the DeepMind Control and Atari benchmarks without any data augmentation or auxiliary losses. Edoardo Cetin, Philip J. Ball, Stephen J. Roberts, Oya Çeliktutan |
ICML | 4 |
| 2022 | Policy Gradient With Serial Markov Chain ReasoningabstractWe introduce a new framework that performs decision-making in reinforcement learning (RL) as an iterative reasoning process. We model agent behavior as the steady-state distribution of a parameterized reasoning Markov chain (RMC), optimized with a new tractable estimate of the policy gradient. We perform action selection by simulating the RMC for enough reasoning steps to approach its steady-state distribution. We show our framework has several useful properties that are inherently missing from traditional RL. For instance, it allows agent behavior to approximate any continuous distribution over actions by parameterizing the RMC with a simple Gaussian transition function. Moreover, the number of reasoning steps to reach convergence can scale adaptively with the difficulty of each action selection decision and can be accelerated by re-using past solutions. Our resulting algorithm achieves state-of-the-art performance in popular Mujoco and DeepMind Control benchmarks, both for proprioceptive and pixel-based tasks. Edoardo Cetin, Oya Çeliktutan |
NeurIPS | 2 |
| 2022 | Agree or Disagreeƒ Generating Body Gestures from Affective Contextual Cues during Dyadic InteractionsabstractHumans naturally produce nonverbal signals such as facial expressions, body movements, hand gestures, and tone of voice, along with words, to communicate their messages, opinions, and feelings. Considering robots are progressively moving out from research laboratories into human environments, it is increasingly desirable that they develop a similar social intelligence. Therefore, equipping social robots with nonverbal communication skills has been an active research area for decades, where data-driven, end-to-end learning approaches have become predominant in recent years, offering scalability and generalisability. However, most of these approaches consider a single character, modelling intrapersonal dynamics only. In this paper, we propose a method based on conditional Generative Adversarial Networks, intending to generate behaviours for a robot in affective dyadic interactions. Our method takes as an input the audio of a target person together with the nonverbal signals of their interacting partner, modelled by a novel Context Encoder, to generate appropriate body gestures. We evaluate our method on the multimodal JESTKOD dataset that comprises dyadic interactions under agreement and disagreement scenarios. The experimental results show that Context Encoder can better contribute to the prediction of co-speech gestures in agreement situations. Nguyen Tan Viet Tuyen, Oya Çeliktutan |
RO-MAN | 2 |
| 2022 | Analysing Eye Gaze Patterns during Confusion and Errors in Human-Agent CollaborationsabstractAs human–agent collaborations become more prevalent, it is increasingly important for an agent to be able to adapt to their collaborator and explain their own behavior. In order to do so, they need to be able to identify critical states during the interaction that call for proactive clarifications or behavioral adaptations. In this paper, we explore whether the agent could infer such states from the human’s eye gaze for which we compare gaze patterns across different situations in a collaborative task. Our findings show that the human’s gaze patterns significantly differ between times at which the user is confused about the task, times at which the agent makes an error, and times of normal workflow. During errors the amount of gaze towards the agent increases, while during confusion the amount towards the environment increases. We conclude that these signals could tell the agent what and when to explain. Lennart Wachowiak, Peter Tisnikar, Gerard Canal, Andrew Coles, Matteo Leonetti, Oya Çeliktutan |
RO-MAN | 6 |
| 2022 | A Cloud-based Robot System for Long-term Interaction: Principles, Implementation, Lessons LearnedabstractMaking the transition to long-term interaction with social-robot systems has been identified as one of the main challenges in human-robot interaction. This article identifies four design principles to address this challenge and applies them in a real-world implementation: cloud-based robot control, a modular design, one common knowledge base for all applications, and hybrid artificial intelligence for decision making and reasoning. The control architecture for this robot includes a common Knowledge-base (ontologies), Data-base, “Hybrid Artificial Brain” (dialogue manager, action selection and explainable AI), Activities Centre (Timeline, Quiz, Break and Sort, Memory, Tip of the Day, \( \ldots \) ), Embodied Conversational Agent (ECA, i.e., robot and avatar), and Dashboards (for authoring and monitoring the interaction). Further, the ECA is integrated with an expandable set of (mobile) health applications. The resulting system is a Personal Assistant for a healthy Lifestyle (PAL), which supports diabetic children with self-management and educates them on health-related issues (48 children, aged 6–14, recruited via hospitals in the Netherlands and in Italy). It is capable of autonomous interaction “in the wild” for prolonged periods of time without the need for a “Wizard-of-Oz” (up until 6 months online). PAL is an exemplary system that provides personalised, stable and diverse, long-term human-robot interaction. Frank Kaptein, Bernd Kiefer, Antoine Cully, Oya Çeliktutan, Bert P. B. Bierman, Rifca Rijgersberg-Peters, Joost Broekens, Willeke van Vught, Michael van Bekkum, Yiannis Demiris, Mark A. Neerincx |
ACM Trans. Hum. Robot Interact. | 4 |
| 2021 | GROWL: Group Detection With Link PredictionabstractInteraction group detection has been previously addressed with bottom-up approaches which relied on the position and orientation information of individuals. These approaches were primarily based on pairwise affinity matrices and were limited to static, third-person views. This problem can greatly benefit from a holistic approach based on Graph Neural Networks (GNNs) beyond pairwise relationships, due to the inherent spatial configuration that exists between individuals who form interaction groups. Our proposed method, GROup detection With Link prediction (GROWL), demonstrates the effectiveness of a GNN based approach. GROWL predicts the link between two individuals by generating a feature embedding based on their neighbourhood in the graph and determines whether they are connected with a shallow binary classification method such as Multi-layer Perceptrons (MLPs). We test our method against other state-of-the-art group detection approaches on both a third-person view dataset and a robocentric (i.e., egocentric) dataset. In addition, we propose a multimodal approach based on RGB and depth data to calculate a representation GROWL can utilise as input. Our results show that a GNN based approach can significantly improve accuracy across different camera views, i.e., third-person and egocentric views. Viktor Schmuck, Oya Çeliktutan |
FG | 2 |
| 2021 | Domain-Robust Visual Imitation Learning with Mutual Information Constraints
Edoardo Cetin, Oya Çeliktutan |
ICLR | 2 |
| 2021 | Socially Informed AI for Healthcare: Understanding and Generating Multimodal Nonverbal CuesabstractAdvances in the areas of face and gesture analysis, computational paralinguistics, multimodal interaction, and human-computer interaction have all played a major role in shaping research into assistive technologies over the last decade. This has resulted in a breadth of practical applications ranging from diagnosis and treatment tools to social companion technologies. From an analytical perspective, nonverbal cues provide understanding into the assessment of mental health and wellbeing (i.e., detecting depression and pain) and the detection of developmental and neurological conditions such as autism, dementia, and schizophrenia. From both a synthesis and generative perspective, it is necessary that assistive technologies, either disembodied or embodied, are capable of generating engaging, interactive behaviours and interventions that are personalised and adapted to user’s needs, profiles, and preferences. While nonverbal cues play an essential role, there are still many key issues to overcome, which affect both the development and the deployment of multimodal technologies in real-world settings. The key aim of this multidisciplinary workshop is to foster cross-pollination by bringing together computer scientists and social psychologists to discuss innovative ideas, challenges and opportunities for understanding and generating multimodal nonverbal cues within the scope of healthcare applications1. Oya Çeliktutan, Alexandra Livia Georgescu, Nicholas Cummins |
ICMI | 1 |
| 2021 | Learning Routines for Effective Off-Policy Reinforcement LearningabstractThe performance of reinforcement learning depends upon designing an appropriate action space, where the effect of each action is measurable, yet, granular enough to permit flexible behavior. So far, this process involved non-trivial user choices in terms of the available actions and their execution frequency. We propose a novel framework for reinforcement learning that effectively lifts such constraints. Within our framework, agents learn effective behavior over a routine space: a new, higher-level action space, where each routine represents a set of ’equivalent’ sequences of granular actions with arbitrary length. Our routine space is learned end-to-end to facilitate the accomplishment of underlying off-policy reinforcement learning objectives. We apply our framework to two state-of-the-art off-policy algorithms and show that the resulting agents obtain relevant performance improvements while requiring fewer interactions with the environment per episode, improving computational efficiency. Edoardo Cetin, Oya Çeliktutan |
ICML | 2 |
| 2021 | Audio-Driven Robot Upper-Body Motion SynthesisabstractBody language is an important aspect of human communication, which an effective human-robot interaction interface should mimic well. Human beings exchange information and convey their thoughts and feelings through gaze, facial expressions, body language, and tone of voice along with spoken words, and infer 65% of the meaning of the communicated messages from these nonverbal cues. Modern robotic platforms are, however, limited in their ability to automatically generate behaviors that align with their speech. In this article, we develop a neural-network-based system that takes audio from a user as an input and generates upper-body gestures, including head, hand, and torso movements of the user on a humanoid robot, namely, Softbank Robotics' Pepper. Our system was evaluated quantitatively as well as qualitatively using Web surveys when driven by natural speech and synthetic speech. We compare the impact of generic and person-specific neural-network models on the quality of synthesized movements. We further investigate the relationships between quantitative and qualitative evaluations and examine how the speaker's personality traits affect the synthesized movements. Jan Ondras, Oya Çeliktutan, Paul Bremner, Hatice Gunes |
IEEE Trans. Cybern. | 2 |
| 2020 | Robocentric Conversational Group DiscoveryabstractDetecting people interacting and conversing with each other is essential to equipping social robots with autonomous navigation and service capabilities in crowded social scenes. In this paper, we introduced a method for unsupervised conversational group detection in images captured from a mobile robot's perspective. To this end, we collected a novel dataset called Robocentric Indoor Crowd Analysis (RICA). The RICA dataset features over 100,000 RGB, depth, and wide- angle camera images as well as LIDAR readings, recorded during a social event where the robot navigated between participants and captured interactions among groups using its on-board sensors. Using the RICA dataset, we implemented an unsupervised group detection method based on agglomerative hierarchical clustering. Our results show that incorporating the depth modality and using normalised features in the clustering algorithm improved group detection accuracy by a margin of 3% on average. Viktor Schmuck, Tingran Sheng, Oya Çeliktutan |
RO-MAN | 3 |
| 2019 | Multimodal Human-Human-Robot Interactions (MHHRI) Dataset for Studying Personality and EngagementabstractIn this paper we introduce a novel dataset, the Multimodal Human-Human-Robot-Interactions (MHHRI) dataset, with the aim of studying personality simultaneously in human-human interactions (HHI) and human-robot interactions (HRI) and its relationship with engagement. Multimodal data was collected during a controlled interaction study where dyadic interactions between two human participants and triadic interactions between two human participants and a robot took place with interactants asking a set of personal questions to each other. Interactions were recorded using two static and two dynamic cameras as well as two biosensors, and meta-data was collected by having participants fill in two types of questionnaires, for assessing their own personality traits and their perceived engagement with their partners (self labels) and for assessing personality traits of the other participants partaking in the study (acquaintance labels). As a proof of concept, we present baseline results for personality and engagement classification. Our results show that (i) trends in personality classification performance remain the same with respect to the self and the acquaintance labels across the HHI and HRI settings; (ii) for extroversion, the acquaintance labels yield better results as compared to the self labels; (iii) in general, multi-modality yields better performance for the classification of personality traits. Oya Çeliktutan, Efstratios Skordos, Hatice Gunes |
IEEE Trans. Affect. Comput. | 1 |
| 2017 | Automatic replication of teleoperator head movements and facial expressions on a humanoid robotabstractRobotic telepresence aims to create a physical presence for a remotely located human (teleoperator) by reproducing their verbal and nonverbal behaviours (e.g. speech, gestures, facial expressions) on a robotic platform. In this work, we propose a novel teleoperation system that combines the replication of facial expressions of emotions (neutral, disgust, happiness, and surprise) and head movements on the fly on the humanoid robot Nao. Robots' expression of emotions is constrained by their physical and behavioural capabilities. As the Nao robot has a static face, we use the LEDs located around its eyes to reproduce the teleoperator expressions of emotions. Using a web camera, we computationally detect the facial action units and measure the head pose of the operator. The emotion to be replicated is inferred from the detected action units by a neural network. Simultaneously, the measured head motion is smoothed and bounded to the robot's physical limits by applying a constrained-state Kalman filter. In order to evaluate the proposed system, we conducted a user study by asking 28 participants to use the replication system by displaying facial expressions and head movements while being recorded by a web camera. Subsequently, 18 external observers viewed the recorded clips via an online survey and assessed the quality of the robot's replication of the participants' behaviours. Our results show that the proposed teleoperation system can successfully communicate emotions and head movements, resulting in a high agreement among the external observers (ICC_E = 0.91, ICC_HP = 0.72). Jan Ondras, Oya Çeliktutan, Evangelos Sariyanidi, Hatice Gunes |
RO-MAN | 2 |
| 2017 | Automatic Prediction of Impressions in Time and across Varying Context: Personality, Attractiveness and LikeabilityabstractIn this paper, we propose a novel multimodal framework for automatically predicting the impressions of extroversion, agreeableness, conscientiousness, neuroticism , openness, attractiveness and likeability continuously in time and across varying situational contexts. Differently from the existing works, we obtain visual-only and audio-only annotations continuously in time for the same set of subjects, for the first time in the literature, and compare them to their audio-visual annotations. We propose a time-continuous prediction approach that learns the temporal relationships rather than treating each time instant separately. Our experiments show that the best prediction results are obtained when regression models are learned from audio-visual annotations and visual cues, and from audio-visual annotations and visual cues combined with audio cues at the decision level. Continuously generated annotations have the potential to provide insight into better understanding which impressions can be formed and predicted more dynamically, varying with situational context, and which ones appear to be more static and stable over time. Oya Çeliktutan, Hatice Gunes |
IEEE Trans. Affect. Comput. | 1 |
| 2016 | Personality Perception of Robot Avatar Tele-operatorsabstractNowadays a significant part of human-human interaction takes place over distance. Tele-operated robot avatars, in which an operator's behaviours are portrayed by a robot proxy, have the potential to improve distance interaction, e.g., improving social presence and trust. However, having communication mediated by a robot changes the perception of the operator's appearance and behaviour, which have been shown to be used alongside vocal cues in judging personality. In this paper we present a study that investigates how robot mediation affects the way the personality of the operator is perceived. More specifically, we aim to investigate if judges of personality can be consistent in assessing personality traits, can agree with one another, can agree with operators' self-assessed personality, and shift their perceptions to incorporate characteristics associated with the robot's appearance. Our experiments show that (i) judges utilise robot appearance cues along with operator vocal cues to make their judgements, (ii) operators' arm gestures reproduced on the robot aid personality judgements, and (iii) how personality cues are perceived and evaluated through speech, gesture and robot appearance is highly operator-dependent. We discuss the implications of these results for both tele-operated and autonomous robots that aim to portray personality. Paul Bremner, Oya Çeliktutan, Hatice Gunes |
HRI | 2 |
| 2015 | Computational analysis of human-robot interactions through first-person vision: Personality and interaction experienceabstractIn this paper, we analyse interactions with Nao, a small humanoid robot, from the viewpoint of human participants through an ego-centric camera placed on their forehead. We focus on human participants' and robot's personalities and their impact on the human-robot interactions. We automatically extract nonverbal cues (e.g., head movement) from first-person perspective and explore the relationship of nonverbal cues with participants' self-reported personality and their interaction experience. We generate two types of behaviours for the robot (i.e., extroverted vs. introverted) and examine how robot's personality and behaviour affect the findings. Significant correlations are obtained between the extroversion and agreeable-ness traits of the participants and the perceived enjoyment with the extroverted robot. Plausible relationships are also found between the measures of interaction experience and personality and the first-person vision features. We then use computational models to automatically predict the participants' personality traits from these features. Promising results are achieved for the traits of agreeableness, conscientiousness and extroversion. Oya Çeliktutan, Hatice Gunes |
RO-MAN | 1 |
| 2014 | Continuous prediction of perceived traits and social dimensions in space and timeabstractDeveloping automatic personality predictors requires generating reliable annotations, i.e., ground truth. To date, researchers have relied on the overall ratings provided for a whole video sequence, either obtained by self-assessment or provided by external observers. In this paper, we propose a novel personality assessment approach, where we ask external observers to continuously provide ratings along multiple dimensions ranging from 0 to 100 along time, and we generate continuous annotations in space and time. In addition to the widely used Big Five personality dimensions, we introduce three more dimensions that have the potential to gauge the reliability of the perceived social and trait judgements in the context of varying situational interactions between a human subject and virtual characters. Our results demonstrate the viability of the proposed approach and the plausible relationship between the extracted features and perceived trait and social dimensions. Annotations obtained continuously in time and in trait-social dimensional space showed that a number of dimensions appear to be more static and stable over time while other dimensions appear to be more dynamic. Oya Çeliktutan, Hatice Gunes |
ICIP | 1 |
| 2014 | MAPTRAITS 2014 - The First Audio/Visual Mapping Personality Traits Challenge - An Introduction: Perceived Personality and Social DimensionsabstractThe Audio/Visual Mapping Personality Challenge and Workshop (MAPTRAITS) is a competition event that is organised to facilitate the development of signal processing and machine learning techniques for the automatic analysis of personality traits and social dimensions. MAPTRAITS includes two sub-challenges, the continuous space-time sub-challenge and the quantised space-time sub-challenge. The continuous sub-challenge evaluated how systems predict the variation of perceived personality traits and social dimensions in time, whereas the quantised challenge evaluated the ability of systems to predict the overall perceived traits and dimensions in shorter video clips. To analyse the effect of audio and visual modalities on personality perception, we compared systems under three different settings: visual-only, audio-only and audio-visual. With MAPTRAITS we aimed at improving the knowledge on the automatic analysis of personality traits and social dimensions by producing a benchmarking protocol and encouraging the participation of various research groups from different backgrounds. Oya Çeliktutan, Florian Eyben, Evangelos Sariyanidi, Hatice Gunes, Björn W. Schuller |
ICMI | 1 |
| 2014 | Evaluation of video activity localizations integrating quality and quantity measurements
Christian Wolf 0001, Eric Lombardi, Julien Mille, Oya Çeliktutan, Mingyuan Jiu, Emre Dogan, Gonen Eren, Moez Baccouche, Emmanuel Dellandréa, Charles-Edmond Bichot, Christophe Garcia, Bülent Sankur |
Comput. Vis. Image Underst. | 4 |
| 2008 | Multi-attribute robust facial feature localizationabstractIn this paper, we focus on the reliable detection of facial fiducial points, such as eye, eyebrow and mouth corners. The proposed algorithm aims to improve automatic land-marking performance in challenging realistic face scenarios subject to pose variations, high-valence facial expressions and occlusions. We explore the potential of several feature modalities, namely, gabor wavelets, independent component analysis (ICA), non-negative matrix factorization (NMF), and discrete cosine transform (DCT), both singly and jointly. We show that the selection of the highest scoring face patch as the corresponding landmark is not always the best, but that there is considerable room for improvement with the cooperation among several high scoring candidates and also using a graph-based post-processing method. We present our experimental results on Bosphorus face database, a new challenging database. Oya Çeliktutan, Hatice Çinar Akakin, Bülent Sankur |
FG | 1 |
| 2008 | Blind Identification of Source Cell-Phone ModelabstractThe various image-processing stages in a digital camera pipeline leave telltale footprints, which can be exploited as forensic signatures. These footprints consist of pixel defects, of unevenness of the responses in the CCD sensor, black current noise, and may originate from proprietary interpolation algorithms involved in color filter array [CFA]. Various imaging device (camera, scanner etc.) identification methods are based on the analysis of these artifacts. In this work, we set to explore three sets of forensic features, namely binary similarity measures, image quality measures and higher order wavelet statistics in conjunction with SVM classifier to identify the originating camera. We demonstrate that our camera model identification algorithm achieves more accurate identification, and that it can be made robust to a host of image manipulations. The algorithm has potential to discriminate camera units within the same model. Oya Çeliktutan, Bülent Sankur, Ismail Avcibas |
IEEE Trans. Inf. Forensics Secur. | 1 |