EDBT 2026 Demo / reviewers in the wild / expert
Ognjen Rudovic
dblp:53/21 · also Oggi Rudovic, Ognjen (Oggi) Rudovic
· DBLP profile ↗
43ranked-venue papers
11as first author
12since 2021 · last 2025
0000-0003-1165-6075ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 27 · 7 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
Hyung-Gun Chi, Zakaria Aldeneh, Tatiana Likhomanenko, Ognjen Rudovic, Takuya Higuchi, Shinji Watanabe 0001, Ahmed Hussen Abdelaziz |
INTERSPEECH | 4 |
| 2025 | Adaptive Knowledge Distillation for Device-Directed Speech Detection
Hyung-Gun Chi, Florian Pesce, Wonil Chang, Ognjen Rudovic, Arturo Argueta, Vineet Garg, Ahmed Hussen Abdelaziz |
INTERSPEECH | 4 |
| 2024 | Modality Drop-Out for Multimodal Device Directed Speech Detection Using Verbal and Non-Verbal FeaturesabstractDevice-directed speech detection (DDSD) is the binary classification task of distinguishing between queries directed at a voice assistant versus side conversation or background speech. State-of-the-art DDSD systems use verbal cues, e.g acoustic, text and/or automatic speech recognition system (ASR) features, to classify speech as device-directed or otherwise, and often have to contend with one or more of these modalities being unavailable when deployed in real-world settings. In this paper, we investigate fusion schemes for DDSD systems that can be made more robust to missing modalities. Concurrently, we study the use of non-verbal cues, specifically prosody features, in addition to verbal cues for DDSD. We present different approaches to combine scores and embeddings from prosody with the corresponding verbal cues, finding that prosody improves DDSD performance by upto 8.5% in terms of false acceptance rate (FA) at a given fixed operating point via non-linear intermediate fusion, while our use of modality dropout techniques improves the performance of these models by 7.4% in terms of FA when evaluated with missing modalities during inference time. Gautam Krishna, Sameer Dharur, Ognjen Rudovic, Pranay Dighe, Saurabh Adya, Ahmed Hussen Abdelaziz, Ahmed H. Tewfik |
ICASSP | 3 |
| 2024 | Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness
Sai Srujana Buddi, Satyam Kumar 0001, Utkarsh Oggy Sarawgi, Vineet Garg, Shivesh Ranjan, Ognjen Rudovic, Ahmed Hussen Abdelaziz, Saurabh Adya |
INTERSPEECH | 6 |
| 2024 | Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
Shruti Palaskar, Ognjen Rudovic, Sameer Dharur, Florian Pesce, Gautam Krishna, Aswin Sivaraman, Jack Berkowitz, Ahmed Hussen Abdelaziz, Saurabh Adya, Ahmed H. Tewfik |
INTERSPEECH | 2 |
| 2023 | Audio-to-Intent Using Acoustic-Textual Subword Representations from End-to-End ASRabstractAccurate prediction of the user intent to interact with a voice assistant (VA) on a device (e.g. a smartphone) is critical for achieving naturalistic, engaging, and privacy-centric interactions with the VA. To this end, we present a novel approach to predict the user intention (whether the user is speaking to the device or not) directly from acoustic and textual information encoded at subword tokens which are obtained via an end-to-end (E2E) ASR model. Modeling directly the subword tokens, compared to modeling of the phonemes and/or full words, has at least two advantages: (i) it provides a unique vocabulary representation, where each token has a semantic meaning, in contrast to the phoneme-level representations, (ii) each subword token has a reusable "sub"-word acoustic pattern (that can be used to construct multiple full words), resulting in a largely reduced vocabulary space than of the full words. To learn the subword representations for the audio-to-intent classification, we extract: (i) acoustic information from an E2E-ASR model, which provides frame-level CTC posterior probabilities for the subword tokens, and (ii) textual information from a pretrained continuous bag-of-words model capturing the semantic meaning of the subword tokens. The key to our approach is that it combines acoustic subword-level posteriors with text information using the notion of positional-encoding to account for multiple ASR hypotheses simultaneously. We show that the proposed approach learns robust representations for audio-to-intent classification and correctly mitigates 93.3% of unintended user audio from invoking the VA at 99% true positive rate. Pranay Dighe, Prateeth Nayak, Ognjen Rudovic, Erik Marchi, Xiaochuan Niu, Ahmed H. Tewfik |
ICASSP | 3 |
| 2023 | Less Is More: A Unified Architecture for Device-Directed Speech Detection with Multiple Invocation TypesabstractSuppressing unintended invocation of the device because of the speech that sounds like wake-word, or accidental button presses, is critical for a good user experience, and is referred to as False-Trigger-Mitigation (FTM). In case of multiple invocation options, the traditional approach to FTM is to use invocation-specific models, or a single model for all invocations. Both approaches are sub-optimal: the memory cost for the former approach grows linearly with the number of invocation options, which is prohibitive for on-device deployment, and does not take advantage of shared training data; while the latter is unable to accurately capture acoustic differences across different invocation types. To this end, we propose a Unified Acoustic Detector (UAD) for FTM when multiple invocation options are available on device. The proposed UAD is trained using a multi-task learning framework, where a jointly trained acoustic encoder model is augmented with invocation-specific classification layers. In the context of the FTM task, we show for the first time that using the shared model architecture across invocations (thus, keeping the model size similar to that of a monolithic model used for a single invocation type), we can not only match but largely improve the accuracy of the invocation-specific models. In particular, in the challenging case of touch-based invocation, we obtain 50% and 35% relative improvement in false positive rate at 99% true positive rate, when compared with a singleoutput model for both invocations, and separate models per invocation, respectively. Furthermore, we propose streaming and non-streaming variants of the UAD, and show that they both outperform a traditional ASR-based approach to FTM. Ognjen Rudovic, Wonil Chang, Vineet Garg, Pranay Dighe, Pramod Simha, Jack Berkowitz, Ahmed Hussen Abdelaziz, Sachin Kajarekar, Erik Marchi, Saurabh Adya |
ICASSP | 1 |
| 2022 | DeepFN: Towards Generalizable Facial Action Unit Recognition with Deep Face NormalizationabstractDeployment of facial action unit recognition models has been impeded due to their limited generalization to unseen people and demographics. This work conducts an in-depth generalization analysis across several sources of variance: individuals (40 subjects), genders (male and female), skin types (darker and lighter), and databases (BP4D and DISFA). To help suppress the variance in data, we propose using self-supervised denoising autoencoders to transfer facial expressions of different people onto a common facial template which is then used to train and evaluate each of the models. We show that person-independent models yielded significantly lower performance (55% average F1 and accuracy across 40 subjects) than person-dependent models (60.3 %), leading to a generalization gap of 5.3%. However, normalizing the data with the proposed method significantly increased the performance of person-independent models (59.6%). Similarly, the proposed method was able to significantly reduce the generalization gap when considering gender (2.4%), skin type (5.3%), and dataset (9.4%). These findings represent an important step towards the creation of more generalizable facial action unit recognition systems. Javier Hernandez, Daniel McDuff, Ognjen Rudovic, Alberto Fung, Mary Czerwinski |
ACII | 3 |
| 2022 | Streaming on-Device Detection of Device Directed Speech from Voice and Touch-Based InvocationabstractWhen interacting with smart devices such as mobile-phones or wearables, the user typically invokes a virtual assistant (VA) by saying a keyword or by pressing a button on the device. However, in many cases, the VA can accidentally be invoked by the keyword-like speech or accidental button press, which may have implications on user experience and privacy. To this end, we propose an acoustic false-trigger-mitigation (FTM) approach for on-device device-directed speech detection that simultaneously handles the voice-trigger and touch-based invocation. To facilitate the model deployment on-device, we introduce a new streaming decision layer, derived using the notion of temporal convolutional networks (TCN) [1], known for their computational efficiency. To the best of our knowledge, this is the first approach that can detect device-directed speech from more than one invocation type in a streaming fashion. We compare this approach with streaming alternatives based on vanilla Average layer, and canonical LSTMs, and show: (i) that all the models show only a small degradation in accuracy compared with the invocation-specific models, and (ii) that the newly introduced streaming TCN consistently performs better or comparable with the alternatives, while mitigating device-undirected speech faster in time, and with (relative) reduction in runtime peak-memory over the LSTM-based approach of 33% vs. 7%, when compared to a non-streaming counterpart. Ognjen Rudovic, Akanksha Bindal, Vineet Garg, Pramod Simha, Pranay Dighe, Sachin Kajarekar |
ICASSP | 1 |
| 2022 | Device-Directed Speech Detection: Regularization via Distillation for Weakly-Supervised ModelsabstractWe address the problem of detecting speech directed to a device that does not contain a specific wake-word.Specifically, we focus on audio coming from a touch-based invocation.Mitigating virtual assistants (VAs) activation due to accidental button presses is critical for user experience.While the majority of approaches to false trigger mitigation (FTM) are designed to detect the presence of a target keyword, inferring user intent in absence of keyword is difficult.This also poses a challenge when creating the training/evaluation data for such systems due to inherent ambiguity in the user's data.To this end, we propose a novel FTM approach that uses weakly-labeled training data obtained with a newly introduced data sampling strategy.While this sampling strategy reduces data annotation efforts, the data labels are noisy as the data are not annotated manually.We use these data to train an acoustics-only model for the FTM task by regularizing its loss function via knowledge distillation from an ASR-based (LatticeRNN) model.This improves the model decisions, resulting in 66% gain in accuracy, as measured by equal-error-rate (EER), over the base acoustics-only model.We also show that the ensemble of the LatticeRNN and acousticdistilled models brings further accuracy improvement of 20%. Vineet Garg, Ognjen Rudovic, Pranay Dighe, Ahmed Hussen Abdelaziz, Erik Marchi, Saurabh Adya, Chandra Dhir, Ahmed H. Tewfik |
INTERSPEECH | 2 |
| 2022 | Toward Personalized Affect-Aware Socially Assistive Robot Tutors for Long-Term Interventions with Children with AutismabstractAffect-aware socially assistive robotics (SAR) has shown great potential for augmenting interventions for children with autism spectrum disorders (ASD). However, current SAR cannot yet perceive the unique and diverse set of atypical cognitive-affective behaviors from children with ASD in an automatic and personalized fashion in long-term (multi-session) real-world interactions. To bridge this gap, this work designed and validated personalized models of arousal and valence for children with ASD using a multi-session in-home dataset of SAR interventions. By training machine learning (ML) algorithms with supervised domain adaptation (s-DA), the personalized models were able to tradeoff between the limited individual data and the more abundant less personal data pooled from other study participants. We evaluated the effects of personalization on a long-term multimodal dataset consisting of four children with ASD with a total of 19 sessions, and derived inter-rater reliability (IR) scores for binary arousal (IR = 83%) and valence (IR = 81%) labels between human annotators. Our results show that personalized Gradient Boosted Decision Trees (XGBoost) models with s-DA outperformed two non-personalized individualized and generic model baselines not only on the weighted average of all sessions, but also statistically ( p < .05) across individual sessions. This work paves the way for the development of personalized autonomous SAR systems tailored toward individuals with atypical cognitive-affective and socio-emotional needs. Zhonghao Shi, Thomas R. Groechel, Shomik Jain, Kourtney Chima, Ognjen Rudovic, Maja J. Mataric |
ACM Trans. Hum. Robot Interact. | 5 |
| 2021 | Special Issue on Automated Perception of Human Affect from Longitudinal Behavioral DataabstractThe papers in this special section are aimed at contributions from computational neuroscience and psychology, artificial intelligence, machine learning, and affective computing, challenging and expanding current research on interpretation and estimation of human affective behavior from longitudinal data, i.e., single or multiple modalities captured over extended periods of time allowing efficient representation of behavior and inference in terms of affect and other socio-cognitive dimensions. Pablo V. A. Barros, Stefan Wermter, Ognjen Rudovic, Hatice Gunes |
IEEE Trans. Affect. Comput. | 3 |
| 2020 | Unsupervised Multi-Target Domain Adaptation: An Information Theoretic ApproachabstractUnsupervised domain adaptation (uDA) models focus on pairwise adaptation settings where there is a single, labeled, source and a single target domain. However, in many real-world settings one seeks to adapt to multiple, but somewhat similar, target domains. Applying pairwise adaptation approaches to this setting may be suboptimal, as they fail to leverage shared information among multiple domains. In this work, we propose an information theoretic approach for domain adaptation in the novel context of multiple target domains with unlabeled instances and one source domain with labeled instances. Our model aims to find a shared latent space common to all domains, while simultaneously accounting for the remaining private, domain-specific factors. Disentanglement of shared and private information is accomplished using a unified information-theoretic approach, which also serves to establish a stronger link between the latent representations and the observed data. The resulting model, accompanied by an efficient optimization algorithm, allows simultaneous adaptation from a single source to multiple target domains. We test our approach on three challenging publicly-available datasets, showing that it outperforms several popular domain adaptation methods. Behnam Gholami, Pritish Sahu, Ognjen Rudovic, Konstantinos Bousmalis, Vladimir Pavlovic 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Multi-modal Active Learning From Human Data: A Deep Reinforcement Learning ApproachabstractHuman behavior expression and experience are inherently multimodal, and characterized by vast individual and contextual heterogeneity. To achieve meaningful human-computer and human-robot interactions, multi-modal models of the user’s states (e.g., engagement) are therefore needed. Most of the existing works that try to build classifiers for the user’s states assume that the data to train the models are fully labeled. Nevertheless, data labeling is costly and tedious, and also prone to subjective interpretations by the human coders. This is even more pronounced when the data are multi-modal (e.g., some users are more expressive with their facial expressions, some with their voice). Thus, building models that can accurately estimate the user’s states during an interaction is challenging. To tackle this, we propose a novel multi-modal active learning (AL) approach that uses the notion of deep reinforcement learning (RL) to find an optimal policy for active selection of the user’s data, needed to train the target (modality-specific) models. We investigate different strategies for multi-modal data fusion, and show that the proposed model-level fusion coupled with RL outperforms the feature-level and modality-specific models, and the naïve AL strategies such as random sampling, and the standard heuristics such as uncertainty sampling. We show the benefits of this approach on the task of engagement estimation from real-world child-robot interactions during an autism therapy. Importantly, we show that the proposed multi-modal AL approach can be used to efficiently personalize the engagement classifiers to the target user using a small amount of actively selected user’s data. Ognjen Rudovic, Meiru Zhang, Björn W. Schuller, Rosalind W. Picard |
ICMI | 1 |
| 2019 | Copula Ordinal Regression Framework for Joint Estimation of Facial Action Unit IntensityabstractJoint modeling of the intensity of multiple facial action units (AUs) from face images is challenging due to the large number of AUs (30+) and their intensity levels (6). This is in part due to the lack of suitable models that can efficiently handle such a large number of outputs/classes simultaneously, but also due to the lack of suitable data the models on. For this reason, majority of the methods resort to independent classifiers for the AU intensity. This is suboptimal for at least two reasons: the facial appearance of some AUs changes depending on the intensity of other AUs, and some AUs co-occur more often than others. To this end, we propose the Copula regression approach for modeling multivariate ordinal variables. Our model accounts for ordinal structure in output variables and their non-linear dependencies via copula functions modeled as cliques of a conditional random fields. The copula ordinal regression model achieves the joint learning and inference of intensities of multiple AUs, while being computationally tractable. We demonstrate the effectiveness of our approach on three challenging datasets of naturalistic facial expressions and we show that the estimation of target AU intensities improves especially in the case of (a) noisy image features, (b) head-pose variations and (c) imbalanced training data. Lastly, we show that the proposed approach consistently outperforms (i) independent modeling of AU intensities and (ii) the state-of-the-art approach for the target task and (iii) deep convolutional neural networks. Robert Walecki, Ognjen Rudovic, Vladimir Pavlovic 0001, Maja Pantic |
IEEE Trans. Affect. Comput. | 2 |
| 2018 | CultureNet: A Deep Learning Approach for Engagement Intensity Estimation from Face Images of Children with AutismabstractMany children on autism spectrum have atypical behavioral expressions of engagement compared to their neu-rotypical peers. In this paper, we investigate the performance of deep learning models in the task of automated engagement estimation from face images of children with autism. Specifically, we use the video data of 30 children with different cultural backgrounds (Asia vs. Europe) recorded during a single session of a robot-assisted autism therapy. We perform a thorough evaluation of the proposed deep architectures for the target task, including within- and across-culture evaluations, as well as when using the child-independent and child-dependent settings. We also introduce a novel deep learning model, named CultureNet, which efficiently leverages the multi-cultural data when performing the adaptation of the proposed deep architecture to the target culture and child. We show that due to the highly heterogeneous nature of the image data of children with autism, the child-independent models lead to overall poor estimation of target engagement levels. On the other hand, when a small amount of data of target children is used to enhance the model learning, the estimation performance on the held-out data from those children increases significantly. This is the first time that the effects of individual and cultural differences in children with autism have empirically been studied in the context of deep learning performed directly from face images. Ognjen Rudovic, Yuria Utsumi, Jaeryoung Lee, Javier Hernandez, Eduardo Castelló Ferrer, Björn W. Schuller, Rosalind W. Picard |
IROS | 1 |
| 2018 | Multi-Instance Dynamic Ordinal Random Fields for Weakly Supervised Facial Behavior AnalysisabstractWe propose a Multi-Instance-Learning (MIL) approach for weakly-supervised learning problems, where a training set is formed by bags (sets of feature vectors or instances) and only labels at bag-level are provided. Specifically, we consider the Multi-Instance Dynamic-Ordinal-Regression (MI-DOR) setting, where the instance labels are naturally represented as ordinal variables and bags are structured as temporal sequences. To this end, we propose Multi-Instance Dynamic Ordinal Random Fields (MI-DORF). In this framework, we treat instance-labels as temporally-dependent latent variables in an Undirected Graphical Model. Different MIL assumptions are modelled via newly introduced high-order potentials relating bag and instance-labels within the energy function of the model. We also extend our framework to address the Partially-Observed MI-DOR problems, where a subset of instance labels are available during training.We show on the tasks of weakly-supervised facial behavior analysis, Facial Action Unit (DISFA dataset) and Pain (UNBC dataset) Intensity estimation, that the proposed framework outperforms alternative learning approaches. Furthermore, we show that MIDORF can be employed to reduce the data annotation efforts in this context by large-scale. Adria Ruiz, Ognjen Rudovic, Xavier Binefa, Maja Pantic |
IEEE Trans. Image Process. | 2 |
| 2017 | GIFGIF+: Collecting emotional animated GIFs with clustered multi-task learningabstractAnimated GIFs are widely used on the Internet to express emotions, but their automatic analysis is largely unexplored. Existing GIF datasets with emotion labels are too small for training contemporary machine learning models, so we propose a semi-automatic method to collect emotional animated GIFs from the Internet with the least amount of human labor. The method trains weak emotion recognizers on labeled data, and uses them to sort a large quantity of unlabeled GIFs. We found that by exploiting the clustered structure of emotions, the number of GIFs a labeler needs to check can be greatly reduced. Using the proposed method, a dataset called GIFGIF+ with 23,544 GIFs over 17 emotions was created, which provides a promising platform for affective computing research. Weixuan 'Vincent' Chen, Ognjen Rudovic, Rosalind W. Picard |
ACII | 2 |
| 2017 | Deep Structured Learning for Facial Action Unit Intensity EstimationabstractWe consider the task of automated estimation of facial expression intensity. This involves estimation of multiple output variables (facial action units - AUs) that are structurally dependent. Their structure arises from statistically induced co-occurrence patterns of AU intensity levels. Modeling this structure is critical for improving the estimation performance, however, this performance is bounded by the quality of the input features extracted from face images. The goal of this paper is to model these structures and estimate complex feature representations simultaneously by combining conditional random field (CRF) encoded AU dependencies with deep learning. To this end, we propose a novel Copula CNN deep learning approach for modeling multivariate ordinal variables. Our model accounts for ordinal structure in output variables and their non-linear dependencies via copula functions modeled as cliques of a CRF. These are jointly optimized with deep CNN feature encoding layers using a newly introduced balanced batch iterative training algorithm. We demonstrate the effectiveness of our approach on the task of AU intensity estimation on two benchmark datasets. We show that joint learning of the deep features and the target output structure results in significant performance gains compared to existing structured deep models and deep models for analysis of facial expressions. Robert Walecki, Ognjen Rudovic, Vladimir Pavlovic 0001, Björn W. Schuller, Maja Pantic |
CVPR | 2 |
| 2017 | PUnDA: Probabilistic Unsupervised Domain Adaptation for Knowledge Transfer Across Visual CategoriesabstractThis paper introduces a probabilistic latent variable model to address unsupervised domain adaptation problems. Specifically, we tackle the task of categorization of visual input from different domains by learning projections from each domain to a latent (shared) space jointly with the classifier in the latent space, which simultaneously minimizes the domain disparity while maximizing the classifier's discriminative power. Furthermore, the non-parametric nature of our adaptation model makes it possible to infer the latent space dimension automatically from data. We also develop a novel regularized Variational Bayes (VB) algorithm for efficient estimation of the model parameters. We compare the proposed model with the state-of-the-art methods for the tasks of visual domain adaptation using both handcrafted and deep-net features. Our experiments show that even with a simple softmax classifier, our model outperforms several state-of-the-art methods that take advantage of more sophisticated classification schemes. Behnam Gholami, Ognjen Rudovic, Vladimir Pavlovic 0001 |
ICCV | 2 |
| 2017 | DeepCoder: Semi-Parametric Variational Autoencoders for Automatic Facial Action CodingabstractHuman face exhibits an inherent hierarchy in its representations (i.e., holistic facial expressions can be encoded via a set of facial action units (AUs) and their intensity). Variational (deep) auto-encoders (VAE) have shown great results in unsupervised extraction of hierarchical latent representations from large amounts of image data, while being robust to noise and other undesired artifacts. Potentially, this makes VAEs a suitable approach for learning facial features for AU intensity estimation. Yet, most existing VAE-based methods apply classifiers learned separately from the encoded features. By contrast, the non-parametric. (probabilistic) approaches, such as Gaussian Processes (GPs), typically outperform their parametric counterparts, but cannot deal easily with large amounts of data. To this end, we propose a novel VAE semi-parametric modeling framework, named DeepCoder, which combines the modeling power of parametric (convolutional) and non-parametric. (ordinal GPs) VAEs, for joint learning of(l) latent representations at multiple levels in a task hierarchy1, and (2) classification of multiple ordinal outputs. We show on benchmark datasets for AU intensity estimation that the proposed DeepCoder outperforms the state-of-the-art approaches, and related VAEs and deep learning models. Dieu Linh Tran, Robert Walecki, Ognjen Rudovic, Stefanos Eleftheriadis, Björn W. Schuller, Maja Pantic |
ICCV | 3 |
| 2017 | Variable-state Latent Conditional Random Field models for facial expression analysis
Robert Walecki, Ognjen Rudovic, Vladimir Pavlovic 0001, Maja Pantic |
Image Vis. Comput. | 2 |
| 2017 | Gaussian Process Domain Experts for Modeling of Facial AffectabstractMost of existing models for facial behavior analysis rely on generic classifiers, which fail to generalize well to previously unseen data. This is because of inherent differences in source (training) and target (test) data, mainly caused by variation in subjects' facial morphology, camera views, and so on. All of these account for different contexts in which target and source data are recorded, and thus, may adversely affect the performance of the models learned solely from source data. In this paper, we exploit the notion of domain adaptation and propose a data efficient approach to adapt already learned classifiers to new unseen contexts. Specifically, we build upon the probabilistic framework of Gaussian processes (GPs), and introduce domain-specific GP experts (e.g., for each subject). The model adaptation is facilitated in a probabilistic fashion, by conditioning the target expert on the predictions from multiple source experts. We further exploit the predictive variance of each expert to define an optimal weighting during inference. We evaluate the proposed model on three publicly available data sets for multi-class (MultiPIE) and multi-label (DISFA, FERA2015) facial expression analysis by performing adaptation of two contextual factors: "where" (view) and "who" (subject). In our experiments, the proposed approach consistently outperforms: 1) both source and target classifiers, while using a small number of target examples during the adaptation and 2) related state-of-the-art approaches for supervised domain adaptation. Stefanos Eleftheriadis, Ognjen Rudovic, Marc Peter Deisenroth, Maja Pantic |
IEEE Trans. Image Process. | 2 |
| 2016 | Variational Gaussian Process Auto-Encoder for Ordinal Prediction of Facial Action Units
Stefanos Eleftheriadis, Ognjen Rudovic, Marc Peter Deisenroth, Maja Pantic |
ACCV (2) | 2 |
| 2016 | Multi-Instance Dynamic Ordinal Random Fields for Weakly-Supervised Pain Intensity Estimation
Adria Ruiz, Ognjen Rudovic, Xavier Binefa, Maja Pantic |
ACCV (2) | 2 |
| 2016 | Copula Ordinal Regression for Joint Estimation of Facial Action Unit IntensityabstractJoint modeling of the intensity of facial action units (AUs) from face images is challenging due to the large number of AUs (30+) and their intensity levels (6). This is in part due to the lack of suitable models that can efficiently handle such a large number of outputs/classes simultaneously, but also due to the lack of labelled target data. For this reason, majority of the methods proposed so far resort to independent classifiers for the AU intensity. This is suboptimal for at least two reasons: the facial appearance of some AUs changes depending on the intensity of other AUs, and some AUs co-occur more often than others. Encoding this is expected to improve the estimation of target AU intensities, especially in the case of noisy image features, head-pose variations and imbalanced training data. To this end, we introduce a novel modeling framework, Copula Ordinal Regression (COR), that leverages the power of copula functions and CRFs, to detangle the probabilistic modeling of AU dependencies from the marginal modeling of the AU intensity. Consequently, the COR model achieves the joint learning and inference of intensities of multiple AUs, while being computationally tractable. We show on two challenging datasets of naturalistic facial expressions that the proposed approach consistently outperforms (i) independent modeling of AU intensities, and (ii) the state-of the-art approach for the target task. Robert Walecki, Ognjen Rudovic, Vladimir Pavlovic 0001, Maja Pantic |
CVPR | 2 |
| 2016 | Multi-modal Neural Conditional Ordinal Random Fields for agreement level estimationabstractThe ability to automatically detect the extent of agreement or disagreement a person expresses is an important indicator of inter-personal relations and emotion expression. Most of existing methods for automated analysis of human agreement from audio-visual data perform agreement detection using either audio or visual modality of human interactions. However, this is suboptimal as expression of different agreement levels is composed of various facial and vocal cues specific to the target level. To this end, we propose the first approach for multi-modal estimation of agreement intensity levels. Specifically, our model leverages the feature representation power of Multi-modal Neural Networks (NN) and discriminative power of Conditional Ordinal Random Fields (CORF) to achieve dynamic classification of agreement levels from videos. We show on the MAHNOB-Mimicry database of dyadic human interactions that the proposed approach outperforms its uni-modal and linear counterparts, and related models that can be applied to the target task. Nemanja Rakicevic, Ognjen Rudovic, Stavros Petridis, Maja Pantic |
ICPR | 2 |
| 2016 | Joint Facial Action Unit Detection and Feature Fusion: A Multi-Conditional Learning ApproachabstractAutomated analysis of facial expressions can benefit many domains, from marketing to clinical diagnosis of neurodevelopmental disorders. Facial expressions are typically encoded as a combination of facial muscle activations, i.e., action units. Depending on context, these action units co-occur in specific patterns, and rarely in isolation. Yet, most existing methods for automatic action unit detection fail to exploit dependencies among them, and the corresponding facial features. To address this, we propose a novel multi-conditional latent variable model for simultaneous fusion of facial features and joint action unit detection. Specifically, the proposed model performs feature fusion in a generative fashion via a low-dimensional shared subspace, while simultaneously performing action unit detection using a discriminative classification approach. We show that by combining the merits of both approaches, the proposed methodology outperforms existing purely discriminative/generative methods for the target task. To reduce the number of parameters, and avoid overfitting, a novel Bayesian learning approach based on Monte Carlo sampling is proposed, to integrate out the shared subspace. We validate the proposed method on posed and spontaneous data from three publicly available datasets (CK+, DISFA and Shoulder-pain), and show that both feature fusion and joint learning of action units leads to improved performance compared to the state-of-the-art methods for the task. Stefanos Eleftheriadis, Ognjen Rudovic, Maja Pantic |
IEEE Trans. Image Process. | 2 |
| 2015 | Neural conditional ordinal random fields for agreement level estimationabstractWe present a novel approach to automated estimation of agreement intensity levels from facial images. To this end, we employ the MAHNOB Mimicry database of subjects recorded during dyadic interactions, where the facial images are annotated in terms of agreement intensity levels using the Likert scale (strong disagreement, disagreement, neutral, agreement and strong agreement). Dynamic modelling of the agreement levels is accomplished by means of a Conditional Ordinal Random Field model. Specifically, we propose a novel Neural Conditional Ordinal Random Field model that performs non-linear feature extraction from face images using the notion of Neural Networks, while also modelling temporal and ordinal relationships between the agreement levels. We show in our experiments that the proposed approach outperforms existing methods for modelling of sequential data. The preliminary results obtained on five subjects demonstrate that the intensity of agreement can successfully be estimated from facial images (39% F1 score) using the proposed method. Nemanja Rakicevic, Ognjen Rudovic, Stavros Petridis, Maja Pantic |
ACII | 2 |
| 2015 | Sentiment apprehension in human-robot interaction with NAOabstractThe ability of robots to interact in a socially intelligent manner with humans is the core of human-robot interaction (HRI). The quality of this interaction is typically measured in terms of how it is engaging to the users either reflected in duration of time users spend interacting with a robot, or their self-reports on engagement during the interaction. In contrast to existing studies that analyze the influence of robots' ability to mimic affective states (happy or sad) of users on their engagement, in this paper we study the influence of sentiment apprehension by robots (i.e., robot's ability to reason about the user's attitudes such as judgment / liking) on the user engagement. Specifically, we present the findings from our pilot study on the effect of sentiment apprehension in HRI using NAO robot. In this study, we analyzed two versions of mimicry game: in the first, NAO was solely mimicking facial expressions of the users, while in the second he was also providing a feedback based on the sentiment apprehension. A total of 32 participants (7 female, 25 male) were recruited for this experiment, and the results show that the participants in the second group spent more time interacting with the robot and played more rounds of the mimicry game. After experiencing both versions of the game, ratings given by the participants indicate (with 99% confidence) that the game with sentiment apprehension is more engaging than the baseline version. Jie Shen 0008, Ognjen Rudovic, Shiyang Cheng 0001, Maja Pantic |
ACII | 2 |
| 2015 | Multi-conditional Latent Variable Model for Joint Facial Action Unit DetectionabstractWe propose a novel multi-conditional latent variable model for simultaneous facial feature fusion and detection of facial action units. In our approach we exploit the structure-discovery capabilities of generative models such as Gaussian processes, and the discriminative power of classifiers such as logistic function. This leads to superior performance compared to existing classifiers for the target task that exploit either the discriminative or generative property, but not both. The model learning is performed via an efficient, newly proposed Bayesian learning strategy based on Monte Carlo sampling. Consequently, the learned model is robust to data overfitting, regardless of the number of both input features and jointly estimated facial action units. Extensive qualitative and quantitative experimental evaluations are performed on three publicly available datasets (CK+, Shoulder-pain and DISFA). We show that the proposed model outperforms the state-of-the-art methods for the target task on (i) feature fusion, and (ii) multiple facial action unit detection. Stefanos Eleftheriadis, Ognjen Rudovic, Maja Pantic |
ICCV | 2 |
| 2015 | Context-Sensitive Dynamic Ordinal Regression for Intensity Estimation of Facial Action UnitsabstractModeling intensity of facial action units from spontaneously displayed facial expressions is challenging mainly because of high variability in subject-specific facial expressiveness, head-movements, illumination changes, etc. These factors make the target problem highly context-sensitive. However, existing methods usually ignore this context-sensitivity of the target problem. We propose a novel Conditional Ordinal Random Field (CORF) model for context-sensitive modeling of the facial action unit intensity, where the W5+ (who, when, what, where, why and how) definition of the context is used. While the proposed model is general enough to handle all six context questions, in this paper we focus on the context questions: who (the observed subject), how (the changes in facial expressions), and when (the timing of facial expressions and their intensity). The context questions who and howare modeled by means of the newly introduced context-dependent covariate effects, and the context question when is modeled in terms of temporal correlation between the ordinal outputs, i.e., intensity levels of action units. We also introduce a weighted softmax-margin learning of CRFs from data with skewed distribution of the intensity levels, which is commonly encountered in spontaneous facial data. The proposed model is evaluated on intensity estimation of pain and facial action units using two recently published datasets (UNBC Shoulder Pain and DISFA) of spontaneously displayed facial expressions. Our experiments show that the proposed model performs significantly better on the target tasks compared to the state-of-the-art approaches. Furthermore, compared to traditional learning of CRFs, we show that the proposed weighted learning results in more robust parameter estimation from the imbalanced intensity data. Ognjen Rudovic, Vladimir Pavlovic 0001, Maja Pantic |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Discriminative Shared Gaussian Processes for Multiview and View-Invariant Facial Expression RecognitionabstractImages of facial expressions are often captured from various views as a result of either head movements or variable camera position. Existing methods for multiview and/or view-invariant facial expression recognition typically perform classification of the observed expression using either classifiers learned separately for each view or a single classifier learned for all views. However, these approaches ignore the fact that different views of a facial expression are just different manifestations of the same facial expression. By accounting for this redundancy, we can design more effective classifiers for the target task. To this end, we propose a discriminative shared Gaussian process latent variable model (DS-GPLVM) for multiview and view-invariant classification of facial expressions from multiple views. In this model, we first learn a discriminative manifold shared by multiple views of a facial expression. Subsequently, we perform facial expression classification in the expression manifold. Finally, classification of an observed facial expression is carried out either in the view-invariant manner (using only a single view of the expression) or in the multiview manner (using multiple views of the expression). The proposed model can also be used to perform fusion of different facial features in a principled manner. We validate the proposed DS-GPLVM on both posed and spontaneously displayed facial expressions from three publicly available datasets (MultiPIE, labeled face parts in the wild, and static facial expressions in the wild). We show that this model outperforms the state-of-the-art methods for multiview and view-invariant facial expression classification, and several state-of-the-art methods for multiview learning and feature fusion. Stefanos Eleftheriadis, Ognjen Rudovic, Maja Pantic |
IEEE Trans. Image Process. | 2 |
| 2014 | Corrigendum to "Hierarchical On-line Appearance-Based Tracking for 3D Head Pose, Eyebrows, Lips, Eyelids and Irises" [Image Vision Comput. (2013) 322-340]
Javier Orozco, Ognjen Rudovic, Jordi Gonzàlez 0001, Maja Pantic |
Image Vis. Comput. | 2 |
| 2013 | Bimodal log-linear regression for fusion of audio and visual featuresabstractOne of the most commonly used audiovisual fusion approaches is feature-level fusion where the audio and visual features are concatenated. Although this approach has been successfully used in several applications, it does not take into account interactions between the features, which can be a problem when one and/or both modalities have noisy features. In this paper, we investigate whether feature fusion based on explicit modelling of interactions between audio and visual features can enhance the performance of the classifier that performs feature fusion using simple concatenation of the audio-visual features. To this end, we propose a log-linear model, named Bimodal Log-linear regression, which accounts for interactions between the features of the two modalities. The performance of the target classifiers is measured in the task of laughter-vs-speech discrimination, since both laughter and speech are naturally audiovisual events. Our experiments on the MAHNOB laughter database suggest that feature fusion based on explicit modelling of interactions between the audio-visual features leads to an improvement of 3\% over the standard feature concatenation approach, when log-linear model is used as the base classifier. Finally, the most and least influential features can be easily identified by observing their interactions. Ognjen Rudovic, Stavros Petridis, Maja Pantic |
ACM Multimedia | 1 |
| 2013 | Hierarchical On-line Appearance-Based Tracking for 3D head pose, eyebrows, lips, eyelids and irises
Javier Orozco, Ognjen Rudovic, Jordi Gonzàlez 0001, Maja Pantic |
Image Vis. Comput. | 2 |
| 2013 | Coupled Gaussian Processes for Pose-Invariant Facial Expression RecognitionabstractWe propose a method for head-pose invariant facial expression recognition that is based on a set of characteristic facial points. To achieve head-pose invariance, we propose the Coupled Scaled Gaussian Process Regression (CSGPR) model for head-pose normalization. In this model, we first learn independently the mappings between the facial points in each pair of (discrete) nonfrontal poses and the frontal pose, and then perform their coupling in order to capture dependences between them. During inference, the outputs of the coupled functions from different poses are combined using a gating function, devised based on the head-pose estimation for the query points. The proposed model outperforms state-of-the-art regression-based approaches to head-pose normalization, 2D and 3D Point Distribution Models (PDMs), and Active Appearance Models (AAMs), especially in cases of unknown poses and imbalanced training data. To the best of our knowledge, the proposed method is the first one that is able to deal with expressive faces in the range from -45° to +45° pan rotation and -30° to +30° tilt rotation, and with continuous changes in head pose, despite the fact that training was conducted on a small set of discrete poses. We evaluate the proposed method on synthetic and real images depicting acted and spontaneously displayed facial expressions. Ognjen Rudovic, Maja Pantic, Ioannis Patras |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Multi-output Laplacian dynamic ordinal regression for facial expression recognition and intensity estimationabstractAutomated facial expression recognition has received increased attention over the past two decades. Existing works in the field usually do not encode either the temporal evolution or the intensity of the observed facial displays. They also fail to jointly model multidimensional (multi-class) continuous facial behaviour data; binary classifiers - one for each target basic-emotion class - are used instead. In this paper, intrinsic topology of multidimensional continuous facial affect data is first modeled by an ordinal manifold. This topology is then incorporated into the Hidden Conditional Ordinal Random Field (H-CORF) framework for dynamic ordinal regression by constraining H-CORF parameters to lie on the ordinal manifold. The resulting model attains simultaneous dynamic recognition and intensity estimation of facial expressions of multiple emotions. To the best of our knowledge, the proposed method is the first one to achieve this on both deliberate as well as spontaneous facial affect data. Ognjen Rudovic, Vladimir Pavlovic 0001, Maja Pantic |
CVPR | 1 |
| 2011 | Shape-constrained Gaussian process regression for facial-point-based head-pose normalizationabstractGiven the facial points extracted from an image of a face in an arbitrary pose, the goal of facial-point-based head-pose normalization is to obtain the corresponding facial points in a predefined pose (e.g., frontal). This involves inference of complex and high-dimensional mappings due to the large number of the facial points employed, and due to differences in head-pose and facial expression. Most regression-based approaches for learning such mappings focus on modeling correlations only between the inputs (i.e., the facial points in a non-frontal pose) and the outputs (i.e., the facial points in the frontal pose), but not within the inputs and the outputs of the model. This makes these models prone to errors due to noise and outliers in test data, often resulting in anatomically impossible facial configurations formed by their predictions. To address this, we propose Shape-constrained Gaussian Process (SC-GP) regression for facial-point-based head-pose normalization. Specifically, a deformable face-shape model is used to learn a face-shape prior, which is placed on both the input and the output of GP regression in order to constrain the model predictions to anatomically feasible facial configurations. Our extensive experiments on both synthetic and real image data show that the proposed approach generalizes well across poses and handles successfully noise and outliers in test data. In addition, the proposed model outperforms previously proposed approaches to facial-point-based head-pose normalization. Ognjen Rudovic, Maja Pantic |
ICCV | 1 |
| 2010 | Coupled Gaussian Process Regression for Pose-Invariant Facial Expression Recognition
Ognjen Rudovic, Ioannis Patras, Maja Pantic |
ECCV (2) | 1 |
| 2010 | Regression-Based Multi-view Facial Expression RecognitionabstractWe present a regression-based scheme for multi-view facial expression recognition based on 2D geometric features. We address the problem by mapping facial points (e.g. mouth corners) from non-frontal to frontal view where further recognition of the expressions can be performed using a state-of-the-art facial expression recognition method. To learn the mapping functions we investigate four regression models: Linear Regression (LR), Support Vector Regression (SVR), Relevance Vector Regression (RVR) and Gaussian Process Regression (GPR). Our extensive experiments on the CMU Multi-PIE facial expression database show that the proposed scheme outperforms view-specific classifiers by utilizing considerably less training data. Ognjen Rudovic, Ioannis Patras, Maja Pantic |
ICPR | 1 |
| 2008 | View-invariant human-body detection with extension to human action recognition using component-wise HMM of body partsabstractThis paper presents a technique for view invariant human detection and extending this idea to recognize basic human actions like walking, jogging, hand waving and boxing etc. To achieve this goal we detect the human in its body parts and then learn the changes of those body parts for action recognition. Human-body part detection in different views is an extremely challenging problem due to drastic change of 3D-pose of human body, self occlusions etc while performing actions. In order to cope with these problems we have designed three example-based detectors that are trained to find separately three components of the human body, namely the head, legs and arms. We incorporate 10 sub-classifiers for the head, arms and the leg detection. Each sub-classifier detects body parts under a specific range of viewpoints. Then, view-invariance is fulfilled by combining the results of these sub classifiers. Subsequently, we extend this approach to recognize actions based on component-wise hidden Markov models (HMM). This is achieved by designing a HMM for each action, which is trained based on the detected body parts. Consequently, we are able to distinguish between similar actions by only considering the body parts which has major contributions to those actions e.g. legs for walking, running etc; hands for boxing, waving etc. Bhaskar Chakraborty, Ognjen Rudovic, Jordi Gonzàlez 0001 |
FG | 2 |
| 2008 | Confidence assessment on eyelid and eyebrow expression recognitionabstractIn this paper, we address the recognition of subtle facial expressions by reasoning on the classification confidence. Psychological evidences have determined that eyelids and eyebrows are significant for the recognition of subtle facial expressions and the early perception of human emotions. This early perception results in a more complex problem, which requires a confidence assessment for any provided solution. Thus, traditional score-based classifiers (e.g. k-NN and NN) are not able to produce confident estimates. Instead, we first present five confidence estimators and a confidence classification assessment for Case-Based Reasoning (CBR). Second, we improve the expression retrieval from the database by learning the neighbourhood's dimensions for the expected classification confidences. Third, we reuse the previous classified expressions and the confidence assessment to improve the classification achieved by k-NN. Fourth, we improve the database for generalization with new subjects by learning thresholds to minimize misclassification with low confidence, maximize correct classifications with high confidence and re-arrange misclassification with high confidence. The proposed system represents an effective contribution for both subtle expression recognition and CBR methodology. It achieves an average recognition of 97% plusmn 1% with a confidence of 96% plusmn 2% for expressiveness between 20% and 100%. Javier Orozco, Ognjen Rudovic, F. Xavier Roca, Jordi Gonzàlez 0001 |
FG | 2 |