VLDB 2026 Research / reviewers in the wild / expert
Albert Ali Salah
dblp:75/4848
· DBLP profile ↗
68ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0001-6342-428XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 13 · 1 first-author · 7 since 2021Security and privacy · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explicitly modeling trajectories and correlations for video analysisabstractVideo analysis tasks such as action recognition and sign language recognition require that both the appearance and dynamics of objects and people of interest are captured. Extraction of distinctive local appearance changes of regions under motion, such as finger movement while gesturing, is often essential to make correct classifications. Without disentangling gross and fine spatio-temporal patterns, both will be confounded, leading to reduced accuracy. We introduce two innovations to tackle this issue. First, we explicitly model motion of image patches by temporally aligning visual tokens using optical flow. In addition to obtaining a coarse motion representation, for a query token, we apply self-attention along the trajectory. This essentially cancels out gross movement and makes possible the extraction of distinctive local patterns of regions under motion. Our second innovation is a dynamic attention mechanism that filters out irrelevant frame regions. It assigns dynamic key–value tokens from correlated regions to each query to focus on coordinated appearance changes such as joint hand and mouth movements while gesturing. We combine these innovations in a novel Trajectory and Correlation (TC) block, a hybrid network that effectively models spatio-temporal information from trajectories and correlated regions. We experiment on four sign language (PHOENIX14, PHOENIX14-T, CSL, and CSL-Daily) and two action recognition (Kinetics-400 and Something-Something V2) datasets. Using TC blocks in different backbones consistently achieves improved performance with a modest increase in parameters, and 30% additional computational cost. Albert Ali Salah, Ronald Poppe |
Image Vis. Comput. | 2 |
| 2026 | Improving the generalization of ViTs for action understanding with VLM pre-trainingabstractOwing to their ability to extract powerful video embeddings, Vision Transformers (ViTs) are currently the best performing models in video action understanding. However, when these models are frozen and applied to downstream tasks, their performance drops significantly, revealing limited generalization. In this paper, we describe the Four-Tiered Prompts (FTP) framework that introduces feature processors to transform the ViT’s output. In a pre-training stage, each feature processor is trained using contrastive learning to align the ViT’s visual embeddings with a vision language model’s (VLM) textual embeddings. We use four feature processors, each linked to the output of a VLM prompt that reflects the fundamental aspects of human action: category, components, description, and context. With the FTP framework, we increase the ViT’s generalization ability by forcing the visual encoder to incorporate relevant, semantic information. Importantly, we only employ the VLM during training. Subsequently, inference incurs a limited computation cost. For video action recognition and detection, employing the FTP framework consistently yields state-of-the-art performance after fine-tuning. Extensive experiments demonstrate how different choices contribute to the overall increase in performance. 1 Albert Ali Salah, Ronald Poppe |
Pattern Recognit. | 2 |
| 2025 | Revisiting Representation Learning and Identity Adversarial Training for Facial Behavior UnderstandingabstractFacial Action Unit (AU) detection has gained significant attention as it enables the breakdown of complex facial expressions into individual muscle movements. In this paper, we revisit two fundamental factors in AU detection: diverse and large-scale data and subject identity regularization. Motivated by recent advances in foundation models, we highlight the importance of data and introduce Face 9 M, a diverse dataset comprising 9 million facial images from multiple public sources. Pretraining a masked autoencoder on Face9M yields strong performance in AU detection and facial expression tasks. More importantly, we emphasize that the Identity Adversarial Training (IAT) has not been well explored in AU tasks. To fill this gap, we first show that subject identity in AU datasets creates shortcut learning for the model and leads to suboptimal solutions to AU predictions. Secondly, we demonstrate that strong IAT regularization is necessary to learn identityinvariant features. Finally, we elucidate the design space of IAT and empirically show that IAT circumvents the identity-based shortcut learning and results in a better solution. Our proposed methods, Facial Masked Autoencoder (FMAE) and IAT, are simple, generic and effective. Remarkably, the proposed FMAEIAT approach achieves new state-of-the-art F1 scores on BP4D ($67.1 \%$), BP4D+ ($66.8 \%$), and DISFA ($70.1 \%$) databases, significantly outperforming previous work. We release the code and model at https://github.com/forever208/FMAE-IAT. Mang Ning, Albert Ali Salah, Itir Önal |
FG | 2 |
| 2025 | Snakes and Ladders: Two Steps Up for VideoMamba
Albert Ali Salah, Ronald Poppe |
ICCV | 2 |
| 2025 | DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT SpaceabstractThis paper explores image modeling from the frequency space and introduces DCTdiff, an end-to-end diffusion generative paradigm that efficiently models images in the discrete cosine transform (DCT) space. We investigate the design space of DCTdiff and reveal the key design factors. Experiments on different frameworks (UViT, DiT), generation tasks, and various diffusion samplers demonstrate that DCTdiff outperforms pixel-based diffusion models regarding generative quality and training efficiency. Remarkably, DCTdiff can seamlessly scale up to 512$\times$512 resolution without using the latent diffusion paradigm and beats latent diffusion (using SD-VAE) with only 1/4 training cost. Finally, we illustrate several intriguing properties of DCT image modeling. For example, we provide a theoretical proof of why `image diffusion can be seen as spectral autoregression', bridging the gap between diffusion and autoregressive models. The effectiveness of DCTdiff and the introduced properties suggest a promising direction for image modeling in the frequency space. The code is at https://github.com/forever208/DCTdiff. Mang Ning, Mingxiao Li 0002, Jianlin Su, Haozhe Jia, Lanmiao Liu, Martin Benes 0001, Wenshuo Chen, Albert Ali Salah, Itir Önal |
ICML | 8 |
| 2024 | TCNet: Continuous Sign Language Recognition from Trajectories and Correlated RegionsabstractA key challenge in continuous sign language recognition (CSLR) is to efficiently capture long-range spatial interactions over time from the video input. To address this challenge, we propose TCNet, a hybrid network that effectively models spatio-temporal information from Trajectories and Correlated regions. TCNet's trajectory module transforms frames into aligned trajectories composed of continuous visual tokens. This facilitates extracting region trajectory patterns. In addition, for a query token, self-attention is learned along the trajectory. As such, our network can also focus on fine-grained spatio-temporal patterns, such as finger movement, of a region in motion. TCNet's correlation module utilizes a novel dynamic attention mechanism that filters out irrelevant frame regions. Additionally, it assigns dynamic key-value tokens from correlated regions to each query. Both innovations significantly reduce the computation cost and memory. We perform experiments on four large-scale datasets: PHOENIX14, PHOENIX14-T, CSL, and CSL-Daily. Our results demonstrate that TCNet consistently achieves state-of-the-art performance. For example, we improve over the previous state-of-the-art by 1.5\% and 1.0\% word error rate on PHOENIX14 and PHOENIX14-T, respectively. Code is available at https://github.com/hotfinda/TCNet Albert Ali Salah, Ronald Poppe |
AAAI | 2 |
| 2024 | Fairness in AI-Based Mental Health: Clinician Perspectives and Bias MitigationabstractThere is limited research on fairness in automated decision-making systems in the clinical domain, particularly in the mental health domain. Our study explores clinicians' perceptions of AI fairness through two distinct scenarios: violence risk assessment and depression phenotype recognition using textual clinical notes. We engage with clinicians through semi-structured interviews to understand their fairness perceptions and to identify appropriate quantitative fairness objectives for these scenarios. Then, we compare a set of bias mitigation strategies developed to improve at least one of the four selected fairness objectives. Our findings underscore the importance of carefully selecting fairness measures, as prioritizing less relevant measures can have a detrimental rather than a beneficial effect on model behavior in real-world clinical use. Gizem Sogancioglu, Pablo Mosteiro, Albert Ali Salah, Floor Scheepers, Heysem Kaya |
AIES (1) | 3 |
| 2024 | Compensation Sampling for Improved Convergence in Diffusion Models
Albert Ali Salah, Ronald Poppe |
ECCV (61) | 2 |
| 2024 | WelcomeabstractIt was our pleasure and privilege to welcome you to Istanbul for the 18th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2024). We hope your experience at FG was rewarding both professionally and personally! Hazim Kemal Ekenel, Albert Ali Salah, Arun Ross, Vitomir Struc, Lale Akarun, Xilin Chen 0001, Shaun J. Canavan |
FG | 2 |
| 2024 | Survey of Automated Methods for Nonverbal Behavior Analysis in Parent-Child InteractionsabstractSocial interactions are fundamental for human beings, motivating the abundance of studies into the behavioral correlates of constructs such as personality and relationship. The primary drivers of this research are video-taped recordings of interactions. Recent advancements in automatic behavior analysis provide a cost-effective and more objective alternative to manual coding by trained experts. Still, the use of automated analysis is far from trivial. In this literature survey, we discuss the current state-of-the-art in automated parent-child interaction analysis, and critically assess opportunities and limitations. We focus on parent-child interactions as they reflect various aspects of a child's development, and provide distinct challenges for the automated measurement and interpretation of the interactive behavior. We briefly discuss single-person and dyadic nonverbal measurements, and identify measurement challenges. We then provide an overview of various developmental constructs that can be measured through the classification of extracted cues. Finally, we outline persistent limitations of the current state-of-the-art, and we highlight promising directions to bridge the gap between manual and automated measurements. Berfu Karaca, Albert Ali Salah, Jaap Denissen, Ronald Poppe, Sonja M. C. de Zwarte |
FG | 2 |
| 2024 | Elucidating the Exposure Bias in Diffusion ModelsabstractDiffusion models have demonstrated impressive generative capabilities, but their exposure bias problem, described as the input mismatch between training and sampling, lacks in-depth exploration. In this paper, we investigate the exposure bias problem in diffusion models by first analytically modelling the sampling distribution, based on which we then attribute the prediction error at each sampling step as the root cause of the exposure bias issue. Furthermore, we discuss potential solutions to this issue and propose an intuitive metric for it. Along with the elucidation of exposure bias, we propose a simple, yet effective, training-free method called Epsilon Scaling to alleviate the exposure bias. We show that Epsilon Scaling explicitly moves the sampling trajectory closer to the vector field learned in the training phase by scaling down the network output, mitigating the input mismatch between training and sampling. Experiments on various diffusion frameworks (ADM, DDIM, EDM, LDM, DiT, PFGM++) verify the effectiveness of our method. Remarkably, our ADM-ES, as a state-of-the-art stochastic sampler, obtains 2.17 FID on CIFAR-10 under 100-step unconditional generation. The code is at https://github.com/forever208/ADM-ES Mang Ning, Mingxiao Li 0002, Jianlin Su, Albert Ali Salah, Itir Önal |
ICLR | 4 |
| 2024 | Decoding Contact: Automatic Estimation of Contact Signatures in Parent-Infant Free Play InteractionsabstractIn parent-child interactions (PCIs), there is frequent physical contact between the two actors. Quantifying this contact provides valuable input to assess the nature of the interaction or the relation between parent and child. Here, we explore the application of vision-based techniques to automatically detect contact signatures at each frame of video recordings of playful parent-infant interactions. We employ two separate models: (i) a multimodal convolutional neural network (CNN) that integrates 2D pose and body part information, and (ii) a unimodal graph convolutional neural network (GCN) that utilizes only 2D pose. We showcase the potential and limitations of automatic contact signature estimation through quantitative and qualitative assessments using a parent-infant free play interaction dataset consisting of 100 parent-child dyadic interactions, covering 20 hours. Additionally, our experiments provide insights into various design choices through systematic experimentation. By releasing our annotations and code, we aim to enable further research in the automatic contact signature estimation during free play interactions between parents and infants. Metehan Doyran, Albert Ali Salah, Ronald Poppe |
ICMI | 2 |
| 2023 | Automated Emotional Valence Estimation in Infants with Stochastic and Strided Temporal SamplingabstractWe propose the first automated approach to estimate the emotional valence of infants from their facial behavior. We use the state-of-the-art transformer-based video masked autoencoder (VideoMAE) that is pre-trained on a large video dataset as a backbone, and finetune it on two large, well-annotated infant video datasets (SIBSMILE and MODELING). To augment the limited data, we propose a novel video temporal augmentation method called Stochastic and Strided Temporal Sampling (SSTS). We demonstrate the effectiveness of our approach for infant valence estimation by achieving 0.671 Concordance Correlation Coefficient (CCC) on SIBSMILE and MODELING. The experiments show that SSTS remarkably accelerates the training speed by 8 times while gaining the best valence estimation performance. Lastly, we suggest that face detection and cropping (coarse registration) is a promising alternative to landmark-based registration (i.e. fine registration) in data pre-processing when accurate infant facial landmark detectors are inaccessible. Mang Ning, Itir Önal, Daniel S. Messinger, Jeffrey F. Cohn, Albert Ali Salah |
ACII | 5 |
| 2023 | The effects of gender bias in word embeddings on patient phenotyping in the mental health domainabstractWord embeddings, renowned for their role as superior semantic feature vector representation in diverse NLP tasks, can exhibit an undesired bias for stereotypical categories. The bias arises from the statistical and societal biases within the datasets used for training. In this study, we analyze the gender bias in four different pre-trained word embeddings for a range of affective computing tasks in the mental health domain including the detection of psychiatric disorders such as depression, and alcohol/substance abuse. We incorporate both contextual and non-contextual embeddings, which are trained not just on general domain data but also on data specific to the clinical domain. Our findings indicate that the bias in embeddings is towards different gender groups, depending on the type of embeddings and the training dataset. Furthermore, we highlight how these existing associations transfer to subsequent tasks and might even be amplified during supervised training for patient phenotyping. We also show that a simple method of data augmentation- swapping gender words - noticeably reduces bias in these subsequent tasks. The scripts to reproduce the results are available at: https:llgithub.comlgizemsogancioglulgender-bias-mental-health. Gizem Sogancioglu, Heysem Kaya, Albert Ali Salah |
ACII | 3 |
| 2023 | Video-based estimation of pain indicators in dogsabstractDog owners are typically capable of recognizing behavioral cues that reveal subjective states of their dogs, such as pain. But automatic recognition of the pain state is very challenging. This paper proposes a novel video-based, two-stream deep neural network approach for this problem. We extract and preprocess body keypoints, and compute features from both keypoints and the RGB representation over the video. We propose an approach to deal with self-occlusions and missing keypoints. We also present a unique video-based dog behavior dataset, collected by veterinary professionals, and annotated for presence of pain, and report good classification results with the proposed approach. This study is one of the first works on machine learning based estimation of dog pain state. Code is available at https://github.con/s04240051/pain_detection Yasemin Salgirli, Pinar Can, Durmus Atilgan, Albert Ali Salah |
ACII | 5 |
| 2023 | SCFormer: Integrating hybrid Features in Vision TransformersabstractHybrid modules that combine self-attention and convolution operations can benefit from the advantages of both, and consequently achieve higher performance than either operation alone. However, current hybrid modules do not capitalize directly on the intrinsic relation between self-attention and convolution, but rather introduce external mechanisms that come with increased computation cost. In this paper, we propose a new hybrid vision transformer called Shift and Concatenate Transformer(SCFormer), which benefits from the intrinsic relationship between convolution and self-attention. SCFormer roots in the Shift and Concatenate Attention (SCA) block, that integrates convolution and self-attention features. We propose a shifting mechanism and corresponding aggregation rules for the feature integration of SCA blocks such that generated features more closely approximate the optimal output features. Extensive experiments show that, with comparable computational complexity, SCFormer consistently achieves improved results over competitive baselines on image recognition and downstream tasks. Our code is available at: https://github.com/hotfinda/SCFormer. Ronald Poppe, Albert Ali Salah |
ICME | 3 |
| 2023 | LA-layer: General local attention layer for full attention networksabstractAttention layers have contributed to state-of-the-art results on vision tasks. Still, they leave room for improvement because position information is used in a fixed manner, and the computation cost is typically high. To mitigate both issues, we propose a convolution-style local attention layer (LA-layer) as a replacement for traditional attention layers. LA-layers not only encode the position information of pixels in a convolutional manner, but also produce position offsets following a novel constrained rule so that keys will deform and result in larger receptive fields. Query and keys are processed by a novel aggregation function that outputs attention weights for the values. In our experiments with different types of ResNets, we replace convolutional layers with LA-layers and address image recognition, object detection and instance segmentation tasks. We consistently demonstrate performance gains, despite having fewer FLOPs and training parameters. Our code is available at: https://github.com/hotfinda/LA-layer. Ronald Poppe, Albert Ali Salah |
ICME | 3 |
| 2023 | Going Deeper than Tracking: A Survey of Computer-Vision Based Recognition of Animal Pain and EmotionsabstractAbstract Advances in animal motion tracking and pose recognition have been a game changer in the study of animal behavior. Recently, an increasing number of works go ‘deeper’ than tracking, and address automated recognition of animals’ internal states such as emotions and pain with the aim of improving animal welfare, making this a timely moment for a systematization of the field. This paper provides a comprehensive survey of computer vision-based research on recognition of pain and emotional states in animals, addressing both facial and bodily behavior analysis. We summarize the efforts that have been presented so far within this topic—classifying them across different dimensions, highlight challenges and research gaps, and provide best practice recommendations for advancing the field, and some future directions for research. Sofia Broomé, Marcelo Feighelstein, Anna Zamansky, Gabriel Carreira Lencioni, Pia Haubro Andersen, Francisca Pessanha, Marwa Mahmoud, Hedvig Kjellström, Albert Ali Salah |
Int. J. Comput. Vis. | 9 |
| 2023 | A survey on computer vision based human analysis in the COVID-19 era
Fevziye Irem Eyiokur, Alperen Kantarci, Mustafa Ekrem Erakin, Naser Damer, Ferda Ofli, Muhammad Imran 0002, Janez Krizaj, Albert Ali Salah, Alex Waibel, Vitomir Struc, Hazim Kemal Ekenel |
Image Vis. Comput. | 8 |
| 2023 | Facial Image-Based Automatic Assessment of Equine PainabstractRecognition of pain in animals is essential for their welfare. However, since there is no verbal communication, this assessment depends solely on the ability of the observer to locate visible or audible signs of pain. The use of grimace scales is proven to be efficient in detecting the pain visually, but the assessment quality depends on the level of training of the assessor and the validity is not easily ensured. There is a clear need for automating the pain assessment process. This work provides a system for pain prediction in horses, based on grimace scales. The pipeline automatically determines the quantitative pose of the equine head and finds facial landmarks before classification, proposing a novel scale-normalisation approach for equine heads. The pain estimation is achieved for each facial region of interest separately, following the clinical pain estimation procedure. We introduce a database of horse images, annotated by professional veterinarians for training and assessment. We also propose a data augmentation method to alleviate the data scarcity issues, which relies on generating realistic 3D equine face models based on 2D annotated images. We show that the data augmentation method improves the performance of both quantitative pose estimation and landmark detection. Our results establish a strong baseline for automatic equine pain estimation. Francisca Pessanha, Albert Ali Salah, Thijs J. P. A. M. van Loon, Remco C. Veltkamp |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Towards using Breathing Features for Multimodal Estimation of Depression SeverityabstractBreathing patterns are shown to have strong correlations with emotional states, and hence have promise for automatic mood order prediction and analysis. An essential challenge here is the lack of ground truth for breathing sounds, especially for medical and archival datasets. In this study, we provide a cross-dataset approach for breathing pattern prediction and analyse the contribution of predicted breath signals for the detection of depressive states, using the DAIC-WOZ corpus. We use interpretable features in our models to provide actionable insights. Our experimental evaluation shows that in participants with higher depression scores (as indicated by the eight-item Patient Health Questionnaire, PHQ-8), breathing events tend to be shallow or slow. We furthermore tested linear and non-linear regression models with breathing, linguistic sentiment and conversational features, and show that these simple models outperform the AVEC17 Real-life Depression Recognition Sub-challenge baseline. Francisca Pessanha, Heysem Kaya, Almila Akdag Salah, Albert Ali Salah |
ICMI | 4 |
| 2022 | Federated learning for violence incident prediction in a simulated cross-institutional psychiatric settingabstractInpatient violence is a common and severe problem within psychiatry. Knowing who might become violent can influence staffing levels and mitigate severity. Predictive machine learning models can assess each patient’s likelihood of becoming violent based on clinical notes. Yet, while machine learning models benefit from having more data, data availability is limited as hospitals typically do not share their data for privacy preservation. Federated Learning (FL) can overcome the problem of data limitation by training models in a decentralised manner, without disclosing data between collaborators. However, although several FL approaches exist, none of these train Natural Language Processing models on clinical notes. In this work, we investigate the application of Federated Learning to clinical Natural Language Processing, applied to the task of Violence Risk Assessment by simulating a cross-institutional psychiatric setting. We train and compare four models: two local models, a federated model and a data-centralised model. Our results indicate that the federated model outperforms the local models and has similar performance as the data-centralised model. These findings suggest that Federated Learning can be used successfully in a cross-institutional setting and is a step towards new applications of Federated Learning based on clinical notes. Thomas Borger, Pablo Mosteiro, Heysem Kaya, Emil Rijcken, Albert Ali Salah, Floor Scheepers, Marco Spruit |
Expert Syst. Appl. | 5 |
| 2022 | A Multimodal Approach for Mania Level Prediction in Bipolar DisorderabstractBipolar disorder is a mental health disorder that causes mood swings that range from depression to mania. Clinical diagnosis of bipolar disorder is based on patient interviews and reports obtained from the relatives of the patients. Subsequently, the diagnosis depends on the experience of the expert, and there is co-morbidity with other mental disorders. Automated processing in the diagnosis of bipolar disorder can help providing quantitative indicators, and allow easier observations of the patients for longer periods. In this paper, we create a multimodal decision system for three level mania classification based on recordings of the patients in acoustic, linguistic, and visual modalities. The system is evaluated on the Turkish Bipolar Disorder corpus we have recently introduced to the scientific community. Comprehensive analysis of unimodal and multimodal systems, as well as fusion techniques, are performed. Using acoustic, linguistic, and visual features in a multimodal fusion system, we achieved a 64.8% unweighted average recall score, which advances the state-of-the-art performance on this dataset. Pinar Baki, Heysem Kaya, Elvan Çiftçi, Hüseyin Güleç, Albert Ali Salah |
IEEE Trans. Affect. Comput. | 5 |
| 2022 | Modeling, Recognizing, and Explaining Apparent Personality From VideosabstractExplainability and interpretability are two critical aspects of decision support systems. Despite their importance, it is only recently that researchers are starting to explore these aspects. This paper provides an introduction to explainability and interpretability in the context of apparent personality recognition. To the best of our knowledge, this is the first effort in this direction. We describe a challenge we organized on explainability in first impressions analysis from video. We analyze in detail the newly introduced data set, evaluation protocol, proposed solutions and summarize the results of the challenge. We investigate the issue of bias in detail. Finally, derived from our study, we outline research opportunities that we foresee will be relevant in this area in the near future. Hugo Jair Escalante, Heysem Kaya, Albert Ali Salah, Sergio Escalera, Yagmur Güçlütürk, Umut Güçlü, Xavier Baró, Isabelle Guyon, Júlio C. S. Jacques Júnior, Meysam Madadi, Stéphane Ayache, Evelyne Viegas, Furkan Gürpinar, Achmadnoer Sukma Wicaksana, Cynthia C. S. Liem, Marcel van Gerven, Rob van Lier |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | Can mood primitives predict apparent personality?abstractFirst impressions play a critical role in shaping social interactions and consequently have a high impact on people’s lives. This study presents an explainable system that models apparent personality traits that influence first impressions as a function of automatically predicted arousal, valence and likeability (AVL) scores. To this end, we enrich the ChaLearn Looking at People - First Impressions (LAP-FI) dataset by annotating a portion of it for the AVL dimensions and carry out extensive uni-modal and multimodal experiments by using state-of-the-art acoustic, visual and linguistic features. We propose to use a glass-box model, namely, Explainable Boosting Machine, to model the Big Five personality traits. Our results demonstrate that personality trait impressions can be effectively predicted through the mood and likeability scores of a given video. We show that the proposed model, which is trained on only a few features, not only provides more meaningful explanations but also yields competitive performance (with a 0.09 Mean Absolute Error) compared to the state-of-the-art methods. The annotated benchmark dataset and the scripts to reproduce the results are available at: https://github.com/gizemsogancioglu/mood-project. Gizem Sogancioglu, Heysem Kaya, Albert Ali Salah |
ACII | 3 |
| 2020 | Automatic Pain Detection on Horse and Donkey FacesabstractRecognition of pain in equines (such as horses and donkeys) is essential for their welfare. However, this assessment depends solely on the ability of the observer to locate visible signs of pain since there is no verbal communication. The use of Grimace scales is proven to be efficient in detecting pain but is time-consuming and also dependent on the level of training of the annotators and, therefore, validity is not easily ensured. There is a need for automation of this process to help training. This work provides a system for pain prediction in horses, based on Grimace scales. The pipeline automatically finds landmarks on horse faces before classification. Our experiments show that using different classifiers for different poses of the horse is necessary, and fusion of different features improves results. We furthermore investigate the transfer of horse-based models for donkeys and illustrate the loss of accuracy in automatic landmark detection and subsequent pain prediction. Hilde I. Hummel, Francisca Pessanha, Albert Ali Salah, Thijs J. P. A. M. van Loon, Remco C. Veltkamp |
FG | 3 |
| 2020 | Video2Report: A Video Database for Automatic Reporting of Medical Consultancy SessionsabstractThe regulation of medical consultations for some countries, such as the Netherlands, dictates the general practitioners to prepare a detailed report for each consultation, for accountability purposes. Automatic report generation during medical consultations can simplify this time-consuming procedure. Action recognition for automatic reporting of medical actions is not a well-researched area, and there are no publicly available medical video databases. We present in this paper Video2Report, the first publicly available medical consultancy video database involving interactions between a general practitioner and one patient. After reviewing the standard medical procedures for general practitioners, we select the most important actions to record, and have an actual medical professional perform the actions and train further actors to create a resource. The actions, as well as the area of investigation during the actions are annotated separately. In this paper, we describe the collection setup, provide several action recognition baselines with OpenPose feature extraction, and make the database, evaluation protocol and all annotations publicly available. The database contains 192 sessions recorded with up to three cameras, with 332 single action clips and 119 multiple action sequences. While the dataset size is too small for end-to-end deep learning, we believe it will be useful for developing approaches to investigate doctor-patient interactions and for medical action recognition. Laura Schiphorst, Metehan Doyran, Sabine Molenaar, Albert Ali Salah, Sjaak Brinkkemper |
FG | 4 |
| 2020 | Message from the General and Program Chairs FG 2020abstractWelcome to the 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020). FG is the premier international conference on vision-based automatic face and body behavior analysis and applications. Since its first meeting in Zurich in 1994, the conference has been held fourteen times throughout the world. This is the 15th conference. Juan P. Wachs, Sergio Escalera, Jeffrey F. Cohn, Albert Ali Salah, Arun Ross |
FG | 4 |
| 2020 | Is Everything Fine, Grandma? Acoustic and Linguistic Modeling for Robust Elderly Speech Emotion RecognitionabstractAcoustic and linguistic analysis for elderly emotion recognition is an under-studied and challenging research direction, but essential for the creation of digital assistants for the elderly, as well as unobtrusive telemonitoring of elderly in their residences for mental healthcare purposes. This paper presents our contribution to the INTERSPEECH 2020 Computational Paralinguistics Challenge (ComParE) - Elderly Emotion Sub-Challenge, which is comprised of two ternary classification tasks for arousal and valence recognition. We propose a bi-modal framework, where these tasks are modeled using state-of-the-art acoustic and linguistic features, respectively. In this study, we demonstrate that exploiting task-specific dictionaries and resources can boost the performance of linguistic models, when the amount of labeled data is small. Observing a high mismatch between development and test set performances of various models, we also propose alternative training and decision fusion strategies to better estimate and improve the generalization performance. Gizem Sogancioglu, Oxana Verkholyak, Heysem Kaya, Dmitrii Fedotov, Tobias Cadèe, Albert Ali Salah, Alexey Karpov 0001 |
INTERSPEECH | 6 |
| 2020 | Special Issue on Generating Realistic Visual Data of Human Behavior
Xavier Alameda-Pineda, Elisa Ricci 0001, Albert Ali Salah, Nicu Sebe, Shuicheng Yan |
Int. J. Comput. Vis. | 3 |
| 2019 | Video and Text-Based Affect Analysis of Children in Play TherapyabstractPlay therapy is an approach to psychotherapy where a child is engaging in play activities. Because of the strong affective component of play, it provides a natural setting to analyze feelings and coping strategies of the child. In this paper, we investigate an approach to track the affective state of a child during a play therapy session. We assume a simple, camera-based sensor setup, and describe the challenges of this application scenario. We use fine-tuned off-the-shelf deep convolutional neural networks for the processing of the child’s face during sessions to automatically extract valence and arousal dimensions of affect, as well as basic emotional expressions. We further investigate text-based and body-movement based affect analysis. We evaluate these modalities separately and in conjunction with play therapy videos in natural sessions, discussing the results of such analysis and how it aligns with the professional clinicians’ assessments. Metehan Doyran, Batikan Türkmen, Eda Aydin Oktay, Sibel Halfon, Albert Ali Salah |
ICMI | 5 |
| 2019 | Affordable person detection in omnidirectional cameras using radial integral channel features
Baris Evrim Demiröz, Albert Ali Salah, Yalin Bastanlar, Lale Akarun |
Mach. Vis. Appl. | 2 |
| 2018 | Continuous Real-Time Vehicle Driver Authentication Using Convolutional Neural Network Based Face RecognitionabstractContinuous driver authentication is useful in the prevention of car thefts, fraudulent switching of designated drivers, and driving beyond a designated amount of time for a single driver. In this paper, we propose a deep neural network based approach for real time and continuous authentication of vehicle drivers. Features extracted from pre-trained neural network models are classified with support vector classifiers. In order to examine realistic conditions, we collect 130 in-car driving videos from 52 different subjects. We investigate the conditions under which current face recognition technology will allow commercialization of continuous driver authentication. Ekberjan Derman, Albert Ali Salah |
FG | 2 |
| 2018 | Feature Selection and Multimodal Fusion for Estimating Emotions Evoked by Movie ClipsabstractPerceptual understanding of media content has many applications, including content-based retrieval, marketing, content optimization, psychological assessment, and affect-based learning. In this paper, we model audio visual features extracted from videos via machine learning approaches to estimate the affective responses of the viewers. We use the LIRIS-ACCEDE dataset and the MediaEval 2017 Challenge setting to evaluate the proposed methods. This dataset is composed of movies of professional or amateur origin, annotated with viewers' arousal, valence, and fear scores. We extract a number of audio features, such as Mel-frequency Cepstral Coefficients, and visual features, such as dense SIFT, hue-saturation histogram, and features from a deep neural network trained for object recognition. We contrast two different approaches in the paper, and report experiments with different fusion and smoothing strategies. We demonstrate the benefit of feature selection and multimodal fusion on estimating affective responses to movie segments. Yasemin Timar, Nihan Karslioglu, Heysem Kaya, Albert Ali Salah |
ICMR | 4 |
| 2017 | Emotion, age, and gender classification in children's speech by humans and machines
Heysem Kaya, Albert Ali Salah, Alexey Karpov 0001, Olga V. Frolova, Aleksei Grigorev, Elena E. Lyakso |
Comput. Speech Lang. | 2 |
| 2017 | Video-based emotion recognition in the wild using deep transfer learning and score fusion
Heysem Kaya, Furkan Gürpinar, Albert Ali Salah |
Image Vis. Comput. | 3 |
| 2017 | Automatic Classification of Player Complaints in Social GamesabstractArtificial intelligence and machine learning techniques are not only useful for creating plausible behaviors for interactive game elements, but also for the analysis of the players to provide a better gaming environment. In this paper, we propose a novel framework for automatic classification of player complaints in a social gaming platform. We use features that describe both parties of the complaint (namely, the accuser and the suspect), as well as interaction features of the game itself. The proposed classification approach, based on gradient boosting machines, is tested on the COPA Database of 100 000 unique users and 800 000 individual games. We advance the state of the art in this challenging problem. Koray Balci, Albert Ali Salah |
IEEE Trans. Comput. Intell. AI Games | 2 |
| 2016 | ERM4CT 2016: 2nd international workshop on emotion representations and modelling for companion systems (workshop summary)abstractIn this paper the organisers present a brief overview of the 2nd International Workshop on Emotion Representations and Modelling for Companion Systems (ERM4CT). The ERM4CT 2016 Workshop is held in conjunction with the 18th ACM International Conference on Multimodal Interaction (ICMI 2016) taking place Tokyo, Japan. The ERM4CT is the follow-up of three previous workshops on emotion modelling for affective human-computer interaction and companion systems. Apart from its usual focus on emotion representations and models, this year's ERM4CT puts special emphasis on how to model adequate affective system behaviour. For the first time, this year's ERM4CT gave out a dataset, which all attendees could investigate to jointly discuss their findings. Kim Hartmann, Ingo Siegert, Albert Ali Salah, Khiet P. Truong |
ICMI | 3 |
| 2016 | Multimodal fusion of audio, scene, and face features for first impression estimationabstractAffective computing, particularly emotion and personality trait recognition, is of increasing interest in many research disciplines. The interplay of emotion and personality shows itself in the first impression left on other people. Moreover, the ambient information, e.g. the environment and objects surrounding the subject, also affect these impressions. In this work, we employ pre-trained Deep Convolutional Neural Networks to extract facial emotion and ambient information from images for predicting apparent personality. We also investigate Local Gabor Binary Patterns from Three Orthogonal Planes video descriptor and acoustic features extracted via the popularly used openSMILE tool. We subsequently propose classifying features using a Kernel Extreme Learning Machine and fusing their predictions. The proposed system is applied to the ChaLearn Challenge on First Impression Recognition, achieving the winning test set accuracy of 0.913, averaged over the “Big Five” personality traits. Furkan Gürpinar, Heysem Kaya, Albert Ali Salah |
ICPR | 3 |
| 2016 | Robust Acoustic Emotion Recognition Based on Cascaded Normalization and Extreme Learning Machines
Heysem Kaya, Alexey Karpov 0001, Albert Ali Salah |
ISNN | 3 |
| 2016 | Seventh International Workshop on Human Behavior Understanding (HBU 2016)abstractWith advances in pattern recognition and multimedia computing, it becomes possible to analyze human behavior via multimodal sensors at varying time-scales, levels of analysis, and meaning. This ability opens up far-ranging possibilities for multimedia and multimodal interaction. Research has the, potential to endow computers with the capacity to detect and understand people's actions and activities and infer their attitudes, preferences, personality, and social relationships. This workshop brings together researchers in this rapidly emerging area and especially those concerned with behavior analysis and multimedia in children. Mohamed Chetouani, Jeffrey F. Cohn, Albert Ali Salah |
ACM Multimedia | 3 |
| 2015 | Contrasting and Combining Least Squares Based Learners for Emotion Recognition in the WildabstractThis paper presents our contribution to ACM ICMI 2015 Emotion Recognition in the Wild Challenge (EmotiW 2015). We participate in both static facial expression (SFEW) and audio-visual emotion recognition challenges. In both challenges, we use a set of visual descriptors and their early and late fusion schemes. For AFEW, we also exploit a set of popularly used spatio-temporal modeling alternatives and carry out multi-modal fusion. For classification, we employ two least squares regression based learners that are shown to be fast and accurate on former EmotiW Challenge corpora. Specifically, we use Partial Least Squares Regression (PLS) and Kernel Extreme Learning Machines (ELM), which is closely related to Kernel Regularized Least Squares. We use a General Procrustes Analysis (GPA) based alignment for face registration. By employing different alignments, descriptor types, video modeling strategies and classifiers, we diversify learners to improve the final fusion performance. Test set accuracies reached in both challenges are relatively 25% above the respective baselines. Heysem Kaya, Furkan Gürpinar, Sadaf Afshar, Albert Ali Salah |
ICMI | 4 |
| 2015 | Looking at Mondrian's Victory Boogie-Woogie: What Do I Feel?
Andreza Sartori, Yan Yan 0002, Gözde Özbal, Almila Akdag Salah, Albert Ali Salah, Nicu Sebe |
IJCAI | 5 |
| 2015 | Fisher vectors with cascaded normalization for paralinguistic analysisabstractComputational Paralinguistics has several unresolved issues, one of which is coping with large variability due to speakers, spoken content and corpora. In this paper, we address the variability compensation issue by proposing a novel method composed of i) Fisher vector encoding of low level descrip-tors extracted from the signal, ii) speaker z-normalization ap-plied after speaker clustering iii) non-linear normalization of features and iv) classification based on Kernel Extreme Learn-ing Machines and Partial Least Squares regression. For ex-perimental validation, we apply the proposed method on IN-TERSPEECH 2015 Computational Paralinguistics Challenge (ComParE 2015), Eating Condition sub-challenge, which is a seven-class classification task. In our preliminary experiments, the proposed method achieves an Unweighted Average Recall (UAR) score of 83.1%, outperforming the challenge test set baseline UAR (65.9%) by a large margin. Heysem Kaya, Alexey Karpov 0001, Albert Ali Salah |
INTERSPEECH | 3 |
| 2015 | Efficient large-scale action recognition in videos using extreme learning machines
Gül Varol, Albert Ali Salah |
Expert Syst. Appl. | 2 |
| 2015 | Random Discriminative Projection Based Feature Selection with Application to Conflict RecognitionabstractComputational paralinguistics deals with underlying meaning of the verbal messages, which is of interest in manifold applications ranging from intelligent tutoring systems to affect sensitive robots. The state-of-the-art pipeline of paralinguistic speech analysis utilizes brute-force feature extraction, and the features need to be tailored according to the relevant task. In this work, we extend a recent discriminative projection based feature selection method using the power of stochasticity to overcome local minima and to reduce the computational complexity. The proposed approach assigns weights both to groups and to features individually in many randomly selected contexts and then combines them for a final ranking. The efficacy of the proposed method is shown in a recent paralinguistic challenge corpus to detect level of conflict in dyadic and group conversations. We advance the state-of-the-art in this corpus using the INTERSPEECH 2013 Challenge protocol. Heysem Kaya, Tugçe Özkaptan, Albert Ali Salah, Fikret S. Gürgen |
IEEE Signal Process. Lett. | 3 |
| 2015 | Brief Introduction to the Special Issue on Behavior Understanding for Arts and EntertainmentabstractThis editorial introduction describes the aims and scope of the special issue of the ACM Transactions on Interactive Intelligent Systems on Behavior Understanding for Arts and Entertainment, which is being published in issues 2 and 3 of volume 5 of the journal. Here we offer a brief introduction to the use of behavior analysis for interactive systems that involve creativity in either the creator or the consumer of a work of art. We then characterize each of the five articles included in this first part of the special issue, which span a wide range of applications. Albert Ali Salah, Hayley Hung, Oya Aran, Hatice Gunes, Matthew Turk 0001 |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2015 | Behavior Understanding for Arts and EntertainmentabstractThis editorial introduction complements the shorter introduction to the first part of the two-part special issue on Behavior Understanding for Arts and Entertainment. It offers a more expansive discussion of the use of behavior analysis for interactive systems that involve creativity, either for the producer or the consumer of such a system. We first summarise the two articles that appear in this second part of the special issue. We then discuss general questions and challenges in this domain that were suggested by the entire set of seven articles of the special issue and by the comments of the reviewers of these articles. Albert Ali Salah, Hayley Hung, Oya Aran, Hatice Gunes, Matthew Turk 0001 |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2015 | Combining Facial Dynamics With Appearance for Age EstimationabstractEstimating the age of a human from the captured images of his/her face is a challenging problem. In general, the existing approaches to this problem use appearance features only. In this paper, we show that in addition to appearance information, facial dynamics can be leveraged in age estimation. We propose a method to extract and use dynamic features for age estimation, using a person's smile. Our approach is tested on a large, gender-balanced database with 400 subjects, with an age range between 8 and 76. In addition, we introduce a new database on posed disgust expressions with 324 subjects in the same age range, and evaluate the reliability of the proposed approach when used with another expression. State-of-the-art appearance-based age estimation methods from the literature are implemented as baseline. We demonstrate that for each of these methods, the addition of the proposed dynamic features results in statistically significant improvement. We further propose a novel hierarchical age estimation architecture based on adaptive age grouping. We test our approach extensively, including an exploration of spontaneous versus posed smile dynamics, and gender-specific age estimation. We show that using spontaneity information reduces the mean absolute error by up to 21%, advancing the state of the art for facial age estimation. Hamdi Dibeklioglu, Fares Alnajar, Albert Ali Salah, Theo Gevers |
IEEE Trans. Image Process. | 3 |
| 2015 | Recognition of Genuine SmilesabstractAutomatic distinction between genuine (spontaneous) and posed expressions is important for visual analysis of social signals. In this paper, we describe an informative set of features for the analysis of face dynamics, and propose a completely automatic system to distinguish between genuine and posed enjoyment smiles. Our system incorporates facial landmarking and tracking, through which features are extracted to describe the dynamics of eyelid, cheek, and lip corner movements. By fusing features over different regions, as well as over different temporal phases of a smile, we obtain a very accurate smile classifier. We systematically investigate age and gender effects, and establish that age-specific classification significantly improves the results, even when the age is automatically estimated. We evaluate our system on the 400-subject UvA-NEMO database we have recently collected, as well as on three other smile databases from the literature . Through an extensive experimental evaluation, we show that our system improves the state of the art in smile classification and provides useful insights in smile psychophysics. Hamdi Dibeklioglu, Albert Ali Salah, Theo Gevers |
IEEE Trans. Multim. | 2 |
| 2014 | CCA based feature selection with application to continuous depression recognition from acoustic speech featuresabstractIn this study we make use of Canonical Correlation Analysis (CCA) based feature selection for continuous depression recognition from speech. Besides its common use in multi-modal/multi-view feature extraction, CCA can be easily employed as a feature selector. We introduce several novel ways of CCA based filter (ranking) methods, showing their relations to previous work. We test the suitability of proposed methods on the AVEC 2013 dataset under the ACM MM 2013 Challenge protocol. Using 17% of features, we obtained a relative improvement of 30% on the challenge's test-set baseline Root Mean Square Error. Heysem Kaya, Florian Eyben, Albert Ali Salah, Björn W. Schuller |
ICASSP | 3 |
| 2014 | Combining Modality-Specific Extreme Learning Machines for Emotion Recognition in the WildabstractThis paper presents our contribution to ACM ICMI 2014 Emotion Recognition in the Wild Challenge and Workshop. The proposed system utilizes Extreme Learning Machines (ELM) for modeling modality-specific features and combines the scores for final prediction. The state-of-the-art results in acoustic and visual emotion recognition are obtained either using deep Neural Networks (DNN) or Support Vector Machines (SVM). The ELM paradigm is proposed as a fast and accurate alternative to these two popular machine learning methods. Benefiting from fast learning advantage of ELM, we carry out extensive tests on the data using moderate computational resources. In the video modality, we test combination of regional visual features obtained from the inner face. In the audio modality, we carry out tests to enhance training via other emotional corpora. We further investigate the suitability of several recently proposed feature selection approaches to prune the acoustic features. In our study, the best results for both modalities are obtained with Kernel ELM compared to basic ELM. On the challenge test set, we obtain 37.84%, 39.07% and 44.23% classification accuracies for audio, video and multimodal fusion, respectively. Heysem Kaya, Albert Ali Salah |
ICMI | 2 |
| 2014 | Canonical correlation analysis and local fisher discriminant analysis based multi-view acoustic feature reduction for physical load predictionabstractIn this study we present our system for INTERSPEECH 2014 Computational Paralinguistics Challenge (ComParE 2014), Physical Load Sub-challenge (PLS). Our contribution is twofold. First, we propose using Low Level Descriptor (LLD) information as hints, so as to partition the feature space into meaningful subsets called views. We also show the virtue of commonly employed feature projections, such as Canoni-cal Correlation Analysis (CCA) and Local Fisher Discriminant Analysis (LFDA) as ranking feature selectors. Results indicate the superiority of multi-view feature reduction approach to its single-view counterpart. Moreover, the discriminative projec-tion matrices are observed to provide valuable information for feature selection, which generalize better than the projection it-self. In our preliminary experiments we reached 75.35 % Un-weighted Average Recall (UAR) on PLS test set, using CCA based multi-view feature selection. Heysem Kaya, Tugçe Özkaptan, Albert Ali Salah, Fikret S. Gürgen |
INTERSPEECH | 3 |
| 2014 | Eyes Whisper Depression: A CCA based Multimodal ApproachabstractThis paper presents our work on ACM MM Audio Visual Emotion Corpus 2013 (AVEC 2013) depression recognition sub-challenge using the baseline features in accordance with the challenge protocol. We use Canonical Correlation Analysis for audio-visual fusion as well as covariate extraction for the target task. The video baseline provides histograms of local phase quantization features extracted from 4x4=16 regions of the detected face. We summarize the video features over segments of length 20 seconds using mode and range functionals. We observe that features of range functional that measure the variance tendency provides statistically significantly higher canonical correlation than mode functional features that measure the mean tendency. Moreover, when audio-visual features are used with varying number of covariates per region, the regions that were consistently found the best are the ones corresponding to two eyes and the right part of the mouth. Heysem Kaya, Albert Ali Salah |
ACM Multimedia | 2 |
| 2014 | Auto-evaluation of motion imitation in a child-robot imitation game for upper arm rehabilitationabstractThe purpose of this study is fusing play-like child robot interaction with physiotherapy in order to achieve upper arm rehabilitation by motivating the child. The proposed system is not intended to substitute for the physiotherapist, but to assist them in their therapeutic tasks by encouraging the child's participation in the activity. Recognizing the imitation performance of the child and supporting him/her with feedback for drawing the child's attention and motivating the child to imitate the robot is crucial. This study concentrates on automatically evaluating the upper body actions of the child during an imitation based physical therapy. For quantifying the performance of the child, two measures were considered: Range of Motion (RoM) and Dynamic Time Warping (DTW) distance. In our initial experiments, eight healthy children were asked to stand in front of a Kinect sensor and to mimic the actions of the humanoid robot Nao, which consist of shoulder abduction, shoulder vertical flexion&extension and elbow flexion. The proposed evaluation measure is verified as a reliable measurement according to Intraclass Correlation Coefficient (ICC) through comparison with evaluations of five physiotherapists as ground truth. The degree of consistency among our ratings and the physiotherapist ratings is between %76 and %96 for different motions. Arzu Güneysu, Recep Doga Siyli, Albert Ali Salah |
RO-MAN | 3 |
| 2013 | Like Father, Like Son: Facial Expression Dynamics for Kinship VerificationabstractKinship verification from facial appearance is a difficult problem. This paper explores the possibility of employing facial expression dynamics in this problem. By using features that describe facial dynamics and spatio-temporal appearance over smile expressions, we show that it is possible to improve the state of the art in this problem, and verify that it is indeed possible to recognize kinship by resemblance of facial expressions. The proposed method is tested on different kin relationships. On the average, 72.89% verification accuracy is achieved on spontaneous smiles. Hamdi Dibeklioglu, Albert Ali Salah, Theo Gevers |
ICCV | 2 |
| 2013 | Fourth international workshop on human behavior understanding (HBU 2013)abstractWith advances in pattern recognition and multimedia computing, it became possible to analyze human behavior via multimodal sensors, at different time-scales and at different levels of interaction and interpretation. This ability opens up enormous possibilities for multimedia and multimodal interaction, with a potential of endowing the computers with a capacity to attribute meaning to users' attitudes, preferences, personality, social relationships, etc., as well as to understand what people are doing, the activities they have been engaged in, their routines and lifestyles. This workshop gathers researchers dealing with the problem of modeling human behavior under its multiple facets with particular attention to interactions in arts, creativity, entertainment and edutainment. Albert Ali Salah, Hayley Hung, Oya Aran, Hatice Gunes |
ACM Multimedia | 1 |
| 2013 | Joint Attention by Gaze Interpolation and SaliencyabstractJoint attention, which is the ability of coordination of a common point of reference with the communicating party, emerges as a key factor in various interaction scenarios. This paper presents an image-based method for establishing joint attention between an experimenter and a robot. The precise analysis of the experimenter's eye region requires stability and high-resolution image acquisition, which is not always available. We investigate regression-based interpolation of the gaze direction from the head pose of the experimenter, which is easier to track. Gaussian process regression and neural networks are contrasted to interpolate the gaze direction. Then, we combine gaze interpolation with image-based saliency to improve the target point estimates and test three different saliency schemes. We demonstrate the proposed method on a human-robot interaction scenario. Cross-subject evaluations, as well as experiments under adverse conditions (such as dimmed or artificial illumination or motion blur), show that our method generalizes well and achieves rapid gaze estimation for establishing joint attention. Zeynep Yücel, Albert Ali Salah, Çetin Meriçli, Tekin Meriçli, Roberto Valenti, Theo Gevers |
IEEE Trans. Cybern. | 2 |
| 2012 | Are You Really Smiling at Me? Spontaneous versus Posed Enjoyment Smiles
Hamdi Dibeklioglu, Albert Ali Salah, Theo Gevers |
ECCV (3) | 2 |
| 2012 | A smile can reveal your age: enabling facial dynamics in age estimationabstractEstimation of a person's age from the facial image has many applications, ranging from biometrics and access control to cosmetics and entertainment. Many image-based methods have been proposed for this problem. In this paper, we propose a method for the use of dynamic features in age estimation, and show that 1) the temporal dynamics of facial features can be used to improve image-based age estimation; 2) considered alone, static image-based features are more accurate than dynamic features. We have collected and annotated an extensive database of face videos from 400 subjects with an age range between 8 and 76, which allows us to extensively analyze the relevant aspects of the problem. The proposed system, which fuses facial appearance and expression dynamics, performs with a mean absolute error of 4.81 (4.87) years. This represents a significant improvement of accuracy in comparison to the sole use of appearance-based features. Hamdi Dibeklioglu, Theo Gevers, Albert Ali Salah, Roberto Valenti |
ACM Multimedia | 3 |
| 2012 | A Statistical Method for 2-D Facial LandmarkingabstractMany facial-analysis approaches rely on robust and accurate automatic facial landmarking to correctly function. In this paper, we describe a statistical method for automatic facial-landmark localization. Our landmarking relies on a parsimonious mixture model of Gabor wavelet features, computed in coarse-to-fine fashion and complemented with a shape prior. We assess the accuracy and the robustness of the proposed approach in extensive cross-database conditions conducted on four face data sets (Face Recognition Grand Challenge, Cohn-Kanade, Bosphorus, and BioID). Our method has 99.33% accuracy on the Bosphorus database and 97.62% accuracy on the BioID database on the average, which improves the state of the art. We show that the method is not significantly affected by low-resolution images, small rotations, facial expressions, and natural occlusions such as beard and mustache. We further test the goodness of the landmarks in a facial expression recognition application and report landmarking-induced improvement over baseline on two separate databases for video-based expression recognition (Cohn-Kanade and BU-4DFE). Hamdi Dibeklioglu, Albert Ali Salah, Theo Gevers |
IEEE Trans. Image Process. | 2 |
| 2011 | Recent developments in social signal processingabstractSocial signal processing has the ambitious goal of bridging the social intelligence gap between computers and humans. Nowadays, computers are not only the new interaction partners of humans, but also a privileged interaction medium for social exchange between humans. Consequently, enhancing machine abilities to interpret and reproduce social signals is a crucial requirement for improving computer-mediated communication and interaction. Furthermore, automated analysis of such signals creates a host of new applications and improvements to existing applications. The study of social signals benefits a wide range of domains, including human-computer interaction, interaction design, entertainment technology, ambient intelligence, health-care, and psychology. This paper briefly introduces the field and surveys its latest developments. Albert Ali Salah, Maja Pantic, Alessandro Vinciarelli |
SMC | 1 |
| 2010 | Eyes do not lie: spontaneous versus posed smilesabstractAutomatic detection of spontaneous versus posed facial expressions received a lot of attention in recent years. However, almost all published work in this area use complex facial features or multiple modalities, such as head pose and body movements with facial features. Besides, the results of these studies are not given on public databases. In this paper, we focus on eyelid movements to classify spontaneous versus posed smiles and propose distance-based and angular features for eyelid movements. We assess the reliability of these features with continuous HMM, k-NN and naive Bayes classifiers on two different public datasets. Experimentation shows that our system provides classification rates up to 91 per cent for posed smiles and up to 80 per cent for spontaneous smiles by using only eyelid movements. We additionally compare the discrimination power of movement features from different facial regions for the same task. Hamdi Dibeklioglu, Roberto Valenti, Albert Ali Salah, Theo Gevers |
ACM Multimedia | 3 |
| 2010 | An Evaluation of Video-to-Video Face VerificationabstractPerson recognition using facial features, e.g., mug-shot images, has long been used in identity documents. However, due to the widespread use of web-cams and mobile devices embedded with a camera, it is now possible to realize facial video recognition, rather than resorting to just still images. In fact, facial video recognition offers many advantages over still image recognition; these include the potential of boosting the system accuracy and deterring spoof attacks. This paper presents an evaluation of person identity verification using facial video data, organized in conjunction with the International Conference on Biometrics (ICB 2009). It involves 18 systems submitted by seven academic institutes. These systems provide for a diverse set of assumptions, including feature representation and preprocessing variations, allowing us to assess the effect of adverse conditions, usage of quality information, query selection, and template construction for video-to-video face authentication. Norman Poh, Chi-Ho Chan, Josef Kittler, Sébastien Marcel, Chris McCool, Enrique Argones-Rúa, José Luis Alba-Castro, Mauricio Villegas, Roberto Paredes, Vitomir Struc, Nikola Pavesic, Albert Ali Salah, Hui Fang 0003, Nicholas Costen |
IEEE Trans. Inf. Forensics Secur. | 12 |
| 2009 | Benchmarking quality-dependent and cost-sensitive score-level multimodal biometric fusion algorithmsabstractAutomatically verifying the identity of a person by means of biometrics (e.g., face and fingerprint) is an important application in our day-to-day activities such as accessing banking services and security control in airports. To increase the system reliability, several biometric devices are often used. Such a combined system is known as a multimodal biometric system. This paper reports a benchmarking study carried out within the framework of the BioSecure DS2 (Access Control) evaluation campaign organized by the University of Surrey, involving face, fingerprint, and iris biometrics for person authentication, targeting the application of physical access control in a medium-size establishment with some 500 persons. While multimodal biometrics is a well-investigated subject in the literature, there exists no benchmark for a fusion algorithm comparison. Working towards this goal, we designed two sets of experiments: quality-dependent and cost-sensitive evaluation. The quality-dependent evaluation aims at assessing how well fusion algorithms can perform under changing quality of raw biometric images principally due to change of devices. The cost-sensitive evaluation, on the other hand, investigates how well a fusion algorithm can perform given restricted computation and in the presence of software and hardware failures, resulting in errors such as failure-to-acquire and failure-to-match. Since multiple capturing devices are available, a fusion algorithm should be able to handle this nonideal but nevertheless realistic scenario. In both evaluations, each fusion algorithm is provided with scores from each biometric comparison subsystem as well as the quality measures of both the template and the query data. The response to the call of the evaluation campaign proved very encouraging, with the submission of 22 fusion systems. To the best of our knowledge, this campaign is the first attempt to benchmark quality-based multimodal fusion algorithms. In the presence of changing image quality which may be due to a change of acquisition devices and/or device capturing configurations, we observe that the top performing fusion algorithms are those that exploit automatically derived quality measurements. Our evaluation also suggests that while using all the available biometric sensors can definitely increase the fusion performance, this comes at the expense of increased cost in terms of acquisition time, computation time, the physical cost of hardware, and its maintenance cost. As demonstrated in our experiments, a promising solution which minimizes the composite cost is sequential fusion, where a fusion algorithm sequentially uses match scores until a desired confidence is reached, or until all the match scores are exhausted, before outputting the final combined score. Norman Poh, Thirimachos Bourlai, Josef Kittler, Lorène Allano, Fernando Alonso-Fernandez, Onkar Ambekar, John P. Baker, Bernadette Dorizzi, Omolara Fatukasi, Julian Fierrez, Harald Ganster, Javier Ortega-Garcia, Donald E. Maurer, Albert Ali Salah, Tobias Scheidat, Claus Vielhauer |
IEEE Trans. Inf. Forensics Secur. | 14 |
| 2008 | Empowering the end-user in biometricsabstractUser empowerment in the context of information and communication technologies (ICT) means enabling the end-users to set up and/or tailor ICT solutions according to their own requirements. In this paper we argue that user empowerment is essential in biometrics for the acceptance and widespread use of biometrical applications, concepts, and technology. A key issue is a user's experience of being in control, and how interfaces, modes and modalities of interaction can support and empower the user, whether by direct interaction and control of devices, or via delegation. More importantly, the practices suggested in the paper are steps towards practical and legal regularization of use and circulation o.f biometric information. Ben A. M. Schouten, Albert Ali Salah |
ICARCV | 2 |
| 2007 | Sensor Networks for Ambient IntelligenceabstractDue to rapid advances in networking and sensing technology we are witnessing a growing interest in sensor networks, in which a variety of sensors are connected to each other and to computational devices capable of multimodal signal processing and data analysis. Such networks are seen to play an increasingly important role as key enablers in emerging pervasive computing technologies. In the first part of this paper we give an overview of recent developments in the area of multimodal sensor networks, paying special attention to ambient intelligence applications. In the second part, we discuss how the time series generated by data streams emanating from the sensors can be mined for temporal patterns, indicating cross-sensor signal correlations. Eric J. Pauwels, Albert Ali Salah, Romain Tavenard |
MMSP | 2 |
| 2002 | A Selective Attention-Based Method for Visual Pattern Recognition with Application to Handwritten Digit Recognition and Face RecognitionabstractParallel pattern recognition requires great computational resources; it is NP-complete. From an engineering point of view it is desirable to achieve good performance with limited resources. For this purpose, we develop a serial model for visual pattern recognition based on the primate selective attention mechanism. The idea in selective attention is that not all parts of an image give us information. If we can attend only to the relevant parts, we can recognize the image more quickly and using less resources. We simulate the primitive, bottom-up attentive level of the human visual system with a saliency scheme and the more complex, top-down, temporally sequential associative level with observable Markov models. In between, there is a neural network that analyses image parts and generates posterior probabilities as observations to the Markov model. We test our model first on a handwritten numeral recognition problem and then apply it to a more complex face recognition problem. Our results indicate the promise of this approach in complicated vision applications. Albert Ali Salah, Ethem Alpaydin, Lale Akarun |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |