Ali N. Salman

dblp:277/6464 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
15since 2021 · last 2025
0000-0003-1395-0382ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Multi-task Pedestrian Attribute Classification Using ConvNeXt with Advanced Data Augmentation
Magzhan Kairanbay, Ali N. Salman
CAIP (1)2
2025 Noise-Robust Speech Emotion Recognition Using Shared Self-Supervised Representations with Integrated Speech Enhancement
abstract
Recent studies have demonstrated the effectiveness of fine-tuning self-supervised speech representation models for speech emotion recognition (SER). However, applying SER in real-world environments remains challenging due to pervasive noise. Relying on low-accuracy predictions due to noisy speech can undermine the user’s trust. This paper proposes a unified self-supervised speech representation framework for enhanced speech emotion recognition designed to increase noise robustness in SER while generating enhanced speech. Our framework integrates speech enhancement (SE) and SER tasks, leveraging shared self-supervised learning (SSL)-derived features to improve emotion classification performance in noisy environments. This strategy encourages the SE module to enhance discriminative information for SER tasks. Additionally, we introduce a cascade unfrozen training strategy, where the SSL model is gradually unfrozen and fine-tuned alongside the SE and SER heads, ensuring training stability and preserving the generalizability of SSL representations. This approach demonstrates improvements in SER performance under unseen noisy conditions without compromising SE quality. When tested at a 0 dB signal-to-noise ratio (SNR) level, our proposed method outperforms the original baseline by 3.7% in F1-Macro and 2.7% in F1-Micro scores, where the differences are statistically significant.
Jing-Tong Tzeng, Seong-Gyun Leem, Ali N. Salman, Chi-Chun Lee, Carlos Busso
ICASSP3
2025 Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition
abstract
Cross-corpus speech emotion recognition (SER) plays a vital role in numerous practical applications. Traditional approaches to cross-corpus emotion transfer often concentrate on adapting acoustic features to align with different corpora, domains, or labels. However, acoustic features are inherently variable and error-prone due to factors like speaker differences, domain shifts, and recording conditions. To address these challenges, this study adopts a novel contrastive approach by focusing on emotion-specific articulatory gestures as the core elements for analysis. By shifting the emphasis on the more stable and consistent articulatory gestures, we aim to enhance emotion transfer learning in SER tasks. Our research leverages the CREMA-D and MSP-IMPROV corpora as benchmarks and it reveals valuable insights into the commonality and reliability of these articulatory gestures. The findings highlight mouth articulatory gesture potential as a better constraint for improving emotion recognition across different settings or domains.
Shreya G. Upadhyay, Ali N. Salman, Carlos Busso, Chi-Chun Lee
ICASSP2
2025 The Interspeech 2025 Challenge on Speech Emotion Recognition in Naturalistic Conditions
Abinay Reddy Naini, Lucas Goncalves, Ali N. Salman, Pravin Mote, Ismail Rasim Ülgen, Thomas Thebaud, Laureano Moro-Velázquez, L. Paola García-Perera, Najim Dehak, Berrak Sisman, Carlos Busso
INTERSPEECH3
2025 Advancing Pediatric ASR: The Role of Voice Generation in Disordered Speech
Karen Rosero, Ali N. Salman, Shreeram Suresh Chandra, Berrak Sisman, Cortney Van't Slot, Alex A. Kane, Rami R. Hallac, Carlos Busso
INTERSPEECH2
2025 Minority Views Matter: Evaluating Speech Emotion Classifiers With Human Subjective Annotations by an All-Inclusive Aggregation Rule
abstract
When selecting test data for subjective tasks, most studies define ground truth labels using aggregation methods such as the majority or plurality rules. These methods discard data points without consensus, making the test set easier than practical tasks where a prediction is needed for each sample. However, the discarded data points often express ambiguous cues that elicit coexisting traits perceived by annotators. This paper addresses the importance of considering all the annotations and samples in the data, highlighting that only showing the model's performance on an incomplete test set selected by using the majority or plurality rules can lead to bias in the models’ performances. We focus onspeech-emotion recognition(SER) tasks. We observe that traditional aggregation rules have a data loss ratio ranging from 5.63% to 89.17%. From this observation, we propose a flexible method named the all-inclusive aggregation rule to evaluate SER systems on the complete test data. We contrast traditional single-label formulations with a multi-label formulation to consider the coexistence of emotions. We show that training an SER model with the data selected by the all-inclusive aggregation rule shows consistently higher macro-F1 scores when tested in the entire test set, including ambiguous samples without agreement.
Huang-Cheng Chou, Lucas Goncalves, Seong-Gyun Leem, Ali N. Salman, Chi-Chun Lee, Carlos Busso
IEEE Trans. Affect. Comput.4
2024 Enhanced Facial Landmarks Detection for Patients with Repaired Cleft Lip and Palate
abstract
Cleft lip and palate (CLP) is a congenital condition causing deformities in the oral and labial tissues. Post-surgery, patients often experience residual issues like facial asymmetry, and speech disorders. Tracking points in the orofacial area using a facial landmark detector (FLD) contributes to the assessment of speech development and movement impairments. However, off-the-shelf FLDs fail at delineating the lips of patients with repaired CLP. To address this need, our study introduces the CLP-Trans strategy, a domain transfer solution that employs tailor-made affine transformations to modify facial images sourced from publicly available datasets, which constitute our source domain, whereas images of patients with repaired CLP form our target domain. We aim to reduce distribution disparities between the source and target domains for FLD by simulating common outcomes of CLP repair surgery. The system utilizes a deep convolutional neural network (CNN) to learn from transformed images, therefore, preserving the privacy and facilitating the reproducibility of the findings. The strategy achieves statistically significant improvements in the normalized mean square error (NMSE), reducing it from 2.417 to 2.086 (i.e., 13.7% error reduction) by using the proposed strategy when evaluating images of patients with CLP.
Karen Rosero, Ali N. Salman, Berrak Sisman, Rami R. Hallac, Carlos Busso
FG2
2024 Lip Abnormality Detection for Patients with Repaired Cleft Lip and Palate: A Lip Normalization Approach
abstract
The cleft lip condition arises from the incomplete fusion of oral and labial structures during fetal development, impacting vital functions. After surgical closure, patients commonly present with abnormal lip shape, which may require secondary revision surgery for both aesthetic and functional improvement. However, a lack of standardized evaluation methods complicates decision-making for secondary surgery. To address this limitation, we propose a transformer-based lip normalization approach that filters out abnormalities and achieves a standardized appearance while preserving individual anatomy. An innovation of our approach is a lip transformation method using available face datasets to mimic repaired cleft lip shapes, enabling the training of deep learning models without using patients’ data. We employ a Siamese convolutional neural network that processes pre- and post-normalization images to detect lip abnormalities with an accuracy of 89%. We compare our approach with a single-branch model without lip normalization, which reached an accuracy of 60%. Our approach has the potential to provide an impartial view to determine the need for revision surgery while also assisting in the selection of healthcare tools specialized for patients with repaired cleft lip. The code for this work is available in our anonymous repository 1.
Karen Rosero, Ali N. Salman, Rami R. Hallac, Carlos Busso
ICMI2
2024 MSP-GEO Corpus: A Multimodal Database for Understanding Video-Learning Experience
abstract
Video-based learning has become a popular, scalable, and effective approach for students to learn new skills. Many of the challenges for video-based learning can be addressed with machine learning models. However, the available datasets often lack the rich source of data that is needed to accurately predict students’ learning experiences and outcomes. To address this limitation, we introduce the MSP-GEO corpus, a new multimodal database that contains detailed demographic and educational data, recordings of the students and their screens, and meta-data about the lecture during the learning experience. The MSP-GEO corpus was collected using a quasi-experimental pre-test/post-test design. It consists of more than 39,600 seconds (11 hours) of continuous facial footage from 76 participants watching one of three experimental videos on the topic of fossil formation, resulting in over one million facial images. The data collected includes 21 gaze synchronization points, webcam and monitor recordings, and metadata for pauses, plays, and timeline navigation. Additionally, we annotated the recordings for engagement, boredom, and confusion using human evaluators. The MSP-GEO corpus has the potential to improve the accuracy of video-based learning outcomes and experience predictions, facilitate research on the psychological processes of video-based learning, inform the design of instructional videos, and advance the development of learning analytics methods.
Ali N. Salman, Ning Wang 0104, Luz Martinez-Lucas, Andrea Vidal, Carlos Busso
ICMI1
2024 Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline
Ali N. Salman, Zongyang Du, Shreeram Suresh Chandra, Ismail Rasim Ülgen, Carlos Busso, Berrak Sisman
INTERSPEECH1
2024 Embracing Ambiguity And Subjectivity Using The All-Inclusive Aggregation Rule For Evaluating Multi-Label Speech Emotion Recognition Systems
abstract
Speech Emotion Recognition (SER) faces a distinct challenge compared to other speech-related tasks because the annotations will show the subjective emotional perceptions of different annotators. Previous SER studies often view the subjectivity of emotion perception as noise by using the majority rule or plurality rule to obtain the consensus labels. However, these standard approaches overlook the valuable information of labels that do not agree with the consensus and make it easier for the test set. Emotion perception can have co-occurring emotions in realistic conditions, and it is unnecessary to regard the disagreement between raters as noise. To bridge the SER into a multi-label task, we introduced an “all-inclusive rule,” which considers all available data, ratings, and distributional labels as multi-label targets and a complete test set. We demonstrated that models trained with multi-label targets generated by the proposed AR outperform conventional single-label methods across incomplete and complete test sets.
Huang-Cheng Chou, Lucas Goncalves, Seong-Gyun Leem, Ali N. Salman, Carlos Busso, Hung-yi Lee, Chi-Chun Lee
SLT5
2023 Analyzing the Effect of Affective Priming on Emotional Annotations
abstract
In the field of affective computing, emotional annotations are highly important for both the recognition and synthesis of human emotions. Researchers must ensure that these emotional labels are adequate for modeling general human perception. An unavoidable part of obtaining such labels is that human annotators are exposed to known and unknown stimuli before and during the annotation process that can affect their perception. Emotional stimuli cause an affective priming effect, which is a pre-conscious phenomenon in which previous emotional stimuli affect the emotional perception of a current target stimulus. In this paper, we use sequences of emotional annotations during a perceptual evaluation to study the effect of affective priming on emotional ratings of speech. We observe that previous emotional sentences with extreme emotional content push annotations of current samples to the same extreme. We create a sentence-level bias metric to study the effect of affective priming on speech emotion recognition(SER) modeling. The metric is used to identify subsets in the database with more affective priming bias intentionally creating biased datasets. We train and test SER models using the full and biased datasets. Our results show that although the biased datasets have low inter-evaluator agreements, SER models for arousal and dominance trained with those datasets perform the best. For valence, the models trained with the less-biased datasets perform the best.
Luz Martinez-Lucas, Ali N. Salman, Seong-Gyun Leem, Shreya G. Upadhyay, Chi-Chun Lee, Carlos Busso
ACII2
2023 An Intelligent Infrastructure Toward Large Scale Naturalistic Affective Speech Corpora Collection
abstract
The field of speech emotion recognition (SER) aims to create scientifically rigorous systems that can reliably characterize emotional behaviors expressed in speech. A key aspect for building SER systems is to obtain emotional data that is both reliable and reproducible for practitioners. However, academic researchers encounter difficulties in accessing or collecting naturalistic large-scale, reliable emotional recordings. Also, the best practices for data collection are not necessarily described or shared when presenting emotional corpora. To address this issue, the paper proposes the creation of an affective naturalistic database consortium (AndC) that can encourage multidisciplinary cooperation among researchers and practitioners in the field of affective computing. This paper’s contribution is twofold. First, it proposes the design of the AndC with a customizable-standard framework for intelligently-controlled emotional data collection. The focus is on leveraging naturalistic spontaneous recordings available on audio-sharing websites. Second, it presents as a case study the development of a naturalistic large-scale Taiwanese Mandarin podcast corpus using the customizable-standard intelligently-controlled framework. The AndC will enable research groups to effectively collect data using the provided pipeline and to contribute with alternative algorithms or data collection protocols.
Shreya G. Upadhyay, Woan-Shiuan Chien, Bo-Hao Su, Lucas Goncalves, Ya-Tse Wu, Ali N. Salman, Carlos Busso, Chi-Chun Lee
ACII6
2023 Preference Learning Labels by Anchoring on Consecutive Annotations
Abinay Reddy Naini, Ali N. Salman, Carlos Busso
INTERSPEECH2
2022 Privacy Preserving Personalization for Video Facial Expression Recognition Using Federated Learning
abstract
The increased ubiquitousness of small smart devices, such as cellphones, tablets, smart watches and laptops, has led to unique user data, which can be locally processed. The sensors (e.g., microphones and webcam) and improved hardware of the new devices have allowed running deep learning models that 20 years ago would have been exclusive to high-end expensive machines. In spite of this progress, state-of-the-art algorithms for facial expression recognition (FER) rely on architectures that cannot be implemented on these devices due to computational and memory constraints. Alternatives involving cloud-based solutions impose privacy barriers that prevent their adoption or user acceptance in wide range of applications. This paper proposes a lightweight model that can run in real-time for image facial expression recognition (IFER) and video facial expression recognition (VFER). The approach relies on a personalization mechanism locally implemented for each subject by fine-tuning a central VFER model with unlabeled videos from a target subject. We train the IFER model to generate pseudo labels and we select the videos with the highest confident predictions to be used for adaptation. The adaptation is performed by implementing a federated learning strategy where the weights of the local model are averaged and used by the central VFER model. We demonstrate that this approach can improve not only the performance on the edge device providing personalized models to the users, but also the central VFER model. We implement a federated learning strategy where the weights of the local models are averaged and used by the central VFER. Within corpus and cross-corpus evaluations on two emotional databases demonstrate that edge models adapted with our personalization strategy achieve up to 13.1% gains in F1-scores. Furthermore, the federated learning implementation improves the mean micro F1-score of the central VFER model by up to 3.4%. The proposed lightweight solution is ideal for interactive user interfaces that preserve the data of the users.
Ali N. Salman, Carlos Busso
ICMI1
2020 Dynamic versus Static Facial Expressions in the Presence of Speech
abstract
Face analysis is an important area in affective computing. While studies have reported important progress in detecting emotions from still images, an open challenge is to determine emotions from videos, leveraging the dynamic nature in the externalization of emotions. A common approach in earlier studies is to individually process each frame of a video, aggregating the results obtained across frames. This study questions this approach, especially when the subjects are speaking. Speech articulation affects the face appearance, which may lead to misleading emotional perceptions when the isolated frames are taken out-of-context. The analysis in this study explores the similarities and differences in emotion perceptions between (1) videos of speaking segments (without audio), and (2) isolated frames from the same videos evaluated out-of-context. We consider the emotions happiness, sadness, anger and neutral state, and emotional attributes valence, arousal, and dominance using the MSP-IMPROV corpus. The results consistently reveal that the emotional perception of static representations of emotion in isolated frames is significantly different from the overall emotional perception of dynamic representation in videos in the presence of speech. The results reveal the intrinsic limitations of the common frame-by-frame analysis of videos, highlighting the importance of explicitly modeling temporal and lexical information in face emotion recognition from videos.
Ali N. Salman, Carlos Busso
FG1
2020 Style Extractor For Facial Expression Recognition in the Presence of Speech
abstract
The performance of facial expression recognition (FER) systems has improved with recent advances in machine learning. While studies have reported impressive accuracies in detecting emotion from posed expressions in static images, there are still important challenges in developing FER systems for videos, especially in the presence of speech. Speech articulation modulates the orofacial area, changing the facial appearance. These facial movements induced by speech introduce noise, reducing the performance of an FER system. Solving this problem is important if we aim to study more naturalistic environment or applications in the wild. We propose a novel approach to compensate for lexical information that does not require phonetic information during inference. The approach relies on a style extractor model, which creates emotional-to-neutral transformations. The transformed facial representations are spatially contrasted with the original faces, highlighting the emotional information conveyed in the video. The results demonstrate that adding the proposed style extractor model to a dynamic FER system improves the performance by 7% (absolute) compared to a similar model with no style extractor. This novel feature representation also improves the generalization of the model.
Ali N. Salman, Carlos Busso
ICIP1
2020 MSP-Face Corpus: A Natural Audiovisual Emotional Database
abstract
Expressive behaviors conveyed during daily interactions are difficult to determine, because they often consist of a blend of different emotions. The complexity in expressive human communication is an important challenge to build and evaluate automatic systems that can reliably predict emotions. Emotion recognition systems are often trained with limited databases, where the emotions are either elicited or recorded by actors. These approaches do not necessarily reflect real emotions, creating a mismatch when the same emotion recognition systems are applied to practical applications. Developing rich emotional databases that reflect the complexity in the externalization of emotion is an important step to build better models to recognize emotions. This study presents the MSP-Face database, a natural audiovisual database obtained from video-sharing websites, where multiple individuals discuss various topics expressing their opinions and experiences. The natural recordings convey a broad range of emotions that are difficult to obtain with other alternative data collection protocols. A feature of the corpus is the addition of two sets. The first set includes videos that have been annotated with emotional labels using a crowd-sourcing protocol (9,370 recordings -- 24 hrs, 41 m). The second set includes similar videos without emotional labels (17,955 recordings -- 45 hrs, 57 m), offering the perfect infrastructure to explore semi-supervised and unsupervised machine-learning algorithms on natural emotional videos. This study describes the process of collecting and annotating the corpus. It also provides baselines over this new database using unimodal (audio, video) and multimodal emotional recognition systems.
Andrea Vidal, Ali N. Salman, Carlos Busso
ICMI2