Heysem Kaya

dblp:77/8212 · DBLP profile ↗
← Back
38ranked-venue papers
16as first author
17since 2021 · last 2025
0000-0001-7947-5508ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 9 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 8 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 11 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Using Emotionally Rich Speech Segments for Depression Prediction
abstract
The automatic assessment of depression through human voice has gained increasing interest due to its cost-effectiveness and non-invasiveness. This paper employs acoustic embeddings from emotionally rich speech segments for depression prediction, given the strong connection between depression and emotion expression. We leverage a large pre-trained model to predict depression. The public dimensional emotion model (PDEM) used in this study is fine-tuned for recognizing arousal, valence and dominance. We use PDEM for both extracting embeddings and selecting emotionally rich speech segments based on its arousal, valence, and dominance predictions. We advance the state-of-the-art performance on both the Androids corpus (Interview task) for depression detection, following the predetermined protocol, and the E-DAIC corpus for acoustic-based depression severity prediction, adhering to the 2019 Audio Visual Emotion Challenge (AVEC) protocol. The analysis demonstrates that emotionally rich speech segments contain more depression-related cues compared to emotion-neutral segments.
Heysem Kaya
ICASSP2
2025 Multi-Modal Multi-Task Affective States Recognition Based on Label Encoder Fusion
abstract
Despite recent advances in multi-modal approaches, recognizing the full range of human affective states, including emotions and sentiments, remains challenging due to complex interactions between different modalities and the hierarchical nature of affective states. This work presents a novel approach for multi-modal multi-task emotion and sentiment recognition that integrates audio, video, and text data. We introduce a Label Encoder Fusion Strategy, which produces and processes uni-modal emotion and sentiment predictions, which are used alongside modality-specific features during the fusion process to provide additional contextual information. We conduct elaborate multi-corpus experiments on the RAMAS, MELD, and CMU-MOSEI corpora. The proposed approach achieves state-of-the-art performance in both affective tasks. On MELD, we achieve a macro F1 (MF) of 40.9% and 67.02% for emotion and sentiment recognition. On CMU-MOSEI, the mean MF is 62.30% and MF is 62.00% for the same tasks.
Maxim Markitantov, Elena Ryumina, Heysem Kaya, Alexey Karpov 0001
INTERSPEECH3
2025 Privacy constrained fairness estimation for decision trees
abstract
Abstract The protection of sensitive data becomes more vital, as data increases in value and potency. Furthermore, the pressure increases from regulators and society on model developers to make their Artificial Intelligence (AI) models non-discriminatory. To boot, there is a need for interpretable, transparent AI models for high-stakes tasks. In general, measuring the fairness of any AI model requires the sensitive attributes of the individuals in the dataset, thus raising privacy concerns. In this work, the trade-offs between fairness (in terms of Statistical Parity (SP)), privacy (quantified with a budget), and interpretability are further explored in the context of Decision Trees (DTs) as intrinsically interpretable models. We propose a novel method, dubbed Privacy-Aware Fairness Estimation of Rules (PAFER), that can estimate SP in a Differential Privacy (DP)-aware manner for DTs. Our method is the first to assess algorithmic fairness on a rule-level, providing insight into sources of discrimination for policy makers. DP, making use of a third-party legal entity that securely holds this sensitive data, guarantees privacy by adding noise to the sensitive data. We experimentally compare several DP mechanisms. We show that using the Laplacian mechanism, the method is able to estimate SP with low error while guaranteeing the privacy of the individuals in the dataset with high certainty. We further show experimentally and theoretically that the method performs better for those DTs that humans generally find easier to interpret. Graphical abstract
Florian van der Steen, Fré Vink, Heysem Kaya
Appl. Intell.3
2024 Fairness in AI-Based Mental Health: Clinician Perspectives and Bias Mitigation
abstract
There is limited research on fairness in automated decision-making systems in the clinical domain, particularly in the mental health domain. Our study explores clinicians' perceptions of AI fairness through two distinct scenarios: violence risk assessment and depression phenotype recognition using textual clinical notes. We engage with clinicians through semi-structured interviews to understand their fairness perceptions and to identify appropriate quantitative fairness objectives for these scenarios. Then, we compare a set of bias mitigation strategies developed to improve at least one of the four selected fairness objectives. Our findings underscore the importance of carefully selecting fairness measures, as prioritizing less relevant measures can have a detrimental rather than a beneficial effect on model behavior in real-world clinical use.
Gizem Sogancioglu, Pablo Mosteiro, Albert Ali Salah, Floor Scheepers, Heysem Kaya
AIES (1)5
2023 The effects of gender bias in word embeddings on patient phenotyping in the mental health domain
abstract
Word embeddings, renowned for their role as superior semantic feature vector representation in diverse NLP tasks, can exhibit an undesired bias for stereotypical categories. The bias arises from the statistical and societal biases within the datasets used for training. In this study, we analyze the gender bias in four different pre-trained word embeddings for a range of affective computing tasks in the mental health domain including the detection of psychiatric disorders such as depression, and alcohol/substance abuse. We incorporate both contextual and non-contextual embeddings, which are trained not just on general domain data but also on data specific to the clinical domain. Our findings indicate that the bias in embeddings is towards different gender groups, depending on the type of embeddings and the training dataset. Furthermore, we highlight how these existing associations transfer to subsequent tasks and might even be amplified during supervised training for patient phenotyping. We also show that a simple method of data augmentation- swapping gender words - noticeably reduces bias in these subsequent tasks. The scripts to reproduce the results are available at: https:llgithub.comlgizemsogancioglulgender-bias-mental-health.
Gizem Sogancioglu, Heysem Kaya, Albert Ali Salah
ACII2
2023 Multimodal Personality Traits Assessment (MuPTA) Corpus: The Impact of Spontaneous and Read Speech
abstract
Automatic personality traits assessment (PTA) provides high-level, intelligible predictive inputs for subsequent critical downstream tasks, such as job interview recommendations and mental healthcare monitoring. In this work, we introduce a novel Multimodal Personality Traits Assessment (MuPTA) corpus. Our MuPTA corpus is unique in that it contains both spontaneous and read speech collected in the midly-resourced Russian language. We present a novel audio-visual approach for PTA that is used in order to set up baseline results on this corpus. We further analyze the impact of both spontaneous and read speech types on the PTA predictive performance. We find that for the audio modality, the PTA predictive performances on short signals are almost equal regardless of the speech type, while PTA using video modality is more accurate with spontaneous speech compared to read one regardless of the signal length.
Elena Ryumina, Dmitry Ryumin, Maxim Markitantov, Heysem Kaya, Alexey Karpov 0001
INTERSPEECH4
2022 3rd ICMI Workshop on Bridging Social Sciences and AI for Understanding Child Behaviour
abstract
Child behaviour is a topic of great scientific interest across a wide range of disciplines, including social sciences and artificial intelligence (AI). Knowledge in these different fields is not yet integrated to its full potential. The aim of this workshop was to bring researchers from these fields together. The first two workshops had a significant impact. In this workshop, we discussed topics such as the use of AI techniques to better examine and model interactions and children’s emotional development, analyzing head movement patterns with respect to child age. This workshop was a successful new step towards the objective of bridging social sciences and AI, attracting contributions from various academic fields on child behaviour analysis. This document summarizes the accepted papers.
Anika van der Klis, Heysem Kaya, Maryam Najafian, Saeid Safavi 0001
ICMI2
2022 Towards using Breathing Features for Multimodal Estimation of Depression Severity
abstract
Breathing patterns are shown to have strong correlations with emotional states, and hence have promise for automatic mood order prediction and analysis. An essential challenge here is the lack of ground truth for breathing sounds, especially for medical and archival datasets. In this study, we provide a cross-dataset approach for breathing pattern prediction and analyse the contribution of predicted breath signals for the detection of depressive states, using the DAIC-WOZ corpus. We use interpretable features in our models to provide actionable insights. Our experimental evaluation shows that in participants with higher depression scores (as indicated by the eight-item Patient Health Questionnaire, PHQ-8), breathing events tend to be shallow or slow. We furthermore tested linear and non-linear regression models with breathing, linguistic sentiment and conversational features, and show that these simple models outperform the AVEC17 Real-life Depression Recognition Sub-challenge baseline.
Francisca Pessanha, Heysem Kaya, Almila Akdag Salah, Albert Ali Salah
ICMI2
2022 Text-based Interpretable Depression Severity Modeling via Symptom Predictions
abstract
Mood disorders in general and depression in particular are common and their impact on individuals and society is high. Roughly 5% of adults worldwide suffer from depression. Commonly, depression diagnosis involves using questionnaires, either clinician-rated or self-reported. Due to the subjectivity in questionnaire methods and high human-related costs involved, there are ongoing efforts to find more objective and easily attainable depression markers. As is the case with recent audio, visual and linguistic applications, state-of-the-art approaches for automated depression severity prediction heavily depend on deep learning and black box modeling without explainability and interpretability considerations. However, for reasons ranging from regulations to understanding the extent and limitations of the model, the clinicians need to understand the decision making process of the model to confidently form their decisions. In this work, we focus on text-based depression severity level prediction on DAIC-WOZ corpus and benefit from PHQ-8 questionnaire items to predict the symptoms as interpretable high level features. We show that using a multi-task regression approach with state-of-the-art text-based features to predict the depression symptoms, it is possible to reach a viable test set Concordance Correlation Coefficient performance comparable to the state-of-the-art systems.
Floris Van Steijn, Gizem Sogancioglu, Heysem Kaya
ICMI3
2022 Complex Paralinguistic Analysis of Speech: Predicting Gender, Emotions and Deception in a Hierarchical Framework
abstract
In this paper, we present a hierarchical framework for complex paralinguistic analysis of speech including gender, emotions and deception recognition. The main idea of the framework is built upon the research on interrelation between various paralinguistic phenomena. It uses gender information to predict emotional states, and the outcome of the emotion recognition to predict the truthfulness of the speech. We use multiple datasets (aGender, Ruslana, EmoDB and DSD) to perform within-corpus and cross-corpus experiments using various performance measures. The experimental results reveal that gender-specific models improve the effectiveness of automatic speech emotion recognition in terms of Unweighted Average Recall up to an absolute 5.7%, and the integration of emotion predictions improves the F-score of automatic deception detection compared to our baseline by an absolute 4.7%. The obtained cross-validation results of 88.4 +/- 1.5% for deception detection beat the existing state-of-the-art by an absolute 2.8%.
Alena Velichko, Maxim Markitantov, Heysem Kaya, Alexey Karpov 0001
INTERSPEECH3
2022 Federated learning for violence incident prediction in a simulated cross-institutional psychiatric setting
abstract
Inpatient violence is a common and severe problem within psychiatry. Knowing who might become violent can influence staffing levels and mitigate severity. Predictive machine learning models can assess each patient’s likelihood of becoming violent based on clinical notes. Yet, while machine learning models benefit from having more data, data availability is limited as hospitals typically do not share their data for privacy preservation. Federated Learning (FL) can overcome the problem of data limitation by training models in a decentralised manner, without disclosing data between collaborators. However, although several FL approaches exist, none of these train Natural Language Processing models on clinical notes. In this work, we investigate the application of Federated Learning to clinical Natural Language Processing, applied to the task of Violence Risk Assessment by simulating a cross-institutional psychiatric setting. We train and compare four models: two local models, a federated model and a data-centralised model. Our results indicate that the federated model outperforms the local models and has similar performance as the data-centralised model. These findings suggest that Federated Learning can be used successfully in a cross-institutional setting and is a step towards new applications of Federated Learning based on clinical notes.
Thomas Borger, Pablo Mosteiro, Heysem Kaya, Emil Rijcken, Albert Ali Salah, Floor Scheepers, Marco Spruit
Expert Syst. Appl.3
2022 A Multimodal Approach for Mania Level Prediction in Bipolar Disorder
abstract
Bipolar disorder is a mental health disorder that causes mood swings that range from depression to mania. Clinical diagnosis of bipolar disorder is based on patient interviews and reports obtained from the relatives of the patients. Subsequently, the diagnosis depends on the experience of the expert, and there is co-morbidity with other mental disorders. Automated processing in the diagnosis of bipolar disorder can help providing quantitative indicators, and allow easier observations of the patients for longer periods. In this paper, we create a multimodal decision system for three level mania classification based on recordings of the patients in acoustic, linguistic, and visual modalities. The system is evaluated on the Turkish Bipolar Disorder corpus we have recently introduced to the scientific community. Comprehensive analysis of unimodal and multimodal systems, as well as fusion techniques, are performed. Using acoustic, linguistic, and visual features in a multimodal fusion system, we achieved a 64.8% unweighted average recall score, which advances the state-of-the-art performance on this dataset.
Pinar Baki, Heysem Kaya, Elvan Çiftçi, Hüseyin Güleç, Albert Ali Salah
IEEE Trans. Affect. Comput.2
2022 Modeling, Recognizing, and Explaining Apparent Personality From Videos
abstract
Explainability and interpretability are two critical aspects of decision support systems. Despite their importance, it is only recently that researchers are starting to explore these aspects. This paper provides an introduction to explainability and interpretability in the context of apparent personality recognition. To the best of our knowledge, this is the first effort in this direction. We describe a challenge we organized on explainability in first impressions analysis from video. We analyze in detail the newly introduced data set, evaluation protocol, proposed solutions and summarize the results of the challenge. We investigate the issue of bias in detail. Finally, derived from our study, we outline research opportunities that we foresee will be relevant in this area in the near future.
Hugo Jair Escalante, Heysem Kaya, Albert Ali Salah, Sergio Escalera, Yagmur Güçlütürk, Umut Güçlü, Xavier Baró, Isabelle Guyon, Júlio C. S. Jacques Júnior, Meysam Madadi, Stéphane Ayache, Evelyne Viegas, Furkan Gürpinar, Achmadnoer Sukma Wicaksana, Cynthia C. S. Liem, Marcel van Gerven, Rob van Lier
IEEE Trans. Affect. Comput.2
2021 Can mood primitives predict apparent personality?
abstract
First impressions play a critical role in shaping social interactions and consequently have a high impact on people’s lives. This study presents an explainable system that models apparent personality traits that influence first impressions as a function of automatically predicted arousal, valence and likeability (AVL) scores. To this end, we enrich the ChaLearn Looking at People - First Impressions (LAP-FI) dataset by annotating a portion of it for the AVL dimensions and carry out extensive uni-modal and multimodal experiments by using state-of-the-art acoustic, visual and linguistic features. We propose to use a glass-box model, namely, Explainable Boosting Machine, to model the Big Five personality traits. Our results demonstrate that personality trait impressions can be effectively predicted through the mood and likeability scores of a given video. We show that the proposed model, which is trained on only a few features, not only provides more meaningful explanations but also yields competitive performance (with a 0.09 Mean Absolute Error) compared to the state-of-the-art methods. The annotated benchmark dataset and the scripts to reproduce the results are available at: https://github.com/gizemsogancioglu/mood-project.
Gizem Sogancioglu, Heysem Kaya, Albert Ali Salah
ACII2
2021 2nd ICMI Workshop on Bridging Social Sciences and AI for Understanding Child Behaviour
abstract
Child behavior is a topic of great scientific interest across a wide range of disciplines, including social and behavioral sciences, as well as artificial intelligence (AI). The first workshop had a significant impact, and in this workshop, we aimed to bring together researchers from these fields to discuss topics such as using AI to better understand and model child behavioral and developmental processes, challenges and opportunities for AI in large-scale child behavior analysis, and implementing explainable ML/AI on sensitive child data. The workshop was a successful second step toward this objective, attracting contributions from many academic fields on child behavior analysis. This document summarizes the workshop’s events as well as the accepted papers and abstracts.
Saeid Safavi 0001, Heysem Kaya, Roy S. Hessels, Maryam Najafian, Sandra Hanekamp
ICMI2
2021 The INTERSPEECH 2021 Computational Paralinguistics Challenge: COVID-19 Cough, COVID-19 Speech, Escalation & Primates
abstract
The INTERSPEECH 2021 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the COVID-19 Cough and COVID-19 Speech Sub-Challenges, a binary classification on COVID-19 infection has to be made based on coughing sounds and speech; in the Escalation SubChallenge, a three-way assessment of the level of escalation in a dialogue is featured; and in the Primates Sub-Challenge, four species vs background need to be classified. We describe the Sub-Challenges, baseline feature extraction, and classifiers based on the 'usual' COMPARE and BoAW features as well as deep unsupervised representation learning using the AuDeep toolkit, and deep feature extraction from pre-trained CNNs using the Deep Spectrum toolkit; in addition, we add deep end-to-end sequential modelling, and partially linguistic analysis.
Björn W. Schuller, Anton Batliner, Christian Bergler, Cecilia Mascolo, Jing Han 0010, Iulia Lefter, Heysem Kaya, Shahin Amiriparian, Alice Baird, Lukas Stappen, Sandra Ottl, Maurice Gerczuk, Panagiotis Tzirakis, Chloë Siegele-Brown, Jagmohan Chauhan, Andreas Grammenos, Apinan Hasthanasombat, Dimitris Spathis, Tong Xia, Pietro Cicuta, Léon J. M. Rothkrantz, Joeri A. Zwerts, Jelle Treep, Casper S. Kaandorp
Interspeech7
2021 Introducing a Central African Primate Vocalisation Dataset for Automated Species Classification
abstract
Automated classification of animal vocalisations is a potentially powerful wildlife monitoring tool. Training robust classifiers requires sizable annotated datasets, which are not easily recorded in the wild. To circumvent this problem, we recorded four primate species under semi-natural conditions in a wildlife sanctuary in Cameroon with the objective to train a classifier capable of detecting species in the wild. Here, we introduce the collected dataset, describe our approach and initial results of classifier development. To increase the efficiency of the annotation process, we condensed the recordings with an energy/change based automatic vocalisation detection. Segmenting the annotated chunks into training, validation and test sets, initial results reveal up to 82% unweighted average recall (UAR) test set performance in four-class primate species classification.
Joeri A. Zwerts, Jelle Treep, Casper S. Kaandorp, Floor Meewis, Amparo C. Koot, Heysem Kaya
Interspeech6
2020 Bridging Social Sciences and AI for Understanding Child Behaviour
abstract
Child behaviour is a topic of wide scientific interest among many different disciplines, including social and behavioural sciences and artificial intelligence (AI). In this workshop, we aimed to connect researchers from these fields to address topics such as the usage of AI to better understand and model child behavioural and developmental processes, challenges and opportunities for AI in large-scale child behaviour analysis and implementing explainable ML/AI on sensitive child data. The workshop served as a successful first step towards this goal and attracted contributions from different research disciplines on the analysis of child behaviour. This paper provides a summary of the activities of the workshop and the accepted papers and abstracts.
Heysem Kaya, Roy S. Hessels, Maryam Najafian, Sandra Hanekamp, Saeid Safavi 0001
ICMI1
2020 Ensembling End-to-End Deep Models for Computational Paralinguistics Tasks: ComParE 2020 Mask and Breathing Sub-Challenges
abstract
This paper describes deep learning approaches for the Mask and Breathing Sub-Challenges (SCs), which are addressed by the INTERSPEECH 2020 Computational Paralinguistics Challenge. Motivated by outstanding performance of state-of-the-art end-to-end (E2E) approaches, we explore and compare effectiveness of different deep Convolutional Neural Network (CNN) architectures on raw data, log Mel-spectrograms, and Mel-Frequency Cepstral Coefficients. We apply a transfer learning approach to improve model’s efficiency and convergence speed. In the Mask SC, we conduct experiments with several pretrained CNN architectures on log-Mel spectrograms, as well as Support Vector Machines on baseline features. For the Breathing SC, we propose an ensemble deep learning system that exploits E2E learning and sequence prediction. The E2E model is based on 1D CNN operating on raw speech signals and is coupled with Long Short-Term Memory layers for sequence modeling. The second model works with log-Mel features and is based on a pretrained 2D CNN model stacked to Gated Recurrent Unit layers. To increase performance of our models in both SCs, we use ensembles of the best deep neural models obtained from N-fold cross-validation on combined challenge training and development datasets. Our results markedly outperform the challenge test set baselines in both SCs.
Maxim Markitantov, Denis Dresvyanskiy, Danila Mamontov, Heysem Kaya, Wolfgang Minker, Alexey Karpov 0001
INTERSPEECH4
2020 Is Everything Fine, Grandma? Acoustic and Linguistic Modeling for Robust Elderly Speech Emotion Recognition
abstract
Acoustic and linguistic analysis for elderly emotion recognition is an under-studied and challenging research direction, but essential for the creation of digital assistants for the elderly, as well as unobtrusive telemonitoring of elderly in their residences for mental healthcare purposes. This paper presents our contribution to the INTERSPEECH 2020 Computational Paralinguistics Challenge (ComParE) - Elderly Emotion Sub-Challenge, which is comprised of two ternary classification tasks for arousal and valence recognition. We propose a bi-modal framework, where these tasks are modeled using state-of-the-art acoustic and linguistic features, respectively. In this study, we demonstrate that exploiting task-specific dictionaries and resources can boost the performance of linguistic models, when the amount of labeled data is small. Observing a high mismatch between development and test set performances of various models, we also propose alternative training and decision fusion strategies to better estimate and improve the generalization performance.
Gizem Sogancioglu, Oxana Verkholyak, Heysem Kaya, Dmitrii Fedotov, Tobias Cadèe, Albert Ali Salah, Alexey Karpov 0001
INTERSPEECH3
2019 Hierarchical Two-level Modelling of Emotional States in Spoken Dialog Systems
abstract
Emotions occur in complex social interactions, and thus processing of isolated utterances may not be sufficient to grasp the nature of underlying emotional states. Dialog speech provides useful information about context that explains nuances of emotions and their transitions. Context can be defined on different levels; this paper proposes a hierarchical context modelling approach based on RNN-LSTM architecture, which models acoustical context on the frame level and partner's emotional context on the dialog level. The method is proved effective together with cross-corpus training setup and domain adaptation technique in a set of speaker independent cross-validation experiments on IEMOCAP corpus for three levels of activation and valence classification. As a result, the state-of-the-art on this corpus is advanced for both dimensions using only acoustic modality.
Oxana Verkholyak, Dmitrii Fedotov, Heysem Kaya, Alexey Karpov 0001
ICASSP3
2018 LSTM Based Cross-corpus and Cross-task Acoustic Emotion Recognition
Heysem Kaya, Dmitrii Fedotov, Ali Yesilkanat, Oxana Verkholyak, Alexey Karpov 0001
INTERSPEECH1
2018 Feature Selection and Multimodal Fusion for Estimating Emotions Evoked by Movie Clips
abstract
Perceptual understanding of media content has many applications, including content-based retrieval, marketing, content optimization, psychological assessment, and affect-based learning. In this paper, we model audio visual features extracted from videos via machine learning approaches to estimate the affective responses of the viewers. We use the LIRIS-ACCEDE dataset and the MediaEval 2017 Challenge setting to evaluate the proposed methods. This dataset is composed of movies of professional or amateur origin, annotated with viewers' arousal, valence, and fear scores. We extract a number of audio features, such as Mel-frequency Cepstral Coefficients, and visual features, such as dense SIFT, hue-saturation histogram, and features from a deep neural network trained for object recognition. We contrast two different approaches in the paper, and report experiments with different fusion and smoothing strategies. We demonstrate the benefit of feature selection and multimodal fusion on estimating affective responses to movie segments.
Yasemin Timar, Nihan Karslioglu, Heysem Kaya, Albert Ali Salah
ICMR3
2018 Efficient and effective strategies for cross-corpus acoustic emotion recognition
Heysem Kaya, Alexey Karpov 0001
Neurocomputing1
2017 Introducing Weighted Kernel Classifiers for Handling Imbalanced Paralinguistic Corpora: Snoring, Addressee and Cold
Heysem Kaya, Alexey Karpov 0001
INTERSPEECH1
2017 Emotion, age, and gender classification in children's speech by humans and machines
Heysem Kaya, Albert Ali Salah, Alexey Karpov 0001, Olga V. Frolova, Aleksei Grigorev, Elena E. Lyakso
Comput. Speech Lang.1
2017 Video-based emotion recognition in the wild using deep transfer learning and score fusion
Heysem Kaya, Furkan Gürpinar, Albert Ali Salah
Image Vis. Comput.1
2016 Multimodal fusion of audio, scene, and face features for first impression estimation
abstract
Affective computing, particularly emotion and personality trait recognition, is of increasing interest in many research disciplines. The interplay of emotion and personality shows itself in the first impression left on other people. Moreover, the ambient information, e.g. the environment and objects surrounding the subject, also affect these impressions. In this work, we employ pre-trained Deep Convolutional Neural Networks to extract facial emotion and ambient information from images for predicting apparent personality. We also investigate Local Gabor Binary Patterns from Three Orthogonal Planes video descriptor and acoustic features extracted via the popularly used openSMILE tool. We subsequently propose classifying features using a Kernel Extreme Learning Machine and fusing their predictions. The proposed system is applied to the ChaLearn Challenge on First Impression Recognition, achieving the winning test set accuracy of 0.913, averaged over the “Big Five” personality traits.
Furkan Gürpinar, Heysem Kaya, Albert Ali Salah
ICPR2
2016 Fusing Acoustic Feature Representations for Computational Paralinguistics Tasks
Heysem Kaya, Alexey Karpov 0001
INTERSPEECH1
2016 Robust Acoustic Emotion Recognition Based on Cascaded Normalization and Extreme Learning Machines
Heysem Kaya, Alexey Karpov 0001, Albert Ali Salah
ISNN1
2015 Contrasting and Combining Least Squares Based Learners for Emotion Recognition in the Wild
abstract
This paper presents our contribution to ACM ICMI 2015 Emotion Recognition in the Wild Challenge (EmotiW 2015). We participate in both static facial expression (SFEW) and audio-visual emotion recognition challenges. In both challenges, we use a set of visual descriptors and their early and late fusion schemes. For AFEW, we also exploit a set of popularly used spatio-temporal modeling alternatives and carry out multi-modal fusion. For classification, we employ two least squares regression based learners that are shown to be fast and accurate on former EmotiW Challenge corpora. Specifically, we use Partial Least Squares Regression (PLS) and Kernel Extreme Learning Machines (ELM), which is closely related to Kernel Regularized Least Squares. We use a General Procrustes Analysis (GPA) based alignment for face registration. By employing different alignments, descriptor types, video modeling strategies and classifiers, we diversify learners to improve the final fusion performance. Test set accuracies reached in both challenges are relatively 25% above the respective baselines.
Heysem Kaya, Furkan Gürpinar, Sadaf Afshar, Albert Ali Salah
ICMI1
2015 Fisher vectors with cascaded normalization for paralinguistic analysis
abstract
Computational Paralinguistics has several unresolved issues, one of which is coping with large variability due to speakers, spoken content and corpora. In this paper, we address the variability compensation issue by proposing a novel method composed of i) Fisher vector encoding of low level descrip-tors extracted from the signal, ii) speaker z-normalization ap-plied after speaker clustering iii) non-linear normalization of features and iv) classification based on Kernel Extreme Learn-ing Machines and Partial Least Squares regression. For ex-perimental validation, we apply the proposed method on IN-TERSPEECH 2015 Computational Paralinguistics Challenge (ComParE 2015), Eating Condition sub-challenge, which is a seven-class classification task. In our preliminary experiments, the proposed method achieves an Unweighted Average Recall (UAR) score of 83.1%, outperforming the challenge test set baseline UAR (65.9%) by a large margin.
Heysem Kaya, Alexey Karpov 0001, Albert Ali Salah
INTERSPEECH1
2015 Random Discriminative Projection Based Feature Selection with Application to Conflict Recognition
abstract
Computational paralinguistics deals with underlying meaning of the verbal messages, which is of interest in manifold applications ranging from intelligent tutoring systems to affect sensitive robots. The state-of-the-art pipeline of paralinguistic speech analysis utilizes brute-force feature extraction, and the features need to be tailored according to the relevant task. In this work, we extend a recent discriminative projection based feature selection method using the power of stochasticity to overcome local minima and to reduce the computational complexity. The proposed approach assigns weights both to groups and to features individually in many randomly selected contexts and then combines them for a final ranking. The efficacy of the proposed method is shown in a recent paralinguistic challenge corpus to detect level of conflict in dyadic and group conversations. We advance the state-of-the-art in this corpus using the INTERSPEECH 2013 Challenge protocol.
Heysem Kaya, Tugçe Özkaptan, Albert Ali Salah, Fikret S. Gürgen
IEEE Signal Process. Lett.1
2014 CCA based feature selection with application to continuous depression recognition from acoustic speech features
abstract
In this study we make use of Canonical Correlation Analysis (CCA) based feature selection for continuous depression recognition from speech. Besides its common use in multi-modal/multi-view feature extraction, CCA can be easily employed as a feature selector. We introduce several novel ways of CCA based filter (ranking) methods, showing their relations to previous work. We test the suitability of proposed methods on the AVEC 2013 dataset under the ACM MM 2013 Challenge protocol. Using 17% of features, we obtained a relative improvement of 30% on the challenge's test-set baseline Root Mean Square Error.
Heysem Kaya, Florian Eyben, Albert Ali Salah, Björn W. Schuller
ICASSP1
2014 Speaker- and Corpus-Independent Methods for Affect Classification in Computational Paralinguistics
abstract
The analysis of spoken emotions is of increasing interest in human computer interaction, in order to drive the machine communication into a humane manner. It has manifold applications ranging from intelligent tutoring systems to affect sensitive robots, from smart call centers to patient telemonitoring. In general the study of computational paralinguistics, which covers the analysis of speaker states and traits, faces with real life challenges of inter-speaker and inter-corpus variability. In this paper, a brief summary of the progress and future directions of my PhD study titled Adaptive Mixture Models for Speech Emotion Recognition that targets these challenges are given. An automatic mixture model selection method for Mixture of Factor Analyzers is proposed for modeling high dimensional data. To provide the mentioned statistical method a compact set of potent features, novel feature selection methods based on Canonical Correlation Analysis are introduced.
Heysem Kaya
ICMI1
2014 Combining Modality-Specific Extreme Learning Machines for Emotion Recognition in the Wild
abstract
This paper presents our contribution to ACM ICMI 2014 Emotion Recognition in the Wild Challenge and Workshop. The proposed system utilizes Extreme Learning Machines (ELM) for modeling modality-specific features and combines the scores for final prediction. The state-of-the-art results in acoustic and visual emotion recognition are obtained either using deep Neural Networks (DNN) or Support Vector Machines (SVM). The ELM paradigm is proposed as a fast and accurate alternative to these two popular machine learning methods. Benefiting from fast learning advantage of ELM, we carry out extensive tests on the data using moderate computational resources. In the video modality, we test combination of regional visual features obtained from the inner face. In the audio modality, we carry out tests to enhance training via other emotional corpora. We further investigate the suitability of several recently proposed feature selection approaches to prune the acoustic features. In our study, the best results for both modalities are obtained with Kernel ELM compared to basic ELM. On the challenge test set, we obtain 37.84%, 39.07% and 44.23% classification accuracies for audio, video and multimodal fusion, respectively.
Heysem Kaya, Albert Ali Salah
ICMI1
2014 Canonical correlation analysis and local fisher discriminant analysis based multi-view acoustic feature reduction for physical load prediction
abstract
In this study we present our system for INTERSPEECH 2014 Computational Paralinguistics Challenge (ComParE 2014), Physical Load Sub-challenge (PLS). Our contribution is twofold. First, we propose using Low Level Descriptor (LLD) information as hints, so as to partition the feature space into meaningful subsets called views. We also show the virtue of commonly employed feature projections, such as Canoni-cal Correlation Analysis (CCA) and Local Fisher Discriminant Analysis (LFDA) as ranking feature selectors. Results indicate the superiority of multi-view feature reduction approach to its single-view counterpart. Moreover, the discriminative projec-tion matrices are observed to provide valuable information for feature selection, which generalize better than the projection it-self. In our preliminary experiments we reached 75.35 % Un-weighted Average Recall (UAR) on PLS test set, using CCA based multi-view feature selection.
Heysem Kaya, Tugçe Özkaptan, Albert Ali Salah, Fikret S. Gürgen
INTERSPEECH1
2014 Eyes Whisper Depression: A CCA based Multimodal Approach
abstract
This paper presents our work on ACM MM Audio Visual Emotion Corpus 2013 (AVEC 2013) depression recognition sub-challenge using the baseline features in accordance with the challenge protocol. We use Canonical Correlation Analysis for audio-visual fusion as well as covariate extraction for the target task. The video baseline provides histograms of local phase quantization features extracted from 4x4=16 regions of the detected face. We summarize the video features over segments of length 20 seconds using mode and range functionals. We observe that features of range functional that measure the variance tendency provides statistically significantly higher canonical correlation than mode functional features that measure the mean tendency. Moreover, when audio-visual features are used with varying number of covariates per region, the regions that were consistently found the best are the ones corresponding to two eyes and the right part of the mouth.
Heysem Kaya, Albert Ali Salah
ACM Multimedia1