Alexey Karpov 0001

dblp:10/300 · also Alexey A. Karpov · DBLP profile ↗
← Back
44ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0003-3424-652XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 7 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Audio-visual occlusion-robust gender recognition and age estimation approach based on multi-task cross-modal attention
Maxim Markitantov, Elena Ryumina, Alexey Karpov 0001
Expert Syst. Appl.3
2025 A Multilingual Telegram Chatbot for Mental Health Data Collection
Danila Mamontov, Alexey Karpov 0001, Wolfgang Minker
ICMI2
2025 Multi-Modal Multi-Task Affective States Recognition Based on Label Encoder Fusion
abstract
Despite recent advances in multi-modal approaches, recognizing the full range of human affective states, including emotions and sentiments, remains challenging due to complex interactions between different modalities and the hierarchical nature of affective states. This work presents a novel approach for multi-modal multi-task emotion and sentiment recognition that integrates audio, video, and text data. We introduce a Label Encoder Fusion Strategy, which produces and processes uni-modal emotion and sentiment predictions, which are used alongside modality-specific features during the fusion process to provide additional contextual information. We conduct elaborate multi-corpus experiments on the RAMAS, MELD, and CMU-MOSEI corpora. The proposed approach achieves state-of-the-art performance in both affective tasks. On MELD, we achieve a macro F1 (MF) of 40.9% and 67.02% for emotion and sentiment recognition. On CMU-MOSEI, the mean MF is 62.30% and MF is 62.00% for the same tasks.
Maxim Markitantov, Elena Ryumina, Heysem Kaya, Alexey Karpov 0001
INTERSPEECH4
2025 Multi-corpus emotion recognition method based on cross-modal gated attention fusion
Elena Ryumina, Dmitry Ryumin, Alexandr Axyonov, Denis Ivanko, Alexey Karpov 0001
Pattern Recognit. Lett.5
2024 Audio-Visual Speech Recognition In-The-Wild: Multi-Angle Vehicle Cabin Corpus and Attention-Based Method
abstract
Audio-visual speech recognition (AVSR) gains increasing attention as an important part of human-machine interaction. However, the publicly available corpora are limited, particularly in driving conditions with prevalent background noise. Research so far has been collected in constrained environments, and thus cannot reflect the true performance of AVSR systems in real-world scenarios. Moreover, data for languages other than English is often unavailable. To meet the request for research on AVSR in unconstrained driving conditions, this paper presents a corpus collected ‘in-the-wild’. We propose a cross-modal attention method enhancing multi-angle AVSR for vehicles, leveraging visual context to improve accuracy and noise robustness. Our proposed model achieves state-of-the-art (SOTA) results with 98.65% accuracy in recognizing driver voice commands. For more details, visit our project page1.
Alexandr Axyonov, Dmitry Ryumin, Denis Ivanko, Alexey M. Kashevnik, Alexey Karpov 0001
ICASSP5
2024 OCEAN-AI: open multimodal framework for personality traits assessment and HR-processes automatization
Elena Ryumina, Dmitry Ryumin, Alexey Karpov 0001
INTERSPEECH3
2024 Audio-visual speech recognition based on regulated transformer and spatio-temporal fusion strategy for driver assistive systems
Dmitry Ryumin, Alexandr Axyonov, Elena Ryumina, Denis Ivanko, Alexey M. Kashevnik, Alexey Karpov 0001
Expert Syst. Appl.6
2024 OCEAN-AI framework with EmoFormer cross-hemiface attention approach for personality traits assessment
Elena Ryumina, Maxim Markitantov, Dmitry Ryumin, Alexey Karpov 0001
Expert Syst. Appl.4
2024 Gated Siamese Fusion Network based on multimodal deep and hand-crafted features for personality traits assessment
Elena Ryumina, Maxim Markitantov, Dmitry Ryumin, Alexey Karpov 0001
Pattern Recognit. Lett.4
2023 Multimodal Personality Traits Assessment (MuPTA) Corpus: The Impact of Spontaneous and Read Speech
abstract
Automatic personality traits assessment (PTA) provides high-level, intelligible predictive inputs for subsequent critical downstream tasks, such as job interview recommendations and mental healthcare monitoring. In this work, we introduce a novel Multimodal Personality Traits Assessment (MuPTA) corpus. Our MuPTA corpus is unique in that it contains both spontaneous and read speech collected in the midly-resourced Russian language. We present a novel audio-visual approach for PTA that is used in order to set up baseline results on this corpus. We further analyze the impact of both spontaneous and read speech types on the PTA predictive performance. We find that for the audio modality, the PTA predictive performances on short signals are almost equal regardless of the speech type, while PTA using video modality is more accurate with spontaneous speech compared to read one regardless of the signal length.
Elena Ryumina, Dmitry Ryumin, Maxim Markitantov, Heysem Kaya, Alexey Karpov 0001
INTERSPEECH5
2022 MIDriveSafely: Multimodal Interaction for Drive Safely
abstract
In this paper, we present a novel multimodal interaction application to help car drivers and increase their road safety. MIDriveSafely is a mobile application that provides the following functions: (1) detect dangerous situations based on video information from a smartphone front-facing camera, such as drowsiness/sleepiness, phone usage while driving, eating, smoking, unfastened seat belt, etc.; gives a feedback to the driver (2) provide entertainment (e.g. rock-paper-scissors game, based on automatic speech recognition), (3) provide voice control capabilities to navigation/multimedia systems of a smartphone (potentially vehicle systems such as lighting conditions/climate control). Speech recognition in driving conditions is highly challenging due to acoustic noises, active head turns, pose variations, distance to recording devices, etc. MIDriveSafely incorporates driver's audio-visual speech recognition (DAVIS) system and uses it for multimodal interaction. Along with this, the original DriveSafely system is used for dangerous state detection. MIDriveSafely improves upon existing driver monitoring applications using multimodal (mainly audio-visual) information. MIDriveSafely motivates people to drive in a safer manner by providing the feedback to the drivers and by creating a fun user experience.
Denis Ivanko, Alexey M. Kashevnik, Dmitry Ryumin, Andrey Kitenko, Alexandr Axyonov, Igor Lashkov, Alexey Karpov 0001
ICMI7
2022 DAVIS: Driver's Audio-Visual Speech recognition
Denis Ivanko, Dmitry Ryumin, Alexey M. Kashevnik, Alexandr Axyonov, Andrey Kitenko, Igor Lashkov, Alexey Karpov 0001
INTERSPEECH7
2022 Biometric Russian Audio-Visual Extended MASKS (BRAVE-MASKS) Corpus: Multimodal Mask Type Recognition Task
Maxim Markitantov, Elena Ryumina, Dmitry Ryumin, Alexey Karpov 0001
INTERSPEECH4
2022 Complex Paralinguistic Analysis of Speech: Predicting Gender, Emotions and Deception in a Hierarchical Framework
abstract
In this paper, we present a hierarchical framework for complex paralinguistic analysis of speech including gender, emotions and deception recognition. The main idea of the framework is built upon the research on interrelation between various paralinguistic phenomena. It uses gender information to predict emotional states, and the outcome of the emotion recognition to predict the truthfulness of the speech. We use multiple datasets (aGender, Ruslana, EmoDB and DSD) to perform within-corpus and cross-corpus experiments using various performance measures. The experimental results reveal that gender-specific models improve the effectiveness of automatic speech emotion recognition in terms of Unweighted Average Recall up to an absolute 5.7%, and the integration of emotion predictions improves the F-score of automatic deception detection compared to our baseline by an absolute 4.7%. The obtained cross-validation results of 88.4 +/- 1.5% for deception detection beat the existing state-of-the-art by an absolute 2.8%.
Alena Velichko, Maxim Markitantov, Heysem Kaya, Alexey Karpov 0001
INTERSPEECH4
2022 RUSAVIC Corpus: Russian Audio-Visual Speech in Cars
abstract
We present a new audio-visual speech corpus (RUSAVIC) recorded in a car environment and designed for noise-robust speech recognition. Our goal was to produce a speech corpus which is natural (recorded in real driving conditions), controlled (providing different SNR levels by windows open/closed, moving/parked vehicle, etc.), and adequate size (the amount of data is enough to train state-of-the-art NN approaches). We focus on the problem of audio-visual speech recognition: with the use of automated lip-reading to improve the performance of audio-based speech recognition in the presence of severe acoustic noise caused by road traffic. We also describe the equipment and procedures used to create RUSAVIC corpus. Data are collected in a synchronous way through several smartphones located at different angles and equipped with FullHD video camera and microphone. The corpus includes the recordings of 20 drivers with minimum of 10 recording sessions for each. Besides providing a detailed description of the dataset and its collection pipeline, we evaluate several popular audio and visual speech recognition methods and present a set of baseline recognition results. At the moment RUSAVIC is a unique audio-visual corpus for the Russian language that is recorded in-the-wild condition and we make it publicly available.
Denis Ivanko, Alexandr Axyonov, Dmitry Ryumin, Alexey M. Kashevnik, Alexey Karpov 0001
LREC5
2022 In search of a robust facial expressions recognition model: A large-scale visual cross-corpus study
Elena Ryumina, Denis Dresvyanskiy, Alexey Karpov 0001
Neurocomputing3
2021 Annotation Confidence vs. Training Sample Size: Trade-Off Solution for Partially-Continuous Categorical Emotion Recognition
Elena Ryumina, Oxana Verkholyak, Alexey Karpov 0001
Interspeech3
2021 Ensemble-Within-Ensemble Classification for Escalation Prediction from Speech
Oxana Verkholyak, Denis Dresvyanskiy, Anastasia Dvoynikova, Denis Kotov, Elena Ryumina, Alena Velichko, Danila Mamontov, Wolfgang Minker, Alexey Karpov 0001
Interspeech9
2020 Ensembling End-to-End Deep Models for Computational Paralinguistics Tasks: ComParE 2020 Mask and Breathing Sub-Challenges
abstract
This paper describes deep learning approaches for the Mask and Breathing Sub-Challenges (SCs), which are addressed by the INTERSPEECH 2020 Computational Paralinguistics Challenge. Motivated by outstanding performance of state-of-the-art end-to-end (E2E) approaches, we explore and compare effectiveness of different deep Convolutional Neural Network (CNN) architectures on raw data, log Mel-spectrograms, and Mel-Frequency Cepstral Coefficients. We apply a transfer learning approach to improve model’s efficiency and convergence speed. In the Mask SC, we conduct experiments with several pretrained CNN architectures on log-Mel spectrograms, as well as Support Vector Machines on baseline features. For the Breathing SC, we propose an ensemble deep learning system that exploits E2E learning and sequence prediction. The E2E model is based on 1D CNN operating on raw speech signals and is coupled with Long Short-Term Memory layers for sequence modeling. The second model works with log-Mel features and is based on a pretrained 2D CNN model stacked to Gated Recurrent Unit layers. To increase performance of our models in both SCs, we use ensembles of the best deep neural models obtained from N-fold cross-validation on combined challenge training and development datasets. Our results markedly outperform the challenge test set baselines in both SCs.
Maxim Markitantov, Denis Dresvyanskiy, Danila Mamontov, Heysem Kaya, Wolfgang Minker, Alexey Karpov 0001
INTERSPEECH6
2020 Is Everything Fine, Grandma? Acoustic and Linguistic Modeling for Robust Elderly Speech Emotion Recognition
abstract
Acoustic and linguistic analysis for elderly emotion recognition is an under-studied and challenging research direction, but essential for the creation of digital assistants for the elderly, as well as unobtrusive telemonitoring of elderly in their residences for mental healthcare purposes. This paper presents our contribution to the INTERSPEECH 2020 Computational Paralinguistics Challenge (ComParE) - Elderly Emotion Sub-Challenge, which is comprised of two ternary classification tasks for arousal and valence recognition. We propose a bi-modal framework, where these tasks are modeled using state-of-the-art acoustic and linguistic features, respectively. In this study, we demonstrate that exploiting task-specific dictionaries and resources can boost the performance of linguistic models, when the amount of labeled data is small. Observing a high mismatch between development and test set performances of various models, we also propose alternative training and decision fusion strategies to better estimate and improve the generalization performance.
Gizem Sogancioglu, Oxana Verkholyak, Heysem Kaya, Dmitrii Fedotov, Tobias Cadèe, Albert Ali Salah, Alexey Karpov 0001
INTERSPEECH7
2020 TheRuSLan: Database of Russian Sign Language
abstract
In this paper, a new Russian sign language multimedia database TheRuSLan is presented. The database includes lexical units (single words and phrases) from Russian sign language within one subject area, namely, “food products at the supermarket”, and was collected using MS Kinect 2.0 device including both FullHD video and the depth map modes, which provides new opportunities for the lexicographical description of the Russian sign language vocabulary and enhances research in the field of automatic gesture recognition. Russian sign language has an official status in Russia, and over 120,000 deaf people in Russia and its neighboring countries use it as their first language. Russian sign language has no writing system, is poorly described and belongs to the low-resource languages. The authors formulate the basic principles of annotation of sign words, based on the collected data, and reveal the content of the collected database. In the future, the database will be expanded and comprise more lexical units. The database is explicitly made for the task of creating an automatic system for Russian sign language recognition.
Ildar Kagirov, Denis Ivanko, Dmitry Ryumin, Alexander A. Petrovsky, Alexey Karpov 0001
LREC5
2020 Class-based LSTM Russian Language Model with Linguistic Information
abstract
In the paper, we present class-based LSTM Russian language models (LMs) with classes generated with the use of both word frequency and linguistic information data, obtained with the help of the “VisualSynan” software from the AOT project. We have created LSTM LMs with various numbers of classes and compared them with word-based LM and class-based LM with word2vec class generation in terms of perplexity, training time, and WER. In addition, we performed a linear interpolation of LSTM language models with the baseline 3-gram language model. The LSTM language models were used for very large vocabulary continuous Russian speech recognition at an N-best list rescoring stage. We achieved significant progress in training time reduction with only slight degradation in recognition accuracy comparing to the word-based LM. In addition, our LM with classes generated using linguistic information outperformed LM with classes generated using word2vec. We achieved WER of 14.94 % at our own speech corpus of continuous Russian speech that is 15 % relative reduction with respect to the baseline 3-gram model.
Irina S. Kipyatkova, Alexey Karpov 0001
LREC2
2019 Hierarchical Two-level Modelling of Emotional States in Spoken Dialog Systems
abstract
Emotions occur in complex social interactions, and thus processing of isolated utterances may not be sufficient to grasp the nature of underlying emotional states. Dialog speech provides useful information about context that explains nuances of emotions and their transitions. Context can be defined on different levels; this paper proposes a hierarchical context modelling approach based on RNN-LSTM architecture, which models acoustical context on the frame level and partner's emotional context on the dialog level. The method is proved effective together with cross-corpus training setup and domain adaptation technique in a set of speaker independent cross-validation experiments on IEMOCAP corpus for three levels of activation and valence classification. As a result, the state-of-the-art on this corpus is advanced for both dimensions using only acoustic modality.
Oxana Verkholyak, Dmitrii Fedotov, Heysem Kaya, Alexey Karpov 0001
ICASSP5
2019 Cross-Corpus Data Augmentation for Acoustic Addressee Detection
abstract
Acoustic addressee detection (AD) is a modern paralinguistic and dialogue challenge that especially arises in voice assistants.In the present study, we distinguish addressees in two settings (a conversation between several people and a spoken dialogue system, and a conversation between several adults and a child) and introduce the first competitive baseline (unweighted average recall equals 0.891) for the Voice Assistant Conversation Corpus that models the first setting.We jointly solve both classification problems, using three models: a linear support vector machine dealing with acoustic functionals and two neural networks utilising raw waveforms alongside with acoustic low-level descriptors.We investigate how different corpora influence each other, applying the mixup approach to data augmentation.We also study the influence of various acoustic context lengths on AD.Two-second speech fragments turn out to be sufficient for reliable AD.Mixup is shown to be beneficial for merging acoustic data (extracted features but not raw waveforms) from different domains that allows us to reach a higher classification performance on human-machine AD and also for training a multipurpose neural network that is capable of solving both human-machine and adult-child AD problems.
Oleg Akhtiamov, Ingo Siegert, Alexey Karpov 0001, Wolfgang Minker
SIGdial3
2018 LSTM Based Cross-corpus and Cross-task Acoustic Emotion Recognition
Heysem Kaya, Dmitrii Fedotov, Ali Yesilkanat, Oxana Verkholyak, Alexey Karpov 0001
INTERSPEECH6
2018 Efficient and effective strategies for cross-corpus acoustic emotion recognition
Heysem Kaya, Alexey Karpov 0001
Neurocomputing2
2017 Speech and Text Analysis for Multimodal Addressee Detection in Human-Human-Computer Interaction
Oleg Akhtiamov, Maxim Sidorov, Alexey Karpov 0001, Wolfgang Minker
INTERSPEECH3
2017 Introducing Weighted Kernel Classifiers for Handling Imbalanced Paralinguistic Corpora: Snoring, Addressee and Cold
Heysem Kaya, Alexey Karpov 0001
INTERSPEECH2
2017 Emotion, age, and gender classification in children's speech by humans and machines
Heysem Kaya, Albert Ali Salah, Alexey Karpov 0001, Olga V. Frolova, Aleksei Grigorev, Elena E. Lyakso
Comput. Speech Lang.3
2016 Fusing Acoustic Feature Representations for Computational Paralinguistics Tasks
Heysem Kaya, Alexey Karpov 0001
INTERSPEECH2
2016 Robust Acoustic Emotion Recognition Based on Cascaded Normalization and Extreme Learning Machines
Heysem Kaya, Alexey Karpov 0001, Albert Ali Salah
ISNN2
2016 Language Models with RNNs for Rescoring Hypotheses of Russian ASR
Irina S. Kipyatkova, Alexey Karpov 0001
ISNN2
2015 Fisher vectors with cascaded normalization for paralinguistic analysis
abstract
Computational Paralinguistics has several unresolved issues, one of which is coping with large variability due to speakers, spoken content and corpora. In this paper, we address the variability compensation issue by proposing a novel method composed of i) Fisher vector encoding of low level descrip-tors extracted from the signal, ii) speaker z-normalization ap-plied after speaker clustering iii) non-linear normalization of features and iv) classification based on Kernel Extreme Learn-ing Machines and Partial Least Squares regression. For ex-perimental validation, we apply the proposed method on IN-TERSPEECH 2015 Computational Paralinguistics Challenge (ComParE 2015), Eating Condition sub-challenge, which is a seven-class classification task. In our preliminary experiments, the proposed method achieves an Unweighted Average Recall (UAR) score of 83.1%, outperforming the challenge test set baseline UAR (65.9%) by a large margin.
Heysem Kaya, Alexey Karpov 0001, Albert Ali Salah
INTERSPEECH2
2014 Audio-visual signal processing in a multimodal assisted living environment
abstract
https://doi.org/10.21437/interspeech.2014-267
Alexey Karpov 0001, Lale Akarun, Hulya Yalcin, Alexander L. Ronzhin, Baris Evrim Demiröz, Aysun Çoban, Milos Zelezný
INTERSPEECH1
2014 Introduction to the special issue on processing under-resourced languages
Laurent Besacier, Etienne Barnard, Alexey Karpov 0001, Tanja Schultz
Speech Commun.3
2014 Automatic speech recognition for under-resourced languages: A survey
Laurent Besacier, Etienne Barnard, Alexey Karpov 0001, Tanja Schultz
Speech Commun.3
2014 Large vocabulary Russian speech recognition using syntactico-statistical language modeling
Alexey Karpov 0001, Konstantin Markov, Irina S. Kipyatkova, Daria Vazhenina, Andrey Ronzhin
Speech Commun.1
2012 Analysis of Long-distance Word Dependencies and Pronunciation Variability at Conversational Russian Speech Recognition
Irina S. Kipyatkova, Alexey Karpov 0001, Vasilisa Verkhodanova, Milos Zelezný
FedCSIS2
2011 Very Large Vocabulary ASR for Spoken Russian with Syntactic and Morphemic Analysis
Alexey Karpov 0001, Irina S. Kipyatkova, Andrey Ronzhin
INTERSPEECH1
2010 Multimodal Human Computer Interaction with MIDAS Intelligent Infokiosk
abstract
In this paper, we present an intelligent information kiosk called MIDAS (Multimodal Interactive-Dialogue Automaton for Self-service), including its hardware and software architecture, stages of deployment of speech recognition and synthesis technologies. MIDAS uses the methodology Wizard of Oz (WOZ) that allows an expert to correct speech recognition results and control the dialogue flow. User statistics of the multimodal human computer interaction (HCI) have been analyzed for the operation of the kiosk in the automatic and automated modes. The infokiosk offers information about the structure and staff of laboratories, the location and phones of departments and employees of the institution. The multimodal user interface is provided with a touch screen, natural speech input and head and manual gestures, both for ordinary and physically handicapped users.
Alexey Karpov 0001, Andrey Ronzhin, Irina S. Kipyatkova, Alexander L. Ronzhin, Lale Akarun
ICPR1
2010 Viseme-dependent weight optimization for CHMM-based audio-visual speech recognition
abstract
The aim of the present study is to investigate some key challenges of the audio-visual speech recognition technology, such as asynchrony modeling of multimodal speech, estimation of auditory and visual speech significance, as well as stream weight optimization. Our research shows that the use of viseme-dependent significance weights improves the performance of state asynchronous CHMM-based speech recognizer. In addition, for a state synchronous MSHMMbased recognizer, fewer errors can be achieved using stationary time delays of visual data with respect to the corresponding audio signal. Evaluation experiments showed that individual audio-visual stream weights for each visemephoneme pair lead to relative reduction of WER by 20%. Index Terms: multimodal speech, audio-visual processing, Hidden Markov Models, asynchrony, significance weights
Alexey Karpov 0001, Andrey Ronzhin, Konstantin Markov, Milos Zelezný
INTERSPEECH1
2009 Audio-visual speech asynchrony modeling in a talking head
abstract
V tomto článku je navržen systém audiovizuální syntézy řeči obsahující modelování asynchronie mezi zvukovou a vizuální modalitou řeči. Studie reálných nahrávek obsažených v řečových databázích nám poskytují požadované údaje k pochopení problému modalit asynchronie, která je částečně způsobena koartikulací. Byl vypracován soubor kontextově závislých pravidel časování a doporučení zajišťující synchronizaci zvukové a vizuální řeči tak, že animace mluvící hlavy je více přirozená. Kognitivní ohodnocení systému mluvící hlavy, který je nastaven pro Ruštinu a implementující původní model asynchronie, ukazuje vysokou srozumitelnost a přirozenost syntetizované audiovizuální řeči.
Alexey Karpov 0001, Liliya Tsirulnik, Zdenek Krnoul, Andrey Ronzhin, Boris Lobanov, Milos Zelezný
INTERSPEECH1
2006 Combined Gesture-Speech Analysis and Speech Driven Gesture Synthesis
abstract
Multimodal speech and speaker modeling and recognition are widely accepted as vital aspects of state of the art human-machine interaction systems. While correlations between speech and lip motion as well as speech and facial expressions are widely studied, relatively little work has been done to investigate the correlations between speech and gesture. Detection and modeling of head, hand and arm gestures of a speaker have been studied extensively and these gestures were shown to carry linguistic information. A typical example is the head gesture while saying "yes/no". In this study, correlation between gestures and speech is investigated. In speech signal analysis, keyword spotting and prosodic accent event detection has been performed. In gesture analysis, hand positions and parameters of global head motion are used as features. The detection of gestures is based on discrete pre-designated symbol sets, which are manually labeled during the training phase. The gesture-speech correlation is modeled by examining the co-occurring speech and gesture patterns. This correlation can be used to fuse gesture and speech modalities for edutainment applications (i.e. video games, 3-D animations) where natural gestures of talking avatars are animated from speech. A speech driven gesture animation example has been implemented for demonstration
Mehmet Emre Sargin, Oya Aran, Alexey Karpov 0001, Ferda Ofli, Yelena Yasinnik, Engin Erzin, Yücel Yemez, A. Murat Tekalp
ICME3
2006 Multi-modal system ICANDO: intellectual computer assistant for disabled operators
Alexey Karpov 0001, Andrey Ronzhin, Alexandre Cadiou
INTERSPEECH1