EDBT 2026 Demo / reviewers in the wild / expert
Athanasios Katsamanis
dblp:26/7051 · also Athanassios Katsamanis, Nassos Katsamanis
· DBLP profile ↗
73ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0002-2642-2354ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 55 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 40 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Medusa: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions
Georgios Chatzichristodoulou, Despoina Kosmopoulou, Antonios Kritikos, Anastasia Poulopoulou, Efthymios Georgiou, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos |
INTERSPEECH | 6 |
| 2024 | RobuSER: A Robustness Benchmark for Speech Emotion RecognitionabstractThe recent surge in deep learning has improved Speech Emotion Recognition (SER) model performance; however, ensuring robustness across diverse scenarios beyond the training dataset remains a problem. This challenge becomes pronounced in real-world situations characterized by noisy conditions, where model adaptability to unclean data is crucial. Despite ongoing efforts to develop noise-robust models, the lack of standardized evaluation protocols hampers fair comparisons among different models. This paper tackles this issue by introducing Robuser, a benchmarking procedure designed specifically for evaluating the robustness of SER models under noise. Robuser is a comprehensive open-source benchmark that can be applied to any speech dataset, focusing on diverse corruption types in two pivotal dimensions: additive background noise and various signal distortion corruptions, each in varying levels of severity. Furthermore, through the evaluation of a state-of-the-art SER model against this benchmark, we offer quantitative insights into the impact of the different corruption types and severity levels on performance. The baseline model reveals a notable performance degradation of up to 22.77% in Unweighted Accuracy (UA) and 20.32% in Weighted Accuracy (WA) on corrupted IEMOCAP, underscoring the substantial room for improvement in this domain. Our code is openly available at the following URL: https://github.com/BehavioralSignalTechnologies/ser_robustness.git Antonia Petrogianni, Lefteris Kapelonis, Nikolaos Antoniou, Sofia Eleftheriou, Petros Mitseas, Dimitris Sgouropoulos, Athanasios Katsamanis, Theodoros Giannakopoulos, Shri Narayanan |
ACII | 7 |
| 2024 | Emotion-Aware Speech Popularity Prediction: A Use-Case on TED TalksabstractIn the context of the ever-growing influence of social media, understanding and predicting the popularity of content has become crucial for creators and marketers alike. Our research addresses this need by introducing a method to forecast the success of oral presentations, focusing on the nuanced use of paralinguistic features and insights derived from speech emotion recognition models. This innovative approach is designed to enhance verbal communication skills by providing public speakers with targeted feedback. We leverage a dataset of 2,462 TED talk videos, complete with metadata such as user comments, tags, and views, to establish a set of four objective metrics for determining presentation popularity. These metrics form the foundation of our analysis, enabling us to evaluate the efficacy of our predictive methodology. By integrating audio-based emotional cues with text-based content analysis we showcase the capability of the proposed speech analytics system to capture user assessments of presentation quality. This research highlights the role of emotional expression in speech as a component of content's appeal, advocating for a broader analytical perspective beyond just text-only analysis. It suggests new directions for improving the impact of public speaking and calls for further investigation into multimodal content analysis, aiming to deepen our understanding of audience engagement on social media and content delivery platforms. Dimitris Sgouropoulos, Petros Mitseas, Sofia Eleftheriou, Theodoros Giannakopoulos, Antonia Petrogianni, Lefteris Kapelonis, Nikolaos Antoniou, Athanasios Katsamanis, Shri Narayanan |
ACII | 8 |
| 2024 | Beam-search SIEVE for low-memory speech recognitionabstractA capacity to recognize speech offline eliminates privacy concerns and the need for an internet connection. Despite efforts to reduce the memory demands of speech recognition systems, these demands remain formidable and thus popular tools such as Kaldi run best via cloud computing. The key bottleneck arises form the fact that a bedrock of such tools, the Viterbi algorithm, requires memory that grows linearly with utterance length even when contained via beam search. A recent recasting of the Viterbi algorithm, SIEVE, eliminates the path length factor from space complexity, but with a significant practical runtime overhead. In this paper, we develop a variant of SIEVE that lessens this runtime overhead via beam search, retains the decoding quality of standard beam search, and waives its linearly growing memory bottleneck. This space-complexity reduction is orthogonal to decoding quality and complementary to memory savings in model representation and training. Martino Ciaperoni, Athanasios Katsamanis, Aristides Gionis, Panagiotis Karras |
INTERSPEECH | 2 |
| 2024 | The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
Georgios Paraskevopoulos, Chara Tsoukala, Athanasios Katsamanis, Vassilis Katsouros |
INTERSPEECH | 3 |
| 2024 | Sample-Efficient Unsupervised Domain Adaptation of Speech Recognition Systems: A Case Study for Modern GreekabstractModern speech recognition systems exhibit rapid performance degradation under domain shift. This issue is especially prevalent in data-scarce settings, such as low-resource languages, where the diversity of training data is limited. In this work, we propose M2DS2, a simple and sample-efficient fine-tuning strategy for large pre-trained speech models, based on mixed source and target domain self-supervision. We find that including source domain self-supervision stabilizes training and avoids mode collapse of the latent representations. For evaluation, we collect HParl, a 120-hour speech corpus for Greek, consisting of plenary sessions in the Greek Parliament. We merge HParl with two popular Greek corpora to create GREC-MD, a test-bed for multi-domain evaluation of Greek ASR systems. In our experiments, we find that, while other Unsupervised Domain Adaptation baselines fail in this resource-constrained environment, M2DS2 yields significant improvements for cross-domain adaptation, even when only a few hours of in-domain audio are available. When we relax the problem in a weakly supervised setting, we find that independent adaptation for audio using M2DS2 and language using simple LM augmentation techniques is particularly effective, yielding word error rates comparable to the fully supervised baselines. Georgios Paraskevopoulos, Theodoros Kouzelis, Georgios Rouvalis, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Designing and Evaluating Speech Emotion Recognition Systems: A Reality Check Case Study with IEMOCAPabstractThere is an imminent need for guidelines and standard test sets to allow direct and fair comparisons of speech emotion recognition (SER). While resources, such as the Interactive Emotional Dyadic Motion Capture (IEMOCAP) database, have emerged as widely-adopted reference corpora for researchers to develop and test models for SER, published work reveals a wide range of assumptions and variety in its use that challenge reproducibility and generalization. Based on a critical review of the latest advances in SER using IEMOCAP as the use case, our work aims at two contributions: First, using an analysis of the recent literature, including assumptions made and metrics used therein, we provide a set of SER evaluation guidelines. Second, using recent publications with open-sourced implementations, we focus on reproducibility assessment in SER. Nikolaos Antoniou, Athanasios Katsamanis, Theodoros Giannakopoulos, Shri Narayanan |
ICASSP | 2 |
| 2023 | Exploring Language-Agnostic Speech Representations Using Domain Knowledge for Detecting Alzheimer's DementiaabstractWe explore ways to use speech data to screen for indications of Alzheimer’s dementia (AD). In particular, we describe our approach to the ICASSP 2023 Signal Processing Grand Challenge, which involves extrapolating from models learned from English speech samples, to Greek speech samples, to determine which subjects have AD. By using acoustic and linguistic features, inspired by clinical research on AD, our top-performing classification model achieves 69% accuracy in distinguishing AD patients from healthy controls, and our regression model attains an RMSE of 4.8 for inferring cognitive testing scores. These outcomes underscore the potential of our explainable model for detecting cognitive decline in AD patients via speech, and its applicability in clinical settings. Zehra Shah, Shiang Qi, Fei Wang 0062, Mahtab Farrokh, Mashrura Tasnim, Eleni Stroulia, Russell Greiner, Manos Plitsis, Athanasios Katsamanis |
ICASSP | 9 |
| 2023 | Weakly-supervised forced alignment of disfluent speech using phoneme-level modeling
Theodoros Kouzelis, Georgios Paraskevopoulos, Athanasios Katsamanis, Vassilis Katsouros |
INTERSPEECH | 3 |
| 2023 | Cross-Lingual Features for Alzheimer's Dementia Detection from Speech
Thomas Melistas, Lefteris Kapelonis, Nikolaos Antoniou, Petros Mitseas, Dimitris Sgouropoulos, Theodoros Giannakopoulos, Athanasios Katsamanis, Shri Narayanan |
INTERSPEECH | 7 |
| 2022 | Audio and ASR-based Filled Pause DetectionabstractFilled pauses (or fillers) are the most common form of speech disfluencies and they can be recognized as hesitation markers (“um”, “uh” and “er”) made by speakers, usually to gain extra time while thinking their next words. Filled pauses are very frequent in spontaneous speech. Their detection is therefore rather important for two basic reasons: (a) their existence influences the performance of individual components, like Automatic Speech Recognition system (ASR), in human-machine interaction and (b) their frequency can characterize the overall speech quality of a particular speaker, as it can be strongly associated with the speaker's confidence. Despite that, only limited work has been published for the detection of filled pauses in speech, especially through audio. In this work, we propose a framework for filled pause detection using both audio and textual information. For the audio modality, we transfer knowledge from a plethora of supervised tasks, such as emotion or speaking rate, using Convolutional Neural Networks (CNNs). For the text modality, we develop a temporal Recurrent Neural Network (RNN) method that takes into account textual information derived from an ASR system. In addition, the proposed transfer learning approach for the audio classifier leads to better results when benchmarked on our internal dataset for which the text is not transcribed but estimated by an ASR system. In this case, a simple late fusion approach boosts the performance even further. This proves that the audio approach is suitable for real-world applications where the transcribed text is not available and has to leverage imperfect ASR results, or even the absence of textual information (to reduce computational cost). Aggelina Chatziagapi, Dimitris Sgouropoulos, Constantinos Karouzos, Thomas Melistas, Theodoros Giannakopoulos, Athanasios Katsamanis, Shri Narayanan |
ACII | 6 |
| 2022 | Zero-Shot Cross-lingual Aphasia Detection using Automatic Speech Recognition
Gerasimos Chatzoudis, Manos Plitsis, Spyridoula Stamouli, Athanasia-Lida Dimou, Athanasios Katsamanis, Vassilis Katsouros |
INTERSPEECH | 5 |
| 2022 | SIEVE: A Space-Efficient Algorithm for Viterbi DecodingabstractCan we get speech recognition tools to work on limited-memory devices? The Viterbi algorithm is a classic dynamic programming (DP) solution used to find the most likely sequence of hidden states in a Hidden Markov Model (HMM). While the algorithm finds universal application ranging from communication systems to speech recognition to bioinformatics, its scalability has been scarcely addressed, stranding it to a space complexity that grows with the number of observations. Martino Ciaperoni, Aristides Gionis, Athanasios Katsamanis, Panagiotis Karras |
SIGMOD Conference | 3 |
| 2022 | Regotron: Regularizing the Tacotron2 Architecture Via Monotonic Alignment LossabstractDeep learning Text-to-Speech (TTS) systems have achieved impressive generated speech quality, close to human parity. However, they suffer from training stability issues and in-correct alignment between the intermediate acoustic representation and the text input. In this work, we propose Regotron, a regularized Tacotron2 version which alleviates the training issues by augmenting the objective function with an additional term, which penalizes non-monotonic alignments in the location-sensitive attention mechanism. By introducing this regularization term we demonstrate its effectiveness to stabilize the training process, produce a monotonic attention quicker (13% of the total number of epochs compared to Tacotron2) and reduce the alignment errors during inference. Moreover, Regotron has minimal additional computational overhead, reduces common TTS mistakes and at the same time achieves improved speech naturalness according to subjective mean opinion scores (MOS) collected from 50 evaluators. Efthymios Georgiou, Kosmas Kritsis, Georgios Paraskevopoulos, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos |
SLT | 4 |
| 2019 | Using Oliver API for emotion-aware movie content characterizationabstractThis paper demonstrates the utilization of Oliver11https://behavioralsignals.com/oliver/, the speech emotion recognition (SER) API created by Behavioral Signals, in the context of a movie content visualization application. Oliver API provides an emotion recognition as-a-service solution that can be accessed via a Web API. In this work, we demonstrate how one can send sound recordings from famous movies, retrieve respective emotional descriptors and use simple aggregations on these descriptors to visualize movie content. We have compiled a dataset of 60 movies, categorized over 8 directors. The classification examples included in this paper indicate the ability of simple emotion aggregations to discriminate between movie directors. In order for others to also experiment with the output of both the API's Emotional and Automatic Speech Recognition, the responses are provided as JSON files in this link: https://tinyurl.com/yxeqvvy2. Theodoros Giannakopoulos, Spiros Dimopoulos, Georgios Pantazopoulos, Aggelina Chatziagapi, Dimitris Sgouropoulos, Athanasios Katsamanis, Alexandros Potamianos, Shri Narayanan |
CBMI | 6 |
| 2019 | Data Augmentation Using GANs for Speech Emotion Recognition
Aggelina Chatziagapi, Georgios Paraskevopoulos, Dimitris Sgouropoulos, Georgios Pantazopoulos, Malvina Nikandrou, Theodoros Giannakopoulos, Athanasios Katsamanis, Alexandros Potamianos, Shri Narayanan |
INTERSPEECH | 7 |
| 2019 | A behaviorally inspired fusion approach for computational audiovisual saliency modeling
Antigoni Tsiami, Petros Koutras, Athanasios Katsamanis, Argiro Vatakis, Petros Maragos |
Signal Process. Image Commun. | 3 |
| 2018 | Multi-View Audio-Articulatory Features for Phonetic Recognition on RTMRI-TIMIT DatabaseabstractIn this paper, we investigate the use of articulatory information, and more specifically real time Magnetic Resonance Imaging (rtMRI) data of the vocal tract, to improve speech recognition performance. For the purpose of our experiments, we use data from the rtMRI-TIMIT database. Firstly, Scale Invariant Feature Transform (SIFT) features are extracted for each video frame. Afterwards, the SIFT descriptors of each frame are transformed to a single histogram per picture, by using the Bag of Visual Words methodology. Since this kind of articulatory information is difficult to acquire in typical speech recognition setups we only consider it to be available in the training phase. Thus, we use a multi-view setup approach by applying Canonical Correlation Analysis (CCA) to visual and audio data. By using the transformation matrix, acquired during the training stage, we transform both train and test audio data to produce MFCC-articulatory features, which form the input for the recognition system. Experimental results demonstrate improvements in phone recognition in comparison with the audio-based baseline. Ioannis K. Douros, Athanasios Katsamanis, Petros Maragos |
ICASSP | 2 |
| 2017 | Photorealistic adaptation and interpolation of facial expressions using HMMS and AAMS for audio-visual speech synthesisabstractIn this paper, motivated by the continuously increasing presence of intelligent agents in everyday life, we address the problem of expressive photorealistic audio-visual speech synthesis, with a strong focus on the visual modality. Emotion constitutes one of the main driving factors of social life and it is expressed mainly through facial expressions. Synthesis of a talking head capable of expressive audio-visual speech is challenging due to the data overhead that arises when considering the vast number of emotions we would like the talking head to express. In order to tackle this challenge, we propose the usage of two methods, namely Hidden Markov Model (HMM) adaptation and interpolation, with HMMs modeling visual parameters via an Active Appearance Model (AAM) of the face. We show that through HMM adaptation we can successfully adapt a “neutral” talking head to a target emotion with a small amount of adaptation data, as well as that through HMM interpolation we can robustly achieve different levels of intensity for an emotion. Panagiotis Paraskevas Filntisis, Athanasios Katsamanis, Petros Maragos |
ICIP | 2 |
| 2017 | Demonstration of an HMM-based photorealistic expressive audio-visual speech synthesis systemabstractSummary form only given. The usage of conversational agents is rapidly increasing in everyday life (cortana, siri, etc.). It has been shown that the inclusion of a talking face, increases the intelligibility of speech and the naturalness of human-computer interaction. Furthermore, an agent capable of expressing emotions has a stronger appeal to the human party and affects the interlocutor's emotional state. The proposed demonstration is a Hidden Markov Model (HMM) based photorealistic audio-visual speech synthesis system, capable of expressing emotions [1, 2]. The system is capable of generating a talking head speaking in three emotions: happiness, anger, and sadness, plus in neutral speaking style. Further capabilities of the system include 1) the usage of HMM interpolation [3] in order to generate speech with mixtures of the original emotions (e.g., both anger and happiness), and speech with different levels of expressiveness (by mixing with the neutral emotion), 2) the usage of HMM adaptation [4], in order to adapt to a target emotion using only a few number of sentences. Equipment In order to showcase our system we will use a laptop and speakers. The system will run fully on the laptop. Demonstration Experience During the demonstration, viewers will have the opportunity to: 1. Watch videos of the talking head speaking in 3 different emotions (plus neutral) and see how the expressive talking head feels more natural compared to the talking head speaking in neutral style. 2. Watch the talking head speaking in two or more emotions at the same time, and see how the weights assigned to each emotion affects the outcome. It will also be of great interest to see which emotion each viewer perceives. In addition, through interpolation with the neutral emotion, viewers will be able to watch the talking head speak in different expressiveness levels for each emotion. 3. See how the neutral talking head can be adapted to speak in another emotion using only a few sentences, and how the number of sentences used affects the expressiveness of the resulting talking head. Panagiotis Paraskevas Filntisis, Athanasios Katsamanis, Petros Maragos |
ICIP | 2 |
| 2017 | Room-localized spoken command recognition in multi-room, multi-microphone environments
Isidoros Rodomagoulakis, Athanasios Katsamanis, Gerasimos Potamianos, Panagiotis Giannoulis, Antigoni Tsiami, Petros Maragos |
Comput. Speech Lang. | 2 |
| 2017 | Video-realistic expressive audio-visual speech synthesis for the Greek language
Panagiotis Paraskevas Filntisis, Athanasios Katsamanis, Pirros Tsiakoulis, Petros Maragos |
Speech Commun. | 2 |
| 2017 | Multiple Instance Learning for Behavioral CodingabstractWe propose a computational methodology for automatically estimating human behavioral patterns using the multiple instance learning (MIL) paradigm. We describe the incremental diverse density algorithm, a particular formulation of multiple instance learning, and discuss its suitability for behavioral coding. We use a rich multi-modal corpus comprised of chronically distressed married couples having problem-solving discussions as a case study to experimentally evaluate our approach. In the multiple instance learning framework, we treat each discussion as a collection of short-term behavioral expressions which are manifested in the acoustic, lexical, and visual channels. We experimentally demonstrate that this approach successfully learns representations that carry relevant information about the behavioral coding task. Furthermore, we employ this methodology to gain novel insights into human behavioral data, such as the local versus global nature of behavioral constructs as well as the level of ambiguity in the expression of behaviors through each respective modality. Finally, we assess the success of each modality for behavioral classification and compare schemes for multimodal fusion within the proposed framework. James Gibson, Athanasios Katsamanis, Francisco Romero, Bo Xiao 0003, Panayiotis G. Georgiou, Shri Narayanan |
IEEE Trans. Affect. Comput. | 2 |
| 2016 | Multimodal human action recognition in assistive human-robot interactionabstractWithin the context of assistive robotics we develop an intelligent interface that provides multimodal sensory processing capabilities for human action recognition. Human action is considered in multimodal terms, containing inputs such as audio from microphone arrays, and visual inputs from high definition and depth cameras. Exploring state-of-the-art approaches from automatic speech recognition, and visual action recognition, we multimodally recognize actions and commands. By fusing the unimodal information streams, we obtain the optimum multimodal hypothesis which is to be further exploited by the active mobility assistance robot in the framework of the MOBOT EU research project. Evidence from recognition experiments shows that by integrating multiple sensors and modalities, we increase multimodal recognition performance in the newly acquired challenging dataset involving elderly people while interacting with the assistive robot. Isidoros Rodomagoulakis, Nikolaos Kardaris, Vassilis Pitsikalis, E. Mavroudi, Athanasios Katsamanis, Antigoni Tsiami, Petros Maragos |
ICASSP | 5 |
| 2016 | Towards a behaviorally-validated computational audiovisual saliency modelabstractComputational saliency models aim at predicting, in a bottom-up fashion, where human attention is drawn in the presented (visual, auditory or audiovisual) scene and have been proven useful in applications like robotic navigation, image compression and movie summarization. Despite the fact that well-established auditory and visual saliency models have been validated in behavioral experiments, e.g., by means of eye-tracking, there is no established computational audiovisual saliency model validated in the same way. In this work, building on biologically-inspired models of visual and auditory saliency, we present a joint audiovisual saliency model and introduce the validation approach we follow to show that it is compatible with recent findings of psychology and neuroscience regarding multimodal integration and attention. In this direction, we initially focus on the "pip and pop" effect which has been observed in behavioral experiments and indicates that visual search in sequences of cluttered images can be significantly aided by properly timed non-spatial auditory signals presented alongside the target visual stimuli. Antigoni Tsiami, Athanasios Katsamanis, Petros Maragos, Argiro Vatakis |
ICASSP | 2 |
| 2016 | FMRI-based perceptual validation of a computational model for visual and auditory saliency in videosabstractIn this study, we make use of brain activation data to investigate the perceptual plausibility of a visual and an auditory model for visual and auditory saliency in video processing. These models have already been successfully employed in a number of applications. In addition, we experiment with parameters, modifications and suitable fusion schemes. As part of this work, fMRI data from complex video stimuli were collected, on which we base our analysis and results. The core part of the analysis involves the use of well-established methods for the manipulation of fMRI data and the examination of variability across brain responses of different individuals. Our results indicate a success in confirming the value of these saliency models in terms of perceptual plausibility. Georgia Panagiotaropoulou, Petros Koutras, Athanasios Katsamanis, Petros Maragos, Athanasia Zlatintsi, Athanassios Protopapas, Eustratios Karavasilis, Nikolaos Smyrnis |
ICIP | 3 |
| 2016 | A Phase-Based Time-Frequency Masking for Multi-Channel Speech Enhancement in Domestic Environments
Alessio Brutti, Antigoni Tsiami, Athanasios Katsamanis, Petros Maragos |
INTERSPEECH | 3 |
| 2015 | Context-sensitive learning for enhanced audiovisual emotion classification (Extended abstract)abstractHuman emotional expression tends to evolve in a structured manner in the sense that certain emotional evolution patterns, i.e., anger to anger, are more probable than others, e.g., anger to happiness. Furthermore the perception of an emotional display can be affected by recent emotional displays. Therefore, the emotional content of past and future observations could offer relevant temporal context when classifying the emotional content of an observation. In this work, we focus on audio-visual recognition of the emotional content of improvised emotional interactions at the utterance level. We examine context-sensitive schemes for emotion recognition within a multimodal, hierarchical approach: bidirectional Long Short-Term Memory (BLSTM) neural networks, hierarchical Hidden Markov Model classifiers (HMMs) and hybrid HMM/BLSTM classifiers are considered for modeling emotion evolution within an utterance and between utterances over the course of a dialog. Overall, our experimental results indicate that incorporating long-term temporal context is beneficial for emotion recognition systems that encounter a variety of emotional manifestations. Angeliki Metallinou, Athanasios Katsamanis, Martin Wöllmer, Florian Eyben, Björn W. Schuller, Shri Narayanan |
ACII | 2 |
| 2015 | Predicting audio-visual salient events based on visual, audio and text modalities for movie summarizationabstractIn this paper, we present a new and improved synergistic approach to the problem of audio-visual salient event detection and movie summarization based on visual, audio and text modalities. Spatio-temporal visual saliency is estimated through a perceptually inspired frontend based on 3D (space, time) Gabor filters and frame-wise features are extracted from the saliency volumes. For the auditory salient event detection we extract features based on Teager-Kaiser Energy Operator, while text analysis incorporates part-of-speech tagging and affective modeling of single words on the movie subtitles. For the evaluation of the proposed system, we employ an elementary and non-parametric classification technique like KNN. Detection results are reported on the MovSum database, using objective evaluations against ground-truth denoting the perceptually salient events, and human evaluations of the movie summaries. Our evaluation verifies the appropriateness of the proposed methods compared to our baseline system. Finally, our newly proposed summarization algorithm produces summaries that consist of salient and meaningful events, also improving the comprehension of the semantics. Petros Koutras, Athanasia Zlatintsi, Elias Iosif, Athanasios Katsamanis, Petros Maragos, Alexandros Potamianos |
ICIP | 4 |
| 2015 | Multimodal gesture recognition via multiple hypotheses rescoring
Vassilis Pitsikalis, Athanasios Katsamanis, Stavros Theodorakis, Petros Maragos |
J. Mach. Learn. Res. | 2 |
| 2014 | Robust far-field spoken command recognition for home automation combining adaptation and multichannel processingabstractThe paper presents our approach to speech-controlled home automation. We are focusing on the detection and recognition of spoken commands preceded by a key-phrase as recorded in a voice-enabled apartment by a set of multiple microphones installed in the rooms. For both problems we investigate robust modeling, environmental adaptation and multichannel processing to cope with a) insufficient training data and b) the far-field effects and noise in the apartment. The proposed integrated scheme is evaluated in a challenging and highly realistic corpus of simulated audio recordings and achieves F-measure close to 0.70 for key-phrase spotting and word accuracy close to 98% for the command recognition task. Athanasios Katsamanis, Isidoros Rodomagoulakis, Gerasimos Potamianos, Petros Maragos, Antigoni Tsiami |
ICASSP | 1 |
| 2014 | Kinect-based multimodal gesture recognition using a two-pass fusion schemeabstractWe present a new framework for multimodal gesture recognition that is based on a two-pass fusion scheme. In this, we deal with a demanding Kinect-based multimodal dataset, which was introduced in a recent gesture recognition challenge. We employ multiple modalities, i.e., visual cues, such as colour and depth images, as well as audio, and we specifically extract feature descriptors of the hands' movement, handshape, and audio spectral properties. Based on these features, we statistically train separate unimodal gesture-word models, namely hidden Markov models, explicitly accounting for the dynamics of each modality. Multimodal recognition of unknown gesture sequences is achieved by combining these models in a late, two-pass fusion scheme that exploits a set of unimodally generated n-best recognition hypotheses. The proposed scheme achieves 88.2% gesture recognition accuracy in the Kinect-based multimodal dataset, outperforming all recently published approaches on the same challenging multimodal gesture recognition task. Georgios Pavlakos, Stavros Theodorakis, Vassilis Pitsikalis, Athanasios Katsamanis, Petros Maragos |
ICIP | 4 |
| 2014 | The DIRHA-GRID corpus: baseline and tools for multi-room distant speech recognition using distributed microphonesabstractDistant speech recognition in real-world environments is still a challenging problem and a particularly interesting topic is the investigation of multi-channel processing in case of distributed microphones in home environments. This paper presents an initiative oriented to address the challenges of such a scenario; an experimental recognition framework comprising a multi-room, multi-channel corpus and the accompanying evaluation tools is made publicly available. The overall goal is to represent a common platform for comparing state-of-the-art algorithms, share ideas of different research communities and integrate several components in a realistic distant-talking recognition chain, e.g., voice activity detection, speech/feature enhancement, channel selection and fusion, model Marco Matassoni, Ramón Fernandez Astudillo, Athanasios Katsamanis, Mirco Ravanelli |
INTERSPEECH | 3 |
| 2014 | ATHENA: a Greek multi-sensory database for home automation control uthor: isidoros rodomagoulakis (NTUA, Greece)
Antigoni Tsiami, Isidoros Rodomagoulakis, Panagiotis Giannoulis, Athanasios Katsamanis, Gerasimos Potamianos, Petros Maragos |
INTERSPEECH | 4 |
| 2014 | Computing vocal entrainment: A signal-derived PCA-based quantification scheme with application to affect analysis in married couple interactions
Chi-Chun Lee, Athanasios Katsamanis, Matthew Black, Brian R. Baucom, Andrew Christensen, Panayiotis G. Georgiou, Shri Narayanan |
Comput. Speech Lang. | 2 |
| 2013 | Multi-band long-term signal variability features for robust voice activity detectionabstractIn this paper, we propose robust features for the problem of voice activity detection (VAD). In particular, we extend the long term signal variability (LTSV) feature to accommodate multiple spectral bands. The motivation of the multi-band approach stems from the non-uniform frequency scale of speech phonemes and noise characteristics. Our analysis shows that the multi-band approach offers advantages over the single band LTSV for voice activity detection. In terms of classification accuracy, we show 0.3%-61.2% relative improvement over the best accuracy of the baselines considered for 7 out 8 different noisy channels. Experimental results, and error analysis, are reported on the DARPA RATS corpora of noisy speech. Index Terms: noisy speech data, voice activity detection, robust feature extraction Andreas Tsiartas, Theodora Chaspari, Athanasios Katsamanis, Prasanta Kumar Ghosh, Ming Li 0026, Maarten Van Segbroeck, Alexandros Potamianos, Shri Narayanan |
INTERSPEECH | 3 |
| 2013 | Tracking continuous emotional trends of participants during affective dyadic interactions using body language and speech information
Angeliki Metallinou, Athanasios Katsamanis, Shri Narayanan |
Image Vis. Comput. | 2 |
| 2013 | Toward automating a human behavioral coding system for married couples' interactions using speech acoustic features
Matthew Black, Athanasios Katsamanis, Brian R. Baucom, Chi-Chun Lee, Adam C. Lammert, Andrew Christensen, Panayiotis G. Georgiou, Shri Narayanan |
Speech Commun. | 2 |
| 2012 | An acoustic analysis of shared enjoyment in ECA interactions of children with autismabstractThe quality of shared enjoyment in interactions is a key aspect related to Autism Spectrum Disorders (ASD). This paper discusses two types of enjoyment: the first refers to humorous events and is associated with one's positive affective state and the second is used to facilitate social interactions between people. These types of shared enjoyment are objectively specified by their proximity to a voiced and unvoiced laughter instance, respectively. The goal of this work is to study the acoustic differences of areas surrounding the two kinds of shared enjoyment instances, called “social zones”, using data collected from children with autism, and their parents, interacting with an Embodied Conversational Agent (ECA). A classification task was performed to predict whether a “social zone” surrounds a voiced or an unvoiced laughter instance. Our results indicate that humorous events are more easily recognized than events acting as social facilitators and that related speech patterns vary more across children compared to other interlocutors. Theodora Chaspari, Emily Mower Provost, Athanasios Katsamanis, Shri Narayanan |
ICASSP | 3 |
| 2012 | A hierarchical framework for modeling multimodality and emotional evolution in affective dialogsabstractIncorporating multimodal information and temporal context from speakers during an emotional dialog can contribute to improving performance of automatic emotion recognition systems. Motivated by these issues, we propose a hierarchical framework which models emotional evolution within and between emotional utterances, i.e., at the utterance and dialog level respectively. Our approach can incorporate a variety of generative or discriminative classifiers at each level and provides flexibility and extensibility in terms of multimodal fusion; facial, vocal, head and hand movement cues can be included and fused according to the modality and the emotion classification task. Our results using the multimodal, multi-speaker IEMOCAP database indicate that this framework is well-suited for cases where emotions are expressed multimodally and in context, as in many real-life situations. Angeliki Metallinou, Athanasios Katsamanis, Shri Narayanan |
ICASSP | 2 |
| 2012 | Analyzing the memory of BLSTM Neural Networks for enhanced emotion classification in dyadic spoken interactionsabstractRecent studies indicate that bidirectional Long Short-Term Memory (BLSTM) recurrent neural networks are well-suited for automatic emotion recognition systems and may lead to better results than systems applying other widely used classifiers such as Support Vector Machines or feedforward Neural Networks. The good performance of BLSTM emotion recognition systems could be attributed to their ability to model and exploit contextual information self-learned via recurrently connected memory blocks which allows them to incorporate information about how emotion evolves over time. However, the actual amount of bidirectional context that a BLSTM classifier takes into account when classifying an observation has not been investigated so far. This paper presents a methodology to systematically investigate the number of past and future utterance-level observations that are considered to generate an emotion prediction for a given utterance, and to examine to what extent this temporal bidirectional context contributes to the overall BLSTM performance. Martin Wöllmer, Angeliki Metallinou, Athanasios Katsamanis, Björn W. Schuller, Shri Narayanan |
ICASSP | 3 |
| 2012 | Based on Isolated Saliency or Causal Integration? Toward a Better Understanding of Human Annotation Process using Multiple Instance Learning and Sequential Probability Ratio TestabstractHuman perception is capable of integrating local events to generate an overall impression at the global level; this is evident in daily life and is utilized repeatedly in behavioral science studies to bring objective measures into studies of human behavior. In this work, we explore two hypotheses considering whether it is the isolated-saliency or the causal-integration of information that can trigger the global perceptual behavioral ratings as trained annotators engage in tasks of observational coding. We carry out analyses using Multiple Instance Learning and Sequential Probability Ratio Test in a corpus of real and spontaneous distressed couples’ interaction with global sessionlevel abstract behavioral coding done by trained human annotators. We present various analyses based on different behavioral detection schemes demonstrating the potential of utilizing these algorithms in bringing insights into the human annotation process. We further show that while annotating behaviors with more positive impression, annotators gather information throughout the session compared to behaviors with more negative impression, where a single salient instance is enough to trigger the final global decision. Chi-Chun Lee, Athanasios Katsamanis, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 2 |
| 2012 | Ada and Grace: Direct Interaction with Museum Visitors
David R. Traum, Priti Aggarwal, Ron Artstein, Susan Foutz, Jillian Gerten, Athanasios Katsamanis, Anton Leuski, Dan Noren, William R. Swartout |
IVA | 6 |
| 2012 | The Twins Corpus of Museum Visitor Questions
Priti Aggarwal, Ron Artstein, Jillian Gerten, Athanasios Katsamanis, Shri Narayanan, Angela Nazarian, David R. Traum |
LREC | 4 |
| 2012 | Context-Sensitive Learning for Enhanced Audiovisual Emotion ClassificationabstractHuman emotional expression tends to evolve in a structured manner in the sense that certain emotional evolution patterns, i.e., anger to anger, are more probable than others, e.g., anger to happiness. Furthermore, the perception of an emotional display can be affected by recent emotional displays. Therefore, the emotional content of past and future observations could offer relevant temporal context when classifying the emotional content of an observation. In this work, we focus on audio-visual recognition of the emotional content of improvised emotional interactions at the utterance level. We examine context-sensitive schemes for emotion recognition within a multimodal, hierarchical approach: bidirectional Long Short-Term Memory (BLSTM) neural networks, hierarchical Hidden Markov Model classifiers (HMMs), and hybrid HMM/BLSTM classifiers are considered for modeling emotion evolution within an utterance and between utterances over the course of a dialog. Overall, our experimental results indicate that incorporating long-term temporal context is beneficial for emotion recognition systems that encounter a variety of emotional manifestations. Context-sensitive approaches outperform those without context for classification tasks such as discrimination between valence levels or between clusters in the valence-activation space. The analysis of emotional transitions in our database sheds light into the flow of affective expressions, revealing potentially useful patterns. Angeliki Metallinou, Martin Wöllmer, Athanasios Katsamanis, Florian Eyben, Björn W. Schuller, Shri Narayanan |
IEEE Trans. Affect. Comput. | 3 |
| 2011 | Multiple Instance Learning for Classification of Human Behavior Observations
Athanasios Katsamanis, James Gibson, Matthew Black, Shri Narayanan |
ACII (1) | 1 |
| 2011 | Affective State Recognition in Married Couples' Interactions Using PCA-Based Vocal Entrainment Measures with Multiple Instance Learning
Chi-Chun Lee, Athanasios Katsamanis, Matthew Black, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan |
ACII (2) | 2 |
| 2011 | Tracking changes in continuous emotion states using body language and prosodic cuesabstractHuman expressive interactions are characterized by an ongoing unfolding of verbal and nonverbal cues. Such cues convey the interlocutor's emotional state which is continuous and of variable intensity and clarity over time. In this paper, we examine the emotional content of body language cues describing a participant's posture, relative position and approach/withdraw behaviors during improvised affective interactions, and show that they reflect changes in the participant's activation and dominance levels. Furthermore, we describe a framework for tracking changes in emotional states during an interaction using a statistical mapping between the observed audiovisual cues and the underlying user state. Our approach shows promising results for tracking changes in activation and dominance. Angeliki Metallinou, Athanasios Katsamanis, Shri Narayanan |
ICASSP | 2 |
| 2011 | Estimation of ordinal approach-avoidance labels in dyadic interactions: Ordinal logistic regression approachabstractBehavioral Signal Processing aims at automating behavioral coding schemes such as those prevalent in psychology and mental health research. This paper describes a method to automatically quantify the approach-and-avoidance (AA) behavior, described by ordinal labels manually assigned by experts using either video-only or video-with-audio. We propose a novel ordinal regression (OR) algorithm and its hidden Markov model (HMM) extension for estimation of AA labels from visual motion capture based and acoustic features. The proposed algorithm transforms the OR to multiple binary classification problems, solves them by independent score-outputting classifiers and fits the cumulative logit logistic regression model with proportional odds (CLLRMP) to vectors of the classifier scores. The time series extension treats labels as states of the HMM with a likelihood function derived from the probabilistic CLLRMP output. We compare performances of the proposed algorithm applying the weighted binary SVMs in the second step (SVM-OLR), its time-series extension (HMM-SVM-OLR) and the baseline multi-class SVM. On the used dyadic interaction dataset the HMM-SVM-OLR achieves the highest estimation accuracies 71.6 % and 65.7 % for AA labels assigned respectively using video-only and video-with-audio. Viktor Rozgic, Bo Xiao 0003, Athanasios Katsamanis, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 3 |
| 2011 | "You made me do it": Classification of Blame in Married Couples' Interactions by Fusing Automatically Derived Speech and Language InformationabstractOne of the goals of behavioral signal processing is the automatic prediction of relevant high-level human behaviors from complex, realistic interactions. In this work, we analyze dyadic discussions of married couples and try to classify extreme instances (low/high) of blame expressed from one spouse to another. Since blame can be conveyed through various communicative channels (e.g., speech, language, gestures), we compare two different classification methods in this paper. The first classifier is trained with the conventional static acoustic features and models “how” the spouses spoke. The second is a novel automatic speech recognition-derived classifier, which models “what” the spouses said. We get the best classification performance (82% accuracy) by exploiting the complementarity of these acoustic and lexical information sources through scorelevel fusion of the two classification methods. Index Terms: behavioral signal processing (BSP), couple therapy, blame, acoustic features, lexical features, fusion Matthew Black, Panayiotis G. Georgiou, Athanasios Katsamanis, Brian R. Baucom, Shri Narayanan |
INTERSPEECH | 3 |
| 2011 | Automatic Identification of Salient Acoustic Instances in Couples' Behavioral Interactions Using Diverse Density Support Vector MachinesabstractBehavioral coding focuses on deriving higher-level behavioral annotations using observational data of human interactions. Automatically identifying salient events in the observed signal data could lead to a deeper understanding of how specific events in an interaction correspond to the perceived high-level behaviors of the subjects. In this paper, we analyze a corpus of married couples’ interactions, in which a number of relevant behaviors, e.g., level of acceptance, were manually coded at the sessionlevel. We propose a multiple instance learning approach called Diverse Density Support Vector Machines, trained with acoustic features, to classify extreme cases of these behaviors, e.g., low acceptance vs. high acceptance. This method has the benefit of identifying salient behavioral events within the interactions, which is demonstrated by comparable classification performance to traditional SVMs while using only a subset of the events from the interactions for classification. Index Terms: behavioral signal processing, multiple instance learning, diverse density, support vector machines James Gibson, Athanasios Katsamanis, Matthew Black, Shri Narayanan |
INTERSPEECH | 2 |
| 2011 | Validating rt-MRI Based Articulatory Representations via Articulatory RecognitionabstractThe large corpus of real time magnetic resonance image sequences of the vocal tract during speech production that was recently acquired and can be referred to as MRI-TIMIT, provides us with a unique platform for systematically studying articulatory dynamics. Compared to previously collected articulatory datasets, e.g., using articulography or X-rays, MRI-TIMIT is a rich source of information for the entire vocal tract and not only for certain articulatory landmarks and further has the potential to continue increasing in size covering a large variety of speakers and speaking styles. In this work, we investigate an articulatory representation based on full vocal tract shapes. We employ an articulatory recognition framework in MRI-TIMIT to analyze its merits and drawbacks. We argue that articulatory recognition can serve as a general validation tool for real-time MRI based articulatory representations. Index Terms: vocal tract shape, articulation, real-time MRI, articulatory recognition Athanasios Katsamanis, Erik Bresch, Vikram Ramanarayanan, Shri Narayanan |
INTERSPEECH | 1 |
| 2011 | Morphological Variation in the Adult Vocal Tract: A Modeling Study of its Potential Acoustic ImpactabstractIn order to fully understand inter-speaker variability in the acoustical and articulatory domains, morphological variability must be considered, as well. Human vocal tracts display substantial morphological differences, all of which have the potential to impact a speaker’s acoustic output. The palate and rear pharyngeal wall, in particular, vary widely and have the potential to strongly impact the resonant properties of the vocal tract. To gain a better understanding of this impact, we combine an examination of morphological variation with acoustic modeling experiments. The goal is to show the theoretical acoustic effect of common inter-speaker differences for a set of English vowels. Modeling results indicate that the effect is indeed strong, but also surprisingly complex and context-specific, even when morphology varies in relatively straightforward ways. Index Terms: speech production, vocal tract morphology, inter-speaker variability, acoustic modeling, speaker modeling Adam C. Lammert, Michael I. Proctor, Athanasios Katsamanis, Shri Narayanan |
INTERSPEECH | 3 |
| 2011 | An Analysis of PCA-Based Vocal Entrainment Measures in Married Couples' Affective Spoken InteractionsabstractEntrainment has played a crucial role in analyzing marital couples interactions. In this work, we introduce a novel technique for quantifying vocal entrainment based on Principal Component Analysis (PCA). The entrainment measure, as we define in this work, is the amount of preserved variability of one interlocutor’s speaking characteristic when projected onto representing space of the other’s speaking characteristics. Our analysis on real couples interactions shows that when a spouse is rated as having positive emotion, he/she has a higher value of vocal entrainment compared when rated as having negative emotion. We further performed various statistical analyses on the strength and the directionality of vocal entrainment under different affective interaction conditions to bring quantitative insights into the entrainment phenomenon. These analyses along with a baseline prediction model demonstrate the validity and utility of the proposed PCA-based vocal entrainment measure. Index Terms: vocal entrainment, couples therapy, behavioral signal processing, principal component analysis Chi-Chun Lee, Athanasios Katsamanis, Matthew Black, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 2 |
| 2011 | A Multimodal Real-Time MRI Articulatory Corpus for Speech ResearchabstractWe present MRI-TIMIT: a large-scale database of synchronized audio and real-time magnetic resonance imaging (rtMRI) data for speech research. The database currently consists of speech data acquired from two male and two female speakers of Amer-ican English. Subjects ’ upper airways were imaged in the mid-sagittal plane while reading the same 460 sentence corpus used in the MOCHA-TIMIT corpus [1]. Accompanying acoustic recordings were phonemically transcribed using forced align-ment. Vocal tract tissue boundaries were automatically identi-fied in each video frame, allowing for dynamic quantification of each speaker’s midsagittal articulation. The database and com-panion toolset provide a unique resource with which to examine articulatory-acoustic relationships in speech production. Index Terms: speech production, speech corpora, real-time MRI, multi-modal database, large-scale phonetic tools Shri Narayanan, Erik Bresch, Prasanta Kumar Ghosh, Louis Goldstein, Athanasios Katsamanis, Adam C. Lammert, Michael I. Proctor, Vikram Ramanarayanan, Yinghua Zhu |
INTERSPEECH | 5 |
| 2011 | Direct Estimation of Articulatory Kinematics from Real-Time Magnetic Resonance Image SequencesabstractA method of rapid, automatic extraction of consonantal artic-ulatory trajectories from real-time magnetic resonance image sequences is described. Constriction location targets are esti-mated by identifying regions of maximally-dynamic correlated pixel activity along the palate, the alveolar ridge, and at the lips. Tissue movement into and out of the constriction location is es-timated by calculating the change in mean pixel intensity in a circle located at the center of the region of interest. Closure and release gesture timings are estimated from landmarks in the ve-locity profile derived from the smoothed intensity function. We demonstrate the utility of the technique in the analysis of Italian intervocalic consonant production. Index Terms: speech production, real-time MRI, consonant ar-ticulation, tongue shaping, articulatory phonology Michael I. Proctor, Adam C. Lammert, Athanasios Katsamanis, Louis Goldstein, Christina Hagedorn, Shri Narayanan |
INTERSPEECH | 3 |
| 2011 | Automatic Data-Driven Learning of Articulatory Primitives from Real-Time MRI Data Using Convolutive NMF with Sparseness ConstraintsabstractWe present a procedure to automatically derive inter-pretable dynamic articulatory primitives in a data-driven man-ner from image sequences acquired through real-time magnetic resonance imaging (rt-MRI). More specifically, we propose a convolutive Nonnegative Matrix Factorization algorithm with sparseness constraints (cNMFsc) to decompose a given set of image sequences into a set of basis image sequences and an acti-vation matrix. We use a recently-acquired rt-MRI corpus of read speech (460 sentences from 4 speakers) as a test dataset for this procedure. We choose the free parameters of the algorithm em-pirically by analyzing algorithm performance for different pa-rameter values. We then validate the extracted basis sequences using an articulatory recognition task and finally present an in-terpretation of the extracted basis set of image sequences in a gesture-based Articulatory Phonology framework. Vikram Ramanarayanan, Athanasios Katsamanis, Shri Narayanan |
INTERSPEECH | 2 |
| 2011 | Acoustic and Visual Cues of Turn-Taking Dynamics in Dyadic InteractionsabstractIn this paper we introduce an empirical study of multimodal cues of turn-taking dynamics in a social interaction context. We first identify pauses, gaps and overlapped speech segments in the dyadic conversation dataset. Second, we define two types of measurements, Mean Equalized Energy (MEE) and Animation Level (AL) on the audio and video channels, respectively. Then, we verify the hypothesis that the speaker with higher MEE or AL is more likely to take the floor after silence or overlapped speech. The results suggest that both the vocal and visual movement energy offer useful cues towards inferring the intention of the interlocutor to grab the floor. Index Terms: turn-taking, cues, equalized energy, motion vector Bo Xiao 0003, Viktor Rozgic, Athanasios Katsamanis, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 3 |
| 2010 | Automatic classification of married couples' behavior using audio featuresabstractIn this work, we analyzed a 96-hour corpus of married couples spontaneously interacting about a problem in their relationship. Each spouse was manually coded with relevant session-level perceptual observations (e.g., level of blame toward other spouse, global positive affect), and our goal was to classify the spouses’ behavior using features derived from the audio signal. Based on automatic segmentation, we extracted prosodic/spectral features to capture global acoustic properties for each spouse. We then trained gender-specific classifiers to predict the behavior of each spouse for six codes. We compare performance for the various factors (across codes, gender, classifier type, and feature type) and discuss future work for this novel and challenging corpus. Index Terms: behavioral signal processing, human behavior analysis, couples therapy, prosody, emotion recognition Matthew Black, Athanasios Katsamanis, Chi-Chun Lee, Adam C. Lammert, Brian R. Baucom, Andrew Christensen, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 2 |
| 2010 | Statistical multi-stream modeling of real-time MRI articulatory speech dataabstractThis paper investigates different statistical modeling frameworks for articulatory speech data obtained using real-time (RT) magnetic resonance imaging (MRI). To quantitatively capture the spatio-temporal shaping process of the human vocal tract during speech production a multi-dimensional stream of direct image features is extracted automatically from the MRI recordings. The features are closely related, though not identical, to the tract variables commonly defined in the articulatory phonology theory. The modeling of the shaping process aims at decomposing the articulatory data streams into primitives by segmentation. A variety of approaches are investigated for carrying out the segmentation task including vector quantizers, Gaussian Mixture Models, Hidden Markov Models, and a coupled Hidden Markov Model. We evaluate the performance of the different segmentation schemes qualitatively with the help of a well understood data set which was used in an earlier study of inter-articulatory timing phenomena of American English nasal sounds. Index Terms: speech production, articulatory modeling, realtime magnetic resonance imaging Erik Bresch, Athanasios Katsamanis, Louis Goldstein, Shri Narayanan |
INTERSPEECH | 2 |
| 2010 | Quantification of prosodic entrainment in affective spontaneous spoken interactions of married couplesabstractInteraction synchrony among interlocutors happens naturally as people adapt their speaking style gradually to promote efficient communication. In this work, we quantify one aspect of interaction synchrony prosodic entrainment, specifically pitch and energy, in married couples’ problem-solving interactions using speech signal-derived measures. Statistical testings demonstrate that some of these measures capture useful information; they show higher values in interactions with couples having high positive attitude compared to high negative attitude. Further, by using quantized entrainment measures employed with statistical symbol sequence matching in a maximum likelihood framework, we obtained 76% accuracy in predicting positive affect vs. negative affect. Chi-Chun Lee, Matthew Black, Athanasios Katsamanis, Adam C. Lammert, Brian R. Baucom, Andrew Christensen, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 3 |
| 2010 | Rapid semi-automatic segmentation of real-time magnetic resonance images for parametric vocal tract analysisabstractA method of rapid semi-automatic segmentation of real-time magnetic resonance image data for parametric analysis of vo-cal tract shaping is described. Tissue boundaries are identified by seeking pixel intensity thresholds along tract-normal grid-lines. Airway contours are constrained with respect to a tract centerline defined as an optimal path over the graph of all in-tensity minima between the glottis and lips. The method allows for superimposition of reference boundaries to guide automatic segmentation of anatomical features which are poorly imaged using magnetic resonance – dentition and the hard palate – re-sulting in more accurate sagittal sections than those produced by fully automatic segmentation. We demonstrate the utility of the technique in the dynamic analysis of tongue shaping in Tamil liquid consonants. Index Terms: speech production, vocal tract segmentation, MRI, tongue shaping, articulatory analysis Michael I. Proctor, Daniel Bone, Athanasios Katsamanis, Shri Narayanan |
INTERSPEECH | 3 |
| 2010 | A new multichannel multi modal dyadic interaction databaseabstractIn this work we present a new multi-modal database for analysis of participant behaviors in dyadic interactions. This database contains multiple channels with closeand far-field audio, a high definition camera array and motion capture data. Presence of the motion capture allows precise analysis of the body language low-level descriptors and its comparison with similar descriptors derived from video data. Data is manually labeled by multiple human annotators using psychologyinformed guides. This work also presents an initial analysis of approach-avoidance (A-A) behavior. Two sets of annotations are provided, one based on video only and the other obtained by using both the audio and video channels. Additionally, we describe the statistics of interaction descriptors and A-A labels on participants’ roles. Finally we provide an analysis of relations between various non-verbal features and approach/avoidance labels. Viktor Rozgic, Bo Xiao 0003, Athanasios Katsamanis, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 3 |
| 2009 | Product-HMMs for automatic sign language recognitionabstractWe address multistream sign language recognition and focus on efficient multistream integration schemes. Alternative approaches are investigated and the application of Product-HMMs (PHMM) is proposed. The PHMM is a variant of the general multistream HMM that also allows for partial asynchrony between the streams. Experiments in classification and isolated sign recognition for the Greek sign language using different fusion methods, show that the PHMMs perform the best. Fusing movement and shape information with the PHMMs has increased sign classification performance by 1,2% in comparison to the Parallel HMM fusion model. Isolated sign recognition rate increased by 8,3% over movement only models and by 1,5% over movement-shape models using multistream HMMs. Stavros Theodorakis, Athanasios Katsamanis, Petros Maragos |
ICASSP | 2 |
| 2009 | Tongue tracking in Ultrasound images with Active Appearance ModelsabstractTongue Ultrasound imaging is widely used for human speech production analysis and modeling. In this paper, we propose a novel method to automatically detect and track the tongue contour in Ultrasound (US) videos. Our method is built on a variant of Active Appearance Modeling. It incorporates shape prior information and can estimate the entire tongue contour robustly and accurately in a sequence of US frames. Experimental evaluation demonstrates the effectiveness of our approach and its improved performance compared to previously proposed tongue tracking techniques. Anastasios Roussos, Athanasios Katsamanis, Petros Maragos |
ICIP | 2 |
| 2009 | Face Active Appearance Modeling and Speech Acoustic Information to Recover ArticulationabstractWe are interested in recovering aspects of vocal tract's geometry and dynamics from speech, a problem referred to as speech inversion. Traditional audio-only speech inversion techniques are inherently ill-posed since the same speech acoustics can be produced by multiple articulatory configurations. To alleviate the ill-posedness of the audio-only inversion process, we propose an inversion scheme which also exploits visual information from the speaker's face. The complex audiovisual-to-articulatory mapping is approximated by an adaptive piecewise linear model. Model switching is governed by a Markovian discrete process which captures articulatory dynamic information. Each constituent linear mapping is effectively estimated via canonical correlation analysis. In the described multimodal context, we investigate alternative fusion schemes which allow interaction between the audio and visual modalities at various synchronization levels. For facial analysis, we employ active appearance models (AAMs) and demonstrate fully automatic face tracking and visual feature extraction. Using the AAM features in conjunction with audio features such as Mel frequency cepstral coefficients (MFCCs) or line spectral frequencies (LSFs) leads to effective estimation of the trajectories followed by certain points of interest in the speech production system. We report experiments on the QSMT and MOCHA databases which contain audio, video, and electromagnetic articulography data recorded in parallel. The results show that exploiting both audio and visual modalities in a multistream hidden Markov model based scheme clearly improves performance relative to either audio or visual-only estimation. Athanasios Katsamanis, George Papandreou, Petros Maragos |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Adaptive Multimodal Fusion by Uncertainty Compensation With Application to Audiovisual Speech RecognitionabstractWhile the accuracy of feature measurements heavily depends on changing environmental conditions, studying the consequences of this fact in pattern recognition tasks has received relatively little attention to date. In this paper, we explicitly take feature measurement uncertainty into account and show how multimodal classification and learning rules should be adjusted to compensate for its effects. Our approach is particularly fruitful in multimodal fusion scenarios, such as audiovisual speech recognition, where multiple streams of complementary time-evolving features are integrated. For such applications, provided that the measurement noise uncertainty for each feature stream can be estimated, the proposed framework leads to highly adaptive multimodal fusion rules which are easy and efficient to implement. Our technique is widely applicable and can be transparently integrated with either synchronous or asynchronous multimodal sequence integration architectures. We further show that multimodal fusion methods relying on stream weights can naturally emerge from our scheme under certain assumptions; this connection provides valuable insights into the adaptivity properties of our multimodal uncertainty compensation approach. We show how these ideas can be practically applied for audiovisual speech recognition. In this context, we propose improved techniques for person-independent visual feature extraction and uncertainty estimation with active appearance models, and also discuss how enhanced audio features along with their uncertainty estimates can be effectively computed. We demonstrate the efficacy of our approach in audiovisual speech recognition experiments on the CUAVE database using either synchronous or asynchronous multimodal integration models. George Papandreou, Athanasios Katsamanis, Vassilis Pitsikalis, Petros Maragos |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Audiovisual-to-articulatory speech inversion using Active Appearance Models for the face and Hidden Markov Models for the dynamicsabstractWe are interested in recovering aspects of vocal tract's geometry and dynamics from auditory and visual speech cues. We approach the problem in a statistical framework based on Hidden Markov Models and demonstrate effective estimation of the trajectories followed by certain points of interest in the speech production system. Alternative fusion schemes are investigated to account for asynchrony between the modalities and allow independent modeling of the dynamics of the involved streams. Visual cues are extracted from the speaker's face by means of active appearance modeling. We report experiments on the QSMT database which contains audio, video, and electromagnetic articulography data recorded in parallel. The results show that exploiting both audio and visual modalities in a multistream HMM based scheme clearly improves performance relative to either audio or visual-only estimation. Athanasios Katsamanis, George Papandreou, Petros Maragos |
ICASSP | 1 |
| 2008 | Multisensor multiband cross-energy tracking for feature extraction and recognitionabstractIn this paper, we present a multisensor multiband energy tracking scheme for robust feature extraction in noisy environments. We introduce a multisensor feature extraction algorithm which combines both the spatial and frequency information incorporated in the speech signals captured by a microphone array. This is based on the estimation of cross-energies over multiple sensors and minimization of an error term due to noise. The relevant noise-analysis is given. Automatic speech recognition (ASR) experiments at various SNR levels demonstrate that the newly proposed frontend performs better than alternative schemes, especially in noisy conditions. Stamatios Lefkimmiatis, Petros Maragos, Athanasios Katsamanis |
ICASSP | 3 |
| 2007 | Audiovisual-to-Articulatory Speech Inversion Using HMMsabstractWe address the problem of audiovisual speech inversion, namely recovering the vocal tract's geometry from auditory and visual speech cues. We approach the problem in a statistical framework, combining ideas from multistream Hidden Markov Models and canonical correlation analysis, and demonstrate effective estimation of the trajectories followed by certain points of interest in the speech production system. Our experiments show that exploiting both audio and visual modalities clearly improves performance relative to either audio-only or visual-only estimation. We report experiments on the QSMT database which contains audio, video, and electromagnetic articulography data recorded in parallel. Athanasios Katsamanis, George Papandreou, Petros Maragos |
MMSP | 1 |
| 2007 | Multimodal Fusion and Learning with Uncertain Features Applied to Audiovisual Speech RecognitionabstractWe study the effect of uncertain feature measurements and show how classification and learning rules should be adjusted to compensate for it. Our approach is particularly fruitful in multimodal fusion scenarios, such as audio-visual speech recognition, where multiple streams of complementary features whose reliability is time-varying are integrated. For such applications, by taking the measurement noise uncertainty of each feature stream into account, the proposed framework leads to highly adaptive multimodal fusion rules for classification and learning which are widely applicable and easy to implement. We further show that previous multimodal fusion methods relying on stream weights fall under our scheme under certain assumptions; this provides novel insights into their applicability for various tasks and suggests new practical ways for estimating the stream weights adaptively. The potential of our approach is demonstrated in audio-visual speech recognition experiments. George Papandreou, Athanasios Katsamanis, Vassilis Pitsikalis, Petros Maragos |
MMSP | 2 |
| 2006 | Adaptive multimodal fusion by uncertainty compensationabstractWhile the accuracy of feature measurements heavily depends on changing environmental conditions, studying the consequences of this fact in pattern recognition tasks has received relatively little attention to date. In this work we explicitly take into account feature measurement uncertainty and we show how classification rules should be adjusted to compensate for its effects. Our approach is particularly fruitful in multimodal fusion scenarios, such as audio-visual speech recognition, where multiple streams of complementary time-evolving features are integrated. For such applications, provided that the measurement noise uncertainty for each feature stream can be estimated, the proposed framework leads to highly adaptive multimodal fusion rules which are widely applicable and easy to implement. We further show that previous multimodal fusion methods relying on stream weights fall under our scheme under certain assumptions; this provides novel insights into their applicability for various tasks and suggests new practical ways for estimating the stream weights adaptively. The potential of our approach is demonstrated in audio-visual speech recognition using either synchronous or asynchronous models. Vassilis Pitsikalis, Athanasios Katsamanis, George Papandreou, Petros Maragos |
INTERSPEECH | 2 |
| 2005 | Advances in statistical estimation and tracking of AM-FM speech componentsabstractIn this paper we present two extensions of a statistical framework to demodulate speech resonances, which are modeled as AM-FM signals. The first approach utilizes bandpass filtering and a standard demodulation algorithm which regularizes instantaneous amplitude and frequency estimates. The second employs particle filtering techniques to allow temporal variations of the parameters that are connected with spectral characteristics of the analyzed signal. Results are presented on both synthetic and real speech signals and improved performance is demonstrated. Both approaches appear to cope quite satisfactorily with the nonstationarity of speech signals. 1. Athanasios Katsamanis, Petros Maragos |
INTERSPEECH | 1 |