Alexandros Potamianos

dblp:17/2202 · DBLP profile ↗
← Back
131ranked-venue papers
18as first author
15since 2021 · last 2025
0009-0007-1532-5288ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 91 · 15 first-author · 8 since 2021Artificial intelligence and machine learning · 81 · 9 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Databases, data management, data science and information retrieval · 3Computer networks · 2
YearPublicationVenuePosition
2025 Medusa: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions
Georgios Chatzichristodoulou, Despoina Kosmopoulou, Antonios Kritikos, Anastasia Poulopoulou, Efthymios Georgiou, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos
INTERSPEECH8
2025 MSDA: Combining Pseudo-labeling and Self-Supervision for Unsupervised Domain Adaptation in ASR
Dimitrios Damianos, Georgios Paraskevopoulos, Alexandros Potamianos
INTERSPEECH3
2025 Aggregation Artifacts in Subjective Tasks Collapse Large Language Models' Posteriors
abstract
Georgios Chochlakis, Alexandros Potamianos, Kristina Lerman, Shrikanth Narayanan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Georgios Chochlakis, Alexandros Potamianos, Kristina Lerman, Shri Narayanan
NAACL (Long Papers)2
2025 Auto-Compressing Networks
abstract
Deep neural networks with short residual connections have demonstrated remarkable success across domains, but increasing depth often introduces computational redundancy without corresponding improvements in representation quality. We introduce Auto-Compressing Networks (ACNs), an architectural variant where additive long feedforward connections from each layer to the output replace traditional short residual connections. By analyzing the distinct dynamics induced by this modification, we reveal a unique property we coin as *auto-compression*—the ability of a network to organically compress information during training with gradient descent, through architectural design alone. Through auto-compression, information is dynamically "pushed" into early layers during training, enhancing their representational quality and revealing potential redundancy in deeper ones. We theoretically show that this property emerges from layer-wise training patterns found only in ACNs, where layers are dynamically utilized during training based on task requirements. We also find that ACNs exhibit enhanced noise robustness compared to residual networks, superior performance in low-data settings, improved transfer learning capabilities, and mitigate catastrophic forgetting suggesting that they learn representations that generalize better despite using fewer parameters. Our results demonstrate up to 18\% reduction in catastrophic forgetting and 30-80\% architectural compression while maintaining accuracy across vision transformers, MLP-mixers, and BERT architectures. These findings establish ACNs as a practical approach to developing efficient neural architectures that automatically adapt their computational footprint to task complexity, while learning robust representations suitable for noisy real-world tasks and continual learning scenarios.
Vaggelis Dorovatas, Georgios Paraskevopoulos, Alexandros Potamianos
NeurIPS3
2024 The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition
abstract
In-context Learning (ICL) has emerged as a power-ful paradigm for performing natural language tasks with Large Language Models (LLM) without updating the models' parameters, in contrast to the traditional gradient-based finetuning. The promise of ICL is that the LLM can adapt to perform the present task at a competitive or state-of-the-art level at a fraction of the cost. The ability of LLMs to perform tasks in this few-shot manner relies on their background knowledge of the task (or task priors). However, recent work has found that, unlike traditional learning, LLMs are unable to fully integrate information from demonstrations that contrast task priors. This can lead to performance saturation at suboptimal levels, especially for subjective tasks such as emotion recognition, where the mapping from text to emotions can differ widely due to variability in human annotations. In this work, we design experiments and propose measurements to explicitly quantify the consistency of proxies of LLM priors and their pull on the posteriors. We show that LLMs have strong yet inconsistent priors in emotion recognition that ossify their predictions. We also find that the larger the model, the stronger these effects become. Our results suggest that caution is needed when using ICL with larger LLMs for affect-centered tasks outside their pretraining domain and when interpreting ICL results.1
Georgios Chochlakis, Alexandros Potamianos, Kristina Lerman, Shri Narayanan
ACII2
2024 $\mathcal {P}$owMix: A Versatile Regularizer for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) leverages heterogeneous data sources to interpret the complex nature of human sentiments. Despite significant progress in multimodal architecture design, the field lacks comprehensive regularization methods. This paper introduces$\mathcal {P}$owMix, a versatile embedding space regularizer that builds upon the strengths of unimodal mixing-based regularization approaches and introduces novel algorithmic components that are specifically tailored to multimodal tasks.$\mathcal {P}$owMix is integrated before the fusion stage of multimodal architectures and facilitates intra-modal mixing, such as mixing text with text, to act as a regularizer.$\mathcal {P}$owMix consists of five components: 1) a varying number of generated mixed examples, 2) mixing factor reweighting, 3) anisotropic mixing, 4) dynamic mixing, and 5) cross-modal label mixing. Extensive experimentation across benchmark MSA datasets and a broad spectrum of diverse architectural designs demonstrate the efficacy of$\mathcal {P}$owMix, as evidenced by consistent performance improvements over baselines and existing mixing methods. An in-depth ablation study highlights the critical contribution of each$\mathcal {P}$owMix component and how they synergistically enhance performance. Furthermore, algorithmic analysis demonstrates how$\mathcal {P}$owMix behaves in different scenarios, particularly comparing early versus late fusion architectures. Notably,$\mathcal {P}$owMix enhances overall performance without sacrificing model robustness or magnifying text dominance. It also retains its strong performance in situations of limited data. Our findings position$\mathcal {P}$owMix as a promising versatile regularization strategy for MSA.
Efthymios Georgiou, Yannis Avrithis, Alexandros Potamianos
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 Sample-Efficient Unsupervised Domain Adaptation of Speech Recognition Systems: A Case Study for Modern Greek
abstract
Modern speech recognition systems exhibit rapid performance degradation under domain shift. This issue is especially prevalent in data-scarce settings, such as low-resource languages, where the diversity of training data is limited. In this work, we propose M2DS2, a simple and sample-efficient fine-tuning strategy for large pre-trained speech models, based on mixed source and target domain self-supervision. We find that including source domain self-supervision stabilizes training and avoids mode collapse of the latent representations. For evaluation, we collect HParl, a 120-hour speech corpus for Greek, consisting of plenary sessions in the Greek Parliament. We merge HParl with two popular Greek corpora to create GREC-MD, a test-bed for multi-domain evaluation of Greek ASR systems. In our experiments, we find that, while other Unsupervised Domain Adaptation baselines fail in this resource-constrained environment, M2DS2 yields significant improvements for cross-domain adaptation, even when only a few hours of in-domain audio are available. When we relax the problem in a weakly supervised setting, we find that independent adaptation for audio using M2DS2 and language using simple LM augmentation techniques is particularly effective, yielding word error rates comparable to the fully supervised baselines.
Georgios Paraskevopoulos, Theodoros Kouzelis, Georgios Rouvalis, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos
IEEE ACM Trans. Audio Speech Lang. Process.6
2023 Adapted Multimodal Bert with Layer-Wise Fusion for Sentiment Analysis
abstract
Multimodal learning pipelines have benefited from the success of pretrained language models. However, this comes at the cost of increased model parameters. In this work, we propose Adapted Multimodal BERT (AMB), a BERT-based architecture for multimodal tasks that uses a combination of adapter modules and intermediate fusion layers. The adapter adjusts the pretrained language model for the task at hand, while the fusion layers perform task-specific, layer-wise fusion of audio-visual information with textual BERT representations. During the adaptation process the pre-trained language model parameters remain frozen, allowing for fast, parameter-efficient training. In our ablations we see that this approach leads to efficient models, that can outperform their fine-tuned counterparts and are robust to input noise. Our experiments on sentiment analysis with CMU-MOSEI show that AMB outperforms the current state-of-the-art across metrics, with 3.4% relative reduction in the resulting error and 2.1% relative improvement in 7−class classification accuracy.
Odysseas S. Chlapanis, Georgios Paraskevopoulos, Alexandros Potamianos
ICASSP3
2023 A Zero-Shot Approach for Multi-User Task-Oriented Dialog Generation
abstract
Prior art investigating task-oriented dialog and automatic generation of such dialogs have focused on single-user dialogs between a single user and an agent.However, there is limited study on adapting such AI agents to multiuser conversations (involving multiple users and an agent).Multi-user conversations are richer than single-user conversations containing social banter and collaborative decision making.The most significant challenge impeding such studies is the lack of suitable multiuser task-oriented dialogs with annotations of user belief states and system actions.One potential solution is multi-user dialog generation from single-user data.Many single-user dialogs datasets already contain dialog state information (intents, slots), thus making them suitable candidates.In this work, we propose a novel approach for expanding single-user taskoriented dialogs (e.g.MultiWOZ) to multiuser dialogs in a zero-shot setting.
Shiv Surya, Yohan Jo, Arijit Biswas, Alexandros Potamianos
INLG4
2022 Mmlatch: Bottom-Up Top-Down Fusion For Multimodal Sentiment Analysis
abstract
Current deep learning approaches for multimodal fusion rely on bottom-up fusion of high and mid-level latent modality representations (late/mid fusion) or low level sensory inputs (early fusion). Models of human perception highlight the importance of top-down fusion, where high-level representations affect the way sensory inputs are perceived, i.e. cognition affects perception. These top-down interactions are not captured in current deep learning models. In this work we propose a neural architecture that captures top-down cross-modal interactions, using a feedback mechanism in the forward pass during network training. The proposed mechanism extracts high-level representations for each modality and uses these representations to mask the sensory inputs, allowing the model to perform top-down feature masking. We apply the proposed model for multimodal sentiment recognition on CMU-MOSEI. Our method shows consistent improvements over the well established MulT and over our strong late fusion baseline, achieving state-of-the-art results.
Georgios Paraskevopoulos, Efthymios Georgiou, Alexandros Potamianos
ICASSP3
2022 A Multi-Task BERT Model for Schema-Guided Dialogue State Tracking
abstract
Task-oriented dialogue systems often employ a Dialogue State Tracker (DST) to successfully complete conversations.Recent state-of-the-art DST implementations rely on schemata of diverse services to improve model robustness and handle zeroshot generalization to new domains [1], however such methods [2, 3] typically require multiple large scale transformer models and long input sequences to perform well.We propose a single multi-task BERT-based model that jointly solves the three DST tasks of intent prediction, requested slot prediction and slot filling.Moreover, we propose an efficient and parsimonious encoding of the dialogue history and service schemata that is shown to further improve performance.Evaluation on the SGD dataset shows that our approach outperforms the baseline SGP-DST by a large margin and performs well compared to the state-of-the-art, while being significantly more computationally efficient.Extensive ablation studies are performed to examine the contributing factors to the success of our model.
Eleftherios Kapelonis, Efthymios Georgiou, Alexandros Potamianos
INTERSPEECH3
2022 Extending Compositional Attention Networks for Social Reasoning in Videos
abstract
We propose a novel deep architecture for the task of reasoning about social interactions in videos. We leverage the multi-step reasoning capabilities of Compositional Attention Networks (MAC), and propose a multimodal extension (MAC-X). MAC-X is based on a recurrent cell that performs iterative mid-level fusion of input modalities (visual, auditory, text) over multiple reasoning steps, by use of a temporal attention mechanism. We then combine MAC-X with LSTMs for temporal input processing in an end-to-end architecture. Our ablation studies show that the proposed MAC-X architecture can effectively leverage multimodal input cues using mid-level fusion mechanisms. We apply MAC-X to the task of Social Video Question Answering in the Social IQ dataset and obtain a 2.5% absolute improvement in terms of binary accuracy over the current state-of-the-art.
Christina Sartzetaki, Georgios Paraskevopoulos, Alexandros Potamianos
INTERSPEECH3
2022 Regotron: Regularizing the Tacotron2 Architecture Via Monotonic Alignment Loss
abstract
Deep learning Text-to-Speech (TTS) systems have achieved impressive generated speech quality, close to human parity. However, they suffer from training stability issues and in-correct alignment between the intermediate acoustic representation and the text input. In this work, we propose Regotron, a regularized Tacotron2 version which alleviates the training issues by augmenting the objective function with an additional term, which penalizes non-monotonic alignments in the location-sensitive attention mechanism. By introducing this regularization term we demonstrate its effectiveness to stabilize the training process, produce a monotonic attention quicker (13% of the total number of epochs compared to Tacotron2) and reduce the alignment errors during inference. Moreover, Regotron has minimal additional computational overhead, reduces common TTS mistakes and at the same time achieves improved speech naturalness according to subjective mean opinion scores (MOS) collected from 50 evaluators.
Efthymios Georgiou, Kosmas Kritsis, Georgios Paraskevopoulos, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos
SLT6
2021 M3: MultiModal Masking Applied to Sentiment Analysis
Efthymios Georgiou, Georgios Paraskevopoulos, Alexandros Potamianos
Interspeech3
2021 UDALM: Unsupervised Domain Adaptation through Language Modeling
abstract
Constantinos Karouzos, Georgios Paraskevopoulos, Alexandros Potamianos. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Constantinos Karouzos, Georgios Paraskevopoulos, Alexandros Potamianos
NAACL-HLT3
2020 Affective Conditioning on Hierarchical Attention Networks Applied to Depression Detection from Transcribed Clinical Interviews
abstract
In this work we propose a machine learning model for depression detection from transcribed clinical interviews. Depression is a mental disorder that impacts not only the subject's mood but also the use of language. To this end we use a Hierarchical Attention Network to classify interviews of depressed subjects. We augment the attention layer of our model with a conditioning mechanism on linguistic features, extracted from affective lexica. Our analysis shows that individuals diagnosed with depression use affective language to a greater extent than not-depressed. Our experiments show that external affective information improves the performance of the proposed architecture in the General Psychotherapy Corpus and the DAIC-WoZ 2017 depression datasets, achieving state-of-the-art 71.6 and 68.6 F1 scores respectively.
Danai Xezonaki, Georgios Paraskevopoulos, Alexandros Potamianos, Shri Narayanan
INTERSPEECH3
2019 Attention-based Conditioning Methods for External Knowledge Integration
abstract
In this paper, we present a novel approach for incorporating external knowledge in Recurrent Neural Networks (RNNs).We propose the integration of lexicon features into the self-attention mechanism of RNN-based architectures.This form of conditioning on the attention distribution, enforces the contribution of the most salient words for the task at hand.We introduce three methods, namely attentional concatenation, feature-based gating and affine transformation.Experiments on six benchmark datasets show the effectiveness of our methods.Attentional feature-based gating yields consistent performance improvement across tasks.Our approach is implemented as a simple add-on module for RNN-based models with minimal computational overhead and can be adapted to any deep neural architecture.
Aikaterini Margatina, Christos Baziotis, Alexandros Potamianos
ACL (1)3
2019 Using Oliver API for emotion-aware movie content characterization
abstract
This paper demonstrates the utilization of Oliver11https://behavioralsignals.com/oliver/, the speech emotion recognition (SER) API created by Behavioral Signals, in the context of a movie content visualization application. Oliver API provides an emotion recognition as-a-service solution that can be accessed via a Web API. In this work, we demonstrate how one can send sound recordings from famous movies, retrieve respective emotional descriptors and use simple aggregations on these descriptors to visualize movie content. We have compiled a dataset of 60 movies, categorized over 8 directors. The classification examples included in this paper indicate the ability of simple emotion aggregations to discriminate between movie directors. In order for others to also experiment with the output of both the API's Emotional and Automatic Speech Recognition, the responses are provided as JSON files in this link: https://tinyurl.com/yxeqvvy2.
Theodoros Giannakopoulos, Spiros Dimopoulos, Georgios Pantazopoulos, Aggelina Chatziagapi, Dimitris Sgouropoulos, Athanasios Katsamanis, Alexandros Potamianos, Shri Narayanan
CBMI7
2019 Data Augmentation Using GANs for Speech Emotion Recognition
Aggelina Chatziagapi, Georgios Paraskevopoulos, Dimitris Sgouropoulos, Georgios Pantazopoulos, Malvina Nikandrou, Theodoros Giannakopoulos, Athanasios Katsamanis, Alexandros Potamianos, Shri Narayanan
INTERSPEECH8
2019 Deep Hierarchical Fusion with Application in Sentiment Analysis
Efthymios Georgiou, Charilaos Papaioannou, Alexandros Potamianos
INTERSPEECH3
2019 Unsupervised Low-Rank Representations for Speech Emotion Recognition
abstract
We examine the use of linear and non-linear dimensionality reduction algorithms for extracting low-rank feature representations for speech emotion recognition. Two feature sets are used, one based on low-level descriptors and their aggregations (IS10) and one modeling recurrence dynamics of speech (RQA), as well as their fusion. We report speech emotion recognition (SER) results for learned representations on two databases using different classification methods. Classification with low-dimensional representations yields performance improvement in a variety of settings. This indicates that dimensionality reduction is an effective way to combat the curse of dimensionality for SER. Visualization of features in two dimensions provides insight into discriminatory abilities of reduced feature sets.
Georgios Paraskevopoulos, Efthymios Tzinis, Nikolaos Ellinas, Theodoros Giannakopoulos, Alexandros Potamianos
INTERSPEECH5
2018 Neural Activation Semantic Models: Computational lexical semantic models of localized neural activations
abstract
Neural activation models have been proposed in the literature that use a set of example words for which fMRI measurements are available in order to find a mapping between word semantics and localized neural activations. Successful mappings let us expand to the full lexicon of concrete nouns using the assumption that similarity of meaning implies similar neural activation patterns. In this paper, we propose a computational model that estimates semantic similarity in the neural activation space and investigates the relative performance of this model for various natural language processing tasks. Despite the simplicity of the proposed model and the very small number of example words used to bootstrap it, the neural activation semantic model performs surprisingly well compared to state-of-the-art word embeddings. Specifically, the neural activation semantic model performs better than the state-of-the-art for the task of semantic similarity estimation between very similar or very dissimilar words, while performing well on other tasks such as entailment and word categorization. These are strong indications that neural activation semantic models can not only shed some light into human cognition but also contribute to computation models for certain tasks.
Nikos Athanasiou, Elias Iosif, Alexandros Potamianos
COLING3
2018 Integrating Recurrence Dynamics for Speech Emotion Recognition
abstract
We investigate the performance of features that can capture nonlinear recurrence dynamics embedded in the speech signal for the task of Speech Emotion Recognition (SER). Reconstruction of the phase space of each speech frame and the computation of its respective Recurrence Plot (RP) reveals complex structures which can be measured by performing Recurrence Quantification Analysis (RQA). These measures are aggregated by using statistical functionals over segment and utterance periods. We report SER results for the proposed feature set on three databases using different classification methods. When fusing the proposed features with traditional feature sets, we show an improvement in unweighted accuracy of up to 5.7% and 10.7% on Speaker-Dependent (SD) and Speaker-Independent (SI) SER tasks, respectively, over the baseline. Following a segment-based approach we demonstrate state-of-the-art performance on IEMOCAP using a Bidirectional Recurrent Neural Network.
Efthymios Tzinis, Georgios Paraskevopoulos, Christos Baziotis, Alexandros Potamianos
INTERSPEECH4
2018 Speech understanding for spoken dialogue systems: From corpus harvesting to grammar rule induction
Elias Iosif, Ioannis Klasinas, Georgia Athanasopoulou, Elisavet Palogiannidi, Spiros Georgiladakis, Katerina Louka, Alexandros Potamianos
Comput. Speech Lang.7
2017 Segment-based speech emotion recognition using recurrent neural networks
abstract
Recently, Recurrent Neural Networks (RNNs) have produced state-of-the-art results for Speech Emotion Recognition (SER). The choice of the appropriate time-scale for Low Level Descriptors (LLDs) (local features) and statistical functionals (global features) is key for a high performing SER system. In this paper, we investigate both local and global features and evaluate the performance at various time-scales (frame, phoneme, word or utterance). We show that for RNN models, extracting statistical functionals over speech segments that roughly correspond to the duration of a couple of words produces optimal accuracy. We report state-of-the-art SER performance on the IEMOCAP corpus at a significantly lower model and computational complexity.
Efthymios Tzinis, Alexandros Potamianos
ACII2
2017 Engagement detection for children with Autism Spectrum Disorder
abstract
Children with Autism Spectrum Disorder (ASD) face several difficulties in social communication. Hence, analyzing social interaction can provide insight on their social and cognitive skills. In this paper, we investigate the degree of engagement of children in interactions with their parents. Features derived from both participants including acoustic, linguistic and dialogue act features are explored. The effect of visual cues is also investigated. We experimented on the task of engagement detection using video-recorded sessions consisting of interactions of typically developing (TD) and ASD children. Results show that engagement is easier to predict for TD children than for ASD children, and that the parent's actions/movements are better predictors of the child's degree of engagement.
Arodami Chorianopoulou, Efthymios Tzinis, Elias Iosif, Asimenia Papoulidi, Christina Papailiou, Alexandros Potamianos
ICASSP6
2016 Speech Emotion Recognition Using Affective Saliency
abstract
We investigate an affective saliency approach for speech emotion recognition of spoken dialogue utterances that estimates the amount of emotional information over time. The proposed saliency approach uses a regression model that combines features extracted from the acoustic signal and the posteriors of a segment-level classifier to obtain frame or segment-level ratings. The affective saliency model is trained using a minimum classification error (MCE) criterion that learns the weights by optimizing an objective loss function related to the classification error rate of the emotion recognition system. Affective saliency scores are then used to weight the contribution of frame-level posteriors and/or features to the speech emotion classification decision. The algorithm is evaluated for the task of anger detection on four call-center datasets for two languages, Greek and English, with good results.
Arodami Chorianopoulou, Polychronis Koutsakis, Alexandros Potamianos
INTERSPEECH3
2016 Root Cause Analysis of Miscommunication Hotspots in Spoken Dialogue Systems
abstract
A major challenge in Spoken Dialogue Systems (SDS) is the detection of problematic communication (hotspots), as well as the classification of these hotspots into different types (root cause analysi ...
Spiros Georgiladakis, Georgia Athanasopoulou, Raveesh Meena, José Lopes 0001, Arodami Chorianopoulou, Elisavet Palogiannidi, Elias Iosif, Gabriel Skantze, Alexandros Potamianos
INTERSPEECH9
2016 Audio-Based Distributional Representations of Meaning Using a Fusion of Feature Encodings
Giannis Karamanolakis, Elias Iosif, Athanasia Zlatintsi, Aggelos Pikrakis, Alexandros Potamianos
INTERSPEECH5
2016 Cognitively Motivated Distributional Representations of Meaning
Elias Iosif, Spiros Georgiladakis, Alexandros Potamianos
LREC3
2016 Crossmodal Network-Based Distributional Semantic Models
Elias Iosif, Alexandros Potamianos
LREC2
2016 The SpeDial datasets: datasets for Spoken Dialogue Systems analytics
José Lopes 0001, Arodami Chorianopoulou, Elisavet Palogiannidi, Helena Moniz, Alberto Abad, Katerina Louka, Elias Iosif, Alexandros Potamianos
LREC8
2016 Affective Lexicon Creation for the Greek Language
Elisavet Palogiannidi, Polychronis Koutsakis, Elias Iosif, Alexandros Potamianos
LREC4
2015 Predicting audio-visual salient events based on visual, audio and text modalities for movie summarization
abstract
In this paper, we present a new and improved synergistic approach to the problem of audio-visual salient event detection and movie summarization based on visual, audio and text modalities. Spatio-temporal visual saliency is estimated through a perceptually inspired frontend based on 3D (space, time) Gabor filters and frame-wise features are extracted from the saliency volumes. For the auditory salient event detection we extract features based on Teager-Kaiser Energy Operator, while text analysis incorporates part-of-speech tagging and affective modeling of single words on the movie subtitles. For the evaluation of the proposed system, we employ an elementary and non-parametric classification technique like KNN. Detection results are reported on the MovSum database, using objective evaluations against ground-truth denoting the perceptually salient events, and human evaluations of the movie summaries. Our evaluation verifies the appropriateness of the proposed methods compared to our baseline system. Finally, our newly proposed summarization algorithm produces summaries that consist of salient and meaningful events, also improving the comprehension of the semantics.
Petros Koutras, Athanasia Zlatintsi, Elias Iosif, Athanasios Katsamanis, Petros Maragos, Alexandros Potamianos
ICIP6
2015 Valence, arousal and dominance estimation for English, German, Greek, Portuguese and Spanish lexica using semantic models
abstract
We propose and evaluate the use of an affective-semantic model to expand the affective lexica of German, Greek, English, Spanish and Portuguese. Motivated by the assumption that semantic similarity implies affective similarity, we use word level semantic similarity scores as semantic features to estimate their corresponding affective scores. Various context-based semantic similarity metrics are investigated using contextual features that include both words and character n-grams. The model produces continuous affective ratings in three dimensions (valence, arousal and dominance) for all five languages, achieving consistent performance. We achieve classification accuracy (valence polarity task) between 85% and 91% for all five languages. For morphologically rich languages the proposed use of character n-grams is shown to improve performance.
Elisavet Palogiannidi, Elias Iosif, Polychronis Koutsakis, Alexandros Potamianos
INTERSPEECH4
2015 Similarity computation using semantic networks created from web-harvested data
abstract
Abstract We investigate language-agnostic algorithms for the construction of unsupervised distributional semantic models using web-harvested corpora. Specifically, a corpus is created from web document snippets, and the relevant semantic similarity statistics are encoded in a semantic network. We propose the notion of semantic neighborhoods that are defined using co-occurrence or context similarity features. Three neighborhood-based similarity metrics are proposed, motivated by the hypotheses of attributional and maximum sense similarity. The proposed metrics are evaluated against human similarity ratings achieving state-of-the-art results.
Elias Iosif, Alexandros Potamianos
Nat. Lang. Eng.2
2014 Low-Dimensional Manifold Distributional Semantic Models
Georgia Athanasopoulou, Elias Iosif, Alexandros Potamianos
COLING3
2014 Affective language model adaptation via corpus selection
abstract
Motivated by methods used in language modeling and grammar induction, we propose the use of pragmatic constraints and perplexity as criteria to filter the unlabeled data used to generate the semantic similarity model. We investigate unsupervised adaptation algorithms of the semantic-affective models proposed in [1, 2]. Affective ratings at the utterance level are generated based on an emotional lexicon, which in turn is created using a semantic (similarity) model estimated over raw, unlabeled text. The proposed adaptation method creates task-dependent semantic similarity models and task-dependent word/term affective ratings. The proposed adaptation algorithms are tested on anger/distress detection of transcribed speech data and sentiment analysis in tweets showing significant relative classification error reduction of up to 10%.
Nikos Malandrakis, Alexandros Potamianos, Kean J. Hsu, Kalina N. Babeva, Michelle C. Feng, Gerald C. Davison, Shri Narayanan
ICASSP2
2014 Spoken dialogue grammar induction from crowdsourced data
abstract
We design and evaluate various crowdsourcing tasks for eliciting spoken dialogue data. Task design is based on an array of parameters that quantify the basic characteristics of the elicitation questions, e.g., how open-ended is a question. The crowdsourced data are used for and evaluated on the unsupervised induction of semantic classes for speech understanding grammars. We show that grammar induction performance is significantly affected by the crowdsourcing task parameters, e.g., paraphrasing tasks prime high lexical entrain-ment and result in poor corpus/grammar quality. The task parameters along with perplexity filters are used for corpus selection achieving grammar induction performance that is comparable to that of using in-domain spoken dialogue data.
Elisavet Palogiannidi, Ioannis Klasinas, Alexandros Potamianos, Elias Iosif
ICASSP3
2014 An investigation of vocal arousal dynamics in child-psychologist interactions using synchrony measures and a conversation-based model
abstract
Researchers from various disciplines are concerned with the study of affective phenomena, especially arousal. Expressed affective modulations, which reflect both an individual’s in-ternal state and external factors, are central to the commu-nicative process. Bone et al. developed a robust, unsuper-vised (rule-based) method which provides a scale-continuous, bounded arousal rating from the vocal signal. In this study, we investigate the joint-dynamics of child and psychologist vocal arousal in autism spectrum disorder (ASD) diagnostic interac-tions. Arousal synchrony is assessed with multiple methods. Results indicate that children with higher ASD severity tend to lead the arousal dynamics more, seemingly because the children aren’t as responsive to the psychologist’s affective modulations. A vocal arousal model is also proposed which incorporates so-cial and conversational constructs. The model captures conver-sational signal relations, and is able to distinguish between high and low ASD severity at accuracies well-above chance.
Daniel Bone, Chi-Chun Lee, Alexandros Potamianos, Shri Narayanan
INTERSPEECH3
2014 Fusion of knowledge-based and data-driven approaches to grammar induction
abstract
Using different sources of information for grammar induction results in grammars that vary in coverage and precision. Fusing such grammars with a strategy that exploits their strengths while minimizing their weaknesses is expected to produce grammars with superior performance. We focus on the fusion of grammars produced using a knowledge-based approach using lexicalized ontologies and a data-driven approach using semantic similarity clustering. We propose various algorithms for finding the map-ping between the (non-terminal) rules generated by each gram-mar induction algorithm, followed by rule fusion. Three fusion approaches are investigated: early, mid and late fusion. Results show that late fusion provides the best relative F-measure per-formance improvement by 20%. Index Terms: spoken dialogue systems, corpus-based grammar induction, ontology-based grammar induction, grammar fusion
Spiros Georgiladakis, Christina Unger, Elias Iosif, Sebastian Walter 0001, Philipp Cimiano, Euripides G. M. Petrakis, Alexandros Potamianos
INTERSPEECH7
2014 Classification of cognitive load from speech using an i-vector framework
abstract
The goal in this work is to automatically classify speakers ’ level of cognitive load (low, medium, high) from a standard battery of reading tasks requiring varying levels of working memory. This is a challenging machine learning problem because of the inherent difficulty in defining/measuring cognitive load and due to intra-/inter-speaker differences in how their effects are man-ifested in behavioral cues. We experimented with a number of static and dynamic features extracted directly from the audio signal (prosodic, spectral, voice quality) and from automatic speech recognition hypotheses (lexical information, speaking rate). Our approach to classification addressed the wide vari-ability and heterogeneity through speaker normalization and by adopting an i-vector framework that affords a systematic way to factorize the multiple sources of variability. Index Terms: computational paralinguistics, behavioral signal processing (BSP), prosody, ASR, i-vector, cognitive load
Maarten Van Segbroeck, Ruchir Travadi, Colin Vaz, Jangwon Kim, Matthew Black, Alexandros Potamianos, Shri Narayanan
INTERSPEECH6
2014 Word Semantic Similarity for Morphologically Rich Languages
Kalliopi Zervanou, Elias Iosif, Alexandros Potamianos
LREC3
2014 Using lexical, syntactic and semantic features for non-terminal grammar rule induction in Spoken Dialogue Systems
abstract
In this work, we propose an algorithm for the automatic induction of non-terminal grammar rules for Spoken Dialogue Systems (SDS). Initially, a grammar developer provides the system with a minimal set of rules that serve as seeding examples. Using these seed rules and (optionally) a seed corpus, in-domain data are harvested and filtered from the web. A challenging task is identifying relevant chunks (phrases) in the web-harvested corpus that are good candidates for enhancing the seed grammar. We propose and evaluate rule-based and statistical classification algorithms for this purpose that use lexical, syntactic and semantic features. Induced grammars are evaluated in terms of accuracy of the proposed rules for two spoken dialogue domains. Results show up to four times absolute precision improvement compared to the naive grammar induction approach using semantic phrase similarity.
Georgia Athanasopoulou, Ioannis Klasinas, Spiros Georgiladakis, Elias Iosif, Alexandros Potamianos
SLT5
2013 Continuous models of affect from text using n-grams
abstract
We propose a method of affective text analysis and modeling that is capable of generating continuous valence ratings at the sentence level starting from word and multi-word term valence ratings. Motivated from the language modeling literature, a back-off algorithm is employed to efficiently fuse the valence of single-word and multi-word terms. Specifically, a term detection criterion is used to select the appropriate n-gram terms, starting with bigrams and potentially backing off to unigrams. Term affective ratings are generated by a lexicon expansion method, using semantic similarity estimates computed on a large web corpus. The proposed framework provides state-of-the art results in the sentence level SemEval'07 task of news headline polarity detection, reaching an accuracy of 75%.
Nikos Malandrakis, Alexandros Potamianos, Shri Narayanan
ICASSP2
2013 Instantaneous frequency and bandwidth estimation using filterbank arrays
abstract
Accurate estimation of the instantaneous frequency of speech resonances is a hard problem mainly due to phase discontinuities in the speech signal associated with excitation instants. We review a variety of approaches for enhanced frequency and bandwidth estimation in the time-domain and propose a new cognitively motivated approach using filterbank arrays. We show that by filtering speech resonances using filters of different center frequency, bandwidth and shape, the ambiguity in instantaneous frequency estimation associated with amplitude envelope minima and phase discontinuities can be significantly reduced. The novel estimators are shown to perform well on synthetic speech signals with frequency and bandwidth micro-modulations (i.e., modulations within a pitch period), as well as on real speech signals. Filterbank arrays, when applied to frequency and bandwidth modulation index estimation, are shown to reduce the estimation error variance by 85% and 70% respectively.
Pirros Tsiakoulis, Alexandros Potamianos, Dimitrios Dimitriadis
ICASSP2
2013 Web data harvesting for speech understanding grammar induction
abstract
The development of a grammar for a spoken dialogue system can be greatly accelerated by using a corpus describing the application. However the development of such a corpus is a slow and expensive process. This paper proposes unsupervised methods for finding relevant corpora in the Web and mining the most informative parts. We show that by utilizing perplexity we are able to increase the in-domainess (precision) of the mined corpora, while by utilizing the rank of the web search engine we can increase the generalizability (recall). The results show that using only unsupervised and language independent methods we can compete with corpora created manually with expert knowledge.
Ioannis Klasinas, Alexandros Potamianos, Elias Iosif, Spiros Georgiladakis, Gianluca Mameli
INTERSPEECH2
2013 Affective evaluation of multimodal dialogue games for preschoolers using physiological signals
abstract
In this pilot study, we investigate the differences in the electroencephalography (EEG) signal patterns of children and adults while interacting with a multimodal dialogue computer game. The gaming application is designed for preschoolers, implements five popular learning tasks and has variable levels of difficulty. In this pilot, to simplify the data collection process for young children, we use the NeuroSky MindSet device which is a single forehead dry sensor device. The raw signals and the estimated attention, meditation and arousal signals are analyzed during the interaction and compared for adult and children user populations. Results show consistent variations as a function of modality used (speech vs mouse input), difficulty level and task success. The physiological signal pattern within an interaction turn is also estimated and analyzed. Overall, children and adults demonstrated very similar physiological signal patterns during multimodal interaction.
Vassiliki Kouloumenta, Manolis Perakakis, Alexandros Potamianos
INTERSPEECH3
2013 Affective classification of generic audio clips using regression models
abstract
We investigate acoustic modeling, feature extraction and feature selection for the problem of affective content recognition of generic, non-speech, non-music sounds. We annotate and analyze a database of generic sounds containing a subset of the BBC sound effects library. We use regression models, longterm features and wrapper-based feature selection to model affect in the continuous 3-D (arousal, valence, dominance) emotional space. The frame-level features for modeling are extracted from each audio clip and combined with functionals to estimate long term temporal patterns over the duration of the clip. Experimental results show that the regression models provide similar categorical performance as the more popular Gaussian Mixture Models. They are also capable of predicting accurate affective ratings on continuous scales, achieving 62-67% 3-class accuracy and 0.69-0.75 correlation with human ratings, higher than comparable numbers in literature.
Nikos Malandrakis, Shiva Sundaram, Alexandros Potamianos
INTERSPEECH3
2013 Multi-band long-term signal variability features for robust voice activity detection
abstract
In this paper, we propose robust features for the problem of voice activity detection (VAD). In particular, we extend the long term signal variability (LTSV) feature to accommodate multiple spectral bands. The motivation of the multi-band approach stems from the non-uniform frequency scale of speech phonemes and noise characteristics. Our analysis shows that the multi-band approach offers advantages over the single band LTSV for voice activity detection. In terms of classification accuracy, we show 0.3%-61.2% relative improvement over the best accuracy of the baselines considered for 7 out 8 different noisy channels. Experimental results, and error analysis, are reported on the DARPA RATS corpora of noisy speech. Index Terms: noisy speech data, voice activity detection, robust feature extraction
Andreas Tsiartas, Theodora Chaspari, Athanasios Katsamanis, Prasanta Kumar Ghosh, Ming Li 0026, Maarten Van Segbroeck, Alexandros Potamianos, Shri Narayanan
INTERSPEECH7
2013 Distributional Semantic Models for Affective Text Analysis
abstract
We present an affective text analysis model that can directly estimate and combine affective ratings of multi-word terms, with application to the problem of sentence polarity/semantic orientation detection. Starting from a hierarchical compositional method for generating sentence ratings, we expand the model by adding multi-word terms that can capture non-compositional semantics. The method operates similarly to a bigram language model, using bigram terms or backing off to unigrams based on a (degree of) compositionality criterion. The affective ratings for n-gram terms of different orders are estimated via a corpus-based method using distributional semantic similarity metrics between unseen words and a set of seed words. N-gram ratings are then combined into sentence ratings via simple algebraic formulas. The proposed framework produces state-of-the-art results for word-level tasks in English and German and the sentence-level news headlines classification SemEval'07-Task14 task. The inclusion of bigram terms to the model provides significant performance improvement, even if no term selection is applied.
Nikos Malandrakis, Alexandros Potamianos, Elias Iosif, Shri Narayanan
IEEE Trans. Speech Audio Process.2
2013 Toward the Automatic Extraction of Policy Networks Using Web Links and Documents
abstract
Policy networks are widely used by political scientists and economists to explain various financial and social phenomena, such as the development of partnerships between political entities or institutions from different levels of governance. The analysis of policy networks demands a series of arduous and time-consuming manual steps including interviews and questionnaires. In this paper, we estimate the strength of relations between actors in policy networks using features extracted from data harvested from the web. Features include webpage counts, outlinks, and lexical information extracted from web documents or web snippets. The proposed approach is automatic and does not require any external knowledge source, other than the specification of the word forms that correspond to the political actors. The features are evaluated both in isolation and jointly for both positive and negative (antagonistic) actor relations. The proposed algorithms are evaluated on two EU policy networks from the political science literature. Performance is measured in terms of correlation and mean square error between the human rated and the automatically extracted relations. Correlation of up to 0.74 is achieved for positive relations. The extracted networks are validated by political scientists and useful conclusions about the evolution of the networks over time are drawn.
Theodosis Moschopoulos, Elias Iosif, Leeda Demetropoulou, Alexandros Potamianos, Shri Narayanan
IEEE Trans. Knowl. Data Eng.4
2013 Multimodal Saliency and Fusion for Movie Summarization Based on Aural, Visual, and Textual Attention
abstract
Multimodal streams of sensory information are naturally parsed and integrated by humans using signal-level feature extraction and higher level cognitive processes. Detection of attention-invoking audiovisual segments is formulated in this work on the basis of saliency models for the audio, visual, and textual information conveyed in a video stream. Aural or auditory saliency is assessed by cues that quantify multifrequency waveform modulations, extracted through nonlinear operators and energy tracking. Visual saliency is measured through a spatiotemporal attention model driven by intensity, color, and orientation. Textual or linguistic saliency is extracted from part-of-speech tagging on the subtitles information available with most movie distributions. The individual saliency streams, obtained from modality-depended cues, are integrated in a multimodal saliency curve, modeling the time-varying perceptual importance of the composite video stream and signifying prevailing sensory events. The multimodal saliency representation forms the basis of a generic, bottom-up video summarization algorithm. Different fusion schemes are evaluated on a movie database of multimodal saliency annotations with comparative results provided across modalities. The produced summaries, based on low-level features and content-independent fusion and selection, are of subjectively high aesthetic and informative quality.
Georgios Evangelopoulos, Athanasia Zlatintsi, Alexandros Potamianos, Petros Maragos, Konstantinos Rapantzikos, Georgios Skoumas, Yannis Avrithis
IEEE Trans. Multim.3
2012 Associative and Semantic Features Extracted From Web-Harvested Corpora
Elias Iosif, Maria Giannoudaki, Eric Fosler-Lussier, Alexandros Potamianos
LREC4
2012 SemSim: Resources for Normalized Semantic Similarity Computation Using Lexical Networks
Elias Iosif, Alexandros Potamianos
LREC2
2012 Affective evaluation of a mobile multimodal dialogue system using brain signals
abstract
We propose the use of affective metrics such as excitement, frustration and engagement for the evaluation of multimodal dialogue systems. The affective metrics are elicited from the ElectroEncephaloGraphy (EEG) signals using the Emotiv EPOC neuroheadset device. The affective metrics are used in conjunction with traditional evaluation metrics (turn duration, input modality) to investigate the effect of speech recognition errors and modality usage patterns in a multimodal (touch and speech) dialogue form-filling application for the iPhone mobile device. Results show that: (1) engagement is higher for touch input, while excitement and frustration is higher for speech input, and (2) speech recognition errors and associated repairs correspond to specific dynamic patters of excitement and frustration. Use of such physiological channels and their elaborated interpretation is a challenging but also a potentially rewarding direction towards emotional and cognitive assessment of multimodal interaction design.
Manolis Perakakis, Alexandros Potamianos
SLT2
2011 A supervised approach to movie emotion tracking
abstract
In this paper, we present experiments on continuous time, continuous scale affective movie content recognition (emotion tracking). A major obstacle for emotion research has been the lack of appropriately annotated databases, limiting the potential for supervised algorithms. To that end we develop and present a database of movie affect, an notated in continuous time, on a continuous valence-arousal scale. Supervised learning methods are proposed to model the continuous affective response using hidden Markov Models (independent) in each dimension. These models classify each video frame into one of seven discrete categories (in each dimension); the discrete-valued curves are then converted to continuous values via spline interpolation. A variety of audio-visual features are investigated and an optimal feature set is selected. The potential of the method is experimentally verified on twelve 30-minute movie clips with good precision at a macroscopic level.
Nikos Malandrakis, Alexandros Potamianos, Georgios Evangelopoulos, Athanasia Zlatintsi
ICASSP2
2011 Kernel Models for Affective Lexicon Creation
abstract
Emotion recognition algorithms for spoken dialogue applications typically employ lexical models that are trained on labeled in-domain data. In this paper, we propose a domainindependent approach to affective text modeling that is based on the creation of an affective lexicon. Starting from a small set of manually annotated seed words, continuous valence ratings for new words are estimated using semantic similarity scores and a kernel model. The parameters of the model are trained using least mean squares estimation. Word level scores are combined to produce sentence-level scores via simple linear and non-linear fusion. The proposed method is evaluated on the SemEval news headline polarity task and on the ChIMP politeness and frustration detection dialogue task, achieving state-of-theart results on both. For politeness detection, best results are obtained when the affective model is adapted using in domain data. For frustration detection, the domain-independent model and non-linear fusion achieve the best performance. Index Terms: language understanding, emotion, affect, affective lexicon
Nikos Malandrakis, Alexandros Potamianos, Elias Iosif, Shri Narayanan
INTERSPEECH2
2011 Detecting emotional state of a child in a conversational computer game
Serdar Yildirim, Shri Narayanan, Alexandros Potamianos
Comput. Speech Lang.3
2011 On the Effects of Filterbank Design and Energy Computation on Robust Speech Recognition
abstract
In this paper, we examine how energy computation and filterbank design contribute to the overall front-end robustness, especially when the investigated features are applied to noisy speech signals, in mismatched training-testing conditions. In prior work (“Auditory Teager energy cepstrum coefficients for robust speech recognition,” D. Dimitriadis, P. Maragos, and A. Potamianos, in Proc. Eurospeech'05, Sep. 2005), a novel feature set called “Teager energy cepstrum coefficients” (TECCs) has been proposed, employing a dense, smooth filterbank and alternative energy computation schemes. TECCs were shown to be more robust to noise and exhibit improved performance compared to the widely used Mel frequency cepstral coefficients (MFCCs). In this paper, we attempt to interpret these results using a combined theoretical and experimental analysis framework. Specifically, we investigate in detail the connection between the filterbank design, i.e., the filter shape and bandwidth, the energy estimation scheme and the automatic speech recognition (ASR) performance under a variety of additive and/or convolutional noise conditions. For this purpose: 1) the performance of filterbanks using triangular, Gabor, and Gammatone filters with various bandwidths and filter positions are examined under different noisy speech recognition tasks, and 2) the squared amplitude and Teager-Kaiser energy operators are compared as two alternative approaches of computing the signal energy. Our end-goal is to understand how to select the most efficient filterbank and energy computation scheme that are maximally robust under both clean and noisy recording conditions. Theoretical and experimental results show that: 1) the filter bandwidth is one of the most important factors affecting speech recognition performance in noise, while the shape of the filter is of secondary importance, and 2) the Teager-Kaiser operator outperforms (on the average and for most noise types) the squared amplitude energy computation scheme for speech recognition in noisy conditions, especially, for large filter bandwidths. Experimental results show that selecting the appropriate filterbank and energy computation scheme can lead to significant error rate reduction over both MFCC and perceptual linear predicion (PLP) features for a variety of speech recognition tasks. A relative error rate reduction of up to ~ 30% for MFCCs and ~ 39% for PLPs is shown for the Aurora-3 Spanish Task.
Dimitrios Dimitriadis, Petros Maragos, Alexandros Potamianos
IEEE Trans. Speech Audio Process.3
2010 On the effect of fundamental frequency on amplitude and frequency modulation patterns in speech resonances
abstract
Amplitude modulation (AM) and frequency modulation (FM) in speech signals are believed to reflect various non-linear phenomena during the speech production process. In this paper, the amplitude and frequency modulation patterns are analyzed for the first three speech resonances in relation to the fundamental frequency (F0). The formant tracks are estimated, and the resonant signals are extracted and demodulated. The Amplitude Modulation Index (AMI) and Frequency Modulation Index (FMI) are computed, and examined in relation to the F0 value, as well as the relation between F0 and the first formant value (F1). Both AMI and FMI are significantly affected by pitch, with modulations being more frequently present in low F0 conditions. Evidence of non-linear interaction between the glottal source and the vocal tract is found in the dependence of the modulation patterns on the ratio of F1 over F0. AMI is amplified when pitch harmonics coincide with F1, while FMI shows complementary behavior. © 2010 ISCA.
Pirros Tsiakoulis, Alexandros Potamianos
INTERSPEECH2
2010 BabyExp: Constructing a Huge Multimodal Resource to Acquire Commonsense Knowledge Like Children Do
Massimo Poesio, Marco Baroni, Oswald Lanz, Alessandro Lenci, Alexandros Potamianos, Hinrich Schütze, Sabine Schulte im Walde, Luca Surian
LREC5
2010 Spectral Moment Features Augmented by Low Order Cepstral Coefficients for Robust ASR
abstract
We propose a novel Automatic Speech Recognition (ASR) front-end, that consists of the first central Spectral Moment time-frequency distribution Augmented by low order Cepstral coefficients (SMAC). We prove that the first central spectral moment is proportional to the spectral derivative with respect to the filter's central frequency. Consequently, the spectral moment is an estimate of the frequency domain derivative of the speech spectrum. However information related to the entire speech spectrum, such as the energy and the spectral tilt, is not adequately modeled. We propose adding this information with few cepstral coefficients. Furthermore, we use a mel-spaced Gabor filterbank with 70% frequency overlap in order to overcome the sensitivity to pitch harmonics. The novel SMAC front-end was evaluated for the speech recognition task for a variety of recording conditions. The experimental results have shown that SMAC performs at least as well as the standard MFCC front-end in clean conditions, and significantly outperforms MFCCs in noisy conditions.
Pirros Tsiakoulis, Alexandros Potamianos, Dimitrios Dimitriadis
IEEE Signal Process. Lett.2
2010 Batch and Adaptive PARAFAC-Based Blind Separation of Convolutive Speech Mixtures
abstract
We present a frequency-domain technique based on PARAllel FACtor (PARAFAC) analysis that performs multichannel blind source separation (BSS) of convolutive speech mixtures. PARAFAC algorithms are combined with a dimensionality reduction step to significantly reduce computational complexity. The identifiability potential of PARAFAC is exploited to derive a BSS algorithm for the under-determined case (more speakers than microphones), combining PARAFAC analysis with time-varying Capon beamforming. Finally, a low-complexity adaptive version of the BSS algorithm is proposed that can track changes in the mixing environment. Extensive experiments with realistic and measured data corroborate our claims, including the under-determined case. Signal-to-interference ratio improvements of up to 6 dB are shown compared to state-of-the-art BSS algorithms, at an order of magnitude lower computational complexity.
Dimitri Nion, Kleanthis N. Mokios, Nicholas D. Sidiropoulos, Alexandros Potamianos
IEEE Trans. Speech Audio Process.4
2010 Unsupervised Semantic Similarity Computation between Terms Using Web Documents
abstract
In this work, Web-based metrics that compute the semantic similarity between words or terms are presented and compared with the state of the art. Starting from the fundamental assumption that similarity of context implies similarity of meaning, relevant Web documents are downloaded via a Web search engine and the contextual information of words of interest is compared (context-based similarity metrics). The proposed algorithms work automatically, do not require any human-annotated knowledge resources, e.g., ontologies, and can be generalized and applied to different languages. Context-based metrics are evaluated both on the Charles-Miller data set and on a medical term data set. It is shown that context-based similarity metrics significantly outperform co-occurrence-based metrics, in terms of correlation with human judgment, for both tasks. In addition, the proposed unsupervised context-based similarity computation algorithms are shown to be competitive with the state-of-the-art supervised semantic similarity algorithms that employ language-specific knowledge resources. Specifically, context-based metrics achieve correlation scores of up to 0.88 and 0.74 for the Charles-Miller and medical data sets, respectively. The effect of stop word filtering is also investigated for word and term similarity computation. Finally, the performance of context-based term similarity metrics is evaluated as a function of the number of Web documents used and for various feature weighting schemes.
Elias Iosif, Alexandros Potamianos
IEEE Trans. Knowl. Data Eng.2
2009 Transition features for CRF-based speech recognition and boundary detection
abstract
In this paper, we investigate a variety of spectral and time domain features for explicitly modeling phonetic transitions in speech recognition. Specifically, spectral and energy distance metrics, as well as, time derivatives of phonological descriptors and MFCCs are employed. The features are integrated in an extended Conditional Random Fields statistical modeling framework that supports general-purpose transition models. For evaluation purposes, we measure both phonetic recognition task accuracy and precision/recall of boundary detection. Results show that when transition features are used in a CRF-based recognition framework, recognition performance improves significantly due to the reduction of phone deletions. The boundary detection performance also improves mainly for transitions among silence, stop, and fricative phonetic classes.
Spiros Dimopoulos, Eric Fosler-Lussier, Alexandros Potamianos
ASRU4
2009 Short-time instantaneous frequency and bandwidth features for speech recognition
abstract
In this paper, we investigate the performance of modulation related features and normalized spectral moments for automatic speech recognition. We focus on the short-time averages of the amplitude weighted instantaneous frequencies and bandwidths, computed at each subband of a mel-spaced filterbank. Similar features have been proposed in previous studies, and have been successfully combined with MFCCs for speech and speaker recognition. Our goal is to investigate the stand-alone performance of these features. First, it is experimentally shown that the proposed features are only moderately correlated in the frequency domain, and, unlike MFCCs, they do not require a transformation to the cepstral domain. Next, the filterbank parameters (number of filters and filter overlap) are investigated for the proposed features and compared with those of MFCCs. Results show that frequency related features perform at least as well as MFCCs for clean conditions, and yield superior results for noisy conditions; up to 50% relative error rate reduction for the AURORA3 Spanish task.
Pirros Tsiakoulis, Alexandros Potamianos, Dimitrios Dimitriadis
ASRU2
2009 Multiple time resolution analysis of speech signal using MCE training with application to speech recognition
abstract
In this paper, we propose two methods of multiple time-resolution analysis of speech and their application to automatic speech recognition (ASR). Constant frame-rate multi-scale analysis is proposed based on a box of multi-scale features. Then a variable rate analysis is proposed based on the selection of the optimal temporal resolution on the fly by a properly trained non-linear classifier unit. The classifier's parameters are trained using the discriminative method of minimum classification error (MCE) training. We use the recently proposed conditional random fields (CRF) phonetic recognition system that effectively combines highly correlated features. Results are reported on a frame-wise classification task and also on TIMIT phone recognition task. Results show that (i) CRFs can effectively combine multi-scale features and (ii) MCE trained variable rate CRFs are competitive with the ldquoboxrdquo combination method.
Spiros Dimopoulos, Alexandros Potamianos, Eric Fosler-Lussier
ICASSP2
2009 Video event detection and summarization using audio, visual and text saliency
abstract
Detection of perceptually important video events is formulated here on the basis of saliency models for the audio, visual and textual information conveyed in a video stream. Audio saliency is assessed by cues that quantify multifrequency waveform modulations, extracted through nonlinear operators and energy tracking. Visual saliency is measured through a spatiotemporal attention model driven by intensity, color and motion. Text saliency is extracted from part-of-speech tagging on the subtitles information available with most movie distributions. The various modality curves are integrated in a single attention curve, where the presence of an event may be signified in one or multiple domains. This multimodal saliency curve is the basis of a bottom-up video summarization algorithm, that refines results from unimodal or audiovisual-based skimming. The algorithm performs favorably for video summarization in terms of informativeness and enjoyability.
Georgios Evangelopoulos, Athanasia Zlatintsi, Georgios Skoumas, Konstantinos Rapantzikos, Alexandros Potamianos, Petros Maragos, Yannis Avrithis
ICASSP5
2009 Statistical analysis of amplitude modulation in speech signals using an AM-FM model
abstract
Several studies have been dedicated to the analysis and modeling of AM-FM modulations in speech and different algorithms have been proposed for the exploitation of modulations in speech applications. This paper details a statistical analysis of amplitude modulations using a multiband AM-FM analysis framework. The aim of this study is to analyze the phonetic- and speaker-dependency of modulations in the amplitude envelope of speech resonances. The analysis focuses on the dependence of such modulations on acoustic features such as, fundamental frequency, formant proximity, phone identity, as well as, speaker identity and contextual features. The results show that the amplitude modulation index of a speech resonance is mainly a function of the speaker-s average fundamental frequency, the phone identity, and the proximity between neighboring formant resonances. The results are especially relevant for speech and speaker recognition application employing modulation features.
Pirros Tsiakoulis, Alexandros Potamianos
ICASSP2
2009 Towards adapting fantasy, curiosity and challenge in multimodal dialogue systems for preschoolers
abstract
We investigate how fantasy, curiosity and challenge contribute to the user experience in multimodal dialogue computer games for preschool children. For this purpose, an on-line multimodal platform has been designed, implemented and used as a starting point to develop web-based speech-enabled applications for children. Five task oriented games suitable for preschoolers have been implemented with varying levels of fantasy and curiosity elements, as well as, variable difficulty levels. Nine preschool children, ages 4-6, were asked to play these games in three sessions; in each session only one of the fantasy, curiosity or challenge factor was evaluated. Both objective and subjective criteria were used to evaluate the factors and applications. Results show that fantasy and curiosity are correlated with children's entertainment, while the level of difficulty seems to depend on each child's individual preferences and capabilities. In addition, high speech usage and high curiosity levels in the application correlate well with task completion, showing that preschoolers become more engaged when multimodal interfaces are speech enabled and contain curiosity elements.
Theofanis Kannetis, Alexandros Potamianos
ICMI2
2009 Unsupervised Stream-Weights Computation in Classification and Recognition Tasks
abstract
In this paper, we provide theoretical results on the problem of optimal stream weight selection for the two stream classification problem. It is shown that in the presence of estimation or modeling errors using stream weights can decrease the total classification error. Specifically, we show that stream weights should be selected to be proportional to the feature stream reliability and informativeness. Next, we turn our attention to the problem of unsupervised stream weights computation in real tasks. Based on the theoretical results we propose to use models and ldquoanti-modelsrdquo (class-specific background models) to estimate stream weights. A nonlinear function of the ratio of the inter- to intra-class distance is proposed for stream weight estimation. The resulting unsupervised stream weight estimation algorithm is evaluated on both artificial data and on the problem of audiovisual speech classification. Finally, the proposed algorithm is extended to the problem of audiovisual speech recognition. It is shown that the proposed algorithms achieve results comparable to the supervised minimum-error training approach for classification tasks under most testing conditions.
Eduardo Sánchez-Soto, Alexandros Potamianos, Khalid Daoudi
IEEE Trans. Speech Audio Process.2
2008 On the effectiveness of PARAFAC-based estimation for blind speech separation
abstract
This work establishes the effectiveness of parallel factor (PARAFAC) analysis in blind speech separation (BSS) problems. The BSS problem is formulated as a conjugate-symmetric PARAFAC model that is fitted optimally, using an efficient alternating least-squares algorithm that converges monotonically. The identifiability properties of the model are also presented, revealing the much broader identifiability potential of joint-diagonalization- based BSS methods. In order to focus on estimation performance, perfect resolution of the permutation ambiguity is assumed. Simulations under varying reverberation conditions and comparison with previous estimation methods that are widely used in BSS problems demonstrate significant performance gains. Signal-to- interference (SIR) ratio improvement of over 27 dB is achieved using PARAFAC. Average SIR gains of 2.5 and 6.3 dB are achieved compared to state-of-the-art FastICA[2] and FDSOS (Parra's)[5] estimation algorithms, respectively.
Kleanthis N. Mokios, Alexandros Potamianos, Nicholas D. Sidiropoulos
ICASSP2
2008 Movie summarization based on audiovisual saliency detection
abstract
Based on perceptual and computational attention modeling studies, we formulate measures of saliency for an audiovisual stream. Audio saliency is captured by signal modulations and related multi-frequency band features, extracted through nonlinear operators and energy tracking. Visual saliency is measured by means of a spatiotemporal attention model driven by various feature cues (intensity, color, motion). Audio and video curves are integrated in a single attention curve, where events may be enhanced, suppressed or vanished. The presence of salient events is signified on this audiovisual curve by geometrical features such as local extrema, sharp transition points and level sets. An audiovisual saliency-based movie summarization algorithm is proposed and evaluated. The algorithm is shown to perform very well in terms of summary informativeness and enjoyability for movie clips of various genres.
Georgios Evangelopoulos, Konstantinos Rapantzikos, Alexandros Potamianos, Petros Maragos, Nancy Zlatintsi, Yannis Avrithis
ICIP3
2008 Multimodal system evaluation using modality efficiency and synergy metrics
abstract
In this paper, we propose two new objective metrics, relative modality efficiency and multimodal synergy, that can provide valuable information and identify usability problems during the evaluation of multimodal systems. Relative modality efficiency (when compared with modality usage) can identify suboptimal use of modalities due to poor interface design or information asymmetries. Multimodal synergy measures the added value from efficiently combining multiple input modalities, and can be used as a single measure of the quality of modality fusion and fission in a multimodal system. The proposed metrics are used to evaluate two multimodal systems that combine pen/speech and mouse/keyboard modalities respectively. The results provide much insight into multimodal interface usability issues, and demonstrate how multimodal systems should adapt to maximize modalities synergy resulting in efficient, natural, and intelligent multimodal interfaces.
Manolis Perakakis, Alexandros Potamianos
ICMI2
2008 Region-based vocal tract length normalization for ASR
abstract
In this paper, we propose a Region-based multi-parametric Vocal Tract Length Normalization (R-VTLN) algorithm for the problem of automatic speech recognition (ASR). The proposed algorithm extends the well-established mono-parametric utterance-based VTLN algorithm of Lee and Rose [1] by dividing the speech frames of a test utterance into regions and by warping independently the features corresponding to each region using a maximum likelihood criterion. We propose two algorithms for classifying frames into regions: (i) an unsupervised clustering algorithm based on spectral distance, and (ii) an unsupervised algorithm assigning frames to regions based on phonetic-class labels obtained from the first recognition pass. We also investigate the ability of various mono-parametric and multiparametric warping functions to reduce the spectral distance between two speakers, as a function of phone. R-VTLN is shown to significantly outperform mono-parametric VTLN in terms of word accuracy for the AURORA4 database.
Michail G. Maragakis, Alexandros Potamianos
INTERSPEECH2
2008 A Study in Efficiency and Modality Usage in Multimodal Form Filling Systems
abstract
The usage patterns of speech and visual input modes are investigated as a function of relative input mode efficiency for both desktop and personal digital assistant (PDA) working environments. For this purpose the form-filling part of a multimodal dialogue system is implemented and evaluated; three multimodal modes of interaction are implemented: ldquoClick-to-Talk,rdquo ldquoOpen-Mike,rdquo and ldquoModality-Selection.rdquo ldquoModality-Selectionrdquo implements an adaptive interface where the system selects the most efficient input mode at each turn, effectively alternating between a ldquoClick-to-Talkrdquo and ldquoOpen-Mikerdquo interaction style as proposed in ldquoModality tracking in the multimodal Bell Labs Communicator,rdquo in Proceedings of the Automatic Speech Recognition and Understanding Workshop, by A. Potamianos, , 2003. The multimodal systems are evaluated and compared with the unimodal systems. Objective and subjective measures used include task completion, task duration, turn duration, and overall user satisfaction. Turn duration is broken down into interaction time and inactivity time to better measure the efficiency of each input mode. Duration statistics and empirical probability density functions are computed as a function of interaction context and user. Results show that the multimodal systems outperform the unimodal systems in terms of objective and subjective criteria. Also, users tend to use the most efficient input mode at each turn; however, biases towards the default input modality and a general bias towards the speech modality also exists. Results demonstrate that although users exploit some of the available synergies in multimodal dialogue interaction, further efficiency gains can be achieved by designing adaptive interfaces that fully exploit these synergies.
Manolis Perakakis, Alexandros Potamianos
IEEE Trans. Speech Audio Process.2
2007 Unsupervised Stream Weight Estimation using Anti-Models
abstract
In this paper, a novel solution to the problem of unsupervised stream weight estimation for multi-stream classification tasks is proposed. Our work is based on theoretical results in A. Potamianos et al. (2006) for the two-class problem were the optimal stream weights are shown to be inversely proportional to the single stream misclassification error. These two-class results are applied to the multi-class problem by using models and "anti-models" (class-specific background models) thus posing the multi-class problem as multiple two-class problems. A nonlinear function of the ratio of the inter- to intra-class distance is proposed as an estimate for single stream classification error and used for stream weight estimation. The proposed unsupervised stream weight estimation algorithm is evaluated on both artificial data and on the problem of audio-visual speech recognition. It is shown that the proposed algorithm achieves results comparable to the supervised minimum-error training approach under most testing conditions.
Eduardo Sánchez-Soto, Alexandros Potamianos, Khalid Daoudi
ICASSP (4)2
2007 The effect of input mode on inactivity and interaction times of multimodal systems
abstract
In this paper, the efficiency and usage patterns of input modes in multimodal dialogue systems is investigated for both desktop and personal digital assistant (PDA) working environments. For this purpose a form-filling travel reservation application is evaluated that combines the speech and visual modalities; three multimodal modes of interaction are implemented, namely: "Click-To-Talk", "Open-Mike" and "Modality-Selection". The three multimodal systems are evaluated and compared with the "GUI-Only" and "Speech-Only" unimodal systems. Mode and duration statistics are computed for each system, for each turn and for each attribute in the form. Turn time is decomposed in interaction and inactivity time and the statistics for each input modeare computed. Results show that multimodal and adaptive interfaces are superior in terms of interaction time, but not always in terms of inactivity time. Also users tend to use themost efficient input mode, although our experiments show abias towards the speech modality.
Manolis Perakakis, Alexandros Potamianos
ICMI2
2007 Advanced front-end for robust speech recognition in extremely adverse environments
abstract
In this paper, a unified approach to speech enhancement, feature extraction and feature normalization for speech recognition in adverse recording conditions is presented. The proposed frontend system consists of several different, independent, processing modules. Each of the algorithms contained in these modules has been independently applied to the problem of speech recognition in noise, significantly improving the recognition rates. In this work, these algorithms are merged in a single front-end and their combined performance is demonstrated. Specifically, the proposed advanced front-end extracts noise-invariant features via the following modules: Wiener filtering, voice-activity detection, robust feature extraction (nonlinear modulation or fractal features), parameter equalization and frame-dropping. The advanced front-end is applied to extremely adverse environments where most feature extraction schemes fail. We show that by combining speech enhancement, robust feature extraction and feature normalization up to a fivefold error rate reduction can be achieved for certain tasks.
Dimitrios Dimitriadis, José C. Segura, Luz García 0001, Alexandros Potamianos, Petros Maragos, Vassilis Pitsikalis
INTERSPEECH4
2007 A soft-clustering algorithm for automatic induction of semantic classes
abstract
In this paper, we propose a soft-decision, unsupervised clus-tering algorithm that generates semantic classes automatically using the probability of class membership for each word, rather than deterministically assigning a word to a semantic class. Se-mantic classes are induced using an unsupervised, automatic procedure that uses a context-based similarity distance to mea-sure semantic similarity between words. The proposed soft-decision algorithm is compared with various “hard ” clustering algorithms, e.g., [1], and it is shown to improve semantic class induction performance in terms of both precision and recall for a travel reservation corpus. It is also shown that additional perfor-mance improvement is achieved by combining (auto-induced) semantic with lexical information to derive the semantic simi-larity distance. Index Terms: semantic classes, unsupervised clustering 1.
Elias Iosif, Alexandros Potamianos
INTERSPEECH2
2007 A review of the acoustic and linguistic properties of children's speech
abstract
In this paper, we review the acoustic and linguistic properties of children's speech for both read and spontaneous speech. First, the effect of developmental changes on the absolute values and variability of acoustic correlates is presented for read speech for children ages 6 and up. Then, verbal child-machine spontaneous interaction is reviewed and results from recent studies are presented. Age trends of acoustic, linguistic and interaction parameters are discussed, such as sentence duration, filled pauses, politeness and frustration markers, and modality usage. Some differences between child-machine and human-human interaction are pointed out. The implications for acoustic modeling, linguistic modeling and spoken dialogue systems design for children are discussed.
Alexandros Potamianos, Shri Narayanan
MMSP1
2007 Multimodal User Interface for Augmented Assembly
abstract
In this paper, a multimodal system for augmented reality aided assembly work is designed and implemented. The multimodal interface allows for speech and gestural input. The system emulates a simplified assembly task in a factory. A 3D puzzle is used to study how to implement the augmented assembly system to a real setting in a factory. The system is used as a demonstrator and as a test-bed to evaluate different input modalities for augmented assembly setups. Preliminary system evaluation results are presented, the user experience is discussed, and some directions for future work are given.
Sanni Siltanen, Mika Hakkarainen, Otto Korkalo, Tapio Salonen, Juha Sääski, Charles Woodward, Theofanis Kannetis, Manolis Perakakis, Alexandros Potamianos
MMSP9
2007 Unsupervised Semantic Similarity Computation usingWeb Search Engines
abstract
In this paper, we propose two novel web-based metrics for semantic similarity computation between words. Both metrics use a web search engine in order to exploit the retrieved information for the words of interest. The first metric considers only the page counts returned by a search engine, based on the work of [1]. The second downloads a number of the top ranked documents and applies "widecontext" and "narrow-context" metrics. The proposed metrics work automatically, without consulting any human annotated knowledge resource. The metrics are compared with WordNet-based methods. The metrics' performance is evaluated in terms of correlation with respect to the pairs of the commonly used Charles - Miller dataset. The proposed "wide-context" metric achieves 71% correlation, which is the highest score achieved among the fully unsupervised metrics in the literature up to date.
Elias Iosif, Alexandros Potamianos
Web Intelligence2
2007 Information Seeking Spoken Dialogue Systems- Part I: Semantics and Pragmatics
abstract
In this paper, the semantic and pragmatic modules of a spoken dialogue system development platform are presented and evaluated. The main goal of this research is to create spoken dialogue system modules that are portable across applications domains and interaction modalities. We propose a hierarchical semantic representation that encodes all information supplied by the user over multiple dialogue turns and can efficiently represent and be used to argue with ambiguous or conflicting information. Implicit in this semantic representation is a pragmatic module, consisting of context tracking, pragmatic analysis and pragmatic scoring submodules, which computes pragmatic confidence scores for all system beliefs. These pragmatic scores are obtained by combining semantic and pragmatic evidence from the various sub-modules (taking into account the modality of input) and are used to rank-order attribute-value pairs in the semantic representation, as well as identifying and resolving ambiguities. These modules were implemented and evaluated within a travel reservation dialogue system under the auspices of the DARPA Communicator project, as well as for a movie information application. Formal evaluation of the semantic and pragmatic modules has shown that by incorporating pragmatic analysis and scoring, the quality of the system improves for over 20% of the dialogue fragments examined
Egbert Ammicht, Eric Fosler-Lussier, Alexandros Potamianos
IEEE Trans. Multim.3
2007 Information Seeking Spoken Dialogue Systems- Part II: Multimodal Dialogue
abstract
For pt.1see ibid., vol. 9, p. 3 (2007). In this paper, the task and user interface modules of a multimodal dialogue system development platform are presented. The main goal of this work is to provide a simple, application-independent solution to the problem of multimodal dialogue design for information seeking applications. The proposed system architecture clearly separates the task and interface components of the system. A task manager is designed and implemented that consists of two main submodules: the electronic form module that handles the list of attributes that have to be instantiated by the user, and the agenda module that contains the sequence of user and system tasks. Both the electronic forms and the agenda can be dynamically updated by the user. Next a spoken dialogue module is designed that implements the speech interface for the task manager. The dialogue manager can handle complex error correction and clarification user input, building on the semantics and pragmatic modules presented in Part I of this paper. The spoken dialogue system is evaluated for a travel reservation task of the DARPA Communicator research program and shown to yield over 90% task completion and good performance for both objective and subjective evaluation metrics. Finally, a multimodal dialogue system which combines graphical and speech interfaces, is designed, implemented and evaluated. Minor modifications to the unimodal semantic and pragmatic modules were required to build the multimodal system. It is shown that the multimodal system significantly outperforms the unimodal speech-only system both in terms of efficiency (task success and time to completion) and user satisfaction for a travel reservation task
Alexandros Potamianos, Eric Fosler-Lussier, Egbert Ammicht, Manolis Perakakis
IEEE Trans. Multim.1
2006 Blind Speech Separation Using Parafac Analysis and Integer Least Squares
abstract
We propose a new two-step frequency domain algorithm for blind speech separation (BSS) for unknown channel order. This new approach employs parallel factor analysis (PARAFAC) to separate the speech signals and a novel integer-least-squares-based method for matching the arbitrary permutations in the frequency domain. The proposed algorithm offers guaranteed convergence and good separation performance, measured both quantitatively and in subjective tests. Performance gains in signal-to-interference ratio of up to 10 db are achieved for certain source-sensor geometries
Kleanthis N. Mokios, Nicholas D. Sidiropoulos, Alexandros Potamianos
ICASSP (5)3
2006 Stream Weight Computation for Multi-Stream Classifiers
abstract
In this paper, we provide theoretical results on the problem of optimal stream weight selection for the multi-stream classification problem. It is shown, that in the presence of estimation or modeling errors using stream weights can decrease the total classification error. The stream weights that minimize classification estimation error are shown to be inversely proportional to the single-stream pdf estimation error. It is also shown that under certain conditions, the optimal stream weights are inversely proportional to the single-stream classification error. We apply these results to the problem of audio-visual speech recognition and experimentally verify our claims. The applicability of the results to the problem of unsupervised stream weight estimation is also discussed
Alexandros Potamianos, Eduardo Sánchez-Soto, Khalid Daoudi
ICASSP (1)1
2006 Unsupervised Combination of Metrics for Semantic Class Induction
abstract
In this paper, unsupervised algorithms for combining semantic similarity metrics are proposed for the problem of automatic class induction. The automatic class induction algorithm is based on the work of Pargellis et al,. The semantic similarity metrics that are evaluated and combined are based on narrow- and wide-context vector- product similarity. The metrics are combined using linear weights that are computed 'on the fly' and are updated at each iteration of the class induction algorithm, forming a corpus-independent metric. Specifically, the weight of each metric is selected to be inversely proportional to the inter-class similarity of the classes induced by that metric and for the current iteration of the algorithm. The proposed algorithms are evaluated on two corpora: a semantically heterogeneous news domain (HR-Net) and an application-specific travel reservation corpus (ATIS). It is shown, that the (unsupervised) adaptive weighting scheme outperforms the (supervised) fixed weighting scheme. Up to 50% relative error reduction is achieved by the adaptive weighting scheme.
Elias Iosif, Athanasios Tegos, Apostolos Pangos, Eric Fosler-Lussier, Alexandros Potamianos
SLT5
2006 Blending speech and Visual Input in Multimodal Dialogue Systems
abstract
In this paper the efficiency and usage patterns of input modes in multimodal dialogue systems is investigated for desktop and personal digital assistant (PDA) working environments. For this purpose a form-filling travel reservation system is designed and implemented that efficiently combines the speech and visual modalities; three multimodal modes of interaction are implemented, namely: "click-to-talk", "open-mike" and "modality-selection". The three multimodal systems are evaluated and compared with the "GUI-only" and "speech-only" unimodal systems. User interface evaluation includes both objective and subjective metrics and shows that all three multimodal systems outperform the unimodal systems on the PDA environment. For the desktop environment the multimodal systems score better than the "speech-only" system but worse than the "GUI-only" system. In all evaluation experiments, the synergy between the visual and speech modality was significant: the multimodal interface was better than the sum of its (unimodal) parts. Results also show that users tend to use the most efficient input mode.
Manolis Perakakis, Michail Toutoudakis, Alexandros Potamianos
SLT3
2005 Auditory Teager energy cepstrum coefficients for robust speech recognition
abstract
In this paper, a feature extraction algorithm for robust speech recognition is introduced. The feature extraction algorithm is motivated by the human auditory processing and the nonlinear Teager-Kaiser energy operator that estimates the true energy of the source of a resonance. The proposed features are labeled as Teager Energy Cepstrum Coefficients (TECCs). TECCs are computed by first filtering the speech signal through a dense non constant-Q Gammatone filterbank and then by estimating the "true" energy of the signal's source, i.e., the short-time average of the output of the Teager-Kaiser energy operator. Error analysis and speech recognition experiments show that the TECCs and the mel frequency cepstrum coefficients (MFCCs) perform similarly for clean recording conditions; while the TECCs perform significantly better than the MFCCs for noisy recognition tasks. Specifically, relative word error rate improvement of 60% over the MFCC baseline is shown for the Aurora-3 database for the high-mismatch condition. Absolute error rate improvement ranging from 5% to 20% is shown for a phone recognition task in (various types of additive) noise.
Dimitrios Dimitriadis, Petros Maragos, Alexandros Potamianos
INTERSPEECH3
2005 Detecting Politeness and frustration state of a child in a conversational computer game
abstract
In this study, we investigate politeness and frustration behavior of children during their spoken interaction with computer characters in a game. We focus on automatically detecting frustrated, polite and neutral attitudes from the child’s speech (acoustic and language) communication cues and study their differences as a function of age and gender. The study is based on a Wizard-of-Oz dialog corpus of 103 children playing a voice activated computer game. Statistical analysis revealed that there was a significant gender effect on politeness with girls in this data exhibiting more explicit politeness markers. The analysis also showed that there is a positive correlation between frustration and the number of dialog turns reflecting the fact that longer time spent solving the puzzle of the game led to a more frustrated child. By combining acoustic and language cues for the task of automatic detection of politeness and frustration, we obtain average accuracy of 84.7 % and 71.3%, respectively, by using age dependent models and 85 % and 72%, respectively, for gender dependent models. 1.
Serdar Yildirim, Chul Min Lee, Sungbok Lee, Alexandros Potamianos, Shri Narayanan
INTERSPEECH4
2005 Robust AM-FM Features for Speech Recognition
abstract
In this letter, a nonlinear AM-FM speech model is used to extract robust features for speech recognition. The proposed features measure the amount of amplitude and frequency modulation that exists in speech resonances and attempt to model aspects of the speech acoustic information that the commonly used linear source-filter model fails to capture. The robustness and discriminability of the AM-FM features is investigated in combination with mel cepstrum coefficients (MFCCs). It is shown that these hybrid features perform well in the presence of noise, both in terms of phoneme-discrimination (J-measure) and in terms of speech recognition performance in several different tasks. Average relative error rate reduction up to 11% for clean and 46% for mismatched noisy conditions is achieved when AM-FM features are combined with MFCCs.
Dimitrios Dimitriadis, Petros Maragos, Alexandros Potamianos
IEEE Signal Process. Lett.3
2005 Adaptive categorical understanding for spoken dialogue systems
abstract
In this paper, the speech understanding problem in the context of a spoken dialogue system is formalized in a maximum likelihood framework. Off-line adaptation of stochastic language models that interpolate dialogue state specific and general application-level language models is proposed. Word and dialogue-state n-grams are used for building categorical understanding and dialogue models, respectively. Acoustic confidence scores are incorporated in the understanding formulation. Problems due to data sparseness and out-of-vocabulary words are discussed. The performance of the speech recognition and understanding language models are evaluated with the "Carmen Sandiego" multimodal computer game corpus. Incorporating dialogue models reduces relative understanding error rate by 15%-25%, while acoustic confidence scores achieve a further 10% error reduction for this computer gaming application.
Alexandros Potamianos, Shri Narayanan, Giuseppe Riccardi
IEEE Trans. Speech Audio Process.1
2004 Auto-induced semantic classes
Andrew N. Pargellis, Eric Fosler-Lussier, Alexandros Potamianos, Augustine Tsai
Speech Commun.4
2003 Robust recognition of children's speech
abstract
Developmental changes in speech production introduce age-dependent spectral and temporal variability in the speech signal produced by children. Such variabilities pose challenges for robust automatic recognition of children's speech. Through an analysis of age-related acoustic characteristics of children's speech in the context of automatic speech recognition (ASR), effects such as frequency scaling of spectral envelope parameters are demonstrated. Recognition experiments using acoustic models trained from adult speech and tested against speech from children of various ages clearly show performance degradation with decreasing age. On average, the word error rates are two to five times worse for children speech than for adult speech. Various techniques for improving ASR performance on children's speech are reported. A speaker normalization algorithm that combines frequency warping and model transformation is shown to reduce acoustic variability and significantly improve ASR performance for children speakers (by 25-45% under various model training and testing conditions). The use of age-dependent acoustic models further reduces word error rate by 10%. The potential of using piece-wise linear and phoneme-dependent frequency warping algorithms for reducing the variability in the acoustic feature space of children is also investigated.
Alexandros Potamianos, Shri Narayanan
IEEE Trans. Speech Audio Process.1
2002 Modulation features for speech recognition
abstract
Automatic speech recognition (ASR) systems can benefit from including into their acoustic processing part new features that account for various nonlinear and time-varying phenomena during speech production. In this paper, we develop robust methods to extract novel acoustic features from speech signals of the modulation type based on time-varying models for speech analysis. Further, we integrate the new speech features with the standard linear ones (mel-frequency cesptrum) to develop a augmented set of acoustic features and demonstrate its efficacy by showing significant improvements in HMM-based word recognition over the TIMIT database.
Dimitrios Dimitriadis, Petros Maragos, Alexandros Potamianos
ICASSP3
2002 Adaptive language models for spoken dialogue systems
abstract
In this paper, we investigate both generative and statistical approaches for language modeling in spoken dialogue systems. Semantic class-based finite state and n-gram grammars are used for improving coverage and modeling accuracy when little training data is available. We have implemented dialogue-state specific language model adaptation to reduce perplexity and improve the efficiency of grammars for spoken dialogue systems. A novel algorithm for combining state-independent n-gram and state-dependent finite state grammars using acoustic confidence scores is proposed. Using this combination strategy, a relative word error reduction of 12% is achieved for certain dialogue states within a travel reservation task. Finally, semantic class multigrams are proposed and briefly evaluated for language modeling in dialogue systems.
Roger Argiles Solsona, Eric Fosler-Lussier, Hong-Kwang Jeff Kuo, Alexandros Potamianos, Imed Zitouni
ICASSP4
2002 DARPA communicator evaluation: progress from 2000 to 2001
abstract
This paper describes the evaluation methodology and results of the DARPA Communicator spoken dialog system evaluation experiments in 2000 and 2001. Nine spoken dialog systems in the travel planning domain participated in the experiments resulting in a total corpus of 1904 dialogs. We describe and compare the experimental design of the 2000 and 2001 DARPA evaluations. We describe how we established a performance baseline in 2001 for complex tasks. We present our overall approach to data collection, the metrics collected, and the application of PARADISE to these data sets. We compare the results we achieved in 2000 for a number of core metrics with those for 2001. These results demonstrate large performance improvements from 2000 to 2001 and show that the Communicator program goal of conversational interaction for complex tasks has been achieved.
Marilyn A. Walker, Alexander I. Rudnicky, John S. Aberdeen, Elizabeth Owen Bratt, John S. Garofolo, Helen Hastie, Audrey N. Le, Bryan L. Pellom, Alexandros Potamianos, Rebecca J. Passonneau, Rashmi Prasad, Salim Roukos, Gregory A. Sanders, Stephanie Seneff, David Stallard
INTERSPEECH9
2002 DARPA communicator: cross-system results for the 2001 evaluation
abstract
This paper describes the evaluation methodology and results of the 2001 DARPA Communicator evaluation. The experiment spanned 6 months of 2001 and involved eight DARPA Communicator systems in the travel planning domain. It resulted in a corpus of 1242 dialogs which include many more dialogues for complex tasks than the 2000 evaluation. We describe the experimental design, the approach to data collection, and the results. We compare the results by the type of travel plan and by system. The results demonstrate some large differences across sites and show that the complex trips are clearly more difficult.
Marilyn A. Walker, Alexander I. Rudnicky, Rashmi Prasad, John S. Aberdeen, Elizabeth Owen Bratt, John S. Garofolo, Helen Hastie, Audrey N. Le, Bryan L. Pellom, Alexandros Potamianos, Rebecca J. Passonneau, Salim Roukos, Gregory A. Sanders, Stephanie Seneff, David Stallard
INTERSPEECH10
2002 Creating conversational interfaces for children
abstract
Creating conversational interfaces for children is challenging in several respects. These include acoustic modeling for automatic speech recognition (ASR), language and dialog modeling, and multimodal-multimedia user interface design. First, issues in ASR of children's speech are introduced by an analysis of developmental changes in the spectral and temporal characteristics of the speech signal using data obtained from 456 children, ages five to 18 years. Acoustic modeling adaptation and vocal tract normalization algorithms that yielded state-of-the-art ASR performance on children's speech are described. Second, an experiment designed to better understand how children interact with machines using spoken language is described. Realistic conversational multimedia interaction data were obtained from 160 children who played a voice-activated computer game in a Wizard of Oz (WoZ) scenario. Results of using these data in developing novel language and dialog models as well as in a unified maximum likelihood framework for acoustic decoding in ASR and semantic classification for spoken language understanding are described. Leveraging the lessons learned from the WoZ study and a concurrent user experience evaluation, a multimedia personal agent prototype for children was designed. Details of the architecture and application details are described. Informal evaluation by children was found positive especially for the animated agent and the speech interface.
Shri Narayanan, Alexandros Potamianos
IEEE Trans. Speech Audio Process.2
2002 An error-protected speech recognition system for wireless communications
abstract
Future wireless multimedia terminals will have a variety of applications that require speech recognition capabilities. We consider a robust distributed speech recognition system where representative parameters of the speech signal are extracted at the wireless terminal and transmitted to a centralized automatic speech recognition (ASR) server. We propose two unequal error protection schemes for the ASR bit stream and demonstrate the satisfactory performance of these schemes for typical wireless cellular channels. In addition, a "soft-feature" error concealment strategy is introduced at the ASR server that uses "soft-outputs" from the channel decoder to compute the marginal distribution of only the reliable features during likelihood computation at the speech recognizer. This soft-feature error concealment technique reduces the ASR error rate by more than a factor of 2.5 for certain channels. Also considered is a channel decoding technique with source information that improves ASR performance.
Vijitha Weerackody, Wolfgang Reichl, Alexandros Potamianos
IEEE Trans. Wirel. Commun.3
2001 Soft-feature decoding for speech recognition over wireless channels
abstract
A distributed automatic speech recognition (ASR) system is considered where features of the speech signal are extracted at the wireless terminal and transmitted to a centralized ASR server. An unequal error protection scheme is used for the quantized ASR feature stream. At the receiver, coherent demodulation is performed and the probability of error for each bit is computed using the max-log MAP algorithm. A 'soft-feature' decoding strategy is introduced at the ASR server that uses the marginal distribution of only the reliable features during likelihood computation. Alternatively, the confidence of each feature is computed from the bit error probabilities and each feature in the probability computation is weighted as a function of the feature confidence. The performance of the proposed soft-feature algorithms is evaluated over typical cellular wireless channels and it is shown to reduce ASR error rate by over 50% for certain channels at a small additional computational cost.
Alexandros Potamianos, Vijitha Weerackody
ICASSP1
2001 Speech recognition for wireless applications
abstract
Future wireless multimedia terminals will have a variety of applications that require speech recognition capabilities. We consider a robust distributed speech recognition system where representative parameters of the speech signal are extracted at the wireless terminal and transmitted to a centralized automatic speech recognition (ASR) server. We propose several unequal error protection schemes for the ASR bit stream and demonstrate the satisfactory performance of these schemes for typical wireless cellular channels. In addition, a "soft-feature" error concealment strategy is introduced at the ASR server that uses "soft-outputs" from the channel decoder. This soft-feature error concealment techniques reduces the ASR error rate by up to four times for certain channels. Also considered is a channel decoding technique with source information that improves ASR performance.
Vijitha Weerackody, Wolfgang Reichl, Alexandros Potamianos
ICC3
2001 Ambiguity representation and resolution in spoken dialogue systems
abstract
Spoken natural language often contains ambiguities that must be addressed by a spoken dialogue system. In this work, we present the internal semantic representation and resolution strategy of a dialogue system designed to understand ambiguous input. These mechanisms are domain independent; almost all task-specific knowledge is represented in parameterizable data structures. The system derives candidate descriptions of what the user said from raw input data, context-tracking and scoring. These candidates are chosen on the basis of a pragmatic analysis of responses elicited by an extensive implicit and explicit confirmation dialogue strategy, combined with specific error correction capabilities available to the user. This new ambiguity resolution strategy greatly improves dialogue interaction, eliminating about half of the errors in dialogues from a travel reservation task.
Egbert Ammicht, Alexandros Potamianos, Eric Fosler-Lussier
INTERSPEECH2
2001 Hybrid natural language generation for spoken dialogue systems
abstract
The natural language generation component of most dialogue systems is based on templates. Template-based generators are hard to maintain and reuse, and the sentences they produce lack the variability and robustness needed by conversational systems. In this paper, a flexible and domain-independent natural language generator for spoken dialogue systems is proposed which combines fixed surface expressions with freely generated text. The generation algorithm follows a hybrid approach, combining finite state machine (FSM) grammars and corpus-based language models. In this approach, the FSM grammar (a reversible parser grammar) is constrained by a word and concept Ò-gram that takes terminals and non-terminal co-occurrences into account. The Ò-gram grammar helps prevent inappropriate derivations, therefore improving the quality of the generated texts. The proposed algorithm achieves faster than real-time performance because of the limited number of derivations.
Michel Galley, Eric Fosler-Lussier, Alexandros Potamianos
INTERSPEECH3
2001 Metrics for measuring domain independence of semantic classes
abstract
The design of dialogue systems for a new domain requires se-mantic classes (concepts) to be identified and defined. This process could be made easier by importing relevant concepts from previously studied domains to the new one. We pro-pose two methodologies, based on comparison of semantic classes across domains, for determining which concepts are domain-independent, and which are specific to the new task. The concept-comparison technique uses a context-dependent Kullback-Leibler distance measurement to compare all pairwise combinations of semantic classes, one from each domain. The concept-projection method uses a similar metric to project a sin-gle semantic class from one domain into the lexical environment of another. Initial results show that both methods are good in-dicators of the degree of domain independence for a wide range of concepts, manually generated for three different tasks: Car-men (children’s game), Movie (information retrieval) and Travel (flight reservations). 1.
Andrew N. Pargellis, Eric Fosler-Lussier, Alexandros Potamianos
INTERSPEECH3
2001 DARPA communicator dialog travel planning systems: the june 2000 data collection
abstract
This paper describes results of an experiment with 9 different DARPA Communicator Systems who participated in the June 2000 data collection. All systems supported travel planning and utilized some form of mixed-initiative interaction. However they varied in several critical dimensions: (1) They targeted different back-end databases for travel information; (2) The used different modules for ASR,NLU,TTS and dialog management. We describe the experimental design, the approach to data collection, the metrics collected, and results comparing the systems. 1.
Marilyn A. Walker, John S. Aberdeen, Julie E. Boland, Elizabeth Owen Bratt, John S. Garofolo, Lynette Hirschman, Audrey N. Le, Sungbok Lee, Shri Narayanan, Kishore Papineni, Bryan L. Pellom, Joseph Polifroni, Alexandros Potamianos, P. Prabhu, Alexander I. Rudnicky, Gregory A. Sanders, Stephanie Seneff, David Stallard, Steve Whittaker 0001
INTERSPEECH13
2001 Time-frequency distributions for automatic speech recognition
abstract
The use of general time-frequency distributions as features for automatic speech recognition (ASR) is discussed in the context of hidden Markov classifiers. Short-time averages of quadratic operators, e.g., energy spectrum, generalized first spectral moments, and short-time averages of the instantaneous frequency, are compared to the standard front end features, and applied to ASR. Theoretical and experimental results indicate a close relationship among these feature sets.
Alexandros Potamianos, Petros Maragos
IEEE Trans. Speech Audio Process.1
2000 Cross-domain classification using generalized domain acts
Andrew N. Pargellis, Alexandros Potamianos
INTERSPEECH2
2000 Dialogue management in the Bell Labs communicator system
Alexandros Potamianos, Egbert Ammicht, Hong-Kwang Jeff Kuo
INTERSPEECH1
2000 Statistical recursive finite state machine parsing for speech understanding
Alexandros Potamianos, Hong-Kwang Jeff Kuo
INTERSPEECH1
1999 Multimodal systems for children: building a prototype
Shri Narayanan, Alexandros Potamianos, Haohong Wang
EUROSPEECH2
1999 Speaker adaptation for audio-visual speech recognition
Gerasimos Potamianos, Alexandros Potamianos
EUROSPEECH2
1999 Categorical understanding using statistical ngram models
Alexandros Potamianos, Giuseppe Riccardi, Shri Narayanan
EUROSPEECH1
1999 Speech analysis and synthesis using an AM-FM modulation model
Alexandros Potamianos, Petros Maragos
Speech Commun.1
1998 Multi-band speech recognition in noisy environments
abstract
This paper presents a new approach for multi-band based automatic speech recognition (ASR). Previous work by Bourlard et al. (see Proc. Int. Conf. on Spoken Language Processing, Philadelphia, p.426-9, 1996) and Hermansky et al. (see Proc. Int. Conf. on Spoken Language Processing, Philadelphia, p.1579-82, 1996) suggests that multi-band ASR gives a more accurate recognition, especially in noisy acoustic environments, by combining the likelihoods of different frequency bands. Here we evaluate this likelihood recombination (LC) approach to multi-band ASR, and propose an alternative method, namely feature recombination (FC). In the FC system, after different acoustic analyzers are applied to each sub-band individually, a vector is composed by combining the sub-band features. The speech classifier then calculates the likelihood from the single vector. Thus, band-limited noise affects only a few of the feature components, as in the multi-band LC system, but, at the same time, all feature components are jointly modeled, as in conventional ASR. The experimental results show that the FC system can yield better performance than both the conventional ASR and the LC strategy for noisy speech.
Shigeki Okawa, Enrico Bocchieri, Alexandros Potamianos
ICASSP3
1998 Spoken dialog systems for children
abstract
We outline the main issues when designing interactive multimedia systems for children and propose a unified approach of acoustic, linguistic, and dialog modeling to system development. The acoustic, linguistic and dialog data collected in a Wizard of Oz experiment from 160 children ages 8-14 playing an interactive computer game are analyzed and children-specific modeling issues are presented. Age-dependent and modality-dependent dialog flow patterns are identified. Furthermore, extraneous speech patterns, linguistic variability and disfluencies are investigated in spontaneous children's speech, and important new results are reported. Finally, baseline automatic speech recognition results are presented for various tasks using simple acoustic and language models.
Alexandros Potamianos, Shri Narayanan
ICASSP1
1998 Language model adaptation for spoken language systems
abstract
1. ABSTRACT In a human-machine interaction (dialog) the statistical language variations are large among different stages of the dialog and across different speakers. Moreover, spoken dialog systems require extensive training data for training adaptive language models. In this paper we address the problem of open-vocabulary language models allowing the user for any possible response at each stage of the dialog. We propose a novel off-line adaptation of stochastic language models effective for their generalization (openvocabulary) and selective (dialog context) properties. We outline the integration of the finite state dialog model and the language model adaptation algorithm. The performance of the speech recognition and understanding language models are evaluated with the Carmen Sandiego multimodal computer game. The new language models give an overall understanding error rate reduction of 44% over the baseline system.
Giuseppe Riccardi, Alexandros Potamianos, Shri Narayanan
ICSLP2
1997 On combining frequency warping and spectral shaping in HMM based speech recognition
abstract
Frequency warping approaches to speaker normalization have been proposed and evaluated on various speech recognition tasks. These techniques have been found to significantly improve performance even for speaker independent recognition from short utterances over the telephone network. In maximum likelihood (ML) based model adaptation a linear transformation is estimated and applied to the model parameters in order to increase the likelihood of the input utterance. The purpose of this paper is to demonstrate that significant advantage can be gained by performing frequency warping and ML speaker adaptation in a unified framework. A procedure is described which compensates utterances by simultaneously scaling the frequency axis and reshaping the spectral energy contour. This procedure is shown to reduce the error rate in a telephone based connected digit recognition task by 30-40%.
Alexandros Potamianos, Richard C. Rose
ICASSP1
1997 Analysis of children's speech: duration, pitch and formants
Sungbok Lee, Alexandros Potamianos, Shri Narayanan
EUROSPEECH2
1997 On using fractal features of speech sounds in automatic speech recognition
Petros Maragos, Alexandros Potamianos
EUROSPEECH2
1997 Speech analysis and synthesis using an AM-FM modulation model
abstract
In this paper, the AM‐FM modulation model is applied to speech analysis, synthesis and coding. The AM‐FM model represents the speech signal as the sum of formant resonance signals each of which contains amplitude and frequency modulation. Multiband filtering and demodulation using the energy separation algorithm are the basic tools used for speech analysis. First, multiband demodulation analysis (MDA) is applied to the problem of fundamental frequency estimation using the average instantaneous frequency as estimates of pitch harmonics. The MDA pitch tracking algorithm is shown to produce smooth and accurate fundamental frequency contours. Next, the AM‐FM modulation vocoder is introduced, which represents speech as the sum of resonance signals. A time-varying filterbank is used to extract the formant bands and then the energy separation algorithm is used to demodulate the resonance signals into the amplitude envelope and instantaneous frequency signals. EAcient modeling and coding (at 4.8‐9.6 kbits/sec) algorithms are proposed for the amplitude envelope and instantaneous frequency of speech resonances. Finally, the perceptual importance of modulations in speech resonances is investigated and it is shown that amplitude modulation patterns are both speaker and phone dependent. ” 1999 Elsevier Science B.V. All rights reserved. Zusammenfassung
Alexandros Potamianos, Petros Maragos
EUROSPEECH1
1997 Automatic speech recognition for children
abstract
In this paper, the acoustic and linguistic characteristics of children speech are investigated in the context of automatic speech recognition. Acoustic variability is identi ed as a major hurdle in building high performance ASR applications for children. A simple speaker normalization algorithm combining frequency warping and spectral shaping introduced in [5] is shown to reduce acoustic variability and signi cantly improve recognition performance for children speakers (by 25{ 45%). Age-dependent acoustic modeling further reduces word error rate by 10%. Piece-wise linear and phoneme-dependent frequency warping algorithms are proposed for reducing acoustic mismatch between the children and adult acoustic spaces.
Alexandros Potamianos, Shri Narayanan, Sungbok Lee
EUROSPEECH1
1997 Unsupervised HMM adaptation based on speech-silence discrimination
Ilija Zeljkovic, Shri Narayanan, Alexandros Potamianos
EUROSPEECH3
1995 Speech formant frequency and bandwidth tracking using multiband energy demodulation
abstract
The AM-FM modulation model and a multiband analysis/demodulation scheme is applied to speech formant frequency and bandwidth tracking. Filtering is performed by a bank of Gabor bandpass filters. Each band is demodulated to amplitude envelope and instantaneous frequency signals using the energy separation algorithm. Short-time formant frequency and bandwidth estimates are obtained from the instantaneous amplitude and frequency signals and their merits are presented. The estimates are used to determine the formant locations and bandwidths. Performance and computational issues (frequency domain implementation) are discussed. Overall, the multiband demodulation approach to formant tracking is easy to implement, provides accurate formant frequency and realistic bandwidth estimates, and performs well in the presence of nasalization.
Alexandros Potamianos, Petros Maragos
ICASSP1
1995 A feature-space transformation for telephone based speech recognition
Alexandros Potamianos, Li Lee, Richard C. Rose
EUROSPEECH1
1995 Higher order differential energy operators
abstract
Instantaneous signal operators /spl Upsi//sub k/(x)=x/spl dot/x/sup (k-1)/-xx/sup (k)/ of integer orders k are proposed to measure the cross energy between a signal x and its derivatives. These higher order differential energy operators contain as a special case, for k=2, the Teager-Kaiser (1990) operator. When applied to (possibly modulated) sinusoids, they yield several new energy measurements useful for parameter estimation or AM-FM demodulation. Applying them to sampled signals involves replacing derivatives with differences that lead to several useful discrete energy operators defined on an extremely short window of samples.>
Petros Maragos, Alexandros Potamianos
IEEE Signal Process. Lett.2
1994 A comparison of the energy operator and the Hilbert transform approach to signal and speech demodulation
Alexandros Potamianos, Petros Maragos
Signal Process.1
1994 A system for finding speech formants and modulations via energy separation
abstract
This correspondence presents an experimental system that uses an energy-tracking operator and a related energy separation algorithm to automatically find speech formants and amplitude/frequency modulations in voiced speech segments. Initial estimates of formant center frequencies are provided by either LPC or morphological spectral peak picking. These estimates are then shown to be improved by a combination of bandpass filtering and iterative application of energy separation.>
Helen M. Hanson, Petros Maragos, Alexandros Potamianos
IEEE Trans. Speech Audio Process.3
1993 Finding speech formants and modulations via energy separation: with application to a vocoder
Helen M. Hanson, Petros Maragos, Alexandros Potamianos
ICASSP (2)3