VLDB 2026 Research / reviewers in the wild / expert
Dimitrios Dimitriadis
dblp:05/3143
· DBLP profile ↗
59ranked-venue papers
15as first author
14since 2021 · last 2025
0000-0001-8483-0105ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 48 · 13 first-author · 6 since 2021Artificial intelligence and machine learning · 35 · 9 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized SettingsabstractMixture-of-Experts (MoEs) achieve scalability by dynamically activating subsets of their components.
Yet, understanding how expertise emerges through joint training of gating mechanisms and experts remains incomplete, especially in scenarios without clear task partitions. Motivated by inference costs and data heterogeneity, we study how joint training of gating functions and experts can dynamically allocate domain-specific expertise across multiple underlying data distributions.
As an outcome of our framework, we develop an instance tailored specifically to decentralized training scenarios, introducing *Dynamically Decentralized Orchestration of MoEs* or *DDOME*. *DDOME* leverages heterogeneity emerging from distributional shifts across decentralized data sources to specialize experts dynamically. By integrating a pretrained common expert to inform a gating function, *DDOME* achieves personalized expert subset selection on-the-fly, facilitating just-in-time personalization.
We empirically validate *DDOME* within a Federated Learning (FL) context: *DDOME* attains from 4\% up to an 24\% accuracy improvement over state-of-the-art FL baselines in image and text classification tasks, while maintaining competitive zero-shot generalization capabilities. Furthermore, we provide theoretical insights confirming that the joint gating-experts training is critical for achieving meaningful expert specialization. Yehya Farhat, Hamza ElMokhtar Shili, Fangshuo Liao, Chen Dun, Mirian Hipolito Garcia, Guoqing Zheng, Ahmed Awadallah 0001, Robert Sim, Dimitrios Dimitriadis, Anastasios Kyrillidis |
NeurIPS | 9 |
| 2024 | MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMsabstractYavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao, Dimitrios Dimitriadis, Salman Avestimehr. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao, Dimitrios Dimitriadis, Amir Salman Avestimehr |
ACL (1) | 5 |
| 2024 | Invariant Aggregator for Defending against Federated Backdoor Attacks
Dimitrios Dimitriadis, Oluwasanmi Koyejo, Shruti Tople |
AISTATS | 2 |
| 2024 | CroMo-Mixup: Augmenting Cross-Model Representations for Continual Self-Supervised Learning
Erum Mushtaq, Duygu Nur Yaldiz, Yavuz Faruk Bakman, Jie Ding 0002, Chenyang Tao, Dimitrios Dimitriadis, Amir Salman Avestimehr |
ECCV (80) | 6 |
| 2024 | Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
Tiantian Feng, Dimitrios Dimitriadis, Shri Narayanan |
INTERSPEECH | 2 |
| 2024 | Predicting Uncertainty of Generative LLMs with MARS: Meaning-Aware Response ScoringabstractGenerative Large Language Models (LLMs) have recently been widely utilized for their unprecedented capabil-ities across many tasks. Considering their use in high-stakes environments and for mission-critical applications, the fact that LLMs often can generate inaccurate or misleading results can be potentially harmful, which motivates us to study the correctness of generative LLM outputs. Uncertainty Estimation (UE) in generative LLMs is a developing area, with state-of-the-art probability-based techniques frequently using length-normalized scoring. As an alternative to length-normalized scoring in UE, in this work, we propose Meaning-Aware Response Scoring (MARS). The key idea of MARS is to consider the semantic contribution of each token of the generated sequence to the context of the question during UE. Through extensive experiments on three question-answering datasets across five pretrained LLMs, we show that utilizing MARS during UE results in a universal and significant improvement in UE performance. Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Amir Salman Avestimehr, Chenyang Tao, Dimitrios Dimitriadis |
ISIT | 6 |
| 2024 | Personalized Federated Learning With Adaptive Batchnorm for HealthcareabstractThere is a growing interest in applying machine learning techniques to healthcare. Recently, federated machine learning (FL) is gaining popularity since it allows researchers to train powerful models without compromising data privacy and security. However, the performance of existing FL approaches often deteriorates when encountering non-iid situations where there exist distribution gaps among clients, and few previous efforts focus on personalization in healthcare. In this article, we propose FedAP to tackle domain shifts and obtain personalized models for local clients. FedAP learns the similarity between clients via the statistics of the batch normalization layers while preserving the specificity of each client with different local batch normalization. Comprehensive experiments on five healthcare benchmarks demonstrate that FedAP achieves better accuracy compared to state-of-the-art methods (e.g., 10%+ accuracy improvement for PAMAP2) with faster convergence speed. Wang Lu 0003, Jindong Wang 0001, Yiqiang Chen 0001, Renjun Xu, Dimitrios Dimitriadis, Tao Qin 0001 |
IEEE Trans. Big Data | 6 |
| 2023 | Efficient and Light-Weight Federated Learning via Asynchronous Distributed DropoutabstractAsynchronous learning protocols have regained attention lately, especially in the Federated Learning (FL) setup, where slower clients can severely impede the learning process. Herein, we propose AsyncDrop, a novel asynchronous FL framework that utilizes dropout regularization to handle device heterogeneity in distributed settings. Overall, AsyncDrop achieves better performance compared to state of the art asynchronous methodologies, while resulting in less communication and training time overheads. The key idea revolves around creating “submodels” out of the global model, and distributing their training to workers, based on device heterogeneity. We rigorously justify that such an approach can be theoretically characterized. We implement our approach and compare it against other asynchronous baselines, both by design and by adapting existing synchronous FL algorithms to asynchronous scenarios. Empirically, AsyncDrop reduces the communication cost and training time, while matching or improving the final test accuracy in diverse non-i.i.d. FL scenarios. Chen Dun, Mirian Hipolito Garcia, Chris Jermaine, Dimitrios Dimitriadis, Anastasios Kyrillidis |
AISTATS | 4 |
| 2023 | Local or Global: Selective Knowledge Assimilation for Federated Learning with Limited LabelsabstractMany existing FL methods assume clients with fully-labeled data, while in realistic settings, clients have limited labels due to the expensive and laborious process of labeling. Limited labeled local data of the clients often leads to their local model having poor generalization abilities to their larger unlabeled local data, such as having class-distribution mismatch with the unlabeled data. As a result, clients may instead look to benefit from the global model trained across clients to leverage their unlabeled data, but this also becomes difficult due to data heterogeneity across clients. In our work, we propose FedLabel where clients selectively choose the local or global model to pseudo-label their unlabeled data depending on which is more of an expert of the data. We further utilize both the local and global models’ knowledge via global-local consistency regularization which minimizes the divergence between the two models’ outputs when they have identical pseudo-labels for the unlabeled data. Unlike other semi-supervised FL baselines, our method does not require additional experts other than the local or global model, nor require additional parameters to be communicated. We also do not assume any server-labeled data or fully labeled clients. For both cross-device and cross-silo settings, we show that FedLabel outperforms other semi-supervised FL baselines by 8-24%, and even outperforms standard fully supervised FL baselines (100% labeled data) with only 5-20% of labeled data. Yae Jee Cho, Gauri Joshi, Dimitrios Dimitriadis |
ICCV | 3 |
| 2022 | Heterogeneous Ensemble Knowledge Transfer for Training Large Models in Federated LearningabstractFederated learning (FL) enables edge-devices to collaboratively learn a model without disclosing their private data to a central aggregating server. Most existing FL algorithms require models of identical architecture to be deployed across the clients and server, making it infeasible to train large models due to clients' limited system resources. In this work, we propose a novel ensemble knowledge transfer method named Fed-ET in which small models (different in architecture) are trained on clients, and used to train a larger model at the server. Unlike in conventional ensemble learning, in FL the ensemble can be trained on clients' highly heterogeneous data. Cognizant of this property, Fed-ET uses a weighted consensus distillation scheme with diversity regularization that efficiently extracts reliable consensus from the ensemble while improving generalization by exploiting the diversity within the ensemble. We show the generalization bound for the ensemble of weighted models trained on heterogeneous datasets that supports the intuition of Fed-ET. Our experiments on image and language tasks show that Fed-ET significantly outperforms other state-of-the-art FL algorithms with fewer communicated parameters, and is also robust against high data-heterogeneity. Yae Jee Cho, Andre Manoel, Gauri Joshi, Robert Sim, Dimitrios Dimitriadis |
IJCAI | 5 |
| 2022 | UserIdentifier: Implicit User Representations for Simple and Effective Personalized Sentiment AnalysisabstractFatemehsadat Mireshghallah, Vaishnavi Shrivastava, Milad Shokouhi, Taylor Berg-Kirkpatrick, Robert Sim, Dimitrios Dimitriadis. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Niloofar Mireshghallah, Vaishnavi Shrivastava, Milad Shokouhi, Taylor Berg-Kirkpatrick, Robert Sim, Dimitrios Dimitriadis |
NAACL-HLT | 6 |
| 2022 | A review of speaker diarization: Recent advances with deep learning
Tae Jin Park, Naoyuki Kanda, Dimitrios Dimitriadis, Kyu Jeong Han, Shinji Watanabe 0001, Shri Narayanan |
Comput. Speech Lang. | 3 |
| 2021 | Ensemble Combination between Different Time SegmentationsabstractHypothesis-level combination between multiple models can often yield gains in speech recognition. However, all models in the ensemble are usually restricted to use the same audio segmentation times. This paper proposes to generalise hypothesis-level combination, allowing the use of different audio segmentation times between the models, by splitting and re-joining the hypothesised N-best lists in time. A hypothesis tree method is also proposed to distribute hypothesis posteriors among the constituent words, to facilitate such splitting when per-word scores are not available. The approach is assessed on a Microsoft meeting transcription task, by performing combination between a streaming first-pass recognition and an offline second-pass recognition. The experimental results show that the proposed approach can yield gains when combining over different segmentation times. Furthermore, the results also show that a combination between a hybrid model and an end-to-end neural network model yields a greater improvement than a combination between two hybrid models. Jeremy H. M. Wong, Dimitrios Dimitriadis, Ken'ichi Kumatani, Yashesh Gaur, George Polovets, Partha Parthasarathy, Eric Sun, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 2 |
| 2021 | One-Shot Voice Conversion with Speaker-Agnostic StarGAN
Sefik Emre Eskimez, Dimitrios Dimitriadis, Ken'ichi Kumatani, Robert Gmyr |
Interspeech | 2 |
| 2020 | A Memory Augmented Architecture for Continuous Speaker Identification in MeetingsabstractWe introduce and analyze a novel approach to the problem of speaker identification in multi-party recorded meetings. Given a speech segment and a set of available candidate profiles, a data-driven approach is proposed learning the distance relations between them, aiming at identifying the correct speaker label corresponding to that segment. A recurrent, memory-based architecture is employed, since this class of neural networks has been shown to yield improved performance in problems requiring relational reasoning. The proposed encoding of distance relations is shown to outperform traditional distance metrics, such as the cosine distance. Additional improvements are reported when the temporal continuity of the audio signals and the speaker changes is modeled in. In this paper, the proposed method is evaluated in two different tasks, i.e. scripted and real-world business meeting scenarios, where a relative reduction in speaker error rate of 39.28% and 51.84%, respectively, is reported when compared with the baseline. Nikolaos Flemotomos, Dimitrios Dimitriadis |
ICASSP | 2 |
| 2020 | Combining Acoustics, Content and Interaction Features to Find Hot Spots in MeetingsabstractInvolvement hot spots have been proposed as a useful concept for meeting analysis and studied off and on for over 15 years. These are regions of meetings that are marked by high participant involvement, as judged by human annotators. However, prior work was either not conducted in a formal machine learning setting, or focused on only a subset of possible meeting features or downstream applications (such as summarization). In this paper we investigate to what extent various acoustic, linguistic and pragmatic aspects of the meetings, both in isolation and jointly, can help detect hot spots. In this context, the openSMILE toolkit [1] is to used to extract features based on acoustic-prosodic cues, BERT word embeddings [2] are used for encoding the lexical content, and a variety of statistics based on speech activity are used to describe the verbal interaction among participants. In experiments on the annotated ICSI meeting corpus, we find that the lexical model is the most informative, with incremental contributions from interaction and acoustic-prosodic model components. Dave Makhervaks, William Hinthorn, Dimitrios Dimitriadis, Andreas Stolcke |
ICASSP | 3 |
| 2020 | A Federated Approach in Training Acoustic ModelsabstractIn this paper, a novel platform for Acoustic Model training based on Federated Learning (FL) is described. This is the first attempt to introduce Federated Learning techniques in Speech Recognition (SR) tasks. Besides the novelty of the task, the paper describes an easily generalizable FL platform and presents the design decisions used for this task. Amongst the novel algorithms introduced is a hierarchical optimization scheme employing pairs of optimizers and an algorithm for gradient selection, leading to improvements in training time and SR performance. The experimental validation of the proposed system is based on the LibriSpeech task, presenting a speed-up of x1.5 and 6% WERR. The proposed Federated Learning system appears to outperform the golden standard of distributed training in both convergence speed and overall model performance. Further improvements have been experienced in internal tasks. Dimitrios Dimitriadis, Ken'ichi Kumatani, Robert Gmyr, Yashesh Gaur, Sefik Emre Eskimez |
INTERSPEECH | 1 |
| 2020 | GAN-Based Data Generation for Speech Emotion RecognitionabstractIn this work, we propose a GAN-based method to generate synthetic data for speech emotion recognition. Specifically, we investigate the usage of GANs for capturing the data manifold when the data is eyes-off, i.e., where we can train networks using the data but cannot copy it from the clients. We propose a CNN-based GAN with spectral normalization on both the generator and discriminator, both of which are pre-trained on large unlabeled speech corpora. We show that our method provides better speech emotion recognition performance than a strong baseline. Furthermore, we show that even after the data on the client is lost, our model can generate similar data that can be used for model bootstrapping in the future. Although we evaluated our method for speech emotion recognition, it can be applied to other tasks. Sefik Emre Eskimez, Dimitrios Dimitriadis, Robert Gmyr, Kenichi Kumanati |
INTERSPEECH | 2 |
| 2020 | Sequence-Level Self-Learning with Multiple HypothesesabstractIn this work, we develop new self-learning techniques with an attention-based sequence-to-sequence (seq2seq) model for automatic speech recognition (ASR). For untranscribed speech data, the hypothesis from an ASR system must be used as a label. However, the imperfect ASR result makes unsupervised learning difficult to consistently improve recognition performance especially in the case that multiple powerful teacher models are unavailable. In contrast to conventional unsupervised learning approaches, we adopt the \emph{multi-task learning} (MTL) framework where the $n$-th best ASR hypothesis is used as the label of each task. The seq2seq network is updated through the MTL framework so as to find the common representation that can cover multiple hypotheses. By doing so, the effect of the \emph{hard-decision} errors can be alleviated. We first demonstrate the effectiveness of our self-learning methods through ASR experiments in an accent adaptation task between the US and British English speech. Our experiment results show that our method can reduce the WER on the British speech data from 14.55\% to 10.36\% compared to the baseline model trained with the US English data only. Moreover, we investigate the effect of our proposed methods in a federated learning scenario. Ken'ichi Kumatani, Dimitrios Dimitriadis, Yashesh Gaur, Robert Gmyr, Sefik Emre Eskimez, Jinyu Li 0001, Michael Zeng 0001 |
INTERSPEECH | 2 |
| 2019 | Advances in Online Audio-Visual Meeting TranscriptionabstractThis paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its ability to handle overlapped speech, which has been an unsolved problem in realistic settings for over a decade. We show that this problem can be addressed by using a continuous speech separation approach. In addition, we describe an online audio-visual speaker diarization method that leverages face tracking and identification, sound source localization, speaker identification, and, if available, prior speaker information for robustness to various real world challenges. All components are integrated in a meeting transcription framework called SRD, which stands for “separate, recognize, and diarize”. Experimental results using recordings of natural meetings involving up to 11 attendees are reported. The continuous speech separation improves a word error rate (WER) by 16.1% compared with a highly tuned beamformer. When a complete list of meeting attendees is available, the discrepancy between WER and speaker-attributed WER is only 1.0%, indicating accurate word-to-speaker association. This increases marginally to 1.6% when 50% of the attendees are unknown to the system. Takuya Yoshioka, Yan Huang 0028, Aviv Hurvitz, Sharon Koubi, Eyal Krupka, Ido Leichter, Changliang Liu, Partha Parthasarathy, Alon Vinnikov, Lingfeng Wu, Igor Abramovski, Wayne Xiong, Huaming Wang, Jun Zhang 0066, Yong Zhao 0008, Tianyan Zhou, Cem Aksoylar, Zhuo Chen 0006, Moshe David, Dimitrios Dimitriadis, Yifan Gong 0001, Ilya Gurvich, Xuedong Huang 0001 |
ASRU | 23 |
| 2019 | Acoustic and Lexical Sentiment Analysis for Customer Service CallsabstractWe describe the development of a sentiment analysis system for customer service calls, starting with the data acquisition and labeling, and proceeding to the algorithmic information extraction and modeling process from both spoken words and their acoustic expression. The proposed system is based on the combination of multiple acoustic and lexical models in a late fusion approach. Acoustic aspects of sentiment are captured by utterance-level features based on aggregated openSMILE and raw cepstral features, and further augmented with an energy contour model. Lexical aspects are captured by back-off n-gram language models. These models are found to combine effectively, showing different strengths as pertains to positive and negative sentiment detection. Bryan Li, Dimitrios Dimitriadis, Andreas Stolcke |
ICASSP | 2 |
| 2019 | Single-channel Speech Extraction Using Speaker Inventory and Attention NetworkabstractNeural network-based speech separation has received a surge of interest in recent years. Previously proposed methods either are speaker independent or extract a target speaker's voice by using his or her voice snippet. In applications such as home devices or office meeting transcriptions, a possible speaker list is available, which can be leveraged for speech separation. This paper proposes a novel speech extraction method that utilizes an inventory of voice snippets of possible interfering speakers, or speaker enrollment data, in addition to that of the target speaker. Furthermore, an attention-based network architecture is proposed to form time-varying masks for both the target and other speakers during the separation process. This architecture does not reduce the enrollment audio of each speaker into a single vector, thereby allowing each short time frame of the input mixture signal to be aligned and accurately compared with the enrollment signals. We evaluate the proposed system on a speaker extraction task derived from the Libri corpus and show the effectiveness of the method. Zhuo Chen 0006, Takuya Yoshioka, Hakan Erdogan, Changliang Liu, Dimitrios Dimitriadis, Jasha Droppo, Yifan Gong 0001 |
ICASSP | 6 |
| 2019 | Low-latency Speaker-independent Continuous Speech SeparationabstractSpeaker independent continuous speech separation (SI-CSS) is a task of converting a continuous audio stream, which may contain overlapping voices of unknown speakers, into a fixed number of continuous signals each of which contains no overlapping speech segment. A separated, or cleaned, version of each utterance is generated from one of SI-CSS's output channels nondeterministically without being split up and distributed to multiple channels. A typical application scenario is transcribing multi-party conversations, such as meetings, recorded with microphone arrays. The output signals can be simply sent to a speech recognition engine because they do not include speech overlaps. The previous SI-CSS method uses a neural network trained with permutation invariant training and a data-driven beamformer and thus requires much processing latency. This paper proposes a low-latency SI-CSS method whose performance is comparable to that of the previous method in a microphone array-based meeting transcription task. This is achieved (1) by using a new speech separation network architecture combined with a double buffering scheme and (2) by performing enhancement with a set of fixed beamformers followed by a neural post-filter. Takuya Yoshioka, Zhuo Chen 0006, Changliang Liu, Hakan Erdogan, Dimitrios Dimitriadis |
ICASSP | 6 |
| 2019 | Meeting Transcription Using Asynchronous Distant MicrophonesabstractWe describe a system that generates speaker-annotated transcripts of meetings by using multiple asynchronous distant microphones. The system is composed of continuous audio stream alignment, blind beamforming, speech recognition, speaker diarization, and system combination. While the idea of improving the meeting transcription accuracy by leveraging multiple recordings has been investigated in certain specific technology areas such as beamforming, our objective is to assess the feasibility of a complete system with a set of mobile devices and conduct a detailed analysis. With seven input audio streams, our system achieves a word error rate (WER) of 22.3% and a speaker-attributed WER (SAWER) of 26.7%, and comes within 3% of the close-talking microphone WER on non-overlapping speech. The relative gains in SAWER over a single-device system are 14.8%, 20.3%, and 22.4% for three, five, and seven microphones, respectively. The full system achieves a 13.6% diarization error rate, 10% of which are due to overlapped speech. Takuya Yoshioka, Dimitrios Dimitriadis, Andreas Stolcke, William Hinthorn, Zhuo Chen 0006, Michael Zeng 0001, Xuedong Huang 0001 |
INTERSPEECH | 2 |
| 2018 | Improving End-of-Turn Detection in Spoken Dialogues by Detecting Speaker Intentions as a Secondary TaskabstractThis work focuses on the use of acoustic cues for modeling turn-taking in dyadic spoken dialogues. Previous work has shown that speaker intentions (e.g., asking a question, uttering a backchannel, etc.) can influence turn-taking behavior and are good predictors of turn-transitions in spoken dialogues. However, speaker intentions are not readily available for use by automated systems at run-time; making it difficult to use this information to anticipate a turn-transition. To this end, we propose a multi-task neural approach for predicting turn-transitions and speaker intentions simultaneously. Our results show that adding the auxiliary task of speaker intention prediction improves the performance of turn-transition prediction in spoken dialogues, without relying on additional input features during run-time. Zakaria Aldeneh, Dimitrios Dimitriadis, Emily Mower Provost |
ICASSP | 2 |
| 2017 | Speaker diarization: A perspective on challenges and opportunities from theory to practiceabstractThis paper discusses some challenges and opportunities in developing a speaker diarization system for operation on real world call center telephony data. We contrast some of the differences between a standard data set akin to NIST evaluations and those found in call centers. In exploring these differences we discovered vulnerabilities and proposed changes to address them. In moving from theory into practice we introduce two tasks in which speaker diarization and recognition can be leveraged. First, we show that speaker diarization and recognition systems can be integrated to find the common speaker (the call center agent) across multiple calls and consequently their role. Furthermore, once the role is determined the corresponding speech recognition output can be analyzed to determine the type of support call. Kenneth Church 0001, Weizhong Zhu, Josef Vopicka, Jason W. Pelecanos, Dimitrios Dimitriadis, Petr Fousek |
ICASSP | 5 |
| 2017 | Pooling acoustic and lexical features for the prediction of valenceabstractIn this paper, we present an analysis of different multimodal fusion approaches in the context of deep learning, focusing on pooling intermediate representations learned for the acoustic and lexical modalities. Traditional approaches to multimodal feature pooling include: concatenation, element-wise addition, and element-wise multiplication. We compare these traditional methods to outer-product and compact bilinear pooling approaches, which consider more comprehensive interactions between features from the two modalities. We also study the influence of each modality on the overall performance of a multimodal system. Our experiments on the IEMOCAP dataset suggest that: (1) multimodal methods that combine acoustic and lexical features outperform their unimodal counterparts; (2) the lexical modality is better for predicting valence than the acoustic modality; (3) outer-product-based pooling strategies outperform other pooling strategies. Zakaria Aldeneh, Soheil Khorram, Dimitrios Dimitriadis, Emily Mower Provost |
ICMI | 3 |
| 2017 | Developing On-Line Speaker Diarization System
Dimitrios Dimitriadis, Petr Fousek |
INTERSPEECH | 1 |
| 2017 | Progressive Neural Networks for Transfer Learning in Emotion RecognitionabstractMany paralinguistic tasks are closely related and thus representations learned in one domain can be leveraged for another. In this paper, we investigate how knowledge can be transferred between three paralinguistic tasks: speaker, emotion, and gender recognition. Further, we extend this problem to cross-dataset tasks, asking how knowledge captured in one emotion dataset can be transferred to another. We focus on progressive neural networks and compare these networks to the conventional deep learning method of pre-training and fine-tuning. Progressive neural networks provide a way to transfer knowledge and avoid the forgetting effect present when pre-training neural networks on different tasks. Our experiments demonstrate that: (1) emotion recognition can benefit from using representations originally learned for different paralinguistic tasks and (2) transfer learning can effectively leverage additional datasets to improve the performance of emotion recognition systems. John Gideon, Soheil Khorram, Zakaria Aldeneh, Dimitrios Dimitriadis, Emily Mower Provost |
INTERSPEECH | 4 |
| 2017 | Capturing Long-Term Temporal Dependencies with Convolutional Networks for Continuous Emotion RecognitionabstractThe goal of continuous emotion recognition is to assign an emotion value to every frame in a sequence of acoustic features. We show that incorporating long-term temporal dependencies is critical for continuous emotion recognition tasks. To this end, we first investigate architectures that use dilated convolutions. We show that even though such architectures outperform previously reported systems, the output signals produced from such architectures undergo erratic changes between consecutive time steps. This is inconsistent with the slow moving ground-truth emotion labels that are obtained from human annotators. To deal with this problem, we model a downsampled version of the input signal and then generate the output signal through upsampling. Not only does the resulting downsampling/upsampling network achieve good performance, it also generates smooth output trajectories. Our method yields the best known audio-only performance on the RECOLA dataset. Soheil Khorram, Zakaria Aldeneh, Dimitrios Dimitriadis, Melvin G. McInnis, Emily Mower Provost |
INTERSPEECH | 3 |
| 2017 | English Conversational Telephone Speech Recognition by Humans and MachinesabstractOne of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major speech recognition improvements on the representative Switchboard conversational corpus. Word error rates that just a few years ago were 14% have dropped to 8.0%, then 6.6% and most recently 5.8%, and are now believed to be within striking range of human performance. This then raises two issues - what IS human performance, and how far down can we still drive speech recognition error rates? A recent paper by Microsoft suggests that we have already achieved human performance. In trying to verify this statement, we performed an independent set of human performance measurements on two conversational tasks and found that human performance may be considerably better than what was earlier reported, giving the community a significantly harder goal to achieve. We also report on our own efforts in this area, presenting a set of acoustic and language modeling techniques that lowered the word error rate of our own English conversational telephone LVCSR system to the level of 5.5%/10.3% on the Switchboard/CallHome subsets of the Hub5 2000 evaluation, which - at least at the writing of this paper - is a new performance milestone (albeit not at what we measure to be human performance!). On the acoustic side, we use a score fusion of three models: one LSTM with multiple feature inputs, a second LSTM trained with speaker-adversarial multi-task learning and a third residual net (ResNet) with 25 convolutional layers and time-dilated convolutions. On the language modeling side, we use word and character LSTMs and convolutional WaveNet-style language models. George Saon, Gakuto Kurata, Tom Sercu, Kartik Audhkhasi, Samuel Thomas 0001, Dimitrios Dimitriadis, Bhuvana Ramabhadran, Michael Picheny, Lynn-Li Lim, Bergul Roomi, Phil Hall |
INTERSPEECH | 6 |
| 2016 | On the importance of event detection for ASRabstractThe performance of modern large vocabulary continuous speech recognition (LVCSR) systems is heavily affected by segment boundaries, proper speaker identification of the segments, as well as removal of spurious data. We propose to use Long Short Term Memory (LSTM) recurrent neural networks to partition audio into speech segments as well as track speaker turns. Additionally, we train an LSTM to also identify music segments. We show that the accurate detection of events, along with removal of silence and music, using our LSTM yields a 9-10% relative improvement in ASR performance. Secondary processing by speaker clustering provides an additional boost in accuracy. Event detection accuracy of the LSTM approach is also described. David Haws, Dimitrios Dimitriadis, George Saon, Samuel Thomas 0001, Michael Picheny |
ICASSP | 2 |
| 2016 | CNMF-based acoustic features for noise-robust ASRabstractWe present an algorithm using convolutive non-negative matrix factorization (CNMF) to create noise-robust features for automatic speech recognition (ASR). Typically in noise-robust ASR, CNMF is used to remove noise from noisy speech prior to feature extraction. However, we find that denoising introduces distortion and artifacts, which can degrade ASR performance. Instead, we propose using the time-activation matrices from CNMF as acoustic model features. In this paper, we describe how to create speech and noise dictionaries that generate noise-robust time-activation matrices from noisy speech. Using the time-activation matrices created by our proposed algorithm, we achieve a 11.8% relative improvement in the word error rate on the Aurora 4 corpus compared to using log-mel filterbank energies. Furthermore, we attain a 13.8% relative improvement over log-mel filterbank energies when we combine them with our proposed features, indicating that our features contain complementary information to log-mel features. Colin Vaz, Dimitrios Dimitriadis, Samuel Thomas 0001, Shri Narayanan |
ICASSP | 2 |
| 2016 | An Investigation on the Use of i-Vectors for Robust ASRabstractIn this paper we propose two different i-vector representations that improve the noise robustness of automatic speech recognition (ASR). The first kind of i-vectors is derived from ``noise only'' components of speech provided by an adaptive denoising algorithm, the second variant is extracted from mel filterbank energies containing both speech and noise. The effectiveness of both these representations is shown by combining them with two different kinds of spectral features - the commonly used log-mel filterbank energies and Teager energy spectral coefficients (TESCs). Using two different DNN architectures for acoustic modeling - a standard state-of-the-art sigmoid-based DNN and an advanced architecture using leaky ReLUs, dropout and resealing, we demonstrate the benefit of the proposed representations. On the Aurora-4 multi-condition training task the proposed front-end improves ASR performance by 4%. Dimitrios Dimitriadis, Samuel Thomas 0001, Sriram Ganapathy |
INTERSPEECH | 1 |
| 2015 | Investigating factor analysis features for deep neural networks in noisy speech recognitionabstractThe problem of speaker and channel adaptation in deep neural network (DNN) based automatic speech recognition (ASR) sys-tems is of substantial interest in advancing the performance of these systems. Recently, the speaker identity vectors (i-vectors) have shown improvements for ASR systems in matched condi-tions. In this paper, we propose the application of the general factor analysis framework for noisy speech recognition tasks. Several methods for deriving speaker and channel factors are explored including joint factor analysis (JFA) and i-vectors de-rived from DNN posteriors instead of the traditional Universal background model (UBM) approach. We also experiment with the late fusion of i-vector features with bottleneck (BN) fea-tures obtained from a previously trained convolutional neural network (CNN) system. The ASR experiments are performed on the Aspire challenge test data which contains noisy far-field speech while the acoustic models are trained with conversa-tional telephone speech (CTS) data from the Fisher corpus. In these experiments, we show that the factor analysis based meth-ods provide significant improvements in the word error rate (relative improvements of about 11 % compared to the baseline DNN system trained with speaker adapted features). Sriram Ganapathy, Samuel Thomas 0001, Dimitrios Dimitriadis, Steven J. Rennie |
INTERSPEECH | 3 |
| 2015 | Use of Micro-Modulation Features in Large Vocabulary Continuous Speech Recognition TasksabstractMost of the state-of-the-art ASR systems take as input a single type of acoustic features, dominated by the traditional feature schemes, i.e., MFCCs or PLPs. However, these features cannot model rapid, intra-frame phenomena present in the actual speech signals. On the other hand, micro-modulation components, inspired by the AM-FM speech model, can capture these important characteristics of spoken speech, resulting in significant performance improvements, as previously shown in small-vocabulary ASR tasks. Yet, they have limited use in large vocabulary ASR applications, where feature post-processing schemes are usually employed. To enable the successful application of these frequency measures in real-life tasks, we investigate their combination with the traditional Cepstral features when employing linear, e.g., HDA, and nonlinear, i.e., bottleneck neural net (BN), feature transforms. This feature combination is investigated in the context of the hybrid DNN-HMM framework, as well. The experimental results reveal that the integration of micro-modulation and Cepstral features, using neural nets, can greatly improve the ASR performance with respect to using the Cepstral features alone. We apply this novel feature extraction approach on different tasks, i.e., a clean speech task (DARPA-WSJ), the Aurora-4 task and a real-life, open-vocabulary, mobile search task, the Speak4it, always reporting improved performance, while the obtained relative word error reduction ranges between 7%-21% depending on the task, e.g., a relative WER improvement of 18% for the Speak4it task, and similar improvements, up to 21%, for the WSJ task are reported. Dimitrios Dimitriadis, Enrico Bocchieri |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Enhancing audio source separability using spectro-temporal regularization with NMF
Colin Vaz, Dimitrios Dimitriadis, Shri Narayanan |
INTERSPEECH | 2 |
| 2013 | Investigating deep neural network based transforms of robust audio features for LVCSRabstractMicro-modulation components such as the formant frequencies are very important characteristics of spoken speech that have allowed great performance improvements in small-vocabulary ASR tasks. Yet they have limited use in large vocabulary ASR applications. To enable the successful application, in real-life tasks, of these frequency measures, we investigate their combination with traditional features (MFCC's and PLP's) by linear (e.g. HDA), and non-linear (bottleneck MLP) feature transforms. Our experiments show that such integration, using non-linear MLP-based transforms, of micro-modulation and cepstral features greatly improves the ASR with respect to the cepstral features alone. We have applied this novel feature extraction scheme onto two very different tasks, i.e. a clean speech task (DARPA-WSJ) and a real-life, open-vocabulary, mobile search task (Speak4itSM), always reporting improved performance. We report relative error rate reduction of 15% for the Speak4itSMtask, and similar improvements, up to 21%, for the WSJ task. Enrico Bocchieri, Dimitrios Dimitriadis |
ICASSP | 2 |
| 2013 | Instantaneous frequency and bandwidth estimation using filterbank arraysabstractAccurate estimation of the instantaneous frequency of speech resonances is a hard problem mainly due to phase discontinuities in the speech signal associated with excitation instants. We review a variety of approaches for enhanced frequency and bandwidth estimation in the time-domain and propose a new cognitively motivated approach using filterbank arrays. We show that by filtering speech resonances using filters of different center frequency, bandwidth and shape, the ambiguity in instantaneous frequency estimation associated with amplitude envelope minima and phase discontinuities can be significantly reduced. The novel estimators are shown to perform well on synthetic speech signals with frequency and bandwidth micro-modulations (i.e., modulations within a pitch period), as well as on real speech signals. Filterbank arrays, when applied to frequency and bandwidth modulation index estimation, are shown to reduce the estimation error variance by 85% and 70% respectively. Pirros Tsiakoulis, Alexandros Potamianos, Dimitrios Dimitriadis |
ICASSP | 3 |
| 2013 | On the improvement of multimodal voice activity detection
Matt Burlick, Dimitrios Dimitriadis, Eric Zavesky |
INTERSPEECH | 2 |
| 2013 | Robust speech enhancement techniques for ASR in non-stationary noise and dynamic environmentsabstractIn the current ASR systems the presence of competing speakers greatly degrades the recognition performance. This phenomenon is getting even more prominent in the case of hands-free, far-field ASR systems like the “Smart-TV” systems, where reverberation and non-stationary noise pose additional challenges. Furthermore, speakers are, most often, not standing still while speaking. To address these issues, we propose a cascaded system that includes Time Differences of Arrival estimation, multi-channel Wiener Filtering, nonnegative matrix factorization (NMF), multi-condition training, and robust feature extraction, whereas each of them additively improves the overall performance. The final cascaded system presents an average of 50% and 45% relative improvement in ASR word accuracy for the CHiME 2011(non-stationary noise) and CHiME 2012 (non-stationary noise plus speaker head movement) tasks, respectively. Dimitrios Dimitriadis, Enrico Bocchieri |
INTERSPEECH | 2 |
| 2013 | Incremental emotion recognitionabstractMost emotion recognition systems do not perform real-time emotion recognition due to latencies caused by phrase segmentation and resource-intensive feature acquisition, etc. To address this issue, we present an emotion recognition approach that can estimate speaker emotions with much lower latency. The proposed approach does not rely on phrase-level features to recognize speaker emotion; rather, it estimates the speaker’s emotional state over the course of the utterance incrementally, using a shifting n-word window on the basis of easily computable features. These features are obtained from three information streams, i.e. cepstral, prosodic and textual, at the wordlevel and combined at decision-level using a statistical framework. Our work shows that combining the three information streams yields higher emotion recognition accuracy than any single information stream. Using features extracted from n-word sequences rather than phrases provides for the low-latency capabilities of the proposed system, without any loss in utterance-level emotion recognition accuracy. The performance of the proposed system on a binary utterance-level emotion recognition task using an in-house database shows a relative improvement of 41 % over chance, compared to a relative improvement of 31.82 % shown by the baseline phrase-level emotion recognition approach. Taniya Mishra, Dimitrios Dimitriadis |
INTERSPEECH | 2 |
| 2012 | Dominant spatio-temporal modulations and energy tracking in videos: Application to interest point detection for action recognitionabstractThe presence of multiband amplitude and frequency modulations (AM-FM) in wideband signals, such as textured images or speech, has led to the development of efficient multicomponent modulation models for low-level image and sound analysis. Moreover, compact yet descriptive representations have emerged by tracking, through non-linear energy operators, the dominant model components across time, space or frequency. In this paper, we propose a generalization of such approaches in the 3D spatio-temporal domain and explore the potential of incorporating the Dominant Component Analysis scheme for interest point detection and human action recognition in videos. Within this framework, actions are implicitly considered as manifestations of spatio-temporal oscillations in the dynamic visual stream. Multiband filtering and energy operators are applied to track the source energy in both spatial and temporal frequency bands. A new measure for extracting keypoint locations is formulated as the temporal dominant energy computed over the spatial dominant components, in terms of their modulation energy, of input video frames. Theoretical formulation is supported by evaluation and comparisons in human action classification, which demonstrate the potential of the proposed spatio-temporal detector. Christos Georgakis 0001, Petros Maragos, Georgios Evangelopoulos, Dimitrios Dimitriadis |
ICIP | 4 |
| 2011 | Speech recognition modeling advances for mobile voice searchabstractThis paper reports on the development and advances in automatic speech recognition for the AT&T Speak4it®voice-search application. With Speak4it as real-life example, we show the effectiveness of acoustic model (AM) and language model (LM) estimation (adaptation and training) on relatively small amounts of application field-data. We then introduce algorithmic improvements concerning the use of sentence length in LM, of non-contextual features in AM decision-trees, and of the Teager energy in the acoustic front-end. The combination of these algorithms, integrated into the AT&T Watson recognizer, yields substantial accuracy improvements. LM and AM estimation on field-data samples increases the word accuracy from 66.4% to 77.1%, a relative word error reduction of 32%. The algorithmic improvements increase the accuracy to 79.7%, an additional 11.3% relative error reduction. Enrico Bocchieri, Diamantino Caseiro, Dimitrios Dimitriadis |
ICASSP | 3 |
| 2011 | An alternative front-end for the AT&T WATSON LV-CSR systemabstractIn previously published work, we have proposed a novel feature extraction algorithm, based on the Teager-Kaiser energy estimates, that approximates human auditory characteristics and that is more robust to sub-band noise than the mean-square estimates of standard MFCCs. We refer to the novel features as Teager energy cepstrum coefficients (TECC). Herein, we study the TECC performance under additive noise and suggest how to predict the noisy TECC deviations by estimating the subband SNR values. Then, we report on the effectiveness of the TECCs when they are used hi the acoustic front-end of the state-of-the-art AT&T WATSON large-vocabulary recognizer. The TECC front-end is tested in the real-life voice-search Speak4it application for mobile devices. It provides a 6% relative word error rate reduction w.r.t. the MFCC front-end, using the same high performance language model, lexicon and acoustic model training. Dimitrios Dimitriadis, Enrico Bocchieri, Diamantino Caseiro |
ICASSP | 1 |
| 2011 | Combining Frame and Segment Level Processing via Temporal Pooling for Phonetic ClassificationabstractWe propose a simple, yet novel, multi-layer model for the problem of phonetic classification. Our model combines the frame level transformation of the acoustic signal with the segment level transformation via a temporal pooling architecture to compute class conditional probabilities of phones. Without the use of any phonetic knowledge, our model achieved the state-ofthe-art performance on the TIMIT phone classification task. The flexibility of our model allows us to mix a variety of pooling architectures, leading to further significant performance improvements. Index Terms: deep networks, connectionist networks, multilayer models, ensemble methods, phone classification Sumit Chopra, Patrick Haffner, Dimitrios Dimitriadis |
INTERSPEECH | 3 |
| 2011 | On the Effects of Filterbank Design and Energy Computation on Robust Speech RecognitionabstractIn this paper, we examine how energy computation and filterbank design contribute to the overall front-end robustness, especially when the investigated features are applied to noisy speech signals, in mismatched training-testing conditions. In prior work (“Auditory Teager energy cepstrum coefficients for robust speech recognition,” D. Dimitriadis, P. Maragos, and A. Potamianos, in Proc. Eurospeech'05, Sep. 2005), a novel feature set called “Teager energy cepstrum coefficients” (TECCs) has been proposed, employing a dense, smooth filterbank and alternative energy computation schemes. TECCs were shown to be more robust to noise and exhibit improved performance compared to the widely used Mel frequency cepstral coefficients (MFCCs). In this paper, we attempt to interpret these results using a combined theoretical and experimental analysis framework. Specifically, we investigate in detail the connection between the filterbank design, i.e., the filter shape and bandwidth, the energy estimation scheme and the automatic speech recognition (ASR) performance under a variety of additive and/or convolutional noise conditions. For this purpose: 1) the performance of filterbanks using triangular, Gabor, and Gammatone filters with various bandwidths and filter positions are examined under different noisy speech recognition tasks, and 2) the squared amplitude and Teager-Kaiser energy operators are compared as two alternative approaches of computing the signal energy. Our end-goal is to understand how to select the most efficient filterbank and energy computation scheme that are maximally robust under both clean and noisy recording conditions. Theoretical and experimental results show that: 1) the filter bandwidth is one of the most important factors affecting speech recognition performance in noise, while the shape of the filter is of secondary importance, and 2) the Teager-Kaiser operator outperforms (on the average and for most noise types) the squared amplitude energy computation scheme for speech recognition in noisy conditions, especially, for large filter bandwidths. Experimental results show that selecting the appropriate filterbank and energy computation scheme can lead to significant error rate reduction over both MFCC and perceptual linear predicion (PLP) features for a variety of speech recognition tasks. A relative error rate reduction of up to ~ 30% for MFCCs and ~ 39% for PLPs is shown for the Aurora-3 Spanish Task. Dimitrios Dimitriadis, Petros Maragos, Alexandros Potamianos |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Spectral Moment Features Augmented by Low Order Cepstral Coefficients for Robust ASRabstractWe propose a novel Automatic Speech Recognition (ASR) front-end, that consists of the first central Spectral Moment time-frequency distribution Augmented by low order Cepstral coefficients (SMAC). We prove that the first central spectral moment is proportional to the spectral derivative with respect to the filter's central frequency. Consequently, the spectral moment is an estimate of the frequency domain derivative of the speech spectrum. However information related to the entire speech spectrum, such as the energy and the spectral tilt, is not adequately modeled. We propose adding this information with few cepstral coefficients. Furthermore, we use a mel-spaced Gabor filterbank with 70% frequency overlap in order to overcome the sensitivity to pitch harmonics. The novel SMAC front-end was evaluated for the speech recognition task for a variety of recording conditions. The experimental results have shown that SMAC performs at least as well as the standard MFCC front-end in clean conditions, and significantly outperforms MFCCs in noisy conditions. Pirros Tsiakoulis, Alexandros Potamianos, Dimitrios Dimitriadis |
IEEE Signal Process. Lett. | 3 |
| 2009 | Short-time instantaneous frequency and bandwidth features for speech recognitionabstractIn this paper, we investigate the performance of modulation related features and normalized spectral moments for automatic speech recognition. We focus on the short-time averages of the amplitude weighted instantaneous frequencies and bandwidths, computed at each subband of a mel-spaced filterbank. Similar features have been proposed in previous studies, and have been successfully combined with MFCCs for speech and speaker recognition. Our goal is to investigate the stand-alone performance of these features. First, it is experimentally shown that the proposed features are only moderately correlated in the frequency domain, and, unlike MFCCs, they do not require a transformation to the cepstral domain. Next, the filterbank parameters (number of filters and filter overlap) are investigated for the proposed features and compared with those of MFCCs. Results show that frequency related features perform at least as well as MFCCs for clean conditions, and yield superior results for noisy conditions; up to 50% relative error rate reduction for the AURORA3 Spanish task. Pirros Tsiakoulis, Alexandros Potamianos, Dimitrios Dimitriadis |
ASRU | 3 |
| 2009 | GridNews: A distributed automatic Greek broadcast transcription systemabstractIn this paper, a distributed system storing and retrieving broadcast news data recorded from the Greek television is presented. These multimodal data are processed in a grid computational environment interconnecting distributed data storage and processing subsystems. The innovative element of this system is the implementation of the signal processing algorithms in this grid environment, offering additional flexibility and computational power. Among the developed signal processing modules are: the Segmentor, cutting up the original videos into shorter ones, the classifier, recognizing whether these short videos contain speech or not, the Greek large-vocabulary speech recognizer, transcribing speech into written text, and finally the text search engine and the video retriever. All the processed data are stored and retrieved in geographically distributed storage elements. A user-friendly, Web-based interface is developed, facilitating the transparent import and storage of new multimodal data, their off-line processing and finally, their search and retrieval. Dimitrios Dimitriadis, A. Metallinou, Ioannis Konstantinou, Georgios I. Goumas, Petros Maragos, Nectarios Koziris |
ICASSP | 1 |
| 2007 | Multiband, multisensor robust features for noisy speech recognitionabstractThis paper presents a novel feature extraction scheme tak-ing advantage of both the nonlinear modulation speech model and the spatial diversity of speech and noise signals in a mul-tisensor environment. Herein, we propose applying robust fea-tures to speech signals captured by a multisensor array mini-mizing a noise energy criterion over multiple frequency bands. We show that we can achieve improved recognition perfor-mance by minimizing the Teager-Kaiser energy of the noise-corrupted signals in different frequency bands. These Multi-band, Multisensor Cepstral (MBSC) features are inspired by similar ones already been applied to single-microphone noisy Speech Recognition tasks with significantly improved results. The recognition results show that the proposed features can per-form better than the widely-used MFCC features. Dimitrios Dimitriadis, Petros Maragos, Stamatios Lefkimmiatis |
INTERSPEECH | 1 |
| 2007 | Advanced front-end for robust speech recognition in extremely adverse environmentsabstractIn this paper, a unified approach to speech enhancement, feature extraction and feature normalization for speech recognition in adverse recording conditions is presented. The proposed frontend system consists of several different, independent, processing modules. Each of the algorithms contained in these modules has been independently applied to the problem of speech recognition in noise, significantly improving the recognition rates. In this work, these algorithms are merged in a single front-end and their combined performance is demonstrated. Specifically, the proposed advanced front-end extracts noise-invariant features via the following modules: Wiener filtering, voice-activity detection, robust feature extraction (nonlinear modulation or fractal features), parameter equalization and frame-dropping. The advanced front-end is applied to extremely adverse environments where most feature extraction schemes fail. We show that by combining speech enhancement, robust feature extraction and feature normalization up to a fivefold error rate reduction can be achieved for certain tasks. Dimitrios Dimitriadis, José C. Segura, Luz García 0001, Alexandros Potamianos, Petros Maragos, Vassilis Pitsikalis |
INTERSPEECH | 1 |
| 2006 | An optimum microphone array post-filter for speech applicationsabstractThis paper proposes a post-filtering estimation scheme for mul-tichannel noise reduction. The proposed method extends and im-proves the existing Zelinski’s and, the most general and prominent, McCowan’s post-filtering methods that use the auto- and cross-spectral densities of the multichannel input signals to estimate the transfer function of the Wiener post-filter. A major drawback of these two speech enhancement algorithms is that the noise power spectrum at the beamformer’s output is over-estimated and there-fore the derived filters are sub-optimal in the Wiener sense. The proposed method deals with this problem and can be considered as an optimal post-filter that is appropriate for a wide variety of different noise fields. In experiments over real-noise multichannel recordings, the proposed technique is shown to obtain a significant headstart over the other methods in terms of signal-to-noise ratio and speech degradation measures. In addition it is used for ASR experiments where promising preliminary results are presented. Stamatios Lefkimmiatis, Dimitrios Dimitriadis, Petros Maragos |
INTERSPEECH | 2 |
| 2006 | Continuous energy demodulation methods and application to speech analysis
Dimitrios Dimitriadis, Petros Maragos |
Speech Commun. | 1 |
| 2005 | Auditory Teager energy cepstrum coefficients for robust speech recognitionabstractIn this paper, a feature extraction algorithm for robust speech recognition is introduced. The feature extraction algorithm is motivated by the human auditory processing and the nonlinear Teager-Kaiser energy operator that estimates the true energy of the source of a resonance. The proposed features are labeled as Teager Energy Cepstrum Coefficients (TECCs). TECCs are computed by first filtering the speech signal through a dense non constant-Q Gammatone filterbank and then by estimating the "true" energy of the signal's source, i.e., the short-time average of the output of the Teager-Kaiser energy operator. Error analysis and speech recognition experiments show that the TECCs and the mel frequency cepstrum coefficients (MFCCs) perform similarly for clean recording conditions; while the TECCs perform significantly better than the MFCCs for noisy recognition tasks. Specifically, relative word error rate improvement of 60% over the MFCC baseline is shown for the Aurora-3 database for the high-mismatch condition. Absolute error rate improvement ranging from 5% to 20% is shown for a phone recognition task in (various types of additive) noise. Dimitrios Dimitriadis, Petros Maragos, Alexandros Potamianos |
INTERSPEECH | 1 |
| 2005 | Robust AM-FM Features for Speech RecognitionabstractIn this letter, a nonlinear AM-FM speech model is used to extract robust features for speech recognition. The proposed features measure the amount of amplitude and frequency modulation that exists in speech resonances and attempt to model aspects of the speech acoustic information that the commonly used linear source-filter model fails to capture. The robustness and discriminability of the AM-FM features is investigated in combination with mel cepstrum coefficients (MFCCs). It is shown that these hybrid features perform well in the presence of noise, both in terms of phoneme-discrimination (J-measure) and in terms of speech recognition performance in several different tasks. Average relative error rate reduction up to 11% for clean and 46% for mismatched noisy conditions is achieved when AM-FM features are combined with MFCCs. Dimitrios Dimitriadis, Petros Maragos, Alexandros Potamianos |
IEEE Signal Process. Lett. | 1 |
| 2003 | Robust energy demodulation based on continuous models with application to speech recognitionabstractIn this paper, we develop improved schemes for simultaneous speech interpolation and demodulation based on continuous-time models. This leads to robust algorithms to estimate the instantaneous amplitudes and frequencies of the speech resonances and extract novel acoustic features for ASR. The continous-time models retain the excellent time resolution of the ESAs based on discrete energy operators and perform better in the presence of noise. We also introduce a robust algorithm based on the ESAs for amplitude compensation of the filtered signals. Furthermore, we use robust nonlinear modulation features to enhance the classic cepstrum-based features and use the augmented feature set for ASR applications. ASR experiments show promising evidence that the robust modulation features improve recognition. 1. Dimitrios Dimitriadis, Petros Maragos |
INTERSPEECH | 1 |
| 2002 | Modulation features for speech recognitionabstractAutomatic speech recognition (ASR) systems can benefit from including into their acoustic processing part new features that account for various nonlinear and time-varying phenomena during speech production. In this paper, we develop robust methods to extract novel acoustic features from speech signals of the modulation type based on time-varying models for speech analysis. Further, we integrate the new speech features with the standard linear ones (mel-frequency cesptrum) to develop a augmented set of acoustic features and demonstrate its efficacy by showing significant improvements in HMM-based word recognition over the TIMIT database. Dimitrios Dimitriadis, Petros Maragos, Alexandros Potamianos |
ICASSP | 1 |
| 2001 | An improved energy demodulation algorithm using splinesabstractA new algorithm is proposed for demodulating discrete-time AM-FM signals, which first interpolates the signals with smooth splines and then uses the continuous-time energy separation algorithm (ESA) based on the Teager-Kaiser energy operator. This spline-based ESA retains the excellent time resolution of the ESA based on discrete energy operators but performs better in the presence of noise. Further, its dependence on smooth splines allows some optimal trade-off between data fitting versus smoothing. Dimitrios Dimitriadis, Petros Maragos |
ICASSP | 1 |