EDBT 2026 Demo / reviewers in the wild / expert
Athanasios Mouchtaris
dblp:11/5225
· DBLP profile ↗
76ranked-venue papers
11as first author
33since 2021 · last 2025
0000-0001-7583-0189ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 59 · 8 first-author · 28 since 2021Artificial intelligence and machine learning · 32 · 4 first-author · 18 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MaZO: Masked Zeroth-Order Optimization for Multi-Task Fine-Tuning of Large Language ModelsabstractLarge language models have demonstrated exceptional capabilities across diverse tasks, but their fine-tuning demands significant memory, posing challenges for resource-constrained environments.Zeroth-order (ZO) optimization provides a memory-efficient alternative by eliminating the need for backpropagation.However, ZO optimization suffers from high gradient variance, and prior research has largely focused on single-task learning, leaving its application to multi-task learning unexplored.Multi-task learning is crucial for leveraging shared knowledge across tasks to improve generalization, yet it introduces unique challenges under ZO settings, such as amplified gradient variance and collinearity.In this paper, we present MaZO, the first framework specifically designed for multi-task LLM fine-tuning under ZO optimization.MaZO tackles these challenges at the parameter level through two key innovations: a weight importance metric to identify critical parameters and a multi-task weight update mask to selectively update these parameters, reducing the dimensionality of the parameter space and mitigating task conflicts.Experiments demonstrate that MaZO achieves state-of-the-art performance, surpassing even multi-task learning methods designed for firstorder optimization. Kai Zhen, Nathan Susanj, Athanasios Mouchtaris, Siegfried Kunzmann |
EMNLP | 5 |
| 2025 | QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language ModelsabstractJiajun Zhou, Yifan Yang, Kai Zhen, Ziyue Liu, Yequan Zhao, Ershad Banijamali, Athanasios Mouchtaris, Ngai Wong, Zheng Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jiajun Zhou 0004, Kai Zhen, Ziyue Liu 0003, Yequan Zhao, Seyed Ershad Banijamali, Athanasios Mouchtaris, Ngai Wong 0001, Zheng Zhang 0005 |
EMNLP | 7 |
| 2025 | Context-aware Dynamic Pruning for Speech Foundation ModelsabstractFoundation models, such as large language models, have achieved remarkable success in natural language processing and are evolving into models capable of handling multiple modalities.
Listening ability, in particular, is crucial for many applications, leading to research on building speech foundation models. However, the high computational cost of these large models presents a significant challenge for real-world applications. Although substantial efforts have been made to reduce computational costs, such as through pruning techniques, the majority of these approaches are applied primarily during the training phase for specific downstream tasks. In this study, we hypothesize that optimal pruned networks may vary based on contextual factors such as speaker characteristics, languages, and tasks. To address this, we propose a dynamic pruning technique that adapts to these contexts during inference without altering the underlying model. We demonstrated that we could successfully reduce inference time by approximately 30\% while maintaining accuracy in multilingual/multi-task scenarios. We also found that the obtained pruned structure offers meaningful interpretations based on the context, e.g., task-related information emerging as the dominant factor for efficient pruning. Masao Someki, Yifan Peng 0003, Siddhant Arora, Athanasios Mouchtaris, Grant P. Strimel, Shinji Watanabe 0001 |
ICLR | 5 |
| 2024 | AdaZeta: Adaptive Zeroth-Order Tensor-Train Adaption for Memory-Efficient Large Language Models Fine-TuningabstractFine-tuning large language models (LLMs) has achieved remarkable performance across various natural language processing tasks, yet it demands more and more memory as model sizes keep growing.To address this issue, the recently proposed Memory-efficient Zerothorder (MeZO) methods attempt to fine-tune LLMs using only forward passes, thereby avoiding the need for a backpropagation graph.However, significant performance drops and a high risk of divergence have limited their widespread adoption.In this paper, we propose the Adaptive Zeroth-order Tensor-Train Adaption (AdaZeta) framework, specifically designed to improve the performance and convergence of the ZO methods.To enhance dimension-dependent ZO estimation accuracy, we introduce a fast-forward, low-parameter tensorized adapter.To tackle the frequently observed divergence issue in large-scale ZO finetuning tasks, we propose an adaptive query number schedule that guarantees convergence.Detailed theoretical analysis and extensive experimental results on Roberta-Large and Llama-2-7B models substantiate the efficacy of our AdaZeta framework in terms of accuracy, memory efficiency, and convergence speed. 1 Kai Zhen, Seyed Ershad Banijamali, Athanasios Mouchtaris, Zheng Zhang 0005 |
EMNLP | 4 |
| 2024 | Max-Margin Transducer Loss: Improving Sequence-Discriminative Training Using a Large-Margin Learning StrategyabstractIn this work, we propose a novel sequence-discriminative training criterion for automatic speech recognition (ASR) based on the Conformer Transducer. Inspired by the large-margin classifier framework, we separate the "good" and the "bad" hypotheses in an N-best list produced from a pre-trained transducer model by a margin (τ), hence the term, Max-Margin Transducer (MMT) loss. It is observed that fine-tuning with the proposed loss achieves significant improvement over baseline transducer loss but does not outperform the state-of-the-art minimum word error rate (MWER) training. However, combining the proposed MMT loss with MWER surpasses the performance of either losses suggesting the complimentary nature of MWER and MMT losses. With the combined losses, we obtained 7.44% and 7.68% relative WER improvements on Librispeech test-clean and test-other sets, respectively, and up to 8.9% relative improvement on Multi-lingual Librispeech test sets. Rupak Vignesh Swaminathan, Grant P. Strimel, Ariya Rastrow, Sri Harish Reddy Mallidi, Kai Zhen, Hieu Duy Nguyen, Nathan Susanj, Athanasios Mouchtaris |
ICASSP | 8 |
| 2023 | Gated Contextual Adapters For Selective Contextual Biasing In Neural TransducersabstractNeural contextual biasing for end-to-end neural ASR transducers has shown significant improvements in the recognition of named entities, such as contact names or device names. However, it comes with the cost of increased compute, as the biasing layers (which are usually based on cross-attention) add complexity to the neural transducers. In this paper, we propose gated contextual biasing models that can estimate at runtime when contextual biasing is needed and can toggle it on or off. That way, contextual biasing does not run on every audio frame, but only on the frames where it can be helpful for correct ASR recognition. We show that our gated contextual biasing models can maintain all the performance improvements of contextual biasing while offering significant compute-cost saving, as the contextual biasing needs to be executed for fewer than 15% of the audio frames. Anastasios Alexandridis, Kanthashree Mysore Sathyendra, Grant P. Strimel, Feng-Ju Chang, Ariya Rastrow, Nathan Susanj, Athanasios Mouchtaris |
ICASSP | 7 |
| 2023 | Robust Acoustic And Semantic Contextual Biasing In Neural Transducers For Speech RecognitionabstractAttention-based contextual biasing approaches have shown significant improvements in the recognition of generic and/or personal rare-words in End-to-End Automatic Speech Recognition (E2E ASR) systems like neural transducers. These approaches employ crossattention to bias the model towards specific contextual entities injected as bias-phrases to the model. Prior approaches typically relied on subword encoders for encoding the bias phrases. However, subword tokenizations are coarse and fail to capture granular pronunciation information which is crucial for biasing based on acoustic similarity. In this work, we propose to use lightweight character representations to encode fine-grained pronunciation features to improve contextual biasing guided by acoustic similarity between the audio and the contextual entities (termed acoustic biasing). We further integrate pretrained neural language model (NLM) based encoders to encode the utterance's semantic context along with contextual entities to perform biasing informed by the utterance’s semantic context (termed semantic biasing). Experiments using a Conformer Transducer model on the Librispeech dataset show a 4.62% - 9.26% relative WER improvement on different biasing list sizes over the baseline contextual model when incorporating our proposed acoustic and semantic biasing approach. On a large-scale in-house dataset, we observe 7.91% relative WER improvement compared to our baseline model. On tail utterances, the improvements are even more pronounced with 36.80% and 23.40% relative WER improvements on Librispeech rare words and an in-house testset respectively. Xuandi Fu, Kanthashree Mysore Sathyendra, Ankur Gandhe, Grant P. Strimel, Ross McGowan, Athanasios Mouchtaris |
ICASSP | 7 |
| 2023 | Multilingual End-To-End Spoken Language Understanding For Ultra-Low Footprint ApplicationsabstractTiny Signal-to-Interpretation (TinyS2I) has been recently introduced as an ultra low-footprint end-to-end spoken language understanding (SLU) model. This architecture is capable of running in ultra resource constrained environments like voice assistant devices, while at the same time reducing latency. In this work, we propose an extension to TinyS2I and train a multilingual system supporting several languages. Multilingual TinyS2I models show little to no degradation compared to their monolingual counterparts. Increasing the network size in width and depth improves the classification accuracy for mono- and multilingual setups, with the multilingual one improving beyond the monolingual accuracy. This enables users to interact with the device in the language of their choice and dynamically switch between languages without an explicit language setting or accuracy degradation. Anastasios Alexandridis, Zach Trozenski, Joel Whiteman, Grant P. Strimel, Nathan Susanj, Athanasios Mouchtaris, Siegfried Kunzmann |
ICASSP | 7 |
| 2023 | Dual-Attention Neural Transducers for Efficient Wake Word Spotting in Speech RecognitionabstractWe present dual-attention neural biasing, an architecture designed to boost Wake Words (WW) recognition and improve inference time latency on speech recognition tasks. This architecture enables a dynamic switch for its runtime compute paths by exploiting WW spotting to select which branch of its attention networks to execute for an input audio frame. With this approach, we effectively improve WW spotting accuracy while saving runtime compute cost as defined by floating point operations (FLOPs). Using an in-house de-identified dataset, we demonstrate that the proposed dual-attention network can reduce the compute cost by 90% for WW audio frames, with only 1% increase in the number of parameters. This architecture improves WW F1 score by 16% relative and improves generic rare word error rate by 3% relative compared to the baselines. Saumya Y. Sahai, Thejaswi Muniyappa, Kanthashree Mysore Sathyendra, Anastasios Alexandridis, Grant P. Strimel, Ross McGowan, Ariya Rastrow, Feng-Ju Chang, Athanasios Mouchtaris, Siegfried Kunzmann |
ICASSP | 10 |
| 2023 | Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural TransducersabstractStreaming speech recognition architectures are employed for low-latency, real-time applications. Such architectures are often characterized by their causality. Causal architectures emit tokens at each frame, relying only on current and past signal, while non-causal models are exposed to a window of future frames at each step to increase predictive accuracy. This dichotomy amounts to a trade-off for real-time Automatic Speech Recognition (ASR) system design: profit from the low-latency benefit of strictly-causal architectures while accepting predictive performance limitations, or realize the modeling benefits of future-context models accompanied by their higher latency penalty. In this work, we relax the constraints of this choice and present the Adaptive Non-Causal Attention Transducer (ANCAT). Our architecture is non-causal in the traditional sense, but executes in a low-latency, streaming manner by dynamically choosing when to rely on future context and to what degree within the audio stream. The resulting mechanism, when coupled with our novel regularization algorithms, delivers comparable accuracy to non-causal configurations while improving significantly upon latency, closing the gap with their causal counterparts. We showcase our design experimentally by reporting comparative ASR task results with measures of accuracy and latency on both publicly accessible and production-scale, voice-assistant datasets. Grant P. Strimel, Brian John King, Martin Radfar, Ariya Rastrow, Athanasios Mouchtaris |
ICML | 6 |
| 2023 | Conmer: Streaming Conformer Without Self-attention for Interactive Voice Assistants
Martin Radfar, Paulina Lyskawa, Brandon Trujillo, Kai Zhen, Jahn Heymann, Denis Filimonov, Grant P. Strimel, Nathan Susanj, Athanasios Mouchtaris |
INTERSPEECH | 10 |
| 2022 | Tie Your Embeddings Down: Cross-Modal Latent Spaces for End-to-end Spoken Language UnderstandingabstractEnd-to-end (E2E) spoken language understanding (SLU) systems can infer the semantics of a spoken utterance directly from an audio signal. However, training an E2E system remains a challenge, largely due to the scarcity of paired audio-semantics data. In this paper, we consider an E2E system as a multi-modal model, with audio and text functioning as its two modalities, and use a cross-modal latent space (CMLS) architecture, where a shared latent space is learned between the ‘acoustic’ and ‘text’ embeddings. We propose using different multi-modal losses to explicitly align the acoustic embedding to the text embeddings (obtained via a semantically powerful pre-trained BERT model) in the latent space. We train the CMLS model on two publicly available E2E datasets and one internal dataset, across different cross-modal losses. Our proposed triplet loss function achieves the best performance. It achieves a relative improvement of 22.1% over an E2E model without a cross-modal space and a relative improvement of 2.8% over a previously published CMLS model using L2loss on our internal dataset. Bhuvan Agrawal, Samridhi Choudhary, Martin Radfar, Athanasios Mouchtaris, Ross McGowan, Nathan Susanj, Siegfried Kunzmann |
ICASSP | 5 |
| 2022 | Caching Networks: Capitalizing on Common Speech for ASRabstractWe introduce Caching Networks (CachingNets), a speech recognition network architecture capable of delivering faster, more accurate decoding by leveraging common speech patterns. By explicitly incorporating select sentences unique to each user into the network’s design, we show how to train the model as an extension of the popular sequence transducer architecture through a multitask learning procedure. We further propose and experiment with different phrase caching policies, which are effective for virtual voice-assistant (VA) applications, to complement the architecture. Our results demonstrate that by pivoting between different inference strategies on the fly, CachingNets can deliver significant performance improvements. Specifically, on an industrial-scale, VA ASR task, we observe up to 7.4% relative word error rate (WER) and 11% sentence error rate (SER) improvements with accompanied latency gains. Anastasios Alexandridis, Grant P. Strimel, Ariya Rastrow, Pavel Kveton, Maurizio Omologo, Siegfried Kunzmann, Athanasios Mouchtaris |
ICASSP | 8 |
| 2022 | TINYS2I: A Small-Footprint Utterance Classification Model with Contextual Support for On-Device SLUabstractOn-device spoken language understanding (SLU) offers the potential for significant latency savings compared to cloud-based processing, as the audio stream does not need to be transmitted to a server. We present Tiny Signal-to-interpretation (TinyS2I), an end-to-end on-device SLU approach which is focused on heavily resource constrained devices. TinyS2I brings latency reduction without accuracy degradation, by exploiting use cases when the distribution of utterances that users speak to a device is largely heavy-tailed. The model is tailored to process on-device frequent utterances with support for dynamic contextual content, while deferring all other requests to the cloud. Compared to a powerful baseline, we demonstrate that TinyS2I achieves comparable performance, while offering latency gains due to local processing. Anastasios Alexandridis, Kanthashree Mysore Sathyendra, Grant P. Strimel, Pavel Kveton, Athanasios Mouchtaris |
ICASSP | 6 |
| 2022 | Contextual Adapters for Personalized Speech Recognition in Neural TransducersabstractPersonal rare word recognition in end-to-end Automatic Speech Recognition (E2E ASR) models is a challenge due to the lack of training data. A standard way to address this issue is with shallow fusion methods at inference time. However, due to their dependence on external language models and the deterministic approach to weight boosting, their performance is limited. In this paper, we propose training neural contextual adapters for personalization in neural transducer based ASR models. Our approach can not only bias towards user-defined words, but also has the flexibility to work with pretrained ASR models. Using an in-house dataset, we demonstrate that contextual adapters can be applied to any general purpose pretrained ASR model to improve personalization. Our method outperforms shallow fusion, while retaining functionality of the pretrained models by not altering any of the model weights. We further show that the adapter style training is superior to full-fine-tuning of the ASR models on datasets with user-defined content. Kanthashree Mysore Sathyendra, Thejaswi Muniyappa, Feng-Ju Chang, Jinru Su, Grant P. Strimel, Athanasios Mouchtaris, Siegfried Kunzmann |
ICASSP | 7 |
| 2022 | A Neural Prosody Encoder for End-to-End Dialogue Act ClassificationabstractDialogue act classification (DAC) is a critical task for spoken language understanding in dialogue systems. Prosodic features such as energy and pitch have been shown to be useful for DAC. Despite their importance, little research has explored neural approaches to integrate prosodic features into end-to-end (E2E) DAC models which infer dialogue acts directly from audio signals. In this work, we propose an E2E neural architecture that takes into account the need for characterizing prosodic phenomena co-occurring at different levels inside an utterance. A novel part of this architecture is a learnable gating mechanism that assesses the importance of prosodic features and selectively retains core information necessary for E2E DAC. Our proposed model improves DAC accuracy by 1.07% absolute across three publicly available benchmark datasets. Dillon Knox, Martin Radfar, Grant P. Strimel, Nathan Susanj, Athanasios Mouchtaris, Maurizio Omologo |
ICASSP | 8 |
| 2022 | Knowledge Distillation via Module Replacing for Automatic Speech Recognition with Recurrent Neural Network Transducer
Kaiqi Zhao 0002, Animesh Jain, Nathan Susanj, Athanasios Mouchtaris, Lokesh Gupta, Ming Zhao 0002 |
INTERSPEECH | 5 |
| 2022 | ConvRNN-T: Convolutional Augmented Recurrent Neural Network Transducers for Streaming Speech Recognition
Martin Radfar, Rohit Barnwal, Rupak Vignesh Swaminathan, Feng-Ju Chang, Grant P. Strimel, Nathan Susanj, Athanasios Mouchtaris |
INTERSPEECH | 7 |
| 2022 | Compute Cost Amortized Transformer for Streaming ASRabstractWe present a streaming, Transformer-based end-to-end automatic speech recognition (ASR) architecture which achieves efficient neural inference through compute cost amortization. Our architecture creates sparse computation pathways dynamically at inference time, resulting in selective use of compute resources throughout decoding, enabling significant reductions in compute with minimal impact on accuracy. The fully differentiable architecture is trained end-to-end with an accompanying lightweight arbitrator mechanism operating at the frame-level to make dynamic decisions on each input while a tunable loss function is used to regularize the overall level of compute against predictive performance. We report empirical results from experiments using the compute amortized Transformer-Transducer (T-T) model conducted on LibriSpeech data. Our best model can achieve a 60% compute cost reduction with only a 3% relative word error rate (WER) increase. Jon Macoskey, Martin Radfar, Feng-Ju Chang, Brian John King, Ariya Rastrow, Athanasios Mouchtaris, Grant P. Strimel |
INTERSPEECH | 7 |
| 2022 | Sub-8-Bit Quantization Aware Training for 8-Bit Neural Network Accelerator with On-Device Speech Recognition
Kai Zhen, Hieu Duy Nguyen, Raviteja Chinta, Nathan Susanj, Athanasios Mouchtaris, Tariq Afzal, Ariya Rastrow |
INTERSPEECH | 5 |
| 2022 | Accelerator-Aware Training for Transducer-Based Speech RecognitionabstractMachine learning model weights and activations are represented in full-precision during training. This leads to performance degradation in runtime when deployed on neural network accelerator (NNA) chips, which leverage highly parallelized fixed-point arithmetic to improve runtime memory and latency. In this work, we replicate the NNA operators during the training phase, accounting for the degradation due to low-precision inference on the NNA in back-propagation. Our proposed method efficiently emulates NNA operations, thus foregoing the need to transfer quantization error-prone data to the Central Processing Unit (CPU), ultimately reducing the user perceived latency (UPL). We apply our approach to Recurrent Neural Network-Transducer (RNN-T), an attractive architecture for on-device streaming speech recognition tasks. We train and evaluate models on 270K hours of English data and show a 5-7% improvement in engine latency while saving up to 10% relative degradation in WER. Suhaila M. Shakiah, Rupak Vignesh Swaminathan, Hieu Duy Nguyen, Raviteja Chinta, Tariq Afzal, Nathan Susanj, Athanasios Mouchtaris, Grant P. Strimel, Ariya Rastrow |
SLT | 7 |
| 2022 | Sub-8-Bit Quantization for On-Device Speech Recognition: A Regularization-Free ApproachabstractFor on-device automatic speech recognition (ASR), quantization aware training (QAT) is ubiquitous to achieve the trade-off between model predictive performance and efficiency. Among existing QAT methods, one major drawback is that the quantization centroids have to be predetermined and fixed. To overcome this limitation, we introduce a regularization-free, “soft-to-hard” compression mechanism with self-adjustable centroids in a$\mu$-Law constrained space, resulting in a simpler yet more versatile quantization scheme, called General Quantizer (GQ). We apply GQ to ASR tasks using Recurrent Neural Network Transducer (RNN-T) and Conformer architectures on both LibriSpeech and de-identified far-field datasets. Without accuracy degradation, GQ can compress both RNN-T and Conformer into sub-8-bit, and for some RNN-T layers, to 1-bit for fast and accurate inference. We observe a 30.73% memory footprint saving and 31.75% user-perceived latency reduction compared to 8-bit QAT via physical device benchmarking. Kai Zhen, Martin Radfar, Hieu Duy Nguyen, Grant P. Strimel, Nathan Susanj, Athanasios Mouchtaris |
SLT | 6 |
| 2021 | Context-Aware Transformer Transducer for Speech RecognitionabstractEnd-to-end (E2E) automatic speech recognition (ASR) systems often have difficulty recognizing uncommon words, that appear infrequently in the training data. One promising method, to improve the recognition accuracy on such rare words, is to latch onto personalized/contextual information at inference. In this work, we present a novel context-aware transformer transducer (CATT) network that improves the state-of-the-art transformer-based ASR system by taking advantage of such contextual signals. Specifically, we propose a multi-head attention-based context-biasing network, which is jointly trained with the rest of the ASR sub-networks. We explore different techniques to encode contextual data and to create the final attention context vectors. We also leverage both BLSTM and pretrained BERT based models to encode contextual data and guide the network training. Using an in-house far-field dataset, we show that CATT, using a BERT based context encoder, improves the word error rate of the baseline transformer transducer and outperforms an existing deep contextual model by 24.2% and 19.4% respectively. Feng-Ju Chang, Martin Radfar, Athanasios Mouchtaris, Maurizio Omologo, Ariya Rastrow, Siegfried Kunzmann |
ASRU | 4 |
| 2021 | In Pursuit of Babel - Multilingual End-to-End Spoken Language UnderstandingabstractEnd-to-end spoken language understanding (E2E SLU) systems predict the utterance semantics directly from speech. So far, to the best of our knowledge, E2E models have only been trained to recognize the semantics for a single language. In this work we introduce the first multilingual E2E SLU system and present results across three languages - English, Spanish and French. We propose a transformer-based, multilingual acoustic encoder to predict intents, that leverages pre-training for both acoustic and linguistic modalities of the SLU model. It learns a robust, cross-modal latent space using a pre-trained multilingual BERT as a semantic teacher. The best performing model achieves relative improvements of 7.2% in a single language setting, 5-6% in two, and 4-6% in three language settings. An intent-wise analysis shows that semantic supervision becomes more important for shorter utterances, while providing an explicit language identifier at the input leads to lower intent classification errors. Samridhi Choudhary, Clement Chung, Athanasios Mouchtaris, Siegfried Kunzmann |
ASRU | 4 |
| 2021 | End-to-End Multi-Channel Transformer for Speech RecognitionabstractTransformers are powerful neural architectures that allow integrating different modalities using attention mechanisms. In this paper, we leverage the neural transformer architectures for multi-channel speech recognition systems, where the spectral and spatial information collected from different microphones are integrated using attention layers. Our multi-channel transformer network mainly consists of three parts: channel-wise self attention layers (CSA), cross-channel attention layers (CCA), and multi-channel encoder-decoder attention layers (EDA). The CSA and CCA layers encode the contextual relationship "within" and "between" channels and across time, respectively. The channel-attended outputs from CSA and CCA are then fed into the EDA layers to help decode the next token given the preceding ones. The experiments show that in a far-field in-house dataset, our method outperforms the baseline single-channel transformer, as well as the super-directive and neural beamformers cascaded with the transformers. Feng-Ju Chang, Martin Radfar, Athanasios Mouchtaris, Brian John King, Siegfried Kunzmann |
ICASSP | 3 |
| 2021 | Joint ASR and Language Identification Using RNN-T: An Efficient Approach to Dynamic Language SwitchingabstractConventional dynamic language switching enables seamless multilingual interactions by running several monolingual ASR systems in parallel and triggering the appropriate downstream components using a standalone language identification (LID) service. Since this solution is neither scalable nor cost- and memory-efficient, especially for on-device applications, we propose end-to-end, streaming, joint ASR-LID architectures based on the recurrent neural network transducer framework. Two key formulations are explored: (1) joint training using a unified output space for ASR and LID vocabularies, and (2) joint training viewed as multi-task optimization. We also evaluate the benefit of using auxiliary language information obtained on-the-fly from an acoustic LID classifier. Experiments with the English-Hindi language pair show that: (a) multi-task architectures perform better overall, and (b) the best joint architecture surpasses monolingual ASR (6.4–9.2% word error rate reduction) and acoustic LID (53.9–56.1% error rate reduction) baselines while reducing the overall memory footprint by up to 46%. Surabhi Punjabi, Harish Arsikere, Zeynab Raeesy, Chander Chandak, Nikhil Bhave, Ankish Bansal, Sergio Murillo, Ariya Rastrow, Andreas Stolcke, Jasha Droppo, Sri Garimella, Roland Maas, Mathieu Hans, Athanasios Mouchtaris, Siegfried Kunzmann |
ICASSP | 15 |
| 2021 | Sparsification via Compressed Sensing for Automatic Speech RecognitionabstractIn order to achieve high accuracy for machine learning (ML) applications, it is essential to employ models with a large number of parameters. Certain applications, such as Automatic Speech Recognition (ASR), however, require real-time interactions with users, hence compelling the model to have as low latency as possible. Deploying large scale ML applications thus necessitates model quantization and compression, especially when running ML models on resource constrained devices. For example, by forcing some of the model weight values into zero, it is possible to apply zero-weight compression, which reduces both the model size and model reading time from the memory. In the literature, such methods are referred to as sparse pruning. The fundamental questions are when and which weights should be forced to zero, i.e. be pruned. In this work, we propose a compressed sensing based pruning (CSP) approach to effectively address those questions. By reformulating sparse pruning as a sparsity inducing and compression-error reduction dual problem, we introduce the classic compressed sensing process into the ML model training process. Using ASR task as an example, we show that CSP consistently outperforms existing approaches in the literature. Kai Zhen, Hieu Duy Nguyen, Feng-Ju Chang, Athanasios Mouchtaris, Ariya Rastrow |
ICASSP | 4 |
| 2021 | Multi-Channel Transformer Transducer for Speech RecognitionabstractMulti-channel inputs offer several advantages over singlechannel, to improve the robustness of on-device speech recognition systems.Recent work on multi-channel transformer, has proposed a way to incorporate such inputs into end-to-end ASR for improved accuracy.However, this approach is characterized by a high computational complexity, which prevents it from being deployed in on-device systems.In this paper, we present a novel speech recognition model, Multi-Channel Transformer Transducer (MCTT), which features end-to-end multi-channel training, low computation cost, and low latency so that it is suitable for streaming decoding in on-device speech recognition.In a far-field in-house dataset, our MCTT outperforms stagewise multi-channel models with transformer-transducer up to 6.01% relative WER improvement (WERR).In addition, MCTT outperforms the multi-channel transformer up to 11.62% WERR, and is 15.8 times faster in terms of inference speed.We further show that we can improve the computational cost of MCTT by constraining the future and previous context in attention computations. Feng-Ju Chang, Martin Radfar, Athanasios Mouchtaris, Maurizio Omologo |
Interspeech | 3 |
| 2021 | Phonetically Induced Subwords for End-to-End Speech Recognition
Vasileios Papadourakis, Athanasios Mouchtaris, Maurizio Omologo |
Interspeech | 4 |
| 2021 | FANS: Fusing ASR and NLU for On-Device SLUabstractSpoken language understanding (SLU) systems translate voice input commands to semantics which are encoded as an intent and pairs of slot tags and values.Most current SLU systems deploy a cascade of two neural models where the first one maps the input audio to a transcript (ASR) and the second predicts the intent and slots from the transcript (NLU).In this paper, we introduce FANS, a new end-to-end SLU model that fuses an ASR audio encoder to a multi-task NLU decoder to infer the intent, slot tags, and slot values directly from a given input audio, obviating the need for transcription.FANS consists of a shared audio encoder and three decoders, two of which are seq-to-seq decoders that predict non null slot tags and slot values in parallel and in an auto-regressive manner.FANS neural encoder and decoders architectures are flexible which allows us to leverage different combinations of LSTM, self-attention, and attenders.Our experiments show compared to the state-of-the-art end-toend SLU models, FANS reduces ICER and IRER errors relatively by 30% and 7%, respectively, when tested on an in-house SLU dataset and by 0.86% and 2% absolute when tested on a public SLU dataset. Martin Radfar, Athanasios Mouchtaris, Siegfried Kunzmann, Ariya Rastrow |
Interspeech | 2 |
| 2021 | End-to-End Spoken Language Understanding for Generalized Voice AssistantsabstractEnd-to-end (E2E) spoken language understanding (SLU) systems predict utterance semantics directly from speech using a single model. Previous work in this area has focused on targeted tasks in fixed domains, where the output semantic structure is assumed a priori and the input speech is of limited complexity. In this work we present our approach to developing an E2E model for generalized SLU in commercial voice assistants (VAs). We propose a fully differentiable, transformer-based, hierarchical system that can be pretrained at both the ASR and NLU levels. This is then fine-tuned on both transcription and semantic classification losses to handle a diverse set of intent and argument combinations. This leads to an SLU system that achieves significant improvements over baselines on a complex internal generalized VA dataset with a 43% improvement in accuracy, while still meeting the 99% accuracy benchmark on the popular Fluent Speech Commands dataset. We further evaluate our model on a hard test set, exclusively containing slot arguments unseen in training, and demonstrate a nearly 20% improvement, showing the efficacy of our approach in truly demanding VA scenarios. Michael Saxon, Samridhi Choudhary, Joseph P. McKenna, Athanasios Mouchtaris |
Interspeech | 4 |
| 2021 | Evaluating the Vulnerability of End-to-End Automatic Speech Recognition Models to Membership Inference Attacks
Muhammad A. Shah, Joseph Szurley, Athanasios Mouchtaris, Jasha Droppo |
Interspeech | 4 |
| 2021 | CoDERT: Distilling Encoder Representations with Co-Learning for Transducer-Based Speech RecognitionabstractWe propose a simple yet effective method to compress an RNN-Transducer (RNN-T) through the well-known knowledge distillation paradigm. We show that the transducer's encoder outputs naturally have a high entropy and contain rich information about acoustically similar word-piece confusions. This rich information is suppressed when combined with the lower entropy decoder outputs to produce the joint network logits. Consequently, we introduce an auxiliary loss to distill the encoder logits from a teacher transducer's encoder, and explore training strategies where this encoder distillation works effectively. We find that tandem training of teacher and student encoders with an inplace encoder distillation outperforms the use of a pre-trained and static teacher transducer. We also report an interesting phenomenon we refer to as implicit distillation, that occurs when the teacher and student encoders share the same decoder. Our experiments show 5.37-8.4% relative word error rate reductions (WERR) on in-house test sets, and 5.05-6.18% relative WERRs on LibriSpeech test sets. Rupak Vignesh Swaminathan, Brian John King, Grant P. Strimel, Jasha Droppo, Athanasios Mouchtaris |
Interspeech | 5 |
| 2020 | Multilingual Grapheme-To-Phoneme Conversion with Byte RepresentationabstractGrapheme-to-phoneme (G2P) models convert a written word into its corresponding pronunciation and are essential components in automatic-speech-recognition and text-to-speech systems. Recently, the use of neural encoder-decoder architectures has substantially improved G2P accuracy for mono- and multi-lingual cases. However, most multilingual G2P studies focus on sets of languages that share similar graphemes, such as European languages. Multilingual G2P for languages from different writing systems, e.g. European and East Asian, remains an understudied area. In this work, we propose a multilingual G2P model with byte-level input representation to accommodate different grapheme systems, along with an attention-based Transformer architecture. We evaluate the performance of both character-level and byte-level G2P using data from multiple European and East Asian locales. Models using byte representation yield 16.2%– 50.2% relative word error rate improvement over character-based counterparts for mono- and multi-lingual use cases. In addition, byte-level models are 15.0%–20.1% smaller in size. Our results show that byte is an efficient representation for multilingual G2P with languages having large grapheme vocabularies. Mingzhi Yu, Hieu Duy Nguyen, Alex Sokolov, Jack Lepird, Kanthashree Mysore Sathyendra, Samridhi Choudhary, Athanasios Mouchtaris, Siegfried Kunzmann |
ICASSP | 7 |
| 2020 | Semantic Complexity in End-to-End Spoken Language UnderstandingabstractEnd-to-end spoken language understanding (SLU) models are a class of model architectures that predict semantics directly from speech. Because of their input and output types, we refer to them as speech-to-interpretation (STI) models. Previous works have successfully applied STI models to targeted use cases, such as recognizing home automation commands, however no study has yet addressed how these models generalize to broader use cases. In this work, we analyze the relationship between the performance of STI models and the difficulty of the use case to which they are applied. We introduce empirical measures of dataset semantic complexity to quantify the difficulty of the SLU tasks. We show that near-perfect performance metrics for STI models reported in the literature were obtained with datasets that have low semantic complexity values. We perform experiments where we vary the semantic complexity of a large, proprietary dataset and show that STI model performance correlates with our semantic complexity measures, such that performance increases as complexity values decrease. Our results show that it is important to contextualize an STI model's performance with the complexity values of its training dataset to reveal the scope of its applicability. Joseph P. McKenna, Samridhi Choudhary, Michael Saxon, Grant P. Strimel, Athanasios Mouchtaris |
INTERSPEECH | 5 |
| 2020 | Quantization Aware Training with Absolute-Cosine Regularization for Automatic Speech Recognition
Hieu Duy Nguyen, Anastasios Alexandridis, Athanasios Mouchtaris |
INTERSPEECH | 3 |
| 2020 | End-to-End Neural Transformer Based Spoken Language UnderstandingabstractSpoken language understanding (SLU) refers to the process of inferring the semantic information from audio signals. While the neural transformers consistently deliver the best performance among the state-of-the-art neural architectures in field of natural language processing (NLP), their merits in a closely related field, i.e., spoken language understanding (SLU) have not beed investigated. In this paper, we introduce an end-to-end neural transformer-based SLU model that can predict the variable-length domain, intent, and slots vectors embedded in an audio signal with no intermediate token prediction architecture. This new architecture leverages the self-attention mechanism by which the audio signal is transformed to various sub-subspaces allowing to extract the semantic context implied by an utterance. Our end-to-end transformer SLU predicts the domains, intents and slots in the Fluent Speech Commands dataset with accuracy equal to 98.1 \%, 99.6 \%, and 99.6 \%, respectively and outperforms the SLU models that leverage a combination of recurrent and convolutional neural networks by 1.4 \% while the size of our model is 25\% smaller than that of these architectures. Additionally, due to independent sub-space projections in the self-attention layer, the model is highly parallelizable which makes it a good candidate for on-device SLU. Martin Radfar, Athanasios Mouchtaris, Siegfried Kunzmann |
INTERSPEECH | 2 |
| 2018 | Normalization of Partly Overlapping Audio Recordings from the Same Event Based on Relative Signal PowersabstractExploiting correlations in the audio, several works in the past have demonstrated the ability to automatically match and synchronize user-generated video or audio files of the same event. Such tools solve for the unknown starting and ending time of each available recording along the event time-line and open the way for collaborative content production approaches. However, a source of difficulty for collaborative processing approaches related to audio is the fact that the different audio recordings may be available at significantly different signal levels. In this paper, we present a normalization approach to automatically define gains for all the recordings so that the variations in the signal levels among different recordings are suppressed. We show that normalization is trivial when all recordings share the same time support but the same process is non-trivial when the recordings partly overlap along time, especially if the acoustic event is characterized by high dynamic variations. We demonstrate the efficiency of the proposed approach under various conditions based on real examples of user-generated audio recordings. Nikolaos Stefanakis, Athanasios Mouchtaris |
ICASSP | 2 |
| 2018 | Multiple Source Location Estimation on a Dataset of Real Recordings in a Wireless Acoustic Sensor NetworkabstractRecently, wireless acoustic sensor networks (WASNs) have received significant attention from the research community and a variety of methods have been proposed for numerous applications, such as location estimation and speech enhancement. The lack of publicly available datasets with signals recorded in WASNs, presents difficulties in obtaining consistent performance indicators across the different approaches. In this paper, we present and release a dataset of real recorded signals in an outdoor WASN comprised of four microphone arrays. Our dataset consists of several speakers recorded at various locations within the WASN and can be used for benchmarking purposes. We also present location estimation results using our real recorded dataset. Our results can serve as a baseline indicator of localization performance of single and multiple sources in a real environment. Anastasios Alexandridis, Anthony Griffin, Athanasios Mouchtaris |
MMSP | 3 |
| 2018 | Multiple Sound Source Location Estimation in Wireless Acoustic Sensor Networks Using DOA Estimates: The Data-Association ProblemabstractIn this paper, we consider the data-association problem for the localization of multiple sound sources in a wireless acoustic sensor network, where each node is a microphone array, using direction of arrival (DOA) estimates. The data-association problem arises because the central node that receives the multiple DOA estimates from the nodes cannot know to which source they belong. Hence, the DOAs from the different nodes that correspond to the same source must be found in order to perform accurate localization. We present a method to identify the correct association of DOAs to the sources and thus accurately estimate their locations. Our method results in high association and localization accuracy in realistic scenarios with missed detections, reverberation, noise, and moving sources and outperforms other recently proposed methods. It also incorporates a bitrate reduction scheme in order to keep the amount of information that needs to be transmitted in the network at low levels without affecting performance. Anastasios Alexandridis, Athanasios Mouchtaris |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Towards wireless acoustic sensor networks for location estimation and counting of multiple speakers in real-life conditionsabstractSpeaker localization and counting in real-life conditions remains a challenging task. The computational burden, transmission usage and synchronization issues pose several limitations. Moreover, the physical characteristics of real speakers in terms of directivity pattern and orientation, as well as restrictions in the microphone array positioning, which commonly have to be placed close to walls, deteriorate the localization performance. In this paper, we propose a localization and counting method that accounts for the adjacent wall reflections and evaluate it using a dataset of real recorded signals of actual speakers that we collected. Our dataset is publicly available to foster further investigation towards localization in real-life scenarios. Anastasios Alexandridis, Nikolaos Stefanakis, Athanasios Mouchtaris |
ICASSP | 3 |
| 2017 | DOA estimation with histogram analysis of spatially constrained active intensity vectorsabstractThe active intensity vector (AIV) is a common descriptor of the sound field. In microphone array processing, AIV is commonly approximated with beamforming operations and utilized as a direction of arrival (DOA) estimator. However, in its original form, it provides inaccurate estimates in sound field conditions where coherent sound sources are simultaneously active. In this work we utilize a higher order intensity-based DOA estimator on spatially-constrained regions (SCR) to overcome such limitations. We then apply 1-dimensional (1D) histogram processing on the noisy estimates for multiple DOA estimation. The performance of the estimator is shown with a 7-channel mobile microphone array, in reverberant conditions and under different signal-to-noise ratios. Symeon Delikaris-Manias, Despoina Pavlidi, Athanasios Mouchtaris, Ville Pulkki |
ICASSP | 3 |
| 2017 | Automatic matching and synchronization of user generated videos from a large scale sport eventabstractExploiting correlations in the audio, several works in the past have demonstrated the ability to automatically match and synchronize User Generated Video (UGV) files of the same event. In this paper, we focus on the challenging acoustic environment of a large scale athletic event. We show that the chanting of the crowd produces an acoustic background common in the audio streams of different UGVs and we design a novel audio fingerprinting method for organizing the UGV collection based on that content. Results presented with recordings from a crowded football match demonstrate that the proposed approach provides significantly better audio matching performance in comparison to three of the most well known audio fingerprinting techniques. Nikolaos Stefanakis, Stavros Chonianakis, Athanasios Mouchtaris |
ICASSP | 3 |
| 2017 | Maximum component elimination in mixing of user generated audio recordingsabstractUser generated content is gradually being recognized for its remarkable potential to enrich the professionally broadcasted content, but also as the means to provide acceptable quality audiovisual content for public events where professional coverage is absent. This potential is particularly interesting with respect to the audio modality, as a multitude of temporally overlapping User Generated audio Recordings (UGRs) may be utilized in order to provide a multichannel recording of the captured acoustic event. In this paper, we formulate a simple audio mixing approach called Maximum Component Elimination (MCE) to process a multiplicity of synchronized UGRs in a collaborative fashion. Operating in the Time-Frequency (TF) domain, MCE relies on the use of binary weights in order to selectively prevent certain TF components from individual UGRs to enter in the final mix. Results from a listening test indicate that the proposed mechanism is very efficient in suppressing foreground speech interference, removing inappropriate content from the audio mix and concealing the identities of individuals whose voices are unintentionally captured by the recording devices. Furthermore, it is shown that audio mixtures produced with MCE improve the user experience compared to the more classical use case where each UGR is consumed individually. Nikolaos Stefanakis, Athanasios Mouchtaris |
MMSP | 2 |
| 2017 | Perpendicular Cross-Spectra Fusion for Sound Source Localization With a Planar Microphone ArrayabstractMultiple sound source localization in reverberant environments stands as one of the most difficult challenges for many applications related to microphone array signal processing. In this paper, we describe perpendicular cross-spectra fusion (PCSF), a new direction-of-arrival (DOA) estimation algorithm, which utilizes an analytic formula for direction estimation in the time-frequency (TF) domain. Inherent to this technique is the presence of multiple direction estimation subsystems which operate in parallel, producing a multiplicity of candidate DOAs at each TF point. We define a metric of coherence, based on the property of divergence of the different DOA estimators, for assessing the reliability of different signal portions, so that only TF bins with a high quality of directional information are exploited for local DOA estimation. The resulting collection of local DOAs is provided as input to a recently proposed histogram processing approach, which is based on matching pursuit. Results based on simulation and real recordings illustrate the advantages of PCSF compared to other DOA estimation techniques subject to the same histogram-based processing, in the context of real-time multiple source localization and counting; improved performance in reverberant conditions and high tolerance to diffuse and common mode noise. Nikolaos Stefanakis, Despoina Pavlidi, Athanasios Mouchtaris |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Corrections to "Perpendicular Cross-Spectra Fusion for Sound Source Localization With a Planar Microphone Array"
Nikolaos Stefanakis, Despoina Pavlidi, Athanasios Mouchtaris |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | A Survey of Sound Source Localization Methods in Wireless Acoustic Sensor NetworksabstractWireless acoustic sensor networks (WASNs) are formed by a distributed group of acoustic-sensing devices featuring audio playing and recording capabilities. Current mobile computing platforms offer great possibilities for the design of audio-related applications involving acoustic-sensing nodes. In this context, acoustic source localization is one of the application domains that have attracted the most attention of the research community along the last decades. In general terms, the localization of acoustic sources can be achieved by studying energy and temporal and/or directional features from the incoming sound at different microphones and using a suitable model that relates those features with the spatial location of the source (or sources) of interest. This paper reviews common approaches for source localization in WASNs that are focused on different types of acoustic features, namely, the energy of the incoming signals, their time of arrival (TOA) or time difference of arrival (TDOA), the direction of arrival (DOA), and the steered response power (SRP) resulting from combining multiple microphone signals. Additionally, we discuss methods not only aimed at localizing acoustic sources but also designed to locate the nodes themselves in the network. Finally, we discuss current challenges and frontiers in this field. Maximo Cobos, Fabio Antonacci, Anastasios Alexandridis, Athanasios Mouchtaris, Bowon Lee |
Wirel. Commun. Mob. Comput. | 4 |
| 2017 | Wireless Acoustic Sensor Networks and Applications
Maximo Cobos, Fabio Antonacci, Athanasios Mouchtaris, Bowon Lee |
Wirel. Commun. Mob. Comput. | 3 |
| 2016 | 3D DOA estimation of multiple sound sources based on spatially constrained beamforming driven by intensity vectorsabstractSound source localization in three dimensions with microphone arrays is an active field of research, applicable in sound enhancement, source separation, and sound field analysis. In this contribution we propose a method for three dimensional multiple sound source localization in reverberant environments. We employ a spatially constrained steered response beamformer on a spherical sector centered at the direction of arrival (DOA) estimates of the intensity vector. Experiments are performed in both simulated and real acoustical environments with a spherical microphone array for multiple sound sources under different reverberation and signal-to-noise ratio (SNR) conditions. The performance of the proposed method is compared with our previously proposed work and a subspace method in the spherical harmonic domain. The results demonstrate a significant improvement in terms of localization accuracy. Despoina Pavlidi, Symeon Delikaris-Manias, Ville Pulkki, Athanasios Mouchtaris |
ICASSP | 4 |
| 2015 | Foreground suppression for capturing and reproduction of crowded acoustic environmentsabstractTraditionally, sensor arrays and spatial filtering aim to enhance individual sources by suppressing ambient noise and reverberation. In this paper, the exactly opposite problem is examined, that of suppressing individual sources in favour of the ambient sound and of the whole acoustic scene in general. We consider a compact circular sensor array which is embedded in a crowded ambient acoustic environment and is at the same time prone to interference from directional speech originating from multiple nearby speakers. We propose a method for suppressing the undesired components and we compare its performance with two established approaches in spatial audio processing, namely, direct-to-diffuse decomposition and Primary-Ambient Extraction (PAE). Experimental results and a listening test which are presented illustrate the superiority of our method. Nikolaos Stefanakis, Athanasios Mouchtaris |
ICASSP | 2 |
| 2015 | Localizing multiple audio sources in a wireless acoustic sensor network
Anthony Griffin, Anastasios Alexandridis, Despoina Pavlidi, Yiannis Mastorakis, Athanasios Mouchtaris |
Signal Process. | 5 |
| 2015 | Speech Analysis and Synthesis with a Computationally Efficient Adaptive Harmonic ModelabstractHarmonic models have to be both precise and fast in order to represent the speech signal adequately and be able to process large amount of data in a reasonable amount of time.For these purposes, the full-band adaptive Harmonic Model (aHM) used by the Adaptive Iterative Refinement (AIR) algorithm has been proposed in order to accurately model the perceived characteristics of a speech signal.Even though aHM-AIR is precise, it lacks the computational efficiency that would make its use convenient for large databases.The Least Squares (LS) solution used in the original aHM-AIR accounts for most of the computational load.In a previous paper, we suggested a Peak Picking (PP) approach as a substitution to the LS solution.In order to integrate the adaptivity scheme of aHM in the PP approach, an adaptive Discrete Fourier Transform (aDFT), whose frequency basis can fully follow the variations of the f0 curve, was also proposed.In this article, we complete the previous publication by evaluating the above methods for the whole analysis process of a speech signal.Evaluations have shown an average time reduction by four times using Peak Picking and aDFT compared to the LS solution.Additionally, based on formal listening tests, when using Peak Picking and aDFT, the quality of the re-synthesis is preserved compared to the original LS-based approach. Veronica Morfi, Gilles Degottex, Athanasios Mouchtaris |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | A computationally efficient refinement of the fundamental frequency estimate for the Adaptive Harmonic ModelabstractThe full-band Adaptive Harmonic Model (aHM) can be used by the Adaptive Iterative Refinement (AIR) algorithm to accurately model the perceived characteristics of a speech recording. However, the Least Squares (LS) solution used in the current aHM-AIR makes the f0refinement in AIR time consuming, limiting the use of this algorithm for large databases. In this paper, a Peak Picking (PP) approach is suggested as a substitution to the LS solution. In order to integrate the adaptivity scheme of aHM in the PP approach, an adaptive Discrete Fourier Transform (aDFT) is also suggested in this paper, whose frequency basis can fully follow the frequency variations of the f0curve. Evaluations have shown an average time reduction of 5.5 times compared to the LS solution approach, while the quality of the resynthesis is preserved compared to the original aHM-AIR. Veronica Morfi, Gilles Degottex, Athanasios Mouchtaris |
ICASSP | 3 |
| 2013 | Directional coding of audio using a circular microphone arrayabstractWe propose a real-time method for coding an acoustic environment based on estimating the Direction-of-Arrival (DOA) and reproducing it using an arbitrary loudspeaker configuration or headphones. We encode the sound field with the use of one audio signal and side-information. The audio signal can be further encoded with an MP3 encoder to reduce the bitrate. We investigate how such coding can affect the spatial impression and sound quality of spatial audio reproduction. Also, we propose a lossless efficient compression scheme for the side-information. Our method is compared with other recently proposed microphone array based methods for directional coding. Listening tests confirm the effectiveness of our method in achieving excellent reconstruction of the sound field while maintaining the sound quality at high levels. Anastasios Alexandridis, Anthony Griffin, Athanasios Mouchtaris |
ICASSP | 3 |
| 2013 | Real-Time Multiple Sound Source Localization and Counting Using a Circular Microphone ArrayabstractIn this work, a multiple sound source localization and counting method is presented, that imposes relaxed sparsity constraints on the source signals. A uniform circular microphone array is used to overcome the ambiguities of linear arrays, however the underlying concepts (sparse component analysis and matching pursuit-based operation on the histogram of estimates) are applicable to any microphone array topology. Our method is based on detecting time-frequency (TF) zones where one source is dominant over the others. Using appropriately selected TF components in these “single-source” zones, the proposed method jointly estimates the number of active sources and their corresponding directions of arrival (DOAs) by applying a matching pursuit-based approach to the histogram of DOA estimates. The method is shown to have excellent performance for DOA estimation and source counting, and to be highly suitable for real-time applications due to its low complexity. Through simulations (in various signal-to-noise ratio conditions and reverberant environments) and real environment experiments, we indicate that our method outperforms other state-of-the-art DOA and source counting methods in terms of accuracy, while being significantly more efficient in terms of computational complexity. Despoina Pavlidi, Anthony Griffin, Matthieu Puigt, Athanasios Mouchtaris |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | Real-time multiple sound source localization using a circular microphone array based on single-source confidence measuresabstractWe propose a novel real-time adaptative localization approach for multiple sources using a circular array, in order to suppress the localization ambiguities faced with linear arrays, and assuming a weak sound source sparsity which is derived from blind source separation methods. Our proposed method performs very well both in simulations and in real conditions at 50% real-time. Despoina Pavlidi, Matthieu Puigt, Anthony Griffin, Athanasios Mouchtaris |
ICASSP | 4 |
| 2011 | Perceptually-Driven Scalable MDCT Enhancement of Compressed Audio Based on Statistical ConversionabstractMany state-of-the-art audio codecs operating in a transform domain provide scalability as a core function by allowing to selectively subtract bits -- usually according to a nonperceptual criterion from the full bit rate data stream. This work presents a different, or even reverse, scalability approach in which a scalable codec can selectively add perceptually significant bits to a low bit rate data stream. The scalable enhancement algorithm presented here operates in the Modified Discrete Cosine Transform domain, which is popular among perceptual audio transform encoders, but its extension on other domains is straightforward. By exploiting the information of an existing low bit rate base layer, the algorithm adds perceptually significant data to the data stream according to a psycho acoustic model, and improves the audio quality at a fraction of the bit rate that would normally be required for the encoding or transmission of the whole audio piece of the same quality. Applications of this can be found in packet retransmission schemes of compressed audio networks and in remote audio enhancement. Demetrios Cantzos, Athanasios Mouchtaris, Chris Kyriakakis |
ISM | 2 |
| 2011 | Single-Channel and Multi-Channel Sinusoidal Audio Coding Using Compressed SensingabstractCompressed sensing (CS) samples signals at a much lower rate than the Nyquist rate if they are sparse in some basis. In this paper, the CS methodology is applied to sinusoidally modeled audio signals. As this model is sparse by definition in the frequency domain (being equal to the sum of a small number of sinusoids), we investigate whether CS can be used to encode audio signals at low bitrates. In contrast to encoding the sinusoidal parameters (amplitude, frequency, phase) as current state-of-the-art methods do, we propose encoding few randomly selected samples of the time-domain description of the sinusoidal component (per signal segment). The potential of applying compressed sensing both to single-channel and multi-channel audio coding is examined. The listening test results are encouraging, indicating that the proposed approach can achieve comparable performance to that of state-of-the-art methods. Given that CS can lead to novel coding systems where the sampling and compression operations are combined into one low-complexity step, the proposed methodology can be considered as an important step towards applying the CS framework to audio coding applications. Anthony Griffin, Toni Hirvonen, Christos Tzagkarakis, Athanasios Mouchtaris, Panagiotis Tsakalides |
IEEE Trans. Speech Audio Process. | 4 |
| 2010 | Top-down strategies in parameter selection of sinusoidal modeling of audioabstractSinusoidal modeling of audio requires the model parameters to be selected by analyzing the original signal spectrum. This paper proposes two improvements in sinusoidal selection by considering how psychoacoustic masking curves can be calculated using a top-down strategy in certain situations. First, a non-iterative component selection method to be used in combination with an added residual signal is presented. Tests indicate computational gain and quality increase when the method is used with a noise-synthesized residual. Secondly, the estimation of the masking curve in binaural listening when signals are panned is considered. Tests show that knowledge of the degree of panning is beneficial when heavy panning is applied to simultaneously rendered audio object signals. Toni Hirvonen, Athanasios Mouchtaris |
ICASSP | 2 |
| 2010 | Sinusoidal spatial audio coding for low-bitrate binaural reproductionabstractA binaural audio synthesis system based on sinusoidal modeling is proposed for spatial, low-bitrate audio coding utilized for example in teleconference applications. The system transmits monaural sinusoidal parameters of a downmix signal, from which the left and right binaural signals are synthesized according to the directional metadata at the receiver. Typical sinusoidal synthesis methods, as well as the effectiveness of a monaural frequency masking model, are evaluated in binaural context. Furthermore, a method for binaural noise residual synthesis and efficiency improvements for HRTF parameter acquisition are suggested. Tests utilizing speech signals indicate that sinusoidal modeling is an attractive technique for applications such as the proposed one. Toni Hirvonen, Athanasios Mouchtaris |
ICASSP | 2 |
| 2009 | Bandwidth extension of low bitrate compressed audio based on statistical conversionabstractAlgorithmic and protocol constraints of most low bitrate compression schemes lead to audio signals of low bandwidth and, inevitably, of low perceptual audio quality. Audio bandwidth extension methods address this problem by reconstructing the high frequency spectrum of a degraded signal based on information from the low frequency part. In this work, a novel audio bandwidth extension method is presented in which high frequency reconstruction is achieved through statistical conversion between the low frequency spectrum of the compressed signal and the high frequency part of the uncompressed signal's spectrum. Even though no psychoacoustic model is used, quality evaluation tests show that the proposed method has similar performance to one of the most recent, state-of-the-art, bandwidth extension schemes. Demetrios Cantzos, Athanasios Mouchtaris, Chris Kyriakakis |
ICME | 2 |
| 2009 | Encoding the sinusoidal model of an audio signal using compressed sensingabstractIn this paper, the compressed sensing (CS) methodology is applied to the harmonic part of sinusoidally-modeled audio signals. As this part of the model is sparse by definition in the frequency domain, we investigate how CS can be used to encode this signal at low bitrates, instead of encoding the sinusoidal parameters (amplitude, frequency, phase) as current state-of-the-art methods do. We extend our previous work by considering an improved system model, by comparing our model to other schemes, and exploring the effect of incorrectly reconstructed frames. We show that encouraging results can be obtained by our approach, although inferior at this point compared to state-of-the-art. Good performance is obtained using 24 bits per sinusoid as indicated by our listening tests. Anthony Griffin, Toni Hirvonen, Athanasios Mouchtaris, Panagiotis Tsakalides |
ICME | 3 |
| 2009 | A Multichannel Sinusoidal Model Applied to Spot Microphone Signals for Immersive AudioabstractIn this paper, a multichannel version of the sinusoids plus noise model (also known as deterministic plus stochastic decomposition) is proposed and applied to spot microphone signals of a music recording. These are the recordings captured by the various microphones placed in a venue, before the mixing process produces the final multichannel audio mix. Coding these microphone signals makes them available to the decoder, allowing for interactive audio reproduction which is a necessary component in immersive audio applications. The proposed model uses a single reference audio signal in order to derive a noise signal per spot microphone. This noise signal can significantly enhance the sinusoidal representation of the corresponding spot signal. The reference can be one of the spot signals or a downmix, depending on the application. Thus, for a collection of multiple spot signals, only the reference is fully encoded (e.g., as an MP3 monophonic signal). For the remaining spot signals, their sinusoidal parameters and corresponding noise spectral envelopes are retained and coded, resulting in bitrates for this side information in the order of 15 kb/s for perceptual performance above the 4.0 grade on the mean opinion score (MOS) scale. Christos Tzagkarakis, Athanasios Mouchtaris, Panagiotis Tsakalides |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Conditional Vector Quantization for Voice ConversionabstractVoice conversion methods have the objective of transforming speech spoken by a particular source speaker, so that it sounds as if spoken by a different target speaker. The majority of voice conversion methods is based on transforming the short-time spectral envelope of the source speaker, based on derived correspondences between the source and target vectors using training speech data from both speakers. These correspondences are usually obtained by segmenting the spectral vectors of one or both speakers into clusters, using soft (GMM-based) or hard (VQ-based) clustering. Here, we propose that voice conversion performance can be improved by taking advantage of the fact that often the relationship between the source and target vectors is one-to-many. In order to illustrate this, we propose that a VQ approach namely constrained vector quantization (CVQ), can be used for voice conversion. Results indicate that indeed such a relationship between the source and target data exists and can be exploited by following a CVQ-based function for voice conversion. Athanasios Mouchtaris, Yannis Agiomyrgiannakis, Yannis Stylianou |
ICASSP (4) | 1 |
| 2007 | Enhanced Multichannel Audio Resynthesis Through Residual Processing and Features AlignmentabstractMultichannel audio refers to a widespread technology that enables audio rendering through multiple channels. Audio reproduction with multiple channels has the advantage of recreating the acoustic scene with unprecedented fidelity and of immersing the listener in an acoustic environment that is virtually indistinguishable from reality. However, one of the greatest challenges of multichannel audio is its high storage and transmission requirements especially since accurate rendering through as many possible channels is the main purpose. Audio resynthesis addresses this issue by enabling us to recreate a set of channels at the receiver end by transmitting only one source channel. We propose a new, enhanced, approach on multichannel audio resynthesis which involves a novel residual processing technique and a features alignment method that significantly increase the resynthesis accuracy. Our results show that this latest method leads to higher audio quality and allows for the robust treatment of any type of multichannel signal set. Demetrios Cantzos, Athanasios Mouchtaris, Chris Kyriakakis |
ICME | 2 |
| 2007 | A Spectral Conversion Approach to Single-Channel Speech EnhancementabstractIn this paper, a novel method for single-channel speech enhancement is proposed, which is based on a spectral conversion feature denoising approach. Spectral conversion has been applied previously in the context of voice conversion, and has been shown to successfully transform spectral features with particular statistical properties into spectral features that best fit (with the constraint of a piecewise linear transformation) different target statistics. This spectral transformation is applied as an initialization step to two well-known single channel enhancement methods, namely the iterative Wiener filter (IWF) and a particular iterative implementation of the Kalman filter. In both cases, spectral conversion is shown here to provide a significant improvement as opposed to initializations using the spectral features directly from the noisy speech. In essence, the proposed approach allows for applying these two algorithms in a user-centric manner, when “clean” speech training data are available from a particular speaker. The extra step of spectral conversion is shown to offer significant advantages regarding output signal-to-noise ratio (SNR) improvement over the conventional initializations, which can reach 2 dB for the IWF and 6 dB for the Kalman filtering algorithm, for low input SNRs and for white and colored noise, respectively. Athanasios Mouchtaris, Jan Van der Spiegel, Paul Mueller, Panagiotis Tsakalides |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Musical Genre Classification VIA Generalized Gaussian and Alpha-Stable ModelingabstractThis paper describes a novel methodology for automatic musical genre classification based on a feature extraction/statistical similarity measurement approach. First, we perform a 1-D wavelet decomposition of the music signal and we model the resulting subband coefficients using the generalized Gaussian density (GGD) and the alpha-stable distribution. Subsequently, the GGD and alpha-stable distribution parameters are estimated during the feature extraction step, while the similarity between two music signals is measured by employing the Kullback-Leibler divergence (KLD) between their corresponding estimated wavelet distributions. We evaluate the performance of the proposed methodology by using a dataset consisting of six different musical genre sets Christos Tzagkarakis, Athanasios Mouchtaris, Panagiotis Tsakalides |
ICASSP (5) | 2 |
| 2006 | Nonparallel training for voice conversion based on a parameter adaptation approachabstractThe objective of voice conversion algorithms is to modify the speech by a particular source speaker so that it sounds as if spoken by a different target speaker. Current conversion algorithms employ a training procedure, during which the same utterances spoken by both the source and target speakers are needed for deriving the desired conversion parameters. Such a (parallel) corpus, is often difficult or impossible to collect. Here, we propose an algorithm that relaxes this constraint, i.e., the training corpus does not necessarily contain the same utterances from both speakers. The proposed algorithm is based on speaker adaptation techniques, adapting the conversion parameters derived for a particular pair of speakers to a different pair, for which only a nonparallel corpus is available. We show that adaptation reduces the error obtained when simply applying the conversion parameters of one pair of speakers to another by a factor that can reach 30%. A speaker identification measure is also employed that more insightfully portrays the importance of adaptation, while listening tests confirm the success of our method. Both the objective and subjective tests employed, demonstrate that the proposed algorithm achieves comparable results with the ideal case when a parallel corpus is available. Athanasios Mouchtaris, Jan Van der Spiegel, Paul Mueller |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | A spectral conversion approach to feature denoising and speech enhancementabstractIn this paper we demonstrate that spectral conversion can be successfully applied to the speech enhancement problem as a feature denoising method. The enhanced spectral features can be used in the context of the Kalman filter for estimating the clean speech signal. In essence, instead of estimating the clean speech features and the clean speech signal using the iterative Kalman filter, we show that is more efficient to initially estimate the clean speech features from the noisy speech features using spectral conversion (using a training speech corpus) and then apply the standard Kalman filter. Our results show an average improvement compared to the iterative Kalman filter that can reach 6 dB in the average segmental output Signal-to-Noise Ratio (SNR), in low input SNR's. Athanasios Mouchtaris, Jan Van der Spiegel, Paul Mueller, Panagiotis Tsakalides |
INTERSPEECH | 1 |
| 2005 | Multichannel audio synthesis by subband-based spectral conversion and parameter adaptationabstractMultichannel audio can immerse a group of listeners in a seamless aural environment. Previously, we proposed a system capable of synthesizing the multiple channels of a virtual multichannel recording from a smaller set of reference recordings. This problem was termed multichannel audio resynthesis and the application was to reduce the excessive transmission requirements of multichannel audio. In this paper, we address the more general problem of multichannel audio synthesis, i.e., how to completely synthesize a multichannel audio recording from a specific stereophonic or monophonic recording, which would significantly enhance the recording's acoustic impression. We approach this problem by extending the model employed for the resynthesis problem. This is accomplished by adapting the resynthesis conversion parameters to the statistical properties of the recording that we wish to enhance. This parameter adaptation is similar to the task adaptation employed in speech recognition, when a specific model is applied to a different environment (speaker, language or channel). One particular approach to this problem is shown here to be quite advantageous toward solving the multichannel audio synthesis problem as well. Athanasios Mouchtaris, Shri Narayanan, Chris Kyriakakis |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Non-parallel training for voice conversion by maximum likelihood constrained adaptationabstractThe objective of voice conversion methods is to modify the speech characteristics of a particular speaker in such manner, as to sound like speech by a different target speaker. Current voice conversion algorithms are based on deriving a conversion function by estimating its parameters through a corpus that contains the same utterances spoken by both speakers. Such a corpus, usually referred to as a parallel corpus, has the disadvantage that many times it is difficult or even impossible to collect. Here, we propose a voice conversion method that does not require a parallel corpus for training, i.e. the spoken utterances by the two speakers need not be the same, by employing speaker adaptation techniques to adapt to a particular pair of source and target speakers, the derived conversion parameters from a different pair of speakers. We show that adaptation reduces the error obtained when simply applying the conversion parameters of one pair of speakers to another by a factor that can reach 30% in many cases, and with performance comparable with the ideal case when a parallel corpus is available. Athanasios Mouchtaris, Jan Van der Spiegel, Paul Mueller |
ICASSP (1) | 1 |
| 2004 | A spectral conversion approach to the iterative Wiener filter for speech enhancementabstractThe iterative Wiener filter (IWF) for speech enhancement in additive noise is an effective and simple algorithm to implement. One of its main disadvantages is the lack of proper criteria for convergence, which has been shown to introduce severe degradation to the estimated clean signal. Here, an improvement of the IWF algorithm is proposed, when additional information is available for the signal to be enhanced. If a small amount of clean speech data is available, spectral conversion techniques can be applied for estimating the clean short-term spectral envelope of the speech signal from the noisy signal, with significant noise reduction. Our results show an average improvement compared to the original IWF that can reach 2 dB in the segmental output signal-to-noise ratio (SNR), in low input SNRs, which is perceptually significant. Athanasios Mouchtaris, Jan Van der Spiegel, Paul Mueller |
ICME | 1 |
| 2002 | Multiresolution spectral conversion for multichannel audio resynthesisabstractMultichannel audio is attracting rapidly increasing popularity in audio reproduction. In most cases, however, its transmission requirements are extremely demanding compared to the available bandwidth. One possible solution to this problem could be to transmit a reference channel and recreate the remaining channels at the receiving end. Such a method is proposed by taking advantage of spectral conversion techniques that have been successfully applied to speech processing. Applications of the proposed system include transmission of multichannel audio over the current Internet infrastructure and, as an extension of the methods proposed here, remastering of existing monophonic and stereophonic recordings for multichannel rendering. Athanasios Mouchtaris, Shri Narayanan, Chris Kyriakakis |
ICME (2) | 1 |
| 2000 | Inverse Filter Design for Immersive Audio Rendering Over LoudspeakersabstractImmersive audio systems can be used to render virtual sound sources in three-dimensional (3-D) space around a listener. This is achieved by simulating the head-related transfer function (HRTF) amplitude and phase characteristics using digital filters. In this paper, we examine certain key signal processing considerations in spatial sound rendering over headphones and loudspeakers. We address the problem of crosstalk inherent in loudspeaker rendering and examine two methods for implementing crosstalk cancellation and loudspeaker frequency response inversion in real time. We demonstrate that it is possible to achieve crosstalk cancellation of 30 dB using both methods, but one of the two (the Fast RLS Transversal Filter Method) offers a significant advantage in terms of computational efficiency. Our analysis is easily extendable to nonsymmetric listening positions and moving listeners. Athanasios Mouchtaris, Panagiotis Reveliotis, Chris Kyriakakis |
IEEE Trans. Multim. | 1 |
| 1999 | Non-minimum phase inverse filter methods for immersive audio renderingabstractImmersive audio systems are being envisioned for applications that include teleconferencing and telepresence; augmented and virtual reality for manufacturing and entertainment; air traffic control, pilot warning, and guidance systems; displays for the visually-impaired; distance learning; and professional sound and picture editing for television and film. The principal function of such systems is to synthesize, manipulate and render sound fields in real time. In this paper we examine several signal processing considerations in spatial sound rendering over loudspeakers. We propose two methods that can be used to implement the necessary filters for generating virtual sound sources based on synthetic head-related transfer functions with the same spectral characteristics as those of the real source. Athanasios Mouchtaris, Panagiotis Reveliotis, Chris Kyriakakis |
ICASSP | 1 |
| 1998 | Head-related transfer function synthesis for immersive audioabstractImmersive audio systems are being envisioned for applications that include teleconferencing and telepresence; augmented and virtual reality for manufacturing and entertainment; air traffic control, pilot warning, and guidance systems; displays for the visually- or aurally-impaired; home entertainment; distance learning; and professional sound and picture editing for television and film. The principal function of such systems is to synthesize, manipulate, and render sound fields in real time. In this paper we examine the limitations that are inherent in spatial sound delivery over loudspeakers and propose a method that generates virtual sound sources based on synthetic head-related transfer functions with the same spectral characteristics as those of the real source. Athanasios Mouchtaris, Jong-soong Lim, Tomlinson Holman, Chris Kyriakakis |
MMSP | 1 |