EDBT 2026 Demo / reviewers in the wild / expert
Aswin Shanmugam Subramanian
dblp:158/4181 · also S. Aswin Shanmugam
· DBLP profile ↗
22ranked-venue papers
7as first author
12since 2021 · last 2025
0000-0003-4446-001XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 14 · 5 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
Aswin Shanmugam Subramanian, Amit Das 0007, Naoyuki Kanda, Jinyu Li 0001, Xiaofei Wang 0007, Yifan Gong 0001 |
INTERSPEECH | 1 |
| 2025 | Length Aware Speech Translation for Video Dubbing
Aswin Shanmugam Subramanian, Harveen Singh Chadha, Vikas Joshi, Shubham Bansal, Rupeshkumar Mehta, Jinyu Li 0001 |
INTERSPEECH | 1 |
| 2024 | Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
Jinyu Li 0001, Jun-Kun Chen, Aswin Shanmugam Subramanian |
INTERSPEECH | 5 |
| 2024 | TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker EmbeddingsabstractSince diarization and source separation of meeting data are closely related tasks, we here propose an approach to perform the two objectives jointly. It builds upon the target-speaker voice activity detection (TS-VAD) diarization approach, which assumes that initial speaker embeddings are available. We replace the final combined speaker activity estimation network of TS-VAD with a network that produces speaker activity estimates at a time-frequency resolution. Those act as masks for source extraction, either via masking or via beamforming. The technique can be applied both for single-channel and multi-channel input and, in both cases, achieves a new state-of-the-art word error rate (WER) on the LibriCSS meeting data recognition task. We further compute speaker-aware and speaker-agnostic WERs to isolate the contribution of diarization errors to the overall WER performance. Christoph Böddeker, Aswin Shanmugam Subramanian, Gordon Wichern, Reinhold Häb-Umbach, Jonathan Le Roux |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Reverberation as Supervision For Speech SeparationabstractThis paper proposes reverberation as supervision (RAS), a novel unsupervised loss function for single-channel reverberant speech separation. Prior methods for unsupervised separation required the synthesis of mixtures of mixtures or assumed the existence of a teacher model, making them difficult to consider as potential methods explaining the emergence of separation abilities in an animal’s auditory system. We assume the availability of two-channel mixtures at training time, and train a neural network to separate the sources given one of the channels as input such that the other channel may be predicted from the separated sources. As the relationship between the room impulse responses (RIRs) of each channel depends on the locations of the sources, which are unknown to the network, the network cannot rely on learning that relationship. Instead, our proposed loss function fits each of the separated sources to the mixture in the target channel via Wiener filtering, and compares the resulting mixture to the ground-truth one. We show that minimizing the scale-invariant signal-to-distortion ratio (SI-SDR) of the predicted right-channel mixture with respect to the ground truth implicitly guides the network towards separating the left-channel sources. On a semi-supervised reverberant speech separation task based on the WHAMR! dataset, using training data where just 5% (resp., 10%) of the mixtures are labeled with associated isolated sources, we achieve 70% (resp., 78%) of the SI-SDR improvement obtained when training with supervision on the full training set, while a model trained only on the labeled data obtains 43% (resp., 45%). Rohith Aralikatti, Christoph Böddeker, Gordon Wichern, Aswin Shanmugam Subramanian, Jonathan Le Roux |
ICASSP | 4 |
| 2023 | Hyperbolic Audio Source SeparationabstractWe introduce a framework for audio source separation using embeddings on a hyperbolic manifold that compactly represent the hierarchical relationship between sound sources and time-frequency features. Inspired by recent successes modeling hierarchical relationships in text and images with hyperbolic embeddings, our algorithm obtains a hyperbolic embedding for each time-frequency bin of a mixture signal and estimates masks using hyperbolic softmax layers. On a synthetic dataset containing mixtures of multiple people talking and musical instruments playing, our hyperbolic model performed comparably to a Euclidean baseline in terms of source to distortion ratio, with stronger performance at low embedding dimensions. Furthermore, we find that time-frequency regions containing multiple overlapping sources are embedded towards the center (i.e., the most uncertain region) of the hyperbolic space, and we can use this certainty estimate to efficiently trade-off between artifact introduction and interference reduction when isolating individual sounds. Darius Petermann, Gordon Wichern, Aswin Shanmugam Subramanian, Jonathan Le Roux |
ICASSP | 3 |
| 2023 | Tackling the Cocktail Fork Problem for Separation and Transcription of Real-World SoundtracksabstractEmulating the human ability to solve the cocktail party problem, i.e., focus on a source of interest in a complex acoustic scene, is a long standing goal of audio source separation research. In this paper, we focus on the cocktail fork problem, which takes a three-pronged approach to source separation by separating an audio mixture such as a movie soundtrack or podcast into the three broad categories of speech, music, and sound effects (SFX - understood to include ambient noise and natural sound events). We evaluate several deep learning-based source separation models on this task using simple objective measures such as signal-to-distortion ratio (SDR) as well as objective metrics that better correlate with human perception. Furthermore, we thoroughly evaluate how source separation can influence the downstream transcription asks of speech recognition for speech and audio tagging for music and SFX. We also investigate the task of activity detection on the three sources as a way to further improve source separation and transcription. While we observe that source separation improves transcription performance in comparison to the original soundtrack, performance is still sub-optimal due to artifacts introduced by the separation process. Therefore, we thoroughly investigate how remixing of the three separated source stems at various relative levels can reduce artifacts and consequently improve transcription performance. We find that remixing music and SFX interferences at a target SNR of 17.5 dB reduces speech recognition word error rate, and similar impact from remixing is observed for tagging music and SFX content. Darius Petermann, Gordon Wichern, Aswin Shanmugam Subramanian, Zhongqiu Wang 0001, Jonathan Le Roux |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Heterogeneous Target Speech SeparationabstractWe introduce a new paradigm for single-channel target source separation where the sources of interest can be distinguished using non-mutually exclusive concepts (e.g., loudness, gender, language, spatial location, etc).Our proposed heterogeneous separation framework can seamlessly leverage datasets with large distribution shifts and learn cross-domain representations under a variety of concepts used as conditioning.Our experiments show that training separation models with heterogeneous conditions facilitates the generalization to new concepts with unseen out-of-domain data while also performing substantially higher than single-domain specialist models.Notably, such training leads to more robust learning of new harder source separation discriminative concepts and can yield improvements over permutation invariant training with oracle source selection.We analyze the intrinsic behavior of source separation training with heterogeneous metadata and propose ways to alleviate emerging problems with challenging separation conditions.We release the collection of preparation recipes for all datasets used to further promote research towards this challenging task. Efthymios Tzinis, Gordon Wichern, Aswin Shanmugam Subramanian, Paris Smaragdis, Jonathan Le Roux |
INTERSPEECH | 3 |
| 2022 | Deep learning based multi-source localization with source splitting and its effectiveness in multi-talker speech recognitionabstractMulti-source localization is an important and challenging technique for multi-talker conversation analysis. This paper proposes a novel supervised learning method using deep neural networks to estimate the direction of arrival (DOA) of all the speakers simultaneously from the audio mixture. At the heart of the proposal is a source splitting mechanism that creates source-specific intermediate representations inside the network. This allows our model to give source-specific posteriors as the output unlike the traditional multi-label classification approach . Existing deep learning methods perform a frame level prediction, whereas our approach performs an utterance level prediction by incorporating temporal selection and averaging inside the network to avoid post-processing. We also experiment with various loss functions and show that a variant of earth mover distance (EMD) is very effective in classifying DOA at a very high resolution by modeling inter-class relationships. In addition to using the prediction error as a metric for evaluating our localization model, we also establish its potency as a frontend with automatic speech recognition (ASR) as the downstream task. We convert the estimated DOAs into a feature suitable for ASR and pass it as an additional input feature to a strong multi-channel and multi-talker speech recognition baseline. This added input feature drastically improves the ASR performance and gives a word error rate (WER) of 6.3% on the evaluation data of our simulated noisy two speaker mixtures, while the baseline which does not use explicit localization input has a WER of 11.5%. We also perform ASR evaluation on real recordings with the overlapped set of the MC-WSJ-AV corpus in addition to simulated mixtures. Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe 0001, Meng Yu 0003, Dong Yu 0001 |
Comput. Speech Lang. | 1 |
| 2021 | An Exploration of Self-Supervised Pretrained Representations for End-to-End Speech RecognitionabstractSelf-supervised pretraining on speech data has achieved a lot of progress. High-fidelity representation of the speech signal is learned from a lot of untranscribed data and shows promising performance. Recently, there are several works focusing on evaluating the quality of self-supervised pretrained representations on various tasks with-out domain restriction, e.g. SUPERB. However, such evaluations do not provide a comprehensive comparison among many ASR benchmark corpora. In this paper, we focus on the general applications of pretrained speech representations, on advanced end-to-end automatic speech recognition (E2E-ASR) models. We select sev-eral pretrained speech representations and present the experimental results on various open-source and publicly available corpora for E2E-ASR. Without any modification of the back-end model archi-tectures or training strategy, some of the experiments with pretrained representations, e.g., WSJ, WSJ0-2mix with HuBERT, reach or out-perform current state-of-the-art (SOTA) recognition performance. Moreover, we further explore more scenarios for whether the pre-training representations are effective, such as the cross-language or overlapped speech. The scripts, configuratons and the trained mod-els have been released in ESPnet to let the community reproduce our experiments and improve them. Xuankai Chang, Takashi Maekaku, Jing Shi 0003, Yen-Ju Lu, Aswin Shanmugam Subramanian, Tianzi Wang, Shu-Wen Yang, Yu Tsao 0001, Hung-yi Lee, Shinji Watanabe 0001 |
ASRU | 6 |
| 2021 | Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source LocalizationabstractThis paper proposes a new paradigm for handling far-field multi-speaker data in an end-to-end (E2E) neural network manner, called directional automatic speech recognition (D-ASR), which explicitly models source speaker locations. In D-ASR, the azimuth angle of the sources with respect to the microphone array is defined as a latent variable. This angle controls the quality of separation, which in turn determines the ASR performance. All three functionalities of D-ASR: localization, separation, and recognition are connected as a single differentiable neural network and trained solely based on ASR error minimization objectives. The advantages of D-ASR over existing methods are threefold: (1) it provides explicit speaker locations, (2) it improves the explainability factor, and (3) it achieves better ASR performance as the process is more streamlined. In addition, D-ASR does not require explicit direction of arrival (DOA) supervision like existing data-driven localization models, which makes it more appropriate for realistic data. For the case of two source mixtures, D-ASR achieves an average DOA prediction error of less than three degrees. It also outperforms a strong far-field multi-speaker end-to-end system in both separation quality and ASR performance. Aswin Shanmugam Subramanian, Chao Weng, Shinji Watanabe 0001, Meng Yu 0003, Yong Xu 0004, Shixiong Zhang 0001, Dong Yu 0001 |
ICASSP | 1 |
| 2021 | ESPnet-SE: End-To-End Speech Enhancement and Separation Toolkit Designed for ASR IntegrationabstractWe present ESPnet-SE, which is designed for the quick development of speech enhancement and speech separation systems in a single framework, along with the optional downstream speech recognition module. ESPnet-SE is a new project which integrates rich automatic speech recognition related models, resources and systems to support and validate the proposed front-end implementation (i.e. speech enhancement and separation).It is capable of processing both single-channel and multi-channel data, with various functionalities including dereverberation, denoising and source separation. We provide all-in-one recipes including data pre-processing, feature extraction, training and evaluation pipelines for a wide range of benchmark datasets. This paper describes the design of the toolkit, several important functionalities, especially the speech recognition integration, which differentiates ESPnet-SE from other open source toolkits, and experimental results with major benchmark datasets. Chenda Li, Jing Shi 0003, Wangyou Zhang, Aswin Shanmugam Subramanian, Xuankai Chang, Naoyuki Kamo, Moto Hira, Tomoki Hayashi, Christoph Böddeker, Zhuo Chen 0006, Shinji Watanabe 0001 |
SLT | 4 |
| 2020 | Attention-Based ASR with Lightweight and Dynamic ConvolutionsabstractEnd-to-end (E2E) automatic speech recognition (ASR) with sequence-to-sequence models has gained attention because of its simple model training compared with conventional hidden Markov model based ASR. Recently, several studies report the state-of-the-art E2E ASR results obtained by Transformer. Compared to recurrent neural network (RNN) based E2E models, training of Transformer is more efficient and also achieves better performance on various tasks. However, self-attention used in Transformer requires computation quadratic in its input length. In this paper, we propose to apply lightweight and dynamic convolution to E2E ASR as an alternative architecture to the self-attention to make the computational order linear. We also propose joint training with connectionist temporal classification, convolution on the frequency axis, and combination with self-attention. With these techniques, the proposed architectures achieve better performance than RNN-based E2E model and performance competitive to state-of-the-art Transformer on various ASR benchmarks including noisy/reverberant tasks. Yuya Fujita, Aswin Shanmugam Subramanian, Motoi Omachi, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2020 | Far-Field Location Guided Target Speech Extraction Using End-to-End Speech Recognition ObjectivesabstractTarget speech extraction is a specific case of source separation where an auxiliary information like the location or some pre-saved anchor speech examples of the target speaker is used to resolve the permutation ambiguity. Traditionally such systems are optimized based on signal reconstruction objectives. Recently end-to-end automatic speech recognition (ASR) methods have enabled to optimize source separation systems with only the transcription based objective. This paper proposes a method to jointly optimize a location guided target speech extraction module along with a speech recognition module only with ASR error minimization criteria. Experimental comparisons with corresponding conventional pipeline systems verify that this task can be realized by end-to-end ASR training objectives without using parallel clean data. We show promising target speech recognition results in mixtures of two speakers and noise, and discuss interesting properties of the proposed system in terms of speech enhancement/separation objectives and word error rates. Finally, we design a system that can take both location and anchor speech as input at the same time and show that the performance can be further improved. Aswin Shanmugam Subramanian, Chao Weng, Meng Yu 0003, Shixiong Zhang 0001, Yong Xu 0004, Shinji Watanabe 0001, Dong Yu 0001 |
ICASSP | 1 |
| 2020 | End-to-End ASR with Adaptive Span Self-Attention
Xuankai Chang, Aswin Shanmugam Subramanian, Shinji Watanabe 0001, Yuya Fujita, Motoi Omachi |
INTERSPEECH | 2 |
| 2020 | End-to-End Far-Field Speech Recognition with Unified Dereverberation and BeamformingabstractDespite successful applications of end-to-end approaches in multi-channel speech recognition, the performance still degrades severely when the speech is corrupted by reverberation. In this paper, we integrate the dereverberation module into the end-to-end multi-channel speech recognition system and explore two different frontend architectures. First, a multi-source mask-based weighted prediction error (WPE) module is incorporated in the frontend for dereverberation. Second, another novel frontend architecture is proposed, which extends the weighted power minimization distortionless response (WPD) convolutional beamformer to perform simultaneous separation and dereverberation. We derive a new formulation from the original WPD, which can handle multi-source input, and replace eigenvalue decomposition with the matrix inverse operation to make the back-propagation algorithm more stable. The above two architectures are optimized in a fully end-to-end manner, only using the speech recognition criterion. Experiments on both spatialized wsj1-2mix corpus and REVERB show that our proposed model outperformed the conventional methods in reverberant scenarios. Wangyou Zhang, Aswin Shanmugam Subramanian, Xuankai Chang, Shinji Watanabe 0001, Yanmin Qian |
INTERSPEECH | 2 |
| 2020 | Significance of spectral cues in automatic speech segmentation for Indian language speech synthesizers
Arun Baby, Jeena J. Prakash, Aswin Shanmugam Subramanian, Hema A. Murthy |
Speech Commun. | 3 |
| 2018 | Building State-of-the-art Distant Speech Recognition Using the CHiME-4 Challenge with a Setup of Speech Enhancement BaselineabstractThis paper describes a new baseline system for automatic speech recognition (ASR) in the CHiME-4 challenge to promote the development of noisy ASR in speech processing communities by providing 1) state-of-the-art system with a simplified single system comparable to the complicated top systems in the challenge, 2) publicly available and reproducible recipe through the main repository in the Kaldi speech recognition toolkit.The proposed system adopts generalized eigenvalue beamforming with bidirectional long short-term memory (LSTM) mask estimation.We also propose to use a time delay neural network (TDNN) based on the lattice-free version of the maximum mutual information (LF-MMI) trained with augmented all six microphones plus the enhanced data after beamforming.Finally, we use a LSTM language model for lattice and n-best re-scoring.The final system achieved 2.74% WER for the real test set in the 6-channel track, which corresponds to the 2nd place in the challenge.In addition, the proposed baseline recipe includes four different speech enhancement measures, short-time objective intelligibility measure (STOI), extended STOI (eSTOI), perceptual evaluation of speech quality (PESQ) and speech distortion ratio (SDR) for the simulation test set.Thus, the recipe also provides an experimental platform for speech enhancement studies with these performance measures. Szu-Jui Chen, Aswin Shanmugam Subramanian, Hainan Xu, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2018 | Student-Teacher Learning for BLSTM Mask-based Speech EnhancementabstractSpectral mask estimation using bidirectional long short-term memory (BLSTM) neural networks has been widely used in various speech enhancement applications, and it has achieved great success when it is applied to multichannel enhancement techniques with a mask-based beamformer.However, when these masks are used for single channel speech enhancement they severely distort the speech signal and make them unsuitable for speech recognition.This paper proposes a studentteacher learning paradigm for single channel speech enhancement.The beamformed signal from multichannel enhancement is given as input to the teacher network to obtain soft masks.An additional cross-entropy loss term with the soft mask target is combined with the original loss, so that the student network with single-channel input is trained to mimic the soft mask obtained with multichannel input through beamforming.Experiments with the CHiME-4 challenge single channel track data shows improvement in ASR performance. Aswin Shanmugam Subramanian, Szu-Jui Chen, Shinji Watanabe 0001 |
INTERSPEECH | 1 |
| 2017 | TBT (Toolkit to Build TTS): A High Performance Framework to Build Multiple Language HTS Voice
Atish Shankar Ghone, Rachana Nerpagar, Pranaw Kumar, Arun Baby, Aswin Shanmugam Subramanian, M. Sasikumar 0001, Hema A. Murthy |
INTERSPEECH | 5 |
| 2016 | Significance of Pseudo-syllables in building better acoustic models for Indian English TTSabstractSignal processing based landmark detection is precise compared to HMM based alignment, primarily because the location of the landmark is not factored in the estimation of parameters. Acoustic cues for syllable boundaries are usually obtained by exploiting the inherent sonority characteristics of a syllable. As syllabification of the text is based on generalized rules or lexicon definitions, there is a mismatch between the acoustical and the lexical segments for non-native syllabification. In this paper, an attempt is made to modify the syllabification rules for Indian English using acoustic cues obtained from syllable boundaries. The modified syllabifier is used to syllabify the text. Embedded re-estimation is performed using forced alignment at the modified syllable level to obtain refined phoneme boundaries. Indian English Text-to-Speech (TTS) systems are built using labels obtained after (i) embedded re-estimation at the sentence level and (ii) the aforementioned procedure. Reduction in the word error rates for both native Aryan and Dravidian speakers (relatively by 54.1% and 52.4% respectively), suggests that there is a significant synthesis quality improvement in the proposed system. S. Rupak Vignesh, Aswin Shanmugam Subramanian, Hema A. Murthy |
ICASSP | 2 |
| 2014 | A hybrid approach to segmentation of speech using group delay processing and HMM based embedded reestimation
Aswin Shanmugam Subramanian, Hema A. Murthy |
INTERSPEECH | 1 |