Lahiru Samarakoon

dblp:175/8860 · DBLP profile ↗
← Back
23ranked-venue papers
15as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 14 first-author · 11 since 2021Artificial intelligence and machine learning · 14 · 7 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Non-Autoregressive Multi-Speaker ASR with Decoupled Speaker Change Detection
abstract
Existing multi-speaker automatic speech recognition (ASR) approaches predominantly rely on autoregressive (AR) decoding. The AR decoding complicates the parallel computation. In this work we explore non-autoregressive (NAR) models for multi-speaker ASR to achieve faster inference. Multi-speaker ASR comprises two primary tasks: speech recognition and speaker change detection (SCD). We introduce a framework that decouples these tasks into separate branches. Inference is carried out in two stages: ASR decoding is performed first, then SCD decoding leverages the ASR output and predicts all the speaker changes in one pass. Experiments show that with NAR decoding, the proposed model outperforms the baseline by a relative cpWER reduction of 30% on LibriSpeechMix and 8% on AMI. Additionally, in scenarios with a 10% overlap ratio, our NAR model achieves a threefold increase in decoding speed with only a 1.2 % reduction in accuracy compared to a strong AR baseline.
Yingke Zhu, Lahiru Samarakoon
ASRU2
2025 Variance-Covariance Regularization for Improved End-to-End Diarization
abstract
End-to-end neural diarization (EEND) methods use just a single neural network. EEND-TA, a recently proposed EEND technique, performs diarization for a flexible number of speakers in a non-autoregressive manner. In this paper, we explore combining EEND-TA with Variance-Covariance Regularization (VCReg) to enhance representation learning. VCReg is designed to promote features with high variance and low covariance. We test several representations from EEND-TA for calculating VCReg loss, including Conversational Summary Vectors (CSVs), outputs from the last Conformer layer, attractor representations, and other intermediate outputs from the Conformer encoder. Our experiments on public datasets show notable improvements over the baseline. Additionally, the proposed method improves the performance over all datasets in our setup, showcasing the generalizability of this approach.
Lahiru Samarakoon, Samuel J. Broughton, Ivan Fung
ICASSP1
2025 Pushing the Limits of End-to-End Diarization
Samuel J. Broughton, Lahiru Samarakoon
INTERSPEECH2
2024 EEND-M2F: Masked-attention mask transformers for speaker diarization
abstract
In this paper, we make the explicit connection between image segmentation methods and end-to-end diarization methods.From these insights, we propose a novel, fully end-to-end diarization model, EEND-M2F, based on the Mask2Former architecture.Speaker representations are computed in parallel using a stack of transformer decoders, in which irrelevant frames are explicitly masked from the cross attention using predictions from previous layers.EEND-M2F is efficient, and truly end-to-end, eliminating the need for additional segmentation models or clustering algorithms.Our model achieves state-of-the-art performance on several public datasets, such as AMI, AliMeeting and RAMC.Most notably our DER of 16.07% on DIHARD-III is the first major improvement upon the challenge winning system.
Marc Härkönen, Samuel J. Broughton, Lahiru Samarakoon
INTERSPEECH3
2023 Robust End-to-End Diarization with Domain Adaptive Training and Multi-Task Learning
abstract
Due to the scarcity of publicly available diarization data, the model performance can be improved by training a single model with data from different domains. In this work, we propose to incorporate domain information to train a single end-to-end diarization model for multiple domains. First, we employ domain adaptive training with parameter-efficient adapters for on-the-fly model reconfiguration. Second, we introduce an auxiliary domain classification task to make the diarization model more domain-aware. For seen domains, the combination of our proposed methods reduces the absolute DER from 17.66% to 16.59% when compared with the baseline. During inference, adapters from ground-truth domains are not available for unseen domains. We demonstrate our model exhibits a stronger generalizability to unseen domains when adapters are removed. For two unseen domains, this improves the DER performance from 39.91% to 23.09% and 25.32% to 18.76% over the baseline, respectively.
Ivan Fung, Lahiru Samarakoon, Samuel J. Broughton
ASRU2
2023 Transformer Attractors for Robust and Efficient End-To-End Neural Diarization
abstract
End-to-end neural diarization with encoder-decoder based attractors (EEND-EDA) is a method to perform diarization in a single neural network. EDA handles the diarization of a flexible number of speakers by using an LSTM-based encoder-decoder that generates a set of speaker-wise attractors in an autoregressive manner. In this paper, we propose to replace EDA with a transformer-based attractor calculation (TA) module. TA is composed of a Combiner block and a Transformer decoder. The main function of the combiner block is to generate conversational dependent (CD) embeddings by incorporating learned conversational information into a global set of embeddings. These CD embeddings will then serve as the input for the transformer decoder. Results on public datasets show that EEND-TA achieves 2.68% absolute DER improvement over EEND-EDA. EEND-TA inference is 1.28 times faster than that of EEND-EDA.
Lahiru Samarakoon, Samuel J. Broughton, Marc Härkönen, Ivan Fung
ASRU1
2023 Improving Non-Autoregressive Speech Recognition with Autoregressive Pretraining
abstract
Autoregressive (AR) automatic speech recognition (ASR) models predict each output token conditioning on the previous ones, which slows down their inference speed. On the other hand, non-autoregressive (NAR) models predict tokens independently and simultaneously within a constant number of decoding iterations, which brings high inference speed. However, NAR models generally have lower accuracy than AR models. In this work, we propose AR pretraining to the NAR encoder to reduce the accuracy gap between AR and NAR models. The experiment results show that our AR-pretrained MaskCTC reaches the same accuracy as AR Conformer on Aishell-1 (both 4.9% CER) and reduce the performance gap with AR Conformer on LibriSpeech by relatively 50%. Moreover, our AR-pretrained MaskCTC only needs single decoding iteration, which reduces inference time by 50%. We also investigate multiple masking strategies in training the masked language model of MaskCTC.
Yanjia Li, Lahiru Samarakoon, Ivan Fung
ICASSP2
2023 Improving End-to-End Neural Diarization Using Conversational Summary Representations
Samuel J. Broughton, Lahiru Samarakoon
INTERSPEECH2
2022 Conformer-Based Speech Recognition with Linear Nyström Attention and Rotary Position Embedding
abstract
Self-attention has become an important component for end-to-end (E2E) automatic speech recognition (ASR). Recently, Convolution-augmented Transformer (Conformer) with relative positional encoding (RPE) achieved state-of-the-art performance. However, the computational and memory complexity of self-attention grows quadratically with the input sequence length. Effect of this can be significant for the Conformer encoder when processing longer sequences. In this work, we propose to replace self-attention with a linear complexity Nyström attention which is a low-rank approximation of the attention scores based on the Nyström method. In addition, we propose to use Rotary Position Embedding (RoPE) with Nyström attention since RPE is of quadratic complexity. Moreover, we show that models can be made even lighter by removing self-attention sub-layers from top encoder layers without any drop in the performance. Furthermore, we demonstrate that Convolutional sub-layers in Conformer can effectively recover the information lost due to the Nyström approximation.
Lahiru Samarakoon, Tsun-Yat Leung
ICASSP1
2022 Untied Positional Encodings for Efficient Transformer-Based Speech Recognition
abstract
Self-attention has become a vital component for end-to-end (E2E) automatic speech recognition (ASR). Convolution-augmented Transformer (Conformer) with relative positional encoding (RPE) achieved state-of-the-art performance. This paper proposes a positional encoding (PE) mechanism called Scaled Untied RPE that unties the feature-position correlations in the self-attention computation, and computes feature correlations and positional correlations separately using different projection matrices. In addition, we propose to scale feature correlations with the positional correlations and the aggressiveness of this multiplicative interaction can be configured using a parameter called amplitude. Moreover, we show that the PE matrix can be sliced to reduce model parameters. Our results on National Speech Corpus (NSC) show that Transformer encoders with Scaled Untied RPE achieves relative improvements of 1.9% in accuracy and up to 50.9% in latency over a Conformer baseline respectively.
Lahiru Samarakoon, Ivan Fung
SLT1
2021 Robust End-to-End Speaker Diarization with Conformer and Additive Margin Penalty
Tsun-Yat Leung, Lahiru Samarakoon
Interspeech2
2019 Incorporating Prior Knowledge into Speaker Diarization and Linking for Identifying Common Speaker
abstract
Speaker Diarization and Linking discovers “who spoke when” across recordings without any speaker enrollment. Diarization is performed on each recording separately, and the linking combines clusters of the same speaker across recordings. It is a two-step approach, however it suffers from propagating the error from diarization step to the linking step. In a situation where a unique speaker appears in a given set of recordings, this paper aims at locating the common speaker using the prior knowledge of his or her existence. That means there is no enrollment data for this common speaker. We propose Pairwise Common Speaker Identification (PCSI) method that takes the existence of a common speaker into account in contrast to the two-step approach. We further show that PCSI can be used to reduce the errors that are introduced in the diarization step of the two-step approach. Our experiments are performed on a corpus synthesised from the AMI corpus and also on a in-house conversational telephony Sichuanese corpus that is mixed with Mandarin. We show up to 7.68% relative improvements of time-weighted equal error rate over a state-of-art x-vector diarization and linking system.
Tsun-Yat Leung, Lahiru Samarakoon, Albert Y. S. Lam
ASRU2
2018 learning Effective Factorized Hidden Layer Bases Using Student-Teacher Training for LSTM Acoustic Model Adaptation
abstract
Factorized Hidden Layer (FHL) has been proposed for the adaptation of deep neural network (DNN) and Long Short-Term Memory (LSTM) based acoustic models (AMs). In FHL, a speaker-dependent (SD) transformation matrix and an SD bias are included in addition to the standard affine transformation. The SD transformation is a linear combination of rank- l matrices whereas the SD bias is a linear combination of vectors. However, the adaptation of LSTMs is challenging and often reports modest gains. In this paper, we propose to use student-teacher training to estimate more efficient FHL bases for LSTM AMs using an FHL adapted DNN as the teacher model. For both AMI IHM and AMI SDM tasks, FHL achieves 3.2% absolute improvement over the frame-level cross entropy trained LSTM baselines. Moreover, FHL results 3.0% and 3.8% absolute improvements over sequentially trained LSTM baselines for the AMI IHM and AMI SDM tasks respectively.
Lahiru Samarakoon, Brian Kan-Wing Mak, Khe Chai Sim
ICASSP1
2018 Domain Adaptation of End-to-end Speech Recognition in Low-Resource Settings
abstract
End-to-end automatic speech recognition (ASR) has simplified the traditional ASR system building pipeline by eliminating the need to have multiple components and also the requirement for expert linguistic knowledge for creating pronunciation dictionaries. Therefore, end-to-end ASR fits well when building systems for new domains. However, one major drawback of end-to-end ASR is that, it is necessary to have a larger amount of labeled speech in comparison to traditional methods. Therefore, in this paper, we explore domain adaptation approaches for end-to-end ASR in low-resource settings. We show that joint domain identification and speech recognition by inserting a symbol for domain at the beginning of the label sequence, factorized hidden layer adaptation and a domain-specific gating mechanism improve the performance for a low-resource target domain. Furthermore, we also show the robustness of proposed adaptation methods to an unseen domain, when only 3 hours of untranscribed data is available with improvements reporting upto 8.7% relative.
Lahiru Samarakoon, Brian Kan-Wing Mak, Albert Y. S. Lam
SLT1
2017 Unsupervised adaptation of student DNNS learned from teacher RNNS for improved ASR performance
abstract
In automatic speech recognition (ASR), adaptation techniques are used to minimize the mismatch between training and testing conditions. Many successful techniques have been proposed for deep neural network (DNN) acoustic model (AM) adaptation. Recently, recurrent neural networks (RNNs) have outperformed DNNs in ASR tasks. However, the adaptation of RNN AMs is challenging and in some cases when combined with adaptation, DNN AMs outperform adapted RNN AMs. In this paper, we combine student-teacher training and unsupervised adaptation to improve ASR performance. First, RNNs are used as teachers to train student DNNs. Then, these student DNNs are adapted in an unsupervised fashion. Experimental results on the AMI IHM and AMI SDM tasks show that student DNNs are adaptable with significant performance improvements for both frame-wise and sequentially trained systems. We also show that the combination of adapted DNNs with teacher RNNs can further improve the performance.
Lahiru Samarakoon, Brian Kan-Wing Mak
ASRU1
2017 An investigation into learning effective speaker subspaces for robust unsupervised DNN adaptation
abstract
Subspace methods are used for deep neural network (DNN)-based acoustic model adaptation. These methods first construct a subspace and then perform the speaker adaptation as a point in the subspace. This paper aims to investigate the effectiveness of subspace methods for robust unsupervised adaptation. For the analysis, we compare two state-of-the-art subspace methods, namely, the singular value decomposition (SVD)-based bottleneck adaptation and the factorized hidden layer (FHL) adaptation. Both of these methods perform speaker adaptation as a linear combination of rank-1 bases. The main difference between the subspace construction is that FHL adaptation constructs a speaker subspace separate from the phoneme classification space while SVD-based bottleneck adaptation shares the same subspace for both the phoneme classification and the speaker adaptation. So far, no direct comparisons between these two methods are reported. In this work, we compare these two methods for their robustness to unsupervised adaptation on Aurora 4, AMI IHM and AMI SDM tasks. Our findings show that the FHL adaptation outperforms the SVD-based bottleneck adaptation especially in challenging conditions where the adaptation data is limited, or the quality of the adaptation alignments are low.
Lahiru Samarakoon, Khe Chai Sim, Brian Kan-Wing Mak
ICASSP1
2017 Learning Factorized Transforms for Unsupervised Adaptation of LSTM-RNN Acoustic Models
abstract
Factorized Hidden Layer (FHL) adaptation has been proposed for speaker adaptation of deep neural network (DNN) based acoustic models. In FHL adaptation, a speaker-dependent (SD) transformation matrix and an SD bias are included in addition to the standard affine transformation. The SD transformation is a linear combination of rank-1 matrices whereas the SD bias is a linear combination of vectors. Recently, the Long Short- Term Memory (LSTM) Recurrent Neural Networks (RNNs) have shown to outperform DNN acoustic models in many Automatic Speech Recognition (ASR) tasks. In this work, we investigate the effectiveness of SD transformations for LSTM-RNN acoustic models. Experimental results show that when combined with scaling of LSTM cell states' outputs, SD transformations achieve 2.3% and 2.1% absolute improvements over the baseline LSTM systems for the AMI IHM and AMI SDM tasks respectively.
Lahiru Samarakoon, Brian Kan-Wing Mak, Khe Chai Sim
INTERSPEECH1
2016 On combining i-vectors and discriminative adaptation methods for unsupervised speaker normalization in DNN acoustic models
abstract
In automatic speech recognition (ASR), adaptation and adaptive training techniques are used to perform speaker normalization. Previous methods mainly focus on using these techniques in isolation. In contrast, this paper investigates two approaches to improve the ASR performance by combining i-vector based speaker adaptive training in deep neural network (DNN) acoustic models with discriminative adaptation techniques. First, we combine these techniques by interpolating the decoding lattices of i-vector based systems with the decoding lattices of a discriminatively adapted model. Then, we combine these methods by discriminatively adapting the i-vector based system in unsupervised fashion. Our experiments on TED-LIUM dataset show that compared with a strong speaker independent baseline, lattice interpolation and adaptation of the i-vector systems achieve 12.0% and 15.6% relative improvements, respectively. Moreover, in comparison to the i-vector based systems, lattice interpolation reported a 4.5% relative improvement while discriminatively adapting the i-vector system reported a 8.3% relative improvement.
Lahiru Samarakoon, Khe Chai Sim
ICASSP1
2016 Subspace LHUC for Fast Adaptation of Deep Neural Network Acoustic Models
Lahiru Samarakoon, Khe Chai Sim
INTERSPEECH1
2016 Multi-Attribute Factorized Hidden Layer Adaptation for DNN Acoustic Models
Lahiru Samarakoon, Khe Chai Sim
INTERSPEECH1
2016 Low-rank bases for factorized hidden layer adaptation of DNN acoustic models
abstract
Recently, the factorized hidden layer (FHL) adaptation method is proposed for speaker adaptation of deep neural network (DNN) acoustic models. An FHL contains a speaker-dependent (SD) transformation matrix using a linear combination of rank-1 matrices and an SD bias using a linear combination of vectors, in addition to the standard affine transformation. On the other hand, full-rank bases are used with a similar DNN adaptation method which is based on cluster adaptive training (CAT). Therefore, it is interesting to investigate the effect of the rank of the bases used for adaptation. The increase of the rank of the bases improves the speaker subspace representation, without increasing the number of learnable speaker parameters. In this work, we investigate the effect of using various ranks for the bases of the SD transformation of FHLs on Aurora 4, AMI IHM and AMI SDM tasks. Experimental results have shown that when one FHL layer is used, it is optimal to use low-ranked bases of rank-50, instead of full-rank bases. Furthermore, when multiple FHLs are used, rank-1 bases are sufficient.
Lahiru Samarakoon, Khe Chai Sim
SLT1
2016 Factorized Hidden Layer Adaptation for Deep Neural Network Based Acoustic Modeling
abstract
In this paper, we propose the factorized hidden layer (FHL) approach to adapt the deep neural network (DNN) acoustic models for automatic speech recognition (ASR). FHL aims at modeling speaker dependent (SD) hidden layers by representing an SD affine transformation as a linear combination of bases. The combination weights are low-dimensional speaker parameters that can be initialized using speaker representations like i-vectors and then reliably refined in an unsupervised adaptation fashion. Therefore, our method provides an efficient way to perform both adaptive training and (test-time) adaptation. Experimental results have shown that the FHL adaptation improves the ASR performance significantly, compared to the standard DNN models, as well as other state-of-the-art DNN adaptation approaches, such as training with the speaker-normalized CMLLR features, speaker-aware training using i-vector and learning hidden unit contributions (LHUC). For Aurora 4, FHL achieves 3.8% and 2.3% absolute improvements over the standard DNNs trained on the LDA + STC and CMLLR features, respectively. It also achieves 1.7% absolute performance improvement over a system that combines the i-vector adaptive training with LHUC adaptation. For the AMI dataset, FHL achieved 1.4% and 1.9% absolute improvements over the sequence-trained CMLLR baseline systems, for the IHM and SDM tasks, respectively.
Lahiru Samarakoon, Khe Chai Sim
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Learning factorized feature transforms for speaker normalization
abstract
This paper proposes an approach to improve automatic speech recognition (ASR) by normalizing the speaker variability of a well trained Deep Neural Network (DNN) acoustic model using i-vectors. Our approach learns a speaker dependent transformation of the acoustic features combined with the standard speaker dependent bias, to minimize the mismatch due to the inter-speaker variability. Speaker normalization experiments on the Aurora 4 task show 10.9% relative improvement over the baseline. Moreover, the proposed approach reported 4.5% relative improvement over the standard i-vector based method where only a speaker dependent bias is used. Furthermore, we report an analysis to compare our approach with the Constrained Maximum Likelihood Linear Regression (CMLLR) method.
Lahiru Samarakoon, Khe Chai Sim
ASRU1