Ivan Fung

dblp:226/5646 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Variance-Covariance Regularization for Improved End-to-End Diarization
abstract
End-to-end neural diarization (EEND) methods use just a single neural network. EEND-TA, a recently proposed EEND technique, performs diarization for a flexible number of speakers in a non-autoregressive manner. In this paper, we explore combining EEND-TA with Variance-Covariance Regularization (VCReg) to enhance representation learning. VCReg is designed to promote features with high variance and low covariance. We test several representations from EEND-TA for calculating VCReg loss, including Conversational Summary Vectors (CSVs), outputs from the last Conformer layer, attractor representations, and other intermediate outputs from the Conformer encoder. Our experiments on public datasets show notable improvements over the baseline. Additionally, the proposed method improves the performance over all datasets in our setup, showcasing the generalizability of this approach.
Lahiru Samarakoon, Samuel J. Broughton, Ivan Fung
ICASSP3
2023 Robust End-to-End Diarization with Domain Adaptive Training and Multi-Task Learning
abstract
Due to the scarcity of publicly available diarization data, the model performance can be improved by training a single model with data from different domains. In this work, we propose to incorporate domain information to train a single end-to-end diarization model for multiple domains. First, we employ domain adaptive training with parameter-efficient adapters for on-the-fly model reconfiguration. Second, we introduce an auxiliary domain classification task to make the diarization model more domain-aware. For seen domains, the combination of our proposed methods reduces the absolute DER from 17.66% to 16.59% when compared with the baseline. During inference, adapters from ground-truth domains are not available for unseen domains. We demonstrate our model exhibits a stronger generalizability to unseen domains when adapters are removed. For two unseen domains, this improves the DER performance from 39.91% to 23.09% and 25.32% to 18.76% over the baseline, respectively.
Ivan Fung, Lahiru Samarakoon, Samuel J. Broughton
ASRU1
2023 Transformer Attractors for Robust and Efficient End-To-End Neural Diarization
abstract
End-to-end neural diarization with encoder-decoder based attractors (EEND-EDA) is a method to perform diarization in a single neural network. EDA handles the diarization of a flexible number of speakers by using an LSTM-based encoder-decoder that generates a set of speaker-wise attractors in an autoregressive manner. In this paper, we propose to replace EDA with a transformer-based attractor calculation (TA) module. TA is composed of a Combiner block and a Transformer decoder. The main function of the combiner block is to generate conversational dependent (CD) embeddings by incorporating learned conversational information into a global set of embeddings. These CD embeddings will then serve as the input for the transformer decoder. Results on public datasets show that EEND-TA achieves 2.68% absolute DER improvement over EEND-EDA. EEND-TA inference is 1.28 times faster than that of EEND-EDA.
Lahiru Samarakoon, Samuel J. Broughton, Marc Härkönen, Ivan Fung
ASRU4
2023 Improving Non-Autoregressive Speech Recognition with Autoregressive Pretraining
abstract
Autoregressive (AR) automatic speech recognition (ASR) models predict each output token conditioning on the previous ones, which slows down their inference speed. On the other hand, non-autoregressive (NAR) models predict tokens independently and simultaneously within a constant number of decoding iterations, which brings high inference speed. However, NAR models generally have lower accuracy than AR models. In this work, we propose AR pretraining to the NAR encoder to reduce the accuracy gap between AR and NAR models. The experiment results show that our AR-pretrained MaskCTC reaches the same accuracy as AR Conformer on Aishell-1 (both 4.9% CER) and reduce the performance gap with AR Conformer on LibriSpeech by relatively 50%. Moreover, our AR-pretrained MaskCTC only needs single decoding iteration, which reduces inference time by 50%. We also investigate multiple masking strategies in training the masked language model of MaskCTC.
Yanjia Li, Lahiru Samarakoon, Ivan Fung
ICASSP3
2022 Untied Positional Encodings for Efficient Transformer-Based Speech Recognition
abstract
Self-attention has become a vital component for end-to-end (E2E) automatic speech recognition (ASR). Convolution-augmented Transformer (Conformer) with relative positional encoding (RPE) achieved state-of-the-art performance. This paper proposes a positional encoding (PE) mechanism called Scaled Untied RPE that unties the feature-position correlations in the self-attention computation, and computes feature correlations and positional correlations separately using different projection matrices. In addition, we propose to scale feature correlations with the positional correlations and the aggressiveness of this multiplicative interaction can be configured using a parameter called amplitude. Moreover, we show that the PE matrix can be sliced to reduce model parameters. Our results on National Speech Corpus (NSC) show that Transformer encoders with Scaled Untied RPE achieves relative improvements of 1.9% in accuracy and up to 50.9% in latency over a Conformer baseline respectively.
Lahiru Samarakoon, Ivan Fung
SLT2
2018 End-To-End Low-Resource Lip-Reading with Maxout Cnn and Lstm
abstract
Lip-reading is the task of recognizing speech solely from the visual movement of the mouth. Although recent works have demonstrated the effectiveness of convolutional neural network (CNN) and long short-term memory (LSTM) recurrent neural network in lip-reading, similar architectures under low-resource scenario have not yet been explored. Our proposed end - to-end deep learning model fuses conventional CNN and bidirectional LSTM (BLSTM) together with max-out activation units (maxout-CNN-BLSTM), and is capable of attaining a word accuracy of 87.6% on the Ouluvs2 corpus, offering an absolute improvement of 3.1 % to the previous state-of-the-art auto-encoder-BLSTM model. To the best of our knowledge, this is the first end - to-end low -resource lip-reading system that does not require any separate feature extraction stage nor pre-training phase with external data resources. This is also the first work that utilizes maxout units in both CNN and LSTM in one single deep neural network.
Ivan Fung, Brian Kan-Wing Mak
ICASSP1