Rogier C. van Dalen

dblp:97/4127 · also Rogier van Dalen · DBLP profile ↗
← Back
33ranked-venue papers
14as first author
13since 2021 · last 2026
0000-0002-9603-5771ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 12 first-author · 11 since 2021Artificial intelligence and machine learning · 24 · 10 first-author · 9 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DP-DyLoRA: Fine-Tuning Transformer-Based Models On-Device under Differentially Private Federated Learning using Dynamic Low-Rank Adaptation
abstract
Federated learning (FL) allows clients to collaboratively train a global model without sharing their local data with a server. However, clients' contributions to the server can still leak sensitive information. Differential privacy (DP) addresses such leakage by providing formal privacy guarantees, with mechanisms that add randomness to the clients' contributions. The randomness makes it infeasible to train large transformer-based models, common in modern federated learning systems. In this work, we empirically evaluate the practicality of fine-tuning large scale on-device transformer-based models with differential privacy in a federated learning system. We conduct comprehensive experiments on various system properties for tasks spanning a multitude of domains: speech recognition, computer vision (CV) and natural language understanding (NLU). Our results show that full fine-tuning under differentially private federated learning (DP-FL) generally leads to huge performance degradation which can be alleviated by reducing the dimensionality of contributions through parameter-efficient fine-tuning (PEFT). Our benchmarks of existing DP-PEFT methods show that DP-Low-Rank Adaptation (DP-LoRA) consistently outperforms other methods. An even more promising approach, DyLoRA, which makes the low rank variable, when naively combined with FL would straightforwardly break differential privacy. We therefore propose an adaptation method that can be combined with differential privacy and call it DP-DyLoRA. Finally, we are able to reduce the accuracy degradation and word error rate (WER) increase due to DP to less than 2% and 7% respectively with 1 million clients and a stringent privacy budget of $ε=2$.
Karthikeyan Saravanan, Rogier C. van Dalen, Haaris Mehmood, David Tuckey, Mete Ozay
ICC3
2025 Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
abstract
Recent work suggests that large language models (LLMs) can improve performance of speech tasks compared to existing systems. To support their claims, results on LibriSpeech and Common Voice are often quoted. However, this work finds that a substantial amount of the LibriSpeech and Common Voice evaluation sets appear in public LLM pretraining corpora. This calls into question the reliability of findings drawn from these two datasets. To measure contamination impact, LLMs trained with/without contamination are compared. A contaminated LLM is more likely to generate test sentences it has seen during training. Then, speech recognisers based on LLMs are compared. They show only subtle error rate differences if the LLM is contaminated, but assign significantly higher probabilities to transcriptions seen during LLM training. Results show that LLM outputs can be biased by tiny amounts of data contamination, highlighting the importance of evaluating LLM-based speech systems with held-out data.
Yuan Tseng, Titouan Parcollet, Rogier C. van Dalen, Shucong Zhang, Sourav Bhattacharya
ASRU3
2025 Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
abstract
Self-attention relies on positional embeddings to encode input order. Relative Position (RelPos) embeddings are widely used in Automatic Speech Recognition (ASR). However, RelPos has quadratic time complexity to input length and is often incompatible with fast GPU implementations of attention. In contrast, Rotary Positional Embedding (RoPE) rotates each input vector based on its absolute position, taking linear time to sequence length, implicitly encoding relative distances through self-attention dot products. Thus, it is usually compatible with efficient attention. However, its use in ASR remains underexplored. This work evaluates RoPE across diverse ASR tasks with training data ranging from 100 to 50,000 hours, covering various speech types (read, spontaneous, clean, noisy) and different accents in both streaming and non-streaming settings. ASR error rates are similar or better than RelPos, while training time is reduced by up to $21 \%$. Code is available via the SpeechBrain toolkit.
Shucong Zhang, Titouan Parcollet, Rogier C. van Dalen, Sourav Bhattacharya
ASRU3
2025 Globally Normalizing the Transducer for Streaming Speech Recognition
abstract
The Transducer (e.g. RNN-Transducer or Conformer-Transducer) generates an output label sequence as it traverses the input sequence. It is straightforward to use in streaming mode, where it generates partial hypotheses before the complete input has been seen. This makes it popular in speech recognition. However, in streaming mode the Transducer has a mathematical flaw which, simply put, restricts the model’s ability to change its mind. The fix is to replace local normalization (e.g. a softmax) with global normalization, but then the loss function becomes impossible to evaluate exactly. A recent paper proposes to solve this by approximating the model, degrading performance. Instead, this paper proposes to approximate the loss function, allowing global normalization to apply to a state-of-the-art streaming model. Global normalization reduces its word error rate by 9–11% relative, closing almost half the gap between streaming and offline mode.
Rogier C. van Dalen
ICASSP1
2025 Linear Time Complexity Conformers with SummaryMixing for Streaming Speech Recognition
abstract
Automatic speech recognition (ASR) with an encoder equipped with self-attention, whether streaming or non-streaming, takes quadratic time in the length of the speech utterance. This slows down training and decoding, increase the cost, and limits the deployment of the ASR in constrained devices. SummaryMixing is a promising linear-time complexity alternative to self-attention for non-streaming speech recognition that, for the first time, preserves or outperforms the accuracy of self-attention models. Unfortunately, the original definition of SummaryMixing is not suited to streaming speech recognition. Hence, this work extends SummaryMixing to a Conformer Transducer that works in both a streaming and an offline mode. It shows that this new linear-time complexity speech encoder outperforms self-attention in both scenarios while requiring less compute and memory during training and decoding.
Titouan Parcollet, Rogier C. van Dalen, Shucong Zhang, Sourav Bhattacharya
ICASSP2
2025 Robust Unsupervised Adaptation of a Speech Recogniser Using Entropy Minimisation and Speaker Codes
Rogier C. van Dalen, Shucong Zhang, Titouan Parcollet, Sourav Bhattacharya
INTERSPEECH1
2025 Loquacious Set: 25, 000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
Titouan Parcollet, Yuan Tseng, Shucong Zhang, Rogier C. van Dalen
INTERSPEECH4
2024 SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition and Understanding
Titouan Parcollet, Rogier C. van Dalen, Shucong Zhang, Sourav Bhattacharya
INTERSPEECH2
2024 Linear-Complexity Self-Supervised Learning for Speech Processing
Shucong Zhang, Titouan Parcollet, Rogier C. van Dalen, Sourav Bhattacharya
INTERSPEECH3
2024 pfl-research: simulation framework for accelerating research in Private Federated Learning
Filip Granqvist, Congzheng Song, Áine Cahill, Rogier C. van Dalen, Martin Pelikan, Yi Sheng Chan, Xiaojun Feng, Natarajan Krishnaswami, Vojta Jina, Mona Chitnis
NeurIPS4
2023 Training Large-Vocabulary Neural Language Models by Private Federated Learning for Resource-Constrained Devices
abstract
Federated Learning (FL) is a technique to train models on distributed edge devices with local data samples. Differential Privacy (DP) can be applied with FL to provide a formal privacy guarantee for sensitive data on device. Our goal is to train a large neural network language model (NNLM) on compute-constrained devices while preserving privacy using FL and DP. However, the noise required to guarantee differential privacy increases as the model size grows, which often prevents convergence. We propose Partial Embedding Updates (PEU), a novel technique to reduce the impact of DP-noise by decreasing payload size. Furthermore, we adopt Low Rank Adaptation (LoRA) and Noise Contrastive Estimation (NCE) to reduce the memory demands of large models on compute-constrained devices. We demonstrate in simulation and with real devices that this combination of techniques makes it possible to train large-vocabulary language models while preserving accuracy and privacy.
Mingbin Xu, Congzheng Song, Neha Agrawal, Filip Granqvist, Rogier C. van Dalen, Arturo Argueta, Shiyi Han, Yaqiao Deng, Leo Liu, Anmol Walia, Alex Jin
ICASSP6
2023 On the (In)Efficiency of Acoustic Feature Extractors for Self-Supervised Speech Representation Learning
abstract
International audience
Titouan Parcollet, Shucong Zhang, Rogier C. van Dalen, Alberto Gil C. P. Ramos, Sourav Bhattacharya
INTERSPEECH3
2023 Real-Time Personalised Speech Enhancement Transformers with Dynamic Cross-attended Speaker Representations
Shucong Zhang, Malcolm Chadwick, Alberto Gil C. P. Ramos, Titouan Parcollet, Rogier C. van Dalen, Sourav Bhattacharya
INTERSPEECH5
2020 Improving On-Device Speaker Verification Using Federated Learning with Privacy
abstract
Information on speaker characteristics can be useful as side information in improving speaker recognition accuracy. However, such information is often private. This paper investigates how privacy-preserving learning can improve a speaker verification system, by enabling the use of privacy-sensitive speaker data to train an auxiliary classification model that predicts vocal characteristics of speakers. In particular, this paper explores the utility achieved by approaches which combine different federated learning and differential privacy mechanisms. These approaches make it possible to train a central model while protecting user privacy, with users' data remaining on their devices. Furthermore, they make learning on a large population of speakers possible, ensuring good coverage of speaker characteristics when training a model. The auxiliary model described here uses features extracted from phrases which trigger a speaker verification system. From these features, the model predicts speaker characteristic labels considered useful as side information. The knowledge of the auxiliary model is distilled into a speaker verification system using multi-task learning, with the side information labels predicted by this auxiliary model being the additional task. This approach results in a 6% relative improvement in equal error rate over a baseline system.
Filip Granqvist, Matt Seigel, Rogier C. van Dalen, Áine Cahill, Matthias Paulik
INTERSPEECH3
2018 Towards automatic assessment of spontaneous spoken English
Yu Wang 0027, Mark J. F. Gales, Kate M. Knill, Konstantinos Kyriakopoulos, Andrey Malinin, Rogier C. van Dalen, M. Rashid
Speech Commun.6
2016 Off-topic Response Detection for Spontaneous Spoken English Assessment
abstract
Automatic spoken language assessment systems are becoming increasingly important to meet the demand for English second language learning.This is a challenging task due to the high error rates of, even state-of-the-art, non-native speech recognition.Consequently current systems primarily assess fluency and pronunciation.However, content assessment is essential for full automation.As a first stage it is important to judge whether the speaker responds on topic to test questions designed to elicit spontaneous speech.Standard approaches to off-topic response detection assess similarity between the response and question based on bag-of-words representations.An alternative framework based on Recurrent Neural Network Language Models (RNNLM) is proposed in this paper.The RNNLM is adapted to the topic of each test question.It learns to associate example responses to questions with points in a topic space constructed using these example responses.Classification is done by ranking the topic-conditional posterior probabilities of a response.The RNNLMs associate a broad range of responses with each topic, incorporate sequence information and scale better with additional training data, unlike standard methods.On experiments conducted on data from the Business Language Testing Service (BULATS) this approach outperforms standard approaches.
Andrey Malinin, Rogier C. van Dalen, Kate M. Knill, Yu Wang 0027, Mark J. F. Gales
ACL (1)2
2016 Towards Using Conversations with Spoken Dialogue Systems in the Automated Assessment of Non-Native Speakers of English
abstract
Diane Litman, Steve Young, Mark Gales, Kate Knill, Karen Ottewell, Rogier van Dalen, David Vandyke. Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2016.
Diane J. Litman, Steve J. Young, Mark J. F. Gales, Kate M. Knill, Karen Ottewell, Rogier C. van Dalen, David Vandyke
SIGDIAL Conference6
2015 Structured discriminative models using deep neural-network features
abstract
State-of-the-art speech recognisers employ neural networks in various configurations. A standard (hybrid) speech recogniser computes the likelihood for one time frame and state, using only one out of thousands of possible neural-network outputs. However, the whole output vector carries information. In this paper, features from state-of-the-art speech recognisers are collected per phone given a particular context, and input to a discriminative log-linear model. The log-linear model is trained with conditional maximum likelihood or a large-margin criterion. A key element is the prior on the parameters of the log-linear model. The mean of the prior is set to the point where the performance of the original systems is attained. The log-linear model then provides an additional increase over the state-of-the-art performance of the individual systems.
Rogier C. van Dalen, Anton Ragni, Chao Zhang 0031, Mark J. F. Gales
ASRU1
2015 Improving multiple-crowd-sourced transcriptions using a speech recogniser
abstract
This paper introduces a method to produce high-quality transcriptions of speech data from only two crowd-sourced transcriptions. These transcriptions, produced cheaply by people on the Internet, for example through Amazon Mechanical Turk, are often of low quality. Often, multiple crowd-sourced transcriptions are combined to form one transcription of higher quality. However, the state of the art is to use essentially a form of majority voting, which requires at least three transcriptions for each utterance. This paper shows how to refine this approach to work with only two transcriptions. It then introduces a method that uses a speech recogniser (bootstrapped on a simple combination scheme) to combine transcriptions. When only two crowd-sourced transcriptions are available, on a noisy data set this improves the word error rate to gold-standard transcriptions by 21% relative.
Rogier C. van Dalen, Kate M. Knill, Pirros Tsiakoulis, Mark J. F. Gales
ICASSP1
2015 Annotating large lattices with the exact word error
abstract
The acoustic model in modern speech recognisers is trained discriminatively, for example with the minimum Bayes risk. This criterion is hard to compute exactly, so that it is normally approximated by a criterion that uses fixed alignments of lat-tice arcs. This approximation becomes particularly problematic with new types of acoustic models that require flexible align-ments. It would be best to annotate lattices with the risk mea-sure of interest, the exact word error. However, the algorithm for this uses finite-state automaton determinisation, which has exponential complexity and runs out of memory for large lat-tices. This paper introduces a novel method for determinis-ing and minimising finite-state automata incrementally. Since it uses less memory, it can be applied to larger lattices. Index Terms: speech recognition, discriminative training, min-imum Bayes risk
Rogier C. van Dalen, Mark J. F. Gales
INTERSPEECH1
2014 Infinite structured support vector machines for speech recognition
abstract
Discriminative models, like support vector machines (SVMs), have been successfully applied to speech recognition and improved performance. A Bayesian non-parametric version of the SVM, the infinite SVM, improves on the SVM by allowing more flexible decision boundaries. However, like SVMs, infinite SVMs model each class separately, which restricts them to classifying one word at a time. A generalisation of the SVM is the structured SVM, whose classes can be sequences of words that share parameters. This paper studies a combination of Bayesian non-parametrics and structured models. One specific instance called infinite structured SVM is discussed in detail, which brings the advantages of the infinite SVM to continuous speech recognition.
Rogier C. van Dalen, Shixiong Zhang 0001, Mark J. F. Gales
ICASSP2
2013 Efficient decoding with generative score-spaces using the expectation semiring
abstract
State-of-the-art speech recognisers are usually based on hidden Markov models (HMMs). They model a hidden symbol sequence with a Markov process, with the observations independent given that sequence. These assumptions yield efficient algorithms, but limit the power of the model. An alternative model that allows a wide range of features, including word- and phone-level features, is a log-linear model. To handle, for example, word-level variable-length features, the original feature vectors must be segmented into words. Thus, decoding must find the optimal combination of segmentation of the utterance into words and word sequence. Features must therefore be extracted for each possible segment of audio. For many types of features, this becomes slow. In this paper, long-span features are derived from the likelihoods of word HMMs. Derivatives of the log-likelihoods, which break the Markov assumption, are appended. Previously, decoding with this model took cubic time in the length of the sequence, and longer for higher-order derivatives. This paper shows how to decode in quadratic time.
Rogier C. van Dalen, Anton Ragni, Mark J. F. Gales
ICASSP1
2013 Infinite support vector machines in speech recognition
abstract
Generative feature spaces provide an elegant way to apply dis-criminative models in speech recognition, and system perfor-mance has been improved by adapting this framework. How-ever, the classes in the feature space may be not linearly sepa-rable. Applying a linear classifier then limits performance. In-stead of a single classifier, this paper applies a mixture of ex-perts. This model trains different classifiers as experts focusing on different regions of the feature space. However, the num-ber of experts is not known in advance. This problem can be bypassed by employing a Bayesian non-parametric model. In this paper, a specific mixture of experts based on the Dirichlet process, namely the infinite support vector machine, is studied. Experiments conducted on the noise-corrupted continuous digit task AURORA 2 show the advantages of this Bayesian non-parametric approach. Index Terms: generative feature space, Bayesian non-parametric, Dirichlet process, mixture of experts, infinite sup-
Rogier C. van Dalen, Mark J. F. Gales
INTERSPEECH2
2013 Importance sampling to compute likelihoods of noise-corrupted speech
Rogier C. van Dalen, Mark J. F. Gales
Comput. Speech Lang.1
2011 A variational perspective on noise-robust speech recognition
abstract
Model compensation methods for noise-robust speech recognition have shown good performance. Predictive linear transformations can approximate these methods to balance computational complexity and compensation accuracy. This paper examines both of these approaches from a variational perspective. Using a matched-pair approximation at the component level yields a number of standard forms of model compensation and predictive linear transformations. However, a tighter bound can be obtained by using variational approximations at the state level. Both model-based and predictive linear transform schemes can be implemented in this framework. Preliminary results show that the tighter bound obtained from the state-level variational approach can yield improved performance over standard schemes.
Rogier C. van Dalen, Mark J. F. Gales
ASRU1
2011 Extended VTS for Noise-Robust Speech Recognition
abstract
Model compensation is a standard way of improving the robustness of speech recognition systems to noise. A number of popular schemes are based on vector Taylor series (VTS) compensation, which uses a linear approximation to represent the influence of noise on the clean speech. To compensate the dynamic parameters, the continuous time approximation is often used. This approximation uses a point estimate of the gradient, which fails to take into account that dynamic coefficients are a function of a number of consecutive static coefficients. In this paper, the accuracy of dynamic parameter compensation is improved by representing the dynamic features as a linear transformation of a window of static features. A modified version of VTS compensation is applied to the distribution of the window of static features and, importantly, their correlations. These compensated distributions are then transformed to distributions over standard static and dynamic features. With this improved approximation, it is also possible to obtain full-covariance corrupted speech distributions. This addresses the correlation changes that occur in noise. The proposed scheme outperformed the standard VTS scheme by 10% to 20% relative on a range of tasks.
Rogier C. van Dalen, Mark J. F. Gales
IEEE Trans. Speech Audio Process.1
2010 Asymptotically exact noise-corrupted speech likelihoods
abstract
Model compensation techniques for noise-robust speech recognition approximate the corrupted speech distribution. This paper introduces a sampling method that, given speech and noise distributions and a mismatch function, in the limit calculates the corrupted speech likelihood exactly. Though it is too slow to compensate a speech recognition system, it enables a more fine-grained assessment of compensation techniques, based on the KL divergence of individual components. This makes it possible to evaluate the impact of approximations that compensation schemes make, such as the form of the mismatch function. Index Terms: speech recognition, noise robustness 1.
Rogier C. van Dalen, Mark J. F. Gales
INTERSPEECH1
2009 Extended VTS for noise-robust speech recognition
abstract
Model compensation is a standard way of improving speech recognisers' robustness to noise. Currently popular schemes are based on vector Taylor series (VTS) compensation. They often use the continuous time approximation to compensate dynamic parameters. In this paper, the accuracy of dynamic parameter compensation is improved by representing the dynamic features as a linear transformation of a window of static features. A modified version of VTS compensation is applied to the distribution of the window of static features and, importantly, their correlations. These compensated distributions are then transformed to standard static and dynamic distributions. The proposed scheme outperformed the standard VTS scheme by about 10% relative.
Rogier C. van Dalen, Mark J. F. Gales
ICASSP1
2009 Transforming features to compensate speech recogniser models for noise
abstract
To make speech recognisers robust to noise, either the features or the models can be compensated. Feature enhancement is often fast; model compensation is often more accurate, because it predicts the corrupted speech distribution. It is therefore able, for example, to take uncertainty about the clean speech into account. This paper re-analyses the recently-proposed predictive linear transformations for noise compensation as minimising the KL divergence between the predicted corrupted speech and the adapted models. New schemes are then introduced which apply observation-dependent transformations in the front-end to adapt the back-end distributions. One applies transforms in the exact same manner as the popular minimum mean square error (MMSE) feature enhancement scheme, and is as fast. The new method performs better on AURORA 2. Index Terms: speech recognition, noise robustness 1.
Rogier C. van Dalen, Federico Flego, Mark J. F. Gales
INTERSPEECH1
2009 Variational dynamic kernels for speaker verification
abstract
An important aspect of SVM-based speaker verification is the choice of dynamic kernel. Recently there has been interest in the use of kernels based on the Kullback-Leibler divergence between GMMs. Since this has no closed-form solution, typically a matched-pair upper bound is used instead. This places significant restrictions on the forms of model structure that may be used. All GMMs must contain the same number of components and must be adapted from a single background model. For many tasks this will not be optimal. In this paper, dynamic kernels are proposed based on alternative, variational approximations to the KL divergence. Unlike the matched-pair bound, these do not restrict the forms of GMM that may be used. Additionally, using a more accurate approximation of the divergence may lead to performance gains. Preliminary results using these kernels are presented on the NIST 2002 SRE dataset.
Chris Longworth, Rogier C. van Dalen, Mark J. F. Gales
INTERSPEECH2
2008 Covariance modelling for noise-robust speech recognition
abstract
Model compensation is a standard way of improving speech recognisers’ robustness to noise. Most model compensation techniques produce diagonal covariances. However, this fails to handle any changes in the feature correlations due to the noise. This paper presents a scheme that allows full-covariance matrices to be estimated. One problem is that full covariance matrix estimation will be more sensitive approximations, those for the dynamic parameters are known to crude. In this paper a linear transformation of a window of consecutive frames is used as the basis for dynamic parameter compensation. A second problem is that the resulting full covariance matrices slow down decoding. This is addressed by using predictive linear transforms that decorrelate the feature space, so that the decoder can then use diagonal covariance matrices. On a noise-corrupted Resource Management task, the proposed scheme outperformed the standard VTS compensation scheme.
Rogier C. van Dalen, Mark J. F. Gales
INTERSPEECH1
2007 Predictive linear transforms for noise robust speech recognition
abstract
It is well known that the addition of background noise alters the correlations between the elements of, for example, the MFCC feature vector. However, standard model-based compensation techniques do not modify the feature-space in which the diagonal covariance matrix Gaussian mixture models are estimated. One solution to this problem, which yields good performance, is joint uncertainty decoding (JUD) with full transforms. Unfortunately, this results in a high computational cost during decoding. This paper contrasts two approaches to approximating full JUD while lowering the computational cost. Both use predictive linear transforms to modify the feature-space: adaptation-based linear transforms, where the model parameters are restricted to be the same as the original clean system; and precision matrix modelling approaches, in particular semi-tied covariance matrices. These predictive transforms are estimated using statistics derived from the full JUD transforms rather than noisy data. The schemes are evaluated on AURORA 2 and a noise-corrupted resource management task.
Mark J. F. Gales, Rogier C. van Dalen
ASRU2
2006 Lexical stress in continuous speech recognition
abstract
Human listeners use lexical stress for word segmentation and disambiguation. We look into using lexical stress for largevocabulary speech recognition for the Dutch language. It appears that beside vowels, consonants should be taken into account. By introducing stressed phonemes, and features for spectral bands and the fundamental frequency, we reduce the word error rate by 2.6 %. Index Terms: speech recognition, lexical stress, Dutch. 1.
Rogier C. van Dalen, Pascal Wiggers, Léon J. M. Rothkrantz
INTERSPEECH1