EDBT 2026 Demo / reviewers in the wild / expert
Bolaji Yusuf
dblp:209/7131
· DBLP profile ↗
18ranked-venue papers
11as first author
13since 2021 · last 2026
0000-0001-9852-3456ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 8 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DiCoW: Diarization-conditioned Whisper for target speaker automatic speech recognition
Alexander Polok, Dominik Klement, Martin Kocour, Jiangyu Han, Federico Landini, Bolaji Yusuf, Matthew Wiesner, Sanjeev Khudanpur, Jan Cernocký, Lukás Burget |
Comput. Speech Lang. | 6 |
| 2025 | DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech RecognitionabstractThis paper presents a simple yet effective regularization for the internal language model induced by the decoder in encoder-decoder ASR models, thereby improving robustness and generalization in both in- and out-of-domain settings. The proposed method, Decoder-Centric Regularization in EncoderDecoder (DeCRED), adds auxiliary classifiers to the decoder, enabling next token prediction via intermediate logits. Empirically, DeCRED reduces the mean internal LM BPE perplexity by 36.6% relative to 11 test sets. Furthermore, this translates into actual WER improvements over the baseline in 5 of 7 in-domain and 3 of 4 out-of-domain test sets, reducing macro WER from 6.4% to 6.3% and 18.2 % to 16.2 %, respectively. On TEDLIUM3, DeCRED achieves 7.0 % WER, surpassing the baseline and encoder-centric InterCTC regularization by 0.6 % and 0.5%, respectively. Finally, we compare DeCRED with OWSM v3.1 and Whisper-medium, showing competitive WERs despite training on much less data with fewer parameters. Alexander Polok, Santosh Kesiraju, Karel Benes, Bolaji Yusuf, Lukás Burget, Jan Cernocký |
ASRU | 4 |
| 2025 | Factors affecting the in-context learning abilities of LLMs for dialogue state tracking
Pradyoth Hegde, Santosh Kesiraju, Jan Svec, Simon Sedlácek, Bolaji Yusuf, Oldrich Plchot, Deepak K. T, Jan Cernocký |
INTERSPEECH | 5 |
| 2025 | Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
Simon Sedlácek, Bolaji Yusuf, Jan Svec, Pradyoth Hegde, Santosh Kesiraju, Oldrich Plchot, Jan Cernocký |
INTERSPEECH | 2 |
| 2024 | Speculative Speech Recognition by Audio-Prefixed Low-Rank Adaptation of Language Models
Bolaji Yusuf, Murali Karthick Baskar, Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 1 |
| 2024 | Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units
Bolaji Yusuf, Jan Cernocký, Murat Saraclar |
INTERSPEECH | 1 |
| 2024 | Written Term Detection Improves Spoken Term DetectionabstractEnd-to-end (E2E) approaches to keyword search (KWS) are considerably simpler in terms of training and indexing complexity when compared to approaches which use the output of automatic speech recognition (ASR) systems. This simplification however has drawbacks due to the loss of modularity. In particular, where ASR-based KWS systems can benefit from external unpaired text via a language model, current formulations of E2E KWS systems have no such mechanism. Therefore, in this paper, we propose a multitask training objective which allows unpaired text to be integrated into E2E KWS without complicating indexing and search. In addition to training an E2E KWS model to retrieve text queries from spoken documents, we jointly train it to retrieve text queries from masked written documents. We show empirically that this approach can effectively leverage unpaired text for KWS, with significant improvements in search performance across a wide variety of languages. We conduct analysis which indicates that these improvements are achieved because the proposed method improves document representations for words in the unpaired text. Finally, we show that the proposed method can be used for domain adaptation in settings where in-domain paired data is scarce or nonexistent. Bolaji Yusuf, Murat Saraclar |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | On-the-Fly Text Retrieval for end-to-end ASR AdaptationabstractEnd-to-end speech recognition models are improved by incorporating external text sources, typically by fusion with an external language model. Such language models have to be retrained whenever the corpus of interest changes. Furthermore, since they store the entire corpus in their parameters, rare words can be challenging to recall. In this work, we propose augmenting a transducer-based ASR model with a retrieval language model, which directly retrieves from an external text corpus plausible completions for a partial ASR hypothesis. These completions are then integrated into subsequent predictions by an adapter, which is trained once, so that the corpus of interest can be switched without incurring the computational overhead of retraining. Our experiments show that the proposed model significantly improves the performance of a transducer baseline on a pair of question-answering datasets. Further, it outperforms shallow fusion on recognition of named entities by about 7% relative; when the two are combined, the relative improvement increases to 13%. Bolaji Yusuf, Aditya Gourav, Ankur Gandhe, Ivan Bulyko |
ICASSP | 1 |
| 2023 | End-to-End Open Vocabulary Keyword Search With Multilingual Neural RepresentationsabstractConventional keyword search systems operate on automatic speech recognition (ASR) outputs, which causes them to have a complex indexing and search pipeline. This has led to interest in ASR-free approaches to simplify the search procedure. We recently proposed a neural ASR-free keyword search model which achieves competitive performance while maintaining an efficient and simplified pipeline, where queries and documents are encoded with a pair of recurrent neural network encoders and the encodings are combined with a dot-product. In this paper, we extend this work with multilingual pretraining and detailed analysis of the model. Our experiments show that the proposed multilingual training significantly improves the model performance and that despite not matching a strong ASR-based conventional keyword search system for short queries and queries comprising in-vocabulary words, the proposed model outperforms the ASR-based system for long queries and queries that do not appear in the training data. Bolaji Yusuf, Jan Cernocký, Murat Saraclar |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Usted: Improving ASR with a Unified Speech and Text Encoder-DecoderabstractImproving end-to-end speech recognition by incorporating external text data has been a longstanding research topic. There has been a recent focus on training E2E ASR models that get the performance benefits of external text data without incurring the extra cost of evaluating an external language model at inference time. In this work, we propose training ASR model jointly with a set of text-to-text auxiliary tasks with which it shares a decoder and parts of the encoder. When we jointly train ASR and masked language model with the 960-hour Librispeech and Opensubtitles data respectively, we observe WER reductions of 16% and 20% on test-other and test-clean respectively over an ASR-only baseline without any extra cost at inference time, and reductions of 6% and 8% compared to a stronger MUTE-L baseline which trains the decoder with the same text data as our model. We achieve further improvements when we train masked language model on Librispeech data or when we use machine translation as the auxiliary task, without significantly sacrificing performance on the task itself. Bolaji Yusuf, Ankur Gandhe, Alex Sokolov |
ICASSP | 1 |
| 2022 | Non-Parametric Bayesian Subspace Models for Acoustic Unit DiscoveryabstractThis work investigates subspace non-parametric models for the task of learning a set of acoustic units from unlabeled speech recordings. We constrain the base-measure of a Dirichlet-Process mixture with a phonetic subspace—estimated from other source languages—to build aneducated prior, thereby forcing the learned acoustic units to resemble phones of known source languages. Two types of models are proposed: (i) the Subspace HMM (SHMM) which assumes that the phonetic subspace is the same for every language, (ii) the Hierarchical-Subspace HMM (H-SHMM) which relaxes this assumption and allows to have a language-specific subspace estimated on the unlabeled target data. These models are applied on 3 languages: English, Yoruba and Mboshi and they are compared with various competitive acoustic units discovery baselines. Experimental results show that both subspace models outperform other systems in terms of clustering quality and segmentation accuracy. Moreover, we observe that the H-SHMM provides results superior to the SHMM supporting the idea that language-specific priors are preferable to language-agnostic priors for acoustic unit discovery. Lucas Ondel Yang, Bolaji Yusuf, Lukás Burget, Murat Saraclar |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | A Hierarchical Subspace Model for Language-Attuned Acoustic Unit DiscoveryabstractIn this work, we propose a hierarchical subspace model for acoustic unit discovery. In this approach, we frame the task as one of learning embeddings on a low-dimensional phonetic subspace, and simultaneously specify the subspace itself as an embedding on a hyper-subspace. We train the hyper-subspace on a set of transcribed languages and transfer it to the target language. In the target language, we infer both the language and unit embeddings in an unsupervised manner, and in so doing, we simultaneously learn a subspace of units specific to that language and the units that dwell on it. We conduct experiments on TIMIT and two low-resource languages: Mboshi and Yoruba. Results show that our model outperforms major acoustic unit discovery techniques, both in terms of clustering quality and segmentation accuracy. Bolaji Yusuf, Lucas Ondel Yang, Lukás Burget, Jan Cernocký, Murat Saraclar |
ICASSP | 1 |
| 2021 | End-to-End Open Vocabulary Keyword SearchabstractRecently, neural approaches to spoken content retrieval have become popular. However, they tend to be restricted in their vocabulary or in their ability to deal with imbalanced test settings. These restrictions limit their applicability in keyword search, where the set of queries is not known beforehand, and where the system should return not just whether an utterance contains a query but the exact location of any such occurrences. In this work, we propose a model directly optimized for keyword search. The model takes a query and an utterance as input and returns a sequence of probabilities for each frame of the utterance of the query having occurred in that frame. Experiments show that the proposed model not only outperforms similar end-to-end models on a task where the ratio of positive and negative trials is artificially balanced, but it is also able to deal with the far more challenging task of keyword search with its inherent imbalance. Furthermore, using our system to rescore the outputs an LVCSR-based keyword search system leads to significant improvements on the latter. Bolaji Yusuf, Alican Gök, Batuhan Gündogdu, Murat Saraclar |
Interspeech | 1 |
| 2020 | Vector Quantized Temporally-Aware Correspondence Sparse Autoencoders for Zero-Resource Acoustic Unit Discovery
Batuhan Gündogdu, Bolaji Yusuf, Mansur Yesilbursa, Murat Saraclar |
INTERSPEECH | 2 |
| 2019 | Temporally-Aware Acoustic Unit Discovery for Zerospeech 2019 Challenge
Bolaji Yusuf, Alican Gök, Batuhan Gündogdu, Öykü Deniz Köse, Murat Saraclar |
INTERSPEECH | 1 |
| 2019 | An Empirical Evaluation of DTW Subsampling Methods for Keyword Search
Bolaji Yusuf, Murat Saraclar |
INTERSPEECH | 1 |
| 2019 | Generative RNNs for OOV Keyword SearchabstractThe modeling of text queries as sequences of embeddings for conducting similarity matching based search within speech features has been recently shown to improve keyword search (KWS) performance, especially for the out-of-vocabulary (OOV) terms. This technique uses a dynamic time warping based search methodology, converting the KWS problem into a pattern search problem by artificially modeling the text queries as pronunciation-based embedding sequences. This query modeling is done by concatenating and repeating frame representations for each phoneme in the keyword's pronunciation. In this letter, we propose a query model that incorporates temporal context information using recurrent neural networks (RNN) trained to generate realistic query representations. With experiments conducted on the IARPA Babel Program's Turkish and Zulu datasets, we show that the proposed RNN-based query generation yields significant improvements over the statistical query models of earlier work, and yields a comparable performance to the state-of-the-art techniques for OOV KWS. Batuhan Gündogdu, Bolaji Yusuf, Murat Saraclar |
IEEE Signal Process. Lett. | 2 |
| 2019 | Low Resource Keyword Search With Synthesized Crosslingual ExemplarsabstractThe transfer of acoustic data across languages has been shown to improve keyword search (KWS) performance in data-scarce settings. In this paper, we propose a way of performing this transfer that reduces the impact of the prevalence of out-of-vocabulary (OOV) terms on KWS in such a setting. We investigate a novel usage of multilingual features for KWS with very little training data in the target languages. The crux of our approach is the use of synthetic phone exemplars to convert the search into a query-by-example task, which we solve with the dynamic time warping algorithm. Using bottleneck features obtained from a network trained multilingually on a set of (source) languages, we train an extended distance metric learner (EDML) for four target languages from the IARPA Babel program (which are distinct from the source languages). Compared with a baseline system that is based on automatic speech recognition (ASR) with a multilingual acoustic model, we observe an average term weighted value improvement of 0.0603 absolute (74% relative) in a setting with only1 h of training data in the target language. When the data scarcity is relaxed to 10 h, we find that phone posteriors obtained by fine-tuning the multilingual network give better EDML systems. In this relaxed setting, the EDML systems still perform better than the baseline on OOV terms. Given their complementary natures, combining the EDML and the ASR-based baseline results in even further performance improvements in all settings. Bolaji Yusuf, Batuhan Gündogdu, Murat Saraclar |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |