EDBT 2026 Demo / reviewers in the wild / expert
Ladislav Mosner
dblp:226/1989
· DBLP profile ↗
27ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0001-8175-2244ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 8 first-author · 17 since 2021Artificial intelligence and machine learning · 17 · 5 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trainable multi-channel front-ends for joint beamforming and speaker embedding extractionabstractMulti-channel speaker verification (SV), employing numerous microphones for capturing enrollment and/or test recordings, gained attention for its benefits in far-field scenarios. While some studies approach the problem by designing multi-channel embedding extractors, we focus on building and thoroughly analyzing a framework integrating beamforming pre-processing paired with single-channel embedding extraction. This strategy benefits from accommodating both multi-channel and single-channel inputs. Furthermore, it provides human-interpretable intermediate output — enhanced speech — that can be independently evaluated and related to SV performance. We first focus on the front-end, taking advantage of deep-learning source separation for direct or indirect mask estimation required by the beamformer. We alternate single-channel network architectures, subsequently extended to multi-channel ones by reference channel attention (RCA). We also analyze the impact of beamformer and network output fusion. Finally, we show improvements brought by end-to-end fine-tuning the entire architecture facilitated by our newly designed multi-channel corpus, MultiSV2, extending our previous MultiSV dataset. Ladislav Mosner, Oldrich Plchot, Lukás Burget, Jan Cernocký, Meng Yu 0003 |
Comput. Speech Lang. | 1 |
| 2025 | State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb DataabstractIn this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source VoxCeleb videos and the names of the celebrities without knowing the time intervals in which they appear in the recording. We experiment with hyperparameters and embedding extractors based on ResNet and WavLM. We show that the method achieves state-of-the-art results in speaker verification, comparable with training the extractors in a standard supervised way on the VoxCeleb dataset. We also extend it by considering segments be-longing to unknown speakers appearing alongside the celebrities, which are typically discarded. Removing the need for speaker timestamps and multimodal alignment, our method unlocks the use of large-scale weakly labeled speech data, enabling direct training of state-of-the-art embedding extractors and offering a visual-free alternative to VoxCeleb-style dataset creation. Sara Barahona, Ladislav Mosner, Themos Stafylakis, Oldrich Plchot, Junyi Peng, Lukás Burget, Jan Cernocký |
ASRU | 2 |
| 2025 | CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker VerificationabstractSelf-supervised learning (SSL) models for speaker verification (SV) have gained significant attention in recent years. However, existing SSL-based SV systems often struggle to capture local temporal dependencies and generalize across different tasks. In this paper, we propose context-aware multi-head factorized attentive pooling (CA-MHFA), a lightweight framework that incorporates contextual information from surrounding frames. CA-MHFA leverages grouped, learnable queries to effectively model contextual dependencies while maintaining efficiency by sharing keys and values across groups. Experimental results on the VoxCeleb dataset show that CA-MHFA achieves EERs of 0.42%, 0.48%, and 0.96% on Vox1-O, Vox1-E, and Vox1-H, respectively, outperforming complex models like WavLM-TDNN with fewer parameters and faster convergence. Additionally, CA-MHFA demonstrates strong generalization across multiple SSL models and tasks, including emotion recognition and anti-spoofing, highlighting its robustness and versatility.1 Junyi Peng, Ladislav Mosner, Lin Zhang 0054, Oldrich Plchot, Themos Stafylakis, Lukás Burget, Jan Cernocký |
ICASSP | 2 |
| 2025 | Analysis of ABC Frontend Audio Systems for the NIST-SRE24abstractSection: Speaker Recognition Sara Barahona, Anna Silnova, Ladislav Mosner, Junyi Peng, Oldrich Plchot, Johan Rohdin, Lin Zhang 0054, Jiangyu Han, Petr Pálka, Federico Landini, Lukás Burget, Themos Stafylakis, Sandro Cumani, Dominik Bobos, Miroslav Hlavácek, Martin Kodovsky, Tomás Pavlícek |
INTERSPEECH | 3 |
| 2025 | Analysis of the ABC Classification Backends for NIST SRE24
Sandro Cumani, Anna Silnova, Sara Barahona, Ladislav Mosner, Oldrich Plchot, Johan Rohdin |
INTERSPEECH | 4 |
| 2024 | Multi-Channel Extension of Pre-trained Models for Speaker VerificationabstractInternational audience Ladislav Mosner, Romain Serizel, Lukás Burget, Oldrich Plchot, Emmanuel Vincent 0001, Junyi Peng, Jan Cernocký |
INTERSPEECH | 1 |
| 2023 | Speech-Based Emotion Recognition with Self-Supervised Models Using Attentive Channel-Wise Correlations and Label SmoothingabstractWhen recognizing emotions from speech, we encounter two common problems: how to optimally capture emotion-relevant information from the speech signal and how to best quantify or categorize the noisy subjective emotion labels. Self-supervised pre-trained representations can robustly capture information from speech enabling state-of-the-art results in many downstream tasks including emotion recognition. However, better ways of aggregating the information across time need to be considered as the relevant emotion information is likely to appear piecewise and not uniformly across the signal. For the labels, we need to take into account that there is a substantial degree of noise that comes from the subjective human annotations. In this paper, we propose a novel approach to attentive pooling based on correlations between the representations’ coefficients combined with label smoothing, a method aiming to reduce the confidence of the classifier on the training labels. We evaluate our proposed approach on the benchmark dataset IEMOCAP, and demonstrate high performance surpassing that in the literature. The code to reproduce the results is available at github.com/skakouros/s3prl_attentive_correlation. Sofoklis Kakouros, Themos Stafylakis, Ladislav Mosner, Lukás Burget |
ICASSP | 3 |
| 2023 | Parameter-Efficient Transfer Learning of Pre-Trained Transformer Models for Speaker Verification Using AdaptersabstractRecently, the pre-trained Transformer models have received a rising interest in the field of speech processing thanks to their great success in various downstream tasks. However, most fine-tuning approaches update all the parameters of the pre-trained model, which becomes prohibitive as the model size grows and sometimes results in over-fitting on small datasets. In this paper, we conduct a comprehensive analysis of applying parameter-efficient transfer learning (PETL) methods to reduce the required learnable parameters for adapting to speaker verification tasks. Specifically, during the fine-tuning process, the pre-trained models are frozen, and only lightweight modules inserted in each Transformer block are trainable (a method known as adapters). Moreover, to boost the performance in a cross-language low-resource scenario, the Transformer model is further tuned on a large intermediate dataset before directly fine-tuning it on a small dataset. With updating fewer than 4% of parameters, (our proposed) PETL-based methods achieve comparable performances with full fine-tuning methods (Vox1-O: 0.55%, Vox1-E: 0.82%, Vox1-H:1.73%). Junyi Peng, Themos Stafylakis, Rongzhi Gu, Oldrich Plchot, Ladislav Mosner, Lukás Burget, Jan Cernocký |
ICASSP | 5 |
| 2023 | Description and Analysis of ABC Submission to NIST LRE 2022
Pavel Matejka, Anna Silnova, Josef Slavícek, Ladislav Mosner, Oldrich Plchot, Michal Klco, Junyi Peng, Themos Stafylakis, Lukás Burget |
INTERSPEECH | 4 |
| 2023 | Multi-Channel Speech Separation with Cross-Attention and Beamforming
Ladislav Mosner, Oldrich Plchot, Junyi Peng, Lukás Burget, Jan Cernocký |
INTERSPEECH | 1 |
| 2023 | Improving Speaker Verification with Self-Pretrained Transformer Models
Junyi Peng, Oldrich Plchot, Themos Stafylakis, Ladislav Mosner, Lukás Burget, Jan Cernocký |
INTERSPEECH | 4 |
| 2022 | Multisv: Dataset for Far-Field Multi-Channel Speaker VerificationabstractMotivated by unconsolidated data situation and the lack of a standard benchmark in the field, we complement our previous efforts and present a comprehensive corpus designed for training and evaluating text-independent multi-channel speaker verification systems. It can be readily used also for experiments with dereverberation, denoising, and speech enhancement. We tackled the ever-present problem of the lack of multi-channel training data by utilizing data simulation on top of clean parts of the Voxceleb corpus. The development and evaluation trials are based on a retransmitted Voices Obscured in Complex Environmental Settings (VOiCES) corpus, which we modified to provide multi-channel trials. We publish full recipes that create the dataset from public sources as the MultiSV dataset, and we provide results with two of our multi-channel speaker verification systems with neural network-based beamforming based either on predicting ideal binary masks or the more recent Conv-TasNet. Ladislav Mosner, Oldrich Plchot, Lukás Burget, Jan Cernocký |
ICASSP | 1 |
| 2022 | Multi-Channel Speaker Verification with Conv-Tasnet Based BeamformerabstractWe focus on the problem of speaker recognition in far-field multichannel data. The main contribution is introducing an alternative way of predicting spatial covariance matrices (SCMs) for a beamformer from the time domain signal. We propose to use ConvTasNet, a well-known source separation model, and we adapt it to perform speech enhancement by forcing it to separate speech and additive noise. We experiment with using the STFT of Conv-TasNet outputs to obtain SCMs of speech and noise, and finally, we fine-tune this multi-channel frontend w.r.t. speaker verification objective. We successfully tackle the problem of the lack of a realistic multichannel training set by using simulated data of MultiSV corpus. The analysis is performed on its retransmitted and simulated test parts. We achieve consistent improvements with a 2.7 times smaller model than the baseline based on a scheme with mask estimating NN. Ladislav Mosner, Oldrich Plchot, Lukás Burget, Jan Cernocký |
ICASSP | 1 |
| 2022 | Probabilistic Spherical Discriminant Analysis: An Alternative to PLDA for length-normalized embeddingsabstractIn speaker recognition, where speech segments are mapped to embeddings on the unit hypersphere, two scoring backends are commonly used, namely cosine scoring or PLDA.Both have advantages and disadvantages, depending on the context.Cosine scoring follows naturally from the spherical geometry, but for PLDA the blessing is mixed-length normalization Gaussianizes the between-speaker distribution, but violates the assumption of a speaker-independent within-speaker distribution.We propose PSDA, an analogue to PLDA that uses Von Mises-Fisher distributions on the hypersphere for both within and between-class distributions.We show how the self-conjugacy of this distribution gives closed-form likelihood-ratio scores, making it a drop-in replacement for PLDA at scoring time.All kinds of trials can be scored, including single-enroll and multienroll verification, as well as more complex likelihood-ratios that could be used in clustering and diarization.Learning is done via an EM-algorithm with closed-form updates.We explain the model and present some first experiments. Niko Brümmer, Albert Swart, Ladislav Mosner, Anna Silnova, Oldrich Plchot, Themos Stafylakis, Lukás Burget |
INTERSPEECH | 3 |
| 2022 | Learnable Sparse Filterbank for Speaker Verification
Junyi Peng, Rongzhi Gu, Ladislav Mosner, Oldrich Plchot, Lukás Burget, Jan Cernocký |
INTERSPEECH | 3 |
| 2022 | Training speaker embedding extractors using multi-speaker audio with unknown speaker boundariesabstractIn this paper, we demonstrate a method for training speaker embedding extractors using weak annotation.More specifically, we are using the full VoxCeleb recordings and the name of the celebrities appearing on each video without knowledge of the time intervals the celebrities appear in the video.We show that by combining a baseline speaker diarization algorithm that requires no training or parameter tuning, a modified loss with aggregation over segments, and a two-stage training approach, we are able to train a competitive ResNet-based embedding extractor.Finally, we experiment with two different aggregation functions and analyze their behaviour in terms of their gradients. Themos Stafylakis, Ladislav Mosner, Oldrich Plchot, Johan Rohdin, Anna Silnova, Lukás Burget, Jan Cernocký |
INTERSPEECH | 2 |
| 2022 | An Attention-Based Backend Allowing Efficient Fine-Tuning of Transformer Models for Speaker VerificationabstractIn recent years, self-supervised learning paradigm has received extensive attention due to its great success in various down-stream tasks. However, the fine-tuning strategies for adapting those pre-trained models to speaker verification task have yet to be fully explored. In this paper, we analyze several feature extraction approaches built on top of a pre-trained model, as well as regularization and a learning rate scheduler to stabilize the fine-tuning process and further boost performance: multi-head factorized attentive pooling is proposed to factorize the comparison of speaker representations into multiple phonetic clusters. We regularize towards the parameters of the pre-trained model and we set different learning rates for each layer of the pre-trained model during fine-tuning. The experimental results show our method can significantly shorten the training time to 4 hours and achieve SOTA performance: 0.59%, 0.79% and 1.77% EER on Vox1-O, Vox1-E and Vox1-H, respectively.11Code is available at https://github.com/JunyiPeng00/IEEE-SLT22-Pretrained-Model-for-SV. Junyi Peng, Oldrich Plchot, Themos Stafylakis, Ladislav Mosner, Lukás Burget, Jan Cernocký |
SLT | 4 |
| 2022 | Extracting Speaker and Emotion Information from Self-Supervised Speech Models via Channel-Wise CorrelationsabstractSelf-supervised learning of speech representations from large amounts of unlabeled data has enabled state-of-the-art results in several speech processing tasks. Aggregating these speech representations across time is typically approached by using descriptive statistics, and in particular, using the first - and second-order statistics of representation coefficients. In this paper, we examine an alternative way of extracting speaker and emotion information from self-supervised trained models, based on the correlations between the coefficients of the representations - correlation pooling. We show improvements over mean pooling and further gains when the pooling methods are combined via fusion. The code is available at github.com/Lamomal/s3prl_correlation. Themos Stafylakis, Ladislav Mosner, Sofoklis Kakouros, Oldrich Plchot, Lukás Burget, Jan Cernocký |
SLT | 2 |
| 2020 | But System for the Second Dihard Speech Diarization ChallengeabstractThis paper describes the winning systems developed by the BUT team for the four tracks of the Second DIHARD Speech Diarization Challenge. For tracks 1 and 2 the systems were mainly based on performing agglomerative hierarchical clustering (AHC) of x-vectors, followed by another x-vector clustering based on Bayes hidden Markov model and variational Bayes inference. We provide a comparison of the improvement given by each step and share the implementation of the core of the system. For tracks 3 and 4 with recordings from the Fifth CHiME Challenge, we explored different approaches for doing multi-channel diarization and our best performance was obtained when applying AHC on the fusion of per channel probabilistic linear discriminant analysis scores. Federico Landini, Shuai Wang 0016, Mireia Díez, Lukás Burget, Pavel Matejka, Katerina Zmolíková, Ladislav Mosner, Anna Silnova, Oldrich Plchot, Ondrej Novotný, Hossein Zeinali, Johan Rohdin |
ICASSP | 7 |
| 2020 | 13 years of speaker recognition research at BUT, with longitudinal analysis of NIST SRE
Pavel Matejka, Oldrich Plchot, Ondrej Glembek, Lukás Burget, Johan Rohdin, Hossein Zeinali, Ladislav Mosner, Anna Silnova, Ondrej Novotný, Mireia Díez, Jan Cernocký |
Comput. Speech Lang. | 7 |
| 2019 | Speaker Verification with Application-Aware BeamformingabstractMultichannel speech processing applications usually employ beamformers as means of speech enhancement through spatial filtering. Beamformers with learnable parameters require training to minimize a loss function that is not necessarily correlated with the final objective. In this paper, we present a framework employing recent neural network based generalized eigenvalue beamformer and application-specific model that allows for optimization of beamformer w.r.t. target application. In our case, the application is speaker verification which utilizes a speaker embedding (x-vector) extractor that conveniently comes with desired loss. We show that application-specific training of the beamformer brings performance improvements over a system trained in the standard way. We perform our analysis on the recently introduced VOiCES corpus which contains multichannel data and allows us to modify the evaluation trials such that enrollment recordings remain single-channel and test utterances are multichannel. Ladislav Mosner, Oldrich Plchot, Johan Rohdin, Lukás Burget, Jan Cernocký |
ASRU | 1 |
| 2019 | Improving Noise Robustness of Automatic Speech Recognition via Parallel Data and Teacher-student LearningabstractFor real-world speech recognition applications, noise robustness is still a challenge. In this work, we adopt the teacher-student (T/S) learning technique using a parallel clean and noisy corpus for improving automatic speech recognition (ASR) performance under multimedia noise. On top of that, we apply a logits selection method which only preserves the k highest values to prevent wrong emphasis of knowledge from the teacher and to reduce bandwidth needed for transferring data. We incorporate up to 8000 hours of untranscribed data for training and present our results on sequence trained models apart from cross entropy trained ones. The best sequence trained student model yields relative word error rate (WER) reductions of approximately 10.1%, 28.7% and 19.6% on our clean, simulated noisy and real test sets respectively comparing to a sequence trained teacher. Ladislav Mosner, Minhua Wu, Anirudh Raju, Sree Hari Krishnan Parthasarathi, Ken'ichi Kumatani, Shiva Sundaram, Roland Maas, Björn Hoffmeister |
ICASSP | 1 |
| 2019 | Analysis of BUT Submission in Far-Field Scenarios of VOiCES 2019 Challenge
Pavel Matejka, Oldrich Plchot, Hossein Zeinali, Ladislav Mosner, Anna Silnova, Lukás Burget, Ondrej Novotný, Ondrej Glembek |
INTERSPEECH | 4 |
| 2019 | Analysis of BUT Submission in Far-Field Scenarios of VOiCES 2019 Challenge
Pavel Matejka, Oldrich Plchot, Hossein Zeinali, Ladislav Mosner, Anna Silnova, Lukás Burget, Ondrej Novotný, Ondrej Glembek |
INTERSPEECH | 4 |
| 2018 | Dereverberation and Beamforming in Far-Field Speaker RecognitionabstractThis paper deals with far-field speaker recognition. On a corpus of NIST SRE 2010 data retransmitted in a real room with multiple microphones, we first demonstrate how room acoustics cause significant degradation of state-of-the-art i-vector based speaker recognition system. We then investigate several techniques to improve the performances ranging from probabilistic linear discriminant analysis (PLDA) re-training, through dereverberation, to beamforming. We found that weighted prediction error (WPE) based dereverberation combined with generalized eigenvalue beamformer with power-spectral density (PSD) weighting masks generated by neural networks (NN) provides results approaching the clean close-microphone setup. Further improvement was obtained by re-training PLDA or the mask-generating NNs on simulated target data. The work shows that a speaker recognition system working robustly in the far-field scenario can be developed. Ladislav Mosner, Pavel Matejka, Ondrej Novotný, Jan Cernocký |
ICASSP | 1 |
| 2018 | BUT System for DIHARD Speech Diarization Challenge 2018
Mireia Díez, Federico Landini, Lukás Burget, Johan Rohdin, Anna Silnova, Katerina Zmolíková, Ondrej Novotný, Karel Veselý, Ondrej Glembek, Oldrich Plchot, Ladislav Mosner, Pavel Matejka |
INTERSPEECH | 11 |
| 2018 | Dereverberation and Beamforming in Robust Far-Field Speaker Recognition
Ladislav Mosner, Oldrich Plchot, Pavel Matejka, Ondrej Novotný, Jan Cernocký |
INTERSPEECH | 1 |