Johan Rohdin

dblp:148/9674 · DBLP profile ↗
← Back
30ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0003-0978-2064ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 19 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Leveraging Self-Supervised Learning for Speaker Diarization
abstract
End-to-end neural diarization has evolved considerably over the past few years, but data scarcity is still a major obstacle for further improvements. Self-supervised learning methods such as WavLM have shown promising performance on several downstream tasks, but their application on speaker diarization is somehow limited. In this work, we explore using WavLM to alleviate the problem of data scarcity for neural diarization training. We use the same pipeline as Pyannote and improve the local end-to-end neural diarization with WavLM and Conformer. Experiments on far-field AMI, AISHELL-4, and AliMeeting datasets show that our method substantially outperforms the Pyannote baseline and achieves new state-of-the-art results on AMI and AISHELL4, respectively. In addition, by analyzing the system performance under different data quantity scenarios, we show that WavLM representations are much more robust against data scarcity than filterbank features, enabling less data hungry training strategies. Furthermore, we found that simulated data, usually used to train end-to-end diarization models, does not help when using WavLM in our experiments. Additionally, we also evaluate our model on the recent CHiME8 NOTSOFAR-1 task where it achieves better performance than the Pyannote baseline. Our source code is publicly available at https://github.com/BUTSpeechFIT/DiariZen.
Jiangyu Han, Federico Landini, Johan Rohdin, Anna Silnova, Mireia Díez, Lukás Burget
ICASSP3
2025 Analysis of ABC Frontend Audio Systems for the NIST-SRE24
abstract
Section: Speaker Recognition
Sara Barahona, Anna Silnova, Ladislav Mosner, Junyi Peng, Oldrich Plchot, Johan Rohdin, Lin Zhang 0054, Jiangyu Han, Petr Pálka, Federico Landini, Lukás Burget, Themos Stafylakis, Sandro Cumani, Dominik Bobos, Miroslav Hlavácek, Martin Kodovsky, Tomás Pavlícek
INTERSPEECH6
2025 Analysis of the ABC Classification Backends for NIST SRE24
Sandro Cumani, Anna Silnova, Sara Barahona, Ladislav Mosner, Oldrich Plchot, Johan Rohdin
INTERSPEECH6
2025 Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
Jiangyu Han, Federico Landini, Johan Rohdin, Anna Silnova, Mireia Díez, Jan Cernocký, Lukás Burget
INTERSPEECH3
2025 Bayesian Learning for Domain-Invariant Speaker Verification and Anti-Spoofing
Man-Wai Mak, Johan Rohdin, Kong-Aik Lee, Hynek Hermansky
INTERSPEECH3
2024 Diacorrect: Error Correction Back-End for Speaker Diarization
abstract
In this work, we propose an error correction framework, named DiaCorrect, to refine the output of a diarization system in a simple yet effective way. This method is inspired by error correction techniques in automatic speech recognition. Our model consists of two parallel convolutional encoders and a transformer-based decoder. By exploiting the interactions between the input recording and the initial system’s outputs, DiaCorrect can automatically correct the initial speaker activities to minimize the diarization errors. Experiments on 2-speaker telephony data show that the proposed DiaCorrect can effectively improve the initial model’s results. Our source code is publicly available at https://github.com/BUTSpeechFIT/diacorrect.
Jiangyu Han, Federico Landini, Johan Rohdin, Mireia Díez, Lukás Burget, Yuhang Cao, Jan Cernocký
ICASSP3
2024 Challenging margin-based speaker embedding extractors by using the variational information bottleneck
Themos Stafylakis, Anna Silnova, Johan Rohdin, Oldrich Plchot, Lukás Burget
INTERSPEECH3
2024 Advancing speaker embedding learning: Wespeaker toolkit for research and production
Shuai Wang 0016, Zhengyang Chen, Bing Han 0008, Chengdong Liang, Xu Xiang, Wen Ding 0005, Johan Rohdin, Anna Silnova, Yanmin Qian, Haizhou Li 0001
Speech Commun.9
2022 Training speaker embedding extractors using multi-speaker audio with unknown speaker boundaries
abstract
In this paper, we demonstrate a method for training speaker embedding extractors using weak annotation.More specifically, we are using the full VoxCeleb recordings and the name of the celebrities appearing on each video without knowledge of the time intervals the celebrities appear in the video.We show that by combining a baseline speaker diarization algorithm that requires no training or parameter tuning, a modified loss with aggregation over segments, and a two-stage training approach, we are able to train a competitive ResNet-based embedding extractor.Finally, we experiment with two different aggregation functions and analyze their behaviour in terms of their gradients.
Themos Stafylakis, Ladislav Mosner, Oldrich Plchot, Johan Rohdin, Anna Silnova, Lukás Burget, Jan Cernocký
INTERSPEECH4
2021 Analysis of the but Diarization System for Voxconverse Challenge
abstract
This paper describes the system developed by the BUT team for the fourth track of the VoxCeleb Speaker Recognition Challenge, focusing on diarization on the VoxConverse dataset. The system consists of signal pre-processing, voice activity detection, speaker embedding extraction, an initial agglomerative hierarchical clustering followed by diarization using a Bayesian hidden Markov model, a reclustering step based on per-speaker global embeddings and overlapped speech detection and handling. We provide comparisons for each of the steps and share the implementation of the most relevant modules of our system. Our system scored second in the challenge in terms of the primary metric (diarization error rate) and first according to the secondary metric (Jaccard error rate).
Federico Landini, Ondrej Glembek, Pavel Matejka, Johan Rohdin, Lukás Burget, Mireia Díez, Anna Silnova
ICASSP4
2021 Speaker Embeddings by Modeling Channel-Wise Correlations
abstract
Speaker embeddings extracted with deep 2D convolutional neural networks are typically modeled as projections of first and second order statistics of channel-frequency pairs onto a linear layer, using either average or attentive pooling along the time axis. In this paper we examine an alternative pooling method, where pairwise correlations between channels for given frequencies are used as statistics. The method is inspired by style-transfer methods in computer vision, where the style of an image, modeled by the matrix of channel-wise correlations, is transferred to another image, in order to produce a new image having the style of the first and the content of the second. By drawing analogies between image style and speaker characteristics, and between image content and phonetic sequence, we explore the use of such channel-wise correlations features to train a ResNet architecture in an end-to-end fashion. Our experiments on VoxCeleb demonstrate the effectiveness of the proposed pooling method in speaker recognition.
Themos Stafylakis, Johan Rohdin, Lukás Burget
Interspeech2
2020 Grading Quality of Color Retinal Images to Assist Fundus Camera Operators
abstract
Suitable image quality is a prerequisite to ensure accurate diagnosis or person recognition by color retinal images. Many factors during image acquisition, transferring and storing can result in poor quality retinal images. Poor quality images not only increase the possibility of wrong diagnosis, false acceptance, or incorrect identification but also increase diagnosis or recognition time. Therefore, retinal image quality assessment has become an important research topic. In general, only one color channel (most of the time either green or grayscale) is used to assess the quality of retinal images ignoring the quality of other channels. However, all image channels carry complementary information. In this paper, we propose a quality assessment approach for a colored retinal image to assist a fundus camera operator to judge the image quality. In our approach, we analyze the histogram of pixel intensity and uniformity of illumination, as well as check the presence of two main anatomical structures, optic disc, and central retinal blood vessels, in all color channels (i.e., red, green and blue) as well as in grayscale format. We show the effectiveness of our approach by grading 3090 color retinal images of five publicly available retinal databases.
Sangeeta Biswas, Johan Rohdin, Andrii Kavetskyi, Martin Drahanský
CBMS2
2020 Investigation of Specaugment for Deep Speaker Embedding Learning
abstract
SpecAugment is a newly proposed data augmentation method for speech recognition. By randomly masking bands in the log Mel spectogram this method leads to impressive performance improvements. In this paper, we investigate the usage of SpecAugment for speaker verification tasks. Two different models, namely 1-D convolutional TDNN and 2-D convolutional ResNet34, trained with either Softmax or AAM-Softmax loss, are used to analyze SpecAugment's effectiveness. Experiments are carried out on the Voxceleb and NIST SRE 2016 dataset. By applying SpecAugment to the original clean data in an on-the-fly manner without complex off-line data augmentation methods, we obtained 3.72% and 11.49% EER for NIST SRE 2016 Cantonese and Tagalog, respectively. For Voxceleb1 evaluation set, we obtained 1.47% EER.
Shuai Wang 0016, Johan Rohdin, Oldrich Plchot, Lukás Burget, Kai Yu 0004, Jan Cernocký
ICASSP2
2020 But System for the Second Dihard Speech Diarization Challenge
abstract
This paper describes the winning systems developed by the BUT team for the four tracks of the Second DIHARD Speech Diarization Challenge. For tracks 1 and 2 the systems were mainly based on performing agglomerative hierarchical clustering (AHC) of x-vectors, followed by another x-vector clustering based on Bayes hidden Markov model and variational Bayes inference. We provide a comparison of the improvement given by each step and share the implementation of the core of the system. For tracks 3 and 4 with recordings from the Fifth CHiME Challenge, we explored different approaches for doing multi-channel diarization and our best performance was obtained when applying AHC on the fusion of per channel probabilistic linear discriminant analysis scores.
Federico Landini, Shuai Wang 0016, Mireia Díez, Lukás Burget, Pavel Matejka, Katerina Zmolíková, Ladislav Mosner, Anna Silnova, Oldrich Plchot, Ondrej Novotný, Hossein Zeinali, Johan Rohdin
ICASSP12
2020 BUT Text-Dependent Speaker Verification System for SdSV Challenge 2020
abstract
In this paper, we present the winning BUT submission for the text-dependent task of the SdSV challenge 2020. Given the large amount of training data available in this challenge, we explore successful techniques from text-independent systems in the text-dependent scenario. In particular, we trained x-vector extractors on both in-domain and out-of-domain datasets and combine them with i-vectors trained on concatenated MFCCs and bottleneck features, which have proven effective for the text-dependent scenario. Moreover, we proposed the use of phrase-dependent PLDA backend for scoring and its combination with a simple phrase recognizer, which brings up to 63% relative improvement on our development set with respect to using standard PLDA. Finally, we combine our different i-vector and x-vector based systems using a simple linear logistic regression score level fusion, which provides 28% relative improvement on the evaluation set with respect to our best single system
Alicia Lozano-Diez, Anna Silnova, Bhargav Pulugundla, Johan Rohdin, Karel Veselý, Lukás Burget, Oldrich Plchot, Ondrej Glembek, Ondrej Novotný, Pavel Matejka
INTERSPEECH4
2020 13 years of speaker recognition research at BUT, with longitudinal analysis of NIST SRE
Pavel Matejka, Oldrich Plchot, Ondrej Glembek, Lukás Burget, Johan Rohdin, Hossein Zeinali, Ladislav Mosner, Anna Silnova, Ondrej Novotný, Mireia Díez, Jan Cernocký
Comput. Speech Lang.5
2020 End-to-end DNN based text-independent speaker recognition for long and short utterances
Johan Rohdin, Anna Silnova, Mireia Díez, Oldrich Plchot, Pavel Matejka, Lukás Burget, Ondrej Glembek
Comput. Speech Lang.1
2019 Speaker Verification with Application-Aware Beamforming
abstract
Multichannel speech processing applications usually employ beamformers as means of speech enhancement through spatial filtering. Beamformers with learnable parameters require training to minimize a loss function that is not necessarily correlated with the final objective. In this paper, we present a framework employing recent neural network based generalized eigenvalue beamformer and application-specific model that allows for optimization of beamformer w.r.t. target application. In our case, the application is speaker verification which utilizes a speaker embedding (x-vector) extractor that conveniently comes with desired loss. We show that application-specific training of the beamformer brings performance improvements over a system trained in the standard way. We perform our analysis on the recently introduced VOiCES corpus which contains multichannel data and allows us to modify the evaluation trials such that enrollment recordings remain single-channel and test utterances are multichannel.
Ladislav Mosner, Oldrich Plchot, Johan Rohdin, Lukás Burget, Jan Cernocký
ASRU3
2019 Speaker Verification Using End-to-end Adversarial Language Adaptation
abstract
In this paper we investigate the use of adversarial domain adaptation for addressing the problem of language mismatch between speaker recognition corpora. In the context of speaker verification, adversarial domain adaptation methods aim at minimizing certain divergences between the distribution that the utterance-level features follow (i.e. speaker embeddings) when drawn from source and target domains (i.e. languages), while preserving their capacity in recognizing speakers. Neural architectures for extracting utterance-level representations enable us to apply adversarial adaptation methods in an end-to-end fashion and train the network jointly with the standard cross-entropy loss. We examine several configurations, such as the use of (pseudo-)labels on the target domain as well as domain labels in the feature extractor, and we demonstrate the effectiveness of our method on the challenging NIST SRE16 and SRE18 benchmarks.
Johan Rohdin, Themos Stafylakis, Anna Silnova, Hossein Zeinali, Lukás Burget, Oldrich Plchot
ICASSP1
2019 How to Improve Your Speaker Embeddings Extractor in Generic Toolkits
abstract
Recently, speaker embeddings extracted with deep neural networks became the state-of-the-art method for speaker verification. In this paper we aim to facilitate its implementation on a more generic toolkit than Kaldi, which we anticipate to enable further improvements on the method. We examine several tricks in training, such as the effects of normalizing input features and pooled statistics, different methods for preventing overfitting as well as alternative non-linearities that can be used instead of Rectifier Linear Units. In addition, we investigate the difference in performance between TDNN and CNN, and between two types of attention mechanism. Experimental results on Speaker in the Wild, SRE 2016 and SRE 2018 datasets demonstrate the effectiveness of the proposed implementation.
Hossein Zeinali, Lukás Burget, Johan Rohdin, Themos Stafylakis, Jan Cernocký
ICASSP3
2019 Bayesian HMM Based x-Vector Clustering for Speaker Diarization
Mireia Díez, Lukás Burget, Shuai Wang 0016, Johan Rohdin, Jan Cernocký
INTERSPEECH4
2019 Self-Supervised Speaker Embeddings
abstract
Contrary to i-vectors, speaker embeddings such as x-vectors are incapable of leveraging unlabelled utterances, due to the classification loss over training speakers. In this paper, we explore an alternative training strategy to enable the use of unlabelled utterances in training. We propose to train speaker embedding extractors via reconstructing the frames of a target speech segment, given the inferred embedding of another speech segment of the same utterance. We do this by attaching to the standard speaker embedding extractor a decoder network, which we feed not merely with the speaker embedding, but also with the estimated phone sequence of the target frame sequence. The reconstruction loss can be used either as a single objective, or be combined with the standard speaker classification loss. In the latter case, it acts as a regularizer, encouraging generalizability to speakers unseen during training. In all cases, the proposed architectures are trained from scratch and in an end-to-end fashion. We demonstrate the benefits from the proposed approach on VoxCeleb and Speakers in the wild, and we report notable improvements over the baseline.
Themos Stafylakis, Johan Rohdin, Oldrich Plchot, Petr Mizera, Lukás Burget
INTERSPEECH2
2019 On the Usage of Phonetic Information for Text-Independent Speaker Embedding Extraction
Shuai Wang 0016, Johan Rohdin, Lukás Burget, Oldrich Plchot, Yanmin Qian, Kai Yu 0004, Jan Cernocký
INTERSPEECH2
2019 Detecting Spoofing Attacks Using VGG and SincNet: BUT-Omilia Submission to ASVspoof 2019 Challenge
abstract
In this paper, we present the system description of the joint efforts of Brno University of Technology (BUT) and Omilia -- Conversational Intelligence for the ASVSpoof2019 Spoofing and Countermeasures Challenge. The primary submission for Physical access (PA) is a fusion of two VGG networks, trained on single and two-channels features. For Logical access (LA), our primary system is a fusion of VGG and the recently introduced SincNet architecture. The results on PA show that the proposed networks yield very competitive performance in all conditions and achieved 86\:\% relative improvement compared to the official baseline. On the other hand, the results on LA showed that although the proposed architecture and training strategy performs very well on certain spoofing attacks, it fails to generalize to certain attacks that are unseen during training.
Hossein Zeinali, Themos Stafylakis, Georgia Athanasopoulou, Johan Rohdin, Ioannis Gkinis, Lukás Burget, Jan Cernocký
INTERSPEECH4
2018 End-to-End DNN Based Speaker Recognition Inspired by I-Vector and PLDA
abstract
Recently, several end-to-end speaker verification systems based on deep neural networks (DNNs) have been proposed. These systems have been proven to be competitive for text-dependent tasks as well as for text-independent tasks with short utterances. However, for text-independent tasks with longer utterances, end-to-end systems are still outperformed by standard i-vector + PLDA systems. In this work, we develop an end-to-end speaker verification system that is initialized to mimic an i-vector + PLDA baseline. The system is then further trained in an end-to-end manner but regularized so that it does not deviate too far from the initial system. In this way we mitigate overfitting which normally limits the performance of end-to-end systems. The proposed system outperforms the i-vector + PLDA baseline on both long and short duration utterances.
Johan Rohdin, Anna Silnova, Mireia Díez, Oldrich Plchot, Pavel Matejka, Lukás Burget
ICASSP1
2018 BUT System for DIHARD Speech Diarization Challenge 2018
Mireia Díez, Federico Landini, Lukás Burget, Johan Rohdin, Anna Silnova, Katerina Zmolíková, Ondrej Novotný, Karel Veselý, Ondrej Glembek, Oldrich Plchot, Ladislav Mosner, Pavel Matejka
INTERSPEECH4
2017 Analysis and Description of ABC Submission to NIST SRE 2016
Oldrich Plchot, Pavel Matejka, Anna Silnova, Ondrej Novotný, Mireia Díez, Johan Rohdin, Ondrej Glembek, Niko Brümmer, Albert Swart, Jesús Jorrín-Prieto, L. Paola García-Perera, Luis Buera, Patrick Kenny, Jahangir Alam 0001, Gautam Bhattacharya
INTERSPEECH6
2016 Robust discriminative training against data insufficiency in PLDA-based speaker verification
Johan Rohdin, Sangeeta Biswas, Koichi Shinoda
Comput. Speech Lang.1
2015 Autonomous selection of i-vectors for PLDA modelling in speaker verification
Sangeeta Biswas, Johan Rohdin, Koichi Shinoda
Speech Commun.2
2014 Constrained discriminative PLDA training for speaker verification
abstract
Many studies have proven the effectiveness of discriminative training for speaker verification based on probabilistic linear discriminative analysis (PLDA) with i-vectors as features. Most of them directly optimize the log-likelihood ratio score function of the PLDA model instead of explicitly train the PLDA model. But this optimization process removes some of the constraints that normally are imposed on the PLDA log likelihood ratio score function. This may deteriorate the verification performance when the amount of training data is limited. In this paper, we first show two constraints which the score function should follow, and then we propose a new constrained discriminative training algorithm which keeps these constraints. Our experiments show that our method obtained significant improvements in the verification performance in the male trials of the telephone speaker verification tasks of NIST SRE08 and SRE10.
Johan Rohdin, Sangeeta Biswas, Koichi Shinoda
ICASSP1