Tobias Cord-Landwehr

dblp:296/4704 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2025
0009-0004-1158-6235ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models
abstract
We propose an approach for simultaneous diarization and separation of meeting data. It consists of a complex Angular Central Gaussian Mixture Model (cACGMM) for speech source separation, and a von-Mises-Fisher Mixture Model (vMFMM) for diarization in a joint statistical framework. Through the integration, both spatial and spectral information are exploited for diarization and separation. We also develop a method for counting the number of active speakers in a segment of a meeting to support block-wise processing. While the total number of speakers in a meeting may be known, it is usually not known on a per-segment level. With the proposed speaker counting, joint diarization and source separation can be done segment-by-segment, and the permutation problem across segments is solved, thus allowing for block-online processing in the future. Experimental results on the LibriCSS meeting corpus show that the integrated approach outperforms a cascaded approach of diarization and speech enhancement in terms of WER, both on a per-segment and on a per-meeting level.
Tobias Cord-Landwehr, Christoph Böddeker, Reinhold Häb-Umbach
ICASSP1
2025 Spatio-Spectral Diarization of Meetings by Combining TDOA-based Segmentation and Speaker Embedding-based Clustering
abstract
We propose a spatio-spectral, combined model-based and data-driven diarization pipeline consisting of TDOA-based segmentation followed by embedding-based clustering. The proposed system requires neither access to multi-channel training data nor prior knowledge about the number or placement of microphones. It works for both a compact microphone array and distributed microphones, with minor adjustments. Due to its superior handling of overlapping speech during segmentation, the proposed pipeline significantly outperforms the single-channel pyannote approach, both in a scenario with a compact microphone array and in a setup with distributed microphones. Additionally, we show that, unlike fully spatial diarization pipelines, the proposed system can correctly track speakers when they change positions.
Tobias Cord-Landwehr, Tobias Gburrek, Marc Deegen, Reinhold Häb-Umbach
INTERSPEECH1
2024 Geodesic Interpolation of Frame-Wise Speaker Embeddings for the Diarization of Meeting Scenarios
abstract
We propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partially overlapping speech. To this end, a geodesic distance loss is used that enforces the embeddings computed from regions with two active speakers to lie on the shortest path on a sphere between the points given by the d-vectors of each of the active speakers. Using those frame-wise speaker embeddings in clustering-based diarization outperforms segment-level clustering-based diarization systems such as VBx and Spectral Clustering. By extending our approach to a mixture-model-based diarization, the performance can be further improved, approaching the diarization error rates of diarization systems that use a dedicated overlap detection, and outperforming these systems when also employing an additional overlap detection.
Tobias Cord-Landwehr, Christoph Böddeker, Catalin Zorila, Rama Sanand Doddipatla, Reinhold Häb-Umbach
ICASSP1
2024 Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment
abstract
Diarization is a crucial component in meeting transcription systems to ease the challenges of speech enhancement and attribute the transcriptions to the correct speaker.Particularly in the presence of overlapping or noisy speech, these systems have problems reliably assigning the correct speaker labels, leading to a significant amount of speaker confusion errors.We propose to add segment-level speaker reassignment to address this issue.By revisiting, after speech enhancement, the speaker attribution for each segment, speaker confusion errors from the initial diarization stage are significantly reduced.Through experiments across different system configurations and datasets, we further demonstrate the effectiveness and applicability in various domains.Our results show that segment-level speaker reassignment successfully rectifies at least 40% of speaker confusion word errors, highlighting its potential for enhancing diarization accuracy in meeting transcription systems.
Christoph Böddeker, Tobias Cord-Landwehr, Reinhold Häb-Umbach
INTERSPEECH2
2023 Frame-Wise and Overlap-Robust Speaker Embeddings for Meeting Diarization
abstract
Using a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact that the student produces sensible speaker embeddings even for segments with speech overlap, the frame-wise embeddings serve as an appropriate representation of the input speech signal for an end-to-end neural meeting diarization (EEND) system. We show in experiments that this representation helps mitigate a well-known problem of EEND systems: when increasing the number of speakers the diarization performance drop is significantly reduced. We also introduce block-wise processing to be able to diarize arbitrarily long meetings.
Tobias Cord-Landwehr, Christoph Böddeker, Catalin Zorila, Rama Sanand Doddipatla, Reinhold Häb-Umbach
ICASSP1
2023 A Teacher-Student Approach for Extracting Informative Speaker Embeddings From Speech Mixtures
abstract
We introduce a monaural neural speaker embeddings extractor that computes an embedding for each speaker present in a speech mixture.To allow for supervised training, a teacherstudent approach is employed: the teacher computes the target embeddings from each speaker's utterance before the utterances are added to form the mixture, and the student embedding extractor is then tasked to reproduce those embeddings from the speech mixture at its input.The system much more reliably verifies the presence or absence of a given speaker in a mixture than a conventional speaker embedding extractor, and even exhibits comparable performance to a multi-channel approach that exploits spatial information for embedding extraction.Further, it is shown that a speaker embedding computed from a mixture can be used to check for the presence of that speaker in another mixture.
Tobias Cord-Landwehr, Christoph Böddeker, Catalin Zorila, Rama Sanand Doddipatla, Reinhold Häb-Umbach
INTERSPEECH1
2022 An Initialization Scheme for Meeting Separation with Spatial Mixture Models
Christoph Böddeker, Tobias Cord-Landwehr, Thilo von Neumann, Reinhold Häb-Umbach
INTERSPEECH2
2021 Contrastive Predictive Coding Supported Factorized Variational Autoencoder For Unsupervised Learning Of Disentangled Speech Representations
abstract
In this work we address disentanglement of style and content in speech signals. We propose a fully convolutional variational autoencoder employing two encoders: a content encoder and a style encoder. To foster disentanglement, we propose adversarial contrastive predictive coding. This new disentanglement method does neither need parallel data nor any supervision. We show that the proposed technique is capable of separating speaker and content traits into the two different representations and show competitive speaker-content disentanglement performance compared to other unsupervised approaches. We further demonstrate an increased robustness of the content representation against a train-test mismatch compared to spectral features, when used for phone recognition.
Janek Ebbers, Michael Kuhlmann, Tobias Cord-Landwehr, Reinhold Häb-Umbach
ICASSP3