Marvin Borsdorf

dblp:313/1682 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0003-1769-1621ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021
YearPublicationVenuePosition
2025 Speech Separation for Low-Resource Languages
abstract
Speech separation aims to equip machines with the human ability of selective listening, i.e. to focus attention on specific information in spoken communication. Studies have shown that the language spoken in a cocktail party scenario matters. While the development of speech separation models can leverage extensive databases, for the majority of languages only very limited data is available. This work presents the very first study on speech separation for low-resource languages. We choose blind source separation as the task to be studied and analyze three strategies to overcome the data scarcity of two low-resource languages from the GlobalPhoneMS2 database. We show that data from other languages can be used to develop models that work for low-resource languages. Finetuning additionally boosts the performance, and training on multiple languages increases both performance and robustness. We show that dynamic mixing in the development helps to find a trade-off between performance and development time.
Marvin Borsdorf, Zexu Pan, Pascal Himmelmann, Haizhou Li 0001, Tanja Schultz
ICASSP1
2024 wTIMIT2mix: A Cocktail Party Mixtures Database to Study Target Speaker Extraction for Normal and Whispered Speech
Marvin Borsdorf, Zexu Pan, Haizhou Li 0001, Tanja Schultz
INTERSPEECH1
2024 Does the Lombard Effect Matter in Speech Separation? Introducing the Lombard-GRID-2mix Dataset
Iva Ewert, Marvin Borsdorf, Haizhou Li 0001, Tanja Schultz
INTERSPEECH2
2024 NeuroHeed: Neuro-Steered Speaker Extraction Using EEG Signals
abstract
Humans possess the remarkable ability to selectively attend to a single speaker amidst competing voices and background noise, known asselective auditory attention. Recent studies in auditory neuroscience indicate a strong correlation between the attended speech signal and the corresponding brain's elicited neuronal activities. In this work, we study such brain activities measured using affordable and non-intrusive electroencephalography (EEG) devices. We present NeuroHeed, a speaker extraction model that leverages the listener's synchronized EEG signals to extract the attended speech signal in a cocktail party scenario, in which the extraction process is conditioned on a neuronal attractor encoded from the EEG signal. We propose both an offline and an online NeuroHeed, with the latter designed for real-time inference. In the online NeuroHeed, we additionally propose an autoregressive speaker encoder, which accumulates past extracted speech signals for self-enrollment of the attended speaker information into an auditory attractor, that retains the attentional momentum over time. Online NeuroHeed extracts the current window of the speech signals with guidance from both attractors. Experimental results on KUL dataset two-speaker scenario demonstrate that NeuroHeed effectively extracts brain-attended speech signals with an average scale-invariant signal-to-noise ratio improvement (SI-SDRi) of 14.3 dB and extraction accuracy of 90.8% in offline settings, and SI-SDRi of 11.2 dB and extraction accuracy of 85.1% in online settings.
Zexu Pan, Marvin Borsdorf, Siqi Cai 0002, Tanja Schultz, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Multi-Head Attention and GRU for Improved Match-Mismatch Classification of Speech Stimulus and EEG Response
abstract
This work is based on the participation by the HyperAttention team in the Auditory EEG Decoding Challenge, 2023 (ICASSP 2023 Signal Processing Grand Challenge) task 1, which deals with the match-mismatch classification of speech stimuli and EEG responses of human listeners. We demonstrate the benefits of using mel-spectrograms instead of speech envelopes as input features as well as the effectiveness of Multi-Head Attention and GRU for EEG and speech processing. With a total score of 79.05 %, we reach the second place in the challenge.
Marvin Borsdorf, Saurav Pahuja, Gabriel Ivucic, Siqi Cai 0002, Haizhou Li 0001, Tanja Schultz
ICASSP1
2023 ImagineNet: Target Speaker Extraction with Intermittent Visual Cue Through Embedding Inpainting
abstract
The speaker extraction technique seeks to single out the voice of a target speaker from the interfering voices in a speech mixture. Typically an auxiliary reference of the target speaker is used to form voluntary attention. Either a pre-recorded utterance or a synchronized lip movement in a video clip can serve as the auxiliary reference. The use of visual cue is not only feasible, but also effective due to its noise robustness, and becoming popular. However, it is difficult to guarantee that such parallel visual cue is always available in real-world applications where visual occlusion or intermittent communication can occur. In this paper, we study the audio-visual speaker extraction algorithms with intermittent visual cue. We propose a joint speaker extraction and visual embedding inpainting framework to explore the mutual benefits. To encourage the interaction between the two tasks, they are performed alternately with an interlacing structure and optimized jointly. We also propose two types of visual inpainting losses and study our proposed method with two types of popularly used visual embeddings. The experimental results show that we outperform the baseline in terms of signal quality, perceptual quality, and intelligibility.
Zexu Pan, Wupeng Wang, Marvin Borsdorf, Haizhou Li 0001
ICASSP3
2023 Speaker Extraction with Detection of Presence and Absence of Target Speakers
Marvin Borsdorf, Zexu Pan, Haizhou Li 0001, Yangjie Wei
INTERSPEECH2
2022 Experts Versus All-Rounders: Target Language Extraction for Multiple Target Languages
abstract
Target language extraction (TLE) is a novel task in the field of selective auditory attention, which seeks to extract all speech signals that are spoken in a target language from other sources in a multilingual cocktail party. In our prior studies, a TLE model was trained to extract a predefined, single target language, referred to as Single-TLE. In this paper, we extend the Single-TLE framework to Multi-TLE. Multi-TLE models can also extract all speech signals of one specific target language, but they are optimized on a set of multiple target languages during training. As such, they learn the characteristics of several target languages and can replace multiple Single-TLE models without retraining. We perform experiments on the GlobalPhoneMCP database and incorporate a dynamic language mixing scheme for training. The Multi-TLE model does not only outperform Single-TLE models, but when given a language ID as additional input, it is also able to extract the speech of a specific target language from a mixture which contains multiple learned target languages.
Marvin Borsdorf, Kevin Scheck, Haizhou Li 0001, Tanja Schultz
ICASSP1
2022 Blind Language Separation: Disentangling Multilingual Cocktail Party Voices by Language
Marvin Borsdorf, Kevin Scheck, Haizhou Li 0001, Tanja Schultz
INTERSPEECH1
2021 Target Language Extraction at Multilingual Cocktail Parties
abstract
Typically, target speaker extraction seeks to extract a target speaker's contribution according to his or her individual voice characteristics. In a “multilingual cocktail party” however, listeners may desire to extract speaker contributions spoken in a particular language, regard-less of the number of contributing speakers. In this paper, we pro-pose a novel task called “target language extraction” (TLE) which extracts voices based on the spoken language rather than on individ-ual speaker characteristics. We introduce a new database for TLE which simulates the multilingual cocktail party problem in mixtures of two and four speakers with German as the target language. The database is derived from the GlobalPhone 2000 Speaker Package and is called “GlobalPhone Multilingual Cocktail Party - German” (GlobalPhoneMCP-GE). Our experimental results show that our approach to TLE achieves very good performance regardless of the number of speakers in the mixture and that TLE generalizes well to unseen speakers and interfering languages. This work represents the first attempt at target language extraction.
Marvin Borsdorf, Haizhou Li 0001, Tanja Schultz
ASRU1
2021 Universal Speaker Extraction in the Presence and Absence of Target Speakers for Speech of One and Two Talkers
Marvin Borsdorf, Chenglin Xu, Haizhou Li 0001, Tanja Schultz
Interspeech1
2021 GlobalPhone Mix-To-Separate Out of 2: A Multilingual 2000 Speakers Mixtures Database for Speech Separation
Marvin Borsdorf, Chenglin Xu, Haizhou Li 0001, Tanja Schultz
Interspeech1