VLDB 2026 Research / reviewers in the wild / expert
Kevin Scheck
dblp:277/3415
· DBLP profile ↗
8ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0001-4679-9239ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiffMV-ETS: Diffusion-based Multi-Voice Electromyography-to-Speech Conversion using Speaker-Independent Speech Training TargetsabstractElectromyography (EMG) signals have been investigated for novel voice prostheses to enable speech communication with silent articulation.In this work, we propose DiffMV-ETS, a multi-voice, diffusion-based EMG-to-speech system that converts EMG signals to speech in selectable voices.We evaluate it for scenarios where no speech of the speaker wearing EMG sensors is used for training.For this purpose, we introduce EMG-VCTK, a dataset containing EMG and audio recordings of sentences from the Voice Conversion Tool Kit corpus.We compare EMG models trained with audio of the same speaker, of auxiliary speakers, and of text-to-speech systems.Experiments indicate that models retain their intelligibility and naturalness when trained with synthetic speech.DiffMV-ETS enhances the speech naturalness and similarity to unseen voices.To the best of our knowledge, this is the first work to train multi-voice EMG-to-speech systems with speaker-independent targets. Kevin Scheck, Tom Dombeck, Zhao Ren, Peter Wu, Michael Wand 0002, Tanja Schultz |
INTERSPEECH | 1 |
| 2023 | Multi-Speaker Speech Synthesis from Electromyographic Signals by Soft Speech Unit PredictionabstractElectromyographic (EMG) signals of articulatory muscles reflect the speech production process even if the user is speaking silently i.e. moving the articulators without producing audible sound. We propose Speech-Unit-based EMG-to-Speech (SU-E2S), a system which relies on EMG to synthesize speech which contains the articulated content but is vocalized in another voice, determined by an acoustic reference utterance. It is based on a Voice Conversion (VC) system which decomposes acoustic speech into continuous soft speech units and a speaker embedding and then reconstructs acoustic features. SU-E2S performs speech synthesis by predicting soft speech units from EMG and using them as input to the VC system. Experiments show that the SU-E2S output is on par in terms of intelligibility of predicting acoustic features directly from EMG, but adds the functionality of synthesizing speech in other voices. Kevin Scheck, Tanja Schultz |
ICASSP | 1 |
| 2023 | STE-GAN: Speech-to-Electromyography Signal Conversion using Generative Adversarial Networks
Kevin Scheck, Tanja Schultz |
INTERSPEECH | 1 |
| 2022 | Experts Versus All-Rounders: Target Language Extraction for Multiple Target LanguagesabstractTarget language extraction (TLE) is a novel task in the field of selective auditory attention, which seeks to extract all speech signals that are spoken in a target language from other sources in a multilingual cocktail party. In our prior studies, a TLE model was trained to extract a predefined, single target language, referred to as Single-TLE. In this paper, we extend the Single-TLE framework to Multi-TLE. Multi-TLE models can also extract all speech signals of one specific target language, but they are optimized on a set of multiple target languages during training. As such, they learn the characteristics of several target languages and can replace multiple Single-TLE models without retraining. We perform experiments on the GlobalPhoneMCP database and incorporate a dynamic language mixing scheme for training. The Multi-TLE model does not only outperform Single-TLE models, but when given a language ID as additional input, it is also able to extract the speech of a specific target language from a mixture which contains multiple learned target languages. Marvin Borsdorf, Kevin Scheck, Haizhou Li 0001, Tanja Schultz |
ICASSP | 2 |
| 2022 | Blind Language Separation: Disentangling Multilingual Cocktail Party Voices by Language
Marvin Borsdorf, Kevin Scheck, Haizhou Li 0001, Tanja Schultz |
INTERSPEECH | 2 |
| 2022 | SmartHelm: User Studies from Lab to Field for Attention ModelingabstractWe present three user studies that gradually prepare our prototype system SmartHelm for use in the field, i.e. supporting cargo cyclists on public roads for cargo delivery. SmartHelm is an attention-sensitive smart helmet that integrates none-invasive brain and eye activity detection with hands-free Augmented Reality (AR) components in a speech-enabled outdoor assistance system. The described studies systematically increased in ecological validity from lab to field. The first study consisted of an Augmented Reality preparation examination in the lab. The second study then investigated simulated attention distraction modeling, whereas the third study examined real-world attention distraction modeling while cycling in traffic. During these three studies, multimodal data (EEG, eye-tracking, video, GPS and speech) has been collected synchronously and analyzed in offline and online experiments. Machine Learning models were trained and optimized for attention modeling.Results: Analyses of self-report and objective data during the simulation study show the plausibility of the simulated internal and external distractions. The analysis of behavioral data captured by multimodal biosignals recorded in the field study further shows that real visual attention distractions can be automatically identified using synchronized video and eye-tracking data. Machine Learning methods based on long short-term memory models (LSTMs) indicate that simulated attention distractions can be automatically detected from EEG data, with the best detection performance for mental distractions. Finally, the self-report data suggest that the comfort of the SmartHelm helmet should be further improved for permanent use in road traffic. Mazen Salous, Dennis Küster, Kevin Scheck, Aytac Dikfidan, Tim Neumann, Felix Putze, Tanja Schultz |
SMC | 3 |
| 2020 | Toward Silent Paralinguistics: Speech-to-EMG - Retrieving Articulatory Muscle Activity from SpeechabstractElectromyographic (EMG) signals recorded during speech production encode information on articulatory muscle activity and also on the facial expression of emotion, thus representing a speech-related biosignal with strong potential for paralinguistic applications.In this work, we estimate the electrical activity of the muscles responsible for speech articulation directly from the speech signal.To this end, we first perform a neural conversion of speech features into electromyographic time domain features, and then attempt to retrieve the original EMG signal from the time domain features.We propose a feed forward neural network to address the first step of the problem (speech features to EMG features) and a neural network composed of a convolutional block and a bidirectional long short-term memory block to address the second problem (true EMG features to EMG signal).We observe that four out of the five originally proposed time domain features can be estimated reasonably well from the speech signal.Further, the five time domain features are able to predict the original speech-related EMG signal with a concordance correlation coefficient of 0.663.We further compare our results with the ones achieved on the inverse problem of generating acoustic speech features from EMG features. Catarina Botelho, Lorenz Diener, Dennis Küster, Kevin Scheck, Shahin Amiriparian, Björn W. Schuller, Tanja Schultz, Alberto Abad, Isabel Trancoso |
INTERSPEECH | 4 |
| 2020 | Towards Silent Paralinguistics: Deriving Speaking Mode and Speaker ID from Electromyographic SignalsabstractSilent Computational Paralinguistics (SCP) -the assessment of speaker states and traits from non-audibly spoken communication -has rarely been targeted in the rich body of either Computational Paralinguistics or Silent Speech Processing.Here, we provide first steps towards this challenging but potentially highly rewarding endeavour: Paralinguistics can enrich spoken language interfaces, while Silent Speech Processing enables confidential and unobtrusive spoken communication for everybody, including mute speakers.We approach SCP by using speech-related biosignals stemming from facial muscle activities captured by surface electromyography (EMG).To demonstrate the feasibility of SCP, we select one speaker trait (speaker identity) and one speaker state (speaking mode).We introduce two promising strategies for SCP: (1) deriving paralinguistic speaker information directly from EMG of silently produced speech versus (2) first converting EMG into an audible speech signal followed by conventional computational paralinguistic methods.We compare traditional feature extraction and decision making approaches to more recent deep representation and transfer learning by convolutional and recurrent neural networks, using openly available EMG data.We find that paralinguistics can be assessed not only from acoustic speech but also from silent speech captured by EMG. Lorenz Diener, Shahin Amiriparian, Catarina Botelho, Kevin Scheck, Dennis Küster, Isabel Trancoso, Björn W. Schuller, Tanja Schultz |
INTERSPEECH | 4 |