Michael Wand 0002

dblp:85/72-2 · DBLP profile ↗
← Back
31ranked-venue papers
13as first author
5since 2021 · last 2026
0000-0003-0966-7824ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 13 first-author · 1 since 2021Artificial intelligence and machine learning · 20 · 9 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Theory of computation · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Interpreting Logical Explanations of Classifying Neural Networks
Fabrizio Leopardi, Faezeh Labbaf, Tomás Kolárik, Michael Wand 0002, Natasha Sharygina
ESANN4
2026 Formally Explaining Neural Network Classification
abstract
Abstract Neural networks (NNs) are the core of AI-based technologies. However, the degree of reliability in performing the task is an open problem. The explainability of a central task of NNs, classification, is of immense importance. While at the rise of AI-based reasoning, explainability of the NN classification has mostly been done using statistical methods, nowadays, a more reliable trend of formal logic-based methods is gaining popularity. The advantage of the formal approach is that it gives strict and provable guarantees of the classification. Formal methods is a mature field that has delivered a number of efficient computational solutions already applied in the analysis of software and hardware systems. Formal explainability methods naturally have the ability to reuse existing techniques and tools for a newly emerging field of formal explainability of NN classification. This paper surveys existing efforts to compute explanations of neural network classification based on logical abductive reasoning. The abduction approach is crucial for generalizing the results, capturing the underlying behavior of the classifier. We present the existing techniques as instances of a general formalization that allows contrasting them against each other. In addition, we discuss the issue of the quality of explanations, focusing on their key metrics and factors. As an illustrative example, the paper also presents a practical framework, SpEXplAIn , which automatically computes Space Explanations, the most general abduction-based explanations for classifying NNs with provable guarantees of the behavior of the network in continuous areas of the input feature space. The tool leverages an SMT solver compatible with a range of flexible Craig interpolation algorithms and unsatisfiable core generation, and is applicable to a wide range of applications.
Tomás Kolárik, Grigory Fedyukovich, Faezeh Labbaf, Fabrizio Leopardi, Natasha Sharygina, Michael Wand 0002
FM (2)6
2025 Space Explanations of Neural Network Classification
abstract
Abstract We present a novel logic-based concept called Space Explanations for classifying neural networks that gives provable guarantees of the behavior of the network in continuous areas of the input feature space. To automatically generate space explanations, we leverage a range of flexible Craig interpolation algorithms and unsatisfiable core generation. Based on real-life case studies, ranging from small to medium to large size, we demonstrate that the generated explanations are more meaningful than those computed by state-of-the-art.
Faezeh Labbaf, Tomás Kolárik, Martin Blicha, Grigory Fedyukovich, Michael Wand 0002, Natasha Sharygina
CAV (3)5
2025 DiffMV-ETS: Diffusion-based Multi-Voice Electromyography-to-Speech Conversion using Speaker-Independent Speech Training Targets
abstract
Electromyography (EMG) signals have been investigated for novel voice prostheses to enable speech communication with silent articulation.In this work, we propose DiffMV-ETS, a multi-voice, diffusion-based EMG-to-speech system that converts EMG signals to speech in selectable voices.We evaluate it for scenarios where no speech of the speaker wearing EMG sensors is used for training.For this purpose, we introduce EMG-VCTK, a dataset containing EMG and audio recordings of sentences from the Voice Conversion Tool Kit corpus.We compare EMG models trained with audio of the same speaker, of auxiliary speakers, and of text-to-speech systems.Experiments indicate that models retain their intelligibility and naturalness when trained with synthetic speech.DiffMV-ETS enhances the speech naturalness and similarity to unseen voices.To the best of our knowledge, this is the first work to train multi-voice EMG-to-speech systems with speaker-independent targets.
Kevin Scheck, Tom Dombeck, Zhao Ren, Peter Wu, Michael Wand 0002, Tanja Schultz
INTERSPEECH5
2021 Improving Stateful Premise Selection with Transformers
Krsto Prorokovic, Michael Wand 0002, Jürgen Schmidhuber
CICM2
2020 Motion Dynamics Improve Speaker-Independent Lipreading
abstract
We present a novel lipreading system that improves on the task of speaker-independent word recognition by decoupling motion and content dynamics. We achieve this by implementing a deep learning architecture that uses two distinct pipelines to process motion and content and subsequently merges them, implementing an end-to-end trainable system that performs fusion of independently learned representations. We obtain a average relative word accuracy improvement of ≈6.8% on unseen speakers and of ≈3.3% on known speakers, with respect to a baseline which uses a standard architecture.
Matteo Riva, Michael Wand 0002, Jürgen Schmidhuber
ICASSP2
2020 Fusion Architectures for Word-Based Audiovisual Speech Recognition
Michael Wand 0002, Jürgen Schmidhuber
INTERSPEECH1
2018 Investigations on End- to-End Audiovisual Fusion
abstract
Audiovisual speech recognition (AVSR) is a method to alleviate the adverse effect of noise in the acoustic signal. Leveraging recent developments in deep neural network-based speech recognition, we present an AVSR neural network architecture which is trained end-to-end, without the need to separately model the process of decision fusion as in conventional (e.g. HMM-based) systems. The fusion system outperforms single-modality recognition under all noise conditions. Investigation of the saliency of the input features shows that the neural network automatically adapts to different noise levels in the acoustic signal.
Michael Wand 0002, Jürgen Schmidhuber, Ngoc Thang Vu
ICASSP1
2018 Domain-Adversarial Training for Session Independent EMG-based Speech Recognition
Michael Wand 0002, Tanja Schultz, Jürgen Schmidhuber
INTERSPEECH1
2017 Improving Speaker-Independent Lipreading with Domain-Adversarial Training
abstract
We present a Lipreading system, i.e. a speech recognition system using only visual features, which uses domain-adversarial training for speaker independence. Domain-adversarial training is integrated into the optimization of a lipreader based on a stack of feedforward and LSTM (Long Short-Term Memory) recurrent neural networks, yielding an end-to-end trainable system which only requires a very small number of frames of untranscribed target data to substantially improve the recognition accuracy on the target speaker. On pairs of different source and target speakers, we achieve a relative accuracy improvement of around 40% with only 15 to 20 seconds of untranscribed target speech data. On multi-speaker training setups, the accuracy improvements are smaller but still substantial.
Michael Wand 0002, Jürgen Schmidhuber
INTERSPEECH1
2017 Biosignal-Based Spoken Communication: A Survey
abstract
Speech is a complex process involving a wide range of biosignals, including but not limited to acoustics. These biosignals-stemming from the articulators, the articulator muscle activities, the neural pathways, and the brain itself-can be used to circumvent limitations of conventional speech processing in particular, and to gain insights into the process of speech production in general. Research on biosignal-based speech processing is a wide and very active field at the intersection of various disciplines, ranging from engineering, computer science, electronics and machine learning to medicine, neuroscience, physiology, and psychology. Consequently, a variety of methods and approaches have been used to investigate the common goal of creating biosignal-based speech processing devices for communication applications in everyday situations and for speech rehabilitation, as well as gaining a deeper understanding of spoken communication. This paper gives an overview of the various modalities, research approaches, and objectives for biosignal-based spoken communication.
Tanja Schultz, Michael Wand 0002, Thomas Hueber, Dean J. Krusienski, Christian Herff, Jonathan S. Brumberg
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Lipreading with long short-term memory
abstract
Lipreading, i.e. speech recognition from visual-only recordings of a speaker's face, can be achieved with a processing pipeline based solely on neural networks, yielding significantly better accuracy than conventional methods. Feedforward and recurrent neural network layers (namely Long Short-Term Memory; LSTM) are stacked to form a single structure which is trained by back-propagating error gradients through all the layers. The performance of such a stacked network was experimentally evaluated and compared to a standard Support Vector Machine classifier using conventional computer vision features (Eigenlips and Histograms of Oriented Gradients). The evaluation was performed on data from 19 speakers of the publicly available GRID corpus. With 51 different words to classify, we report a best word accuracy on held-out evaluation speakers of 79.6% using the end-to-end neural network-based solution (11.6% improvement over the best feature-based solution evaluated).
Michael Wand 0002, Jan Koutník, Jürgen Schmidhuber
ICASSP1
2016 Deep Neural Network Frontend for Continuous EMG-Based Speech Recognition
Michael Wand 0002, Jürgen Schmidhuber
INTERSPEECH1
2015 Biosignal-based spoken communication: welcome and introduction
Matthias Janke, Michael Wand 0002
INTERSPEECH2
2015 Biosignal-based spoken communication: panel and discussion
Matthias Janke, Michael Wand 0002
INTERSPEECH2
2014 Fundamental frequency generation for whisper-to-audible speech conversion
abstract
In this work, we address the issues involved in whisper-to-audible speech conversion. Spectral mapping techniques using Gaussian mixture models or Artificial Neural Networks borrowed from voice conversion have been applied to transform whisper spectral features to normally phonated audible speech. However, the modeling and generation of fundamental frequency (F0) and its contour in the converted speech is a major issue. Whispered speech does not contain explicit voicing characteristics and hence it is hard to derive a suitable F0, making it difficult to generate a natural prosody after conversion. Our work addresses the F0 modeling in whisper-to-speech conversion. We show that F0 contours can be derived from the mapped spectral vectors, which can be used for the synthesis of a speech signal. We also present a hybrid unit selection approach for whisper-to-speech conversion. Unit selection is performed on the spectral vectors, where F0 and its contour can be obtained as a byproduct without any additional modeling.
Matthias Janke, Michael Wand 0002, Till Heistermann, Tanja Schultz, K. Prahallad
ICASSP2
2014 Compensation of recording position shifts for a myoelectric Silent Speech Recognizer
abstract
A myoelectric Silent Speech Recognizer is a system which recognizes speech by capturing the electrical activity of the human articulatory muscles, thus enabling the user to communicate silently. We recently devised a recording setup based on electrode arrays with multiple measuring points. In this study we show that this allows to compensate for shifts of the recording position, which happen when the array is removed and reattached between system training and application. We present a method which determines the amount of recording position shift; compensation is performed by linear interpolation. We evaluate our method by running recognition experiments across recording sessions and obtain a Word Error Rate improvement of 14.3% relative on the development set and 12.9% relative on the evaluation set, compared to using classical session adaptation.
Michael Wand 0002, Christopher Schulte, Matthias Janke, Tanja Schultz
ICASSP1
2014 BioKIT - real-time decoder for biosignal processing
abstract
We introduce BioKIT, a new Hidden Markov Model based toolkit to preprocess, model and interpret biosignals such as speech, motion, muscle and brain activities. The focus of this toolkit is to enable researchers from various communities to pursue their experiments and integrate real-time biosignal interpretation into their applications. BioKIT boosts a flexible two-layer structure with a modular C++ core that interfaces with a Python scripting layer, to facilitate development of new applications. BioKIT employs sequence-level parallelization and memory sharing across threads. Additionally, a fully integrated error blaming component facilitates in-depth analysis. A generic terminology keeps the barrier to entry for researchers from multiple fields to a minimum. We describe our onlinecapable dynamic decoder and report on initial experiments on three different tasks. The presented speech recognition experiments employ Kaldi [1] trained deep neural networks with the results set in relation to the real time factor needed to obtain them.
Dominic Telaar, Michael Wand 0002, Dirk Gehrig, Felix Putze, Christoph Amma, Dominic Heger, Ngoc Thang Vu, Mark Erhardt, Tim Schlippe, Matthias Janke, Christian Herff, Tanja Schultz
INTERSPEECH2
2014 The EMG-UKA corpus for electromyographic speech processing
abstract
This article gives an overview of the EMG-UKA corpus, a corpus of electromyographic (EMG) recordings of articulatory activity enabling speech processing (in particular speech recognition and synthesis) based on EMG signals, with the purpose of building Silent Speech interfaces. Data is available in multiple speaking modes, namely audibly spoken, whispered, and silently articulated speech. Besides the EMG data, synchronous acoustic data was additionally recorded to serve as a reference. The corpus comprises 63 recorded sessions from 8 speakers, the total amount of data is 7:32 hours. A trial subset, consisting of 1:52 hours of data, is freely available for download.
Michael Wand 0002, Matthias Janke, Tanja Schultz
INTERSPEECH1
2014 Towards real-life application of EMG-based speech recognition by using unsupervised adaptation
abstract
This paper deals with a Silent Speech Interface based on Surface Electromyography (EMG), where electrodes capture the electric activity generated by the articulatory muscles from a user’s face in order to decode the underlying speech, allowing speech to be recognized even when no sound is heard or created. So far, most EMG-based speech recognizers described in literature do not allow electrode reattachment between system training and usage, which we consider unsuitable for practical applications. In this study we report on our research on unsupervised session adaptation: A system is pre-trained with data from multiple recording sessions and then adapted towards the current recording session using data accruable during normal use, without requiring a time-consuming specific enrollment phase. We show that considerable accuracy improvements can be achieved with this method, paving the way towards real-life applications of the technology. Index Terms: Silent Speech Interfaces, EMG, EMG-based Speech Recognition, Unsupervised Adaptation
Michael Wand 0002, Tanja Schultz
INTERSPEECH1
2014 Conversion from facial myoelectric signals to speech: a unit selection approach
abstract
This paper reports on our recent research on surface electromyographic (EMG) speech synthesis: a direct conversion of the EMG signals of the articulatory muscle movements to the acoustic speech signal. In this work we introduce a unit selection approach which compares segments of the input EMG signal to a database of simultaneously recorded EMG/audio unit pairs and selects the best matching audio unit based on target and concatenation cost, which will be concatenated to synthesize an acoustic speech output. We show that this approach is feasible to generate a proper speech output from the input EMG signal. We evaluate different properties of the units and investigate what amount of data is necessary for an initial transformation. Prior work on EMG-to-speech conversion used a framebased approach from the voice conversion domain, which struggles with the generation of a natural $F_0$ contour. This problem may also be tackled by our unit selection approach.
Marlene Zahner, Matthias Janke, Michael Wand 0002, Tanja Schultz
INTERSPEECH3
2012 Further investigations on EMG-to-speech conversion
abstract
Our study deals with a Silent Speech Interface based on mapping surface electromyographic (EMG) signals to speech waveforms. Electromyographic signals recorded from the facial muscles capture the activity of the human articulatory apparatus and therefore allow to retrace speech, even when no audible signal is produced. The mapping of EMG signals to speech is done via a Gaussian mixture model (GMM)-based conversion technique. In this paper, we follow the lead of EMG-based speech-to-text systems and apply two major recent technological advances to our system, namely, we consider session-independent systems, which are robust against electrode repositioning, and we show that mapping the EMG signal to whispered speech creates a better speech signal than a mapping to normally spoken speech. We objectively evaluate the performance of our systems using a spectral distortion measure.
Matthias Janke, Michael Wand 0002, Keigo Nakamura, Tanja Schultz
ICASSP2
2011 Estimation of fundamental frequency from surface electromyographic data: EMG-to-F0
abstract
In this paper, we present our recent studies of F0estimation from the surface electromyographic (EMG) data us ing a Gaussian mixture model (GMM)-based voice con version (VC) technique, referred to as EMG-to-F0. In our approach, a support vector machine recognizes individual frames as unvoiced and voiced (U/V), and voiced F0contours are discriminated by the trained GMM based on the manner of minimum mean-square error. EMG-to-F0is experimentally evaluated using three data sets of different speakers. Each data set includes almost 500 utterances. Objective experiments demonstrate that we achieve a correlation coefficient of up to 0.49 between estimated and target F0contours with more than 84% U/V decision accuracy, although the results have large variations.
Keigo Nakamura, Matthias Janke, Michael Wand 0002, Tanja Schultz
ICASSP3
2011 Analysis of phone confusion in EMG-based speech recognition
abstract
In this paper we present a study on phone confusabilities based on phone recognition experiments from facial surface electromyographic (EMG) signals. In our study EMG captures the electrical potentials of the human articulatory muscles. This technology can be used to create Silent Speech Interfaces, where a user can communicate naturally without uttering any sound. This paper investigates to which extent different phone properties can be recognized from an EMG signal, shows which weaknesses have yet to be overcome, and compares the results to acoustic-based recognition of phones.
Michael Wand 0002, Tanja Schultz
ICASSP1
2011 Impact of Different Feedback Mechanisms in EMG-Based Speech Recognition
abstract
This paper reports on our recent research in the feedback effects of Silent Speech. Our technology is based on surface electromyography (EMG) which captures the electrical potentials of the human articulatory muscles rather than the acoustic speech signal. While recognition results are good for loudly articulated speech and when experienced users speak silently, novice users usually achieve far worse results when speaking silently. Since there is no acoustic feedback when speaking silently, we investigate different kinds of feedback modes: no additional feedback except the natural somatosensory feedback (like the touching of the lips), visual feedback using a mirror and indirect acoustic feedback by speaking simultaneously to a previously recorded audio signal. In addition we examine recorded EMG data when the subject speaks audibly and silently in a loud environment to see if the Lombard effect can be observed in Silent Speech, too.
Christian Herff, Matthias Janke, Michael Wand 0002, Tanja Schultz
INTERSPEECH3
2011 Investigations on Speaking Mode Discrepancies in EMG-Based Speech Recognition
abstract
In this paper we present our recent study on the impact of speaking mode variabilities on speech recognition by surface electromyography (EMG). Surface electromyography captures the electric potentials of the human articulatory muscles, which enables a user to communicate naturally without making any audible sound. Our previous experiments have shown that the EMG signal varies greatly between different speaking modes, like audibly uttered speech and silently articulated speech. In this study we extend our previous research and quantify the impact of different speaking modes by investigating the amount of mode-specific leaves in phonetic decision trees. We show that this measure correlates highly with discrepancies in the spectral energy of the EMG signal, as well as with differences in the performance of a recognizer on different speaking modes. We furthermore present how EMG signal adaptation by spectral mapping decreases the effect of the speaking mode.
Michael Wand 0002, Matthias Janke, Tanja Schultz
INTERSPEECH1
2010 Impact of lack of acoustic feedback in EMG-based silent speech recognition
abstract
This paper presents our recent advances in speech recognition based on surface electromyography (EMG). This technology allows for Silent Speech Interfaces since EMG captures the electrical potentials of the human articulatory muscles rather than the acoustic speech signal. Our earlier experiments have shown that the EMG signal is greatly impacted by the mode of speaking. In this study we extend this line of research by comparing EMG signals from audible, whispered, and silent speaking mode. We distinguish between phonetic features like consonants and vowels and show that the lack of acoustic feedback in silent speech implies an increased focus on somatosensoric feedback, which is visible in the EMG signal. Based on this analysis we develop a spectral mapping method to compensate for these differences. Finally, we apply the spectral mapping to the front-end of our speech recognition system and show that recognition rates on silent speech improve by up to 11.59% relative. Index Terms: EMG, EMG-based speech recognition, Silent Speech Interfaces, somatosensoric feedback
Matthias Janke, Michael Wand 0002, Tanja Schultz
INTERSPEECH2
2010 Modeling coarticulation in EMG-based continuous speech recognition
Tanja Schultz, Michael Wand 0002
Speech Commun.2
2009 Synthesizing speech from electromyography using voice transformation techniques
abstract
Surface electromyography (EMG) can be used to record the activation potentials of articulatory muscles while a person speaks. This technique could enable silent speech interfaces, as EMG signals are generated even when people pantomime speech without producing sound. Having effective silent speech interfaces would enable a number of compelling applications, allowing people to communicate in areas where they would not want to be overheard or where the background noise is so prevalent that they could not be heard. In order to use EMG signals in speech interfaces, however, there must be a relatively accurate method to map the signals to speech. Up to this point, it appears that most attempts to use EMG signals for speech interfaces have focused on Automatic Speech Recognition (ASR) based on features derived from EMG signals. Following the lead of other researchers who worked with Electro-Magnetic Articulograph (EMA) data and Non-Audible Murmur (NAM) speech, we explore the alternative idea of using Voice Transformation (VT) techniques to synthesize speech from EMG signals. With speech output, both ASR systems and human listeners can directly use EMG-based systems. We report the results of our preliminary studies, noting the difficulties we encountered and suggesting areas for future work. Index Terms: electromyography, silent speech, voice transformation, speech synthesis
Arthur R. Toth, Michael Wand 0002, Tanja Schultz
INTERSPEECH2
2009 Impact of different speaking modes on EMG-based speech recognition
abstract
We present our recent results on speech recognition by surface electromyography (EMG), which captures the electric potentials that are generated by the human articulatory muscles. This technique can be used to enable Silent Speech Interfaces, since EMG signals are generated even when people only articulate speech without producing any sound. Preliminary experiments have shown that the EMG signals created by audible and silent speech are quite distinct. In this paper we first compare various methods of initializing a silent speech EMG recognizer, showing that the performance of the recognizer substantially varies across different speakers. Based on this, we analyze EMG signals from audible and silent speech, present first results on how discrepancies between these speaking modes affect EMG recognizers, and suggest areas for future work. Index Terms: speech recognition, surface electromyography, silent speech, articulation
Michael Wand 0002, Szu-Chen Stan Jou, Arthur R. Toth, Tanja Schultz
INTERSPEECH1
2007 Wavelet-based front-end for electromyographic speech recognition
abstract
In this paper we present our investigations on the potential of wavelet-based preprocessing for surface electromyographic speech recognition.We implemented several variants of the Discrete Wavelet Transform and applied them to electromyographical data.First we examined different transforms with various filters and decomposition levels and found that the Redundant Discrete Wavelet Transform performs the best among all tested wavelet transforms.Furthermore, we compared the best wavelet transform to our EMG optimized spectral-and timedomain features.The results showed that the best wavelet transform slightly outperforms the optimized features with 30.9% word error rate compared to 32% for the optimized EMG spectral and time-domain features.Both numbers were achieved on a 108 word vocabulary test set using phone based acoustic models trained on continuously spoken speech captured by EMG.
Michael Wand 0002, Szu-Chen Stan Jou, Tanja Schultz
INTERSPEECH1