Mark C. Fuhs

dblp:25/516 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2025 A Hybrid Approach to Combining Role Diarization with ASR for Professional Conversations
Bongjun Kim, Mark C. Fuhs, Anurag Chowdhury, Deblin Bagchi, Monika Woszczyna
INTERSPEECH3
2024 Investigating Confidence Estimation Measures for Speaker Diarization
Anurag Chowdhury, Abhinav Misra, Mark C. Fuhs, Monika Woszczyna
INTERSPEECH3
2024 Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition
abstract
One solution to automatic speech recognition (ASR) of overlapping speakers is to separate speech and then perform ASR on the separated signals.Commonly, the separator produces artefacts which often degrade ASR performance.Addressing this issue typically requires reference transcriptions to jointly train the separation and ASR networks.This is often not viable for training on real-world in-domain audio where reference transcript information is not always available.This paper proposes a transcription-free method for joint training using only audio signals.The proposed method uses embedding differences of pre-trained ASR encoders as a loss with a proposed modification to permutation invariant training (PIT) called guided PIT (GPIT).The method achieves a 6.4% improvement in word error rate (WER) measures over a signal-level loss and also shows enhancement improvements in perceptual measures such as short-time objective intelligibility (STOI).
William Ravenscroft, George Close, Stefan Goetze, Thomas Hain, Mohammad Soleymanpour, Anurag Chowdhury, Mark C. Fuhs
INTERSPEECH7
2022 Low-resource Low-footprint Wake-word Detection using Knowledge Distillation
abstract
As virtual assistants have become more diverse and specialized, so has the demand for application or brand-specific wake words.However, the wake-word-specific datasets typically used to train wake-word detectors are costly to create.In this paper, we explore two techniques to leverage acoustic modeling data for large-vocabulary speech recognition to improve a purpose-built wake-word detector: transfer learning and knowledge distillation.We also explore how these techniques interact with timesynchronous training targets to improve detection latency.Experiments are presented on the open-source "Hey Snips" dataset and a more challenging in-house far-field dataset.Using phonesynchronous targets and knowledge distillation from a large acoustic model, we are able to improve accuracy across dataset sizes for both datasets while reducing latency.
Mark C. Fuhs, Deblin Bagchi, Bahman Farahani, Monika Woszczyna
INTERSPEECH2
2021 Are You Dictating to Me? Detecting Embedded Dictations in Doctor-Patient Conversations
abstract
Medical scribes chart doctor-patient conversations in real time or by listening to an audio recording afterwards. Doctors sometimes dictate during a patient encounter, a highly informative part for a scribe. We introduced a light-weight annotation schema and ana-lyzed recordings of 105 randomly selected doctor-patient encounters from 21 physicians to quantify the frequency and automatically de-tect dictated regions. Dictation behavior of individual doctors was consistent but varied among them. A linguistic analysis is provided to describe differences of doctors speech when talking to a patient or dictating. A description of the data is given, highlighting challenges of segmenting audio into conversation and dictation regions. We in-vestigate different features and methods to segment conversations including keyword spotting, acoustic features and class-conditioned language models. Results are anchored to a majority class base-line. Using only acoustic features allows to predict dictated speech without the need of a speech recognition system performing com-parable to a rule-based approach using lexical features derived from a speech recognition system. Performance is assessed using leave-one-physician-out cross validation and an analysis using a random forest classifier indicates that language model derived features are most useful, and that a combination of acoustic and lexical features performed best.
Thomas Schaaf, Longxiang Zhang, Alireza Bayestehtashk, Mark C. Fuhs, Shahid Durrani, Susanne Burger, Monika Woszczyna, Thomas Polzin
ASRU4
2013 Neighbour selection and adaptation for rapid speaker-dependent ASR
abstract
Speaker dependent (SD) ASR systems have significantly lower word error rates (WER) compared to speaker independent (SI) systems. However, SD systems require sufficient training data from the target speaker, which is impractical to collect in a short time. We present a technique for training SD models using just few minutes of speaker's data. We compensate for the lack of adequate speaker-specific data by selecting neighbours from a database of existing speakers who are acoustically close to the target speaker. These neighbours provide ample training data, which is used to adapt the SI model to obtain an initial SD model for the new speaker with significantly lower WER. We evaluate various neighbour selection algorithms on a large-scale medical transcription task and report significant reduction in WER using only 5 mins of speaker-specific data. We conduct a detailed analysis of various factors such as gender and accent in the neighbour selection. Finally, we study neighbour selection and adaptation in the context of discriminative objective functions.
Udhyakumar Nallasamy, Mark C. Fuhs, Monika Woszczyna, Florian Metze, Tanja Schultz
ASRU2
2013 Accent- and speaker-specific polyphone decision trees for non-native speech recognition
abstract
Acoustic models in state-of-the-art LVCSR systems are typically trained on data from thousands of speakers and then adapted to a speaker using, e.g., various combinations of CMLLR, MLLR and MAP. This adaptation step is particularly important for speakers with accents that are not well represented in the training set. The present study explores how to improve performance on South-Asian-accented speakers (SoA-accented) with the availability of thousands of US-accented, hundreds of SoA-accented, and tens of hours of speaker-specific training data. We employ a decision tree similarity measure to analyze how varying co-articulations across accents and people manifest themselves in the decision tree. Modeling these variations in addition to adapting the GMMs of an existing baseline system to a speaker improved WER for small systems (1k GMMs), but improvement for systems with larger trees (2k, 3k GMMs) was modest. Overall, GMM adaptation/retraining yields significant performance benefits, and training a SoA-accent-specific system is particularly worthwhile when lacking speaker adaptation data.
Dominic Telaar, Mark C. Fuhs
INTERSPEECH2
2009 Detecting bandlimited audio in broadcast television shows
abstract
For TV and radio shows containing narrowband speech, Speech-to-text (STT) accuracy on the narrowband audio can be improved by using an acoustic model trained on acoustically matched data. To selectively apply it, one must first be able to accurately detect which audio segments are narrowband. The present paper explores two different bandwidth classification approaches: a traditional Gaussian mixture model (GMM) approach and a spline-based classifier that categorizes audio segments based on their power spectra. We focus on shows found in the DARPA GALE Mandarin training and test sets, where the ratio of wideband to narrowband shows is very large. In this setting, the spline-based classifier reduces the number of misclassified wideband segments by up to 95% relative to the GMM-based classifier for the same number of misclassified narrowband segments.
Mark C. Fuhs, Qin Jin, Tanja Schultz
ICASSP1
2008 The CMU-interACT 2008 Mandarin transcription system
abstract
We present our Mandarin BN/BC transcription system recently developed for the GALE07 evaluation. The system employs a 3-pass decoding strategy trained with over 1300 hours of quickly transcribed audio. We successfully apply discriminative training, dynamic unsupervised language model adaptation, and system combination techniques in our system. We furthermore achieve improvements by combining an Initial-Final system with a genre dependent phone system. On the GALE07 phase 2 retest evaluation, our system achieves a character error rate(CER) of 13.3 % on dev07 test set and 13.5 % on eval07 unsequestered test set. Our system also allows combination with other sites and in this paper, we investigate different system combination strategies which significantly improve the final recognition performance. Index Terms: Mandarin transcription system, broadcast news, broadcast conversation, GALE evaluation
Roger Hsiao, Mark C. Fuhs, Yik-Cheung Tam, Qin Jin, Tanja Schultz
INTERSPEECH2
2007 Context Learning in the Rodent Hippocampus
abstract
We present a Bayesian statistical theory of context learning in the rodent hippocampus. While context is often defined in an experimental setting in relation to specific background cues or task demands, we advance a single, more general notion of context that suffices for a variety of learning phenomena. Specifically, a context is defined as a statistically stationary distribution of experiences, and context learning is defined as the problem of how to form contexts out of groups of experiences that cluster together in time. The challenge of context learning is solving the model selection problem: How many contexts make up the rodent's world? Solving this problem requires balancing two opposing goals: minimize the variability of the distribution of experiences within a context and minimize the likelihood of transitioning between contexts. The theory provides an understanding of why hippocampal place cell remapping sometimes develops gradually over many days of experience and why even consistent landmark differences may need to be relearned after other environmental changes. The theory provides an explanation for progressive performance improvements in serial reversal learning, based on a clear dissociation between the incremental process of context learning and the relatively abrupt context selection process. The impact of partial reinforcement on reversal learning is also addressed. Finally, the theory explains why alternating sequence learning does not consistently result in unique context-dependent sequence representations in hippocampus.
Mark C. Fuhs, David S. Touretzky
Neural Comput.1
2000 Synaptic learning models of map separation in the hippocampus
Mark C. Fuhs, David S. Touretzky
Neurocomputing1