EDBT 2026 Demo / reviewers in the wild / expert
John B. Harvill
dblp:278/2872
· DBLP profile ↗
12ranked-venue papers
8as first author
11since 2021 · last 2024
0000-0003-3633-6756ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | LI-TTA: Language Informed Test-Time Adaptation for Automatic Speech Recognition
Eunseop Yoon, Hee Suk Yoon, John B. Harvill, Mark Hasegawa-Johnson, Chang Dong Yoo |
INTERSPEECH | 3 |
| 2023 | SPADE: Self-Supervised Pretraining for Acoustic DisentanglementabstractSelf-supervised representation learning approaches have grown in popularity due to the ability to train models on large amounts of unlabeled data and have demonstrated success in diverse fields such as natural language processing, computer vision, and speech. Previous self-supervised work in the speech domain has disentangled multiple attributes of speech such as linguistic content, speaker identity, and rhythm. In this work, we introduce a self-supervised approach to disentangle room acoustics from speech and use the acoustic representation on the downstream task of device arbitration. Our results demonstrate that our proposed approach significantly improves performance over a baseline when labeled training data is scarce, indicating that our pretraining scheme learns to encode room acoustic information while remaining invariant to other attributes of the speech signal. John B. Harvill, Jarred Barber, Arun Nair, Ramin Pishehvar |
ICASSP | 1 |
| 2023 | Mitigating the Exposure Bias in Sentence-Level Grapheme-to-Phoneme (G2P) Transduction
Eunseop Yoon, Hee Suk Yoon, Dhananjaya Gowda, SooHwan Eom, Daehyeok Kim, John B. Harvill, Heting Gao, Mark Hasegawa-Johnson, Chanwoo Kim 0001, Chang Dong Yoo |
INTERSPEECH | 6 |
| 2023 | Quantifying Emotional Similarity in SpeechabstractThis study proposes the novel formulation of measuring emotional similarity between speech recordings. This formulation explores the ordinal nature of emotions by comparing emotional similarities instead of predicting an emotional attribute, or recognizing an emotional category. The proposed task determines which of two alternative samples has the most similar emotional content to the emotion of a given anchor. This task raises some interesting questions. Which is the emotional descriptor that provide the most suitable space to assess emotional similarities Can deep neural networks (DNNs learn representations to robustly quantify emotional similarities We address these questions by exploring alternative emotional spaces created with attribute-based descriptors and categorical emotions. We create the representation using a DNN trained with the triplet loss function, which relies on triplets formed with an anchor, a positive example, and a negative example. We select a positive sample that has similar emotion content to the anchor, and a negative sample that has dissimilar emotion to the anchor. The task of our DNN is to identify the positive sample. The experimental evaluations demonstrate that we can learn a meaningful embedding to assess emotional similarities, achieving higher performance than human evaluators asked to complete the same task. John B. Harvill, Seong-Gyun Leem, Mohammed Abdel-Wahab 0001, Reza Lotfian, Carlos Busso |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Detection of Covid-19 from Joint Time and Frequency Analysis of Speech, Breathing and Cough AudioabstractThe distinct cough sounds produced by a variety of respiratory diseases suggest the potential for the development of a new class of audio bio-markers for the detection of COVID-19. Accurate audio biomarker-based COVID-19 tests would be inexpensive, readily scalable, and non-invasive. Audio biomarker screening could also be utilized in resource-limited settings prior to traditional diagnostic testing. Here we explore the possibility of leveraging three audio modalities: cough, breathing, and speech to determine COVID-19 status. We train a separate neural classification system on each modality, as well as a fused classification system on all three modalities together. Ablation studies are performed to understand the relationship between individual and collective performance of the modalities. Additionally, we analyze the extent to which temporal and spectral features contribute to COVID-19 status information contained in the audio signals. John B. Harvill, Yash R. Wani, Moitreya Chatterjee, Mustafa Alam, David G. Beiser, David Chestek, Mark Hasegawa-Johnson, Narendra Ahuja |
ICASSP | 1 |
| 2022 | Frame-Level Stutter Detection
John B. Harvill, Mark Hasegawa-Johnson, Chang Dong Yoo |
INTERSPEECH | 1 |
| 2022 | Syn2Vec: Synset Colexification Graphs for Lexical Semantic SimilarityabstractIn this paper we focus on patterns of colexification (co-expressions of form-meaning mapping in the lexicon) as an aspect of lexicalsemantic organization, and use them to build large scale synset graphs across BabelNet's typologically diverse set of 499 world languages.We introduce and compare several approaches: monolingual and cross-lingual colexification graphs, popular distributional models, and fusion approaches.The models are evaluated against human judgments on a semantic similarity task for nine languages.Our strong empirical findings also point to the importance of universality of our graph synset embedding representations with no need for any language-specific adaptation when evaluated on the lexical similarity task.The insights of our exploratory investigation of largescale colexification graphs could inspire significant advances in NLP across languages, especially for tasks involving languages which lack dedicated lexical resources, and can benefit from language transfer from large shared cross-lingual semantic spaces. John B. Harvill, Roxana Girju, Mark Hasegawa-Johnson |
NAACL-HLT | 1 |
| 2022 | Seamless equal accuracy ratio for inclusive CTC speech recognitionabstractConcerns have been raised regarding performance disparity in automatic speech recognition (ASR) systems as they provide unequal transcription accuracy for different user groups defined by different attributes that include gender, dialect, and race. In this paper, we propose “equal accuracy ratio”, a novel inclusiveness measure for ASR systems that can be seamlessly integrated into the standard connectionist temporal classification (CTC) training pipeline of an end-to-end neural speech recognizer to increase the recognizer’s inclusiveness. We also create a novel multi-dialect benchmark dataset to study the inclusiveness of ASR, by combining data from existing corpora in seven dialects of English (African American, General American, Latino English, British English, Indian English, Afrikaaner English, and Xhosa English). Experiments on this multi-dialect corpus show that using the equal accuracy ratio as a regularization term along with CTC loss, succeeds in lowering the accuracy gap between user groups and reduces the recognition error rate compared with a non-regularized baseline. Experiments on additional speech corpora that have different user groups also confirm our findings. Heting Gao, Sunghun Kang, Rusty Mina, Dias Issa, John B. Harvill, Leda Sari, Mark Hasegawa-Johnson, Chang Dong Yoo |
Speech Commun. | 6 |
| 2021 | The Twelvefold Way of Non-Sequential Lossless CompressionabstractMany information sources are not just sequences of distinguishable symbols but rather have invariances governed by alternative counting paradigms such as permutations, combinations, and partitions. We consider an entire classification of these invariances called the twelvefold way in enumerative combinatorics and develop a method to characterize lossless compression limits. Explicit computations for all twelve settings are carried out for i.i.d. uniform and Bernoulli distributions. Comparisons among settings provide quantitative insight. Taha Ameen ur Rahman, Alton S. Barbehenn, Xinan Chen 0003, Hassan Dbouk, James A. Douglas, Yuncong Geng, Ian George, John B. Harvill, Sung Woo Jeon, Kartik K. Kansal, Kiwook Lee, Kelly A. Levick, Bochao Li, Yashaswini Murthy, Adarsh Muthuveeru-Subramaniam, S. Yagiz Olmez, Matthew J. Tomei, Tanya Veeravalli, Xuechao Wang, Eric A. Wayman, Fan Wu 0011, Heling Zhang, Sourya Basu, Lav R. Varshney |
DCC | 8 |
| 2021 | Synthesis of New Words for Improved Dysarthric Speech Recognition on an Expanded VocabularyabstractDysarthria is a condition where people experience a reduction in speech intelligibility due to a neuromotor disorder. Previous works in dysarthric speech recognition have focused on accurate recognition of words encountered in training data. Due to the rarity of dysarthria in the general population, a relatively small amount of publicly-available training data exists for dysarthric speech. The number of unique words in these datasets is small, so ASR systems trained with existing dysarthric speech data are limited to recognition of those words. In this paper, we propose a data augmentation method using voice conversion that allows dysarthric ASR systems to accurately recognize words outside of the training set vocabulary. We demonstrate that a small amount of dysarthric speech data can be used to capture the relevant vocal characteristics of a speaker with dysarthria through a parallel voice conversion system. We show that it’s possible to synthesize utterances of new words that were never recorded by speakers with dysarthria, and that these synthesized utterances can be used to train a dysarthric ASR system. John B. Harvill, Dias Issa, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICASSP | 1 |
| 2021 | Classification of COVID-19 from Cough Using Autoregressive Predictive Coding Pretraining and Spectral Data AugmentationabstractSerum and saliva-based testing methods have been crucial to slowing the COVID-19 pandemic, yet have been limited by slow throughput and cost.A system able to determine COVID-19 status from cough sounds alone would provide a low cost, rapid, and remote alternative to current testing methods.We explore the applicability of recent techniques such as pre-training and spectral augmentation in improving the performance of a neural cough classification system.We use Autoregressive Predictive Coding (APC) to pre-train a unidirectional LSTM on the COUGHVID dataset.We then generate our final model by finetuning added BLSTM layers on the DiCOVA challenge dataset.We perform various ablation studies to see how each component impacts performance and improves generalization with a small dataset.Our final system achieves an AUC of 85.35 and places third out of 29 entries in the DiCOVA challenge. John B. Harvill, Yash R. Wani, Mark Hasegawa-Johnson, Narendra Ahuja, David G. Beiser, David Chestek |
Interspeech | 1 |
| 2019 | Retrieving Speech Samples with Similar Emotional Content Using a Triplet Loss FunctionabstractThe ability to identify speech with similar emotional content is valuable to many applications, including speech retrieval, surveillance, and emotional speech synthesis. While current formulations in speech emotion recognition based on classification or regression are not appropriate for this task, solutions based on preference learning offer appealing approaches for this task. This paper aims to find speech samples that are emotionally similar to an anchor speech sample provided as a query. This novel formulation opens interesting research questions. How well can a machine complete this task? How does the accuracy of automatic algorithms compare to the performance of a human performing this task? This study addresses these questions by training a deep learning model using a triplet loss function, mapping the acoustic features into an embedding that is discriminative for this task. The network receives an anchor speech sample and two competing speech samples, and the task is to determine which of the candidate speech sample conveys the closest emotional content to the emotion conveyed by the anchor. By comparing the results from our model with human perceptual evaluations, this study demonstrates that the proposed approach has performance very close to human performance in retrieving samples with similar emotional content. John B. Harvill, Mohammed Abdel-Wahab 0001, Reza Lotfian, Carlos Busso |
ICASSP | 1 |