VLDB 2026 Research / reviewers in the wild / expert
Matthew Perez
dblp:196/3191
· DBLP profile ↗
12ranked-venue papers
6as first author
9since 2021 · last 2025
0009-0005-5348-3020ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multimodal Classroom Diarization with GPT Re-scoring: Teacher or Student?
Matthew Perez, Berk Coker, Kemal Berk Kocabagli, Jessica Vitale, Alyssa Van Camp |
AIED (5) | 1 |
| 2024 | Beyond Binary: Multiclass Paraphasia Detection with Generative Pretrained Transformers and End-to-End Models
Matthew Perez, Aneesha Sampath, Minxue Niu, Emily Mower Provost |
INTERSPEECH | 1 |
| 2023 | Episodic Memory For Domain-Adaptable, Robust Speech Emotion Recognition
James Tavernor, Matthew Perez, Emily Mower Provost |
INTERSPEECH | 2 |
| 2023 | PronScribe: Highly Accurate Multimodal Phonemic Transcription From Speech and TextabstractWe present PronScribe, a novel method for phonemic transcription from speech and text input based on careful finetuning and adaptation of a massive, multilingual, multimodal speech-text pretrained model.We show that our model is capable of phonemically transcribing pronunciations of full utterances with accurate word boundaries in a variety of languages covering diverse phonological phenomena, achieving phoneme error rates in the vicinity of 1-2% which is comparable to human transcribers.We show that PronScribe can effectively learn this task from relatively little training data, making it attractive even in low-resource settings.It learns from text and speech simultaneously in a coherent way, and is better than previous models using speech, text or both.Additionally, the model's good transfer learning characteristics in multilingual settings can effectively boost performance for lower-resourced languages. Matthew Perez, Ankur Bapna, Fadi Haik, Siamak Tazari |
INTERSPEECH | 2 |
| 2022 | Mind the gap: On the value of silence representations to lexical-based speech emotion recognition
Matthew Perez, Mimansa Jaiswal, Minxue Niu, Cristina Gorrostieta, Matthew Roddy, Kye Taylor, Reza Lotfian, Emily Mower Provost |
INTERSPEECH | 1 |
| 2022 | Enabling Off-the-Shelf Disfluency Detection and Categorization for Pathological SpeechabstractA speech disfluency, such as a filled pause, repetition, or revision, disrupts the typical flow of speech. Disfluency modeling has grown as a research area, as recent work has shown that these disfluencies may help in assessing health conditions. For example, for individuals with cognitive impairment, changes in disfluencies may indicate worsening symptoms. However, work on disfluency modeling has focused heavily on detection and less on categorization. Work that has focused on categorization has suffered with two specific classes: repetitions and revisions. In this paper, we evaluate how BERT (Bidirectional Encoder Representations from Transformers) compares to other models on disfluency detection and categorization. We also propose adding a second fine-tuning task where BERT learns to distance repetitions and revisions from their repairs with triplet loss. We find that BERT and BERT with triplet loss outperform previous work on disfluency detection and categorization, particularly for repetitions and revisions. In this paper we present the first analysis of how these models can be fine-tuned on widely available disfluency data, and then used in an off-the-shelf manner on small corpora of pathological speech. Amrit Romana, Minxue Niu, Matthew Perez, Angela Roberts 0001, Emily Mower Provost |
INTERSPEECH | 3 |
| 2021 | Articulatory Coordination for Speech Motor Tracking in Huntington DiseaseabstractHuntington Disease (HD) is a progressive disorder which often manifests in motor impairment. Motor severity (captured via motor score) is a key component in assessing overall HD severity. However, motor score evaluation involves in-clinic visits with a trained medical professional, which are expensive and not always accessible. Speech analysis provides an attractive avenue for tracking HD severity because speech is easy to collect remotely and provides insight into motor changes. HD speech is typically characterized as having irregular articulation. With this in mind, acoustic features that can capture vocal tract movement and articulatory coordination are particularly promising for characterizing motor symptom progression in HD. In this paper, we present an experiment that uses Vocal Tract Coordination (VTC) features extracted from read speech to estimate a motor score. When using an elastic-net regression model, we find that VTC features significantly outperform other acoustic features across varied-length audio segments, which highlights the effectiveness of these features for both short- and long-form reading tasks. Lastly, we analyze the F-value scores of VTC features to visualize which channels are most related to motor score. This work enables future research efforts to consider VTC features for acoustic analyses which target HD motor symptomatology tracking. Matthew Perez, Amrit Romana, Angela Roberts 0001, Noelle Carlozzi, Jennifer Ann Miner, Praveen Dayalu, Emily Mower Provost |
Interspeech | 1 |
| 2021 | Automatically Detecting Errors and Disfluencies in Read Speech to Predict Cognitive Impairment in People with Parkinson's DiseaseabstractParkinson's disease (PD) is a central nervous system disorder that causes motor impairment. Recent studies have found that people with PD also often suffer from cognitive impairment (CI). While a large body of work has shown that speech can be used to predict motor symptom severity in people with PD, much less has focused on cognitive symptom severity. Existing work has investigated if acoustic features, derived from speech, can be used to detect CI in people with PD. However, these acoustic features are general and are not targeted toward capturing CI. Speech errors and disfluencies provide additional insight into CI. In this study, we focus on read speech, which offers a controlled template from which we can detect errors and disfluencies, and we analyze how errors and disfluencies vary with CI. The novelty of this work is an automated pipeline, including transcription and error and disfluency detection, capable of predicting CI in people with PD. This will enable efficient analyses of how cognition modulates speech for people with PD, leading to scalable speech assessments of CI. Amrit Romana, John Bandon, Matthew Perez, Stephanie Gutierrez, Richard Richter, Angela Roberts 0001, Emily Mower Provost |
Interspeech | 3 |
| 2021 | Learning Paralinguistic Features from Audiobooks through Style Voice ConversionabstractZakaria Aldeneh, Matthew Perez, Emily Mower Provost. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zakaria Aldeneh, Matthew Perez, Emily Mower Provost |
NAACL-HLT | 2 |
| 2020 | Aphasic Speech Recognition Using a Mixture of Speech Intelligibility ExpertsabstractRobust speech recognition is a key prerequisite for semantic feature extraction in automatic aphasic speech analysis. However, standard one-size-fits-all automatic speech recognition models perform poorly when applied to aphasic speech. One reason for this is the wide range of speech intelligibility due to different levels of severity (i.e., higher severity lends itself to less intelligible speech). To address this, we propose a novel acoustic model based on a mixture of experts (MoE), which handles the varying intelligibility stages present in aphasic speech by explicitly defining severity-based experts. At test time, the contribution of each expert is decided by estimating speech intelligibility with a speech intelligibility detector (SID). We show that our proposed approach significantly reduces phone error rates across all severity stages in aphasic speech compared to a baseline approach that does not incorporate severity information into the modeling process. Matthew Perez, Zakaria Aldeneh, Emily Mower Provost |
INTERSPEECH | 1 |
| 2018 | Classification of Huntington Disease Using Acoustic and Lexical FeaturesabstractSpeech is a critical biomarker for Huntington Disease (HD), with changes in speech increasing in severity as the disease progresses. Speech analyses are currently conducted using either transcriptions created manually by trained professionals or using global rating scales. Manual transcription is both expensive and time-consuming and global rating scales may lack sufficient sensitivity and fidelity [1]. Ultimately, what is needed is an unobtrusive measure that can cheaply and continuously track disease progression. We present first steps towards the development of such a system, demonstrating the ability to automatically differentiate between healthy controls and individuals with HD using speech cues. The results provide evidence that objective analyses can be used to support clinical diagnoses, moving towards the tracking of symptomatology outside of laboratory and clinical environments. Matthew Perez, Wenyu Jin 0001, Noelle Carlozzi, Praveen Dayalu, Angela Roberts 0001, Emily Mower Provost |
INTERSPEECH | 1 |
| 2017 | Portable mTBI Assessment Using Temporal and Frequency Analysis of SpeechabstractThis paper shows that extraction and analysis of various acoustic features from speech using mobile devices can allow the detection of patterns that could be indicative of neurological trauma. This may pave the way for new types of biomarkers and diagnostic tools. Toward this end, we created a mobile application designed to diagnose mild traumatic brain injuries (mTBI) such as concussions. Using this application, data were collected from youth athletes from 47 high schools and colleges in the Midwestern United States. In this paper, we focus on the design of a methodology to collect speech data, the extraction of various temporal and frequency metrics from that data, and the statistical analysis of these metrics to find patterns that are indicative of a concussion. Our results suggest a strong correlation between certain temporal and frequency features and the likelihood of a concussion. Louis Daudet, Nikhil Yadav, Matthew Perez, Christian Poellabauer, Sandra L. Schneider, Alan Huebner |
IEEE J. Biomed. Health Informatics | 3 |