EDBT 2026 Demo / reviewers in the wild / expert
Jackson Liscombe
dblp:59/3748
· DBLP profile ↗
33ranked-venue papers
5as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Exploration of Interpretable Deep Learning Models for the Assessment of Mild Cognitive ImpairmentabstractEarly diagnosis and intervention are crucial for mild cognitive impairment (MCI), as MCI often progresses to more severe neurodegenerative conditions. In this study, we explore utilizing deep learning for MCI detection without loosing the interpretability provided by feature-based approaches. We used a dataset consisting of 90 MCI patients and 91 controls collected via a remote assessment platform and analyzed the participants' spontaneous speech responses to the Patient Report of Problems (PROP) which asks patients to report their most bothersome general health problems. The proposed deep neural network, which features a bottleneck layer including 13 interpretable symptom domains, achieved an AUC of 0.62, thereby outperforming a set of feature-based classifiers while ensuring interpretability due to the bottleneck layer. We further illustrated the model's interpretability by examining how the predicted PROP domains influence final predictions using Shapley values. Emma C. L. Leschly, Oliver Roesler, Michael Neumann 0001, Jackson Liscombe, Abhishek Hosamath, Lakshmi Arbatti, Line Harder Clemmensen, Melanie Ganz-Benjaminsen, Vikram Ramanarayanan |
INTERSPEECH | 4 |
| 2025 | Accessible Real-time Eye-gaze Tracking for Neurocognitive Health Assessment: A Multimodal Web-based Approach
Daniel Tisdale, Jackson Liscombe, David Pautler, Michael Neumann 0001, Vikram Ramanarayanan |
INTERSPEECH | 2 |
| 2024 | How Consistent are Speech-Based Biomarkers in Remote Tracking of ALS Disease Progression Across Languages? A Case Study of English and DutchabstractPrevious work has demonstrated the utility of speech-based digital biomarkers for remotely tracking longitudinal progression in people with Amyotrophic Lateral Sclerosis (pALS). Here, we investigate the responsiveness of these biomarkers across languages for consistency. We collected audiovisual data using a cloud-based multimodal dialogue platform, where pALS interacted with a virtual guide to perform several speaking exercises. We automatically extracted speech, linguistic and orofacial metrics from 143 English-speaking pALS (36 bulbar onset, 107 non-bulbar onset) and 26 Dutch-speaking pALS (10 bulbar, 16 non-bulbar onset). We used growth curve models to estimate the trajectory of these metrics over time. We observe that for most of these metrics, English-speaking pALS and Dutch-speaking pALS follow similar trajectories, i.e. the slopes are not statistically different from each other, demonstrating the potential of such speech-based biomarkers for remote monitoring across languages. Hardik Kothare, Michael Neumann 0001, Cathy Zhang, Jackson Liscombe, Jordi W. J. van Unnik, Lianne C. M. Botman, Leonard H. van den Berg, Ruben P. A van Eijk, Vikram Ramanarayanan |
INTERSPEECH | 4 |
| 2024 | Multimodal Digital Biomarkers for Longitudinal Tracking of Speech Impairment Severity in ALS: An Investigation of Clinically Important Differences
Michael Neumann 0001, Hardik Kothare, Jackson Liscombe, Emma C. L. Leschly, Oliver Roesler, Vikram Ramanarayanan |
INTERSPEECH | 3 |
| 2024 | Towards Scalable Remote Assessment of Mild Cognitive Impairment Via Multimodal Dialog
Oliver Roesler, Jackson Liscombe, Michael Neumann 0001, Hardik Kothare, Abhishek Hosamath, Lakshmi Arbatti, Doug Habberstad, Christiane Suendermann-Oeft, Meredith Bartlett, Cathy Zhang, Nikhil Sukhdev, Kolja Wilms, Anusha Badathala, Sandrine Istas, Steve Ruhmel, Bryan Hansen, Madeline Hannan, David Henley, Arthur W. Wallace, Ira Shoulson, David Suendermann-Oeft, Vikram Ramanarayanan |
INTERSPEECH | 2 |
| 2023 | Responsiveness, Sensitivity and Clinical Utility of Timing-Related Speech Biomarkers for Remote Monitoring of ALS Disease Progressionabstract= 94). We further evaluated the sensitivity of speech metrics in tracking disease progression in pALS while their ALSFRS-R speech score remained unchanged at 3 out of a total possible score of 4. We observed that timing-related speech metrics showed significant longitudinal changes even after accounting for learning effects. The findings of this study have the potential to inform disease prognosis and functional outcomes of clinical trials. Hardik Kothare, Michael Neumann 0001, Jackson Liscombe, Jordan R. Green, Vikram Ramanarayanan |
INTERSPEECH | 3 |
| 2023 | When Words Speak Just as Loudly as Actions: Virtual Agent Based Remote Health Assessment Integrating What Patients Say with What They Do
Vikram Ramanarayanan, David Pautler, Lakshmi Arbatti, Abhishek Hosamath, Michael Neumann 0001, Hardik Kothare, Oliver Roesler, Jackson Liscombe, Andrew Cornish, Doug Habberstad, Vanessa Richter, David Suendermann-Oeft, Ira Shoulson |
INTERSPEECH | 8 |
| 2022 | Statistical and clinical utility of multimodal dialogue-based speech and facial metrics for Parkinson's disease assessment
Hardik Kothare, Michael Neumann 0001, Jackson Liscombe, Oliver Roesler, William Burke, Andrew Exner, Sandy Snyder, Andrew Cornish, Doug Habberstad, David Pautler, David Suendermann-Oeft, Jessica Huber, Vikram Ramanarayanan |
INTERSPEECH | 3 |
| 2021 | Investigating the Interplay Between Affective, Phonatory and Motoric Subsystems in Autism Spectrum Disorder Using a Multimodal Dialogue AgentabstractAbstract We explore the utility of an on-demand multimodal conversational platform in extracting speech and facial metrics in children with Autism Spectrum Disorder (ASD). We investigate the extent to which these metrics correlate with objective clinical measures, particularly as they pertain to the interplay be-tween the affective, phonatory and motoric subsystems. 22 participants diagnosed with ASD engaged with a virtual agent in conversational affect production tasks designed to elicit facial and vocal affect. We found significant correlations between vocal pitch and loudness extracted by our platform during these tasks and accuracy in recognition of facial and vocal affect, as-sessed via the Diagnostic Analysis of Nonverbal Accuracy-2 (DANVA-2) neuropsychological task. We also found significant correlations between jaw kinematic metrics extracted using our platform and motor speed of the dominant hand assessed via a standardised neuropsychological finger tapping task. These findings offer preliminary evidence for the usefulness of these audiovisual analytic metrics and could help us better model the interplay between different physiological subsystems in individuals with ASD. Hardik Kothare, Vikram Ramanarayanan, Oliver Roesler, Michael Neumann 0001, Jackson Liscombe, William Burke, Andrew Cornish, Doug Habberstad, Alaa Sakallah, Sara Markuson, Seemran Kansara, Afik Faerman, Yasmine Bensidi-Slimane, Laura Fry, Saige Portera, David Suendermann-Oeft, David Pautler, Carly Demopoulos |
Interspeech | 5 |
| 2021 | Investigating the Utility of Multimodal Conversational Technology and Audiovisual Analytic Measures for the Assessment and Monitoring of Amyotrophic Lateral Sclerosis at ScaleabstractWe propose a cloud-based multimodal dialog platform for the remote assessment and monitoring of Amyotrophic Lateral Sclerosis (ALS) at scale. This paper presents our vision, technology setup, and an initial investigation of the efficacy of the various acoustic and visual speech metrics automatically extracted by the platform. 82 healthy controls and 54 people with ALS (pALS) were instructed to interact with the platform and completed a battery of speaking tasks designed to probe the acoustic, articulatory, phonatory, and respiratory aspects of their speech. We find that multiple acoustic (rate, duration, voicing) and visual (higher order statistics of the jaw and lip) speech metrics show statistically significant differences between controls, bulbar symptomatic and bulbar pre-symptomatic patients. We report on the sensitivity and specificity of these metrics using five-fold cross-validation. We further conducted a LASSO-LARS regression analysis to uncover the relative contributions of various acoustic and visual features in predicting the severity of patients' ALS (as measured by their self-reported ALSFRS-R scores). Our results provide encouraging evidence of the utility of automatically extracted audiovisual analytics for scalable remote patient assessment and monitoring in ALS. Michael Neumann 0001, Oliver Roesler, Jackson Liscombe, Hardik Kothare, David Suendermann-Oeft, David Pautler, Indu Navar, Aria Anvar, Jochen Kumm, Raquel Norel, Ernest Fraenkel, Alexander V. Sherman, James D. Berry, Gary L. Pattee, Jun Wang 0037, Jordan R. Green, Vikram Ramanarayanan |
Interspeech | 3 |
| 2020 | Toward Remote Patient Monitoring of Speech, Video, Cognitive and Respiratory Biomarkers Using Multimodal Dialog Technology
Vikram Ramanarayanan, Oliver Roesler, Michael Neumann 0001, David Pautler, Doug Habberstad, Andrew Cornish, Hardik Kothare, Vignesh Murali, Jackson Liscombe, Dirk Schnelle-Walka, Patrick L. Lange, David Suendermann-Oeft |
INTERSPEECH | 9 |
| 2019 | NEMSI: A Multimodal Dialog System for Screening of Neurological or Mental ConditionsabstractWe present NEMSI, a cloud-based multimodal dialog system designed to have naturalistic interactions with individuals for the purpose of screening neurological or mental conditions. The system has been used by thousands of people capturing audio and video responses to open-ended questions and structured health surveys. David Suendermann-Oeft, Amanda Robinson 0002, Andrew Cornish, Doug Habberstad, David Pautler, Dirk Schnelle-Walka, Franziska Haller, Jackson Liscombe, Michael Neumann 0001, Mike Merrill, Oliver Roesler, Renko Geffarth |
IVA | 8 |
| 2011 | Large-Scale Experiments on Data-Driven Design of Commercial Spoken Dialog SystemsabstractThe design of commercial spoken dialog systems is most commonly based on hand-crafting call flows. Voice interaction designers write prompts, predict caller responses, set speech recognition parameters, implement interaction strategies, all based on “best design practices”. Recently, we presented the mathematical framework “Contender” (similar to reinforcement learning) that allows for replacing manual decisions made during system design by data-driven soft decisions made at system run time optimizing the cumulative reward of an application. The current paper reports on the results of 26 Contenders implemented in commercial applications processing a total of about 15 million calls. David Suendermann-Oeft, Jackson Liscombe, Jonathan Bloom, Grace Li, Roberto Pieraccini |
INTERSPEECH | 2 |
| 2010 | Optimize the obvious: Automatic call flow generationabstractIn commercial spoken dialog systems, call flows are built by call flow designers implementing a predefined business logic. While it may appear obvious from this logic how the call flow has to look like, i.e., which pieces of information have to be gathered from the caller or back-end systems and in which sequence, there are, in fact, strong arguments for automating call flow generation: 1) manual generation is time-consuming 2) manual generation is suboptimal and error-prone 3) automatic generation can react on dynamically changing business logic or external factors such as the distribution of callers and call reasons This paper presents a method for automatically deriving a call flow minimizing the average number of user turns given a business logic and a frequency distribution of call reasons. As an example, we applied the method to a call routing application whose manually built call flow is processing about 4 million calls per month and whose call reason distribution served to measure the impact of the automatic call flow generation. David Suendermann-Oeft, Jackson Liscombe, Roberto Pieraccini |
ICASSP | 2 |
| 2010 | Is it possible to predict task completion in automated troubleshooters?abstractThede online prediction of task success in Interactive Voice Response (IVR) systems is a comparatively new field of research. It helps to identify problemantic calls and enables the dialog system to react before the caller gets overly frustrated. This publication investigates, to which extent it is possible to predict task completion and how existing approaches generalize for long dialogs. We compare the performance of two different modeling techniques: linear modeling and n-gram modeling. We show that n-gram modeling outperforms linear modeling significantly at later prediction points. From a comprehensive set of interaction parameters, we identify the relevant ones using the Information Gain Ratio. New interaction parameters are presented and evaluated. The study is based on 41,422 calls from an automated Internet troubleshooter with an average of 21.4 turns per call. Alexander Schmitt, Wolfgang Minker, Jackson Liscombe, David Suendermann-Oeft |
INTERSPEECH | 4 |
| 2010 | Minimally invasive surgery for spoken dialog systemsabstractWe demonstrate three techniques (Escalator, Engager, and EverywhereContender) designed to optimize performance of commercial spoken dialog systems. These techniques have in common that they produce very small or no negative performance impact even during a potential experimental phase. This is because they can either be applied offline to data collected on a deployed system, or they can be incorporated conservatively such that only a low percentage of calls will get affected until the optimal strategy becomes apparent. David Suendermann-Oeft, Jackson Liscombe, Roberto Pieraccini |
INTERSPEECH | 2 |
| 2010 | WITcHCRafT: A Workbench for Intelligent exploraTion of Human ComputeR conversaTions
Alexander Schmitt, Gregor Bertrand, Tobias Heinroth, Wolfgang Minker, Jackson Liscombe |
LREC | 5 |
| 2010 | The Influence of the Utterance Length on the Recognition of Aged Voices
Alexander Schmitt, Tim Polzehl, Wolfgang Minker, Jackson Liscombe |
LREC | 4 |
| 2010 | How to Drink from a Fire Hose: One Person Can Annoscribe One Million Utterances in One Month
David Suendermann-Oeft, Jackson Liscombe, Roberto Pieraccini |
SIGDIAL Conference | 2 |
| 2010 | ContenderabstractContender (or what the academic community would refer to as a light version of reinforcement learning) is a simple technique to experiment with a number of competing paths in a (commercial) spoken dialog system. By randomly routing certain portions of traffic to individual paths and computing average rewards for each of the routes, the goal is to find out which one performs best. This paper is to do away with common uncertainties on how to set up contender weights, how much data needs to be accumulated to draw reliable conclusions, and how this all relates to the notion of statistical significance. David Suendermann-Oeft, Jackson Liscombe, Roberto Pieraccini |
SLT | 2 |
| 2009 | From rule-based to statistical grammars: Continuous improvement of large-scale spoken dialog systemsabstractStatistical Spoken Language Understanding grammars (SSLUs) are often used only at the top recognition contexts of modern large-scale spoken dialog systems. We propose to use SSLUs at every recognition context in a dialog system, effectively replacing conventional, manually written grammars. Furthermore, we present a methodology of continuous improvement in which data are collected at every recognition context over an entire dialog system. These data are then used to automatically generate updated context-specific SSLUs at regular intervals and, in so doing, continually improve system performance over time. We have found that SSLUs significantly and consistently outperform even the most carefully designed rule-based grammars in a wide range of contexts in a corpus of over two million utterances collected for a complex call-routing and troubleshooting dialog system. David Suendermann-Oeft, Keelan Evanini, Jackson Liscombe, Phillip Hunter, Krishna Dayanidhi, Roberto Pieraccini |
ICASSP | 3 |
| 2009 | Localization of speech recognition in spoken dialog systems: how machine translation can make our lives easierabstractThe localization of speech recognition for large-scale spoken dialog systems can be a tremendous exercise. Usually, all in- volved grammars have to be translated by a language expert, and new data has to be collected, transcribed, and annotated for statistical utterance classifiers resulting in a time-consuming and expensive undertaking. Often though, a vast number of transcribed and annotated utterances exists for the source lan- guage. In this paper, we propose to use such data and translate it into the target language using machine translation. The trans- lated utterances and their associated (original) annotations are then used to train statistical grammars for all contexts of the target system. As an example, we localize an English spoken dialog system for Internet troubleshooting to Spanish by trans- lating more than 4 million source utterances without any human intervention. In an application of the localized system to more than 10,000 utterances collected on a similar Spanish Internet troubleshooting system, we show that the overall accuracy was only 5.7% worse than that of the English source system. Index Terms: spoken dialog systems, machine translation, lo- calization David Suendermann-Oeft, Jackson Liscombe, Krishna Dayanidhi, Roberto Pieraccini |
INTERSPEECH | 2 |
| 2009 | On NoMatchs, NoInputs and BargeIns: Do Non-Acoustic Features Support Anger Detection?
Alexander Schmitt, Tobias Heinroth, Jackson Liscombe |
SIGDIAL Conference | 3 |
| 2009 | A Handsome Set of Metrics to Measure Utterance Classification Performance in Spoken Dialog Systems
David Suendermann-Oeft, Jackson Liscombe, Krishna Dayanidhi, Roberto Pieraccini |
SIGDIAL Conference | 2 |
| 2008 | When calls go wrong: how to detect problematic calls based on log-files and emotions?
Ota Herm, Alexander Schmitt, Jackson Liscombe |
INTERSPEECH | 3 |
| 2008 | Caller Experience: A method for evaluating dialog systems and its automatic predictionabstractIn this paper we introduce a subjective metric for evaluating the performance of spoken dialog systems, caller experience (CE). CE is a useful metric for tracking the overall performance of a system in deployment, as well as for isolating individual problematic calls in which the system underperforms. The proposed CE metric differs from most performance evaluation metrics proposed in the past in that it is a) a subjective, qualitative rating of the call, and b) provided by expert, external listeners, not the callers themselves. The results of an experiment in which a set of human experts listened to the same calls three times are presented. The fact that these results show a high level of agreement among different listeners, despite the subjective nature of the task, demonstrates the validity of using CE as a standard metric. Finally, an automated rating system using objective measures is shown to perform at the same high level as the humans. This is an important advance, since it provides a way to reduce the human labor costs associated with producing a reliable CE. Keelan Evanini, Phillip Hunter, Jackson Liscombe, David Suendermann-Oeft, Krishna Dayanidhi, Roberto Pieraccini |
SLT | 3 |
| 2008 | C5abstractThe annotation of hundreds of thousands of utterances for the training of statistical utterance classifiers requires a careful quality assurance procedure to make the data consistent and reliable. In this paper, we present five methods to analyze different aspects of annotated data to ensure their Completeness, Consistency, Correlation, Congruence and to avoid Confusion-collectively referred to as C5. David Suendermann-Oeft, Jackson Liscombe, Keelan Evanini, Krishna Dayanidhi, Roberto Pieraccini |
SLT | 2 |
| 2006 | Detecting question-bearing turns in spoken tutorial dialoguesabstractCurrent speech-enabled Intelligent Tutoring Systems do not model student question behavior the way human tutors do, despite evidence indicating the importance of doing so.Our study examined a corpus of spoken tutorial dialogues collected for development of ITSpoke, an Intelligent Tutoring Spoken Dialogue System.The authors extracted prosodic, lexical, syntactic, and student and task dependent information from student turns.Results of running 5-fold cross validation machine learning experiments using AdaBoosted C4.5 decision trees show prediction of student question-bearing turns at a rate of 79.7%.The most useful features were prosodic, especially the pitch slope of the last 200 milliseconds of the student turn.Student pre-test score was the most-used feature.Findings indicate that using turn-based units is acceptable for incorporating question detection capability into practical Intelligent Tutoring Systems. Jackson Liscombe, Jennifer J. Venditti, Julia Hirschberg |
INTERSPEECH | 1 |
| 2006 | Intonational cues to student questions in tutoring dialogsabstractSuccessful Intelligent Tutoring Systems (ITSs) must be able to recognize when their students are asking a question.They must identify question form as well as function in order to respond appropriately.Our study examines whether intonational features, specifically, F0 height and rise range, are useful cues to student question type in a corpus of 643 American English questions.Results show a quantitative effect of both form and function.In addition, among clarification-seeking questions, we observed differences based on the type of clarification being sought. 1 Jennifer J. Venditti, Julia Hirschberg, Jackson Liscombe |
INTERSPEECH | 3 |
| 2006 | Detecting Emotion in Speech: Experiments in Three Domains
Jackson Liscombe |
HLT-NAACL | 1 |
| 2005 | Detecting certainness in spoken tutorial dialoguesabstractWhat role does affect play in spoken tutorial systems and is it automatically detectable? We investigated the classification of student certainness in a corpus collected for ITSPOKE, a speech-enabled Intelligent Tutorial System (ITS). Our study suggests that tutors respond to indications of student uncertainty differently from student certainty. Results of machine learning experiments indicate that acoustic-prosodic features can distinguish student certainness from other student states. A combination of acoustic-prosodic features extracted at two levels of intonational analysis --- breath groups and turns --- achieves 76.42% classification accuracy, a 15.8% relative improvement over baseline performance. Our results suggest that student certainness can be automatically detected and utilized to create better spoke dialog ITSs. Jackson Liscombe, Julia Hirschberg, Jennifer J. Venditti |
INTERSPEECH | 1 |
| 2005 | Using context to improve emotion detection in spoken dialog systemsabstractMost research that explores the emotional state of users of spoken dialog systems does not fully utilize the contextual nature that the dialog structure provides.This paper reports results of machine learning experiments designed to automatically classify the emotional state of user turns using a corpus of 5,690 dialogs collected with the "How May I Help You SM " spoken dialog system.We show that augmenting standard lexical and prosodic features with contextual features that exploit the structure of spoken dialog and track user state increases classification accuracy by 2.6%. Jackson Liscombe, Giuseppe Riccardi, Dilek Hakkani-Tür |
INTERSPEECH | 1 |
| 2003 | Classifying subject ratings of emotional speech using acoustic featuresabstractThis paper presents results from a study examining emotional speech using acoustic features and their use in automatic machine learning classification.In addition, we propose a classification scheme for the labeling of emotions on continuous scales.Our findings support those of previous research as well as indicate possible future directions utilizing spectral tilt and pitch contour to distinguish emotions in the valence dimension. Emotion Recognition Survey: Sound File 1 of 47not at all a little somewhat quite extremely How frustrated does this person sound?How confident does this person sound?How interested does this person sound?How sad does this person sound?How happy does this person sound?How friendly does this person sound?How angry does this person sound?How anxious does this person sound?How bored does this person sound?How encouraging does this person sound? Jackson Liscombe, Jennifer J. Venditti, Julia Hirschberg |
INTERSPEECH | 1 |