Vikram Ramanarayanan

dblp:33/9231 · DBLP profile ↗
← Back
63ranked-venue papers
21as first author
18since 2021 · last 2025
0000-0001-7810-2769ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 18 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 13 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Multimodal Speech-Based Biomarkers Outperform the ALS Functional Rating Scale in Predicting Individual Disease Progression in ALS
Hardik Kothare, Michael Neumann 0001, Vikram Ramanarayanan
INTERSPEECH3
2025 An Exploration of Interpretable Deep Learning Models for the Assessment of Mild Cognitive Impairment
abstract
Early diagnosis and intervention are crucial for mild cognitive impairment (MCI), as MCI often progresses to more severe neurodegenerative conditions. In this study, we explore utilizing deep learning for MCI detection without loosing the interpretability provided by feature-based approaches. We used a dataset consisting of 90 MCI patients and 91 controls collected via a remote assessment platform and analyzed the participants' spontaneous speech responses to the Patient Report of Problems (PROP) which asks patients to report their most bothersome general health problems. The proposed deep neural network, which features a bottleneck layer including 13 interpretable symptom domains, achieved an AUC of 0.62, thereby outperforming a set of feature-based classifiers while ensuring interpretability due to the bottleneck layer. We further illustrated the model's interpretability by examining how the predicted PROP domains influence final predictions using Shapley values.
Emma C. L. Leschly, Oliver Roesler, Michael Neumann 0001, Jackson Liscombe, Abhishek Hosamath, Lakshmi Arbatti, Line Harder Clemmensen, Melanie Ganz-Benjaminsen, Vikram Ramanarayanan
INTERSPEECH9
2025 Multimodal Speech, Language and Orofacial Analysis for Remote Assessment of Positive, Negative and Cognitive Symptoms in Schizophrenia
Michael Neumann 0001, Hardik Kothare, Beverly Insel, Anzalee Khan, Danyah Nadim, Jean-Pierre Lindenmayer, Vikram Ramanarayanan
INTERSPEECH7
2025 Accessible Real-time Eye-gaze Tracking for Neurocognitive Health Assessment: A Multimodal Web-based Approach
Daniel Tisdale, Jackson Liscombe, David Pautler, Michael Neumann 0001, Vikram Ramanarayanan
INTERSPEECH5
2024 Preliminary Investigation of Psychometric Properties of a Novel Multimodal Dialog Based Affect Production Task in Children and Adolescents with Autism
Carly Demopoulos, Linnea Lampinen, Cristian Preciado, Hardik Kothare, Vikram Ramanarayanan
INTERSPEECH5
2024 How Consistent are Speech-Based Biomarkers in Remote Tracking of ALS Disease Progression Across Languages? A Case Study of English and Dutch
abstract
Previous work has demonstrated the utility of speech-based digital biomarkers for remotely tracking longitudinal progression in people with Amyotrophic Lateral Sclerosis (pALS). Here, we investigate the responsiveness of these biomarkers across languages for consistency. We collected audiovisual data using a cloud-based multimodal dialogue platform, where pALS interacted with a virtual guide to perform several speaking exercises. We automatically extracted speech, linguistic and orofacial metrics from 143 English-speaking pALS (36 bulbar onset, 107 non-bulbar onset) and 26 Dutch-speaking pALS (10 bulbar, 16 non-bulbar onset). We used growth curve models to estimate the trajectory of these metrics over time. We observe that for most of these metrics, English-speaking pALS and Dutch-speaking pALS follow similar trajectories, i.e. the slopes are not statistically different from each other, demonstrating the potential of such speech-based biomarkers for remote monitoring across languages.
Hardik Kothare, Michael Neumann 0001, Cathy Zhang, Jackson Liscombe, Jordi W. J. van Unnik, Lianne C. M. Botman, Leonard H. van den Berg, Ruben P. A van Eijk, Vikram Ramanarayanan
INTERSPEECH9
2024 Multimodal Digital Biomarkers for Longitudinal Tracking of Speech Impairment Severity in ALS: An Investigation of Clinically Important Differences
Michael Neumann 0001, Hardik Kothare, Jackson Liscombe, Emma C. L. Leschly, Oliver Roesler, Vikram Ramanarayanan
INTERSPEECH6
2024 Towards Scalable Remote Assessment of Mild Cognitive Impairment Via Multimodal Dialog
Oliver Roesler, Jackson Liscombe, Michael Neumann 0001, Hardik Kothare, Abhishek Hosamath, Lakshmi Arbatti, Doug Habberstad, Christiane Suendermann-Oeft, Meredith Bartlett, Cathy Zhang, Nikhil Sukhdev, Kolja Wilms, Anusha Badathala, Sandrine Istas, Steve Ruhmel, Bryan Hansen, Madeline Hannan, David Henley, Arthur W. Wallace, Ira Shoulson, David Suendermann-Oeft, Vikram Ramanarayanan
INTERSPEECH22
2024 Bayesian inference of state feedback control parameters for fo perturbation responses in cerebellar ataxia
abstract
Behavioral speech tasks have been widely used to understand the mechanisms of speech motor control in typical speakers as well as in various clinical populations. However, determining which neural functions differ between typical speakers and clinical populations based on behavioral data alone is difficult because multiple mechanisms may lead to the same behavioral differences. For example, individuals with cerebellar ataxia (CA) produce atypically large compensatory responses to pitch perturbations in their auditory feedback, compared to typical speakers, but this pattern could have many explanations. Here, computational modeling techniques were used to address this challenge. Bayesian inference was used to fit a state feedback control (SFC) model of voice fundamental frequency (fo) control to the behavioral pitch perturbation responses of speakers with CA and typical speakers. This fitting process resulted in estimates of posterior likelihood distributions for five model parameters (sensory feedback delays, absolute and relative levels of auditory and somatosensory feedback noise, and controller gain), which were compared between the two groups. Results suggest that the speakers with CA may proportionally weight auditory and somatosensory feedback differently from typical speakers. Specifically, the CA group showed a greater relative sensitivity to auditory feedback than the control group. There were also large group differences in the controller gain parameter, suggesting increased motor output responses to target errors in the CA group. These modeling results generate hypotheses about how CA may affect the speech motor system, which could help guide future empirical investigations in CA. This study also demonstrates the overall proof-of-principle of using this Bayesian inference approach to understand behavioral speech data in terms of interpretable parameters of speech motor control models.
Jessica L. Gaines, Kwang S. Kim, Benjamin Parrell, Vikram Ramanarayanan, Alvincé L. Pongos, Srikantan S. Nagarajan, John F. Houde
PLoS Comput. Biol.4
2023 Responsiveness, Sensitivity and Clinical Utility of Timing-Related Speech Biomarkers for Remote Monitoring of ALS Disease Progression
abstract
= 94). We further evaluated the sensitivity of speech metrics in tracking disease progression in pALS while their ALSFRS-R speech score remained unchanged at 3 out of a total possible score of 4. We observed that timing-related speech metrics showed significant longitudinal changes even after accounting for learning effects. The findings of this study have the potential to inform disease prognosis and functional outcomes of clinical trials.
Hardik Kothare, Michael Neumann 0001, Jackson Liscombe, Jordan R. Green, Vikram Ramanarayanan
INTERSPEECH5
2023 A Multimodal Investigation of Speech, Text, Cognitive and Facial Video Features for Characterizing Depression With and Without Medication
Michael Neumann 0001, Hardik Kothare, Doug Habberstad, Vikram Ramanarayanan
INTERSPEECH4
2023 Combining Multiple Multimodal Speech Features into an Interpretable Index Score for Capturing Disease Progression in Amyotrophic Lateral Sclerosis
abstract
Multiple speech biomarkers have been shown to carry useful information regarding Amyotrophic Lateral Sclerosis (ALS) pathology. We propose a two-step framework to compute optimal linear combinations (indexes) of these biomarkers that are more discriminative and noise-robust than the individual markers, which is important for clinical care and pharmaceutical trial applications. First, we use a hierarchical clustering based method to select representative speech metrics from a dataset comprising 143 people with ALS and 135 age- and sex-matched healthy controls. Second, we analyze three methods of index computation that optimize linear discriminability, Youden Index, and sparsity of logistic regression model weights, respectively, and evaluate their performance with 5-fold cross validation. We find that the proposed indexes are generally more discriminative of bulbar vs non-bulbar onset in ALS than their individual component metrics as well as an equally-weighted baseline.
Michael Neumann 0001, Hardik Kothare, Vikram Ramanarayanan
INTERSPEECH3
2023 When Words Speak Just as Loudly as Actions: Virtual Agent Based Remote Health Assessment Integrating What Patients Say with What They Do
Vikram Ramanarayanan, David Pautler, Lakshmi Arbatti, Abhishek Hosamath, Michael Neumann 0001, Hardik Kothare, Oliver Roesler, Jackson Liscombe, Andrew Cornish, Doug Habberstad, Vanessa Richter, David Suendermann-Oeft, Ira Shoulson
INTERSPEECH1
2023 Remote Assessment for ALS using Multimodal Dialog Agents: Data Quality, Feasibility and Task Compliance
abstract
We investigate the feasibility, task compliance and audiovisual data quality of a multimodal dialog-based solution for remote assessment of Amyotrophic Lateral Sclerosis (ALS). 53 people with ALS and 52 healthy controls interacted with Tina, a cloud-based conversational agent, in performing speech tasks designed to probe various aspects of motor speech function while their audio and video was recorded. We rated a total of 250 recordings for audio/video quality and participant task compliance, along with the relative frequency of different issues observed. We observed excellent compliance (98%) and audio (95.2%) and visual quality rates (84.8%), resulting in an overall yield of 80.8% recordings that were both compliant and of high quality. Furthermore, recording quality and compliance were not affected by level of speech severity and did not differ significantly across end devices. These findings support the utility of dialog systems for remote monitoring of speech in ALS.
Vanessa Richter, Michael Neumann 0001, Jordan R. Green, Brian Richburg, Oliver Roesler, Hardik Kothare, Vikram Ramanarayanan
INTERSPEECH7
2023 Mechanisms of sensorimotor adaptation in a hierarchical state feedback control model of speech
abstract
Upon perceiving sensory errors during movements, the human sensorimotor system updates future movements to compensate for the errors, a phenomenon called sensorimotor adaptation. One component of this adaptation is thought to be driven by sensory prediction errors-discrepancies between predicted and actual sensory feedback. However, the mechanisms by which prediction errors drive adaptation remain unclear. Here, auditory prediction error-based mechanisms involved in speech auditory-motor adaptation were examined via the feedback aware control of tasks in speech (FACTS) model. Consistent with theoretical perspectives in both non-speech and speech motor control, the hierarchical architecture of FACTS relies on both the higher-level task (vocal tract constrictions) as well as lower-level articulatory state representations. Importantly, FACTS also computes sensory prediction errors as a part of its state feedback control mechanism, a well-established framework in the field of motor control. We explored potential adaptation mechanisms and found that adaptive behavior was present only when prediction errors updated the articulatory-to-task state transformation. In contrast, designs in which prediction errors updated forward sensory prediction models alone did not generate adaptation. Thus, FACTS demonstrated that 1) prediction errors can drive adaptation through task-level updates, and 2) adaptation is likely driven by updates to task-level control rather than (only) to forward predictive models. Additionally, simulating adaptation with FACTS generated a number of important hypotheses regarding previously reported phenomena such as identifying the source(s) of incomplete adaptation and driving factor(s) for changes in the second formant frequency during adaptation to the first formant perturbation. The proposed model design paves the way for a hierarchical state feedback control framework to be examined in the context of sensorimotor adaptation in both speech and non-speech effector systems.
Kwang S. Kim, Jessica L. Gaines, Benjamin Parrell, Vikram Ramanarayanan, Srikantan S. Nagarajan, John F. Houde
PLoS Comput. Biol.4
2022 Statistical and clinical utility of multimodal dialogue-based speech and facial metrics for Parkinson's disease assessment
Hardik Kothare, Michael Neumann 0001, Jackson Liscombe, Oliver Roesler, William Burke, Andrew Exner, Sandy Snyder, Andrew Cornish, Doug Habberstad, David Pautler, David Suendermann-Oeft, Jessica Huber, Vikram Ramanarayanan
INTERSPEECH13
2021 Investigating the Interplay Between Affective, Phonatory and Motoric Subsystems in Autism Spectrum Disorder Using a Multimodal Dialogue Agent
abstract
Abstract We explore the utility of an on-demand multimodal conversational platform in extracting speech and facial metrics in children with Autism Spectrum Disorder (ASD). We investigate the extent to which these metrics correlate with objective clinical measures, particularly as they pertain to the interplay be-tween the affective, phonatory and motoric subsystems. 22 participants diagnosed with ASD engaged with a virtual agent in conversational affect production tasks designed to elicit facial and vocal affect. We found significant correlations between vocal pitch and loudness extracted by our platform during these tasks and accuracy in recognition of facial and vocal affect, as-sessed via the Diagnostic Analysis of Nonverbal Accuracy-2 (DANVA-2) neuropsychological task. We also found significant correlations between jaw kinematic metrics extracted using our platform and motor speed of the dominant hand assessed via a standardised neuropsychological finger tapping task. These findings offer preliminary evidence for the usefulness of these audiovisual analytic metrics and could help us better model the interplay between different physiological subsystems in individuals with ASD.
Hardik Kothare, Vikram Ramanarayanan, Oliver Roesler, Michael Neumann 0001, Jackson Liscombe, William Burke, Andrew Cornish, Doug Habberstad, Alaa Sakallah, Sara Markuson, Seemran Kansara, Afik Faerman, Yasmine Bensidi-Slimane, Laura Fry, Saige Portera, David Suendermann-Oeft, David Pautler, Carly Demopoulos
Interspeech2
2021 Investigating the Utility of Multimodal Conversational Technology and Audiovisual Analytic Measures for the Assessment and Monitoring of Amyotrophic Lateral Sclerosis at Scale
abstract
We propose a cloud-based multimodal dialog platform for the remote assessment and monitoring of Amyotrophic Lateral Sclerosis (ALS) at scale. This paper presents our vision, technology setup, and an initial investigation of the efficacy of the various acoustic and visual speech metrics automatically extracted by the platform. 82 healthy controls and 54 people with ALS (pALS) were instructed to interact with the platform and completed a battery of speaking tasks designed to probe the acoustic, articulatory, phonatory, and respiratory aspects of their speech. We find that multiple acoustic (rate, duration, voicing) and visual (higher order statistics of the jaw and lip) speech metrics show statistically significant differences between controls, bulbar symptomatic and bulbar pre-symptomatic patients. We report on the sensitivity and specificity of these metrics using five-fold cross-validation. We further conducted a LASSO-LARS regression analysis to uncover the relative contributions of various acoustic and visual features in predicting the severity of patients' ALS (as measured by their self-reported ALSFRS-R scores). Our results provide encouraging evidence of the utility of automatically extracted audiovisual analytics for scalable remote patient assessment and monitoring in ALS.
Michael Neumann 0001, Oliver Roesler, Jackson Liscombe, Hardik Kothare, David Suendermann-Oeft, David Pautler, Indu Navar, Aria Anvar, Jochen Kumm, Raquel Norel, Ernest Fraenkel, Alexander V. Sherman, James D. Berry, Gary L. Pattee, Jun Wang 0037, Jordan R. Green, Vikram Ramanarayanan
Interspeech17
2020 Effect of Modality on Human and Machine Scoring of Presentation Videos
abstract
We investigate the effect of observed data modality on human and machine scoring of informative presentations in the context of oral English communication training and assessment. Three sets of raters scored the content of three minute presentations by college students on the basis of either the video, the audio or the text transcript using a custom scoring rubric. We find significant differences between the scores assigned when raters view a transcript or listen to audio recordings in comparison to watching a video of the same presentation, and present an analysis of those differences. Using the human scores, we train machine learning models to score a given presentation using text, audio, and video features separately. We analyze the distribution of machine scores against the modality and label bias we observe in human scores, discuss its implications for machine scoring and recommend best practices for future work in this direction. Our results demonstrate the importance of checking and correcting for bias across different modalities in evaluations of multi-modal performances.
Haley Lepp, Chee Wee Leong, Katrina Roohr, Michelle P. Martin-Raugh, Vikram Ramanarayanan
ICMI5
2020 Design and Development of a Human-Machine Dialog Corpus for the Automated Assessment of Conversational English Proficiency
Vikram Ramanarayanan
INTERSPEECH1
2020 Toward Remote Patient Monitoring of Speech, Video, Cognitive and Respiratory Biomarkers Using Multimodal Dialog Technology
Vikram Ramanarayanan, Oliver Roesler, Michael Neumann 0001, David Pautler, Doug Habberstad, Andrew Cornish, Hardik Kothare, Vignesh Murali, Jackson Liscombe, Dirk Schnelle-Walka, Patrick L. Lange, David Suendermann-Oeft
INTERSPEECH1
2019 Native Language Identification from Raw Waveforms Using Deep Convolutional Neural Networks with Attentive Pooling
abstract
Automatic detection of an individual's native language (L1) based on speech data from their second language (L2) can be useful for informing a variety of speech applications such as automatic speech recognition (ASR), speaker recognition, voice biometrics, and computer assisted language learning (CALL). Previously proposed systems for native language identification from L2 acoustic signals rely on traditional feature extraction pipelines to extract relevant features such as mel-filterbanks, cepstral coefficients, i-vectors, etc. In this paper, we present a fully convolutional neural network approach that is trained end-to-end to predict the native language of the speaker directly from the raw waveforms, thereby removing the feature extraction step altogether. Experimental results using this approach on a database of 11 different L1s suggest that the learnable convolutional layers of our proposed attention-based end-to-end model extract meaningful features from raw waveforms. Further, the attentive pooling mechanism in our proposed network enables our model to focus on the most discriminative features leading to improvements over the conventional baseline.
Rutuja Ubale, Vikram Ramanarayanan, Yao Qian, Keelan Evanini, Chee Wee Leong, Chong Min Lee
ASRU2
2019 Scoring Interactional Aspects of Human-Machine Dialog for Language Learning and Assessment using Text Features
abstract
While there has been much work in the language learning and assessment literature on human and automated scoring of essays and short constructed responses, there is little to no work examining text features for scoring of dialog data, particularly interactional aspects thereof, to assess conversational proficiency over and above constructed response skills.Our work bridges this gap by investigating both human and automated approaches towards scoring human-machine text dialog in the context of a real-world language learning application.We collected conversational data of human learners interacting with a cloud-based standards-compliant dialog system, triple-scored these data along multiple dimensions of conversational proficiency, and then analyzed the performance trends.We further examined two different approaches to automated scoring of such data and show that these approaches are able to perform at or above par with human agreement for a majority of dimensions of the scoring rubric.
Vikram Ramanarayanan, Matthew Mulholland, Yao Qian
SIGdial1
2019 The FACTS model of speech motor control: Fusing state estimation and task-based control
abstract
We present a new computational model of speech motor control: the Feedback-Aware Control of Tasks in Speech or FACTS model. FACTS employs a hierarchical state feedback control architecture to control simulated vocal tract and produce intelligible speech. The model includes higher-level control of speech tasks and lower-level control of speech articulators. The task controller is modeled as a dynamical system governing the creation of desired constrictions in the vocal tract, after Task Dynamics. Both the task and articulatory controllers rely on an internal estimate of the current state of the vocal tract to generate motor commands. This estimate is derived, based on efference copy of applied controls, from a forward model that predicts both the next vocal tract state as well as expected auditory and somatosensory feedback. A comparison between predicted feedback and actual feedback is then used to update the internal state prediction. FACTS is able to qualitatively replicate many characteristics of the human speech system: the model is robust to noise in both the sensory and motor pathways, is relatively unaffected by a loss of auditory feedback but is more significantly impacted by the loss of somatosensory feedback, and responds appropriately to externally-imposed alterations of auditory and somatosensory feedback. The model also replicates previously hypothesized trade-offs between reliance on auditory and somatosensory feedback and shows for the first time how this relationship may be mediated by acuity in each sensory domain. These results have important implications for our understanding of the speech motor control system in humans.
Benjamin Parrell, Vikram Ramanarayanan, Srikantan S. Nagarajan, John F. Houde
PLoS Comput. Biol.2
2018 Improvements to an Automated Content Scoring System for Spoken CALL Responses: the ETS Submission to the Second Spoken CALL Shared Task
Keelan Evanini, Matthew Mulholland, Rutuja Ubale, Yao Qian, Robert A. Pugh, Vikram Ramanarayanan, Aoife Cahill
INTERSPEECH6
2018 Game-based Spoken Dialog Language Learning Applications for Young Students
Keelan Evanini, Veronika Timpe-Laughlin, Eugene Tsuprun, Ian Blood, Jeremy Lee, James V. Bruno, Vikram Ramanarayanan, Patrick L. Lange, David Suendermann-Oeft
INTERSPEECH7
2018 FACTS: A Hierarchical Task-based Control Model of Speech Incorporating Sensory Feedback
Benjamin Parrell, Vikram Ramanarayanan, Srikantan S. Nagarajan, John F. Houde
INTERSPEECH2
2018 Toward Scalable Dialog Technology for Conversational Language Learning: Case Study of the TOEFL® MOOC
Vikram Ramanarayanan, David Pautler, Patrick L. Lange, Eugene Tsuprun, Rutuja Ubale, Keelan Evanini, David Suendermann-Oeft
INTERSPEECH1
2018 Leveraging Multimodal Dialog Technology for the Design of Automated and Interactive Student Agents for Teacher Training
abstract
We present a paradigm for interactive teacher training that leverages multimodal dialog technology to puppeteer customdesigned embodied conversational agents (ECAs) in student roles.We used the open-source multimodal dialog system HALEF to implement a small-group classroom math discussion involving Venn diagrams where a human teacher candidate has to interact with two student ECAs whose actions are controlled by the dialog system.Such an automated paradigm has the potential to be extended and scaled to a wide range of interactive simulation scenarios in education, medicine, and business where group interaction training is essential.
David Pautler, Vikram Ramanarayanan, Kirby Cofino, Patrick L. Lange, David Suendermann-Oeft
SIGDIAL Conference2
2018 Automatic Token and Turn Level Language Identification for Code-Switched Text Dialog: An Analysis Across Language Pairs and Corpora
abstract
We examine the efficacy of various feature-learner combinations for language identification in different types of text-based code-switched interactionshuman-human dialog, human-machine dialog, as well as monolog -at both the token and turn levels.In order to examine the generalization of such methods across language pairs and datasets, we analyze ten different datasets of code-switched text.We extract a variety of character-and word-based text features and pass them into multiple learners, including conditional random fields, logistic regressors, and recurrent neural networks.We further examine the efficacy of character-level embedding and GloVe features in improving performance and observe that our best-performing text system significantly outperforms the majority vote baseline across language pairs and datasets.
Vikram Ramanarayanan, Robert Pugh
SIGDIAL Conference1
2018 Analysis of speech production real-time MRI
Vikram Ramanarayanan, Sam Tilsen, Michael I. Proctor, Johannes Töger, Louis Goldstein, Krishna S. Nayak, Shri Narayanan
Comput. Speech Lang.1
2018 Acoustic Denoising Using Dictionary Learning With Spectral and Temporal Regularization
abstract
We present a method for speech enhancement of data collected in extremely noisy environments, such as those obtained during magnetic resonance imaging (MRI) scans. We propose an algorithm based on dictionary learning to perform this enhancement. We use complex nonnegative matrix factorization with intra-source additivity (CMF-WISA) to learn dictionaries of the noise and speech+noise portions of the data and use these to factor the noisy spectrum into estimated speech and noise components. We augment the CMF-WISA cost function with spectral and temporal regularization terms to improve the noise modeling. Based on both objective and subjective assessments, we find that our algorithm significantly outperforms traditional techniques such as Least Mean Squares (LMS) filtering, while not requiring prior knowledge or specific assumptions such as periodicity of the noise waveforms that current state-of-the-art algorithms require.
Colin Vaz, Vikram Ramanarayanan, Shri Narayanan
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Exploring ASR-free end-to-end modeling to improve spoken language understanding in a cloud-based dialog system
abstract
Spoken language understanding (SLU) in dialog systems is generally performed using a natural language understanding (NLU) model based on the hypotheses produced by an automatic speech recognition (ASR) system. However, when new spoken dialog applications are built from scratch in real user environments that often have sub-optimal audio characteristics, ASR performance can suffer due to factors such as the paucity of training data or a mismatch between the training and test data. To address this issue, this paper proposes an ASR-free, end-to-end (E2E) modeling approach to SLU for a cloud-based, modular spoken dialog system (SDS). We evaluate the effectiveness of our approach on crowdsourced data collected from non-native English speakers interacting with a conversational language learning application. Experimental results show that our approach is particularly promising in situations with low ASR accuracy. It can further improve the performance of a sophisticated CNN-based SLU system with more accurate ASR hypotheses by fusing the scores from E2E system, i.e., the overall accuracy of SLU is improved from 85.6% to 86.5%.
Yao Qian, Rutuja Ubale, Vikram Ramanarayanan, Patrick L. Lange, David Suendermann-Oeft, Keelan Evanini, Eugene Tsuprun
ASRU3
2017 A modular, multimodal open-source virtual interviewer dialog agent
abstract
We present an open-source multimodal dialog system equipped with a virtual human avatar interlocutor. The agent, rigged in Blender and developed in Unity with WebGL support, interfaces with the HALEF open-source cloud-based standard-compliant dialog framework. To demonstrate the capabilities of the system, we designed and implemented a conversational job interview scenario where the avatar plays the role of an interviewer and responds to user input in real-time to provide an immersive user experience.
Kirby Cofino, Vikram Ramanarayanan, Patrick L. Lange, David Pautler, David Suendermann-Oeft, Keelan Evanini
ICMI2
2017 Crowdsourcing ratings of caller engagement in thin-slice videos of human-machine dialog: benefits and pitfalls
abstract
We analyze the efficacy of different crowds of naive human raters in rating engagement during human--machine dialog interactions. Each rater viewed multiple 10 second, thin-slice videos of native and non-native English speakers interacting with a computer-assisted language learning (CALL) system and rated how engaged and disengaged those callers were while interacting with the automated agent. We observe how the crowd's ratings compared to callers' self ratings of engagement, and further study how the distribution of these rating assignments vary as a function of whether the automated system or the caller was speaking. Finally, we discuss the potential applications and pitfalls of such crowdsourced paradigms in designing, developing and analyzing engagement-aware dialog systems.
Vikram Ramanarayanan, Chee Wee Leong, David Suendermann-Oeft, Keelan Evanini
ICMI1
2017 Human and Automated Scoring of Fluency, Pronunciation and Intonation During Human-Machine Spoken Dialog Interactions
Vikram Ramanarayanan, Patrick L. Lange, Keelan Evanini, Hillary Molloy, David Suendermann-Oeft
INTERSPEECH1
2017 Rushing to Judgement: How do Laypeople Rate Caller Engagement in Thin-Slice Videos of Human-Machine Dialog?
Vikram Ramanarayanan, Chee Wee Leong, David Suendermann-Oeft
INTERSPEECH1
2017 Jee haan, I'd like both, por favor: Elicitation of a Code-Switched Corpus of Hindi-English and Spanish-English Human-Machine Dialog
Vikram Ramanarayanan, David Suendermann-Oeft
INTERSPEECH1
2017 Database of Volumetric and Real-Time Vocal Tract MRI for Speech Science
Tanner Sorensen, Z.-I. Skordilis, Asterios Toutios, Yoon-Chul Kim, Yinghua Zhu, Jangwon Kim, Adam C. Lammert, Vikram Ramanarayanan, Louis Goldstein, Dani Byrd, Krishna S. Nayak, Shri Narayanan
INTERSPEECH8
2016 Novel features for capturing cooccurrence behavior in dyadic collaborative problem solving tasks
Vikram Ramanarayanan, Saad Khan 0003
EDM1
2016 Noise and Metadata Sensitive Bottleneck Features for Improving Speaker Recognition with Non-Native Speech Input
Yao Qian, Jidong Tao, David Suendermann-Oeft, Keelan Evanini, Alexei V. Ivanov, Vikram Ramanarayanan
INTERSPEECH6
2016 A New Model of Speech Motor Control Based on Task Dynamics and State Feedback
Vikram Ramanarayanan, Benjamin Parrell, Louis Goldstein, Srikantan S. Nagarajan, John F. Houde
INTERSPEECH1
2016 Speaker verification based on the fusion of speech acoustics and inverted articulatory signals
Ming Li 0026, Jangwon Kim, Adam C. Lammert, Prasanta Kumar Ghosh, Vikram Ramanarayanan, Shri Narayanan
Comput. Speech Lang.5
2016 Directly data-derived articulatory gesture-like representations retain discriminatory information about phone categories
Vikram Ramanarayanan, Maarten Van Segbroeck, Shri Narayanan
Comput. Speech Lang.1
2015 Using bidirectional lstm recurrent neural networks to learn high-level abstractions of sequential features for automated scoring of non-native spontaneous speech
abstract
We introduce a new method to grade non-native spoken language tests automatically. Traditional automated response grading approaches use manually engineered time-aggregated features (such as mean length of pauses). We propose to incorporate general time-sequence features (such as pitch) which preserve more information than time-aggregated features and do not require human effort to design. We use a type of recurrent neural network to jointly optimize the learning of high level abstractions from time-sequence features with the time-aggregated features. We first automatically learn high level abstractions from time-sequence features with a Bidirectional Long Short Term Memory (BLSTM) and then combine the high level abstractions with time-aggregated features in a Multilayer Perceptron (MLP)/Linear Regression (LR). We optimize the BLSTM and the MLP/LR jointly. We find such models reach the best performance in terms of correlation with human raters. We also find that when there are limited time-aggregated features available, our model that incorporates time-sequence features improves performance drastically.
Zhou Yu 0005, Vikram Ramanarayanan, David Suendermann-Oeft, Klaus Zechner, Lei Chen 0004, Jidong Tao, Aliaksei Ivanou, Yao Qian
ASRU2
2015 Evaluating Speech, Face, Emotion and Body Movement Time-series Features for Automated Multimodal Presentation Scoring
abstract
We analyze how fusing features obtained from different multimodal data streams such as speech, face, body movement and emotion tracks can be applied to the scoring of multimodal presentations. We compute both time-aggregated and time-series based features from these data streams--the former being statistical functionals and other cumulative features computed over the entire time series, while the latter, dubbed histograms of cooccurrences, capture how different prototypical body posture or facial configurations co-occur within different time-lags of each other over the evolution of the multimodal, multivariate time series. We examine the relative utility of these features, along with curated speech stream features in predicting human-rated scores of multiple aspects of presentation proficiency. We find that different modalities are useful in predicting different aspects, even outperforming a naive human inter-rater agreement baseline for a subset of the aspects analyzed.
Vikram Ramanarayanan, Chee Wee Leong, Lei Chen 0004, Gary Feng, David Suendermann-Oeft
ICMI1
2015 An analysis of time-aggregated and time-series features for scoring different aspects of multimodal presentation data
abstract
We present a technique for automated assessment of public speaking and presentation proficiency based on the analysis of concurrently recorded speech and motion capture data. With respect to Kinect motion capture data, we examine both timeaggregated as well as time-series based features. While the former is based on statistical functionals of body-part position and/or velocity computed over the entire series, the latter feature set, dubbed histograms of cooccurrences, captures how often different broad postural configurations co-occur within different time lags of each other over the evolution of the multimodal time series. We examine the relative utility of these features, along with curated features derived from the speech stream, in predicting human-rated scores of different aspects of public speaking and presentation proficiency. We further show that these features outperform the human inter-rater agreement baseline for a subset of the analyzed aspects.
Vikram Ramanarayanan, Lei Chen 0004, Chee Wee Leong, Gary Feng, David Suendermann-Oeft
INTERSPEECH1
2015 Experimental assessment of the tongue incompressibility hypothesis during speech production
abstract
The human tongue is an important organ for speech production. Its deformation and motion control the shape of the vocal tract significantly and thereby the acoustic properties of the speech signal produced. Thus, much effort in the speech research com-munity has been directed towards its biomechanical modeling. A common assumption incorporated into many models of the human tongue is the tissue incompressibility hypothesis: the tongue is considered a muscular hydrostat and therefore its vol-ume should remain constant regardless of its posture. To the best of our knowledge, experimental assessment of the constant volume hypothesis during actual speech production is limited. In this work, the aim is to experimentally assess the incom-pressibility hypothesis during actual speech production using a dataset of volumetric Magnetic Resonance (MR) images of 17 subjects sustaining contextualized continuants (27 continu-ants per subject). A seeded region growing based algorithm is used to segment the tongue and calculate its volume. Then the intra-subject variability of the tongue volume along the differ-ent tongue postures is examined. Within the accuracy of our tongue volume measurements, our empirical results seem con-sistent with the incompressibility hypothesis. Index Terms: speech production, tongue volume, muscular hy-drostat, tissue incompressibility, volumetric MRI
Z.-I. Skordilis, Vikram Ramanarayanan, Louis Goldstein, Shri Narayanan
INTERSPEECH2
2015 Automated Speech Recognition Technology for Dialogue Interaction with Non-Native Interlocutors
abstract
Alexei V. Ivanov, Vikram Ramanarayanan, David Suendermann-Oeft, Melissa Lopez, Keelan Evanini, Jidong Tao. Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2015.
Alexei V. Ivanov, Vikram Ramanarayanan, David Suendermann-Oeft, Melissa Lopez, Keelan Evanini, Jidong Tao
SIGDIAL Conference2
2015 A distributed cloud-based dialog system for conversational application development
abstract
We have previously presented HALEF-an open-source spoken dialog system-that supports telephonic interfaces and has a distributed architecture.In this paper, we extend this infrastructure to be cloud-based, and thus truly distributed and scalable.This cloud-based spoken dialog system can be accessed both via telephone interfaces as well as through web clients with WebRTC/HTML5 integration, allowing in-browser access to potentially multimodal dialog applications.We demonstrate the versatility of the system with two conversation applications in the educational domain.
Vikram Ramanarayanan, David Suendermann-Oeft, Alexei V. Ivanov, Keelan Evanini
SIGDIAL Conference1
2014 A real-time MRI study of articulatory setting in second language speech
abstract
Previous work has shown that languages differ in their articulatory setting, the postural configuration that the vocal tract articulators tend to adopt when they are not engaged in any active speech gesture, and that this posture might be specified as part of the phonological knowledge speakers have of the language. This study tests whether the articulatory setting of a language can be acquired by non-native speakers. Three native speakers of German who had learned English as a second language were imaged using real-time MRI of the vocal tract while reading passages in German and English, and features that capture vocal tract posture were extracted from the inter-speech pauses in their native and non-native languages. Results show that the speakers exhibit distinct inter-speech postures in each language, with a lower and more retracted tongue in English, consistent with classic descriptions of the differences between the German and the English articulatory settings. This supports the view that non-native speakers may acquire relevant features of the articulatory setting of a second language, and also lends further support to the idea that articulatory setting is part of a speaker’s phonological competence in a language. Index Terms: articulatory setting, speech production, second language speech acquisition, real-time MRI.
Andrés Benítez, Vikram Ramanarayanan, Louis Goldstein, Shri Narayanan
INTERSPEECH2
2014 Motor control primitives arising from a learned dynamical systems model of speech articulation
abstract
We present a method to derive a small number of speech motor control “primitives” that can produce linguisticallyinterpretable articulatory movements. We envision that such a dictionary of primitives can be useful for speech motor control, particularly in finding a low-dimensional subspace for such control. First, we use the iterative Linear Quadratic Gaussian with Learned Dynamics (iLQG-LD) algorithm to derive (for a set of utterances) a set of stochastically optimal control inputs to a learned dynamical systems model of the vocal tract that produces desired movement sequences. Second, we use a convolutive Nonnegative Matrix Factorization with sparseness constraints (cNMFsc) algorithm to find a small dictionary of control input primitives that can be used to reproduce the aforementioned optimal control inputs that produce the observed articulatory movements. The method performs favorably on both qualitative and quantitative evaluations conducted on synthetic data produced by an articulatory synthesizer. Such a primitivesbased framework could help inform theories of speech motor control and coordination. Index Terms: speech motor control, motor primitives, synergies, dynamical systems, iLQG, NMF.
Vikram Ramanarayanan, Louis Goldstein, Shri Narayanan
INTERSPEECH1
2014 Joint filtering and factorization for recovering latent structure from noisy speech data
abstract
We propose a joint filtering and factorization algorithm to re-cover latent structure from noisy speech. We incorporate the minimum variance distortionless response (MVDR) formula-tion within the non-negative matrix factorization (NMF) frame-work to derive a single, unified cost function for both filtering and factorization. Minimizing this cost function jointly opti-mizes three quantities – a filter that removes noise, a basis ma-trix that captures latent structure in the data, and an activation matrix that captures how the elements in the basis matrix can be linearly combined to reconstruct input data. Results show that the proposed algorithm recovers the speech basis matrix from noisy speech significantly better than NMF alone or Wiener fil-tering followed by NMF. Furthermore, PESQ scores show that our algorithm is a viable choice for speech denoising. Index Terms: NMF, MVDR, denoising, filtering. 1.
Colin Vaz, Vikram Ramanarayanan, Shri Narayanan
INTERSPEECH2
2013 Analyzing eye-voice coordination in rapid automatized naming
abstract
Rapid Automatized Naming (RAN) is a powerful tool for pre-dicting future reading skill. A person’s ability to quickly name symbols as they scan a table is related to higher-level reading proficiency in adults and is predictive of future literacy gains in children. However, noticeable differences are present in the strategies or patterns within groups having similar task comple-tion times. Thus, a further stratification of RAN dynamics may lead to better characterization and later intervention to support reading skill acquisition. In this work, we analyze the dynamics of the eyes, voice, and the coordination between the two during performance. It is shown that fast performers are more similar to each other than to slow performers in their patterns, but not vice versa. Further insights are provided about the patterns of more proficient subjects. For instance, fast performers tended to exhibit smoother behavior contours, suggesting a more sta-ble perception-production process.
Daniel Bone, Chi-Chun Lee, Vikram Ramanarayanan, Shri Narayanan, Renske S. Hoedemaker, Peter C. Gordon
INTERSPEECH3
2013 Vocal tract cross-distance estimation from real-time MRI using region-of-interest analysis
abstract
Real-Time Magnetic Resonance Imaging affords speech articu-lation data with good spatial and temporal resolution and com-plete midsagittal views of the moving vocal tract, but also brings many challenges in the domain of image processing and analy-sis. Region-of-interest analysis has previously been proposed for simple, efficient and robust extraction of linguistically-meaningful constriction degree information. However, the ac-curacy of such methods has not been rigorously evaluated, and no method has been proposed to calibrate the pixel intensity values or convert them into absolute measurements of length. This work provides such an evaluation, as well as insights into the placement of regions in the image plane and calibration of the resultant pixel intensity measurements. Measurement errors are shown to be generally at or below the spatial resolution of the imaging protocol with a high degree of consistency across time and overall vocal tract configuration, validating the utility of this method of image analysis. Index Terms: speech production data, real-time mri, analysis tools, vocal tract area functions
Adam C. Lammert, Vikram Ramanarayanan, Michael I. Proctor, Shri Narayanan
INTERSPEECH2
2013 Speaker verification based on fusion of acoustic and articulatory information
abstract
We propose a practical, feature-level fusion approach for com-bining acoustic and articulatory information in speaker ver-ification task. We find that concatenating articulation fea-tures obtained from the measured speech production data with conventional Mel-frequency cepstral coefficients (MFCCs) im-proves the overall speaker verification performance. However, since access to the measured articulatory data is impractical for real world speaker verification applications, we also ex-periment with estimated articulatory features obtained using acoustic-to-articulatory inversion technique. Specifically, we show that augmenting MFCCs with articulatory features ob-tained from subject-independent acoustic-to-articulatory inver-sion technique also significantly enhances the speaker verifi-cation performance. This performance boost could be due to the information about inter-speaker variation present in the es-timated articulatory features, especially at the mean and vari-ance level. Experimental results on the Wisconsin X-Ray Mi-crobeam database show that the proposed acoustic-estimated-articulatory fusion approach significantly outperforms the tra-ditional acoustic-only baseline, providing up to 10 % relative re-duction in Equal Error Rate (EER). We further show that we can achieve an additional 5 % relative reduction in EER after score-level fusion. Index Terms: speech production, speaker verification, articula-tion features, acoustic-to-articulatory inversion, biometrics
Ming Li 0026, Jangwon Kim, Prasanta Kumar Ghosh, Vikram Ramanarayanan, Shri Narayanan
INTERSPEECH4
2013 Articulatory settings facilitate mechanically advantageous motor control of vocal tract articulators
abstract
It was recently shown that vocal tract postures assumed during pauses in read speech are significantly different from those assumed at absolute rest. This paper examines whether the former category of “articulatory settings” are more mechanically advantageous than absolute rest postures with respect to speech articulation. Appropriate task and articulator variables are extracted from real-time Magnetic Resonance Imaging (rtMRI) data of five speakers reading aloud. Locally-weighted regression is then used to calculate Jacobian matrices representing the transformation between articulatory task velocities and postural velocities. A measure of mechanical advantage is proposed based on the obtained Jacobian. Speech-ready postures and postures during inter-speech pauses are observed to be significantly more mechanically advantageous as compared to rest postures. Furthermore, other postures, such as those that occur during the production of different vowels and consonants, are shown to have mechanical advantages that lie in between this continuum. These results could provide insights into understanding postural motor control and other linguistic phenomena, such as sonority hierarchies, in speech production. Index Terms: speech production, real-time MRI, articulatory setting, postural motor control, task dynamics, forward kinematics, vocal tract shaping.
Vikram Ramanarayanan, Adam C. Lammert, Louis Goldstein, Shri Narayanan
INTERSPEECH1
2013 A two-step technique for MRI audio enhancement using dictionary learning and wavelet packet analysis
abstract
We present a method for speech enhancement of data collected in extremely noisy environments, such as those found during magnetic resonance imaging (MRI) scans. We propose a twostep algorithm to perform this noise suppression. First, we use probabilistic latent component analysis to learn dictionaries of the noise and speech+noise portions of the data and use these to factor the noisy spectrum into estimated speech and noise components. Second, we apply a wavelet packet analysis in conjunction with a wavelet threshold that minimizes the KL divergence between the estimated speech and noise to achieve further noise suppression. Based on both objective and subjective assessments, we find that our algorithm significantly outperforms traditional techniques such as nLMS, while not requiring prior knowledge or periodicity of the noise waveforms that current state-of-the-art algorithms require. Index Terms: rtMRI, noise suppression, wavelets, pLCA, dictionary learning.
Colin Vaz, Vikram Ramanarayanan, Shri Narayanan
INTERSPEECH2
2013 The effect of word frequency and lexical class on articulatory-acoustic coupling
abstract
Word frequency and lexical class distinction between function and content words have been shown to significantly influence word production. In this paper, we use real-time magnetic resonance imaging to investigate the effect of word frequency and lexical class on articulatory characteristics (the articulator speed) as well as acoustic characteristics (F0 and short-term en-ergy) in word production. Multiple regression analyses showed that word frequency exhibits significantly higher correlation with articulatory and acoustic factors for content words com-pared to function words. A Granger causality analysis uncov-ered a causal relationship from articulatory speed to F0/energy for low-frequency content words. We further observed, us-ing functional canonical correlation analysis, a tight coupling of articulatory and acoustic characteristics for low-frequency content words. These results support the view that word fre-quency distinctly influences the production of function and con-tent words as manifested in their articulation and acoustics, as well as the dynamic coupling of these temporal streams. Index Terms: word frequency, lexical class, articulatory-acoustic coupling, real-time MRI, speech production.
Vikram Ramanarayanan, Dani Byrd, Shri Narayanan
INTERSPEECH2
2011 Validating rt-MRI Based Articulatory Representations via Articulatory Recognition
abstract
The large corpus of real time magnetic resonance image sequences of the vocal tract during speech production that was recently acquired and can be referred to as MRI-TIMIT, provides us with a unique platform for systematically studying articulatory dynamics. Compared to previously collected articulatory datasets, e.g., using articulography or X-rays, MRI-TIMIT is a rich source of information for the entire vocal tract and not only for certain articulatory landmarks and further has the potential to continue increasing in size covering a large variety of speakers and speaking styles. In this work, we investigate an articulatory representation based on full vocal tract shapes. We employ an articulatory recognition framework in MRI-TIMIT to analyze its merits and drawbacks. We argue that articulatory recognition can serve as a general validation tool for real-time MRI based articulatory representations. Index Terms: vocal tract shape, articulation, real-time MRI, articulatory recognition
Athanasios Katsamanis, Erik Bresch, Vikram Ramanarayanan, Shri Narayanan
INTERSPEECH3
2011 A Multimodal Real-Time MRI Articulatory Corpus for Speech Research
abstract
We present MRI-TIMIT: a large-scale database of synchronized audio and real-time magnetic resonance imaging (rtMRI) data for speech research. The database currently consists of speech data acquired from two male and two female speakers of Amer-ican English. Subjects ’ upper airways were imaged in the mid-sagittal plane while reading the same 460 sentence corpus used in the MOCHA-TIMIT corpus [1]. Accompanying acoustic recordings were phonemically transcribed using forced align-ment. Vocal tract tissue boundaries were automatically identi-fied in each video frame, allowing for dynamic quantification of each speaker’s midsagittal articulation. The database and com-panion toolset provide a unique resource with which to examine articulatory-acoustic relationships in speech production. Index Terms: speech production, speech corpora, real-time MRI, multi-modal database, large-scale phonetic tools
Shri Narayanan, Erik Bresch, Prasanta Kumar Ghosh, Louis Goldstein, Athanasios Katsamanis, Adam C. Lammert, Michael I. Proctor, Vikram Ramanarayanan, Yinghua Zhu
INTERSPEECH9
2011 Automatic Data-Driven Learning of Articulatory Primitives from Real-Time MRI Data Using Convolutive NMF with Sparseness Constraints
abstract
We present a procedure to automatically derive inter-pretable dynamic articulatory primitives in a data-driven man-ner from image sequences acquired through real-time magnetic resonance imaging (rt-MRI). More specifically, we propose a convolutive Nonnegative Matrix Factorization algorithm with sparseness constraints (cNMFsc) to decompose a given set of image sequences into a set of basis image sequences and an acti-vation matrix. We use a recently-acquired rt-MRI corpus of read speech (460 sentences from 4 speakers) as a test dataset for this procedure. We choose the free parameters of the algorithm em-pirically by analyzing algorithm performance for different pa-rameter values. We then validate the extracted basis sequences using an articulatory recognition task and finally present an in-terpretation of the extracted basis set of image sequences in a gesture-based Articulatory Phonology framework.
Vikram Ramanarayanan, Athanasios Katsamanis, Shri Narayanan
INTERSPEECH1
2010 Investigating articulatory setting - pauses, ready position, and rest - using real-time MRI
abstract
We present a novel automatic procedure to analyze ―articulatory setting (AS) ‖ or ―basis of articulation ‖ using realtime magnetic resonance images (rt-MRI) of the human vocal tract recorded for read and spontaneously spoken speech. We extract relevant frames of inter-speech pauses (ISPs) and rest positions from MRI sequences of read and spontaneous speech and use automatically-extracted features to quantify areas of different regions of the vocal tract as well as the angle of the jaw. Significant differences were found between the ASs adopted for ISPs in read and spontaneous speech, as well as those between ISPs and absolute rest positions. We further contrast differences between ASs adopted when the person is ready to speak as opposed to an absolute rest position. Index Terms — speech production, real-time MRI, basis of articulation, articulatory setting, pause articulation, read speech, spontaneous speech. 1.
Vikram Ramanarayanan, Dani Byrd, Louis Goldstein, Shri Narayanan
INTERSPEECH1