VLDB 2026 Research / reviewers in the wild / expert
Alexander I. Rudnicky
dblp:29/5401 · also Alex Rudnicky, Alexander Rudnicky
· DBLP profile ↗
135ranked-venue papers
10as first author
17since 2021 · last 2025
0000-0003-2044-8446ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 90 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 80 · 8 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 21 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Language Models Can be Efficiently Steered via Minimal Embedding Layer TransformationsabstractLarge Language Models (LLMs) are increasingly costly to fine-tune due to their size, with embedding layers alone accounting for up to 20% of model parameters.While Parameter-Efficient Fine-Tuning (PEFT) methods exist, they largely overlook the embedding layer.In this paper, we introduce TinyTE, a novel PEFT approach that steers model behavior via minimal translational transformations in the embedding space.TinyTE modifies input embeddings without altering hidden layers, achieving competitive performance while requiring approximately 0.0001% of the parameters needed for full fine-tuning.Experiments across architectures provide a new lens for understanding the relationship between input representations and model behavior-revealing them to be more flexible at their foundation than previously thought. 1 Diogo Tavares, David Semedo, Alexander I. Rudnicky, João Magalhães |
EMNLP | 3 |
| 2025 | Exploring Prediction Targets in Masked Pre-Training for Speech Foundation ModelsabstractSpeech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use a masked prediction objective, where the model learns to predict information about masked input segments from the unmasked context. The choice of prediction targets in this framework impacts their performance on downstream tasks. For instance, models pre-trained with targets that capture prosody learn representations suited for speaker-related tasks, while those pre-trained with targets that capture phonetics learn representations suited for content-related tasks. Moreover, prediction targets can differ in the level of detail they capture. Models pre-trained with targets that encode fine-grained acoustic features perform better on tasks like denoising, while those pre-trained with targets focused on higher-level abstractions are more effective for content-related tasks. Despite the importance of prediction targets, the design choices that affect them have not been thoroughly studied. This work explores the design choices and their impact on downstream task performance. Our results indicate that the commonly used design choices for HuBERT can be suboptimal. We propose approaches to create more informative prediction targets and demonstrate their effectiveness through improvements across various downstream tasks. Takuya Higuchi, He Bai 0013, Ahmed Hussen Abdelaziz, Shinji Watanabe 0001, Alexander I. Rudnicky, Tatiana Likhomanenko, Barry-John Theobald, Zakaria Aldeneh |
ICASSP | 6 |
| 2025 | A Variational Framework for Improving Naturalness in Generative Spoken Language ModelsabstractThe success of large language models in text processing has inspired their adaptation to speech modeling.
However, since speech is continuous and complex, it is often discretized for autoregressive modeling.
Speech tokens derived from self-supervised models (known as semantic tokens) typically focus on the linguistic aspects of speech but neglect prosodic information.
As a result, models trained on these tokens can generate speech with reduced naturalness.
Existing approaches try to fix this by adding pitch features to the semantic tokens.
However, pitch alone cannot fully represent the range of paralinguistic attributes, and selecting the right features requires careful hand-engineering.
To overcome this, we propose an end-to-end variational approach that automatically learns to encode these continuous speech attributes to enhance the semantic tokens.
Our approach eliminates the need for manual extraction and selection of paralinguistic features.
Moreover, it produces preferred speech continuations according to human raters.
Code, samples and models are available at https://github.com/b04901014/vae-gslm. Takuya Higuchi, Zakaria Aldeneh, Ahmed Hussen Abdelaziz, Alexander I. Rudnicky |
ICML | 5 |
| 2024 | Overview of the Tenth Dialog System Technology Challenge: DSTC10abstractThis article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky |
IEEE ACM Trans. Audio Speech Lang. Process. | 28 |
| 2023 | A Vector Quantized Approach for Text to Speech Synthesis on Real-World Spontaneous SpeechabstractRecent Text-to-Speech (TTS) systems trained on reading or acted corpora have achieved near human-level naturalness. The diversity of human speech, however, often goes beyond the coverage of these corpora. We believe the ability to handle such diversity is crucial for AI systems to achieve human-level communication. Our work explores the use of more abundant real-world data for building speech synthesizers. We train TTS systems using real-world speech from YouTube and podcasts. We observe the mismatch between training and inference alignments in mel-spectrogram based autoregressive models, leading to unintelligible synthesis, and demonstrate that learned discrete codes within multiple code groups effectively resolves this issue. We introduce our MQTTS system whose architecture is designed for multiple code generation and monotonic alignment, along with the use of a clean silence prompt to improve synthesis quality. We conduct ablation analyses to identify the efficacy of our methods. We show that MQTTS outperforms existing TTS systems in several objective and subjective measures. Shinji Watanabe 0001, Alexander I. Rudnicky |
AAAI | 3 |
| 2023 | Dissecting Transformer Length Extrapolation via the Lens of Receptive Field AnalysisabstractLength extrapolation permits training a transformer language model on short sequences that preserves perplexities when tested on substantially longer sequences.A relative positional embedding design, ALiBi, has had the widest usage to date.We dissect ALiBi via the lens of receptive field analysis empowered by a novel cumulative normalized gradient tool.The concept of receptive field further allows us to modify the vanilla Sinusoidal positional embedding to create Sandwich, the first parameter-free relative positional embedding design that truly length information uses longer than the training sequence.Sandwich shares with KERPLE and T5 the same logarithmic decaying temporal bias pattern with learnable relative positional embeddings; these elucidate future extrapolatable positional embedding design. Ta-Chung Chi, Ting-Han Fan, Alexander I. Rudnicky, Peter J. Ramadge |
ACL (1) | 3 |
| 2023 | Exploring Wav2vec 2.0 Fine Tuning for Improved Speech Emotion RecognitionabstractWhile Wav2Vec 2.0 has been proposed for speech recognition (ASR), it can also be used for speech emotion recognition (SER); its performance can be significantly improved using different fine-tuning strategies. Two baseline methods, vanilla fine-tuning (V-FT) and task adaptive pretraining (TAPT) are first presented. We show that V-FT is able to outperform state-of-the-art models on the IEMOCAP dataset. TAPT, an existing NLP fine-tuning strategy, further improves the performance on SER. We also introduce a novel fine-tuning method termed P-TAPT, which modifies the TAPT objective to learn contextualized emotion representations. Experiments show that P-TAPT performs better than TAPT, especially under low-resource settings. Compared to prior works in this literature, our top-line system achieved a 7.4% absolute improvement in unweighted accuracy (UA) over the state-of-the-art performance on IEMOCAP. Our code is publicly available.1 Alexander I. Rudnicky |
ICASSP | 2 |
| 2023 | A Unified One-Shot Prosody and Speaker Conversion System with Self-Supervised Discrete Speech UnitsabstractWe present a unified system to realize one-shot voice conversion (VC) on the pitch, rhythm, and speaker attributes. Existing works generally ignore the correlation between prosody and language content, leading to the degradation of naturalness in converted speech. Additionally, the lack of proper language features prevents these systems from accurately preserving language content after conversion. To address these issues, we devise a cascaded modular system leveraging self-supervised discrete speech units as language representation. These discrete units provide duration information essential for rhythm modeling. Our system first extracts utterance-level prosody and speaker representations from the raw waveform. Given the prosody representation, a prosody predictor estimates pitch, energy, and duration for each discrete unit in the utterance. A synthesizer further reconstructs speech based on the predicted prosody, speaker representation, and discrete units. Experiments show that our system outperforms previous approaches in naturalness, intelligibility, speaker transferability, and prosody transferability. Code and samples are publicly available.1 Shinji Watanabe 0001, Alexander I. Rudnicky |
ICASSP | 3 |
| 2023 | Learning to Ask Questions for Zero-shot Dialogue State TrackingabstractWe present a method for performing zero-shot Dialogue State Tracking (DST) by casting the task as a learning-to-ask-questions framework. The framework learns to pair the best question generation (QG) strategy with in-domain question answering (QA) methods to extract slot values from a dialogue without any human intervention. A novel self-supervised QA pretraining step using in-domain data is essential to learn the structure without requiring any slot-filling annotations. Moreover, we show that QG methods need to be aligned with the same grammatical person used in the dialogue. Empirical evaluation on the MultiWOZ 2.1 dataset demonstrates that our approach, when used alongside robust QA models, outperforms existing zero-shot methods in the challenging task of zero-shot cross domain adaptation-given a comparable amount of domain knowledge during data creation. Finally, we analyze the impact of the types of questions used, and demonstrate that the algorithmic approach outperforms template-based question generation. Diogo Tavares, David Semedo, Alexander I. Rudnicky, João Magalhães |
SIGIR | 3 |
| 2022 | Fine-Grained Style Control In Transformer-Based Text-To-Speech SynthesisabstractIn this paper, we present a novel architecture to realize fine-grained style control on the transformer-based text-to-speech synthesis (TransformerTTS). Specifically, we model the speaking style by extracting a time sequence of local style tokens (LST) from the reference speech. The existing content encoder in TransformerTTS is then replaced by our designed cross-attention blocks for fusion and alignment between content and style. As the fusion is performed along with the skip connection, our cross-attention block provides a good inductive bias to gradually infuse the phoneme representation with a given style. Additionally, we prevent the style embedding from encoding linguistic content by randomly truncating LST during training and using wav2vec 2.0 features. Experiments show that with fine-grained style control, our system performs better in terms of naturalness, intelligibility, and style transferability. Our code and samples are publicly available.1 Alexander I. Rudnicky |
ICASSP | 2 |
| 2022 | Training Discrete Deep Generative Models via Gapped Straight-Through EstimatorabstractWhile deep generative models have succeeded in image processing, natural language processing, and reinforcement learning, training that involves discrete random variables remains challenging due to the high variance of its gradient estimation process. Monte Carlo is a common solution used in most variance reduction approaches. However, this involves time-consuming resampling and multiple function evaluations. We propose a Gapped Straight-Through (GST) estimator to reduce the variance without incurring resampling overhead. This estimator is inspired by the essential properties of Straight-Through Gumbel-Softmax. We determine these properties and show via an ablation study that they are essential. Experiments demonstrate that the proposed GST estimator enjoys better performance compared to strong baselines on two discrete deep generative modeling tasks, MNIST-VAE and ListOps. Ting-Han Fan, Ta-Chung Chi, Alexander I. Rudnicky, Peter J. Ramadge |
ICML | 3 |
| 2022 | KERPLE: Kernelized Relative Positional Embedding for Length ExtrapolationabstractRelative positional embeddings (RPE) have received considerable attention since RPEs effectively model the relative distance among tokens and enable length extrapolation. We propose KERPLE, a framework that generalizes relative position embedding for extrapolation by kernelizing positional differences. We achieve this goal using conditionally positive definite (CPD) kernels, a class of functions known for generalizing distance metrics. To maintain the inner product interpretation of self-attention, we show that a CPD kernel can be transformed into a PD kernel by adding a constant offset. This offset is implicitly absorbed in the Softmax normalization during self-attention. The diversity of CPD kernels allows us to derive various RPEs that enable length extrapolation in a principled way. Experiments demonstrate that the logarithmic variant achieves excellent extrapolation performance on three large language modeling datasets. Our implementation and pretrained checkpoints are released at~\url{https://github.com/chijames/KERPLE.git}. Ta-Chung Chi, Ting-Han Fan, Peter J. Ramadge, Alexander I. Rudnicky |
NeurIPS | 4 |
| 2022 | Structured Dialogue Discourse ParsingabstractDialogue discourse parsing aims to uncover the internal structure of a multi-participant conversation by finding all the discourse links and corresponding relations.Previous work either treats this task as a series of independent multiple-choice problems, in which the link existence and relations are decoded separately, or the encoding is restricted to only local interaction, ignoring the holistic structural information.In contrast, we propose a principled method that improves upon previous work from two perspectives: encoding and decoding.From the encoding side, we perform structured encoding on the adjacency matrix followed by the matrix-tree learning algorithm, where all discourse links and relations in the dialogue are jointly optimized based on latent tree-level distribution.From the decoding side, we perform structured inference using the modified Chiu-Liu-Edmonds algorithm, which explicitly generates the labeled multi-root non-projective spanning tree that best captures the discourse structure.In addition, unlike in previous work, we do not rely on hand-crafted features; this improves the model's robustness.Experiments show that our method achieves new state-of-the-art, surpassing the previous model by 2.3 on STAC and 1.5 on Molweni (F1 scores). 1 Ta-Chung Chi, Alexander I. Rudnicky |
SIGDIAL | 2 |
| 2022 | Spoken language interaction with robots: Recommendations for future researchabstractWith robotics rapidly advancing, more effective human–robot interaction is increasingly needed to realize the full potential of robots for society. While spoken language must be part of the solution, our ability to provide spoken language interaction capabilities is still very limited. In this article, based on the report of an interdisciplinary workshop convened by the National Science Foundation, we identify key scientific and engineering advances needed to enable effective spoken language interaction with robotics. We make 25 recommendations, involving eight general themes: putting human needs first, better modeling the social and interactive aspects of language, improving robustness, creating new methods for rapid adaptation, better integrating speech and language with other communication modalities, giving speech and language components access to rich representations of the robot’s current knowledge and state, making all components operate in real time, and improving research infrastructure and resources. Research and development that prioritizes these topics will, we believe, provide a solid foundation for the creation of speech-capable robots that are easy and effective for humans to work with. Matthew Marge, Carol Y. Espy-Wilson, Nigel G. Ward, Abeer Alwan, Yoav Artzi, Mohit Bansal, Gilmer L. Blankenship, Joyce Y. Chai, Hal Daumé III, Debadeepta Dey, Mary P. Harper, Thomas Howard, Casey Kennington, Ivana Kruijff-Korbayová, Dinesh Manocha, Cynthia Matuszek, Ross Mead, Raymond J. Mooney, Roger K. Moore, Mari Ostendorf, Heather Pon-Barry, Alexander I. Rudnicky, Matthias Scheutz, Robert St. Amant, Stefanie Tellex, David R. Traum, Zhou Yu 0005 |
Comput. Speech Lang. | 22 |
| 2021 | Zero-Shot Dialogue Disentanglement by Self-Supervised Entangled Response SelectionabstractDialogue disentanglement aims to group utterances in a long and multi-participant dialogue into threads.This is useful for discourse analysis and downstream applications such as dialogue response selection, where it can be the first step to construct a clean context/response set.Unfortunately, labeling all reply-to links takes quadratic effort w.r.t the number of utterances: an annotator must check all preceding utterances to identify the one to which the current utterance is a reply.In this paper, we are the first to propose a zero-shot dialogue disentanglement solution.Firstly, we train a model on a multi-participant response selection dataset harvested from the web which is not annotated; we then apply the trained model to perform zero-shot dialogue disentanglement.Without any labeled data, our model can achieve a cluster F1 score of 25.We also fine-tune the model using various amounts of labeled data.Experiments show that with only 10% of the data, we achieve nearly the same performance of using the full dataset 1 . Ta-Chung Chi, Alexander I. Rudnicky |
EMNLP (1) | 2 |
| 2021 | Speech Representation Learning Combining Conformer CPC with Deep Cluster for the ZeroSpeech Challenge 2021abstractWe present a system for the Zero Resource Speech Challenge 2021, which combines a Contrastive Predictive Coding (CPC) with deep cluster. In deep cluster, we first prepare pseudo-labels obtained by clustering the outputs of a CPC network with k-means. Then, we train an additional autoregressive model to classify the previously obtained pseudo-labels in a supervised manner. Phoneme discriminative representation is achieved by executing the second-round clustering with the outputs of the final layer of the autoregressive model. We show that replacing a Transformer layer with a Conformer layer leads to a further gain in a lexical metric. Experimental results show that a relative improvement of 35% in a phonetic metric, 1.5% in the lexical metric, and 2.3% in a syntactic metric are achieved compared to a baseline method of CPC-small which is trained on LibriSpeech 460h data. We achieve top results in this challenge with the syntactic metric. Takashi Maekaku, Xuankai Chang, Yuya Fujita, Shinji Watanabe 0001, Alexander I. Rudnicky |
Interspeech | 6 |
| 2021 | Temporal Context in Speech Emotion Recognition
Yangyang Xia, Alexander I. Rudnicky, Richard M. Stern |
Interspeech | 3 |
| 2020 | Adjusting Image Attributes of Localized Regions with Low-level DialogueabstractNatural Language Image Editing (NLIE) aims to use natural language instructions to edit images. Since novices are inexperienced with image editing techniques, their instructions are often ambiguous and contain high-level abstractions which require complex editing steps. Motivated by this inexperience aspect, we aim to smooth the learning curve by teaching the novices to edit images using low-level command terminologies. Towards this end, we develop a task-oriented dialogue system to investigate low-level instructions for NLIE. Our system grounds language on the level of edit operations, and suggests options for users to choose from. Though compelled to express in low-level terms, user evaluation shows that 25% of users found our system easy-to-use, resonating with our motivation. Analysis shows that users generally adapt to utilizing the proposed low-level language interface. We also identified object segmentation as the key factor to user satisfaction. Our work demonstrates advantages of low-level, direct language-action mapping approach that can be applied to other problem domains beyond image editing such as audio editing or industrial design. Tzu-Hsiang Lin, Alexander I. Rudnicky, Trung Bui, Doo Soon Kim, Jean Oh |
LREC | 2 |
| 2019 | Miscommunication Detection and Recovery in Situated Human-Robot DialogueabstractEven without speech recognition errors, robots may face difficulties interpreting natural-language instructions. We present a method for robustly handling miscommunication between people and robots in task-oriented spoken dialogue. This capability is implemented in TeamTalk, a conversational interface to robots that supports detection and recovery from the situated grounding problems of referential ambiguity and impossible actions. We introduce a representation that detects these problems and a nearest-neighbor learning algorithm that selects recovery strategies for a virtual robot. When the robot encounters a grounding problem, it looks back on its interaction history to consider how it resolved similar situations. The learning method is trained initially on crowdsourced data but is then supplemented by interactions from a longitudinal user study in which six participants performed navigation tasks with the robot. We compare results collected using a general model to user-specific models and find that user-specific models perform best on measures of dialogue efficiency, while the general model yields the highest agreement with human judges. Our overall contribution is a novel approach to detecting and recovering from miscommunication in dialogue by including situated context, namely, information from a robot’s path planner and surroundings. Matthew Marge, Alexander I. Rudnicky |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2018 | SOGO: A Social Intelligent Negotiation Dialogue SystemabstractIn this paper, we propose a semi-automatic social intelligent negotiation dialogue system that interweaves task utterance with conversational strategies to engage human users in negotiation. Our two-phase system operates sequentially in a reasoning-and-generation loop: In the task phase, we leverage an off-the-shelf end-to-end dialogue model for negotiation to build a dialogue manager which decides the next system's task intention. Then, during the social phase, we employ a theory-driven, template-based natural language generator to realize the task intention as a genre of social conversational strategy. Subsequently, a set of conversational strategies are presented to a human expert who decides the final sentence to be uttered by the dialogue system. Compared to the baseline system, our proposed social intelligent dialogue system achieves a higher agreement rate and more "good deals" with humans while building interpersonal rapport. Oscar J. Romero, Alexander I. Rudnicky |
IVA | 3 |
| 2018 | Word Segmentation From Phoneme Sequences Based On Pitman-Yor Semi-Markov Model Exploiting Subword InformationabstractWord segmentation from phoneme sequences is essential to identify unknown words -of-vocabulary; OOV) in spoken dialogues. The Pitman-Yor semi-Markov model (PYSMM) is used for word segmentation that handles dynamic increase in vocabularies. The obtained vocabularies, however, still include meaningless entries due to insufficient cues for phoneme sequences. We focus here on using subword information to capture patterns as “words.” We propose 1) a model based on subword N-gram and subword estimation using a vocabulary set, and 2) posterior fusion of the results of a PYSMM and our model to take advantage of both. Our experiments showed 1) the potential of using subword information for OOV acquisition, and 2) that our method outperformed the PYSMM by 1.53 and 1.07 in terms of the F-measure of the obtained OOV set for English and Japanese corpora, respectively. Ryu Takeda, Kazunori Komatani, Alexander I. Rudnicky |
SLT | 3 |
| 2017 | Learning Conversational Systems that Interleave Task and Non-Task ContentabstractTask-oriented dialog systems have been applied in various tasks, such as automated personal assistants, customer service providers and tutors. These systems work well when users have clear and explicit intentions that are well-aligned to the systems' capabilities. However, they fail if users intentions are not explicit.To address this shortcoming, we propose a framework to interleave non-task content (i.e.everyday social conversation) into task conversations. When the task content fails, the system can still keep the user engaged with the non-task content. We trained a policy using reinforcement learning algorithms to promote long-turn conversation coherence and consistency, so that the system can have smooth transitions between task and non-task content.To test the effectiveness of the proposed framework, we developed a movie promotion dialog system. Experiments with human users indicate that a system that interleaves social and task content achieves a better task success rate and is also rated as more engaging compared to a pure task-oriented system. Zhou Yu 0005, Alexander I. Rudnicky, Alan W. Black |
IJCAI | 2 |
| 2016 | Unsupervised user intent modeling by feature-enriched matrix factorizationabstractSpoken language interfaces are being incorporated into various devices such as smart phones and TVs. However, dialogue systems may fail to respond correctly when users' request functionality is not supported by currently installed apps. This paper proposes a feature-enriched matrix factorization (MF) approach to model open domain intents, which allows a system to dynamically add unexplored domains according to users' requests. First we leverage the structured knowledge from Wikipedia and Freebase to automatically acquire domain-related semantics to enrich features of input utterances, and then MF is applied to model automatically acquired knowledge, published app textual descriptions and users' spoken requests in a joint fashion; this generates latent feature vectors for utterances and user intents without need of prior annotations. Experiments show that the proposed MF models incorporated with rich features significantly improve intent prediction, achieving about 34% of mean average precision (MAP) for both ASR and manual transcripts. Yun-Nung Chen, Ming Sun 0001, Alexander I. Rudnicky, Anatole Gershman |
ICASSP | 3 |
| 2016 | An Intelligent Assistant for High-Level Task UnderstandingabstractPeople are able to interact with domain-specific intelligent assistants (IAs) and get help with tasks. But sometimes user goals are complex and may require interactions with multiple applications. However current IAs are limited to specific applications and users have to directly manage execution spanning multiple applications in order to engage in more complex activities. An ideal personal agent would be able to learn, over time, about tasks that span different resources. This paper addresses the problem of cross-domain task assistance in the context of spoken dialogue systems. We propose approaches to discover users' high-level intentions and using this information to assist users in their task. We collected real-life smartphone usage data from 14 participants and investigated how to extract high-level intents from users' descriptions of their activities. Our experiments show that understanding high-level tasks allows the agent to actively suggest apps relevant to pursuing particular user goals and reduce the cost of users' self-management. Ming Sun 0001, Yun-Nung Chen, Alexander I. Rudnicky |
IUI | 3 |
| 2016 | User Engagement Study with Virtual Agents Under Different Cultural Contexts
Zhou Yu 0005, Xinrui He, Alan W. Black, Alexander I. Rudnicky |
IVA | 4 |
| 2016 | AppDialogue: Multi-App Dialogues for Intelligent Assistants
Ming Sun 0001, Yun-Nung Chen, Zhenhao Hua, Yulian Tamres-Rudnicky, Arnab Dash, Alexander I. Rudnicky |
LREC | 6 |
| 2016 | A Wizard-of-Oz Study on A Non-Task-Oriented Dialog Systems That Reacts to User EngagementabstractIn this paper, we describe a system that reacts to both possible system breakdowns and low user engagement with a set of conversational strategies.These general strategies reduce the number of inappropriate responses and produce better user engagement.We also found that a system that reacts to both possible system breakdowns and low user engagement is rated by both experts and non-experts as having better overall user engagement compared to a system that only reacts to possible system breakdowns.We argue that for non-task-oriented systems we should optimize on both system response appropriateness and user engagement.We also found that apart from making the system response appropriate, funny and provocative responses can also lead to better user engagement.On the other hand, short appropriate responses, such as "Yes" or "No" can lead to decreased user engagement.We will use these findings to further improve our system. Zhou Yu 0005, Leah Nicolich-Henkin, Alan W. Black, Alexander I. Rudnicky |
SIGDIAL Conference | 4 |
| 2016 | Strategy and Policy Learning for Non-Task-Oriented Conversational SystemsabstractWe propose a set of generic conversational strategies to handle possible system breakdowns in non-task-oriented dialog systems.We also design policies to select these strategies according to dialog context.We combine expert knowledge and the statistical findings derived from data in designing these policies.The policy learned via reinforcement learning outperforms the random selection policy and the locally greedy policy in both simulated and real-world settings.In addition, we propose three metrics for conversation quality evaluation which consider both the local and global quality of the conversation. Zhou Yu 0005, Ziyu Xu 0001, Alan W. Black, Alexander I. Rudnicky |
SIGDIAL Conference | 4 |
| 2016 | Weakly supervised user intent detection for multi-domain dialoguesabstractUsers interact with mobile apps with certain intents such as finding a restaurant. Some intents and their corresponding activities are complex and may involve multiple apps; for example, a restaurant app, a messenger app and a calendar app may be needed to plan a dinner with friends. However, activities may be quite personal and third-party developers would not be building apps to specifically handle complex intents (e.g., a DinnerPlanner). Instead we want our intelligent agent to actively learn to understand these intents and provide assistance when needed. This paper proposes a framework to enable the agent to learn an inventory of intents from a small set of task-oriented user utterances. The experiments show that on previously unseen user activities, the agent is able to reliably recognize user intents using graph-based semi-supervised learning methods. The dataset, models, and the system outputs are available to research community. Ming Sun 0001, Aasish Pappu, Yun-Nung Chen, Alexander I. Rudnicky |
SLT | 4 |
| 2015 | Matrix Factorization with Knowledge Graph Propagation for Unsupervised Spoken Language UnderstandingabstractYun-Nung Chen, William Yang Wang, Anatole Gershman, Alexander Rudnicky. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Yun-Nung Chen, William Yang Wang, Anatole Gershman, Alexander I. Rudnicky |
ACL (1) | 4 |
| 2015 | Leveraging Behavioral Patterns of Mobile Applications for Personalized Spoken Language UnderstandingabstractSpoken language interfaces are appearing in various smart devices (e.g. smart-phones, smart-TV, in-car navigating systems) and serve as intelligent assistants (IAs). However, most of them do not consider individual users' behavioral profiles and contexts when modeling user intents. Such behavioral patterns are user-specific and provide useful cues to improve spoken language understanding (SLU). This paper focuses on leveraging the app behavior history to improve spoken dialog systems performance. We developed a matrix factorization approach that models speech and app usage patterns to predict user intents (e.g. launching a specific app). We collected multi-turn interactions in a WoZ scenario; users were asked to reproduce the multi-app tasks that they had performed earlier on their smart-phones. By modeling latent semantics behind lexical and behavioral patterns, the proposed multi-model system achieves about 52% of turn accuracy for intent prediction on ASR transcripts. Yun-Nung Chen, Ming Sun 0001, Alexander I. Rudnicky, Anatole Gershman |
ICMI | 3 |
| 2015 | Learning semantic hierarchy with distributed representations for unsupervised spoken language understandingabstractWe study the problem of unsupervised ontology learning for semantic understanding in spoken dialogue systems, in particular, learning the hierarchical semantic structure from the data. Given unlabelled conversations, we augment a frame-semantic based unsupervised slot induction approach with hierarchical agglomerative clustering to merge topically-related slots (e.g., both slots “direction” and “locale” convey location-related information) for building a coherent semantic hierarchy, and then estimate the slot importance at different levels. The high-level semantic estimation involves not only within-slot but also crossslot relations. The experiments show that high-level semantic information can accurately estimate the prominence of slots, significantly improving the slot induction performance; furthermore, a semantic decoder trained on the data with automatically extracted slots achieves about 68% F-measure, which is close to the one from hand-crafted grammars. Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky |
INTERSPEECH | 3 |
| 2015 | Distributed representation-based spoken word sense inductionabstractSpoken Term Detection (STD) or Keyword Search (KWS) techniques can locate keyword instances but do not differentiate between meanings. Spoken Word Sense Induction (SWSI) differentiates target instances by clustering according to context, providing a more useful result. In this paper we present a fully unsupervised SWSI approach based on distributed representations of spoken utterances. We compare this approach to several others, including the state-of-the-art Hierarchical Dirichlet Process (HDP). To determine how ASR performance affects SWSI, we used three different levels of Word Error Rate (WER), 40%, 20 % and 0%; 40 % WER is representative of online video, 0 % of text. We show that the distributed representation approach outperforms all other approaches, regardless of the WER. Although LDA-based approaches do well on clean data, they degrade significantly with WER. Paradoxically, lower WER does not guarantee better SWSI performance, due to the influence of common locutions. Justin T. Chiu, Yajie Miao, Alan W. Black, Alexander I. Rudnicky |
INTERSPEECH | 4 |
| 2015 | Learning OOV through semantic relatedness in spoken dialog systemsabstract• Speech recognition and language understanding performance can be improved through an OOV expectand-learn procedure. • A limited domain vocabulary can be utilized to effectively acquire OOVs by the word relatedness theory through web knowledge bases. • With data-driven semantic relatedness, both the global and local learning procedures are able to successfully harvest more than 50% of OOVs, leading to better recognition and understanding performance. • This work demonstrates that o OOV learning may benefit dialog system o the proposed expect-and-learn strategy outperforms the traditional detect-and-learn in both higher effectiveness and no human involvement. 1. Linguistically semantic relatedness o Defined by linguistics, e.g., WordNet (WN), Paraphrase Database (PPDB) (Ganitkevitch et al., 2013) 2. Data-driven semantic relatedness o Distributional semantics, e.g., continuous bag-ofword embeddings (CBOW) (Mikolov et al., 2013) Detect-and-Learn (Qin et al., 2011; 2012): o Discover OOV words during the conversation o Example: S: “I heard something like SELF, can you repeat it?” U: “It’s SELFIE.” o Drawbacks • Limited number of new words • Required human efforts to correct spellings and pronunciations Expect-and-Learn (proposed): o Use semantic relatedness to automatically enrich the vocabulary and language model beforehand Ming Sun 0001, Yun-Nung Chen, Alexander I. Rudnicky |
INTERSPEECH | 3 |
| 2015 | Jointly Modeling Inter-Slot Relations by Random Walk on Knowledge Graphs for Unsupervised Spoken Language UnderstandingabstractYun-Nung Chen, William Yang Wang, Alexander Rudnicky. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky |
HLT-NAACL | 3 |
| 2015 | Miscommunication Recovery in Physically Situated DialogueabstractWe describe an empirical study that crowdsourced human-authored recovery strategies for various problems encountered in physically situated dialogue.The purpose was to investigate the strategies that people use in response to requests that are referentially ambiguous or impossible to execute.Results suggest a general preference for including specific kinds of visual information when disambiguating referents, and for volunteering alternative plans when the original instruction was not possible to carry out. Matthew Marge, Alexander I. Rudnicky |
SIGDIAL Conference | 2 |
| 2014 | Two-Stage Stochastic Email SynthesizerabstractThis paper presents the design and im-plementation details of an email synthe-sizer using two-stage stochastic natural language generation, where the first stage structures the emails according to sender style and topic structure, and the second stage synthesizes text content based on the particulars of an email structure element and the goals of a given communication for surface realization. The synthesized emails reflect sender style and the intent of communication, which can be further used as synthetic evidence for developing other applications. 1 Yun-Nung Chen, Alexander I. Rudnicky |
INLG | 2 |
| 2014 | Two-Stage Stochastic Natural Language Generation for Email Synthesis by Modeling Sender Style and Topic StructureabstractThis paper describes a two-stage pro-cess for stochastic generation of email, in which the first stage structures the emails according to sender style and topic struc-ture (high-level generation), and the sec-ond stage synthesizes text content based on the particulars of an email element and the goals of a given communication (surface-level realization). Synthesized emails were rated in a preliminary experi-ment. The results indicate that sender style can be detected. In addition we found that stochastic generation performs better if applied at the word level than at an original-sentence level (“template-based”) in terms of email coherence, sentence flu-ency, naturalness, and preference. 1 Yun-Nung Chen, Alexander I. Rudnicky |
INLG | 2 |
| 2014 | Combination of FST and CN search in spoken term detectionabstractSpoken Term Detection (STD) focuses on finding instances of a particular spoken word or phrase in an audio corpus. Most STD systems have a two-step pipeline, ASR followed by search. Two approaches to search are common, Confusion Network (CN) based search and Finite State Transducer (FST) based search. In this paper, we examine combination of these two different search approaches, using the same ASR output. We find that the CN search performs better on shorter queries, and FST search performs better on longer queries. By combining the different search results from the same ASR decoding, we achieve better performance compared to either search approach on its own. We also find that this improvement is additive to the usual combination of decoder results using different modeling techniques. Justin T. Chiu, Jan Trmal, Daniel Povey, Guoguo Chen, Alexander I. Rudnicky |
INTERSPEECH | 6 |
| 2014 | Learning situated knowledge bases through dialogabstractTo respond to a user's query, dialog agents can use a knowledge base that is either domain specific, commonsense (e.g., NELL, Freebase) or a combination of both. The drawback is that domain-specific knowledge bases will likely be limited and static; commonsense ones are dynamic but contain general information found on the web and will be sparse with respect to a domain. We address this issue through a system that solicits situational information from its users in a domain that provides information on events (seminar talks) to augment its knowledge base (covering an academic field). We find that this knowledge is consistent and useful and that it provides reliable information to users. We show that, in comparison to a base system, users find that retrievals are more relevant when the system uses its informally acquired knowledge to augment their queries. Aasish Pappu, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2014 | Building a vocabulary self-learning speech recognition systemabstractThis paper presents initial studies on building a vocabulary self-learning speech recognition system that can automatically learn unknown words and expand its recognition vocabulary. Our recognizer can detect and recover out-of-vocabulary (OOV) words in speech, then incorporate OOV words into its lexicon and language model (LM). As a result, these unknown words can be correctly recognized when encountered by the recognizer in future. Specifically, we apply the word-fragment hybrid system framework to detect the presence of OOV words. We propose a better phoneme-to-grapheme (P2G) model so as to correctly recover the written form for more OOV words. Furthermore, we estimate LM scores for OOV words using their syntactic and semantic properties. The experimental results show that more than 40% OOV words are successfully learned from the development data, and about 60% learned OOV words are recognized in the testing data. Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2014 | Knowledge Acquisition Strategies for Goal-Oriented Dialog SystemsabstractMany goal-oriented dialog agents are expected to identify slot-value pairs in a spoken query, then perform lookup in a knowledge base to complete the task. When the agent encounters unknown slotvalues, it may ask the user to repeat or reformulate the query. But a robust agent can proactively seek new knowledge from a user, to help reduce subsequent task failures. In this paper, we propose knowledge acquisition strategies for a dialog agent and show their effectiveness. The acquired knowledge can be shown to subsequently contribute to task completion. Aasish Pappu, Alexander I. Rudnicky |
SIGDIAL Conference | 2 |
| 2014 | Dynamically supporting unexplored domains in conversational interactions by enriching semantics with neural word embeddingsabstractSpoken language interfaces are being incorporated into various devices (e.g. smart-phones, smart TVs, etc). However, current technology typically limits conversational interactions to a few narrow predefined domains/topics. For example, dialogue systems for smartphone operation fail to respond when users ask for functions not supported by currently installed applications. We propose to dynamically add application-based domains according to users' requests by using descriptions of applications as a retrieval cue to find relevant applications. The approach uses structured knowledge resources (e.g. Freebase, Wikipedia, FrameNet) to induce types of slots for generating semantic seeds, and enriches the semantics of spoken queries with neural word embeddings, where semantically related concepts can be additionally included for acquiring knowledge that does not exist in the predefined domains. The system can then retrieve relevant applications or dynamically suggest users install applications that support unexplored domains. We find that vendor descriptions provide a reliable source of information for this purpose. Yun-Nung Chen, Alexander I. Rudnicky |
SLT | 2 |
| 2014 | Leveraging frame semantics and distributional semantics for unsupervised semantic slot induction in spoken dialogue systemsabstractDistributional semantics and frame semantics are two representative views on language understanding in the statistical world and the linguistic world, respectively. In this paper, we combine the best of two worlds to automatically induce the semantic slots for spoken dialogue systems. Given a collection of unlabeled audio files, we exploit continuous-valued word embeddings to augment a probabilistic frame-semantic parser that identifies key semantic slots in an unsupervised fashion. In experiments, our results on a real-world spoken dialogue dataset show that the distributional word representations significantly improve the adaptation of FrameNet-style parses of ASR decodings to the target semantic space; that comparing to a state-of-the-art baseline, a 13% relative average precision improvement is achieved by leveraging word vectors trained on two 100-billion words datasets; and that the proposed technology can be used to reduce the costs for designing task-oriented spoken dialogue systems. Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky |
SLT | 3 |
| 2013 | Unsupervised induction and filling of semantic slots for spoken dialogue systems using frame-semantic parsingabstractSpoken dialogue systems typically use predefined semantic slots to parse users' natural language inputs into unified semantic representations. To define the slots, domain experts and professional annotators are often involved, and the cost can be expensive. In this paper, we ask the following question: given a collection of unlabeled raw audios, can we use the frame semantics theory to automatically induce and fill the semantic slots in an unsupervised fashion? To do this, we propose the use of a state-of-the-art frame-semantic parser, and a spectral clustering based slot ranking model that adapts the generic output of the parser to the target semantic space. Empirical experiments on a real-world spoken dialogue dataset show that the automatically induced semantic slots are in line with the reference slots created by domain experts: we observe a mean averaged precision of 69.36% using ASR-transcribed data. Our slot filling evaluations also indicate the promising future of this proposed approach. Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky |
ASRU | 3 |
| 2013 | Using web text to improve keyword spotting in speechabstractFor low resource languages, collecting sufficient training data to build acoustic and language models is time consuming and often expensive. But large amounts of text data, such as online newspapers, web forums or online encyclopedias, usually exist for languages that have a large population of native speakers. This text data can be easily collected from the web and then used to both expand the recognizer's vocabulary and improve the language model. One challenge, however, is normalizing and filtering the web data for a specific task. In this paper, we investigate the use of online text resources to improve the performance of speech recognition specifically for the task of keyword spotting. For the five languages provided in the base period of the IARPA BABEL project, we automatically collected text data from the web using only Limited LP resources. We then compared two methods for filtering the web data, one based on perplexity ranking and the other based on out-of-vocabulary (OOV) word detection. By integrating the web text into our systems, we observed significant improvements in keyword spotting accuracy for four out of the five languages. The best approach obtained an improvement in actual term weighted value (ATWV) of 0.0424 compared to a baseline system trained only on LimitedLP resources. On average, ATWV was improved by 0.0243 across five languages. Ankur Gandhe, Florian Metze, Alexander I. Rudnicky, Ian Lane, Matthias Eck 0001 |
ASRU | 4 |
| 2013 | Learning better lexical properties for recurrent OOV wordsabstractOut-of-vocabulary (OOV) words can appear more than once in a conversation or over a period of time. Such multiple instances of the same OOV word provide valuable information for learning the lexical properties of the word. Therefore, we investigated how to estimate better pronunciation, spelling and part-of-speech (POS) label for recurrent OOV words. We first identified recurrent OOV words from the output of a hybrid decoder by applying a bottom-up clustering approach. Then, multiple instances of the same OOV word were used simultaneously to learn properties of the OOV word. The experimental results showed that the bottom-up clustering approach is very effective at detecting the recurrence of OOV words. Furthermore, by using evidence from multiple instances of the same word, the pronunciation accuracy, recovery rate and POS label accuracy of recurrent OOV words can be substantially improved. Alexander I. Rudnicky |
ASRU | 2 |
| 2013 | An empirical investigation of sparse log-linear models for improved dialogue act classificationabstractPrevious work on dialogue act classification have primarily focused on dense generative and discriminative models. However, since the automatic speech recognition (ASR) outputs are often noisy, dense models might generate biased estimates and overfit to the training data. In this paper, we study sparse modeling approaches to improve dialogue act classification, since the sparse models maintain a compact feature space, which is robust to noise. To test this, we investigate various element-wise frequentist shrinkage models such as lasso, ridge, and elastic net, as well as structured sparsity models and a hierarchical sparsity model that embed the dependency structure and interaction among local features. In our experiments on a real-world dataset, when augmenting N-best word and phone level ASR hypotheses with confusion network features, our best sparse log-linear model obtains a relative improvement of 19.7% over a rule-based baseline, a 3.7% significant improvement over a traditional non-sparse log-linear model, and outperforms a state-of-the-art SVM model by 2.2%. Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky |
ICASSP | 3 |
| 2013 | Using conversational word bursts in spoken term detectionabstractWe describe a language independent word burst feature based on the structure of conversational speech that can be used to improve spoken term detection (STD) performance. Word burst refers to a phenomenon in conversational speech in which particular content words tend to occur in close proximity of each other as a byproduct of the topic under discussion. To take advantage of bursts, we describe a rescoring procedure that can be applied to lattice and confusion network outputs to improve STD performance. This approach is particularly effective when acoustic models are built with limited training data (and ASR performance is relatively poor). We find that word bursts appear in the four languages we examined and that STD performance can be improved for three of them; the remaining language is agglutinative. Justin T. Chiu, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2013 | Finding recurrent out-of-vocabulary wordsabstractOut-of-vocabulary (OOV) words can appear more than once in a conversation or over a period of time. Such multiple instances of the same OOV word provide valuable information for estimating the pronunciation or the part-of-speech (POS) tag of the word. But in a conventional OOV word detection system, each OOV word is recognized and treated individually. We therefore investigated how to identify recurrent OOV words in speech recognition. Specifically, we propose to cluster multiple instances of the same OOV word using a bottom-up approach. Phonetic, acoustic and contextual features were collected to measure the distance between OOV candidates. The experimental results show that the bottom-up clustering approach is very effective at detecting the recurrence of OOV words. We also found that the phonetic feature is better than the acoustic and contextual features, and the best performance is achieved when combining all features. Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2013 | Emotion Recognition Modulating the Behavior of Intelligent SystemsabstractThe paper presents an audio-based emotion recognition system that is able to classify emotions as anger, fear, happy, neutral, sadness or disgust in real time. We use the virtual coach as an application example of how emotion recognition can be used to modulate intelligent systems' behavior. A novel minimum-error feature removal mechanism to reduce bandwidth and increase accuracy of our emotion recognition system has been introduced. A two-stage hierarchical classification approach along with a One-Against-All (OAA) framework are used. We obtained an average accuracy of 82.07% using the OAA approach, and 87.70% with a two-stage hierarchical approach, by pruning the feature set and using Support Vector Machines (SVMs) for classification. Asim Smailagic, Daniel P. Siewiorek, Alexander I. Rudnicky, Sandeep Nallan Chakravarthula, Anshuman Kar, Nivedita Jagdale, Saksham Gautam, Rohit Vijayaraghavan, Shaurya Jagtap |
ISM | 3 |
| 2013 | SiMPE: 8th workshop on speech and sound in mobile and pervasive environmentsabstractThe SiMPE workshop series started in 2006 with the goal of enabling speech processing on mobile and embedded devices. The SiMPE 2012 workshop extended the notion of audio to non-speech "Sounds" and thus the expansion became "Speech and Sound". SiMPE 2010 and 2011 brought together researchers from the speech and the HCI communities. Speech User interaction in cars was a focus area in 2009. Multimodality got more attention in SiMPE 2008. In SiMPE 2007, the focus was on developing regions. Amit Anil Nanavati, Nitendra Rajput, Cumhur Erkut, Antti Jylhä, Alexander I. Rudnicky, Stefania Serafin, Markku Turunen |
Mobile HCI | 6 |
| 2013 | Towards evaluating recovery strategies for situated grounding problems in human-robot dialogueabstractRobots can use information from their surroundings to improve spoken language communication with people. Even when speech recognition is correct, robots face challenges when interpreting human instructions. These situated grounding problems include referential ambiguities and impossible-to-execute instructions. We present an approach to resolving situated grounding problems through spoken dialogue recovery strategies that robots can invoke to repair these problems. We describe a method for evaluating these strategies in human-robot navigation scenarios. Matthew Marge, Alexander I. Rudnicky |
RO-MAN | 2 |
| 2013 | Predicting Tasks in Goal-Oriented Spoken Dialog Systems using Semantic Knowledge Bases
Aasish Pappu, Alexander I. Rudnicky |
SIGDIAL Conference | 2 |
| 2012 | System combination for out-of-vocabulary word detectionabstractThis paper presents a method to improve the out-of-vocabulary (OOV) word detection performance by combining multiple speech recognition systems' outputs. Three different fragment-word hybrid systems, the phone, subword, and graphone systems, were built for detecting OOV words. Then outputs from each individual system were combined using ROVER. Two combination metrics were explored in ROVER, voting by word frequency and voting by both word frequency and word confidence score. The experimental results show that the OOV word detection performance of the ROVER system with confidence scores is better than the ROVER system with only word frequency, as well as any of the individual hybrid systems. Ming Sun 0001, Alexander I. Rudnicky |
ICASSP | 3 |
| 2012 | NeuroDialog: an EEG-enabled spoken dialog interfaceabstractUnderstanding user intent is a difficult problem in Dialog Systems, as they often need to make decisions under uncertainty. Using an inexpensive, consumer grade EEG sensor and a Wizard-of-Oz dialog system, we show that it is possible to detect system misunderstanding even before the user reacts vocally. We also present the design and implementation details of NeuroDialog, a proof-of-concept dialog system that uses an EEG based predictive model to detect system misrecognitions during live interaction. Seshadri Sridharan, Yun-Nung Chen, Kai-min Chang, Alexander I. Rudnicky |
ICMI | 4 |
| 2012 | OOV Word Detection using Hybrid Models with Mixed Types of FragmentsabstractThis paper presents initial studies to improve the out-of-vocabulary (OOV) word detection performance by using mixed types of fragment units in one hybrid system. Three types of fragment units, subwords, syllables, and graphones, were combined in two different ways to build the hybrid lexicon and language model. The experimental results show that hybrid systems with mixed types of fragment units perform better than hybrid systems using only one type of fragment unit. After comparing the OOV word detection performance with the number and length of fragment units of each system, we proposed future work to better utilize mixed types of fragment units in a hybrid system. Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2012 | The Structure and Generality of Spoken Route Instructions
Aasish Pappu, Alexander I. Rudnicky |
SIGDIAL Conference | 2 |
| 2011 | OOV Detection and Recovery Using Hybrid Models with Different FragmentsabstractIn this paper, we address the out-of-vocabulary (OOV) detection and recovery problem by developing three different fragment-word hybrid systems. A fragment language model (LM) and a word LM were trained separately and then combined into a single hybrid LM. Using this hybrid model, the recognizer can recognize any OOVs as fragment sequences. Different types of fragments, such as phones, subwords, and graphones were tested and compared on the WSJ 5k and 20k evaluation sets. The experiment results show that the subword and graphone hybrid systems perform better than the phone hybrid system in both 5k and 20k tasks. Furthermore, given less training data, the subword hybrid system is more preferable than the graphone hybrid system. Ming Sun 0001, Alexander I. Rudnicky |
INTERSPEECH | 3 |
| 2011 | SiMPE: 6th Workshop on Speech in Mobile and Pervasive EnvironmentsabstractWith the proliferation of pervasive devices and the increase in their processing capabilities, client-side speech processing has been emerging as a viable alternative. The SiMPE workshop series started in 2006 [5] with the goal of enabling speech processing on mobile and embedded devices to meet the challenges of pervasive environments (such as noise) and leveraging the context they offer (such as location). SiMPE 2010, the latest in the series brought together, very successfully, researchers from the speech and the HCI communities. We believe this is the beginning. Amit Anil Nanavati, Nitendra Rajput, Alexander I. Rudnicky, Markku Turunen, Andrew L. Kun, Tim Paek, Ivan Tashev |
Mobile HCI | 3 |
| 2010 | Using the Amazon Mechanical Turk for transcription of spoken languageabstractWe investigate whether Amazon's Mechanical Turk (MTurk) service can be used as a reliable method for transcription of spoken language data. Utterances with varying speaker demographics (native and non-native English, male and female) were posted on the MTurk marketplace together with standard transcription guidelines. Transcriptions were compared against transcriptions carefully prepared in-house through conventional (manual) means. We found that transcriptions from MTurk workers were generally quite accurate. Further, when transcripts for the same utterance produced by multiple workers were combined using the ROVER voting scheme, the accuracy of the combined transcript rivaled that observed for conventional transcription methods. We also found that accuracy is not particularly sensitive to payment amount, implying that high quality results can be obtained at a fraction of the cost and turnaround time of conventional methods. Matthew Marge, Satanjeev Banerjee, Alexander I. Rudnicky |
ICASSP | 3 |
| 2010 | The effect of lattice pruning on MMIE trainingabstractIn discriminative training, such as Maximum Mutual Information Estimation (MMIE) training, a word lattice is usually used as a compact representation of many different sentence hypotheses and hence provides an efficient representation of the confusion data. However, in a large vocabulary continuous speech recognition (LVCSR) system trained from hundreds or thousands hours training data, the extended Baum-Welch (EBW) computation on the word lattice is still very expensive. In this paper, we investigated the effect of lattice pruning on MMIE training, where we tested the MMIE performance trained with different lattice complexity. A beam pruning and a posterior probability pruning method were applied to generate different sizes of word lattices. The experimental results show that using the posterior probability lattice pruning algorithm, we can save about 40% of the total computation and get the same or more improvement compared to the baseline MMIE result. Alexander I. Rudnicky |
ICASSP | 2 |
| 2010 | SiMPE: 5th workshop on speech in mobile and pervasive environmentsabstractWith the proliferation of pervasive devices and the increase in their processing capabilities, client-side speech processing has been emerging as a viable alternative. The SiMPE workshop series started in 2006 [5] with the goal of enabling speech processing on mobile and embedded devices to meet the challenges of pervasive environments (such as noise) and leveraging the context they offer (such as location). Amit Anil Nanavati, Nitendra Rajput, Alexander I. Rudnicky, Markku Turunen, Andrew L. Kun, Tim Paek, Ivan Tashev |
Mobile HCI | 3 |
| 2010 | Towards Improving the Naturalness of Social Conversations with Dialogue Systems
Matthew Marge, João Miranda, Alan W. Black, Alexander I. Rudnicky |
SIGDIAL Conference | 4 |
| 2010 | Comparing Spoken Language Route Instructions for Robots across Environment Representations
Matthew Marge, Alexander I. Rudnicky |
SIGDIAL Conference | 2 |
| 2010 | Let's Buy Books: Finding eBooks using voice searchabstractWe describe Let's Buy Books, a dialog system that helps users search for eBook titles. In this paper we compare different vector space approaches to voice search and find that a hybrid approach using a weighted sub-space model smoothed with a general model provides the best performance over different conditions and evaluated using both synthetic queries and queries collected from users through questionnaires. Cheongjae Lee, Alexander I. Rudnicky, Gary Geunbae Lee |
SLT | 2 |
| 2009 | Combining mixture weight pruning and quantization for small-footprint speech recognitionabstractSemi-continuous acoustic models, where the output distributions for all Hidden Markov Model states share a common codebook of Gaussian density functions, are a well-known and proven technique for reducing computation in automatic speech recognition. However, the size of the parameter files, and thus their memory footprint at runtime, can be very large. We demonstrate how non-linear quantization can be combined with a mixture weight distribution pruning technique to halve the size of the models with minimal performance overhead and no increase in error rate. David Huggins-Daines, Alexander I. Rudnicky |
ICASSP | 2 |
| 2009 | SiMPE: Fourth Workshop on Speech in Mobile and Pervasive Environments
Amit Anil Nanavati, Nitendra Rajput, Alexander I. Rudnicky, Markku Turunen, Andrew L. Kun, Tim Paek, Ivan Tashev |
Mobile HCI | 3 |
| 2009 | Detecting the Noteworthiness of Utterances in Human Meetings
Satanjeev Banerjee, Alexander I. Rudnicky |
SIGDIAL Conference | 2 |
| 2009 | The RavenClaw dialog management framework: Architecture and systems
Dan Bohus, Alexander I. Rudnicky |
Comput. Speech Lang. | 2 |
| 2008 | Acquiring Domain-Specific Dialog Information from Task-Oriented Human-Human Interaction through an Unsupervised Learning
Ananlada Chotimongkol, Alexander I. Rudnicky |
EMNLP | 2 |
| 2008 | Automatic Extraction of Briefing Templates
Dipanjan Das 0001, Alexander I. Rudnicky |
IJCNLP | 3 |
| 2008 | SiMPE: third workshop on speech in mobile and pervasive environmentsabstractIn the past, voice-based applications have been accessed using unintelligent telephone devices through Voice Browsers that reside on the server. The proliferation of pervasive devices and the increase in their processing capabilities, clientside speech processing has been emerging as a viable alternative. In SiMPE 2008, the third in the series, we will continue to explore the various possibilities and issues that arise while enabling speech processing on resource-constrained, possibly mobile devices. Amit Anil Nanavati, Nitendra Rajput, Alexander I. Rudnicky, Markku Turunen |
Mobile HCI | 3 |
| 2008 | An extractive-summarization baseline for the automatic detection of noteworthy utterances in multi-party human-human dialogabstractOur goal is to reduce meeting participants' note-taking effort by automatically identifying utterances whose contents meeting participants are likely to include in their notes. Though note-taking is different from meeting summarization, these two problems are related. In this paper we apply techniques developed in extractive meeting summarization research to the problem of identifying noteworthy utterances. We show that these algorithms achieve an f-measure of 0.14 over a 5-meeting sequence of related meetings. The precision - 0.15 - is triple that of the trivial baseline of simply labeling every utterance as noteworthy. We also introduce the concept of ldquoshow-worthyrdquo utterances - utterances that contain information that could conceivably result in a note. We show that such utterances can be recognized with an 81% accuracy (compared to 53% accuracy of a majority classifier). Further, if non-show-worthy utterances are filtered out, the precision of noteworthiness detection improves by 33% relative. Satanjeev Banerjee, Alexander I. Rudnicky |
SLT | 2 |
| 2007 | TeamTalk: A Platform for Multi-Human-Robot Dialog Research in Coherent Real and Virtual Spaces
Thomas K. Harris, Alexander I. Rudnicky |
AAAI | 2 |
| 2007 | Data selection for speech recognitionabstractThis paper presents a strategy for efficiently selecting informative data from large corpora of transcribed speech. We propose to choose data uniformly according to the distribution of some target speech unit (phoneme, word, character, etc). In our experiment, in contrast to the common belief that "there is no data like more data", we found it possible to select a highly informative subset of data that produces recognition performance comparable to a system that makes use of a much larger amount of data. At the same time, our selection process is efficient and fast. Yi Wu 0002, Rong Zhang 0003, Alexander I. Rudnicky |
ASRU | 3 |
| 2007 | Learning from the Report-writing Behavior of Individuals
Nikesh Garera, Alexander I. Rudnicky |
IJCAI | 3 |
| 2007 | Segmenting meetings into agenda items by extracting implicit supervision from human note-takingabstractSplitting a meeting into segments such that each segment contains discussions on exactly one agenda item is useful for tasks such as retrieval and summarization of agenda item discussions. However, accurate topic segmentation of meetings is a difficult task. In this paper, we investigate the idea of acquiring implicit supervision from human meeting participants to solve the segmentation problem. Specifically we have implemented and tested a note taking interface that gives value to users by helping them organize and retrieve their notes easily, but that also extracts a segmentation of the meeting based on note taking behavior. We show that the segmentation so obtained achieves a Pk value of 0.212 which improves upon an unsupervised baseline by 45% relative, and compares favorably with a current state–of–the–art algorithm. Most importantly, we achieve this performance without any features or algorithms in the classic sense. Satanjeev Banerjee, Alexander I. Rudnicky |
IUI | 2 |
| 2006 | Pocketsphinx: A Free, Real-Time Continuous Speech Recognition System for Hand-Held DevicesabstractThe availability of real-time continuous speech recognition on mobile and embedded devices has opened up a wide range of research opportunities in human-computer interactive applications. Unfortunately, most of the work in this area to date has been confined to proprietary software, or has focused on limited domains with constrained grammars. In this paper, we present a preliminary case study on the porting and optimization of CMU Sphinx-11, a popular open source large vocabulary continuous speech recognition (LVCSR) system, to hand-held devices. The resulting system operates in an average 0.87 times real-time on a 206 MHz device, 8.03 times faster than the baseline system. To our knowledge, this is the first hand-held LVCSR system available under an open-source license David Huggins-Daines, Arthur Chan, Alan W. Black, Mosur Ravishankar, Alexander I. Rudnicky |
ICASSP (1) | 6 |
| 2006 | A New Data Selection Approach for Semi-Supervised Acoustic ModelingabstractCurrent approaches to semi-supervised incremental learning prefer to select unlabeled examples predicted with high confidence for model re-training. However, this strategy can degrade the classification performance rather than improve it. We present an analysis for the reasons of this phenomenon, showing that only relying on high confidence for data selection can lead to an erroneous estimate to the true distribution when the confidence annotator is highly correlated with the classifier in the information they use. We propose a new data selection approach to address this problem and apply it to a variety of applications, including machine learning and speech recognition. Encouraging improvements in recognition accuracy are observed in our experiments Rong Zhang 0003, Alexander I. Rudnicky |
ICASSP (1) | 2 |
| 2006 | A Briefing Tool that Learns Individual Report-Writing BehaviorabstractWe describe a briefing system that learns to predict the contents of reports generated by users who create periodic (weekly) reports as part of their normal activity. We address the question whether data derived from the implicit supervision provided by end-users is robust enough to support not only model parameter tuning but also a form of feature discovery. The system was evaluated under realistic conditions, by collecting data in a project-based university course where student group leaders were tasked with preparing weekly reports for the benefit of the instructors, using the material from individual student reports Nikesh Garera, Alexander I. Rudnicky |
ICTAI | 3 |
| 2006 | A texttiling based approach to topic boundary detection in meetingsabstractOur goal is to automatically detect boundaries between discussions of different topics in meetings. Towards this end we adapt the TextTiling algorithm [1] to the context of meetings. Our features include not only the overlapped words between adjacent windows, but also overlaps in the amount of speech contributed by each meeting participant. We evaluate our algorithm by comparing the automatically detected boundaries with the true ones, and computing precision, recall and f–measure. We report average precision of 0.85 and recall of 0.59 when segmenting unseen test meetings. Error analysis of our results shows that although the basic idea of our algorithm is sound, it breaks down when participants stray from typical behavior (such as when they monopolize the conversation for too long). Satanjeev Banerjee, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2006 | A constrained baum-welch algorithm for improved phoneme segmentation and efficient trainingabstractWe describe an extension to the Baum-Welch algorithm for training Hidden Markov Models that uses explicit phoneme segmentation to constrain the forward and backward lattice. The HMMs trained with this algorithm can be shown to improve the accuracy of automatic phoneme segmentation. In addition, this algorithm is significantly more computationally efficient than the full BaumWelch algorithm, while producing models that achieve equivalent accuracy on a standard phoneme recognition task. David Huggins-Daines, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2006 | Investigations of issues for using multiple acoustic models to improve continuous speech recognitionabstractThis paper investigates two important issues in constructing and combining ensembles of acoustic models for reducing recognition errors. First, we investigate the applicability of the AnyBoost algorithm for acoustic model training. AnyBoost is a generalized Boosting method that allows the use of an arbitrary loss function as the training criterion to construct ensemble of classifiers. We choose the MCE discriminative objective function for our experiments. Initial test results on a real-world meeting recognition corpus show that AnyBoost is a competitive alternate to the standard AdaBoost algorithm. Second, we investigate ROVER-based combination, focusing on the technique for selecting correct hypothesized words from aligned WTN. We propose a neural network based insertion detection and word scoring scheme for this. Our approach consistently outperforms the current voting technique used by ROVER in the experiments. Rong Zhang 0003, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2006 | SiMPE: speech in mobile and pervasive environmentsabstractTraditionally, voice-based applications have been accessed using unintelligent telephone devices through Voice Browsers that reside on the server. The proliferation of pervasive devices and the increase in their processing capabilities, client-side speech processing is emerging as a viable alternative. This workshop will explore the various possibilities and issues that arise while enabling speech processing on resource-constrained, possibly mobile devices. The workshop will highlight the many open areas that require research attention, identify key problems that need to be addressed, and also discuss a few approaches for solving some of them - to build the next generation of conversational systems. Amit Anil Nanavati, Nitendra Rajput, Alexander I. Rudnicky, Roberto Sicconi |
Mobile HCI | 3 |
| 2006 | SmartNotes: Implicit Labeling of Meeting Data through User Note-Taking and Browsing
Satanjeev Banerjee, Alexander I. Rudnicky |
HLT-NAACL | 2 |
| 2006 | Online Supervised Learning of Non-Understanding Recovery PoliciesabstractSpoken dialog systems typically use a limited number of non- understanding recovery strategies and simple heuristic policies to engage them (e.g. first ask user to repeat, then give help, then transfer to an operator). We propose a supervised, online method for learning a non-understanding recovery policy over a large set of recovery strategies. The approach consists of two steps: first, we construct runtime estimates for the likelihood of success of each recovery strategy, and then we use these estimates to construct a policy. An experiment with a publicly available spoken dialog system shows that the learned policy produced a 12.5% relative improvement in the non-understanding recovery rate. Dan Bohus, Brian Langner, Antoine Raux, Alan W. Black, Maxine Eskénazi, Alexander I. Rudnicky |
SLT | 6 |
| 2005 | The Necessity of a Meeting Recording and Playback System, and the Benefit of Topic-Level Annotations to Meeting Browsing
Satanjeev Banerjee, Carolyn P. Rosé, Alexander I. Rudnicky |
INTERACT | 3 |
| 2005 | A principled approach for rejection threshold optimization in spoken dialog systemsabstractA common design pattern in spoken dialog systems is to reject an input when the recognition confidence score falls below a preset rejection threshold. However, this introduces a potentially non-optimal tradeoff between various types of errors such as misunderstandings and false rejections. In this paper, we propose a data-driven method for determining the relative costs of these errors, and then use these costs to optimize state-specific rejection thresholds. We illustrate the use of this approach with data from a spoken dialog system that handles conference room reservations. The results obtained confirm our intuitions about the costs of the errors, and are consistent with anecdotal evidence gathered throughout the use of the system. Dan Bohus, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2005 | On improvements to CI-based GMM selectionabstractGaussian Mixture Model (GMM) computation is known to be one of the most computation-intensive components in speech decoding. In our previous work, context-independent model based GMM selection (CIGMMS) was found to be an effective way to reduce the cost of GMM computation without significant loss in recognition accuracy. In this work, we propose three methods to further improve the performance of CIGMMS. Each method brings an additional 5-10% relative speed improvement, with a cumulative improvement up to 37% on some tasks. Detailed analysis and experimental results on three corpora are presented. Arthur Chan, Mosur Ravishankar, Alexander I. Rudnicky |
INTERSPEECH | 3 |
| 2005 | Investigations on ensemble based semi-supervised acoustic model trainingabstractComputer Science Department Rong Zhang 0003, Ziad Al Bawab, Arthur Chan, Ananlada Chotimongkol, David Huggins-Daines, Alexander I. Rudnicky |
INTERSPEECH | 6 |
| 2004 | Segmentation and classification of meetings using multiple information streamsabstractWe present a meeting recorder infrastructure used to record and annotate events that occur in meetings. Multiple data streams are recorded and analyzed in order to infer a higher-level state of the group’s activities. We describe the hardware and software systems used to capture people’s activities as well as the methods used to characterize them. Paul E. Rybski, Satanjeev Banerjee, Fernando De la Torre, Carlos Vallespí, Alexander I. Rudnicky, Manuela M. Veloso |
ICMI | 5 |
| 2004 | Using simple speech-based features to detect the state of a meeting and the roles of the meeting participantsabstractWe introduce a simple taxonomy of meeting states and participant roles. Our goal is to automatically detect the state of a meeting and the role of each meeting participant and to do so concurrent with a meeting. We trained a decision tree classifier that learns to detect these states and roles from simple speech–based features that are easy to compute automatically. This classifier detects meeting states 18% absolute more accurately than a random classifier, and detects participant roles 10% absolute more accurately than a majority classifier. The results imply that simple, easy to compute features can be used for this purpose. Satanjeev Banerjee, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2004 | Four-layer categorization scheme of fast GMM computation techniques in large vocabulary continuous speech recognition systemsabstractLarge vocabulary continuous speech recognition systems are known to be computationally intensive. A major bottleneck is the Gaussian mixture model (GMM) computation and various techniques have been proposed to address this problem. We present a systematic study of fast GMM computation techniques. As there are a large number of these and it is impractical to exhaustively evaluate all of them, we first categorized techniques into four layers and selected representative ones to evaluate in each layer. Based on this framework of study, we provide a detailed analysis and comparison of GMM computation techniques from the four-layer perspective and explore two subtle practical issues, 1) how different techniques can be combined effectively and 2) how beam pruning will affect the performance of GMM computation techniques. All techniques are evaluated in the CMU Communicator domain. We also compare their performance with others reported in the literature. 1. Arthur Chan, Mosur Ravishankar, Alexander I. Rudnicky, Jahanzeb Sherwani |
INTERSPEECH | 3 |
| 2004 | A frame level boosting training scheme for acoustic modelingabstractConventional Boosting algorithms for acoustic modeling have two notable weaknesses. (1) The objective function aims to minimize utterance error rate, though the goal for most speech recognition systems is to reduce word error rate. (2) During Boosting training, an utterance is treated as a unit for resampling and each frame within the same utterance is assigned equal weight. Intuitively, the frames associated with a is classified word should be given more emphasis than others. We propose a frame level Boosting training scheme that addresses these shortcomings and allows each frame to have a different weight. We describe a technique and provide experimental results for this approach. Rong Zhang 0003, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2004 | Optimizing boosting with discriminative criteriaabstractWe describe the use of discriminative criteria to optimize Boosting based ensembles. Boosting algorithms may create hundreds of individual classifiers in order to fit the training data. However, this strategy isn’t feasible and necessary for complex classification problems, such as real-time continuous speech recognition, in which only the combination of a few of acoustic models is practical. How to improve the classification accuracy for small size of ensemble is the focus of this paper. Two discriminative criteria that attempt to minimize the true Bayes error rate are investigated. Improvements are observed over a variety of datasets including image and speech recognition, indicating the prospective utility of these two criteria. Rong Zhang 0003, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2004 | Apply n-best list re-ranking to acoustic model combinations of boosting trainingabstractThe object function for Boosting training method in acoustic modeling aims to reduce utterance level error rate. This is different from the most commonly used performance metric in speech recognition, word error rate. This paper proposes that the combination of N-best list re-ranking and ROVER can partly address this problem. In particular, model combination is applied to re-ranked hypotheses rather than to the original top-1 hypotheses and carried on word level. Improvement of system performance is observed in our experiments. In addition, we describe and evaluate a new confidence feature that measures the correctness of frame level decoding result. Rong Zhang 0003, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2003 | Improving the performance of an LVCSR system through ensembles of acoustic modelsabstractThis paper describes our work on applying ensembles of acoustic models to the problem of large vocabulary continuous speech recognition (LVCSR). We propose three algorithms for constructing ensembles. The first two have their roots in bagging algorithms; however, instead of randomly sampling examples our algorithms construct training sets based on the word error rate. The third one is a boosting style algorithm. Different from other boosting methods which demand large resources for computation and storage, our method present a more efficient solution suitable for acoustic model training. We also investigate a method that seeks optimal combination for models. We report experimental results on a large real world corpus collected from the Carnegie Mellon Communicator dialog system. Significant improvements on system performance are observed in that up to 15.56% relative reduction on word error rate is achieved. Rong Zhang 0003, Alexander I. Rudnicky |
ICASSP (1) | 2 |
| 2003 | Ravenclaw: dialog management using hierarchical task decomposition and an expectation agendaabstractWe describe RavenClaw, a new dialog management framework developed as a successor to the Agenda [1] architecture used in the CMU Communicator. RavenClaw introduces a clear separation between task and discourse behavior specification, and allows rapid development of dialog management components for spoken dialog systems operating in complex, goal-oriented domains. The system development effort is focused entirely on the specification of the dialog task, while a rich set of domain-independent conversational behaviors are transparently generated by the dialog engine. To date, RavenClaw has been applied to five different domains allowing us to draw some preliminary conclusions as to the generality of the approach. We briefly describe our experience in developing these systems. Dan Bohus, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2003 | Comparative study of boosting and non-boosting training for constructing ensembles of acoustic modelsabstractThis paper compares the performance of Boosting and non-Boosting training algorithms in large vocabulary continuous speech recognition (LVCSR) using ensembles of acoustic models. Both algorithms demonstrated significant word error rate reduction on the CMU Communicator corpus. However, both algorithms produced comparable improvements, even though one would expect that the Boosting algorithm, which has a solid theoretic foundation, should work much better than the non-Boosting algorithm. Several voting schemes for hypothesis combining were evaluated, including weighted voting, un-weighted voting and ROVER. 1. Rong Zhang 0003, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2002 | Building voiceXML-based applicationsabstractThe Language Technologies Institute (LTI) at Carnegie Mellon University has, for the past several years, conducted a lab course in building spoken-language dialog systems. In the most recent versions of the course, we have used (commercial) web-based development environments to build systems. This paper describes our experiences and discusses the characteristics of applications that are developed within this framework. Christina L. Bennett, Ariadna Font Llitjós, Stefanie Shriver, Alexander I. Rudnicky, Alan W. Black |
INTERSPEECH | 4 |
| 2002 | The carnegie mellon communicator corpusabstractAs part of the DARPA Communicator program, Carnegie Mellon has, over the past three years, collected a large corpus of speech produced by callers to its Travel Planning system. To date, a total of 180,605 utterances (90.9 hours) have been collected. The data were used for a number of purposes, including acoustic and language modeling and the development of a spoken dialog system. The collection, transcription and annotation of these data prompted us to develop a number of procedures for managing the transcription process and for ensuring accuracy. We describe these, as well as some results based on these data. A portion of this corpus, covering the years 1999-2001, is being published for research purposes. 1. Christina L. Bennett, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2002 | Rapid development of speech-to-speech translation systems
Alan W. Black, Ralf D. Brown, Robert E. Frederking, Kevin A. Lenzo, John Moody, Alexander I. Rudnicky, Rita Singh, Eric Steinbrecher |
INTERSPEECH | 6 |
| 2002 | Automatic concept identification in goal-oriented conversationsabstractWe address the problem of identifying key domain concepts automatically from an unannotated corpus of goal-oriented human-human conversations. We examine two clustering algorithms, one based on mutual information and another one based on Kullback-Liebler distance. In order to compare the results from both techniques quantitatively, we evaluate the outcome clusters against reference concept labels using precision and recall metrics adopted from the evaluation of topic identification task. However, since our system allows more than one cluster to associate with each concept an additional metric, a singularity score, is added to better capture cluster quality. Based on the proposed quality metrics, the results show that Kullback-Liebler-based clustering outperforms mutual information-based clustering for both the optimal quality and the quality achieved using an automatic stopping criterion Ananlada Chotimongkol, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2002 | DARPA communicator evaluation: progress from 2000 to 2001abstractThis paper describes the evaluation methodology and results of the DARPA Communicator spoken dialog system evaluation experiments in 2000 and 2001. Nine spoken dialog systems in the travel planning domain participated in the experiments resulting in a total corpus of 1904 dialogs. We describe and compare the experimental design of the 2000 and 2001 DARPA evaluations. We describe how we established a performance baseline in 2001 for complex tasks. We present our overall approach to data collection, the metrics collected, and the application of PARADISE to these data sets. We compare the results we achieved in 2000 for a number of core metrics with those for 2001. These results demonstrate large performance improvements from 2000 to 2001 and show that the Communicator program goal of conversational interaction for complex tasks has been achieved. Marilyn A. Walker, Alexander I. Rudnicky, John S. Aberdeen, Elizabeth Owen Bratt, John S. Garofolo, Helen Hastie, Audrey N. Le, Bryan L. Pellom, Alexandros Potamianos, Rebecca J. Passonneau, Rashmi Prasad, Salim Roukos, Gregory A. Sanders, Stephanie Seneff, David Stallard |
INTERSPEECH | 2 |
| 2002 | DARPA communicator: cross-system results for the 2001 evaluationabstractThis paper describes the evaluation methodology and results of the 2001 DARPA Communicator evaluation. The experiment spanned 6 months of 2001 and involved eight DARPA Communicator systems in the travel planning domain. It resulted in a corpus of 1242 dialogs which include many more dialogues for complex tasks than the 2000 evaluation. We describe the experimental design, the approach to data collection, and the results. We compare the results by the type of travel plan and by system. The results demonstrate some large differences across sites and show that the complex trips are clearly more difficult. Marilyn A. Walker, Alexander I. Rudnicky, Rashmi Prasad, John S. Aberdeen, Elizabeth Owen Bratt, John S. Garofolo, Helen Hastie, Audrey N. Le, Bryan L. Pellom, Alexandros Potamianos, Rebecca J. Passonneau, Salim Roukos, Gregory A. Sanders, Stephanie Seneff, David Stallard |
INTERSPEECH | 2 |
| 2002 | Improve latent semantic analysis based language model by integrating multiple level knowledgeabstractWe describe an extension to the use of Latent Semantic Analysis (LSA) for language modeling. This technique makes it easier to exploit long distance relationships in natural language for which the traditional n-gram is unsuited. However, with the growth of length, the semantic representation of the history may be contaminated by irrelevant information, increasing the uncertainty in predicting the next word. To address this problem, we propose a multilevel framework dividing the history into three levels corresponding to document, paragraph and sentence. To combine the three levels of information with the n-gram, a Softmax network is used. We further present a statistical scheme that dynamically determines the unit scope in the generalization stage. The combination of all the techniques leads to a 14% perplexity reduction on a subset of Wall Street Journal, compared with the trigram model. Rong Zhang 0003, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2002 | Stochastic natural language generation for spoken dialog systems
Alice Oh, Alexander I. Rudnicky |
Comput. Speech Lang. | 2 |
| 2001 | Is this conversation on track?abstractConfidence annotation allows a spoken dialog system to accurately assess the likelihood of misunderstanding at the utterance level and to avoid breakdowns in interaction. We describe experiments that assess the utility of features from the decoder, parser and dialog levels of processing. We also investigate the effectiveness of various classifiers, including Bayesian Networks, Neural Networks, SVMs, Decision Trees, AdaBoost and Naive Bayes, to combine this information into an utterancelevel confidence metric. We found that a combination of a subset of the features considered produced promising results with several of the classification algorithms considered, e.g., our Bayesian Network classifier produced a 45.7% relative reduction in confidence assessment error and a 29.6% reduction relative to a handcrafted rule. Paul Carpenter 0001, Chun Jin, Rong Zhang 0003, Dan Bohus, Alexander I. Rudnicky |
INTERSPEECH | 6 |
| 2001 | N-best speech hypotheses reordering using linear regressionabstractWe propose a hypothesis reordering technique to improve speech recognition accuracy in a dialog system. For such systems, additional information external to the decoding process itself is available, in particular features derived from the parse and the dialog. Such features can be combined with recognizer features by means of a linear regression model to predict the most likely entry in the hypothesis list. We introduce the use of concept error rate as an alternative accuracy measurement and compare it withy the use of word error rate. The proposed model performs better than human subjects performing the same hypothesis reordering task. Ananlada Chotimongkol, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2001 | Universalizing speech: notes from the USI projectabstractThis paper discusses progress in designing a standardized interface for speech interaction with simple machines – the Universal Speech Interface (USI) project. We discuss the motivation for such a design and issues that must be addressed by such an interface. We present our current proposals for handling these issues, and comment on the usability of these approaches based on user interactions with the system. Finally, we discuss future work and plans for the USI project. 1. Stefanie Shriver, Ronald Rosenfeld, Xiaojin Zhu 0001, Arthur R. Toth, Alexander I. Rudnicky, Markus D. Flückiger |
INTERSPEECH | 5 |
| 2001 | DARPA communicator dialog travel planning systems: the june 2000 data collectionabstractThis paper describes results of an experiment with 9 different DARPA Communicator Systems who participated in the June 2000 data collection. All systems supported travel planning and utilized some form of mixed-initiative interaction. However they varied in several critical dimensions: (1) They targeted different back-end databases for travel information; (2) The used different modules for ASR,NLU,TTS and dialog management. We describe the experimental design, the approach to data collection, the metrics collected, and results comparing the systems. 1. Marilyn A. Walker, John S. Aberdeen, Julie E. Boland, Elizabeth Owen Bratt, John S. Garofolo, Lynette Hirschman, Audrey N. Le, Sungbok Lee, Shri Narayanan, Kishore Papineni, Bryan L. Pellom, Joseph Polifroni, Alexandros Potamianos, P. Prabhu, Alexander I. Rudnicky, Gregory A. Sanders, Stephanie Seneff, David Stallard, Steve Whittaker 0001 |
INTERSPEECH | 15 |
| 2001 | Word level confidence annotation using combinations of featuresabstractThis paper describes the development of a word-level confidence metric suitable for use in a dialog system. Two aspects of the problems are investigated: the identification of useful features and the selection of an effective classifier. We find that two parse-level features, Parsing-Mode and SlotBackoff-Mode, provide annotation accuracy comparable to that observed for decoder-level features. However, both decoderlevel and parse-level features independently contribute to confidence annotation accuracy. In comparing different classification techniques, we found that Support Vector Machines (SVMs) appear to provide the best accuracy. Overall we achieve 39.7% reduction in annotation uncertainty for a binary confidence decision in a travel-planning domain. Rong Zhang 0003, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2000 | Task and domain specific modelling in the Carnegie Mellon communicator systemabstractThe Carnegie Mellon Communicator is a telephone-based dialog system that supports planning in a travel domain. The implementation of such a system requires two complimentary components, an architecture capable of managing interaction and the task, as well as a knowledge base that captures the speech, language and task characteristics specific to the domain. Given a suitable architecture, the principal effort in development in taken up in the acquisition and processing of a domain knowledge base. This paper describes a variety of techniques we have applied to modeling in acoustic, language, task, generation and synthesis components of the system. 1. INTRODUCTION System development involves a great deal of knowledge engineering, which is both time-consuming and requires a variety of experts to participate in the process. Therefore methods that seek to minimize this resource, for example through training based on domain-specific corpora are preferred. Effective use of corpora, however, ... Alexander I. Rudnicky, Christina L. Bennett, Alan W. Black, Ananlada Chotimongkol, Kevin A. Lenzo, Alice Oh, Rita Singh |
INTERSPEECH | 1 |
| 2000 | Language modeling for dialog systemabstractLanguage modeling for speech recognizer in dialog systems can take two forms. Human input can be constrained through a directed dialog, allowing the decoder to use a state-specific language model to improve recognition accuracy. Mixedinitiative systems allow for human input that while domainspecific might not be state-specific. Nevertheless, for the most part human input to a mixed-initiative system is predictable, particularly when given information about the immediately preceding system prompt. The work reported in this paper addresses the problem of balancing state-specific and general language modeling in a mixed-initiative dialog system. By incorporating dialog state adaptation of the language model, we have reduced the recognition error rate by 11.5% Wei Xu 0017, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2000 | Can artificial neural networks learn language models?abstractCurrently, N-gram models are the most common and widely used models for statistical language modeling. In this paper, we investigated an alternative way to build language models, i.e., using artificial neural networks to learn the language model. Our experiment result shows that the neural network can learn a language model that has performance even better than standard statistical methods. 1. Wei Xu 0017, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2000 | Interactive Speech Translation in the Diplomat Project
Robert E. Frederking, Alexander I. Rudnicky, Christopher Hogan, Kevin A. Lenzo |
Mach. Transl. | 2 |
| 1999 | Dialog analysis in the carnegie mellon communicatorabstractIn this paper, we present a formative evaluation procedure that we have applied to the Communicator dialog system. In the system improvement process, we have recognized the need to identify interaction failures through passive observation of system use. By systematizing the process of dialog evaluation, we hope to gain a mechanism for effectively communicating descriptions of interaction failures, specifically for use in system improvement. Additionally, we argue that this process can be taught to and executed by an evaluator external to the system development process, with the same proficiency as someone intimately familiar with the mechanics of the system components. Paul C. Constantinides, Alexander I. Rudnicky |
EUROSPEECH | 2 |
| 1999 | Data collection and processing in the carnegie mellon communicator
Maxine Eskénazi, Alexander I. Rudnicky, Karin Gregory, Paul C. Constantinides, Robert Brennan, Christina L. Bennett, Jwan Allen |
EUROSPEECH | 2 |
| 1999 | Creating natural dialogs in the carnegie mellon communicator system
Alexander I. Rudnicky, Eric H. Thayer, Paul C. Constantinides, Chris Tchou, R. Shern, Kevin A. Lenzo, Alice Oh |
EUROSPEECH | 1 |
| 1999 | A new approach to the translating telephoneabstractThe Translating Telephone has been a major goal of speech translation for many years. Previous approaches have attempted to work from limited-domain, fully-automatic translation towards broad-coverage, fully-automatic translation. We are approaching the problem from a different direction: starting with a broad-coverage but not fully-automatic system, and working towards full automation. We believe that working in this direction will provide us with better feedback, by observing users and collecting language data under realistic conditions, and thus may allow more rapid progress towards the same ultimate goal. Our initial approach relies on the wide-spread availability of Internet connections and web browsers to provide a user interface. We describe our initial work, which is an extension of the Diplomat wearable speech translator. Robert E. Frederking, Christopher Hogan, Alexander I. Rudnicky |
MTSummit | 3 |
| 1998 | A schema based approach to dialog controlabstractFrame-based approaches to spoken language interaction work well for limited tasks such as information access, given that the goal of the interaction is to construct a correct query then execute it. More complex tasks, however, can benefit from more active system participation. We describe two mechanisms that provide this, a modified stack that allows the system to track multiple topics, and form-specific schema that allow the system to deal with tasks that involve completion of multiple forms. Domain-dependent schema specify system behavior and are executed by a domain-independent engine. We describe implementations for a personal calendar system and for an air travel planning system. 1. INTRODUCTION The success of frame-based information access systems, such as for the ATIS domain (e.g., Ward & Issar, 1994) and others (e.g, Goddeau et al., 1996), leads to the question of whether such an approach could be adapted for domains that may require more sophisticated dialog management. Idea... Paul C. Constantinides, Scott Hansma, Chris Tchou, Alexander I. Rudnicky |
ICSLP | 4 |
| 1996 | Speechwear: a mobile speech system
Alexander I. Rudnicky, Stephen Reed, Eric H. Thayer |
ICSLP | 1 |
| 1995 | Speech for Multimedia Information RetrievalabstractNo abstract available. Alex Hauptmann 0001, Michael Witbrock, Alexander I. Rudnicky |
ACM Symposium on User Interface Software and Technology | 3 |
| 1993 | Factors affecting choice of speech over keyboard and mouse in a simple data-retrieval taskabstractThis paper describes some recent experiments that assess user mode selection behavior in a multi-modal environment in which actions can be performed with equivalent effect by speech, keyboard or scroller. Results indicate that users freely choose speech over other modalities, even when it is less efcient in objective terms, such as time-to-completion or input error. Additional evidence indicates that users appear to focus on simple input time in making their choice of mode, in effect minimizing the amount of personal effort expended. Alexander I. Rudnicky |
EUROSPEECH | 1 |
| 1992 | A performance model of system delay and user strategy selectionabstractThis study lays the ground work for a predictive, zero-parameter engineering model that characterizes the relationship between system delay and user performance. This study specifically investigates how system delays affects a user’s selection of task strategy. Strategy selection is hypothesized to be based on a cost function combining two factors: (1) the effort required to synchronize input with system availability and (2) the accuracy level afforded. Results indicate that users, seeking to minimize effort and maximize accuracy, choose among three strategies - automatic performance, pacing, and monitoring. These findings provide a systematic account of the influence of system delay on user performance, based on adaptive strategy choice drive by cost. Steven L. Teal, Alexander I. Rudnicky |
CHI | 2 |
| 1991 | Spoken language interfaces: the OM systemabstractNo abstract available. Jean-Michel Lunati, Alexander I. Rudnicky |
CHI | 2 |
| 1991 | Models for evaluating interaction protocols in speech recognitionabstractRecognitionerrors complicate the assessment of speech systems.This paper presents a new approach to modeling spoken language interaction protocols, based on finite Markov chains.An interaction protocol, prescribed by the interface design, defines a set of primitive transaction steps and the order of their execut ion.The efficiency of an interface depends on the interaction protocol as well as the cost of each different transaction step.Markov chains provide a simple and computationally eflicient method for modeling errorful systems.They allow for detailed comparisons between different interaction protocols and between different modalities.The method is illustrated by application to example protocols. Alexander I. Rudnicky, Alex Hauptmann 0001 |
CHI | 1 |
| 1991 | Spoken language recognition in an office management domainabstractThe authors highlight needs related to a voice interface and describe the implementation of a general-purpose spoken language interface, the Carnegie Mellon Spoken Language Shell (CM-SLS). CM-SLS provides voice interface services to different applications running on the same computer. CM-SLS was used to build the Office Manager, a collection of applications that includes an appointment calendar, a personal database, voice mail, and a calculator. The performance of several system components is described.> Alexander I. Rudnicky, Jean-Michel Lunati, Alexander M. Franz |
ICASSP | 1 |
| 1990 | Spoken language interaction in a goal-directed taskabstractTo study the spoken language interface in the context of a complex problem-solving task, a group of users are asked to perform a spreadsheet task, alternating voice and keyboard input. A total of 40 tasks are performed by each participant, the first 30 in a group (over several days), the remaining ones a month later. The voice spreadsheet program is extensively instrumented to provide detailed information about the components of the interaction. These data, as well as analysis of the participant's utterances and recognizer output, provide a fairly detailed picture of spoken language interaction. Although task completion by voice takes longer than by keyboard, analysis shows that users would be able to perform the spreadsheet task faster by voice, if two key criteria could be met: recognition occurs in real-time, and the error rate is sufficiently low. This initial experience with a spoken language system also allows the identification of several metrics, beyond those traditionally associated with speech recognition, that can be used to characterize system performance.> Alexander I. Rudnicky, Michelle Sakamoto, Joseph Polifroni |
ICASSP | 1 |
| 1990 | Spoken language interaction in a spreadsheet task
Alexander I. Rudnicky, Michelle Sakamoto, Joseph Polifroni |
INTERACT | 1 |
| 1988 | An unanchored matching algorithm for lexical accessabstractDescribes the lexical access component of the Carnegie-Mellon University (CMU) continuous speech recognition system. The word recognition algorithm operates in a left to right fashion, building words as it traverses an input network. Search is initiated at each node in the input network. The score assigned to a word is a function of both arc phone probabilities assigned by the acoustic phonetic module and knowledge of expected phone duration and frequency of occurrence of different word pronunciations. The algorithm also incorporates knowledge-based strategies to control the number of hypotheses generated by the matcher. These strategies use criteria external to the search. Performance characteristics are reported using a 1029 word lexicon built automatically from standard pronunciation base forms by context-dependent phonetic rules. Lexical rules are independent of specific lexicons and are derived by examination of transcribed speech data. The lexical representation now includes juncture rules that model specific inter-word phenomena. A junction validation module is also described, whose task is to evaluate the connectivity of words in the word hypotheses lattice.> Alexander I. Rudnicky, Zongge Li, Joseph Polifroni, Eric H. Thayer, Julia L. Gale |
ICASSP | 1 |
| 1988 | Talking to Computers: An Empirical Investigation
Alex Hauptmann 0001, Alexander I. Rudnicky |
Int. J. Man Mach. Stud. | 2 |
| 1987 | Lexical access with lattice inputabstractThis paper describes an alternative approach to lexical access in the CMU ANGEL speech recognition system. Using this approach, the asynchronous phonetic hypotheses generated by an acoustic-phonetics module are converted to a directed graph. This graph is compared to a pronunciation dictionary. Performance results for this approach and the original CMU approach are similar. An error analysis indicates promising directions for further work. Hy Murveit, Mitch Weintraub, Jared Bernstein, Alexander I. Rudnicky |
ICASSP | 5 |
| 1987 | The lexical access component of the CMU continuous speech recognition systemabstractThe CMU Lexical Access system hypothesizes words from a phonetic lattice, supplemented by a coarse labelling of the speech signal. Word hypotheses are anchored on syllabic nuclei and are generated independently for different parts of the utterance. Junctures between words are resolved separately, on demand from the Parser module. The lexical representation is generated by rule from baseforms, in a completely automatic process. A description of the various components of the system is provided, as well as performance data. Alexander I. Rudnicky, Lynn K. Baumeister, Kevin H. DeGraaf, Eric Lehmann |
ICASSP | 1 |