Alexander I. Rudnicky

dblp:29/5401 · also Alex Rudnicky, Alexander Rudnicky · DBLP profile ↗
← Back
135ranked-venue papers
10as first author
17since 2021 · last 2025
0000-0003-2044-8446ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 90 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 80 · 8 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 21 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Language Models Can be Efficiently Steered via Minimal Embedding Layer Transformations
abstract
Large Language Models (LLMs) are increasingly costly to fine-tune due to their size, with embedding layers alone accounting for up to 20% of model parameters.While Parameter-Efficient Fine-Tuning (PEFT) methods exist, they largely overlook the embedding layer.In this paper, we introduce TinyTE, a novel PEFT approach that steers model behavior via minimal translational transformations in the embedding space.TinyTE modifies input embeddings without altering hidden layers, achieving competitive performance while requiring approximately 0.0001% of the parameters needed for full fine-tuning.Experiments across architectures provide a new lens for understanding the relationship between input representations and model behavior-revealing them to be more flexible at their foundation than previously thought. 1
Diogo Tavares, David Semedo, Alexander I. Rudnicky, João Magalhães
EMNLP3
2025 Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
abstract
Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use a masked prediction objective, where the model learns to predict information about masked input segments from the unmasked context. The choice of prediction targets in this framework impacts their performance on downstream tasks. For instance, models pre-trained with targets that capture prosody learn representations suited for speaker-related tasks, while those pre-trained with targets that capture phonetics learn representations suited for content-related tasks. Moreover, prediction targets can differ in the level of detail they capture. Models pre-trained with targets that encode fine-grained acoustic features perform better on tasks like denoising, while those pre-trained with targets focused on higher-level abstractions are more effective for content-related tasks. Despite the importance of prediction targets, the design choices that affect them have not been thoroughly studied. This work explores the design choices and their impact on downstream task performance. Our results indicate that the commonly used design choices for HuBERT can be suboptimal. We propose approaches to create more informative prediction targets and demonstrate their effectiveness through improvements across various downstream tasks.
Takuya Higuchi, He Bai 0013, Ahmed Hussen Abdelaziz, Shinji Watanabe 0001, Alexander I. Rudnicky, Tatiana Likhomanenko, Barry-John Theobald, Zakaria Aldeneh
ICASSP6
2025 A Variational Framework for Improving Naturalness in Generative Spoken Language Models
abstract
The success of large language models in text processing has inspired their adaptation to speech modeling. However, since speech is continuous and complex, it is often discretized for autoregressive modeling. Speech tokens derived from self-supervised models (known as semantic tokens) typically focus on the linguistic aspects of speech but neglect prosodic information. As a result, models trained on these tokens can generate speech with reduced naturalness. Existing approaches try to fix this by adding pitch features to the semantic tokens. However, pitch alone cannot fully represent the range of paralinguistic attributes, and selecting the right features requires careful hand-engineering. To overcome this, we propose an end-to-end variational approach that automatically learns to encode these continuous speech attributes to enhance the semantic tokens. Our approach eliminates the need for manual extraction and selection of paralinguistic features. Moreover, it produces preferred speech continuations according to human raters. Code, samples and models are available at https://github.com/b04901014/vae-gslm.
Takuya Higuchi, Zakaria Aldeneh, Ahmed Hussen Abdelaziz, Alexander I. Rudnicky
ICML5
2024 Overview of the Tenth Dialog System Technology Challenge: DSTC10
abstract
This article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks.
Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky
IEEE ACM Trans. Audio Speech Lang. Process.28
2023 A Vector Quantized Approach for Text to Speech Synthesis on Real-World Spontaneous Speech
abstract
Recent Text-to-Speech (TTS) systems trained on reading or acted corpora have achieved near human-level naturalness. The diversity of human speech, however, often goes beyond the coverage of these corpora. We believe the ability to handle such diversity is crucial for AI systems to achieve human-level communication. Our work explores the use of more abundant real-world data for building speech synthesizers. We train TTS systems using real-world speech from YouTube and podcasts. We observe the mismatch between training and inference alignments in mel-spectrogram based autoregressive models, leading to unintelligible synthesis, and demonstrate that learned discrete codes within multiple code groups effectively resolves this issue. We introduce our MQTTS system whose architecture is designed for multiple code generation and monotonic alignment, along with the use of a clean silence prompt to improve synthesis quality. We conduct ablation analyses to identify the efficacy of our methods. We show that MQTTS outperforms existing TTS systems in several objective and subjective measures.
Shinji Watanabe 0001, Alexander I. Rudnicky
AAAI3
2023 Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis
abstract
Length extrapolation permits training a transformer language model on short sequences that preserves perplexities when tested on substantially longer sequences.A relative positional embedding design, ALiBi, has had the widest usage to date.We dissect ALiBi via the lens of receptive field analysis empowered by a novel cumulative normalized gradient tool.The concept of receptive field further allows us to modify the vanilla Sinusoidal positional embedding to create Sandwich, the first parameter-free relative positional embedding design that truly length information uses longer than the training sequence.Sandwich shares with KERPLE and T5 the same logarithmic decaying temporal bias pattern with learnable relative positional embeddings; these elucidate future extrapolatable positional embedding design.
Ta-Chung Chi, Ting-Han Fan, Alexander I. Rudnicky, Peter J. Ramadge
ACL (1)3
2023 Exploring Wav2vec 2.0 Fine Tuning for Improved Speech Emotion Recognition
abstract
While Wav2Vec 2.0 has been proposed for speech recognition (ASR), it can also be used for speech emotion recognition (SER); its performance can be significantly improved using different fine-tuning strategies. Two baseline methods, vanilla fine-tuning (V-FT) and task adaptive pretraining (TAPT) are first presented. We show that V-FT is able to outperform state-of-the-art models on the IEMOCAP dataset. TAPT, an existing NLP fine-tuning strategy, further improves the performance on SER. We also introduce a novel fine-tuning method termed P-TAPT, which modifies the TAPT objective to learn contextualized emotion representations. Experiments show that P-TAPT performs better than TAPT, especially under low-resource settings. Compared to prior works in this literature, our top-line system achieved a 7.4% absolute improvement in unweighted accuracy (UA) over the state-of-the-art performance on IEMOCAP. Our code is publicly available.1
Alexander I. Rudnicky
ICASSP2
2023 A Unified One-Shot Prosody and Speaker Conversion System with Self-Supervised Discrete Speech Units
abstract
We present a unified system to realize one-shot voice conversion (VC) on the pitch, rhythm, and speaker attributes. Existing works generally ignore the correlation between prosody and language content, leading to the degradation of naturalness in converted speech. Additionally, the lack of proper language features prevents these systems from accurately preserving language content after conversion. To address these issues, we devise a cascaded modular system leveraging self-supervised discrete speech units as language representation. These discrete units provide duration information essential for rhythm modeling. Our system first extracts utterance-level prosody and speaker representations from the raw waveform. Given the prosody representation, a prosody predictor estimates pitch, energy, and duration for each discrete unit in the utterance. A synthesizer further reconstructs speech based on the predicted prosody, speaker representation, and discrete units. Experiments show that our system outperforms previous approaches in naturalness, intelligibility, speaker transferability, and prosody transferability. Code and samples are publicly available.1
Shinji Watanabe 0001, Alexander I. Rudnicky
ICASSP3
2023 Learning to Ask Questions for Zero-shot Dialogue State Tracking
abstract
We present a method for performing zero-shot Dialogue State Tracking (DST) by casting the task as a learning-to-ask-questions framework. The framework learns to pair the best question generation (QG) strategy with in-domain question answering (QA) methods to extract slot values from a dialogue without any human intervention. A novel self-supervised QA pretraining step using in-domain data is essential to learn the structure without requiring any slot-filling annotations. Moreover, we show that QG methods need to be aligned with the same grammatical person used in the dialogue. Empirical evaluation on the MultiWOZ 2.1 dataset demonstrates that our approach, when used alongside robust QA models, outperforms existing zero-shot methods in the challenging task of zero-shot cross domain adaptation-given a comparable amount of domain knowledge during data creation. Finally, we analyze the impact of the types of questions used, and demonstrate that the algorithmic approach outperforms template-based question generation.
Diogo Tavares, David Semedo, Alexander I. Rudnicky, João Magalhães
SIGIR3
2022 Fine-Grained Style Control In Transformer-Based Text-To-Speech Synthesis
abstract
In this paper, we present a novel architecture to realize fine-grained style control on the transformer-based text-to-speech synthesis (TransformerTTS). Specifically, we model the speaking style by extracting a time sequence of local style tokens (LST) from the reference speech. The existing content encoder in TransformerTTS is then replaced by our designed cross-attention blocks for fusion and alignment between content and style. As the fusion is performed along with the skip connection, our cross-attention block provides a good inductive bias to gradually infuse the phoneme representation with a given style. Additionally, we prevent the style embedding from encoding linguistic content by randomly truncating LST during training and using wav2vec 2.0 features. Experiments show that with fine-grained style control, our system performs better in terms of naturalness, intelligibility, and style transferability. Our code and samples are publicly available.1
Alexander I. Rudnicky
ICASSP2
2022 Training Discrete Deep Generative Models via Gapped Straight-Through Estimator
abstract
While deep generative models have succeeded in image processing, natural language processing, and reinforcement learning, training that involves discrete random variables remains challenging due to the high variance of its gradient estimation process. Monte Carlo is a common solution used in most variance reduction approaches. However, this involves time-consuming resampling and multiple function evaluations. We propose a Gapped Straight-Through (GST) estimator to reduce the variance without incurring resampling overhead. This estimator is inspired by the essential properties of Straight-Through Gumbel-Softmax. We determine these properties and show via an ablation study that they are essential. Experiments demonstrate that the proposed GST estimator enjoys better performance compared to strong baselines on two discrete deep generative modeling tasks, MNIST-VAE and ListOps.
Ting-Han Fan, Ta-Chung Chi, Alexander I. Rudnicky, Peter J. Ramadge
ICML3
2022 KERPLE: Kernelized Relative Positional Embedding for Length Extrapolation
abstract
Relative positional embeddings (RPE) have received considerable attention since RPEs effectively model the relative distance among tokens and enable length extrapolation. We propose KERPLE, a framework that generalizes relative position embedding for extrapolation by kernelizing positional differences. We achieve this goal using conditionally positive definite (CPD) kernels, a class of functions known for generalizing distance metrics. To maintain the inner product interpretation of self-attention, we show that a CPD kernel can be transformed into a PD kernel by adding a constant offset. This offset is implicitly absorbed in the Softmax normalization during self-attention. The diversity of CPD kernels allows us to derive various RPEs that enable length extrapolation in a principled way. Experiments demonstrate that the logarithmic variant achieves excellent extrapolation performance on three large language modeling datasets. Our implementation and pretrained checkpoints are released at~\url{https://github.com/chijames/KERPLE.git}.
Ta-Chung Chi, Ting-Han Fan, Peter J. Ramadge, Alexander I. Rudnicky
NeurIPS4
2022 Structured Dialogue Discourse Parsing
abstract
Dialogue discourse parsing aims to uncover the internal structure of a multi-participant conversation by finding all the discourse links and corresponding relations.Previous work either treats this task as a series of independent multiple-choice problems, in which the link existence and relations are decoded separately, or the encoding is restricted to only local interaction, ignoring the holistic structural information.In contrast, we propose a principled method that improves upon previous work from two perspectives: encoding and decoding.From the encoding side, we perform structured encoding on the adjacency matrix followed by the matrix-tree learning algorithm, where all discourse links and relations in the dialogue are jointly optimized based on latent tree-level distribution.From the decoding side, we perform structured inference using the modified Chiu-Liu-Edmonds algorithm, which explicitly generates the labeled multi-root non-projective spanning tree that best captures the discourse structure.In addition, unlike in previous work, we do not rely on hand-crafted features; this improves the model's robustness.Experiments show that our method achieves new state-of-the-art, surpassing the previous model by 2.3 on STAC and 1.5 on Molweni (F1 scores). 1
Ta-Chung Chi, Alexander I. Rudnicky
SIGDIAL2
2022 Spoken language interaction with robots: Recommendations for future research
abstract
With robotics rapidly advancing, more effective human–robot interaction is increasingly needed to realize the full potential of robots for society. While spoken language must be part of the solution, our ability to provide spoken language interaction capabilities is still very limited. In this article, based on the report of an interdisciplinary workshop convened by the National Science Foundation, we identify key scientific and engineering advances needed to enable effective spoken language interaction with robotics. We make 25 recommendations, involving eight general themes: putting human needs first, better modeling the social and interactive aspects of language, improving robustness, creating new methods for rapid adaptation, better integrating speech and language with other communication modalities, giving speech and language components access to rich representations of the robot’s current knowledge and state, making all components operate in real time, and improving research infrastructure and resources. Research and development that prioritizes these topics will, we believe, provide a solid foundation for the creation of speech-capable robots that are easy and effective for humans to work with.
Matthew Marge, Carol Y. Espy-Wilson, Nigel G. Ward, Abeer Alwan, Yoav Artzi, Mohit Bansal, Gilmer L. Blankenship, Joyce Y. Chai, Hal Daumé III, Debadeepta Dey, Mary P. Harper, Thomas Howard, Casey Kennington, Ivana Kruijff-Korbayová, Dinesh Manocha, Cynthia Matuszek, Ross Mead, Raymond J. Mooney, Roger K. Moore, Mari Ostendorf, Heather Pon-Barry, Alexander I. Rudnicky, Matthias Scheutz, Robert St. Amant, Stefanie Tellex, David R. Traum, Zhou Yu 0005
Comput. Speech Lang.22
2021 Zero-Shot Dialogue Disentanglement by Self-Supervised Entangled Response Selection
abstract
Dialogue disentanglement aims to group utterances in a long and multi-participant dialogue into threads.This is useful for discourse analysis and downstream applications such as dialogue response selection, where it can be the first step to construct a clean context/response set.Unfortunately, labeling all reply-to links takes quadratic effort w.r.t the number of utterances: an annotator must check all preceding utterances to identify the one to which the current utterance is a reply.In this paper, we are the first to propose a zero-shot dialogue disentanglement solution.Firstly, we train a model on a multi-participant response selection dataset harvested from the web which is not annotated; we then apply the trained model to perform zero-shot dialogue disentanglement.Without any labeled data, our model can achieve a cluster F1 score of 25.We also fine-tune the model using various amounts of labeled data.Experiments show that with only 10% of the data, we achieve nearly the same performance of using the full dataset 1 .
Ta-Chung Chi, Alexander I. Rudnicky
EMNLP (1)2
2021 Speech Representation Learning Combining Conformer CPC with Deep Cluster for the ZeroSpeech Challenge 2021
abstract
We present a system for the Zero Resource Speech Challenge 2021, which combines a Contrastive Predictive Coding (CPC) with deep cluster. In deep cluster, we first prepare pseudo-labels obtained by clustering the outputs of a CPC network with k-means. Then, we train an additional autoregressive model to classify the previously obtained pseudo-labels in a supervised manner. Phoneme discriminative representation is achieved by executing the second-round clustering with the outputs of the final layer of the autoregressive model. We show that replacing a Transformer layer with a Conformer layer leads to a further gain in a lexical metric. Experimental results show that a relative improvement of 35% in a phonetic metric, 1.5% in the lexical metric, and 2.3% in a syntactic metric are achieved compared to a baseline method of CPC-small which is trained on LibriSpeech 460h data. We achieve top results in this challenge with the syntactic metric.
Takashi Maekaku, Xuankai Chang, Yuya Fujita, Shinji Watanabe 0001, Alexander I. Rudnicky
Interspeech6
2021 Temporal Context in Speech Emotion Recognition
Yangyang Xia, Alexander I. Rudnicky, Richard M. Stern
Interspeech3
2020 Adjusting Image Attributes of Localized Regions with Low-level Dialogue
abstract
Natural Language Image Editing (NLIE) aims to use natural language instructions to edit images. Since novices are inexperienced with image editing techniques, their instructions are often ambiguous and contain high-level abstractions which require complex editing steps. Motivated by this inexperience aspect, we aim to smooth the learning curve by teaching the novices to edit images using low-level command terminologies. Towards this end, we develop a task-oriented dialogue system to investigate low-level instructions for NLIE. Our system grounds language on the level of edit operations, and suggests options for users to choose from. Though compelled to express in low-level terms, user evaluation shows that 25% of users found our system easy-to-use, resonating with our motivation. Analysis shows that users generally adapt to utilizing the proposed low-level language interface. We also identified object segmentation as the key factor to user satisfaction. Our work demonstrates advantages of low-level, direct language-action mapping approach that can be applied to other problem domains beyond image editing such as audio editing or industrial design.
Tzu-Hsiang Lin, Alexander I. Rudnicky, Trung Bui, Doo Soon Kim, Jean Oh
LREC2
2019 Miscommunication Detection and Recovery in Situated Human-Robot Dialogue
abstract
Even without speech recognition errors, robots may face difficulties interpreting natural-language instructions. We present a method for robustly handling miscommunication between people and robots in task-oriented spoken dialogue. This capability is implemented in TeamTalk, a conversational interface to robots that supports detection and recovery from the situated grounding problems of referential ambiguity and impossible actions. We introduce a representation that detects these problems and a nearest-neighbor learning algorithm that selects recovery strategies for a virtual robot. When the robot encounters a grounding problem, it looks back on its interaction history to consider how it resolved similar situations. The learning method is trained initially on crowdsourced data but is then supplemented by interactions from a longitudinal user study in which six participants performed navigation tasks with the robot. We compare results collected using a general model to user-specific models and find that user-specific models perform best on measures of dialogue efficiency, while the general model yields the highest agreement with human judges. Our overall contribution is a novel approach to detecting and recovering from miscommunication in dialogue by including situated context, namely, information from a robot’s path planner and surroundings.
Matthew Marge, Alexander I. Rudnicky
ACM Trans. Interact. Intell. Syst.2
2018 SOGO: A Social Intelligent Negotiation Dialogue System
abstract
In this paper, we propose a semi-automatic social intelligent negotiation dialogue system that interweaves task utterance with conversational strategies to engage human users in negotiation. Our two-phase system operates sequentially in a reasoning-and-generation loop: In the task phase, we leverage an off-the-shelf end-to-end dialogue model for negotiation to build a dialogue manager which decides the next system's task intention. Then, during the social phase, we employ a theory-driven, template-based natural language generator to realize the task intention as a genre of social conversational strategy. Subsequently, a set of conversational strategies are presented to a human expert who decides the final sentence to be uttered by the dialogue system. Compared to the baseline system, our proposed social intelligent dialogue system achieves a higher agreement rate and more "good deals" with humans while building interpersonal rapport.
Oscar J. Romero, Alexander I. Rudnicky
IVA3
2018 Word Segmentation From Phoneme Sequences Based On Pitman-Yor Semi-Markov Model Exploiting Subword Information
abstract
Word segmentation from phoneme sequences is essential to identify unknown words -of-vocabulary; OOV) in spoken dialogues. The Pitman-Yor semi-Markov model (PYSMM) is used for word segmentation that handles dynamic increase in vocabularies. The obtained vocabularies, however, still include meaningless entries due to insufficient cues for phoneme sequences. We focus here on using subword information to capture patterns as “words.” We propose 1) a model based on subword N-gram and subword estimation using a vocabulary set, and 2) posterior fusion of the results of a PYSMM and our model to take advantage of both. Our experiments showed 1) the potential of using subword information for OOV acquisition, and 2) that our method outperformed the PYSMM by 1.53 and 1.07 in terms of the F-measure of the obtained OOV set for English and Japanese corpora, respectively.
Ryu Takeda, Kazunori Komatani, Alexander I. Rudnicky
SLT3
2017 Learning Conversational Systems that Interleave Task and Non-Task Content
abstract
Task-oriented dialog systems have been applied in various tasks, such as automated personal assistants, customer service providers and tutors. These systems work well when users have clear and explicit intentions that are well-aligned to the systems' capabilities. However, they fail if users intentions are not explicit.To address this shortcoming, we propose a framework to interleave non-task content (i.e.everyday social conversation) into task conversations. When the task content fails, the system can still keep the user engaged with the non-task content. We trained a policy using reinforcement learning algorithms to promote long-turn conversation coherence and consistency, so that the system can have smooth transitions between task and non-task content.To test the effectiveness of the proposed framework, we developed a movie promotion dialog system. Experiments with human users indicate that a system that interleaves social and task content achieves a better task success rate and is also rated as more engaging compared to a pure task-oriented system.
Zhou Yu 0005, Alexander I. Rudnicky, Alan W. Black
IJCAI2
2016 Unsupervised user intent modeling by feature-enriched matrix factorization
abstract
Spoken language interfaces are being incorporated into various devices such as smart phones and TVs. However, dialogue systems may fail to respond correctly when users' request functionality is not supported by currently installed apps. This paper proposes a feature-enriched matrix factorization (MF) approach to model open domain intents, which allows a system to dynamically add unexplored domains according to users' requests. First we leverage the structured knowledge from Wikipedia and Freebase to automatically acquire domain-related semantics to enrich features of input utterances, and then MF is applied to model automatically acquired knowledge, published app textual descriptions and users' spoken requests in a joint fashion; this generates latent feature vectors for utterances and user intents without need of prior annotations. Experiments show that the proposed MF models incorporated with rich features significantly improve intent prediction, achieving about 34% of mean average precision (MAP) for both ASR and manual transcripts.
Yun-Nung Chen, Ming Sun 0001, Alexander I. Rudnicky, Anatole Gershman
ICASSP3
2016 An Intelligent Assistant for High-Level Task Understanding
abstract
People are able to interact with domain-specific intelligent assistants (IAs) and get help with tasks. But sometimes user goals are complex and may require interactions with multiple applications. However current IAs are limited to specific applications and users have to directly manage execution spanning multiple applications in order to engage in more complex activities. An ideal personal agent would be able to learn, over time, about tasks that span different resources. This paper addresses the problem of cross-domain task assistance in the context of spoken dialogue systems. We propose approaches to discover users' high-level intentions and using this information to assist users in their task. We collected real-life smartphone usage data from 14 participants and investigated how to extract high-level intents from users' descriptions of their activities. Our experiments show that understanding high-level tasks allows the agent to actively suggest apps relevant to pursuing particular user goals and reduce the cost of users' self-management.
Ming Sun 0001, Yun-Nung Chen, Alexander I. Rudnicky
IUI3
2016 User Engagement Study with Virtual Agents Under Different Cultural Contexts
Zhou Yu 0005, Xinrui He, Alan W. Black, Alexander I. Rudnicky
IVA4
2016 AppDialogue: Multi-App Dialogues for Intelligent Assistants
Ming Sun 0001, Yun-Nung Chen, Zhenhao Hua, Yulian Tamres-Rudnicky, Arnab Dash, Alexander I. Rudnicky
LREC6
2016 A Wizard-of-Oz Study on A Non-Task-Oriented Dialog Systems That Reacts to User Engagement
abstract
In this paper, we describe a system that reacts to both possible system breakdowns and low user engagement with a set of conversational strategies.These general strategies reduce the number of inappropriate responses and produce better user engagement.We also found that a system that reacts to both possible system breakdowns and low user engagement is rated by both experts and non-experts as having better overall user engagement compared to a system that only reacts to possible system breakdowns.We argue that for non-task-oriented systems we should optimize on both system response appropriateness and user engagement.We also found that apart from making the system response appropriate, funny and provocative responses can also lead to better user engagement.On the other hand, short appropriate responses, such as "Yes" or "No" can lead to decreased user engagement.We will use these findings to further improve our system.
Zhou Yu 0005, Leah Nicolich-Henkin, Alan W. Black, Alexander I. Rudnicky
SIGDIAL Conference4
2016 Strategy and Policy Learning for Non-Task-Oriented Conversational Systems
abstract
We propose a set of generic conversational strategies to handle possible system breakdowns in non-task-oriented dialog systems.We also design policies to select these strategies according to dialog context.We combine expert knowledge and the statistical findings derived from data in designing these policies.The policy learned via reinforcement learning outperforms the random selection policy and the locally greedy policy in both simulated and real-world settings.In addition, we propose three metrics for conversation quality evaluation which consider both the local and global quality of the conversation.
Zhou Yu 0005, Ziyu Xu 0001, Alan W. Black, Alexander I. Rudnicky
SIGDIAL Conference4
2016 Weakly supervised user intent detection for multi-domain dialogues
abstract
Users interact with mobile apps with certain intents such as finding a restaurant. Some intents and their corresponding activities are complex and may involve multiple apps; for example, a restaurant app, a messenger app and a calendar app may be needed to plan a dinner with friends. However, activities may be quite personal and third-party developers would not be building apps to specifically handle complex intents (e.g., a DinnerPlanner). Instead we want our intelligent agent to actively learn to understand these intents and provide assistance when needed. This paper proposes a framework to enable the agent to learn an inventory of intents from a small set of task-oriented user utterances. The experiments show that on previously unseen user activities, the agent is able to reliably recognize user intents using graph-based semi-supervised learning methods. The dataset, models, and the system outputs are available to research community.
Ming Sun 0001, Aasish Pappu, Yun-Nung Chen, Alexander I. Rudnicky
SLT4
2015 Matrix Factorization with Knowledge Graph Propagation for Unsupervised Spoken Language Understanding
abstract
Yun-Nung Chen, William Yang Wang, Anatole Gershman, Alexander Rudnicky. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Yun-Nung Chen, William Yang Wang, Anatole Gershman, Alexander I. Rudnicky
ACL (1)4
2015 Leveraging Behavioral Patterns of Mobile Applications for Personalized Spoken Language Understanding
abstract
Spoken language interfaces are appearing in various smart devices (e.g. smart-phones, smart-TV, in-car navigating systems) and serve as intelligent assistants (IAs). However, most of them do not consider individual users' behavioral profiles and contexts when modeling user intents. Such behavioral patterns are user-specific and provide useful cues to improve spoken language understanding (SLU). This paper focuses on leveraging the app behavior history to improve spoken dialog systems performance. We developed a matrix factorization approach that models speech and app usage patterns to predict user intents (e.g. launching a specific app). We collected multi-turn interactions in a WoZ scenario; users were asked to reproduce the multi-app tasks that they had performed earlier on their smart-phones. By modeling latent semantics behind lexical and behavioral patterns, the proposed multi-model system achieves about 52% of turn accuracy for intent prediction on ASR transcripts.
Yun-Nung Chen, Ming Sun 0001, Alexander I. Rudnicky, Anatole Gershman
ICMI3
2015 Learning semantic hierarchy with distributed representations for unsupervised spoken language understanding
abstract
We study the problem of unsupervised ontology learning for semantic understanding in spoken dialogue systems, in particular, learning the hierarchical semantic structure from the data. Given unlabelled conversations, we augment a frame-semantic based unsupervised slot induction approach with hierarchical agglomerative clustering to merge topically-related slots (e.g., both slots “direction” and “locale” convey location-related information) for building a coherent semantic hierarchy, and then estimate the slot importance at different levels. The high-level semantic estimation involves not only within-slot but also crossslot relations. The experiments show that high-level semantic information can accurately estimate the prominence of slots, significantly improving the slot induction performance; furthermore, a semantic decoder trained on the data with automatically extracted slots achieves about 68% F-measure, which is close to the one from hand-crafted grammars.
Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky
INTERSPEECH3
2015 Distributed representation-based spoken word sense induction
abstract
Spoken Term Detection (STD) or Keyword Search (KWS) techniques can locate keyword instances but do not differentiate between meanings. Spoken Word Sense Induction (SWSI) differentiates target instances by clustering according to context, providing a more useful result. In this paper we present a fully unsupervised SWSI approach based on distributed representations of spoken utterances. We compare this approach to several others, including the state-of-the-art Hierarchical Dirichlet Process (HDP). To determine how ASR performance affects SWSI, we used three different levels of Word Error Rate (WER), 40%, 20 % and 0%; 40 % WER is representative of online video, 0 % of text. We show that the distributed representation approach outperforms all other approaches, regardless of the WER. Although LDA-based approaches do well on clean data, they degrade significantly with WER. Paradoxically, lower WER does not guarantee better SWSI performance, due to the influence of common locutions.
Justin T. Chiu, Yajie Miao, Alan W. Black, Alexander I. Rudnicky
INTERSPEECH4
2015 Learning OOV through semantic relatedness in spoken dialog systems
abstract
• Speech recognition and language understanding performance can be improved through an OOV expectand-learn procedure. • A limited domain vocabulary can be utilized to effectively acquire OOVs by the word relatedness theory through web knowledge bases. • With data-driven semantic relatedness, both the global and local learning procedures are able to successfully harvest more than 50% of OOVs, leading to better recognition and understanding performance. • This work demonstrates that o OOV learning may benefit dialog system o the proposed expect-and-learn strategy outperforms the traditional detect-and-learn in both higher effectiveness and no human involvement. 1. Linguistically semantic relatedness o Defined by linguistics, e.g., WordNet (WN), Paraphrase Database (PPDB) (Ganitkevitch et al., 2013) 2. Data-driven semantic relatedness o Distributional semantics, e.g., continuous bag-ofword embeddings (CBOW) (Mikolov et al., 2013)  Detect-and-Learn (Qin et al., 2011; 2012): o Discover OOV words during the conversation o Example: S: “I heard something like SELF, can you repeat it?” U: “It’s SELFIE.” o Drawbacks • Limited number of new words • Required human efforts to correct spellings and pronunciations  Expect-and-Learn (proposed): o Use semantic relatedness to automatically enrich the vocabulary and language model beforehand
Ming Sun 0001, Yun-Nung Chen, Alexander I. Rudnicky
INTERSPEECH3
2015 Jointly Modeling Inter-Slot Relations by Random Walk on Knowledge Graphs for Unsupervised Spoken Language Understanding
abstract
Yun-Nung Chen, William Yang Wang, Alexander Rudnicky. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky
HLT-NAACL3
2015 Miscommunication Recovery in Physically Situated Dialogue
abstract
We describe an empirical study that crowdsourced human-authored recovery strategies for various problems encountered in physically situated dialogue.The purpose was to investigate the strategies that people use in response to requests that are referentially ambiguous or impossible to execute.Results suggest a general preference for including specific kinds of visual information when disambiguating referents, and for volunteering alternative plans when the original instruction was not possible to carry out.
Matthew Marge, Alexander I. Rudnicky
SIGDIAL Conference2
2014 Two-Stage Stochastic Email Synthesizer
abstract
This paper presents the design and im-plementation details of an email synthe-sizer using two-stage stochastic natural language generation, where the first stage structures the emails according to sender style and topic structure, and the second stage synthesizes text content based on the particulars of an email structure element and the goals of a given communication for surface realization. The synthesized emails reflect sender style and the intent of communication, which can be further used as synthetic evidence for developing other applications. 1
Yun-Nung Chen, Alexander I. Rudnicky
INLG2
2014 Two-Stage Stochastic Natural Language Generation for Email Synthesis by Modeling Sender Style and Topic Structure
abstract
This paper describes a two-stage pro-cess for stochastic generation of email, in which the first stage structures the emails according to sender style and topic struc-ture (high-level generation), and the sec-ond stage synthesizes text content based on the particulars of an email element and the goals of a given communication (surface-level realization). Synthesized emails were rated in a preliminary experi-ment. The results indicate that sender style can be detected. In addition we found that stochastic generation performs better if applied at the word level than at an original-sentence level (“template-based”) in terms of email coherence, sentence flu-ency, naturalness, and preference. 1
Yun-Nung Chen, Alexander I. Rudnicky
INLG2
2014 Combination of FST and CN search in spoken term detection
abstract
Spoken Term Detection (STD) focuses on finding instances of a particular spoken word or phrase in an audio corpus. Most STD systems have a two-step pipeline, ASR followed by search. Two approaches to search are common, Confusion Network (CN) based search and Finite State Transducer (FST) based search. In this paper, we examine combination of these two different search approaches, using the same ASR output. We find that the CN search performs better on shorter queries, and FST search performs better on longer queries. By combining the different search results from the same ASR decoding, we achieve better performance compared to either search approach on its own. We also find that this improvement is additive to the usual combination of decoder results using different modeling techniques.
Justin T. Chiu, Jan Trmal, Daniel Povey, Guoguo Chen, Alexander I. Rudnicky
INTERSPEECH6
2014 Learning situated knowledge bases through dialog
abstract
To respond to a user's query, dialog agents can use a knowledge base that is either domain specific, commonsense (e.g., NELL, Freebase) or a combination of both. The drawback is that domain-specific knowledge bases will likely be limited and static; commonsense ones are dynamic but contain general information found on the web and will be sparse with respect to a domain. We address this issue through a system that solicits situational information from its users in a domain that provides information on events (seminar talks) to augment its knowledge base (covering an academic field). We find that this knowledge is consistent and useful and that it provides reliable information to users. We show that, in comparison to a base system, users find that retrievals are more relevant when the system uses its informally acquired knowledge to augment their queries.
Aasish Pappu, Alexander I. Rudnicky
INTERSPEECH2
2014 Building a vocabulary self-learning speech recognition system
abstract
This paper presents initial studies on building a vocabulary self-learning speech recognition system that can automatically learn unknown words and expand its recognition vocabulary. Our recognizer can detect and recover out-of-vocabulary (OOV) words in speech, then incorporate OOV words into its lexicon and language model (LM). As a result, these unknown words can be correctly recognized when encountered by the recognizer in future. Specifically, we apply the word-fragment hybrid system framework to detect the presence of OOV words. We propose a better phoneme-to-grapheme (P2G) model so as to correctly recover the written form for more OOV words. Furthermore, we estimate LM scores for OOV words using their syntactic and semantic properties. The experimental results show that more than 40% OOV words are successfully learned from the development data, and about 60% learned OOV words are recognized in the testing data.
Alexander I. Rudnicky
INTERSPEECH2
2014 Knowledge Acquisition Strategies for Goal-Oriented Dialog Systems
abstract
Many goal-oriented dialog agents are expected to identify slot-value pairs in a spoken query, then perform lookup in a knowledge base to complete the task. When the agent encounters unknown slotvalues, it may ask the user to repeat or reformulate the query. But a robust agent can proactively seek new knowledge from a user, to help reduce subsequent task failures. In this paper, we propose knowledge acquisition strategies for a dialog agent and show their effectiveness. The acquired knowledge can be shown to subsequently contribute to task completion.
Aasish Pappu, Alexander I. Rudnicky
SIGDIAL Conference2
2014 Dynamically supporting unexplored domains in conversational interactions by enriching semantics with neural word embeddings
abstract
Spoken language interfaces are being incorporated into various devices (e.g. smart-phones, smart TVs, etc). However, current technology typically limits conversational interactions to a few narrow predefined domains/topics. For example, dialogue systems for smartphone operation fail to respond when users ask for functions not supported by currently installed applications. We propose to dynamically add application-based domains according to users' requests by using descriptions of applications as a retrieval cue to find relevant applications. The approach uses structured knowledge resources (e.g. Freebase, Wikipedia, FrameNet) to induce types of slots for generating semantic seeds, and enriches the semantics of spoken queries with neural word embeddings, where semantically related concepts can be additionally included for acquiring knowledge that does not exist in the predefined domains. The system can then retrieve relevant applications or dynamically suggest users install applications that support unexplored domains. We find that vendor descriptions provide a reliable source of information for this purpose.
Yun-Nung Chen, Alexander I. Rudnicky
SLT2
2014 Leveraging frame semantics and distributional semantics for unsupervised semantic slot induction in spoken dialogue systems
abstract
Distributional semantics and frame semantics are two representative views on language understanding in the statistical world and the linguistic world, respectively. In this paper, we combine the best of two worlds to automatically induce the semantic slots for spoken dialogue systems. Given a collection of unlabeled audio files, we exploit continuous-valued word embeddings to augment a probabilistic frame-semantic parser that identifies key semantic slots in an unsupervised fashion. In experiments, our results on a real-world spoken dialogue dataset show that the distributional word representations significantly improve the adaptation of FrameNet-style parses of ASR decodings to the target semantic space; that comparing to a state-of-the-art baseline, a 13% relative average precision improvement is achieved by leveraging word vectors trained on two 100-billion words datasets; and that the proposed technology can be used to reduce the costs for designing task-oriented spoken dialogue systems.
Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky
SLT3
2013 Unsupervised induction and filling of semantic slots for spoken dialogue systems using frame-semantic parsing
abstract
Spoken dialogue systems typically use predefined semantic slots to parse users' natural language inputs into unified semantic representations. To define the slots, domain experts and professional annotators are often involved, and the cost can be expensive. In this paper, we ask the following question: given a collection of unlabeled raw audios, can we use the frame semantics theory to automatically induce and fill the semantic slots in an unsupervised fashion? To do this, we propose the use of a state-of-the-art frame-semantic parser, and a spectral clustering based slot ranking model that adapts the generic output of the parser to the target semantic space. Empirical experiments on a real-world spoken dialogue dataset show that the automatically induced semantic slots are in line with the reference slots created by domain experts: we observe a mean averaged precision of 69.36% using ASR-transcribed data. Our slot filling evaluations also indicate the promising future of this proposed approach.
Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky
ASRU3
2013 Using web text to improve keyword spotting in speech
abstract
For low resource languages, collecting sufficient training data to build acoustic and language models is time consuming and often expensive. But large amounts of text data, such as online newspapers, web forums or online encyclopedias, usually exist for languages that have a large population of native speakers. This text data can be easily collected from the web and then used to both expand the recognizer's vocabulary and improve the language model. One challenge, however, is normalizing and filtering the web data for a specific task. In this paper, we investigate the use of online text resources to improve the performance of speech recognition specifically for the task of keyword spotting. For the five languages provided in the base period of the IARPA BABEL project, we automatically collected text data from the web using only Limited LP resources. We then compared two methods for filtering the web data, one based on perplexity ranking and the other based on out-of-vocabulary (OOV) word detection. By integrating the web text into our systems, we observed significant improvements in keyword spotting accuracy for four out of the five languages. The best approach obtained an improvement in actual term weighted value (ATWV) of 0.0424 compared to a baseline system trained only on LimitedLP resources. On average, ATWV was improved by 0.0243 across five languages.
Ankur Gandhe, Florian Metze, Alexander I. Rudnicky, Ian Lane, Matthias Eck 0001
ASRU4
2013 Learning better lexical properties for recurrent OOV words
abstract
Out-of-vocabulary (OOV) words can appear more than once in a conversation or over a period of time. Such multiple instances of the same OOV word provide valuable information for learning the lexical properties of the word. Therefore, we investigated how to estimate better pronunciation, spelling and part-of-speech (POS) label for recurrent OOV words. We first identified recurrent OOV words from the output of a hybrid decoder by applying a bottom-up clustering approach. Then, multiple instances of the same OOV word were used simultaneously to learn properties of the OOV word. The experimental results showed that the bottom-up clustering approach is very effective at detecting the recurrence of OOV words. Furthermore, by using evidence from multiple instances of the same word, the pronunciation accuracy, recovery rate and POS label accuracy of recurrent OOV words can be substantially improved.
Alexander I. Rudnicky
ASRU2
2013 An empirical investigation of sparse log-linear models for improved dialogue act classification
abstract
Previous work on dialogue act classification have primarily focused on dense generative and discriminative models. However, since the automatic speech recognition (ASR) outputs are often noisy, dense models might generate biased estimates and overfit to the training data. In this paper, we study sparse modeling approaches to improve dialogue act classification, since the sparse models maintain a compact feature space, which is robust to noise. To test this, we investigate various element-wise frequentist shrinkage models such as lasso, ridge, and elastic net, as well as structured sparsity models and a hierarchical sparsity model that embed the dependency structure and interaction among local features. In our experiments on a real-world dataset, when augmenting N-best word and phone level ASR hypotheses with confusion network features, our best sparse log-linear model obtains a relative improvement of 19.7% over a rule-based baseline, a 3.7% significant improvement over a traditional non-sparse log-linear model, and outperforms a state-of-the-art SVM model by 2.2%.
Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky
ICASSP3
2013 Using conversational word bursts in spoken term detection
abstract
We describe a language independent word burst feature based on the structure of conversational speech that can be used to improve spoken term detection (STD) performance. Word burst refers to a phenomenon in conversational speech in which particular content words tend to occur in close proximity of each other as a byproduct of the topic under discussion. To take advantage of bursts, we describe a rescoring procedure that can be applied to lattice and confusion network outputs to improve STD performance. This approach is particularly effective when acoustic models are built with limited training data (and ASR performance is relatively poor). We find that word bursts appear in the four languages we examined and that STD performance can be improved for three of them; the remaining language is agglutinative.
Justin T. Chiu, Alexander I. Rudnicky
INTERSPEECH2
2013 Finding recurrent out-of-vocabulary words
abstract
Out-of-vocabulary (OOV) words can appear more than once in a conversation or over a period of time. Such multiple instances of the same OOV word provide valuable information for estimating the pronunciation or the part-of-speech (POS) tag of the word. But in a conventional OOV word detection system, each OOV word is recognized and treated individually. We therefore investigated how to identify recurrent OOV words in speech recognition. Specifically, we propose to cluster multiple instances of the same OOV word using a bottom-up approach. Phonetic, acoustic and contextual features were collected to measure the distance between OOV candidates. The experimental results show that the bottom-up clustering approach is very effective at detecting the recurrence of OOV words. We also found that the phonetic feature is better than the acoustic and contextual features, and the best performance is achieved when combining all features.
Alexander I. Rudnicky
INTERSPEECH2
2013 Emotion Recognition Modulating the Behavior of Intelligent Systems
abstract
The paper presents an audio-based emotion recognition system that is able to classify emotions as anger, fear, happy, neutral, sadness or disgust in real time. We use the virtual coach as an application example of how emotion recognition can be used to modulate intelligent systems' behavior. A novel minimum-error feature removal mechanism to reduce bandwidth and increase accuracy of our emotion recognition system has been introduced. A two-stage hierarchical classification approach along with a One-Against-All (OAA) framework are used. We obtained an average accuracy of 82.07% using the OAA approach, and 87.70% with a two-stage hierarchical approach, by pruning the feature set and using Support Vector Machines (SVMs) for classification.
Asim Smailagic, Daniel P. Siewiorek, Alexander I. Rudnicky, Sandeep Nallan Chakravarthula, Anshuman Kar, Nivedita Jagdale, Saksham Gautam, Rohit Vijayaraghavan, Shaurya Jagtap
ISM3
2013 SiMPE: 8th workshop on speech and sound in mobile and pervasive environments
abstract
The SiMPE workshop series started in 2006 with the goal of enabling speech processing on mobile and embedded devices. The SiMPE 2012 workshop extended the notion of audio to non-speech "Sounds" and thus the expansion became "Speech and Sound". SiMPE 2010 and 2011 brought together researchers from the speech and the HCI communities. Speech User interaction in cars was a focus area in 2009. Multimodality got more attention in SiMPE 2008. In SiMPE 2007, the focus was on developing regions.
Amit Anil Nanavati, Nitendra Rajput, Cumhur Erkut, Antti Jylhä, Alexander I. Rudnicky, Stefania Serafin, Markku Turunen
Mobile HCI6
2013 Towards evaluating recovery strategies for situated grounding problems in human-robot dialogue
abstract
Robots can use information from their surroundings to improve spoken language communication with people. Even when speech recognition is correct, robots face challenges when interpreting human instructions. These situated grounding problems include referential ambiguities and impossible-to-execute instructions. We present an approach to resolving situated grounding problems through spoken dialogue recovery strategies that robots can invoke to repair these problems. We describe a method for evaluating these strategies in human-robot navigation scenarios.
Matthew Marge, Alexander I. Rudnicky
RO-MAN2
2013 Predicting Tasks in Goal-Oriented Spoken Dialog Systems using Semantic Knowledge Bases
Aasish Pappu, Alexander I. Rudnicky
SIGDIAL Conference2
2012 System combination for out-of-vocabulary word detection
abstract
This paper presents a method to improve the out-of-vocabulary (OOV) word detection performance by combining multiple speech recognition systems' outputs. Three different fragment-word hybrid systems, the phone, subword, and graphone systems, were built for detecting OOV words. Then outputs from each individual system were combined using ROVER. Two combination metrics were explored in ROVER, voting by word frequency and voting by both word frequency and word confidence score. The experimental results show that the OOV word detection performance of the ROVER system with confidence scores is better than the ROVER system with only word frequency, as well as any of the individual hybrid systems.
Ming Sun 0001, Alexander I. Rudnicky
ICASSP3
2012 NeuroDialog: an EEG-enabled spoken dialog interface
abstract
Understanding user intent is a difficult problem in Dialog Systems, as they often need to make decisions under uncertainty. Using an inexpensive, consumer grade EEG sensor and a Wizard-of-Oz dialog system, we show that it is possible to detect system misunderstanding even before the user reacts vocally. We also present the design and implementation details of NeuroDialog, a proof-of-concept dialog system that uses an EEG based predictive model to detect system misrecognitions during live interaction.
Seshadri Sridharan, Yun-Nung Chen, Kai-min Chang, Alexander I. Rudnicky
ICMI4
2012 OOV Word Detection using Hybrid Models with Mixed Types of Fragments
abstract
This paper presents initial studies to improve the out-of-vocabulary (OOV) word detection performance by using mixed types of fragment units in one hybrid system. Three types of fragment units, subwords, syllables, and graphones, were combined in two different ways to build the hybrid lexicon and language model. The experimental results show that hybrid systems with mixed types of fragment units perform better than hybrid systems using only one type of fragment unit. After comparing the OOV word detection performance with the number and length of fragment units of each system, we proposed future work to better utilize mixed types of fragment units in a hybrid system.
Alexander I. Rudnicky
INTERSPEECH2
2012 The Structure and Generality of Spoken Route Instructions
Aasish Pappu, Alexander I. Rudnicky
SIGDIAL Conference2
2011 OOV Detection and Recovery Using Hybrid Models with Different Fragments
abstract
In this paper, we address the out-of-vocabulary (OOV) detection and recovery problem by developing three different fragment-word hybrid systems. A fragment language model (LM) and a word LM were trained separately and then combined into a single hybrid LM. Using this hybrid model, the recognizer can recognize any OOVs as fragment sequences. Different types of fragments, such as phones, subwords, and graphones were tested and compared on the WSJ 5k and 20k evaluation sets. The experiment results show that the subword and graphone hybrid systems perform better than the phone hybrid system in both 5k and 20k tasks. Furthermore, given less training data, the subword hybrid system is more preferable than the graphone hybrid system.
Ming Sun 0001, Alexander I. Rudnicky
INTERSPEECH3
2011 SiMPE: 6th Workshop on Speech in Mobile and Pervasive Environments
abstract
With the proliferation of pervasive devices and the increase in their processing capabilities, client-side speech processing has been emerging as a viable alternative. The SiMPE workshop series started in 2006 [5] with the goal of enabling speech processing on mobile and embedded devices to meet the challenges of pervasive environments (such as noise) and leveraging the context they offer (such as location). SiMPE 2010, the latest in the series brought together, very successfully, researchers from the speech and the HCI communities. We believe this is the beginning.
Amit Anil Nanavati, Nitendra Rajput, Alexander I. Rudnicky, Markku Turunen, Andrew L. Kun, Tim Paek, Ivan Tashev
Mobile HCI3
2010 Using the Amazon Mechanical Turk for transcription of spoken language
abstract
We investigate whether Amazon's Mechanical Turk (MTurk) service can be used as a reliable method for transcription of spoken language data. Utterances with varying speaker demographics (native and non-native English, male and female) were posted on the MTurk marketplace together with standard transcription guidelines. Transcriptions were compared against transcriptions carefully prepared in-house through conventional (manual) means. We found that transcriptions from MTurk workers were generally quite accurate. Further, when transcripts for the same utterance produced by multiple workers were combined using the ROVER voting scheme, the accuracy of the combined transcript rivaled that observed for conventional transcription methods. We also found that accuracy is not particularly sensitive to payment amount, implying that high quality results can be obtained at a fraction of the cost and turnaround time of conventional methods.
Matthew Marge, Satanjeev Banerjee, Alexander I. Rudnicky
ICASSP3
2010 The effect of lattice pruning on MMIE training
abstract
In discriminative training, such as Maximum Mutual Information Estimation (MMIE) training, a word lattice is usually used as a compact representation of many different sentence hypotheses and hence provides an efficient representation of the confusion data. However, in a large vocabulary continuous speech recognition (LVCSR) system trained from hundreds or thousands hours training data, the extended Baum-Welch (EBW) computation on the word lattice is still very expensive. In this paper, we investigated the effect of lattice pruning on MMIE training, where we tested the MMIE performance trained with different lattice complexity. A beam pruning and a posterior probability pruning method were applied to generate different sizes of word lattices. The experimental results show that using the posterior probability lattice pruning algorithm, we can save about 40% of the total computation and get the same or more improvement compared to the baseline MMIE result.
Alexander I. Rudnicky
ICASSP2
2010 SiMPE: 5th workshop on speech in mobile and pervasive environments
abstract
With the proliferation of pervasive devices and the increase in their processing capabilities, client-side speech processing has been emerging as a viable alternative. The SiMPE workshop series started in 2006 [5] with the goal of enabling speech processing on mobile and embedded devices to meet the challenges of pervasive environments (such as noise) and leveraging the context they offer (such as location).
Amit Anil Nanavati, Nitendra Rajput, Alexander I. Rudnicky, Markku Turunen, Andrew L. Kun, Tim Paek, Ivan Tashev
Mobile HCI3
2010 Towards Improving the Naturalness of Social Conversations with Dialogue Systems
Matthew Marge, João Miranda, Alan W. Black, Alexander I. Rudnicky
SIGDIAL Conference4
2010 Comparing Spoken Language Route Instructions for Robots across Environment Representations
Matthew Marge, Alexander I. Rudnicky
SIGDIAL Conference2
2010 Let's Buy Books: Finding eBooks using voice search
abstract
We describe Let's Buy Books, a dialog system that helps users search for eBook titles. In this paper we compare different vector space approaches to voice search and find that a hybrid approach using a weighted sub-space model smoothed with a general model provides the best performance over different conditions and evaluated using both synthetic queries and queries collected from users through questionnaires.
Cheongjae Lee, Alexander I. Rudnicky, Gary Geunbae Lee
SLT2
2009 Combining mixture weight pruning and quantization for small-footprint speech recognition
abstract
Semi-continuous acoustic models, where the output distributions for all Hidden Markov Model states share a common codebook of Gaussian density functions, are a well-known and proven technique for reducing computation in automatic speech recognition. However, the size of the parameter files, and thus their memory footprint at runtime, can be very large. We demonstrate how non-linear quantization can be combined with a mixture weight distribution pruning technique to halve the size of the models with minimal performance overhead and no increase in error rate.
David Huggins-Daines, Alexander I. Rudnicky
ICASSP2
2009 SiMPE: Fourth Workshop on Speech in Mobile and Pervasive Environments
Amit Anil Nanavati, Nitendra Rajput, Alexander I. Rudnicky, Markku Turunen, Andrew L. Kun, Tim Paek, Ivan Tashev
Mobile HCI3
2009 Detecting the Noteworthiness of Utterances in Human Meetings
Satanjeev Banerjee, Alexander I. Rudnicky
SIGDIAL Conference2
2009 The RavenClaw dialog management framework: Architecture and systems
Dan Bohus, Alexander I. Rudnicky
Comput. Speech Lang.2
2008 Acquiring Domain-Specific Dialog Information from Task-Oriented Human-Human Interaction through an Unsupervised Learning
Ananlada Chotimongkol, Alexander I. Rudnicky
EMNLP2
2008 Automatic Extraction of Briefing Templates
Dipanjan Das 0001, Alexander I. Rudnicky
IJCNLP3
2008 SiMPE: third workshop on speech in mobile and pervasive environments
abstract
In the past, voice-based applications have been accessed using unintelligent telephone devices through Voice Browsers that reside on the server. The proliferation of pervasive devices and the increase in their processing capabilities, clientside speech processing has been emerging as a viable alternative. In SiMPE 2008, the third in the series, we will continue to explore the various possibilities and issues that arise while enabling speech processing on resource-constrained, possibly mobile devices.
Amit Anil Nanavati, Nitendra Rajput, Alexander I. Rudnicky, Markku Turunen
Mobile HCI3
2008 An extractive-summarization baseline for the automatic detection of noteworthy utterances in multi-party human-human dialog
abstract
Our goal is to reduce meeting participants' note-taking effort by automatically identifying utterances whose contents meeting participants are likely to include in their notes. Though note-taking is different from meeting summarization, these two problems are related. In this paper we apply techniques developed in extractive meeting summarization research to the problem of identifying noteworthy utterances. We show that these algorithms achieve an f-measure of 0.14 over a 5-meeting sequence of related meetings. The precision - 0.15 - is triple that of the trivial baseline of simply labeling every utterance as noteworthy. We also introduce the concept of ldquoshow-worthyrdquo utterances - utterances that contain information that could conceivably result in a note. We show that such utterances can be recognized with an 81% accuracy (compared to 53% accuracy of a majority classifier). Further, if non-show-worthy utterances are filtered out, the precision of noteworthiness detection improves by 33% relative.
Satanjeev Banerjee, Alexander I. Rudnicky
SLT2
2007 TeamTalk: A Platform for Multi-Human-Robot Dialog Research in Coherent Real and Virtual Spaces
Thomas K. Harris, Alexander I. Rudnicky
AAAI2
2007 Data selection for speech recognition
abstract
This paper presents a strategy for efficiently selecting informative data from large corpora of transcribed speech. We propose to choose data uniformly according to the distribution of some target speech unit (phoneme, word, character, etc). In our experiment, in contrast to the common belief that "there is no data like more data", we found it possible to select a highly informative subset of data that produces recognition performance comparable to a system that makes use of a much larger amount of data. At the same time, our selection process is efficient and fast.
Yi Wu 0002, Rong Zhang 0003, Alexander I. Rudnicky
ASRU3
2007 Learning from the Report-writing Behavior of Individuals
Nikesh Garera, Alexander I. Rudnicky
IJCAI3
2007 Segmenting meetings into agenda items by extracting implicit supervision from human note-taking
abstract
Splitting a meeting into segments such that each segment contains discussions on exactly one agenda item is useful for tasks such as retrieval and summarization of agenda item discussions. However, accurate topic segmentation of meetings is a difficult task. In this paper, we investigate the idea of acquiring implicit supervision from human meeting participants to solve the segmentation problem. Specifically we have implemented and tested a note taking interface that gives value to users by helping them organize and retrieve their notes easily, but that also extracts a segmentation of the meeting based on note taking behavior. We show that the segmentation so obtained achieves a Pk value of 0.212 which improves upon an unsupervised baseline by 45% relative, and compares favorably with a current state–of–the–art algorithm. Most importantly, we achieve this performance without any features or algorithms in the classic sense.
Satanjeev Banerjee, Alexander I. Rudnicky
IUI2
2006 Pocketsphinx: A Free, Real-Time Continuous Speech Recognition System for Hand-Held Devices
abstract
The availability of real-time continuous speech recognition on mobile and embedded devices has opened up a wide range of research opportunities in human-computer interactive applications. Unfortunately, most of the work in this area to date has been confined to proprietary software, or has focused on limited domains with constrained grammars. In this paper, we present a preliminary case study on the porting and optimization of CMU Sphinx-11, a popular open source large vocabulary continuous speech recognition (LVCSR) system, to hand-held devices. The resulting system operates in an average 0.87 times real-time on a 206 MHz device, 8.03 times faster than the baseline system. To our knowledge, this is the first hand-held LVCSR system available under an open-source license
David Huggins-Daines, Arthur Chan, Alan W. Black, Mosur Ravishankar, Alexander I. Rudnicky
ICASSP (1)6
2006 A New Data Selection Approach for Semi-Supervised Acoustic Modeling
abstract
Current approaches to semi-supervised incremental learning prefer to select unlabeled examples predicted with high confidence for model re-training. However, this strategy can degrade the classification performance rather than improve it. We present an analysis for the reasons of this phenomenon, showing that only relying on high confidence for data selection can lead to an erroneous estimate to the true distribution when the confidence annotator is highly correlated with the classifier in the information they use. We propose a new data selection approach to address this problem and apply it to a variety of applications, including machine learning and speech recognition. Encouraging improvements in recognition accuracy are observed in our experiments
Rong Zhang 0003, Alexander I. Rudnicky
ICASSP (1)2
2006 A Briefing Tool that Learns Individual Report-Writing Behavior
abstract
We describe a briefing system that learns to predict the contents of reports generated by users who create periodic (weekly) reports as part of their normal activity. We address the question whether data derived from the implicit supervision provided by end-users is robust enough to support not only model parameter tuning but also a form of feature discovery. The system was evaluated under realistic conditions, by collecting data in a project-based university course where student group leaders were tasked with preparing weekly reports for the benefit of the instructors, using the material from individual student reports
Nikesh Garera, Alexander I. Rudnicky
ICTAI3
2006 A texttiling based approach to topic boundary detection in meetings
abstract
Our goal is to automatically detect boundaries between discussions of different topics in meetings. Towards this end we adapt the TextTiling algorithm [1] to the context of meetings. Our features include not only the overlapped words between adjacent windows, but also overlaps in the amount of speech contributed by each meeting participant. We evaluate our algorithm by comparing the automatically detected boundaries with the true ones, and computing precision, recall and f–measure. We report average precision of 0.85 and recall of 0.59 when segmenting unseen test meetings. Error analysis of our results shows that although the basic idea of our algorithm is sound, it breaks down when participants stray from typical behavior (such as when they monopolize the conversation for too long).
Satanjeev Banerjee, Alexander I. Rudnicky
INTERSPEECH2
2006 A constrained baum-welch algorithm for improved phoneme segmentation and efficient training
abstract
We describe an extension to the Baum-Welch algorithm for training Hidden Markov Models that uses explicit phoneme segmentation to constrain the forward and backward lattice. The HMMs trained with this algorithm can be shown to improve the accuracy of automatic phoneme segmentation. In addition, this algorithm is significantly more computationally efficient than the full BaumWelch algorithm, while producing models that achieve equivalent accuracy on a standard phoneme recognition task.
David Huggins-Daines, Alexander I. Rudnicky
INTERSPEECH2
2006 Investigations of issues for using multiple acoustic models to improve continuous speech recognition
abstract
This paper investigates two important issues in constructing and combining ensembles of acoustic models for reducing recognition errors. First, we investigate the applicability of the AnyBoost algorithm for acoustic model training. AnyBoost is a generalized Boosting method that allows the use of an arbitrary loss function as the training criterion to construct ensemble of classifiers. We choose the MCE discriminative objective function for our experiments. Initial test results on a real-world meeting recognition corpus show that AnyBoost is a competitive alternate to the standard AdaBoost algorithm. Second, we investigate ROVER-based combination, focusing on the technique for selecting correct hypothesized words from aligned WTN. We propose a neural network based insertion detection and word scoring scheme for this. Our approach consistently outperforms the current voting technique used by ROVER in the experiments.
Rong Zhang 0003, Alexander I. Rudnicky
INTERSPEECH2
2006 SiMPE: speech in mobile and pervasive environments
abstract
Traditionally, voice-based applications have been accessed using unintelligent telephone devices through Voice Browsers that reside on the server. The proliferation of pervasive devices and the increase in their processing capabilities, client-side speech processing is emerging as a viable alternative. This workshop will explore the various possibilities and issues that arise while enabling speech processing on resource-constrained, possibly mobile devices. The workshop will highlight the many open areas that require research attention, identify key problems that need to be addressed, and also discuss a few approaches for solving some of them - to build the next generation of conversational systems.
Amit Anil Nanavati, Nitendra Rajput, Alexander I. Rudnicky, Roberto Sicconi
Mobile HCI3
2006 SmartNotes: Implicit Labeling of Meeting Data through User Note-Taking and Browsing
Satanjeev Banerjee, Alexander I. Rudnicky
HLT-NAACL2
2006 Online Supervised Learning of Non-Understanding Recovery Policies
abstract
Spoken dialog systems typically use a limited number of non- understanding recovery strategies and simple heuristic policies to engage them (e.g. first ask user to repeat, then give help, then transfer to an operator). We propose a supervised, online method for learning a non-understanding recovery policy over a large set of recovery strategies. The approach consists of two steps: first, we construct runtime estimates for the likelihood of success of each recovery strategy, and then we use these estimates to construct a policy. An experiment with a publicly available spoken dialog system shows that the learned policy produced a 12.5% relative improvement in the non-understanding recovery rate.
Dan Bohus, Brian Langner, Antoine Raux, Alan W. Black, Maxine Eskénazi, Alexander I. Rudnicky
SLT6
2005 The Necessity of a Meeting Recording and Playback System, and the Benefit of Topic-Level Annotations to Meeting Browsing
Satanjeev Banerjee, Carolyn P. Rosé, Alexander I. Rudnicky
INTERACT3
2005 A principled approach for rejection threshold optimization in spoken dialog systems
abstract
A common design pattern in spoken dialog systems is to reject an input when the recognition confidence score falls below a preset rejection threshold. However, this introduces a potentially non-optimal tradeoff between various types of errors such as misunderstandings and false rejections. In this paper, we propose a data-driven method for determining the relative costs of these errors, and then use these costs to optimize state-specific rejection thresholds. We illustrate the use of this approach with data from a spoken dialog system that handles conference room reservations. The results obtained confirm our intuitions about the costs of the errors, and are consistent with anecdotal evidence gathered throughout the use of the system.
Dan Bohus, Alexander I. Rudnicky
INTERSPEECH2
2005 On improvements to CI-based GMM selection
abstract
Gaussian Mixture Model (GMM) computation is known to be one of the most computation-intensive components in speech decoding. In our previous work, context-independent model based GMM selection (CIGMMS) was found to be an effective way to reduce the cost of GMM computation without significant loss in recognition accuracy. In this work, we propose three methods to further improve the performance of CIGMMS. Each method brings an additional 5-10% relative speed improvement, with a cumulative improvement up to 37% on some tasks. Detailed analysis and experimental results on three corpora are presented.
Arthur Chan, Mosur Ravishankar, Alexander I. Rudnicky
INTERSPEECH3
2005 Investigations on ensemble based semi-supervised acoustic model training
abstract
Computer Science Department
Rong Zhang 0003, Ziad Al Bawab, Arthur Chan, Ananlada Chotimongkol, David Huggins-Daines, Alexander I. Rudnicky
INTERSPEECH6
2004 Segmentation and classification of meetings using multiple information streams
abstract
We present a meeting recorder infrastructure used to record and annotate events that occur in meetings. Multiple data streams are recorded and analyzed in order to infer a higher-level state of the group’s activities. We describe the hardware and software systems used to capture people’s activities as well as the methods used to characterize them.
Paul E. Rybski, Satanjeev Banerjee, Fernando De la Torre, Carlos Vallespí, Alexander I. Rudnicky, Manuela M. Veloso
ICMI5
2004 Using simple speech-based features to detect the state of a meeting and the roles of the meeting participants
abstract
We introduce a simple taxonomy of meeting states and participant roles. Our goal is to automatically detect the state of a meeting and the role of each meeting participant and to do so concurrent with a meeting. We trained a decision tree classifier that learns to detect these states and roles from simple speech–based features that are easy to compute automatically. This classifier detects meeting states 18% absolute more accurately than a random classifier, and detects participant roles 10% absolute more accurately than a majority classifier. The results imply that simple, easy to compute features can be used for this purpose.
Satanjeev Banerjee, Alexander I. Rudnicky
INTERSPEECH2
2004 Four-layer categorization scheme of fast GMM computation techniques in large vocabulary continuous speech recognition systems
abstract
Large vocabulary continuous speech recognition systems are known to be computationally intensive. A major bottleneck is the Gaussian mixture model (GMM) computation and various techniques have been proposed to address this problem. We present a systematic study of fast GMM computation techniques. As there are a large number of these and it is impractical to exhaustively evaluate all of them, we first categorized techniques into four layers and selected representative ones to evaluate in each layer. Based on this framework of study, we provide a detailed analysis and comparison of GMM computation techniques from the four-layer perspective and explore two subtle practical issues, 1) how different techniques can be combined effectively and 2) how beam pruning will affect the performance of GMM computation techniques. All techniques are evaluated in the CMU Communicator domain. We also compare their performance with others reported in the literature. 1.
Arthur Chan, Mosur Ravishankar, Alexander I. Rudnicky, Jahanzeb Sherwani
INTERSPEECH3
2004 A frame level boosting training scheme for acoustic modeling
abstract
Conventional Boosting algorithms for acoustic modeling have two notable weaknesses. (1) The objective function aims to minimize utterance error rate, though the goal for most speech recognition systems is to reduce word error rate. (2) During Boosting training, an utterance is treated as a unit for resampling and each frame within the same utterance is assigned equal weight. Intuitively, the frames associated with a is classified word should be given more emphasis than others. We propose a frame level Boosting training scheme that addresses these shortcomings and allows each frame to have a different weight. We describe a technique and provide experimental results for this approach.
Rong Zhang 0003, Alexander I. Rudnicky
INTERSPEECH2
2004 Optimizing boosting with discriminative criteria
abstract
We describe the use of discriminative criteria to optimize Boosting based ensembles. Boosting algorithms may create hundreds of individual classifiers in order to fit the training data. However, this strategy isn’t feasible and necessary for complex classification problems, such as real-time continuous speech recognition, in which only the combination of a few of acoustic models is practical. How to improve the classification accuracy for small size of ensemble is the focus of this paper. Two discriminative criteria that attempt to minimize the true Bayes error rate are investigated. Improvements are observed over a variety of datasets including image and speech recognition, indicating the prospective utility of these two criteria.
Rong Zhang 0003, Alexander I. Rudnicky
INTERSPEECH2
2004 Apply n-best list re-ranking to acoustic model combinations of boosting training
abstract
The object function for Boosting training method in acoustic modeling aims to reduce utterance level error rate. This is different from the most commonly used performance metric in speech recognition, word error rate. This paper proposes that the combination of N-best list re-ranking and ROVER can partly address this problem. In particular, model combination is applied to re-ranked hypotheses rather than to the original top-1 hypotheses and carried on word level. Improvement of system performance is observed in our experiments. In addition, we describe and evaluate a new confidence feature that measures the correctness of frame level decoding result.
Rong Zhang 0003, Alexander I. Rudnicky
INTERSPEECH2
2003 Improving the performance of an LVCSR system through ensembles of acoustic models
abstract
This paper describes our work on applying ensembles of acoustic models to the problem of large vocabulary continuous speech recognition (LVCSR). We propose three algorithms for constructing ensembles. The first two have their roots in bagging algorithms; however, instead of randomly sampling examples our algorithms construct training sets based on the word error rate. The third one is a boosting style algorithm. Different from other boosting methods which demand large resources for computation and storage, our method present a more efficient solution suitable for acoustic model training. We also investigate a method that seeks optimal combination for models. We report experimental results on a large real world corpus collected from the Carnegie Mellon Communicator dialog system. Significant improvements on system performance are observed in that up to 15.56% relative reduction on word error rate is achieved.
Rong Zhang 0003, Alexander I. Rudnicky
ICASSP (1)2
2003 Ravenclaw: dialog management using hierarchical task decomposition and an expectation agenda
abstract
We describe RavenClaw, a new dialog management framework developed as a successor to the Agenda [1] architecture used in the CMU Communicator. RavenClaw introduces a clear separation between task and discourse behavior specification, and allows rapid development of dialog management components for spoken dialog systems operating in complex, goal-oriented domains. The system development effort is focused entirely on the specification of the dialog task, while a rich set of domain-independent conversational behaviors are transparently generated by the dialog engine. To date, RavenClaw has been applied to five different domains allowing us to draw some preliminary conclusions as to the generality of the approach. We briefly describe our experience in developing these systems.
Dan Bohus, Alexander I. Rudnicky
INTERSPEECH2
2003 Comparative study of boosting and non-boosting training for constructing ensembles of acoustic models
abstract
This paper compares the performance of Boosting and non-Boosting training algorithms in large vocabulary continuous speech recognition (LVCSR) using ensembles of acoustic models. Both algorithms demonstrated significant word error rate reduction on the CMU Communicator corpus. However, both algorithms produced comparable improvements, even though one would expect that the Boosting algorithm, which has a solid theoretic foundation, should work much better than the non-Boosting algorithm. Several voting schemes for hypothesis combining were evaluated, including weighted voting, un-weighted voting and ROVER. 1.
Rong Zhang 0003, Alexander I. Rudnicky
INTERSPEECH2
2002 Building voiceXML-based applications
abstract
The Language Technologies Institute (LTI) at Carnegie Mellon University has, for the past several years, conducted a lab course in building spoken-language dialog systems. In the most recent versions of the course, we have used (commercial) web-based development environments to build systems. This paper describes our experiences and discusses the characteristics of applications that are developed within this framework.
Christina L. Bennett, Ariadna Font Llitjós, Stefanie Shriver, Alexander I. Rudnicky, Alan W. Black
INTERSPEECH4
2002 The carnegie mellon communicator corpus
abstract
As part of the DARPA Communicator program, Carnegie Mellon has, over the past three years, collected a large corpus of speech produced by callers to its Travel Planning system. To date, a total of 180,605 utterances (90.9 hours) have been collected. The data were used for a number of purposes, including acoustic and language modeling and the development of a spoken dialog system. The collection, transcription and annotation of these data prompted us to develop a number of procedures for managing the transcription process and for ensuring accuracy. We describe these, as well as some results based on these data. A portion of this corpus, covering the years 1999-2001, is being published for research purposes. 1.
Christina L. Bennett, Alexander I. Rudnicky
INTERSPEECH2
2002 Rapid development of speech-to-speech translation systems
Alan W. Black, Ralf D. Brown, Robert E. Frederking, Kevin A. Lenzo, John Moody, Alexander I. Rudnicky, Rita Singh, Eric Steinbrecher
INTERSPEECH6
2002 Automatic concept identification in goal-oriented conversations
abstract
We address the problem of identifying key domain concepts automatically from an unannotated corpus of goal-oriented human-human conversations. We examine two clustering algorithms, one based on mutual information and another one based on Kullback-Liebler distance. In order to compare the results from both techniques quantitatively, we evaluate the outcome clusters against reference concept labels using precision and recall metrics adopted from the evaluation of topic identification task. However, since our system allows more than one cluster to associate with each concept an additional metric, a singularity score, is added to better capture cluster quality. Based on the proposed quality metrics, the results show that Kullback-Liebler-based clustering outperforms mutual information-based clustering for both the optimal quality and the quality achieved using an automatic stopping criterion
Ananlada Chotimongkol, Alexander I. Rudnicky
INTERSPEECH2
2002 DARPA communicator evaluation: progress from 2000 to 2001
abstract
This paper describes the evaluation methodology and results of the DARPA Communicator spoken dialog system evaluation experiments in 2000 and 2001. Nine spoken dialog systems in the travel planning domain participated in the experiments resulting in a total corpus of 1904 dialogs. We describe and compare the experimental design of the 2000 and 2001 DARPA evaluations. We describe how we established a performance baseline in 2001 for complex tasks. We present our overall approach to data collection, the metrics collected, and the application of PARADISE to these data sets. We compare the results we achieved in 2000 for a number of core metrics with those for 2001. These results demonstrate large performance improvements from 2000 to 2001 and show that the Communicator program goal of conversational interaction for complex tasks has been achieved.
Marilyn A. Walker, Alexander I. Rudnicky, John S. Aberdeen, Elizabeth Owen Bratt, John S. Garofolo, Helen Hastie, Audrey N. Le, Bryan L. Pellom, Alexandros Potamianos, Rebecca J. Passonneau, Rashmi Prasad, Salim Roukos, Gregory A. Sanders, Stephanie Seneff, David Stallard
INTERSPEECH2
2002 DARPA communicator: cross-system results for the 2001 evaluation
abstract
This paper describes the evaluation methodology and results of the 2001 DARPA Communicator evaluation. The experiment spanned 6 months of 2001 and involved eight DARPA Communicator systems in the travel planning domain. It resulted in a corpus of 1242 dialogs which include many more dialogues for complex tasks than the 2000 evaluation. We describe the experimental design, the approach to data collection, and the results. We compare the results by the type of travel plan and by system. The results demonstrate some large differences across sites and show that the complex trips are clearly more difficult.
Marilyn A. Walker, Alexander I. Rudnicky, Rashmi Prasad, John S. Aberdeen, Elizabeth Owen Bratt, John S. Garofolo, Helen Hastie, Audrey N. Le, Bryan L. Pellom, Alexandros Potamianos, Rebecca J. Passonneau, Salim Roukos, Gregory A. Sanders, Stephanie Seneff, David Stallard
INTERSPEECH2
2002 Improve latent semantic analysis based language model by integrating multiple level knowledge
abstract
We describe an extension to the use of Latent Semantic Analysis (LSA) for language modeling. This technique makes it easier to exploit long distance relationships in natural language for which the traditional n-gram is unsuited. However, with the growth of length, the semantic representation of the history may be contaminated by irrelevant information, increasing the uncertainty in predicting the next word. To address this problem, we propose a multilevel framework dividing the history into three levels corresponding to document, paragraph and sentence. To combine the three levels of information with the n-gram, a Softmax network is used. We further present a statistical scheme that dynamically determines the unit scope in the generalization stage. The combination of all the techniques leads to a 14% perplexity reduction on a subset of Wall Street Journal, compared with the trigram model.
Rong Zhang 0003, Alexander I. Rudnicky
INTERSPEECH2
2002 Stochastic natural language generation for spoken dialog systems
Alice Oh, Alexander I. Rudnicky
Comput. Speech Lang.2
2001 Is this conversation on track?
abstract
Confidence annotation allows a spoken dialog system to accurately assess the likelihood of misunderstanding at the utterance level and to avoid breakdowns in interaction. We describe experiments that assess the utility of features from the decoder, parser and dialog levels of processing. We also investigate the effectiveness of various classifiers, including Bayesian Networks, Neural Networks, SVMs, Decision Trees, AdaBoost and Naive Bayes, to combine this information into an utterancelevel confidence metric. We found that a combination of a subset of the features considered produced promising results with several of the classification algorithms considered, e.g., our Bayesian Network classifier produced a 45.7% relative reduction in confidence assessment error and a 29.6% reduction relative to a handcrafted rule.
Paul Carpenter 0001, Chun Jin, Rong Zhang 0003, Dan Bohus, Alexander I. Rudnicky
INTERSPEECH6
2001 N-best speech hypotheses reordering using linear regression
abstract
We propose a hypothesis reordering technique to improve speech recognition accuracy in a dialog system. For such systems, additional information external to the decoding process itself is available, in particular features derived from the parse and the dialog. Such features can be combined with recognizer features by means of a linear regression model to predict the most likely entry in the hypothesis list. We introduce the use of concept error rate as an alternative accuracy measurement and compare it withy the use of word error rate. The proposed model performs better than human subjects performing the same hypothesis reordering task.
Ananlada Chotimongkol, Alexander I. Rudnicky
INTERSPEECH2
2001 Universalizing speech: notes from the USI project
abstract
This paper discusses progress in designing a standardized interface for speech interaction with simple machines – the Universal Speech Interface (USI) project. We discuss the motivation for such a design and issues that must be addressed by such an interface. We present our current proposals for handling these issues, and comment on the usability of these approaches based on user interactions with the system. Finally, we discuss future work and plans for the USI project. 1.
Stefanie Shriver, Ronald Rosenfeld, Xiaojin Zhu 0001, Arthur R. Toth, Alexander I. Rudnicky, Markus D. Flückiger
INTERSPEECH5
2001 DARPA communicator dialog travel planning systems: the june 2000 data collection
abstract
This paper describes results of an experiment with 9 different DARPA Communicator Systems who participated in the June 2000 data collection. All systems supported travel planning and utilized some form of mixed-initiative interaction. However they varied in several critical dimensions: (1) They targeted different back-end databases for travel information; (2) The used different modules for ASR,NLU,TTS and dialog management. We describe the experimental design, the approach to data collection, the metrics collected, and results comparing the systems. 1.
Marilyn A. Walker, John S. Aberdeen, Julie E. Boland, Elizabeth Owen Bratt, John S. Garofolo, Lynette Hirschman, Audrey N. Le, Sungbok Lee, Shri Narayanan, Kishore Papineni, Bryan L. Pellom, Joseph Polifroni, Alexandros Potamianos, P. Prabhu, Alexander I. Rudnicky, Gregory A. Sanders, Stephanie Seneff, David Stallard, Steve Whittaker 0001
INTERSPEECH15
2001 Word level confidence annotation using combinations of features
abstract
This paper describes the development of a word-level confidence metric suitable for use in a dialog system. Two aspects of the problems are investigated: the identification of useful features and the selection of an effective classifier. We find that two parse-level features, Parsing-Mode and SlotBackoff-Mode, provide annotation accuracy comparable to that observed for decoder-level features. However, both decoderlevel and parse-level features independently contribute to confidence annotation accuracy. In comparing different classification techniques, we found that Support Vector Machines (SVMs) appear to provide the best accuracy. Overall we achieve 39.7% reduction in annotation uncertainty for a binary confidence decision in a travel-planning domain.
Rong Zhang 0003, Alexander I. Rudnicky
INTERSPEECH2
2000 Task and domain specific modelling in the Carnegie Mellon communicator system
abstract
The Carnegie Mellon Communicator is a telephone-based dialog system that supports planning in a travel domain. The implementation of such a system requires two complimentary components, an architecture capable of managing interaction and the task, as well as a knowledge base that captures the speech, language and task characteristics specific to the domain. Given a suitable architecture, the principal effort in development in taken up in the acquisition and processing of a domain knowledge base. This paper describes a variety of techniques we have applied to modeling in acoustic, language, task, generation and synthesis components of the system. 1. INTRODUCTION System development involves a great deal of knowledge engineering, which is both time-consuming and requires a variety of experts to participate in the process. Therefore methods that seek to minimize this resource, for example through training based on domain-specific corpora are preferred. Effective use of corpora, however, ...
Alexander I. Rudnicky, Christina L. Bennett, Alan W. Black, Ananlada Chotimongkol, Kevin A. Lenzo, Alice Oh, Rita Singh
INTERSPEECH1
2000 Language modeling for dialog system
abstract
Language modeling for speech recognizer in dialog systems can take two forms. Human input can be constrained through a directed dialog, allowing the decoder to use a state-specific language model to improve recognition accuracy. Mixedinitiative systems allow for human input that while domainspecific might not be state-specific. Nevertheless, for the most part human input to a mixed-initiative system is predictable, particularly when given information about the immediately preceding system prompt. The work reported in this paper addresses the problem of balancing state-specific and general language modeling in a mixed-initiative dialog system. By incorporating dialog state adaptation of the language model, we have reduced the recognition error rate by 11.5%
Wei Xu 0017, Alexander I. Rudnicky
INTERSPEECH2
2000 Can artificial neural networks learn language models?
abstract
Currently, N-gram models are the most common and widely used models for statistical language modeling. In this paper, we investigated an alternative way to build language models, i.e., using artificial neural networks to learn the language model. Our experiment result shows that the neural network can learn a language model that has performance even better than standard statistical methods. 1.
Wei Xu 0017, Alexander I. Rudnicky
INTERSPEECH2
2000 Interactive Speech Translation in the Diplomat Project
Robert E. Frederking, Alexander I. Rudnicky, Christopher Hogan, Kevin A. Lenzo
Mach. Transl.2
1999 Dialog analysis in the carnegie mellon communicator
abstract
In this paper, we present a formative evaluation procedure that we have applied to the Communicator dialog system. In the system improvement process, we have recognized the need to identify interaction failures through passive observation of system use. By systematizing the process of dialog evaluation, we hope to gain a mechanism for effectively communicating descriptions of interaction failures, specifically for use in system improvement. Additionally, we argue that this process can be taught to and executed by an evaluator external to the system development process, with the same proficiency as someone intimately familiar with the mechanics of the system components.
Paul C. Constantinides, Alexander I. Rudnicky
EUROSPEECH2
1999 Data collection and processing in the carnegie mellon communicator
Maxine Eskénazi, Alexander I. Rudnicky, Karin Gregory, Paul C. Constantinides, Robert Brennan, Christina L. Bennett, Jwan Allen
EUROSPEECH2
1999 Creating natural dialogs in the carnegie mellon communicator system
Alexander I. Rudnicky, Eric H. Thayer, Paul C. Constantinides, Chris Tchou, R. Shern, Kevin A. Lenzo, Alice Oh
EUROSPEECH1
1999 A new approach to the translating telephone
abstract
The Translating Telephone has been a major goal of speech translation for many years. Previous approaches have attempted to work from limited-domain, fully-automatic translation towards broad-coverage, fully-automatic translation. We are approaching the problem from a different direction: starting with a broad-coverage but not fully-automatic system, and working towards full automation. We believe that working in this direction will provide us with better feedback, by observing users and collecting language data under realistic conditions, and thus may allow more rapid progress towards the same ultimate goal. Our initial approach relies on the wide-spread availability of Internet connections and web browsers to provide a user interface. We describe our initial work, which is an extension of the Diplomat wearable speech translator.
Robert E. Frederking, Christopher Hogan, Alexander I. Rudnicky
MTSummit3
1998 A schema based approach to dialog control
abstract
Frame-based approaches to spoken language interaction work well for limited tasks such as information access, given that the goal of the interaction is to construct a correct query then execute it. More complex tasks, however, can benefit from more active system participation. We describe two mechanisms that provide this, a modified stack that allows the system to track multiple topics, and form-specific schema that allow the system to deal with tasks that involve completion of multiple forms. Domain-dependent schema specify system behavior and are executed by a domain-independent engine. We describe implementations for a personal calendar system and for an air travel planning system. 1. INTRODUCTION The success of frame-based information access systems, such as for the ATIS domain (e.g., Ward & Issar, 1994) and others (e.g, Goddeau et al., 1996), leads to the question of whether such an approach could be adapted for domains that may require more sophisticated dialog management. Idea...
Paul C. Constantinides, Scott Hansma, Chris Tchou, Alexander I. Rudnicky
ICSLP4
1996 Speechwear: a mobile speech system
Alexander I. Rudnicky, Stephen Reed, Eric H. Thayer
ICSLP1
1995 Speech for Multimedia Information Retrieval
abstract
No abstract available.
Alex Hauptmann 0001, Michael Witbrock, Alexander I. Rudnicky
ACM Symposium on User Interface Software and Technology3
1993 Factors affecting choice of speech over keyboard and mouse in a simple data-retrieval task
abstract
This paper describes some recent experiments that assess user mode selection behavior in a multi-modal environment in which actions can be performed with equivalent effect by speech, keyboard or scroller. Results indicate that users freely choose speech over other modalities, even when it is less efcient in objective terms, such as time-to-completion or input error. Additional evidence indicates that users appear to focus on simple input time in making their choice of mode, in effect minimizing the amount of personal effort expended.
Alexander I. Rudnicky
EUROSPEECH1
1992 A performance model of system delay and user strategy selection
abstract
This study lays the ground work for a predictive, zero-parameter engineering model that characterizes the relationship between system delay and user performance. This study specifically investigates how system delays affects a user’s selection of task strategy. Strategy selection is hypothesized to be based on a cost function combining two factors: (1) the effort required to synchronize input with system availability and (2) the accuracy level afforded. Results indicate that users, seeking to minimize effort and maximize accuracy, choose among three strategies - automatic performance, pacing, and monitoring. These findings provide a systematic account of the influence of system delay on user performance, based on adaptive strategy choice drive by cost.
Steven L. Teal, Alexander I. Rudnicky
CHI2
1991 Spoken language interfaces: the OM system
abstract
No abstract available.
Jean-Michel Lunati, Alexander I. Rudnicky
CHI2
1991 Models for evaluating interaction protocols in speech recognition
abstract
Recognitionerrors complicate the assessment of speech systems.This paper presents a new approach to modeling spoken language interaction protocols, based on finite Markov chains.An interaction protocol, prescribed by the interface design, defines a set of primitive transaction steps and the order of their execut ion.The efficiency of an interface depends on the interaction protocol as well as the cost of each different transaction step.Markov chains provide a simple and computationally eflicient method for modeling errorful systems.They allow for detailed comparisons between different interaction protocols and between different modalities.The method is illustrated by application to example protocols.
Alexander I. Rudnicky, Alex Hauptmann 0001
CHI1
1991 Spoken language recognition in an office management domain
abstract
The authors highlight needs related to a voice interface and describe the implementation of a general-purpose spoken language interface, the Carnegie Mellon Spoken Language Shell (CM-SLS). CM-SLS provides voice interface services to different applications running on the same computer. CM-SLS was used to build the Office Manager, a collection of applications that includes an appointment calendar, a personal database, voice mail, and a calculator. The performance of several system components is described.>
Alexander I. Rudnicky, Jean-Michel Lunati, Alexander M. Franz
ICASSP1
1990 Spoken language interaction in a goal-directed task
abstract
To study the spoken language interface in the context of a complex problem-solving task, a group of users are asked to perform a spreadsheet task, alternating voice and keyboard input. A total of 40 tasks are performed by each participant, the first 30 in a group (over several days), the remaining ones a month later. The voice spreadsheet program is extensively instrumented to provide detailed information about the components of the interaction. These data, as well as analysis of the participant's utterances and recognizer output, provide a fairly detailed picture of spoken language interaction. Although task completion by voice takes longer than by keyboard, analysis shows that users would be able to perform the spreadsheet task faster by voice, if two key criteria could be met: recognition occurs in real-time, and the error rate is sufficiently low. This initial experience with a spoken language system also allows the identification of several metrics, beyond those traditionally associated with speech recognition, that can be used to characterize system performance.>
Alexander I. Rudnicky, Michelle Sakamoto, Joseph Polifroni
ICASSP1
1990 Spoken language interaction in a spreadsheet task
Alexander I. Rudnicky, Michelle Sakamoto, Joseph Polifroni
INTERACT1
1988 An unanchored matching algorithm for lexical access
abstract
Describes the lexical access component of the Carnegie-Mellon University (CMU) continuous speech recognition system. The word recognition algorithm operates in a left to right fashion, building words as it traverses an input network. Search is initiated at each node in the input network. The score assigned to a word is a function of both arc phone probabilities assigned by the acoustic phonetic module and knowledge of expected phone duration and frequency of occurrence of different word pronunciations. The algorithm also incorporates knowledge-based strategies to control the number of hypotheses generated by the matcher. These strategies use criteria external to the search. Performance characteristics are reported using a 1029 word lexicon built automatically from standard pronunciation base forms by context-dependent phonetic rules. Lexical rules are independent of specific lexicons and are derived by examination of transcribed speech data. The lexical representation now includes juncture rules that model specific inter-word phenomena. A junction validation module is also described, whose task is to evaluate the connectivity of words in the word hypotheses lattice.>
Alexander I. Rudnicky, Zongge Li, Joseph Polifroni, Eric H. Thayer, Julia L. Gale
ICASSP1
1988 Talking to Computers: An Empirical Investigation
Alex Hauptmann 0001, Alexander I. Rudnicky
Int. J. Man Mach. Stud.2
1987 Lexical access with lattice input
abstract
This paper describes an alternative approach to lexical access in the CMU ANGEL speech recognition system. Using this approach, the asynchronous phonetic hypotheses generated by an acoustic-phonetics module are converted to a directed graph. This graph is compared to a pronunciation dictionary. Performance results for this approach and the original CMU approach are similar. An error analysis indicates promising directions for further work.
Hy Murveit, Mitch Weintraub, Jared Bernstein, Alexander I. Rudnicky
ICASSP5
1987 The lexical access component of the CMU continuous speech recognition system
abstract
The CMU Lexical Access system hypothesizes words from a phonetic lattice, supplemented by a coarse labelling of the speech signal. Word hypotheses are anchored on syllabic nuclei and are generated independently for different parts of the utterance. Junctures between words are resolved separately, on demand from the Parser module. The lexical representation is generated by rule from baseforms, in a completely automatic process. A description of the various components of the system is provided, as well as performance data.
Alexander I. Rudnicky, Lynn K. Baumeister, Kevin H. DeGraaf, Eric Lehmann
ICASSP1