Giuseppe Di Fabbrizio

dblp:76/1870 · DBLP profile ↗
← Back
43ranked-venue papers
11as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 9 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 5 first-authorDatabases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4Human-computer interaction and ubiquitous computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Question answering and dialogue systems · 92% Language models and text generation · 8%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 56% Spatial and temporal data management · 44%
Computer graphics and multimedia
1 paper
Multimedia systems and quality of experience · 100%
Human-computer interaction and pervasive computing
1 paper
Interaction techniques and input · 44% User interface design and tools · 44% Usability and user experience research · 13%

Topics — the 6 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems › dialogue understanding
dialogue act classification
0.112008
Learning the Structure of Task-Driven Human-Human Dialogs · IEEE Trans. Speech Audio Process. 2008
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue
0.112008
Learning the Structure of Task-Driven Human-Human Dialogs · IEEE Trans. Speech Audio Process. 2008
Spatial and temporal data management › spatial analysis
location inference
0.112007
Geotracker: geospatial and temporal RSS navigation · WWW 2007
Natural language and speech › Question answering and dialogue systems
dialogue management
0.112006
Learning the Structure of Task-Driven Human-Human Dialogs · ACL 2006
Interaction techniques and input
voice interaction
0.011998
What can I say? Evaluating a Spoken Language Interface to Email · CHI 1998
Usability and user experience research
user preference
0.011998
What can I say? Evaluating a Spoken Language Interface to Email · CHI 1998

Methods — techniques the papers use, named apart from their topics

personalization · 0.2temporal browsing · 0.1geospatial browsing · 0.1dialog act classification · 0.1data-driven modeling · 0.1middleware aggregation · 0.1location inference · 0.1parse-based hierarchical model · 0.1chunk-based model · 0.1experimental study · 0.0dialogue design · 0.0
YearPublicationVenuePosition
2022 From Product Searches to Conversational Agents for E-Commerce
abstract
As consumers' demand for online shopping substantially increased in the last few years, e-commerce companies are still far from providing a high-quality user experience that may compete with in-store experiences. On the one hand, matching search queries with highly relevant products for discovery and browsing is still a challenge within existing search technologies. Available e-commerce solutions hardly provide tools to optimize product search relevance and fail to integrate user behavior signals into the search optimization pipeline. On the other hand, accessing the rich and complex information concealed in an e-commerce catalog through a search bar has not evolved far since its initial adoption. In this talk, we illustrate how the VUI conversational AI platform has been successfully adopted to both improve the user's experience quality with highly relevant search and discovery results and expand the traditional search bar with conversational agents' technology, enriching the user's experience at each stage of the e-commerce product life cycle. We review in depth some of the key deep learning models as part of the query understanding component and discuss the overall conversation architecture as it integrates with an existing e-commerce catalog. We include real-life demonstrations derived from use cases extracted from deployed systems.
Giuseppe Di Fabbrizio
CIKM1
2019 Active Annotation: Bootstrapping Annotation Lexicon and Guidelines for Supervised NLU Learning
abstract
Natural Language Understanding (NLU) models are typically trained in a supervised learning framework. In the case of intent classification, the predicted labels are predefined and based on the designed annotation schema while the labelling process is based on a laborious task where annotators manually inspect each utterance and assign the corresponding label. We propose an Active Annotation (AA) approach where we combine an unsupervised learning method in the embedding space, a human-in-the-loop verification process, and linguistic insights to create lexicons that can be open categories and adapted over time. In particular, annotators define the y-label space on-the-fly during the annotation using an iterative process and without the need for prior knowledge about the input data. We evaluate the proposed annotation paradigm in a real use-case NLU scenario. Results show that our Active Annotation paradigm achieves accurate and higher quality training data, with an annotation speed of an order of magnitude higher with respect to the traditional human-only driven baseline annotation methodology.
Federico Marinelli, Alessandra Cervone, Giuliano Tortoreto, Evgeny A. Stepanov, Giuseppe Di Fabbrizio, Giuseppe Riccardi
INTERSPEECH5
2018 E-commerce Product Query Classification Using Implicit User's Feedback from Clicks
abstract
Query classification (QC) has been widely studied to understand users' search intent. For e-commerce search queries, users typically search for either a specific product or a category of products. In both cases, a query can be associated with a category label that belongs to a taxonomy tree describing the items in the catalog. However, product-related search queries are typically short, ambiguous, and continuously changing depending on seasonal trends and the introduction of new products over time. Traditional supervised approaches to e-commerce QC are not feasible due to the high cost of manual annotation and the high volume of traffic on e-commerce search engines. In this work, we introduce an unsupervised method to collect large amounts of query classification data using user's implicit click feedback. We obtain a large multi-label dataset containing 403,349 unique queries from 2,085 categories. We compare and contrast different state-of-the-art text classifiers and demonstrate that an ensemble of linear SVMs models achieves a micro-F1 score of 0.60 and 0.82 at leaf and top level, respectively.
Yiu-Chang Lin, Ankur Datta, Giuseppe Di Fabbrizio
IEEE BigData3
2017 Web-Scale Language-Independent Cataloging of Noisy Product Listings for E-Commerce
abstract
Pradipto Das, Yandi Xia, Aaron Levine, Giuseppe Di Fabbrizio, Ankur Datta. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Pradipto Das, Yandi Xia, Aaron Levine, Giuseppe Di Fabbrizio, Ankur Datta
EACL (1)4
2016 Large-scale taxonomy categorization for noisy product listings
abstract
E-commerce catalogs include a continuously growing number of products that are constantly updated. Each item in a catalog is characterized by several attributes and identified by a taxonomy label. Categorizing products with their taxonomy labels is fundamental to effectively search and organize listings in a catalog. However, manual and/or rule based approaches to categorization are not scalable. In this paper, we compare several classifiers to product taxonomy categorization of top-level categories. We first investigate a number of feature sets and observe that a combination of word unigrams from product names and navigational breadcrumbs work best for categorization. Secondly, we apply correspondence topic models to detect noisy data and introduce a lightweight manual process to improve dataset quality. Finally, we evaluate linear models, gradient boosted trees (GBTs) and convolutional neural networks (CNNs) with pre-trained word embeddings demonstrating that, compared to other baselines, GBTs and CNNs yield the highest gains in error reduction.
Pradipto Das, Yandi Xia, Aaron Levine, Giuseppe Di Fabbrizio, Ankur Datta
IEEE BigData4
2015 Comment-to-Article Linking in the Online News Domain
abstract
Online commenting to news articles provides a communication channel between media professionals and readers offering a crucial tool for opinion exchange and freedom of expression.Currently, comments are detached from the news article and thus removed from the context that they were written for.In this work, we propose a method to connect readers' comments to the news article segments they refer to.We use similarity features to link comments to relevant article segments and evaluate both word-based and term-based vector spaces.Our results are comparable to state-of-theart topic modeling techniques when used for linking tasks.We demonstrate that article segments and comments representation are relevant to linking accuracy since we achieve better performances when similarity features are computed using similarity between terms rather than words.
Ahmet Aker, Emina Kurtic, Mark Hepple, Robert J. Gaizauskas, Giuseppe Di Fabbrizio
SIGDIAL Conference5
2014 A Hybrid Approach to Multi-document Summarization of Opinions in Reviews
abstract
We present a hybrid method to generate summaries of product and services reviews by combining natural language generation and salient sentence selection techniques. Our system, STARLET-H, receives as input textual reviews with associated rated topics, and produces as output a natural language document summarizing the opinions expressed in the reviews. STARLET-H operates as a hybrid
Giuseppe Di Fabbrizio, Amanda Stent, Robert J. Gaizauskas
INLG1
2014 Forms2Dialog: Automatic dialog generation for Web tasks
abstract
Today, many common tasks (e.g. booking flights, ordering food) can be done by filling out web forms. Automatic processing of Web forms to support interactive speech input is useful for numerous reasons, including ease of use for mobile device users and accessibility for people with visual or print disabilities. In this paper, we propose an automated method to process web forms and convert them into dialog flows for spoken interaction. First we identify relevant information for each form element (including element type, label, values and help messages) and key relationships between form elements (including ordering and dependencies). We then generate two types of dialog flow for each Web form. Experimental results show that the method generates efficient and informative dialog flows for web tasks, a key step for building virtual assistants. An Android application has been realized as a use case of the generated dialog flows.
Nobal B. Niraula, Amanda Stent, Hyuckchul Jung, Giuseppe Di Fabbrizio, I. Dan Melamed, Vasile Rus
SLT4
2013 Emotion Detection in Email Customer Care
abstract
Prompt and knowledgeable responses to customers’ email are critical in maximizing customer satisfaction. Such messages often contain complaints about unfair treatment due to negligence, incompetence, rigid protocols, unfriendly systems, and unresponsive personnel. In this paper, we refer to these email messages as emotional email. They provide valuable feedback to improve contact center efficiency and the quality of the overall customer care experience, which in turn results in increased customer retention. We describe a method that uses salient features to identify emotional email in the customer care domain. Salient features in customer care related email are expressions of customer frustration, dissatisfaction with the business, and threats to either leave, take legal action, and/or report to authorities. Compared to a baseline system using word unigram features, our proposed approach significantly improves emotional email detection performance.
Narendra K. Gupta, Mazin Gilbert, Giuseppe Di Fabbrizio
Comput. Intell.3
2012 Can prosody inform sentiment analysis? Experiments on short spoken reviews
abstract
While most online content is created using textual interfaces, recent improvements in speech recognition accuracy allows the creation of content through speech. This technology allows users to share reviews about entities of interest without any delay, using mobile devices. This paper builds on the previous work on textual sentiment analysis to investigate whether information in the speech signal can be used to predict sentiment from short spoken reviews. For this purpose we collected a short spoken reviews from 84 speakers. Results show that models trained on features characterizing the review's pitch significantly outperform a majority class baseline, without textual information. When taking text-based sentiment predictions into account, our results suggest that prosody can alleviate the effect of speech recognition errors on sentiment detection, however a larger dataset is needed to test whether this can be done without harming performance on low word error rates.
François Mairesse, Joseph Polifroni, Giuseppe Di Fabbrizio
ICASSP3
2012 Building Text-To-Speech Voices in the Cloud
Alistair Conkie, Thomas Okken, Yeon-Jun Kim, Giuseppe Di Fabbrizio
LREC4
2011 Semantic data selection for vertical business voice search
abstract
Local business voice search is a popular application for mobile phones, where hands-free interaction and speed are critical to users. However, speech recognition accuracy is still not satisfactory when the number of businesses and locations is extended nationwide. For mobile users, searching a local business directory is often related to the fulfillment of specific tasks “on-the-move”, such as finding a restaurant, a movie theater, or a retailer chain. Restricting the local search to specific domains improves the quality of search results. In this paper, we present a new approach to data selection for bootstrapping and optimizing language models for vertical business sectors by exploiting semantic knowledge encoded in the business database and in the business category taxonomy. We demonstrate that, in the case of queries in the restaurant domain and without collecting new data, speech recognition word accuracy improves by 9.5% relative when compared with a generic local business language model.
Giuseppe Di Fabbrizio, Diamantino Caseiro, Amanda Stent
ICASSP1
2011 SpeechForms: From Web to Speech and Back
abstract
This paper describes SpeechForms, a system that uses novel techniques to automatically identify form element semantics and form element content, and to semi-automatically generate language models that allow users to fill out each web form element by voice. Preliminary experimental results show that simple per-element language models are faster and may be more accurate than statistical n-gram language models trained on large amounts of web text data. Index Terms: language modeling, form understanding, information retrieval
Luciano Barbosa, Diamantino Caseiro, Giuseppe Di Fabbrizio
INTERSPEECH3
2011 mTalk - A Multimodal Browser for Mobile Services
abstract
The MTALKmultimodal browser is a tool which enables rapid prototyping for research and development of mobile multimodal interfaces combining natural modalities such as speech, touch, and gesture. MTALKintegrates a broad range of open standards for authoring graphical and spoken user interfaces and is supported by a cloud-based multimodal processing architecture. In this paper, we describe MTALKand illustrate its capabilities through examination of a series of sample applications. Index Terms: multimodal, browser, speech, gesture
Michael Johnston, Giuseppe Di Fabbrizio, Simon Urbanek
INTERSPEECH2
2011 AT&T VoiceBuilder: A Cloud-Based Text-to-Speech Voice Builder Tool
Yeon-Jun Kim, Thomas Okken, Alistair Conkie, Giuseppe Di Fabbrizio
INTERSPEECH4
2009 A speech mashup framework for multimodal mobile services
abstract
Amid today's proliferation of Web content and mobile phones with broadband data access, interacting with small-form factor devices is still cumbersome. Spoken interaction could overcome the input limitations of mobile devices, but running an automatic speech recognizer with the limited computational capabilities of a mobile device becomes an impossible challenge when large vocabularies for speech recognition must often be updated with dynamic content. One popular option is to move the speech processing resources into the network by concentrating the heavy computation load onto server farms. Although successful services have exploited this approach, it is unclear how such a model can be generalized to a large range of mobile applications and how to scale it for large deployments. To address these challenges we introduce the AT&T speech mashup architecture, a novel approach to speech services that leverages web services and cloud computing to make it easier to combine web content and speech processing. We show that this new compositional method is suitable for integrating automatic speech recognition and text-to-speech synthesis resources into real multimodal mobile services. The generality of this method allows researchers and speech practitioners to explore a countless variety of mobile multimodal services with a finer grain of control and richer multimedia interfaces. Moreover, we demonstrate that the speech mashup is scalable and particularly optimized to minimize round trips in the mobile network, reducing latency for better user experience.
Giuseppe Di Fabbrizio, Thomas Okken, Jay G. Wilpon
ICMI1
2008 Trainable Speaker-Based Referring Expression Generation
Giuseppe Di Fabbrizio, Amanda Stent, Srinivas Bangalore
CoNLL1
2008 Where do thewords come from? Learning models for word choice and ordering from spoken dialog corpora
abstract
Most existing generation systems for spoken dialog require the system engineer to specify by hand the words to be used in system prompts. However, the existence of corpora of spoken dialog makes it possible to acquire the words and structure of system prompts automatically. In this paper, we construct statistical models for generating system prompts, both for word choice and for word ordering. We evaluate these models using a human-computer dialog multicorpus and a human-human dialog corpus. Our results show that statistical models for word choice can work well, while more work is needed on statistical models for word ordering.
Amanda Stent, Srinivas Bangalore, Giuseppe Di Fabbrizio
ICASSP3
2008 Referring Expression Generation Using Speaker-based Attribute Selection and Trainable Realization (ATTR)
Giuseppe Di Fabbrizio, Amanda Stent, Srinivas Bangalore
INLG1
2008 Bootstrapping spoken dialogue systems by exploiting reusable libraries
abstract
Abstract Building natural language spoken dialogue systems requires large amounts of human transcribed and labeled speech utterances to reach useful operational service performances. Furthermore, the design of such complex systems consists of several manual steps. The User Experience (UE) expert analyzes and defines by hand the system core functionalities: the system semantic scope (call-types) and the dialogue manager strategy that will drive the human–machine interaction. This approach is extensive and error-prone since it involves several nontrivial design decisions that can be evaluated only after the actual system deployment. Moreover, scalability is compromised by time, costs, and the high level of UE know-how needed to reach a consistent design. We propose a novel approach for bootstrapping spoken dialogue systems based on the reuse of existing transcribed and labeled data, common reusable dialogue templates, generic language and understanding models, and a consistent design process. We demonstrate that our approach reduces design and development time while providing an effective system without any application-specific data.
Giuseppe Di Fabbrizio, Gökhan Tür, Dilek Hakkani-Tür, Mazin Gilbert, Bernard Renger, David C. Gibbon, Zhu Liu 0001, Behzad Shahraray
Nat. Lang. Eng.1
2008 Learning the Structure of Task-Driven Human-Human Dialogs
abstract
With the availability of large corpora of spoken dialog, it is now possible to use data-driven techniques to build and use models of task-oriented dialogs. In this paper, we use data-driven techniques to build task structures for individual dialogs, and use the dialog task structures for: dialog act classification, task/subtask classification, task/subtask prediction, and dialog act prediction. We evaluate our approach using a corpus of customer/agent dialogs from a catalog service domain. This paper demonstrates the feasibility of using corpora of human–human conversation to learn dialog models suitable for human–computer dialog applications.
Srinivas Bangalore, Giuseppe Di Fabbrizio, Amanda Stent
IEEE Trans. Speech Audio Process.2
2007 GeoTV: navigating geocoded rss to create an iptv experience
abstract
The Web is rapidly moving towards a platform for mass collaboration in content production and consumption from three screens: computers, mobile phones, and TVs. While there has been a surge of interests in making Web content accessible from mobile devices, there is a significant lack of progress when it comes to making the web experience suitable for viewing on a television. Towards this end, we describe a novel concept, namely GeoTV, where we explore a framework by which web content can be presented or pushed in a meaningful manner to create an entertainment experience for the TV audience. Fresh content on a variety of topics, people, and places is being created and made available on the Web at breathtaking speed. Navigating fresh content effectively on TV demands a new browsing paradigm that requires fewer mouse clicks or user interactions from the remote control. Novel geospatial and temporal browsing techniques are provided in GeoTV that allow users the capability of aggregating and navigating RSS-enabled content in a timely, personalized and automatic manner for viewing in an IPTV environment. This poster is an extension of our previous work on GeoTracker that utilizes both a geospatial representation and a temporal (chronological) presentation to help users spot the most relevant updates quickly within the context of a Web-enabled environment. We demonstrate 1) the usability of such a tool that greatly enhances a user.s ability in locating and browsing videos based on his or her geographical interests and 2) various innovative interface designs for showing RSS-enabled information in an IPTV environment.
Yih-Farn Robin Chen, Giuseppe Di Fabbrizio, David C. Gibbon, Rittwik Jana, Serban Jora, Bernard Renger, Bin Wei 0003
WWW2
2007 Geotracker: geospatial and temporal RSS navigation
abstract
The Web is rapidly moving towards a platform for mass collaboration in content production and consumption. Fresh content on a variety of topics, people, and places is being created and made available on the Web at breathtaking speed. Navigating the content effectively not only requires techniques such as aggregating various RSS-enabled feeds, but it also demands a new browsing paradigm. In this paper, we present novel geospatial and temporal browsing techniques that provide users with the capability of aggregating and navigating RSS-enabled content in a timely, personalized and automatic manner. In particular, we describe a system called GeoTracker that utilizes both a geospatial representation and a temporal (chronological) presentation to help users spot the most relevant updates quickly. Within the context of this work, we provide a middleware engine that supports intelligent aggregation and dissemination of RSS feeds with personalization to desktops and mobile devices. We study the navigation capabilities of this system on two kinds of data sets, namely, 2006 World Cup soccer data collected over two months and breaking news items that occur every day. We also demonstrate that the application of such technologies to the video search results returned by YouTube and Google greatly enhances a user.s ability in locating and browsing videos based on his or her geographical interests. Finally, we demonstrate that the location inference performance of GeoTracker compares well against machine learning techniques used in the natural language processing/information retrieval community. Despite its algorithm simplicity, it preserves high recall percentages.
Yih-Farn Robin Chen, Giuseppe Di Fabbrizio, David C. Gibbon, Rittwik Jana, Serban Jora, Bernard Renger, Bin Wei 0003
WWW2
2006 Learning the Structure of Task-Driven Human-Human Dialogs
abstract
Data-driven techniques have been used for many computational linguistics tasks.Models derived from data are generally more robust than hand-crafted systems since they better reflect the distribution of the phenomena being modeled.With the availability of large corpora of spoken dialog, dialog management is now reaping the benefits of data-driven techniques.In this paper, we compare two approaches to modeling subtask structure in dialog: a chunk-based model of subdialog sequences, and a parse-based, or hierarchical, model.We evaluate these models using customer agent dialogs from a catalog service domain.
Srinivas Bangalore, Giuseppe Di Fabbrizio, Amanda Stent
ACL2
2006 Towards Learning to Converse: Structuring Task-Oriented Human-Human Dialogs
abstract
Data-driven techniques have influenced many aspects of speech and language processing. Models derived from data are generally more robust than hand-crafted systems since they better reflect the distributions of the phenomena being modeled. With the availability of large spoken dialog corpora, dialog management can now reap the benefit of data-driven techniques. In this paper, we present our view of structuring human-human dialogs in order to learn models for human-machine dialogs. We present the problems of dialog segmentation and dialog act labeling, develop a model for predicting and labeling topic segments and dialog acts and evaluate the model on customer-agent dialogs from a catalog service domain
Srinivas Bangalore, Giuseppe Di Fabbrizio, Amanda Stent
ICASSP (1)2
2006 Webtalk: Towards Automatically Building Spoken Dialog Systems Through Miningwebsites
abstract
Web Talk is a system for analyzing unstructured information from company websites to support automtic reation of spoken dialog applications. The goal is to completely automate the process of building, maintaining and deploying dialog applications by leveraging the wealth of information on the World Wide Web. Web Talk employs technologies in web mining, document understanding, question/answering, and speech and language processing. In this paper, we review extensions to these technologies to make them suitable for creating a Web Talk application. e present an evaluation study of a Web Talk spoken dialog system that has been instantiated on a telecom company website. Experiments with 30 different scenarios indicate promising results and provide evidence that such systems can potentially revolutionize the paradigm for creating and scaling spoken dialog services.
Junlan Feng, Dilek Hakkani-Tür, Giuseppe Di Fabbrizio, Mazin Gilbert, Marc C. Beutnagel
ICASSP (1)3
2006 Prompt selection with reinforcement learning in an AT&t call routing application
abstract
Reinforcement Learning (RL) algorithms provide a type of unsupervised learning that is especially well suited for the challenges of spoken dialogue systems (SDS) design. SDS are constantly subjected to new environments in the form of new groups of users, and RL provides an approach for automated learning that can adapt to new environments without costly supervision. In this paper, we describe some results from experiments with RL to select prompts for a call routing application. A simulation of the dialogue outcomes were used to experiment with different scenarios and demonstrate how RL can make a system more robust without supervision or developer intervention. Index Terms: spoken dialogue systems, reinforcement learning, call routing
Charles Lewis, Giuseppe Di Fabbrizio
INTERSPEECH2
2006 Let's Discoh: Collecting an Annotated Open Corpuswith Dialogue Acts and Reward signals for Natural Language Helpdesks
abstract
We motivate and explain the DlSCoH project, which uses a publicly deployed spoken dialogue system for conference services to collect a richly annotated corpus of mixed-initiative human- machine spoken dialogues. System users are able to call a phone number and learn about a conference, including paper submission, program, venue, accommodation options and costs, etc. The collected corpus is (1) usable for training, evaluating and comparing statistical models, (2) naturally spoken and task oriented, (3) extendible / generalizable, (4) collected using state-of-the-art research and commercial technology, (5) freely available to researchers. We explain the principles behind the dialogue context representations and reward signals collected by the system, as well as the overall system design, call types, and call flow. We also present results regarding the initial ASR models and spoken language understanding models. We expect the resulting corpora to be used in advanced dialogue research over the coming years.
Giovanni Andreani, Giuseppe Di Fabbrizio, Mazin Gilbert, Daniel Gillick, Dilek Hakkani-Tür, Oliver Lemon
SLT2
2005 A Clarification Algorithm for Spoken Dialogue Systems
abstract
The paper presents an algorithm for spoken dialogue systems that uses mixed initiative interactions and identification of multiple call-types to clarify the needs of the user efficiently. The representation of the application domain includes the relationships between key topics in the domain, the prompts used to discern between these topics, and the call-types associated with the topics. This representation is used by the algorithm to maintain the state of the conversation. By maintaining a picture of how all of the information conveyed by the user fits into this domain, regardless of whether it was information specifically requested by the system, the algorithm expedites the clarification process.
Charles Lewis, Giuseppe Di Fabbrizio
ICASSP (1)2
2005 Automated wizard-of-oz for spoken dialogue systems
abstract
Designing and building natural language spoken dialogue systems require large amounts of speech utterances, which adequately represent the intended human-machine dialogues.For this purpose, typically, first a "Wizard-of-Oz" data collection is performed, and then the collected data is transcribed and labeled by expert labelers.Finally, the data is used to train both the speech recognizer and the spoken language understanding stochastic models.Data collection and labeling is an expensive and time consuming manual process.In this paper we propose a completely Automated Wizard, which is capable of recognizing and understanding application independent requests reusing the previously labeled and transcribed data from similar domains, and improving the informativeness of the collected data.We demonstrate that, in the context of automated call routing, compared to the existing data collection systems, the Automated Wizard better captures the user intentions and produces substantially shorter interactions resulting in a better user experience and a less intrusive approach.
Giuseppe Di Fabbrizio, Gökhan Tür, Dilek Hakkani-Tür
INTERSPEECH1
2004 Florence: a dialogue manager framework for spoken dialogue systems
abstract
Recent advances in speech and language technology have made spoken dialogue systems mainstream in many industries. They allow customers to engage in natural speech interactions with machines instead of being compelled to navigate menus of options with touch tones inputs. VoiceXML was a major milestone for the process of using automated speech applications to expose business portals to ubiquitous telephone access. By the introduction of a uniform and universally accepted client-server browser model, the VoiceXML programming model greatly simplified previously dominant proprietary computer telephony interfaces. However, natural language spoken dialogue systems entail more complex interactions with the user which, depending upon the application domain, may require computational models that are difficult to express directly in VoiceXML. This paper describes Florence, a dialogue manager with a more general approach that uses an extensible and flexible framework to combine interchangeable and interoperable dialogue strategies as appropriate to the task. Florence’s declarative XML-based language facilitates the development of natural language applications and allows the dialogue author to encapsulate and reuse different algorithms between applications. Moreover, it addresses large-scale natural language issues related to enterprise backend access, logging, distributed deployment, and fail-over support. These issues must be addressed in a modern, industrial-strength application server environment.
Giuseppe Di Fabbrizio, Charles Lewis
INTERSPEECH1
2002 AT&t help desk
Giuseppe Di Fabbrizio, Dawn Dutton, Narendra K. Gupta, Barbara Hollister, Mazin G. Rahim, Giuseppe Riccardi, Robert E. Schapire, Juergen Schroeter
INTERSPEECH1
2001 Towards SMIL as a foundation for multimodal, multimedia applications
abstract
Rich and interactive multimedia applications, where audio, video, graphics and text are precisely synchronized under timing constraints are becoming ubiquitous. Multimodal applications further extend the concept of user interaction combining different modalities, like speech recognition, speech synthesis and gestures. However, authoring dialog-capable multimodal, multimedia services is a very difficult task. Fortunately, the W3C has sponsored the development of SMIL, an elegant notation for multimedia applications, which has been embraced by both Microsoft and RealNetworks. In this paper, we argue that SMIL is an ideal substrate for extending multimedia applications with multimodal facilities. SMIL as it stands is not a general notation for controlling media and input mode resources. We show that all what is needed are few natural extensions to SMIL along with the addition of a simple reactive programming language that we call ReX. Our language is designed to be maximally compatible with existing W3C recommendations through a generic event system based on DOM and an expression language based on XPATH. It is also designed to be simple so that the fundamental notion of seeking time (e.g. going backwards and forwards in presentations) is preserved.
Jennifer L. Beckmann, Giuseppe Di Fabbrizio, Nils Klarlund
INTERSPEECH2
2001 Voice-IF: a mixed-initiative spoken dialogue system for AT&t conference services
Mazin G. Rahim, Giuseppe Di Fabbrizio, Candace A. Kamm, Marilyn A. Walker, A. Pokrovsky, P. Ruscitti, Esther Levin, Sungbok Lee, Ann K. Syrdal, K. Schlosser
INTERSPEECH2
2000 Web-based monitoring, logging and reporting tools for multi-service multi-modal systems
abstract
This paper describes MILER (Multi-modal data Logger for Evaluation and Report), a web-based multi-service monitoring, logging and reporting tool for advanced multi-modal dialog systems.MILER has been designed to directly arrange and synchronize logging data collected from live services and to provide real-time reports about service usage and system performance.Special attention has been given to the architecture design in order to achieve service and access-device independence and reliable synchronization of data from distributed logs.MILER allows researchers to analyze multimodal interactions, analyze the call flow, reconstruct the system/user dialogue turns, play the recorded user utterances, and provide a preliminary dialogue performance evaluation.It also supports labeling and annotation of the dialogue turns for further offline analysis.Once the user inputs (i.e.speech and other input modalities) are manually transcribed and labeled, along with detailed log events from each dialog, MILER derives a set of objective measures, which includes word and concept accuracy, number of attempts per concept, dialog turn counts and duration, and task completion rates.Subjective measures extracted from user's surveys, including perceived task success and ease of use measures, can be combined with the objective measures and the results used later for accuracy computation.
Giuseppe Di Fabbrizio, Shri Narayanan
INTERSPEECH1
2000 The AT&t-DARPA communicator mixed-initiative spoken dialog system
abstract
The design and implementation of the AT&T Communicator mixed-initiative spoken dialog system is described.The Communicator project, sponsored by DARPA and launched in 1999, is a multi-year multi-site project on advanced spoken dialog systems research.The main focus of this paper is on issues related to the design of mixed-initiative systems.In addition to describing our architecture and implementation of the complex travel task, the paper reports on some preliminary evaluation results.
Esther Levin, Shri Narayanan, Roberto Pieraccini, Konstantin Biatov, Enrico Bocchieri, Giuseppe Di Fabbrizio, Wieland Eckert, Sungbok Lee, A. Pokrovsky, Mazin G. Rahim, P. Ruscitti, Marilyn A. Walker
INTERSPEECH6
2000 Effects of dialog initiative and multi-modal presentation strategies on large directory information access
abstract
This paper compares the effects of three different dialog initiative strategies (system initiative, mixed initiative and user initiative) on system performance and user acceptance on a large directory information access task. We used a personnel directory query application that could be accessed from a voice-only (telephony) and a multi-modal (kiosk) interface. Although the user initiative condition resulted in a lower proportion of in-grammar utterances, no significant effects of dialog initiative were observed for concept accuracy, perceived task completion, ease of use or user satisfaction. Dialogs were significantly shorter with the kiosk interface than with the telephony interface, and users preferred the kiosk interface and found it easier to use. 1.
Shri Narayanan, Giuseppe Di Fabbrizio, Candace A. Kamm, James Hubbell, Bruce Buntschuh, P. Ruscitti, Jeremy H. Wright
INTERSPEECH2
2000 A spoken dialogue system for conference/workshop services
abstract
This paper describes our progress towards building a telephony-based spoken dialogue system for workshop/conference services. A mixed-initiative dialogue system has been developed that is engineered to o er users natural interaction with the system, ease-of-use and robustness towards ambiguous requests and machine errors. A prototype system, known as W99, is described in this paper
Mazin G. Rahim, Roberto Pieraccini, Wieland Eckert, Esther Levin, Giuseppe Di Fabbrizio, Giuseppe Riccardi, Candace A. Kamm, Shri Narayanan
INTERSPEECH5
1998 What can I say? Evaluating a Spoken Language Interface to Email
abstract
This paper presents experimental results comparing two different designs for a spoken language interface to entail.We compare a mixed-initiative dialogue style, in which users can flexibly control the dialogue, to a systemiuitiative dialogue style, in which the system controls the dialogue.Our results show that even though the mixedinitiative system is more eflicient, as measured by number of turns, or elapsed time to complete a set of email tasks, users prefer the system-initiative interface.We posit that these preferences arise from the fact that the system initiative inter&e is easier to learn and more predictable.
Marilyn A. Walker, Jeanne C. Fromer, Giuseppe Di Fabbrizio, Craig Mestel, Donald Hindle
CHI3
1998 VPQ: a spoken language interface to large scale directory information
abstract
This paper describes VPQ (Voice Post Query), a dialog system that provides spoken access to the information in the AT&T corporate personnel database (>120,000 entries). An explicit design goal is to have the user’s initial interaction with the system be rather unconstrained and to rely on tighter, prompt constrained, dialog only when absolutely necessary. The purpose of VPQ is both a) to explore and exploit the capabilities of “state of the art” speech recognition systems for this highperplexity task, and b) to develop the natural language understanding and dialog control components necessary for effective and efficient user interactions. The VPQ task spans a wide range of possible dialog scenarios. They range from simple “one-shot” to complex multi-turn interactions. The former correspond to interactions where the initial utterance is unambiguous and the system’s response appropriately terminates the interaction either by providing the desired information or completing a call to the requested person. The more complex interactions occur primarily whenever ambiguities or errors require resolution. Current speech recognition accuracy of 80% is adequate to pursue such an ambitious task. This paper highlights the inherent challenges in such a task, the major components of the system, the rationale for their design, and how they perform. The VPQ project targets a variety of access devices, including telephony, desktop and handheld devices offering multi-modal user interfaces. In this paper we focus on describing the telephony interface.
Bruce Buntschuh, Candace A. Kamm, Giuseppe Di Fabbrizio, Alicia Abella, Mehryar Mohri, Shri Narayanan, Ilija Zeljkovic, R. Doug Sharp, Jeremy H. Wright, S. Marcus, J. Shaffer, R. Duncan, Jay G. Wilpon
ICSLP3
1997 Evaluating competing agent strategies for a voice email agent
abstract
This paper reports experimental results comparing a mixed-initiative to a system-initiative dialog strategy in the context of a personal voice email agent. To independently test the effects of dialog strategy and user expertise, users interact with either the system-initiative or the mixed-initiative agent to perform three successive tasks which are identical for both agents. We report performance comparisons across agent strategies as well as over tasks. This evaluation utilizes and tests the PARADISE evaluation framework, and discusses the performance function derivable from the experimental data.
Marilyn A. Walker, Donald Hindle, Jeanne C. Fromer, Giuseppe Di Fabbrizio, Craig Mestel
EUROSPEECH4
1993 SIRVA - a large speech database collected on the Italian telephone network
Giuseppe Castagneri, Giuseppe Di Fabbrizio, Antonio Massone, Mario Oreglia
EUROSPEECH2
1992 Comparison between two methodologies of testing isolated word speech recognizers
Franco Canavesio, Giuseppe Castagneri, Giuseppe Di Fabbrizio, Francesco Senia
ICSLP3