VLDB 2026 Research / reviewers in the wild / expert
Giuseppe Riccardi
dblp:27/1313
· DBLP profile ↗
138ranked-venue papers
11as first author
10since 2021 · last 2025
0000-0002-0739-8184ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 96 · 8 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 88 · 8 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Investigating the Use and Perception of Blocking Feature in Social Virtual Reality Spaces: A Study on Discussion ForumsabstractThis study explores the use of 'blocking' as a safety tool in social VR platforms, a feature widely implemented to protect users from harassment. Despite its prevalence, little research investigates the reasons users block others and how they perceive the feature.To address this gap, we analyzed discussions from r/VRChat, r/AltspaceVR, r/HorizonWorld, and r/RecRoom, using thematic analysis. Our findings reveal that users block others for various reasons, extending beyond mere disruption avoidance to include personal preferences of avatars and social management such as blocking a friend for a disagreement. Furthermore, the study highlights the empowering roles and the limitations of blocking as a solution to harassment. The study contributes to understanding user behavior in social VR and offers insights for developing more effective safety tools in these immersive environments. Qijia Chen, Seyed Mahed Mousavi, Giuseppe Riccardi, Giulio Jacucci |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2024 | Will LLMs Replace the Encoder-Only Models in Temporal Relation Classification?abstractThe automatic detection of temporal relations among events has been mainly investigated with encoder-only models such as RoBERTa.Large Language Models (LLM) have recently shown promising performance in temporal reasoning tasks such as temporal question answering.Nevertheless, recent studies have tested the LLMs' performance in detecting temporal relations of closed-source models only, limiting the interpretability of those results.In this work, we investigate LLMs' performance and decision process in the Temporal Relation Classification task.First, we assess the performance of seven open and closed-sourced LLMs experimenting with in-context learning and lightweight fine-tuning approaches.Results show that LLMs with in-context learning significantly underperform smaller encoderonly models based on RoBERTa.Then, we delve into the possible reasons for this gap by applying explainable methods.The outcome suggests a limitation of LLMs in this task due to their autoregressive nature, which causes them to focus only on the last part of the sequence.Additionally, we evaluate the word embeddings of these two models to better understand their pre-training differences.The code and the fine-tuned models can be found respectively on GitHub 1 . Gabriel Roccabruna, Massimo Rizzoli, Giuseppe Riccardi |
EMNLP | 3 |
| 2024 | Should We Fine-Tune or RAG? Evaluating Different Techniques to Adapt LLMs for DialogueabstractWe study the limitations of Large Language Models (LLMs) for the task of response generation in human-machine dialogue.Several techniques have been proposed in the literature for different dialogue types (e.g., Open-Domain).However, the evaluations of these techniques have been limited in terms of base LLMs, dialogue types and evaluation metrics.In this work, we extensively analyze different LLM adaptation techniques when applied to different dialogue types.We have selected two base LLMs, Llama2 C and Mistral I , and four dialogue types Open-Domain, Knowledge-Grounded, Task-Oriented, and Question Answering.We evaluate the performance of incontext learning and fine-tuning techniques across datasets selected for each dialogue type.We assess the impact of incorporating external knowledge to ground the generation in both scenarios of Retrieval-Augmented Generation (RAG) and gold knowledge.We adopt consistent evaluation and explainability criteria for automatic metrics and human evaluation protocols.Our analysis shows that there is no universal best-technique for adapting large language models as the efficacy of each technique depends on both the base LLM and the specific type of dialogue.Last but not least, the assessment of the best adaptation technique should include human evaluation to avoid false expectations and outcomes derived from automatic metrics. Simone Alghisi, Massimo Rizzoli, Gabriel Roccabruna, Seyed Mahed Mousavi, Giuseppe Riccardi |
INLG | 5 |
| 2023 | Let's Give a Voice to Conversational Agents in Virtual Reality
Michele Yin, Gabriel Roccabruna, Abhinav Azad, Giuseppe Riccardi |
INTERSPEECH | 4 |
| 2023 | Evaluation of interpretability for deep learning algorithms in EEG emotion recognition: A case study in autism
Juan Manuel Mayor Torres, Sara E. Medina-DeVilliers, Tessa Clarkson, Matthew D. Lerner, Giuseppe Riccardi |
Artif. Intell. Medicine | 5 |
| 2022 | What can Speech and Language Tell us About the Working Alliance in PsychotherapyabstractWe are interested in the problem of conversational analysis and its application to the health domain. Cognitive Behavioral Therapy is a structured approach in psychotherapy, allowing the therapist to help the patient to identify and modify the malicious thoughts, behavior, or actions. This cooperative effort can be evaluated using the Working Alliance Inventory Observer-rated Shortened – a 12 items inventory covering task, goal, and relationship – which has a relevant influence on therapeutic outcomes. In this work, we investigate the relation between this alliance inventory and the spoken conversations (sessions) between the patient and the psychotherapist. We have delivered eight weeks of e-therapy, collected their audio and video call sessions, and manually transcribed them. The spoken conversations have been annotated and evaluated with WAI ratings by professional therapists. We have investigated speech and language features and their association with WAI items. The feature types include turn dynamics, lexical entrainment, and conversational descriptors extracted from the speech and language signals. Our findings provide strong evidence that a subset of these features are strong indicators of working alliance. To the best of our knowledge, this is the first and a novel study to exploit speech and language for characterising working alliance Sebastian P. Bayerl, Gabriel Roccabruna, Shammur Absar Chowdhury, Tommaso Ciulli, Morena Danieli, Korbinian Riedhammer, Giuseppe Riccardi |
INTERSPEECH | 7 |
| 2022 | Multi-source Multi-domain Sentiment Analysis with BERT-based ModelsabstractSentiment analysis is one of the most widely studied tasks in natural language processing. While BERT-based models have achieved state-of-the-art results in this task, little attention has been given to its performance variability across class labels, multi-source and multi-domain corpora. In this paper, we present an improved state-of-the-art and comparatively evaluate BERT-based models for sentiment analysis on Italian corpora. The proposed model is evaluated over eight sentiment analysis corpora from different domains (social media, finance, e-commerce, health, travel) and sources (Twitter, YouTube, Facebook, Amazon, Tripadvisor, Opera and Personal Healthcare Agent) on the prediction of positive, negative and neutral classes. Our findings suggest that BERT-based models are confident in predicting positive and negative examples but not as much with neutral examples. We release the sentiment analysis model as well as a newly financial domain sentiment corpus. Gabriel Roccabruna, Steve Azzolin, Giuseppe Riccardi |
LREC | 3 |
| 2022 | Annotation of Valence Unfolding in Spoken Personal NarrativesabstractPersonal Narrative (PN) is the recollection of individuals’ life experiences, events, and thoughts along with the associated emotions in the form of a story. Compared to other genres such as social media texts or microblogs, where people write about experienced events or products, the spoken PNs are complex to analyze and understand. They are usually long and unstructured, involving multiple and related events, characters as well as thoughts and emotions associated with events, objects, and persons. In spoken PNs, emotions are conveyed by changing the speech signal characteristics as well as the lexical content of the narrative. In this work, we annotate a corpus of spoken personal narratives, with the emotion valence using discrete values. The PNs are segmented into speech segments, and the annotators annotate them in the discourse context, with values on a 5-point bipolar scale ranging from -2 to +2 (0 for neutral). In this way, we capture the unfolding of the PNs events and changes in the emotional state of the narrator. We perform an in-depth analysis of the inter-annotator agreement, the relation between the label distribution w.r.t. the stimulus (positive/negative) used for the elicitation of the narrative, and compare the segment-level annotations to a baseline continuous annotation. We find that the neutral score plays an important role in the agreement. We observe that it is easy to differentiate the positive from the negative valence while the confusion with the neutral label is high. Keywords: Personal Narratives, Emotion Annotation, Segment Level Annotation Aniruddha Tammewar, Franziska Braun, Gabriel Roccabruna, Sebastian P. Bayerl, Korbinian Riedhammer, Giuseppe Riccardi |
LREC | 6 |
| 2021 | Detecting Emotion Carriers by Combining Acoustic and Lexical RepresentationsabstractPersonal narratives (PN) – spoken or written – are recollections of facts, people, events, and thoughts from one's own experience. Emotion recognition and sentiment analysis tasks are usually defined at the utterance or document level. However, in this work, we focus on Emotion Carriers (EC) defined as the segments (speech or text) that best explain the emotional state of the narrator (“loss of father”, “made me choose”). Once extracted, such EC can provide a richer representation of the user state to improve natural language understanding and dialogue modeling. In previous work, it has been shown that EC can be identified using lexical features. However, spoken narratives should provide a richer description of the context and the users' emotional state. In this paper, we leverage word-based acoustic and textual embeddings as well as early and late fusion techniques for the detection of ECs in spoken narratives. For the acoustic word-level representations, we use Residual Neural Networks (ResNet) pretrained on separate speech emotion corpora and fine-tuned to detect EC. Experiments with different fusion and system combination strategies show that late fusion leads to significant improvements for this task. Sebastian P. Bayerl, Aniruddha Tammewar, Korbinian Riedhammer, Giuseppe Riccardi |
ASRU | 4 |
| 2021 | Emotion Carrier Recognition from Personal NarrativesabstractPersonal Narratives (PN) - recollections of facts, events, and thoughts from one's own experience - are often used in everyday conversations. So far, PNs have mainly been explored for tasks such as valence prediction or emotion classification (e.g. happy, sad). However, these tasks might overlook more fine-grained information that could prove to be relevant for understanding PNs. In this work, we propose a novel task for Narrative Understanding: Emotion Carrier Recognition (ECR). Emotion carriers, the text fragments that carry the emotions of the narrator (e.g. loss of a grandpa, high school reunion), provide a fine-grained description of the emotion state. We explore the task of ECR in a corpus of PNs manually annotated with emotion carriers and investigate different machine learning models for the task. We propose evaluation strategies for ECR including metrics that can be appropriate for different tasks. Aniruddha Tammewar, Alessandra Cervone, Giuseppe Riccardi |
Interspeech | 3 |
| 2020 | Annotation of Emotion Carriers in Personal NarrativesabstractWe are interested in the problem of understanding personal narratives (PN) - spoken or written - recollections of facts, events, and thoughts. For PNs, we define emotion carriers as the speech or text segments that best explain the emotional state of the narrator. Such segments may span from single to multiple words, containing for example verb or noun phrases. Advanced automatic understanding of PNs requires not only the prediction of the narrator’s emotional state but also to identify which events (e.g. the loss of a relative or the visit of grandpa) or people (e.g. the old group of high school mates) carry the emotion manifested during the personal recollection. This work proposes and evaluates an annotation model for identifying emotion carriers in spoken personal narratives. Compared to other text genres such as news and microblogs, spoken PNs are particularly challenging because a narrative is usually unstructured, involving multiple sub-events and characters as well as thoughts and associated emotions perceived by the narrator. In this work, we experiment with annotating emotion carriers in speech transcriptions from the Ulm State-of-Mind in Speech (USoMS) corpus, a dataset of PNs in German. We believe this resource could be used for experiments in the automatic extraction of emotion carriers from PN, a task that could provide further advancements in narrative understanding. Aniruddha Tammewar, Alessandra Cervone, Eva-Maria Messner, Giuseppe Riccardi |
LREC | 4 |
| 2020 | Is this Dialogue Coherent? Learning from Dialogue Acts and EntitiesabstractIn this work, we investigate the human perception of coherence in open-domain dialogues.In particular, we address the problem of annotating and modeling the coherence of nextturn candidates while considering the entire history of the dialogue.First, we create the Switchboard Coherence (SWBD-Coh) corpus, a dataset of human-human spoken dialogues annotated with turn coherence ratings, where next-turn candidate utterances ratings are provided considering the full dialogue context.Our statistical analysis of the corpus indicates how turn coherence perception is affected by patterns of distribution of entities previously introduced and the Dialogue Acts used.Second, we experiment with different architectures to model entities, Dialogue Acts and their combination and evaluate their performance in predicting human coherence ratings on SWBD-Coh.We find that models combining both DA and entity information yield the best performances both for response selection and turn coherence rating. Alessandra Cervone, Giuseppe Riccardi |
SIGdial | 2 |
| 2019 | An Incremental Turn-Taking Model for Task-Oriented Dialog SystemsabstractIn a human-machine dialog scenario, deciding the appropriate time for the machine to take the turn is an open research problem. In contrast, humans engaged in conversations are able to timely decide when to interrupt the speaker for competitive or non-competitive reasons. In state-of-the-art turn-by-turn dialog systems the decision on the next dialog action is taken at the end of the utterance. In this paper, we propose a token-by-token prediction of the dialog state from incremental transcriptions of the user utterance. To identify the point of maximal understanding in an ongoing utterance, we a) implement an incremental Dialog State Tracker which is updated on a token basis (iDST) b) re-label the Dialog State Tracking Challenge 2 (DSTC2) dataset and c) adapt it to the incremental turn-taking experimental scenario. The re-labeling consists of assigning a binary value to each token in the user utterance that allows to identify the appropriate point for taking the turn. Finally, we implement an incremental Turn Taking Decider (iTTD) that is trained on these new labels for the turn-taking decision. We show that the proposed model can achieve a better performance compared to a deterministic handcrafted turn-taking algorithm. Andrei Catalin, Koichiro Yoshino, Yukitoshi Murase, Satoshi Nakamura 0001, Giuseppe Riccardi |
INTERSPEECH | 5 |
| 2019 | Active Annotation: Bootstrapping Annotation Lexicon and Guidelines for Supervised NLU LearningabstractNatural Language Understanding (NLU) models are typically trained in a supervised learning framework. In the case of intent classification, the predicted labels are predefined and based on the designed annotation schema while the labelling process is based on a laborious task where annotators manually inspect each utterance and assign the corresponding label. We propose an Active Annotation (AA) approach where we combine an unsupervised learning method in the embedding space, a human-in-the-loop verification process, and linguistic insights to create lexicons that can be open categories and adapted over time. In particular, annotators define the y-label space on-the-fly during the annotation using an iterative process and without the need for prior knowledge about the input data. We evaluate the proposed annotation paradigm in a real use-case NLU scenario. Results show that our Active Annotation paradigm achieves accurate and higher quality training data, with an annotation speed of an order of magnitude higher with respect to the traditional human-only driven baseline annotation methodology. Federico Marinelli, Alessandra Cervone, Giuliano Tortoreto, Evgeny A. Stepanov, Giuseppe Di Fabbrizio, Giuseppe Riccardi |
INTERSPEECH | 6 |
| 2019 | Modeling User Context for Valence Prediction from NarrativesabstractAutomated prediction of valence, one key feature of a person's emotional state, from individuals' personal narratives may provide crucial information for mental healthcare (e.g. early diagnosis of mental diseases, supervision of disease course, etc.). In the Interspeech 2018 ComParE Self-Assessed Affect challenge, the task of valence prediction was framed as a three-class classification problem using 8 seconds fragments from individuals' narratives. As such, the task did not allow for exploring contextual information of the narratives. In this work, we investigate the intrinsic information from multiple narratives recounted by the same individual in order to predict their current state-of-mind. Furthermore, with generalizability in mind, we decided to focus our experiments exclusively on textual information as the public availability of audio narratives is limited compared to text. Our hypothesis is that context modeling might provide insights about emotion triggering concepts (e.g. events, people, places) mentioned in the narratives that are linked to an individual's state of mind. We explore multiple machine learning techniques to model narratives. We find that the models are able to capture inter-individual differences, leading to more accurate predictions of an individual's emotional state, as compared to single narratives. Aniruddha Tammewar, Alessandra Cervone, Eva-Maria Messner, Giuseppe Riccardi |
INTERSPEECH | 4 |
| 2019 | Automatic classification of speech overlaps: Feature representation and algorithmsabstractOverlapping speech is a natural and frequently occurring phenomenon in human–human conversations with an underlying purpose. Speech overlap events may be categorized as competitive and non-competitive. While the former is an attempt to grab the floor, the latter is an attempt to assist the speaker to continue the turn. The presence and distribution of these categories are indicative of the speakers’ states during the conversation. Therefore, understanding these manifestations is crucial for conversational analysis and for modeling human–machine dialogs. The goal of this study is to design computational models to classify overlapping speech segments of dyadic conversations into competitive vs. non-competitive acts using lexical and acoustic cues, as well as their surrounding context. The designed overlap representations are evaluated in both linear – Support Vector Machines (SVM) – and non-linear – feed-forward (FFNN), convolutional (CNN) and long short-term memory (LSTM) neural network – models. We experiment with lexical and acoustic representations and their combinations from both speaker channels in feature and hidden space. We observe that lexical word-embedding features significantly increase the overall F 1 -measure compared to both acoustic and bag-of-ngrams lexical representations , suggesting that lexical information can be utilized as a powerful cue for overlap classification. Our comparative study shows that the best computational architecture is an FFNN along with a combination of word embeddings and acoustic features. Shammur Absar Chowdhury, Evgeny A. Stepanov, Morena Danieli, Giuseppe Riccardi |
Comput. Speech Lang. | 4 |
| 2018 | ISO-Standard Domain-Independent Dialogue Act Tagging for Conversational AgentsabstractDialogue Act (DA) tagging is crucial for spoken language understanding systems, as it provides a general representation of speakers’ intents, not bound to a particular dialogue system. Unfortunately, publicly available data sets with DA annotation are all based on different annotation schemes and thus incompatible with each other. Moreover, their schemes often do not cover all aspects necessary for open-domain human-machine interaction. In this paper, we propose a methodology to map several publicly available corpora to a subset of the ISO standard, in order to create a large task-independent training corpus for DA classification. We show the feasibility of using this corpus to train a domain-independent DA tagger testing it on out-of-domain conversational data, and argue the importance of training on multiple corpora to achieve robustness across different DA categories. Stefano Mezza, Alessandra Cervone, Evgeny A. Stepanov, Giuliano Tortoreto, Giuseppe Riccardi |
COLING | 5 |
| 2018 | HEAL: A Health Analytics Intelligent Agent Platform for the acquisition and analysis of physiological signalsabstractEffectively caring for patients suffering from chronic diseases is a challenging and costly task. There is a need to decrease the cost of treating these chronic patients, while increasing the quality of the care provided to the patients. Mobile health (mHealth) technologies such as remote monitoring, telemedicine, and home-based monitoring have become effective tools shifting the focus towards a more patient-centric healthcare. In this paper we present HEAL, an intelligent healthcare analytics and personal agent platform. The HEAL personal agent enables on-the-go continuous monitoring of covert and overt signals from patients. The HEAL agent consists of a pipeline that aggregates signals from devices and sensor types. The platform collects motion profiles, and user annotation for performing activity recognition and stress level detection. We evaluate the HEAL platform on patients suffering from essential hypertension and show that an intelligent personal mobile agent is capable of monitoring, analyzing and tracking the characteristics of hypertension management. Arindam Ghosh 0004, Evgeny A. Stepanov, Juan Manuel Mayor Torres, Morena Danieli, Giuseppe Riccardi |
HealthCom | 5 |
| 2018 | Depression Severity Estimation from Multiple ModalitiesabstractDepression is a major debilitating disorder which can affect people from all ages. With a continuous increase in the number of annual cases of depression, there is a need to develop automatic techniques for the detection of the presence and its severity. We explore different modalities (speech, behavioral characteristics, language and visual features extracted from face) to design and develop automatic methods for the detection of depression. In psychology literature, the eight-item Patient Health Questionnaire depression scale (PHQ-8) is well established as a tool for measuring the severity of depression. In this paper we aim to automatically predict the total sum of PHQ-8 scores from features extracted from the different modalities. We demonstrate that among the considered modalities, behavioral characteristic features extracted from speech yield the lowest MAE, outperforming the best system at the Audio/Visual Emotion Challenge (AVEC) 2017 depression sub-challenge. Evgeny A. Stepanov, Stéphane Lathuilière, Shammur Absar Chowdhury, Arindam Ghosh 0004, Radu L. Vieriu, Nicu Sebe, Giuseppe Riccardi |
HealthCom | 7 |
| 2018 | Coherence Models for DialogueabstractCoherence across multiple turns is a major challenge for state-of-the-art dialogue models. Arguably the most successful approach to automatically learning text coherence is the entity grid, which relies on modelling patterns of distribution of entities across multiple sentences of a text. Originally applied to the evaluation of automatic summaries and the news genre, among its many extensions, this model has also been successfully used to assess dialogue coherence. Nevertheless, both the original grid and its extensions do not model intents, a crucial aspect that has been studied widely in the literature in connection to dialogue structure. We propose to augment the original grid document representation for dialogue with the intentional structure of the conversation. Our models outperform the original grid representation on both text discrimination and insertion, the two main standard tasks for coherence assessment across three different dialogue datasets, confirming that intents play a key role in modelling dialogue coherence. Alessandra Cervone, Evgeny A. Stepanov, Giuseppe Riccardi |
INTERSPEECH | 3 |
| 2018 | Annotating and modeling empathy in spoken conversations
Firoj Alam, Morena Danieli, Giuseppe Riccardi |
Comput. Speech Lang. | 3 |
| 2017 | A Deep Learning approach to modeling competitiveness in spoken conversationsabstractThe motivation behind the research on overlapping speech has always been dominated by the need to model human-machine interaction for dialog systems and conversation analysis. To have more complex insights of the interlocutors' intentions behind the interaction, we need to understand the type of overlaps. Overlapping speech signals the interlocutor's intention to grab the floor. This act could be a competitive or non-competitive act, which either signals a problem or indicates assistance in communication. In this paper, we present a Deep Learning approach to modeling competitiveness in overlapping speech using acoustic and lexical features and their combination. We compare a fully-connected feed-forward neural network to the Support Vector Machine (SVM) models on real call center human-human conversations. We have observed that feature combination with DNN (significantly) outperforms SVM models, both the individual feature baselines and the feature combination model by 4% and 2% respectively. Shammur Absar Chowdhury, Giuseppe Riccardi |
ICASSP | 2 |
| 2017 | A cross-modal adaptation approach for brain decodingabstractBrain decoding has become a hot topic in many recent brain studies. In a typical neuroimaging experiment, participants are presented with different categories of stimuli while their concurrent brain activity is recorded. Then a classifier is trained on the features extracted from the recorded brain data to discriminate different target stimuli classes. It is a common practice to hypothesize that the stimulus-related information exists in the brain data if the decoder can accurately predict the target stimulus category. However, most of the neuroimaging studies suffer from few and noisy samples. These constraints affects the performance of such decoding systems. In order to cope with this limitation, a dictionary learning approach is used in this paper to transfer knowledge from the multimedia domain to the brain domain. We show that such cross-modal domain adaptation yields better performance of the learning algorithm in the brain domain. This is the first study in the direction of cross-modal adaptation by joint dictionary learning on multimedia and brain modality. Pouya Ghaemmaghami, Moin Nabi, Yan Yan 0002, Giuseppe Riccardi, Nicu Sebe |
ICASSP | 4 |
| 2017 | Towards End-to-End Spoken Dialogue Systems with Turn Embeddings
Ali Orkan Bayer, Evgeny A. Stepanov, Giuseppe Riccardi |
INTERSPEECH | 3 |
| 2016 | How Interlocutors Coordinate with each other within Emotional Segments?abstractIn this paper, we aim to investigate the coordination of interlocutors behavior in different emotional segments. Conversational coordination between the interlocutors is the tendency of speakers to predict and adjust each other accordingly on an ongoing conversation. In order to find such a coordination, we investigated 1) lexical similarities between the speakers in each emotional segments, 2) correlation between the interlocutors using psycholinguistic features, such as linguistic styles, psychological process, personal concerns among others, and 3) relation of interlocutors turn-taking behaviors such as competitiveness. To study the degree of coordination in different emotional segments, we conducted our experiments using real dyadic conversations collected from call centers in which agent’s emotional state include empathy and customer’s emotional states include anger and frustration. Our findings suggest that the most coordination occurs between the interlocutors inside anger segments, where as, a little coordination was observed when the agent was empathic, even though an increase in the amount of non-competitive overlaps was observed. We found no significant difference between anger and frustration segment in terms of turn-taking behaviors. However, the length of pause significantly decreases in the preceding segment of anger where as it increases in the preceding segment of frustration. Firoj Alam, Shammur Absar Chowdhury, Morena Danieli, Giuseppe Riccardi |
COLING | 4 |
| 2016 | Predicting Student Progress from Peer-Assessment Data
Michael Mogessie Ashenafi, Marco Ronchetti, Giuseppe Riccardi |
EDM | 3 |
| 2016 | Discourse connective detection in spoken conversationsabstractDiscourse parsing is an important task in Language Understanding with applications to human-human and human-machine communication modeling. However, most of the research has focused on written text, and parsers heavily rely on syntactic parsers that themselves have low performance on dialog data. In our work, we address the problem of analyzing the semantic relations between discourse units in human-human spoken conversations. In particular, in this paper we focus on the detection of discourse connectives which are the predicate of such relations. The discourse relations are drawn from the Penn Discourse Treebank annotation model and adapted to a domain-specific Italian human-human spoken conversations. We study the relevance of lexical and acoustic context in predicting discourse connectives. We observe that both lexical and acoustic context have mixed effect on the prediction of specific connectives. While the oracle of using lexical and acoustic contextual feature combinations is F1 = 68.53, the lexical context alone significantly outperforms the baseline by more than 10 points with F1= 64.93. Giuseppe Riccardi, Evgeny A. Stepanov, Shammur Absar Chowdhury |
ICASSP | 1 |
| 2016 | Predicting User Satisfaction from Turn-Taking in Spoken Conversations
Shammur Absar Chowdhury, Evgeny A. Stepanov, Giuseppe Riccardi |
INTERSPEECH | 3 |
| 2016 | Multilevel Annotation of Agreement and Disagreement in Italian News Blogs
Fabio Celli, Giuseppe Riccardi, Firoj Alam |
LREC | 2 |
| 2016 | Transfer of Corpus-Specific Dialogue Act Annotation to ISO Standard: Is it worth it?
Shammur Absar Chowdhury, Evgeny A. Stepanov, Giuseppe Riccardi |
LREC | 3 |
| 2016 | Summarizing Behaviours: An Experiment on the Annotation of Call-Centre Conversations
Morena Danieli, A. R. Balamurali, Evgeny A. Stepanov, Benoît Favre, Frédéric Béchet, Giuseppe Riccardi |
LREC | 6 |
| 2016 | Semantic language models with deep neural networks
Ali Orkan Bayer, Giuseppe Riccardi |
Comput. Speech Lang. | 2 |
| 2016 | In the mood for sharing contents: Emotions, personality and interaction styles in the diffusion of news
Fabio Celli, Arindam Ghosh 0004, Firoj Alam, Giuseppe Riccardi |
Inf. Process. Manag. | 4 |
| 2015 | Predicting students' final exam scores from their course activitiesabstractA common approach to the problem of predicting students' exam scores has been to base this prediction on the previous educational history of students. In this paper, we present a model that bases this prediction on students' performance on several tasks assigned throughout the duration of the course. In order to build our prediction model, we use data from a semi-automated peer-assessment system implemented in two undergraduate-level computer science courses, where students ask questions about topics discussed in class, answer questions from their peers, and rate answers provided by their peers. We then construct features that are used to build several multiple linear regression models. We use the Root Mean Squared Error (RMSE) of the prediction models to evaluate their performance. Our final model, which has recorded an RMSE of 2.93 for one course and 3.44 for another on predicting grades on a scale of 18 to 30, is built using 14 features that capture various activities of students. Our work has possible implications in the MOOC arena and in similar online course administration systems. Michael Mogessie Ashenafi, Giuseppe Riccardi, Marco Ronchetti |
FIE | 2 |
| 2015 | Annotating and categorizing competition in overlap speechabstractOverlapping speech is a common and relevant phenomenon in human conversations, reflecting many aspects of discourse dynamics. In this paper, we focus on the pragmatic role of overlaps in turn-in-progress, where it can be categorized as competitive or non-competitive. Previous studies on these two categories have mostly relied on controlled scenarios and small datasets. In our study, we focus on call center data, with customers and operators engaged in problem-solving tasks. We propose and evaluate an annotation scheme for these two overlap categories in the context of spontaneous and in-vivo human conversations. We analyze the distinctive predictive characteristics of a very large set of high-dimensional acoustic feature. We obtained a significant improvement in classification results as well as significant reduction in the feature set size. Shammur Absar Chowdhury, Morena Danieli, Giuseppe Riccardi |
ICASSP | 3 |
| 2015 | Deep semantic encodings for language modelingabstractWord error rate (WER) is not an appropriate metric for spoken language systems (SLS) because lower WER does not necessarily yield better understanding performance. Therefore, language models (LMs) that are used in SLS should be trained to jointly optimize transcription and understanding performance. Semantic LMs (SELMs) are based on the theory of frame semantics and incorporate features of frames and meaning bearing words (target words) as semantic context when training LMs. The performance of SELMs is affected by the errors on the ASR and the semantic parser output. In this paper we address the problem of coping with such noise in the training phase of the neural network-based architecture of LMs. We propose the use of deep autoencoders for the encoding of semantic context while accounting for ASR errors. We investigate the optimization of SELMs both for transcription and understanding by using deep semantic encodings. Deep semantic encodings suppress the noise introduced by the ASR module, and enable SELMs to be optimized adequately. We assess the understanding performance by measuring the errors made on target words and we achieve 3.7% relative improvement over recurrent neural network LMs. Ali Orkan Bayer, Giuseppe Riccardi |
INTERSPEECH | 2 |
| 2015 | Selection and aggregation techniques for crowdsourced semantic annotation task
Shammur Absar Chowdhury, Marcos Calvo, Arindam Ghosh 0004, Evgeny A. Stepanov, Ali Orkan Bayer, Giuseppe Riccardi, Fernando García 0001, Emilio Sanchis Arnal |
INTERSPEECH | 6 |
| 2015 | The role of speakers and context in classifying competition in overlapping speechabstractOverlapping speech is one of the most frequently occurring events in the course of human-human conversations.Understanding the dynamics of overlapping speech is crucial for conversational analysis and for modeling human-machine dialog.Overlapping speech may signal the speaker's intention to grab the floor with a competitive vs non-competitive act.In this paper, we study the role of speakers, whether they initiate (overlapper) or not (overlappee) the overlap, and the context of the event.The speech overlap may be explained and predicted by the dialog context, the linguistic or acoustic descriptors.Our goal is to understand whether the competitiveness of the overlap is best predicted by the overlapper, the overlappee, the context or by their combinations.For each overlap and its context we have extracted acoustic, linguistic, and psycholinguistic features and combined decisions from the best classification models.The evaluation of the classifier has been carried out over call center human-human conversations.The results show that the complete knowledge of speakers' role and context highly contribute to the classification results when using acoustic and psycholinguistic features.Our findings also suggest that the lexical selections of the overlapper are good indicators of speaker's competitive or non-competitive intentions. Shammur Absar Chowdhury, Morena Danieli, Giuseppe Riccardi |
INTERSPEECH | 3 |
| 2015 | Call Centre Conversation Summarization: A Pilot Task at Multiling 2015abstractThis paper describes the results of the Call Centre Conversation Summarization task at Multiling'15.The CCCS task consists in generating abstractive synopses from call centre conversations between a caller and an agent.Synopses are summaries of the problem of the caller, and how it is solved by the agent.Generating them is a very challenging task given that deep analysis of the dialogs and text generation are necessary.Three languages were addressed: French, Italian and English translations of conversations from those two languages.The official evaluation metric was ROUGE-2.Two participants submitted a total of four systems which had trouble beating the extractive baselines.The datasets released for the task will allow more research on abstractive dialog summarization. Benoît Favre, Evgeny A. Stepanov, Jérémy Trione, Frédéric Béchet, Giuseppe Riccardi |
SIGDIAL Conference | 5 |
| 2014 | Fusion of acoustic, linguistic and psycholinguistic features for Speaker Personality Traits recognitionabstractBehavioral analytics is an emerging research area that aims at automatic understanding of human behavior. For the advancement of this research area, we are interested in the problem of learning the personality traits from spoken data. In this study, we investigated the contribution of different types of speech features to the automatic recognition of Speaker Personality Trait (SPT) across diverse speech corpora (broadcast news and spoken conversation). We have extracted acoustic, linguistic, and psycholinguistic features and modeled their combination as input to the classification task. For the classification, we used Sequential Minimal Optimization for Support Vector Machine (SMO) together with Relief feature selection. The present study shows different levels of performance for automatically selected feature sets, and overall improved performance with their combination across diverse corpora. Firoj Alam, Giuseppe Riccardi |
ICASSP | 2 |
| 2014 | Cross-language transfer of semantic annotation via targeted crowdsourcing
Shammur Absar Chowdhury, Arindam Ghosh 0004, Evgeny A. Stepanov, Ali Orkan Bayer, Giuseppe Riccardi, Ioannis Klasinas |
INTERSPEECH | 5 |
| 2014 | The Development of the Multilingual LUNA Corpus for Spoken Language System Porting
Evgeny A. Stepanov, Giuseppe Riccardi, Ali Orkan Bayer |
LREC | 2 |
| 2014 | The Workshop on Computational Personality Recognition 2014abstractThe Workshop on Computational Personality Recognition aims to define the state-of-the-art in the field and to provide tools for future standard evaluations in personality recognition tasks. In the WCPR14 we released two different datasets: one of Youtube Vlogs and one of Mobile Phone interactions. We structured the workshop in two tracks: an open shared task, where participants can do any kind of experiment, and a competition. We also distinguished two tasks: A) personality recognition from multimedia data, and B) personality recognition from text only. In this paper we discuss the results of the workshop. Fabio Celli, Bruno Lepri, Joan-Isaac Biel, Daniel Gatica-Perez, Giuseppe Riccardi, Fabio Pianesi |
ACM Multimedia | 5 |
| 2014 | Recognizing Human Activities from Smartphone Sensor SignalsabstractIn context-aware computing, Human Activity Recognition (HAR) aims to understand the current activity of users from their connected sensors. Smartphones with their various sensors are opening a new frontier in building human-centered applications for understanding users' personal and world contexts. While in-lab and controlled activity recognition systems have yielded very good results, they do not perform well under in-the-wild scenarios. The objective of this paper is to 1) Investigate how audio signal can complement and improve other on-board sensors (accelerometer and gyroscope) for activity recognition; 2) Design and evaluate the fusion of such multiple signal streams to optimize performance and sampling rate. We show that fusion of these signal streams, including audio, achieves high performance even at very low sampling rates; 3) Evaluate the performance of the multi-stream human activity recognition under different real end-user activity conditions. Arindam Ghosh 0004, Giuseppe Riccardi |
ACM Multimedia | 2 |
| 2014 | Semantic language models for Automatic Speech RecognitionabstractWe are interested in the problem of semantics-aware training of language models (LMs) for Automatic Speech Recognition (ASR). Traditional language modeling research have ignored semantic constraints and focused on limited size histories of words. Semantic structures may provide information to capture lexically realized long-range dependencies as well as the linguistic scene of a speech utterance. In this paper, we present a novel semantic LM(SELM) that is based on the theory of frame semantics. Frame semantics analyzes meaning of words by considering their role in the semantic frames they occur and by considering their syntactic properties. We show that by integrating semantic frames and target words into recurrent neural network LMs we can gain significant improvements in perplexity and word error rates. We have evaluated the semantic LM on the publicly available ASR baselines on the Wall Street Journal (WSJ) corpus. SELMs achieve 50% and 64% relative reduction in perplexity compared to n-gram models by using frames and target words respectively. In addition, 12% and 7% relative improvements in word error rates are achieved by SELMs on the Nov'92 and Nov'93 test sets with respect to the baseline tri-gram LM. Ali Orkan Bayer, Giuseppe Riccardi |
SLT | 2 |
| 2014 | A domain-independent statistical methodology for dialog management in spoken dialog systems
David Griol, Zoraida Callejas Carrión, Ramón López-Cózar, Giuseppe Riccardi |
Comput. Speech Lang. | 4 |
| 2013 | On-line adaptation of semantic models for spoken language understandingabstractSpoken language understanding (SLU) systems extract semantic information from speech signals, which is usually mapped onto concept sequences. The distribution of concepts in dialogues are usually sparse. Therefore, general models may fail to model the concept distribution for a dialogue and semantic models can benefit from adaptation. In this paper, we present an instance-based approach for on-line adaptation of semantic models. We show that we can improve the performance of an SLU system on an utterance, by retrieving relevant instances from the training data and using them for on-line adapting the semantic models. The instance-based adaptation scheme uses two different similarity metrics edit distance and n-gram match score on three different to-kenizations; word-concept pairs, words, and concepts. We have achieved a significant improvement (6% relative) in the understanding performance by conducting rescoring experiments on the n-best lists that the SLU outputs. We have also applied a two-level adaptation scheme, where adaptation is first applied to the automatic speech recognizer (ASR) and then to the SLU. Ali Orkan Bayer, Giuseppe Riccardi |
ASRU | 2 |
| 2013 | Language style and domain adaptation for cross-language SLU portingabstractAutomatic cross-language Spoken Language Understanding porting is plagued by two limitations. First, SLU are usually trained on limited domain corpora. Second, language pair resources (e.g. aligned corpora) are scarce or unmatched in style (e.g. news vs. conversation). We present experiments on automatic style adaptation of the input for the translation systems and their output for SLU. We approach the problem of scarce aligned data by adapting the available parallel data to the target domain using limited in-domain and larger web crawled close-to-domain corpora. SLU performance is optimized by reranking its output with Recurrent Neural Network-based joint language model. We evaluate end-to-end SLU porting on close and distant language pairs: Spanish - Italian and Turkish - Italian; and achieve significant improvements both in translation quality and SLU performance. Evgeny A. Stepanov, Ilya Kashkarev, Ali Orkan Bayer, Giuseppe Riccardi, Arindam Ghosh 0004 |
ASRU | 4 |
| 2013 | Comparative study of speaker personality traits recognition in conversational and broadcast news speechabstractNatural human-computer interaction requires, in addition to understand what the speaker is saying, recognition of behavioral descriptors, such as speaker’s personality traits (SPTs). The complexity of this problem depends on the high variability and dimensionality of the acoustic, lexical and situational context manifestations of the SPTs. In this paper, we present a comparative study of automatic speaker personality trait recognition from speech corpora that differ in the source speaking style (broadcast news vs. conversational) and experimental context. We evaluated different feature selection algorithms such as information gain, relief and ensemble classification methods to address the high dimensionality issues. We trained and evaluated ensemble methods to leverage base learners, using three different algorithms such as SMO (Sequential Minimal Optimization for Support Vector Machine), RF (Random Forest) and Adaboost. After that, we combined them using majority voting and stacking methods. Our study shows that, performance of the system greatly benefits from feature selection and ensemble methods across corpora. Firoj Alam, Giuseppe Riccardi |
INTERSPEECH | 2 |
| 2013 | Instance-based on-line language model adaptationabstractTEST 02 - Elsevier's Scopus, the largest abstract and citation database of peer-reviewed literature. Search and access research from the science, technology, medicine, social sciences and arts and humanities fields. Ali Orkan Bayer, Giuseppe Riccardi |
INTERSPEECH | 2 |
| 2013 | Motivational feedback in crowdsourcing: a case study in speech transcription
Giuseppe Riccardi, Arindam Ghosh 0004, Shammur Absar Chowdhury, Ali Orkan Bayer |
INTERSPEECH | 1 |
| 2012 | Kolmogorov-Smirnov test for feature selection in emotion recognition from speechabstractAutomatic emotion recognition from speech is limited by the ability to discover the relevant predicting features. The common approach is to extract a very large set of features over a generally long analysis time window. In this paper we investigate the applicability of two-sample Kolmogorov-Smirnov statistical test (KST) to the problem of segmental speech emotion recognition. We train emotion classifiers for each speech segment within an utterance. The segment labels are then combined to predict the dominant emotion label. Our findings show that KST can be successfully used to extract statistically relevant features. KST criterion is used to optimize the parameters of the statistical segmental analysis, namely the window segment size and shift. We carry out seven binary class emotion classification experiments on the Emo-DB and evaluate the impact of the segmental analysis and emotion-specific feature selection. Alexei V. Ivanov, Giuseppe Riccardi |
ICASSP | 2 |
| 2012 | Improving the Recall of a Discourse Parser by Constraint-based Postprocessing
Sucheta Ghosh, Richard Johansson, Giuseppe Riccardi, Sara Tonelli |
LREC | 3 |
| 2012 | Global Features for Shallow Discourse Parsing
Sucheta Ghosh, Giuseppe Riccardi, Richard Johansson |
SIGDIAL Conference | 2 |
| 2012 | Joint language models for automatic speech recognition and understandingabstractLanguage models (LMs) are one of the main knowledge sources used by automatic speech recognition (ASR) and Spoken Language Understanding (SLU) systems. In ASR systems they are optimized to decode words from speech for a transcription task. In SLU systems they are optimized to map words into concept constructs or interpretation representations. Performance optimization is generally designed independently for ASR and SLU models in terms of word accuracy and concept accuracy respectively. However, the best word accuracy performance does not always yield the best understanding performance. In this paper we investigate how LMs originally trained to maximize word accuracy can be parametrized to account for speech understanding constraints and maximize concept accuracy. Incremental reduction in concept error rate is observed when a LM is trained on word-to-concept mappings. We show how to optimize the joint transcription and understanding task performance in the lexical-semantic relation space. Ali Orkan Bayer, Giuseppe Riccardi |
SLT | 2 |
| 2012 | Combining multiple translation systems for Spoken Language Understanding portabilityabstractWe are interested in the problem of learning Spoken Language Understanding (SLU) models for multiple target languages. Learning such models requires annotated corpora, and porting to different languages would require corpora with parallel text translation and semantic annotations. In this paper we investigate how to learn a SLU model in a target language starting from no target text and no semantic annotation. Our proposed algorithm is based on the idea of exploiting the diversity (with regard to performance and coverage) of multiple translation systems to transfer statistically stable word-to-concept mappings in the case of the romance language pair, French and Spanish. Each translation system performs differently at the lexical level (wrt BLEU). The best translation system performances for the semantic task are gained from their combination at different stages of the portability methodology. We have evaluated the portability algorithms on the French MEDIA corpus, using French as the source language and Spanish as the target language. The experiments show the effectiveness of the proposed methods with respect to the source language SLU baseline. Fernando García 0001, Lluís F. Hurtado, Encarna Segarra, Emilio Sanchis Arnal, Giuseppe Riccardi |
SLT | 5 |
| 2012 | Discriminative Reranking for Spoken Language UnderstandingabstractSpoken language understanding (SLU) is concerned with the extraction of meaning structures from spoken utterances. Recent computational approaches to SLU, e.g., conditional random fields (CRFs), optimize local models by encoding several features, mainly based on simple n-grams. In contrast, recent works have shown that the accuracy of CRF can be significantly improved by modeling long-distance dependency features. In this paper, we propose novel approaches to encode all possible dependencies between features and most importantly among parts of the meaning structure, e.g., concepts and their combination. We rerank hypotheses generated by local models, e.g., stochastic finite state transducers (SFSTs) or CRF, with a global model. The latter encodes a very large number of dependencies (in the form of trees or sequences) by applying kernel methods to the space of all meaning (sub) structures. We performed comparative experiments between SFST, CRF, support vector machines (SVMs), and our proposed discriminative reranking models (DRMs) on representative conversational speech corpora in three different languages: the ATIS (English), the MEDIA (French), and the LUNA (Italian) corpora. These corpora have been collected within three different domain applications of increasing complexity: informational, transactional, and problem-solving tasks, respectively. The results show that our DRMs consistently outperform the state-of-the-art models based on CRF. Marco Dinarelli, Alessandro Moschitti, Giuseppe Riccardi |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Using Syntactic and Semantic Structural Kernels for Classifying Definition Questions in Jeopardy!
Alessandro Moschitti, Jennifer Chu-Carroll, Siddharth Patwardhan, James Fan, Giuseppe Riccardi |
EMNLP | 5 |
| 2011 | Simultaneous dialog act segmentation and classification from human-human spoken conversationsabstractAn accurate identification dialog acts (DAs), which represent the illocutionary aspect of communication, is essential to support the understanding of human conversations. This requires (1) the segmentation of human-human dialogs into turns, (2) the intra-turn segmentation into DA boundaries and (3) the classification of each segment according to a DA tag. This process is particularly challenging when both segmentation and tagging are automated and utterance hypotheses derive from the erroneous results of ASR. In this paper, we use Conditional Random Fields to learn models for simultaneous segmentation and labeling of DAs from whole human-human spoken dialogs. We identify the best performing lexical feature combinations on the LUNA and SWITCHBOARD human-human dialog corpora and compare performances to those of discriminative D classifiers based on manually segmented utterances. Additionally, we assess our models' robustness to recognition errors, showing that DA identification is robust in the presence of high word error rates. Silvia Quarteroni, Alexei V. Ivanov, Giuseppe Riccardi |
ICASSP | 3 |
| 2011 | POMDP concept policies and task structures for hybrid dialog managementabstractWe address several challenges for applying statistical dialog managers based on Partially Observable Markov Models to real world problems: to deal with large numbers of concepts, we use individual POMDP policies for each concept. To control the use of the concept policies, the dialog manager uses explicit task structures. The POMDP policies model the confusability of concepts at the value level. In contrast to previous work, we use explicit confusability statistics including confidence scores based on real world data in the POMDP models. Since data sparseness becomes a key issue for estimating these probabilities, we introduce a form of smoothing the observation probabilities that maintains the overall concept error rate. We evaluated three POMDP-based dialog systems and a rule-based one in a phone-based user evaluation in a tourist domain. The results show that a POMDP that uses confidence scores, in combination with an improved SLU module, achieves the highest concept precision. Sebastian Varges, Giuseppe Riccardi, Silvia Quarteroni, Alexei V. Ivanov |
ICASSP | 2 |
| 2011 | Shallow Discourse Parsing with Conditional Random Fields
Sucheta Ghosh, Richard Johansson, Giuseppe Riccardi, Sara Tonelli |
IJCNLP | 3 |
| 2011 | Collecting Life Logs for Experience-Based Corpora
Fabiano Francesconi, Arindam Ghosh 0004, Giuseppe Riccardi, Marco Ronchetti, Alex Vagin |
INTERSPEECH | 3 |
| 2011 | Recognition of Personality Traits from Human Spoken ConversationsabstractWe are interested in understanding human personality and its manifestations in human interactions. The automatic analysis of such personality traits in natural conversation is quite complex due to the user-profiled corpora acquisition, annotation task and multidimensional modeling. While in the experimental psychology research this topic has been addressed extensively, speech and language scientists have recently engaged in limited experiments. In this paper we describe an automated system for speaker-independent personality prediction in the context of human-human spoken conversations. The evaluation of such system is carried out on the PersIA human-human spoken dialog corpus annotated with user self-assessments of the Big-Five personality traits. The personality predictor has been trained on paralinguistic features and its evaluation on five personality traits shows encouraging results for the conscientiousness and extroversion labels. Index Terms: Automated personality prediction from speech, human–human dialog analysis. Alexei V. Ivanov, Giuseppe Riccardi, Adam J. Sporka, Jakub Franc |
INTERSPEECH | 2 |
| 2011 | Comparing Stochastic Approaches to Spoken Language Understanding in Multiple LanguagesabstractOne of the first steps in building a spoken language understanding (SLU) module for dialogue systems is the extraction of flat concepts out of a given word sequence, usually provided by an automatic speech recognition (ASR) system. In this paper, six different modeling approaches are investigated to tackle the task of concept tagging. These methods include classical, well-known generative and discriminative methods like Finite State Transducers (FSTs), Statistical Machine Translation (SMT), Maximum Entropy Markov Models (MEMMs), or Support Vector Machines (SVMs) as well as techniques recently applied to natural language processing such as Conditional Random Fields (CRFs) or Dynamic Bayesian Networks (DBNs). Following a detailed description of the models, experimental and comparative results are presented on three corpora in different languages and with different complexity. The French MEDIA corpus has already been exploited during an evaluation campaign and so a direct comparison with existing benchmarks is possible. Recently collected Italian and Polish corpora are used to test the robustness and portability of the modeling approaches. For all tasks, manual transcriptions as well as ASR inputs are considered. Additionally to single systems, methods for system combination are investigated. The best performing model on all tasks is based on conditional random fields. On the MEDIA evaluation corpus, a concept error rate of 12.6% could be achieved. Here, additionally to attribute names, attribute values have been extracted using a combination of a rule-based and a statistical approach. Applying system combination using weighted ROVER with all six systems, the concept error rate (CER) drops to 12.0%. Stefan Hahn, Marco Dinarelli, Christian Raymond, Fabrice Lefèvre, Patrick Lehnen, Renato De Mori, Alessandro Moschitti, Hermann Ney, Giuseppe Riccardi |
IEEE Trans. Speech Audio Process. | 9 |
| 2010 | The LUNA Spoken Dialogue System: Beyond utterance classificationabstractWe present a call routing application for complex problem solving tasks. Up to date work on call routing has been mainly dealing with call-type classification. In this paper we take call routing further: Initial call classification is done in parallel with a robust statistical Spoken Language Understanding module. This is followed by a dialogue to elicit further task-relevant details from the user before passing on the call. The dialogue capability also allows us to obtain clarifications of the initial classifier guess. Based on an evaluation, we show that conducting a dialogue significantly improves upon call routing based on call classification alone. We present both subjective and objective evaluation results of the system according to standard metrics on real users. Marco Dinarelli, Evgeny A. Stepanov, Sebastian Varges, Giuseppe Riccardi |
ICASSP | 4 |
| 2010 | Automatic turn segmentation in spoken conversations
Alexei V. Ivanov, Giuseppe Riccardi |
INTERSPEECH | 2 |
| 2010 | Acoustic correlates of meaning structure in conversational speechabstractWe are interested in the problem of extracting meaning structures from spoken utterances in human communication. In Spoken Language Understanding (SLU) systems, parsing of meaning structures is carried over the word hypotheses generated by the Automatic Speech Recognizer (ASR). This approach suffers from high word error rates and ad-hoc conceptual representations. In contrast, in this paper we aim at discovering meaning components from direct measurements of acoustic and non-verbal linguistic features. The meaning structures are taken from the frame semantics model proposed in FrameNet, a consistent and extendable semantic structure resource covering a large set of domains. We give a quantitative analysis of meaning structures in terms of speech features across human–human dialogs from the manually annotated LUNA corpus. We show that the acoustic correlations between pitch, formant trajectories, intensity and harmonicity and meaning features are statistically significant over the whole corpus as well as relevant in classifying the target words evoked by a semantic frame. Alexei V. Ivanov, Giuseppe Riccardi, Sucheta Ghosh, Sara Tonelli, Evgeny A. Stepanov |
INTERSPEECH | 2 |
| 2010 | Combining user intention and error modeling for statistical dialog simulatorsabstractStatistical user simulation is an efficient and effective way to train and evaluate the performance of a (spoken) dialog system.\nIn this paper, we design and evaluate a modular data-driven dialog simulator where we decouple the “intentional” component\nof the User Simulator from the Error Simulator representing different types of ASR/SLU noisy channel distortion. While the\nformer is composed by a Dialog Act Model, a Concept Model and a User Model, the latter is centered around an Error Model. We test different Dialog Act Models and Error Models against a baseline dialog manager and compare results with real dialogs obtained using the same dialog manager. On the grounds of dialog act, task and concept accuracy, our results show that 1) datadriven\nDialog Act Models achieve good accuracy with respect to real user behavior and 2) data-driven Error Models make task\ncompletion times and rates closer to real data. Silvia Quarteroni, Meritxell González, Giuseppe Riccardi, Sebastian Varges |
INTERSPEECH | 3 |
| 2010 | Classifying dialog acts in human-human and human-machine spoken conversations
Silvia Quarteroni, Giuseppe Riccardi |
INTERSPEECH | 2 |
| 2010 | Annotation of Discourse Relations for Conversational Spoken Dialogs
Sara Tonelli, Giuseppe Riccardi, Rashmi Prasad, Aravind K. Joshi |
LREC | 2 |
| 2010 | Cooperative User Models in Statistical Dialog Simulators
Meritxell González, Silvia Quarteroni, Giuseppe Riccardi, Sebastian Varges |
SIGDIAL Conference | 3 |
| 2010 | Investigating Clarification Strategies in a Hybrid POMDP Dialog Manager
Sebastian Varges, Silvia Quarteroni, Giuseppe Riccardi, Alexei V. Ivanov |
SIGDIAL Conference | 3 |
| 2010 | Hypotheses selection for re-ranking semantic annotationsabstractDiscriminative reranking has been successfully used for several tasks of Natural Language Processing (NLP). Recently it has been applied also to Spoken Language Understanding, imrpoving state-of-the-art for some applications. However, such proposed models can be further improved by considering: (i) a better selection of the initial n-best hypotheses to be re-ranked and (ii) the use of a strategy that decides when the reranking model should be used, i.e. in some cases only the basic approach should be applied. In this paper, we apply a semantic inconsistency metric to select the n-best hypotheses from a large set generated by an SLU basic system. Then we apply a state-of-the-art re-ranker based on the Partial Tree Kernel (PTK), which encodes SLU hypotheses in Support Vector Machines (SVM) with complex structured features. Finally, we apply a decision model based on confidence values to select between the first hypothesis provided by the basic SLU model and the first hypothesis provided by the re-ranker. We show the effectiveness of our approach presenting comparative results obtained by reranking hypotheses generated by two very different models: a simple Stochastic Language Model encoded in Finite State Machines (FSM) and a Conditional Random Field (CRF) model. We evaluate our approach on the French MEDIA corpus and on an Italian corpus acquired in the European Project LUNA. The results show a significant improvement with respect to the current state-of-the-art and previous re-ranking models. Marco Dinarelli, Alessandro Moschitti, Giuseppe Riccardi |
SLT | 3 |
| 2009 | Ontology-based grounding of Spoken Language UnderstandingabstractCurrent Spoken Language Understanding models rely on either hand-written semantic grammars or flat attribute-value sequence labeling. In most cases, no relations between concepts are modeled, and both concepts and relations are domain-specific, making it difficult to expand or port the domain model. In contrast, we expand our previous work on a domain model based on an ontology where concepts follow the predicate-argument semantics and domain-independent classical relations are defined on such concepts. We conduct a thorough study on a spoken dialog corpus collected within a customer care problem-solving domain, and we evaluate the coverage and impact of the ontology for the interpretation, grounding and re-ranking of spoken language understanding interpretations. Silvia Quarteroni, Marco Dinarelli, Giuseppe Riccardi |
ASRU | 3 |
| 2009 | The exploration/exploitation trade-off in Reinforcement Learning for dialogue managementabstractConversational systems use deterministic rules that trigger actions such as requests for confirmation or clarification. More recently, reinforcement learning and (partially observable) Markov decision processes have been proposed for this task. In this paper, we investigate action selection strategies for dialogue management, in particular the exploration/exploitation trade-off and its impact on final reward (i.e. the session reward after optimization has ended) and lifetime reward (i.e. the overall reward accumulated over the learner's lifetime). We propose to use interleaved exploitation sessions as a learning methodology to assess the reward obtained from the current policy. The experiments show a statistically significant difference in final reward of exploitation-only sessions between a system that optimizes lifetime reward and one that maximizes the reward of the final policy. Sebastian Varges, Giuseppe Riccardi, Silvia Quarteroni, Alexei V. Ivanov |
ASRU | 2 |
| 2009 | Re-Ranking Models for Spoken Language Understanding
Marco Dinarelli, Alessandro Moschitti, Giuseppe Riccardi |
EACL | 3 |
| 2009 | Re-Ranking Models Based-on Small Training Data for Spoken Language Understanding
Marco Dinarelli, Alessandro Moschitti, Giuseppe Riccardi |
EMNLP | 3 |
| 2009 | Convolution Kernels on Constituent, Dependency and Sequential Structures for Relation Extraction
Truc-Vien T. Nguyen, Alessandro Moschitti, Giuseppe Riccardi |
EMNLP | 3 |
| 2009 | Concept segmentation and labeling for conversational speechabstractSpoken Language Understanding performs automatic concept labeling and segmentation of speech utterances. For this task, many approaches have been proposed based on both generative and discriminative models. While all these methods have shown remarkable accuracy on manual transcription of spoken utterances, robustness to noisy automatic transcription is still an open issue. In this paper we study algorithms for Spoken Language Understanding combining complementary learning models: Stochastic Finite State Transducers produce a list of hypotheses, which are re-ranked using a discriminative algorithm based on kernel methods. Our experiments on two different spoken dialog corpora, MEDIA and LUNA, show that the combined generative-discriminative model reaches the state-of-the-art such as Conditional Random Fields (CRF) on manual transcriptions, and it is robust to noisy automatic transcriptions, outperforming, in some cases, the state-of-the-art. Copyright © 2009 ISCA. Marco Dinarelli, Alessandro Moschitti, Giuseppe Riccardi |
INTERSPEECH | 3 |
| 2009 | A statistical dialog manager for the LUNA projectabstractIn this paper, we present an approach for the development of a statistical dialog manager, in which the system response is selected by means of a classification process which considers all the previous history of the dialog to select the next system response. In particular, we use decision trees for its implementation. The statistical model is automatically learned from training data which are labeled in terms of different SLU features. This methodology has been applied to develop a dialog manager within the framework of the European LUNA project, whose main goal is the creation of a robust natural spoken language understanding system. We present an evaluation of this approach for both human machine and human-human conversations acquired in this project. We demonstrate that a statistical dialog manager developed with the proposed technique and learned from a corpus of human-machine dialogs can successfully infer the task-related topics present in spontaneous humanhuman dialogs. David Griol, Giuseppe Riccardi, Emilio Sanchis Arnal |
INTERSPEECH | 2 |
| 2009 | Learning the structure of human-computer and human-human dialogsabstractWe are interested in the problem of understanding human conversation structure in the context of human-machine and human-human interaction. We present a statistical methodol-ogy for detecting the structure of spoken dialogs based on a generative model learned using decision trees. To evaluate our approach we have used the LUNA corpora, collected from real users engaged in problem solving tasks. The results of the evaluation show that automatic segmentation of spoken dialogs is very effective not only with models built using separately human-machine dialogs or human-human dialogs, but it is also possible to infer the task-related structure of human-human di-alogs with a model learned using only human-machine dialogs. David Griol, Giuseppe Riccardi, Emilio Sanchis Arnal |
INTERSPEECH | 2 |
| 2009 | What's in an ontology for spoken language understandingabstractCurrent Spoken Language Understanding systems rely either on hand-written semantic grammars or on flat attribute-value se-quence labeling. In both approaches, concepts and their rela-tions (when modeled at all) are domain-specific, thus making it difficult to expand or port the domain model. To address this issue, we introduce: 1) a domain model based on an ontology where concepts are classified into either predicative or argumentative; 2) the modeling of relations be-tween such concept classes in terms of classical relations as defined in lexical semantics. We study and analyze our ap-proach on a corpus of customer care data, where we evaluate the coverage and relevance of the ontology for the interpreta-tion of speech utterances. Index Terms: Spoken Language Understanding, domain mod-eling, ontology design, semantic relations Silvia Quarteroni, Giuseppe Riccardi, Marco Dinarelli |
INTERSPEECH | 2 |
| 2009 | Leveraging POMDPs Trained with User Simulations and Rule-based Dialogue Management in a Spoken Dialogue System
Sebastian Varges, Silvia Quarteroni, Giuseppe Riccardi, Alexei V. Ivanov, Pierluigi Roberti |
SIGDIAL Conference | 3 |
| 2008 | Active Annotation in the LUNA Italian Corpus of Spontaneous Dialogues
Christian Raymond, Kepa Joseba Rodríguez, Giuseppe Riccardi |
LREC | 3 |
| 2008 | Semantic annotations for conversational speech: From speech transcriptions to predicate argument structuresabstractIn this paper, we describe the semantic content, which can be automatically generated, for the design of advanced dialog systems. Since the latter will be based on machine learning approaches, we created training data by annotating a corpus with the needed content. Given a sentence of our transcribed corpus, domain concepts and other linguistic levels ranging from basic ones, i.e. part-of-speech tagging and constituent chunking level, to more advanced ones, i.e. syntactic and predicate argument structure (PAS) levels are annotated. In particular, the proposed PAS and taxonomy of dialog acts appear to be promising for the design of more complex dialog systems. Statistics about our semantic annotation are reported. Arianna Bisazza, Marco Dinarelli, Silvia Quarteroni, Sara Tonelli, Alessandro Moschitti, Giuseppe Riccardi |
SLT | 6 |
| 2008 | Automatic framenet-based annotation of conversational speechabstractCurrent Spoken Language Understanding technology is based on a simple concept annotation of word sequences, where the interdependencies between concepts and their compositional semantics are neglected. This prevents an effective handling of language phenomena, with a consequential limitation on the design of more complex dialog systems. In this paper, we argue that shallow semantic representation as formulated in the Berkeley FrameNet Project may be useful to improve the capability of managing more complex dialogs. To prove this, the first step is to show that a FrameNet parser of sufficient accuracy can be designed for conversational speech. We show that exploiting a small set of FrameNet-based manual annotations, it is possible to design an effective semantic parser. Our experiments on an Italian spoken dialog corpus, created within the LUNA project, show that our approach is able to automatically annotate unseen dialog turns with a high accuracy. Bonaventura Coppola, Alessandro Moschitti, Sara Tonelli, Giuseppe Riccardi |
SLT | 4 |
| 2008 | Joint generative and discriminative models for spoken language understandingabstractSpoken Language Understanding aims at mapping a natural language spoken sentence into a semantic representation. In the last decade two main approaches have been pursued: generative and discriminative models. The former is more robust to overfitting whereas the latter is more robust to many irrelevant features. Additionally, the way in which these approaches encode prior knowledge is very different and their relative performance changes based on the task. In this paper we describe a training framework where both models are used: a generative model produces a list of ranked hypotheses whereas a discriminative model, depending on string kernels and Support Vector Machines, re-ranks such list. We tested such approach on a new corpus produced in the European LUNA project. The results show a large improvement on the state-of-the-art in concept segmentation and labeling. Marco Dinarelli, Alessandro Moschitti, Giuseppe Riccardi |
SLT | 3 |
| 2007 | Spoken language understanding with kernels for syntactic/semantic structuresabstractAutomatic concept segmentation and labeling are the fundamental problems of spoken language understanding in dialog systems. Such tasks are usually approached by using generative or discriminative models based on n-grams. As the uncertainty or ambiguity of the spoken input to dialog system increase, we expect to need dependencies beyond n-gram statistics. In this paper, a general purpose statistical syntactic parser is used to detect syntactic/semantic dependencies between concepts in order to increase the accuracy of sentence segmentation and concept labeling. The main novelty of the approach is the use of new tree kernel functions which encode syntactic/semantic structures in discriminative learning models. We experimented with support vector machines and the above kernels on the standard ATIS dataset. The proposed algorithm automatically parses natural language text with off-the-shelf statistical parser and labels the syntactic (sub)trees with concept labels. The results show that the proposed model is very accurate and competitive with respect to state-of-the-art models when combined with n-gram based models. Alessandro Moschitti, Giuseppe Riccardi, Christian Raymond |
ASRU | 2 |
| 2007 | A data-centric architecture for data-driven spoken dialog systemsabstractData is becoming increasingly crucial for training and (self-) evaluation of spoken dialog systems (SDS). Data is used to train models (e.g. acoustic models) and is 'forgotten'. Data is generated on-line from the different components of the SDS system, e.g. the dialog manager, as well as from the world it is interacting with (e.g. news streams, ambient sensors etc.). The data is used to evaluate and analyze conversational systems both on-line and off-line. We need to be able query such heterogeneous data for further processing. In this paper we present an approach with two novel components: first, an architecture for SDSs that takes a data-centric view, ensuring persistency and consistency of data as it is generated. The architecture is centered around a database that stores dialog data beyond the lifetime of individual dialog sessions, facilitating dialog mining, annotation, and logging. Second, we take advantage of the state-fullness of the data-centric architecture by means of a lightweight, reactive and inference-based dialog manager that itself is stateless. The feasibility of our approach has been validated within a prototype of a phone-based university help-desk application. We detail SDS architecture and dialog management, model, and data representation. Sebastian Varges, Giuseppe Riccardi |
ASRU | 2 |
| 2007 | Generative and discriminative algorithms for spoken language understandingabstractSpoken Language Understanding (SLU) for conversational systems (SDS) aims at extracting concept and their relations from spontaneous speech. Previous approaches to SLU have modeled concept relations as stochastic semantic networks ranging from generative approach to discriminative. As spoken dialog systems complexity increases, SLU needs to perform understanding based on a richer set of features ranging from a-priori knowledge, long dependency, dialog history, system belief, etc. This paper studies generative and discriminative approaches to modeling the sentence segmentation and concept labeling. We evaluate algorithms based on Finite State Transducers (FST) as well as discriminative algorithms based on Support Vector Machine sequence classifier based and Conditional Random Fields (CRF). We compare them in terms of concept accuracy, generalization and robustness to annotation ambiguities. We also show how non-local non-lexical features (e.g. a-priori knowledge) can be modeled with CRF which is the best performing algorithm across tasks. The evaluation is carried out on two SLU tasks of different complexity, namely ATIS and MEDIA corpora. Christian Raymond, Giuseppe Riccardi |
INTERSPEECH | 2 |
| 2006 | Beyond ASR 1-best: Using word confusion networks in spoken language understanding
Dilek Hakkani-Tür, Frédéric Béchet, Giuseppe Riccardi, Gökhan Tür |
Comput. Speech Lang. | 3 |
| 2006 | The AT&T spoken language understanding systemabstractSpoken language understanding (SLU) aims at extracting meaning from natural language speech. Over the past decade, a variety of practical goal-oriented spoken dialog systems have been built for limited domains. SLU in these systems ranges from understanding predetermined phrases through fixed grammars, extracting some predefined named entities, extracting users' intents for call classification, to combinations of users' intents and named entities. In this paper, we present the SLU system of VoiceTone/spl reg/ (a service provided by AT&T where AT&T develops, deploys and hosts spoken dialog applications for enterprise customers). The SLU system includes extracting both intents and the named entities from the users' utterances. For intent determination, we use statistical classifiers trained from labeled data, and for named entity extraction we use rule-based fixed grammars. The focus of our work is to exploit data and to use machine learning techniques to create scalable SLU systems which can be quickly deployed for new domains with minimal human intervention. These objectives are achieved by 1) using the predicate-argument representation of semantic content of an utterance; 2) extending statistical classifiers to seamlessly integrate hand crafted classification rules with the rules learned from data; and 3) developing an active learning framework to minimize the human labeling effort for quickly building the classifier models and adapting them to changes. We present an evaluation of this system using two deployed applications of VoiceTone/spl reg/. Narendra K. Gupta, Gökhan Tür, Dilek Hakkani-Tür, Srinivas Bangalore, Giuseppe Riccardi, Mazin Gilbert |
IEEE Trans. Speech Audio Process. | 5 |
| 2005 | The AT&T WATSON Speech RecognizerabstractThis paper describes the AT&T WATSON real-time speech recognizer, the product of several decades of research at AT&T. The recognizer handles a wide range of vocabulary sizes and is based on continuous-density hidden Markov models for acoustic modeling and finite state networks for language modeling. The recognition network is optimized for efficient search. We identify the algorithms used for high-accuracy, real-time and low-latency recognition. We present results for small and large vocabulary tasks taken from the AT&T VoiceTone/sup /spl reg// service, showing word accuracy improvement of about 5% absolute and real-time processing speed-up by a factor between 2 and 3. Vincent Goffin, Cyril Allauzen, Enrico Bocchieri, Dilek Hakkani-Tür, Andrej Ljolje, Sarangarajan Parthasarathy, Mazin G. Rahim, Giuseppe Riccardi, Murat Saraclar |
ICASSP (1) | 8 |
| 2005 | Error Prediction in Spoken Dialog: From Signal-to-Noise Ratio to Semantic Confidence ScoresabstractSpoken dialog systems aim to interpret the meanings of users' utterances and respond to them accordingly. The users' utterances are first recognized by an automatic speech recognizer (ASR) and the intents of the users are extracted by the spoken language understanding (SLU) unit. Both ASR and SLU are noisy and in general their noise statistics are not correlated. Our goal is to exploit the signal-to-noise information and ASR lattice-based and semantic confidence scores for SLU error prediction and prevention of these by rejecting erroneous utterances, or asking confirmation questions. In our experiments, we have shown up to 80% relative decrease in the error rate of the accepted utterances collected using the AT&T How May I Help You/spl trade/ spoken dialog system used for customer care. Dilek Hakkani-Tür, Gökhan Tür, Giuseppe Riccardi, Hong Kook Kim |
ICASSP (1) | 3 |
| 2005 | Using context to improve emotion detection in spoken dialog systemsabstractMost research that explores the emotional state of users of spoken dialog systems does not fully utilize the contextual nature that the dialog structure provides.This paper reports results of machine learning experiments designed to automatically classify the emotional state of user turns using a corpus of 5,690 dialogs collected with the "How May I Help You SM " spoken dialog system.We show that augmenting standard lexical and prosodic features with contextual features that exploit the structure of spoken dialog and track user state increases classification accuracy by 2.6%. Jackson Liscombe, Giuseppe Riccardi, Dilek Hakkani-Tür |
INTERSPEECH | 2 |
| 2005 | Adaptive categorical understanding for spoken dialogue systemsabstractIn this paper, the speech understanding problem in the context of a spoken dialogue system is formalized in a maximum likelihood framework. Off-line adaptation of stochastic language models that interpolate dialogue state specific and general application-level language models is proposed. Word and dialogue-state n-grams are used for building categorical understanding and dialogue models, respectively. Acoustic confidence scores are incorporated in the understanding formulation. Problems due to data sparseness and out-of-vocabulary words are discussed. The performance of the speech recognition and understanding language models are evaluated with the "Carmen Sandiego" multimodal computer game corpus. Incorporating dialogue models reduces relative understanding error rate by 15%-25%, while acoustic confidence scores achieve a further 10% error reduction for this computer gaming application. Alexandros Potamianos, Shri Narayanan, Giuseppe Riccardi |
IEEE Trans. Speech Audio Process. | 3 |
| 2005 | Active learning: theory and applications to automatic speech recognitionabstractWe are interested in the problem of adaptive learning in the context of automatic speech recognition (ASR). In this paper, we propose an active learning algorithm for ASR. Automatic speech recognition systems are trained using human supervision to provide transcriptions of speech utterances. The goal of Active Learning is to minimize the human supervision for training acoustic and language models and to maximize the performance given the transcribed and untranscribed data. Active learning aims at reducing the number of training examples to be labeled by automatically processing the unlabeled examples, and then selecting the most informative ones with respect to a given cost function for a human to label. In this paper we describe how to estimate the confidence score for each utterance through an on-line algorithm using the lattice output of a speech recognizer. The utterance scores are filtered through the informativeness function and an optimal subset of training samples is selected. The active learning algorithm has been applied to both batch and on-line learning scheme and we have experimented with different selective sampling algorithms. Our experiments show that by using active learning the amount of labeled data needed for a given word accuracy can be reduced by more than 60% with respect to random sampling. Giuseppe Riccardi, Dilek Hakkani-Tür |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Mining Spoken Dialogue Corpora for System Evaluation and Modelin
Frédéric Béchet, Giuseppe Riccardi, Dilek Hakkani-Tür |
EMNLP | 2 |
| 2004 | Unsupervised and active learning in automatic speech recognition for call classificationabstractA key challenge in rapidly building spoken natural language dialog applications is minimizing the manual effort required in transcribing and labeling speech data. This task is not only expensive but also time consuming. We present a novel approach that aims at reducing the amount of manually transcribed in-domain data required for building automatic speech recognition (ASR) models in spoken language dialog systems. Our method is based on mining relevant text from various conversational systems and Web sites. An iterative process is employed where the performance of the models can be improved through both unsupervised and active learning of the ASR models. We have evaluated the robustness of our approach on a call classification task that has been selected from AT&T VoiceTone/sup SM/ customer care. Our results indicate that with unsupervised learning it is possible to achieve a call classification performance that is only 1.5% lower than the upper bound set when using all available in-domain transcribed data. Dilek Hakkani-Tür, Gökhan Tür, Mazin G. Rahim, Giuseppe Riccardi |
ICASSP (1) | 4 |
| 2004 | Extending boosting for call classification using word confusion networksabstractWe are interested in the problem of robust understanding from noisy spontaneous speech input. In goal driven human-machine dialog, utterance classification is a key component of the understanding process to determine the intent of the speaker. We propose a novel algorithm for exploiting ASR word confidence scores for better classification of spoken utterances. Word confidence scores for automatic speech recognition (ASR) provide estimates for word error rates. While previous work has focused on straightforward combination of word confidence scores into Bayesian classifiers, we extend the mathematical formulation for boosting classifiers. This extension of the algorithm allows confidence scores to be exploited from a 1-best ASR output or from word confusion networks (WCNs). We present methods for on-line and off-line score combinations. The results we show are for a large database of utterances collected using the AT&T VoiceTone/sup SM/ spoken dialog system. Our experiments show between 5% and 10% reduction in error (1-precision) for a given recall using WCNs compared to ASR output. Gökhan Tür, Dilek Hakkani-Tür, Giuseppe Riccardi |
ICASSP (1) | 3 |
| 2003 | A general algorithm for word graph matrix decompositionabstractIn automatic speech recognition, word graphs (lattices) are commonly used as an approximate representation of the complete word search space. Usually these word lattices are acyclic and have no a-priori structure. More recently a new class of normalized word lattices have been proposed. These word lattices (a.k.a. sausages) are very efficient (space) and they provide a normalization (chunking) of the lattice, by aligning words from all possible hypotheses. We propose a general framework for lattice chunking, the pivot algorithm. There are four important components of the pivot algorithm. First, the time information is not necessary but is beneficial for the overall performance. Second, the algorithm allows the definition of a predefined chunk structure of the final word lattice. Third, the algorithm operates on both weighted and unweighted lattices. Fourth, the labels on the graph are generic, and could be words as well as part of speech tags or parse tags. While the algorithm has applications to many tasks (e.g. parsing, named entity extraction) we present results on the performance of confidence scores for different large vocabulary speech recognition tasks. We compare the results of our algorithms against off-the-shelf methods and show significant improvements. Dilek Hakkani-Tür, Giuseppe Riccardi |
ICASSP (1) | 2 |
| 2003 | Multi-channel sentence classification for spoken dialogue language modelingabstractIn traditional language modeling word prediction is based on the local context (e.g. n-gram). In spoken dialog, language statistics are affected by the multidimensional structure of the human-machine interaction. In this paper we investigate the statistical dependencies of users’ responses with respect to the system’s and user’s channel. The system channel components are the prompts’ text, dialogue history, dialogue state. The user channel components are the Automatic Speech Recognition (ASR) transcriptions, the semantic classifier output and the sentence length. We describe an algorithm for language model rescoring using users’ response classification. The user’s response is first mapped into a multidimensional state and the state specific language model is applied for ASR rescoring. We present perplexity and ASR results on the How May I Help You ? sm 100K spoken dialogs. Frédéric Béchet, Giuseppe Riccardi, Dilek Hakkani-Tür |
INTERSPEECH | 2 |
| 2003 | Active and unsupervised learning for automatic speech recognitionabstractState-of-the-art speech recognition systems are trained using human transcriptions of speech utterances. In this paper, we describe a method to combine active and unsupervised learning for automatic speech recognition (ASR). The goal is to minimize the human supervision for training acoustic and language models and to maximize the performance given the transcribed and untranscribed data. Active learning aims at reducing the number of training examples to be labeled by automatically processing the unlabeled examples, and then selecting the most informative ones with respect to a given cost function. For unsupervised learning, we utilize the remaining untranscribed data by using their ASR output and word confidence scores. Our experiments show that the amount of labeled data needed for a given word accuracy can be reduced by 75% by combining active and unsupervised learning. Giuseppe Riccardi, Dilek Hakkani-Tür |
INTERSPEECH | 1 |
| 2002 | Bootstrapping Bilingual Data using Consensus Translation for a Multilingual Instant Messaging System
Srinivas Bangalore, Vanessa Murdock 0001, Giuseppe Riccardi |
COLING | 3 |
| 2002 | Active learning for automatic speech recognitionabstractState-of-the-art speech recognition systems are trained using transcribed utterances, preparation of which is labor intensive and time-consuming. In this paper, we describe a new method for reducing the transcription effort for training in automatic speech recognition (ASR). Active learning aims at reducing the number of training examples to be labeled by automatically processing the unlabeled examples, and then selecting the most informative ones with respect to a given cost function for a human to label. We automatically estimate a confidence score for each word of the utterance, exploiting the lattice output of a speech recognizer, which was trained on a small set of transcribed data. We compute utterance confidence scores based on these word confidence scores, then selectively sample the utterances to be transcribed using the utterance confidence scores. In our experiments, we show that we reduce the amount of labeled data needed for a given word accuracy by 27%. 1. Dilek Hakkani-Tür, Giuseppe Riccardi, Allen L. Gorin |
ICASSP | 2 |
| 2002 | Combining prior knowledge and boosting for call classification in spoken language dialogueabstractData collection and annotation are major bottlenecks in rapid development of accurate syntactic and semantic models for natural-language dialogue systems. In this paper we show how human knowledge can be used when designing a language understanding system in a manner that would alleviate the dependence on large sets of data. In particular, we extend BoosTexter, a member of the boosting family of algorithms, to combine and balance hand-crafted rules with the statistics of available data. Experiments on two voice-enabled applications for customer care and help desk are presented. Marie Rochery, Robert E. Schapire, Mazin G. Rahim, Narendra K. Gupta, Giuseppe Riccardi, Srinivas Bangalore, Hiyan Alshawi, Shona Douglas |
ICASSP | 5 |
| 2002 | AT&t help desk
Giuseppe Di Fabbrizio, Dawn Dutton, Narendra K. Gupta, Barbara Hollister, Mazin G. Rahim, Giuseppe Riccardi, Robert E. Schapire, Juergen Schroeter |
INTERSPEECH | 6 |
| 2002 | Acoustic and word lattice based algorithms for confidence scoresabstractWord confidence scores are crucial for unsupervised learning in automatic speech recognition. In the last decade there has been a flourish of work on two fundamentally different approaches to compute confidence scores. The first paradigm is acoustic and the second is based on word lattices. The first approach is dataintensive and it requires to explicitly model the acoustic channel. The second approach is suitable for on-line (unsupervised) learning and requires no training. In this paper we present a comparative analysis of off-the-shelf and new algorithms for computing confidence scores, following the acoustic and lattice-based paradigms. We compare the performance of these algorithms across three tasks for small, medium and large vocabulary speech recognition tasks and for two languages (Italian and English). We show that wordlattice based algorithm provides consistent and effective performance across automatic speech recognition tasks. 1. Daniele Falavigna, Roberto Gretter, Giuseppe Riccardi |
INTERSPEECH | 3 |
| 2002 | Improving spoken language understanding using word confusion networksabstractA natural language spoken dialog system includes a large vocabulary automatic speech recognition (ASR) engine, whose output is used as the input of a spoken language understanding component. Two challenges in such a framework are that the ASR component is far from being perfect and the users can say the same thing in very different ways. So, it is very important to be tolerant to recognition errors and some amount of orthographic variability. In this paper, we present our work on developing new methods and investigating various ways of robust recognition and understanding of an utterance. To this end, we exploit word-level confusion networks (sausages), obtained from ASR word graphs (lattices) instead of the ASR 1-best hypothesis. Using sausages with an improved confidence model, we decreased the calltype classification error rate for AT&T's How May I Help You (HMIHY ) natural dialog system by 38%. Gökhan Tür, Jeremy H. Wright, Allen L. Gorin, Giuseppe Riccardi, Dilek Hakkani-Tür |
INTERSPEECH | 4 |
| 2002 | Stochastic Finite-State Models for Spoken Language Machine Translation
Srinivas Bangalore, Giuseppe Riccardi |
Mach. Transl. | 2 |
| 2001 | On-line learning of language models with word error probability distributionsabstractWe are interested in the problem of learning stochastic language models on-line (without speech transcriptions) for adaptive speech recognition and understanding. We propose an algorithm to adapt to variations in the language model distributions based on speech input only and without its true transcription. The on-line probability estimate is defined. as a function of the prior and word error distributions. We show the effectiveness of word-lattice based error probability distributions in terms of receiver operating characteristics (ROC) curves and word accuracy. We apply the new estimates P/sub adapt/(w) to the task of adapting on-line an initial large vocabulary trigram language model and show improvement in word accuracy with respect to the baseline speech recognizer. Roberto Gretter, Giuseppe Riccardi |
ICASSP | 2 |
| 2001 | A Finite-State Approach to Machine Translation
Srinivas Bangalore, Giuseppe Riccardi |
NAACL | 2 |
| 2001 | Robust numeric recognition in spoken language dialogue
Mazin G. Rahim, Giuseppe Riccardi, Lawrence K. Saul, Jeremy H. Wright, Bruce Buntschuh, Allen L. Gorin |
Speech Commun. | 2 |
| 2001 | Integration of utterance verification with statistical language modeling and spoken language understanding
Richard C. Rose, H. Yao, Giuseppe Riccardi, Jeremy H. Wright |
Speech Commun. | 3 |
| 2000 | Finite-state models for lexical reordering in spoken language translation
Srinivas Bangalore, Giuseppe Riccardi |
INTERSPEECH | 2 |
| 2000 | Detecting acoustic morphemes in lattices for spoken language understandingabstractCurrent methods for training statistical language models for recognition and understanding require large annotated corpora. The collection, transcription and labeling of such corpora is a major bottleneck for creating new applications and for refinements of existing ones. Thus, it is of great interest to develop methods for automatically learning vocabulary, grammar and semantics from a speech corpus without transcriptions. In this paper we report on an experiment where acoustic morphemes are automatically acquired from the output of a task-independent phone recognizer. The utility of these units is experimentally evaluated for call-type classification in the 'How may I help you?' task. Detected occurrences of the acoustic morphemes in the lattice output provide the basis for the classification of the test sentences. Using lattices, we achieve a reduction of 59% from the false rejection rate using best paths, albeit with a 5% reduction in the correct classification performance from that baseline. Dijana Petrovska-Delacrétaz, Allen L. Gorin, Jeremy H. Wright, Giuseppe Riccardi |
INTERSPEECH | 4 |
| 2000 | A spoken dialogue system for conference/workshop servicesabstractThis paper describes our progress towards building a telephony-based spoken dialogue system for workshop/conference services. A mixed-initiative dialogue system has been developed that is engineered to o er users natural interaction with the system, ease-of-use and robustness towards ambiguous requests and machine errors. A prototype system, known as W99, is described in this paper Mazin G. Rahim, Roberto Pieraccini, Wieland Eckert, Esther Levin, Giuseppe Di Fabbrizio, Giuseppe Riccardi, Candace A. Kamm, Shri Narayanan |
INTERSPEECH | 6 |
| 2000 | On-line learning of acoustic and lexical units for domain-independent ASR
Giuseppe Riccardi |
INTERSPEECH | 1 |
| 2000 | Stochastic language adaptation over time and state in natural spoken dialog systemsabstractWe are interested in adaptive spoken dialog systems for automated services. Peoples' spoken language usage varies over time for a given task, and furthermore varies depending on the state of the dialog. Thus, it is crucial to adapt automatic speech recognition (ASR) language models to these varying conditions. We characterize and quantify these variations based on a database of 30 K user-transactions with AT&T's experimental How May I Help You? spoken dialog system. We describe a novel adaptation algorithm for language models with time and dialog-state varying parameters. Our language adaptation framework allows for recognizing and understanding unconstrained speech at each stage of the dialog, enabling context-switching and error recovery. These models have been used to train state-dependent ASR language models. We have evaluated their performance with respect to word accuracy and perplexity over time and dialog states. We have achieved a reduction of 40% in perplexity and of 8.4% in word error rate over the baseline system, averaged across all dialog states. Giuseppe Riccardi, Allen L. Gorin |
IEEE Trans. Speech Audio Process. | 1 |
| 1999 | Spoken language variation over time and state in a natural spoken dialog systemabstractWe are interested in adaptive spoken dialog systems for automated services. Peoples' spoken language usage varies over time for a fixed task, and furthermore varies depending on the state of the dialog. We characterize and quantify this variation based on a database of 20 K user-transactions with AT&T's experimental 'How May I Help You?' spoken dialog system. We then report on a language adaptation algorithm which was used to train state-dependent ASR language models, experimentally evaluating their improved performance with respect to word accuracy and perplexity. Allen L. Gorin, Giuseppe Riccardi |
ICASSP | 2 |
| 1999 | Modeling disfluency and background events in ASR for a natural language understanding taskabstractThis paper investigates techniques for minimizing the impact of non-speech events on the performance of large vocabulary continuous speech recognition (LVCSR) systems. An experimental study is presented that investigates whether the careful manual labeling of disfluency and background events in conversational speech can be used to provide an additional level of supervision in training HMM acoustic models and statistical language models. First, techniques are investigated for incorporating explicitly labeled disfluency and background events directly into the acoustic HMM. Second, phrase-based statistical language models are trained from utterance transcriptions which include labeled instances of these events. Finally, it is shown that significant word accuracy and run-time performance improvements are obtained for both sets of techniques on a telephone-based spoken language understanding task. Richard C. Rose, Giuseppe Riccardi |
ICASSP | 2 |
| 1999 | Prosody recognition from speech utterances using acoustic and linguistic based models of prosodic events
Alistair Conkie, Giuseppe Riccardi, Richard C. Rose |
EUROSPEECH | 2 |
| 1999 | Categorical understanding using statistical ngram models
Alexandros Potamianos, Giuseppe Riccardi, Shri Narayanan |
EUROSPEECH | 2 |
| 1999 | Automatic speech recognition using acoustic confidence conditioned language modelsabstractA modi ed decoding algorithm for automatic speech recognition ASR will be described which facilitates a closer coupling between the acoustic and language modeling components of a speech recognition system.This closer coupling is obtained by extracting word level measures of acoustic con dence during decoding, and making coded representations of these con dence measures available to the ASR network during decoding.A simulation of this decoding strategy is implemented using a word lattice rescoring paradigm.A joint acoustic language model will be described where linguistic context is augmented to include the encoded values of acoustic con dence.Finally, the performance of the word lattice based implementation of the decoding algorithm will be evaluated on a large vocabulary natural language understanding task. Richard C. Rose, Giuseppe Riccardi |
EUROSPEECH | 2 |
| 1999 | Grammar Fragment acquisition using syntactic and semantic clustering
Kazuhiro Arai, Jeremy H. Wright, Giuseppe Riccardi, Allen L. Gorin |
Speech Commun. | 3 |
| 1998 | Integration of utterance verification with statistical language modeling and spoken language understandingabstractMethods for utterance verification (UV) and their integration into statistical language modeling and spoken language understanding formalisms for a large vocabulary spoken understanding system are presented. The paper consists of three parts. First, a set of acoustic likelihood ratio based utterance verification techniques are described and applied to the problem of rejecting portions of a hypothesized word string that may have been incorrectly decoded by a large vocabulary continuous speech recognizer. Second, a procedure for integrating the acoustic level confidence measures with the statistical language model is described. Finally, the effect of integrating acoustic level confidence into the spoken language understanding unit (SLU) in a call-type classification task is discussed. These techniques were evaluated on utterances collected from a highly unconstrained call routing task performed over the telephone network. They have been evaluated in terms of their ability to classify utterances into a set of fifteen semantic actions corresponding to call-types that are accepted by the application. Richard C. Rose, H. Yao, Giuseppe Riccardi, Jeremy H. Wright |
ICASSP | 3 |
| 1998 | Grammar fragment acquisition using syntactic and semantic clusteringabstractA new method for automatically acquiring Fragments for understanding fluent speech is proposed. The goal of this method is to generate a collection of Fragments, each representing a set of syntactically and semantically similar phrases. First, phrases observed frequently in the training set are selected as candidates. Each candidate phrase has three associated probability distributions: of following contexts, of preceding contexts, and of associated semantic actions. The similarity between candidate phrases is measured by applying the Kullback-Leibler distance to these three probability distributions. Candidate phrases that are close in all three distances are clustered into a Fragment. Salient sequences of these Fragments are then automatically acquired, and exploited by a spoken language understanding module to classify calls in AT&T's "How May I Help You?" task. These Fragments allow us to generalize unobserved phrases. For instance, they detected 246 phrases in the test-set that we... Kazuhiro Arai, Jeremy H. Wright, Giuseppe Riccardi, Allen L. Gorin |
ICSLP | 3 |
| 1998 | Stochastic language models for speech recognition and understandingabstractStochastic language models for speech recognition have traditionally been designed and evaluated in or-der to optimize word accuracy. In this work, we present a novel framework for training stochastic language models by optimizing two different criteria appropri-ate for speech recognition and language understand-ing. First, the language entropy and salience measure are used for learning the relevant spoken language features (phrases). Secondly, a novel algorithm for training stochastic finite state machines is presented which incorporates the acquired phrase structure into a single stochastic language model. Thirdly, we show the benefit of our novel framework with an end-to-end evaluation of a large vocabulary spoken language system for call routing. 2. Giuseppe Riccardi, Allen L. Gorin |
ICSLP | 1 |
| 1998 | Language model adaptation for spoken language systemsabstract1. ABSTRACT In a human-machine interaction (dialog) the statistical language variations are large among different stages of the dialog and across different speakers. Moreover, spoken dialog systems require extensive training data for training adaptive language models. In this paper we address the problem of open-vocabulary language models allowing the user for any possible response at each stage of the dialog. We propose a novel off-line adaptation of stochastic language models effective for their generalization (openvocabulary) and selective (dialog context) properties. We outline the integration of the finite state dialog model and the language model adaptation algorithm. The performance of the speech recognition and understanding language models are evaluated with the Carmen Sandiego multimodal computer game. The new language models give an overall understanding error rate reduction of 44% over the baseline system. Giuseppe Riccardi, Alexandros Potamianos, Shri Narayanan |
ICSLP | 1 |
| 1997 | A spoken language system for automated call routingabstractWe are interested in the problem of understanding fluently spoken language. In particular, we consider people's responses to the open-ended prompt of "How may I help you?". We then further restrict the problem to classifying and automatically routing such a call, based on the meaning of the user's response. Thus, we aim at extracting a relatively small number of semantic actions from the utterances of a very large set of users who are not trained to the system's capabilities and limitations. In this paper, we describe the main components of our speech understanding system: the large vocabulary recognizer and the language understanding module performing the call-type classification. In particular, we propose automatic algorithms for selecting phrases from a training corpus in order to enhance the prediction power of the standard word n-gram. The phrase language models are integrated into stochastic finite state machines which outperform standard word n-gram language models. From the speech recognizer output we recognize and exploit automatically-acquired salient phrase fragments to make a call-type classification. This system is evaluated on a database of 10 K fluently spoken utterances collected from interactions between users and human agents. Giuseppe Riccardi, Allen L. Gorin, Andrej Ljolje, Michael Riley 0001 |
ICASSP | 1 |
| 1997 | Automatic acquisition of salient grammar fragments for call-type classification
Jeremy H. Wright, Allen L. Gorin, Giuseppe Riccardi |
EUROSPEECH | 3 |
| 1997 | How may I help you?
Allen L. Gorin, Giuseppe Riccardi, Jeremy H. Wright |
Speech Commun. | 2 |
| 1996 | Stochastic automata for language modeling
Giuseppe Riccardi, Roberto Pieraccini, Enrico Bocchieri |
Comput. Speech Lang. | 1 |
| 1995 | Non-deterministic stochastic language models for speech recognitionabstractTraditional stochastic language models for speech recognition (i.e. n-grams) are deterministic, in the sense that there is one and only one derivation for each given sentence. Moreover a fixed temporal window is always assumed in the estimation of the traditional stochastic language models. This paper shows how non-determinism is introduced to effectively approximate a back-off n-gram language model through a finite state network formalism. It also shows that a new flexible and powerful network formalization can be obtained by releasing the assumption of a fixed history size. As a result, a class of automata for language modeling (variable n-gram stochastic automata) is obtained, for which we propose some methods for the estimation of the transition probabilities. VNSAs have been used in a spontaneous speech recognizer for the ATIS task. The accuracy on a standard test set is presented. Giuseppe Riccardi, Enrico Bocchieri, Roberto Pieraccini |
ICASSP | 1 |
| 1995 | State tying of triphone HMM's for the 1994 AT&t ARPA ATIS recognizer
Enrico Bocchieri, Giuseppe Riccardi |
EUROSPEECH | 2 |
| 1994 | A localization property of line spectrum frequenciesabstractThe interlacing property for the line spectrum frequencies (LSFs) is extended to the LSFs associated to successive order predictor polynomials. The corresponding separation theorem gives the most precise lower and upper bounds of the intervals which the LSFs may belong to.> Gian Antonio Mian, Giuseppe Riccardi |
IEEE Trans. Speech Audio Process. | 2 |
| 1993 | An approach to parameter reoptimization in multipulse-based codersabstractAn algorithm for LPC parameter optimization in multipulse (MP)-linear-prediction-coding (LPC)-based speech coders is presented. It is shown that by taking the nature of the MP excitation signal into account in the LPC parameter computation, it is possible to improve the effectiveness of the LPC model. This results in a better quality of the reconstructed signal in terms of both objective and subjective criteria. The implementation details of the algorithm are discussed, and experimental results are presented. In particular, a comparison with standard MP LPC techniques is given.> Marco Fratti, Gian Antonio Mian, Giuseppe Riccardi |
IEEE Trans. Speech Audio Process. | 3 |
| 1992 | On the effectiveness of parameter reoptimization in multipulse based codersabstractAn algorithm for LPC (linear predictive coding) parameter optimization in multipulse (MP)-LPC based speech coders is presented. It is shown that, by taking into account the nature of the MP-excitation signal into LPC parameter computation, it is possible to improve the effectiveness of the LPC model. This results in a better quality of the reconstructed signal in terms both of objective and subjective criteria. The implementation details of the algorithm are discussed and experimental results are presented. In particular a comparison with standard MP-LPC techniques is given.> Marco Fratti, Gian Antonio Mian, Giuseppe Riccardi |
ICASSP | 3 |