Rebecca J. Passonneau

dblp:04/696 · DBLP profile ↗
← Back
75ranked-venue papers
25as first author
10since 2021 · last 2025
0000-0001-8626-811XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 64 · 25 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Human-computer interaction and ubiquitous computing · 6 · 5 since 2021Databases, data management, data science and information retrieval · 3Software engineering, systems software and programming languages · 2Computer networks · 1
YearPublicationVenuePosition
2025 Improving Model Evaluation using SMART Filtering of Benchmark Datasets
abstract
Vipul Gupta, Candace Ross, David Pantoja, Rebecca J. Passonneau, Megan Ung, Adina Williams. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Candace Ross, David Pantoja, Rebecca J. Passonneau, Megan Ung, Adina Williams
NAACL (Long Papers)4
2024 VerAs: Verify Then Assess STEM Lab Reports
Berk Atil, Mahsa Sheikhi Karizaki, Rebecca J. Passonneau
AIED (1)3
2024 How Well Can You Articulate that Idea? Insights from Automated Formative Assessment
Mahsa Sheikhi Karizaki, Dana Gnesdilow, Sadhana Puntambekar, Rebecca J. Passonneau
AIED (2)4
2023 Learning When to Defer to Humans for Short Answer Grading
Yumi Jin, Xuesong Cang, Sadhana Puntambekar, Rebecca J. Passonneau
AIED6
2023 The Sentiment Problem: A Critical Survey towards Deconstructing Sentiment Analysis
abstract
Pranav Venkit, Mukund Srinath, Sanjana Gautam, Saranya Venkatraman, Vipul Gupta, Rebecca Passonneau, Shomir Wilson. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Pranav Venkit, Mukund Srinath, Sanjana Gautam, Saranya Venkatraman, Rebecca J. Passonneau, Shomir Wilson
EMNLP6
2022 CONTaiNER: Few-Shot Named Entity Recognition via Contrastive Learning
abstract
Named Entity Recognition (NER) in Few-Shot setting is imperative for entity tagging in low resource domains.Existing approaches only learn class-specific semantic features and intermediate representations from source domains.This affects generalizability to unseen target domains, resulting in suboptimal performances.To this end, we present CONTAINER, a novel contrastive learning technique that optimizes the inter-token distribution distance for Few-Shot NER.Instead of optimizing class-specific attributes, CONTAINER optimizes a generalized objective of differentiating between token categories based on their Gaussian-distributed embeddings.This effectively alleviates overfitting issues originating from training domains.Our experiments in several traditional test domains (OntoNotes, CoNLL'03, WNUT '17, GUM) and a new large scale Few-Shot NER dataset (Few-NERD) demonstrate that, on average, CONTAINER outperforms previous methods by 3%-13% absolute F1 points while showing consistent performance trends, even in challenging scenarios where previous approaches could not achieve appreciable performance.The source code of CONTAINER will be available at: https://github.com/ psunlpgroup/CONTaiNER.
Sarkar Snigdha Sarathi Das, Arzoo Katiyar, Rebecca J. Passonneau, Rui Zhang 0037
ACL (1)3
2022 Automated Support to Scaffold Students' Written Explanations in Science
Purushartha Singh, Rebecca J. Passonneau, Mohammad Wasih, Xuesong Cang, ChanMin Kim, Sadhana Puntambekar
AIED (1)2
2021 ABCD: A Graph Framework to Convert Complex Sentences to a Covering Set of Simple Sentences
abstract
Yanjun Gao, Ting-Hao Huang, Rebecca J. Passonneau. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yanjun Gao, Ting-Hao 'Kenneth' Huang, Rebecca J. Passonneau
ACL/IJCNLP (1)3
2021 Automated Assessment of Quality and Coverage of Ideas in Students' Source-Based Writing
Yanjun Gao, Rebecca J. Passonneau
AIED (2)2
2021 A Semantic Feature-Wise Transformation Relation Network for Automatic Short Answer Grading
abstract
Automatic short answer grading (ASAG) is the task of assessing students' short natural language responses to objective questions.It is a crucial component of new education platforms, and could support more wide-spread use of constructed response questions to replace cognitively less challenging multiple choice questions.We propose a Semantic Feature-wise transformation Relation Network (SFRN) that exploits the multiple components of ASAG datasets more effectively.SFRN captures relational knowledge among the questions (Q), reference answers or rubrics (R), and labeled student answers (A).A relation network learns vector representations for the elements of QRA triples, then combines the learned representations using learned semantic feature-wise transformations.We apply translation-based data augmentation to address the two problems of limited training data, and high data skew for multi-class ASAG tasks.Our model has up to 11% performance improvement over state-of-the-art results on the benchmark SemEval-2013 datasets, and surpasses custom approaches designed for a Kaggle challenge, demonstrating its generality.
Yajur Tomar, Rebecca J. Passonneau
EMNLP (1)3
2020 Dialogue Policies for Learning Board Games through Multimodal Communication
abstract
This paper presents MDP policy learning for agents to learn strategic behavior-how to play board games-during multimodal dialogues.Policies are trained offline in simulation, with dialogues carried out in a formal language.The agent has a temporary belief state for the dialogue, and a persistent knowledge store represented as an extensive-form game tree.How well the agent learns a new game from a dialogue with a simulated partner is evaluated by how well it plays the game, given its dialoguefinal knowledge state.During policy training, we control for the simulated dialogue partner's level of informativeness in responding to questions.The agent learns best when its trained policy matches the current dialogue partner's informativeness.We also present a novel data collection for training natural language modules.Human subjects who engaged in dialogues with a baseline system rated the system's language skills as above average.Further, results confirm that human dialogue partners also vary in their informativeness.
Maryam Zare, Ali Ayub, Aishan Liu, Sweekar Sudhakara, Albert F. Wagner, Rebecca J. Passonneau
SIGdial6
2019 Automated Pyramid Summarization Evaluation
abstract
Pyramid evaluation was developed to assess the content of paragraph length summaries of source texts.A pyramid lists the distinct units of content found in several reference summaries, weights content units by how many reference summaries they occur in, and produces three scores based on the weighted content of new summaries.We present an automated method that is more efficient, more transparent, and more complete than previous automated pyramid methods.It is tested on a new dataset of student summaries, and historical NIST data from extractive summarizers.
Yanjun Gao, Rebecca J. Passonneau
CoNLL3
2019 Show me how to win: a robot that uses dialog management to learn from demonstrations
abstract
We present an approach for robot learning from demonstration and communication applied to simple board games like Connect Four. In such games, a visual representation of a winning condition on the board can be converted to an extensive form representation that can then support computation of a winning strategy. We present a robot that can learn simple games from responses to visual questions based on synthesized images, or to verbal questions. We illustrate how reliance on both modalities leads to more efficient learning.
Maryam Zare, Ali Ayub, Alan R. Wagner, Rebecca J. Passonneau
FDG4
2018 PyrEval: An Automated Method for Summary Content Analysis
Yanjun Gao, Andrew Warner, Rebecca J. Passonneau
LREC3
2018 Prediction of a hotspot pattern in keyword search results
Axinia Radeva, Chuyao Shen, Shiqi Wang 0002, Qianbo Wang, Rebecca J. Passonneau
Comput. Speech Lang.6
2017 Resource Allocation for Pragmatically-Assisted Quality of Information-Aware Networking
abstract
In this work, we present a framework for handling multiple, simultaneous, natural language queries, in a resource constrained environment, using Quality of Information (QoI). Incoming queries are first parsed into response graphs, tree-like structures designed to formalize a system's understanding of user intent, via a pragmatics toolkit. The system then uses a combination of QoI-awareness, adaptive intent determination, and packing algorithms to maximize the QoI realized by the system. We employ two different methods of evaluation, a one-shot model and an iterative, time-staged model. Under the one-shot model, packed jobs are answered and the rest discarded, and we aim to maximize the total realized QoI. Under the staged model, the system repeatedly packs and offers answers until all jobs are complete; here we aim to maximize time-weighted QoI and minimize completion time. We evaluate the performance of different instantiations of our system through thousands of procedurally-generated simulations.
James Edwards 0002, Rebecca J. Passonneau, Taylor Cassidy, Thomas La Porta
ICCCN2
2017 Distractor Generation with Generative Adversarial Nets for Automatically Creating Fill-in-the-blank Questions
abstract
Distractor generation is a crucial step for fill-in-the-blank question generation. We propose a generative model learned from training generative adversarial nets (GANs) to create useful distractors. Our method utilizes only context information and does not use the correct answer, which is completely different from previous Ontology-based or similarity-based approaches. Trained on the Wikipedia corpus, the proposed model is able to predict Wiki entities as distractors. Our method is evaluated on two biology question datasets collected from Wikipedia and actual college-level exams. Experimental results show that our context-based method achieves comparable performance to a frequently used word2vec-based method for the Wiki dataset. In addition, we propose a second-stage learner to combine the strengths of the two methods, which further improves the performance on both datasets, with 51.7% and 48.4% of generated distractors being acceptable.
Chen Liang 0001, Xiao Yang 0004, Drew Wham, Bart Pursel, Rebecca J. Passonneau, C. Lee Giles
K-CAP5
2016 PEAK: Pyramid Evaluation via Automated Knowledge Extraction
abstract
Evaluating the selection of content in a summary is important both for human-written summaries, which can be a useful pedagogical tool for reading and writing skills, and machine-generated summaries, which are increasingly being deployed in information management. The pyramid method assesses a summary by aggregating content units from the summaries of a wise crowd (a form of crowdsourcing). It has proven highly reliable but has largely depended on manual annotation. We propose PEAK, the first method to automatically assess summary content using the pyramid method that also generates the pyramid content models. PEAK relies on open information extraction and graph algorithms. The resulting scores correlate well with manually derived pyramid scores on both human and machine summaries, opening up the possibility of wide-spread use in numerous applications.
Qian Yang 0003, Rebecca J. Passonneau, Gerard de Melo
AAAI2
2015 Abstractive Multi-Document Summarization via Phrase Selection and Merging
abstract
Lidong Bing, Piji Li, Yi Liao, Wai Lam, Weiwei Guo, Rebecca Passonneau. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Lidong Bing, Piji Li, Wai Lam, Weiwei Guo, Rebecca J. Passonneau
ACL (1)6
2015 Semantic Similarity Graphs of Mathematics Word Problems: Can Terminology Detection Help?
Rogers Jeffrey Leo John, Rebecca J. Passonneau, Thomas S. McTavish
EDM2
2015 Estimation of Discourse Segmentation Labels from Crowd Data
abstract
For annotation tasks involving independent judgments, probabilistic models have been used to infer ground truth labels from data where a crowd of many annotators labels the same items.Such models have been shown to produce results superior to taking the majority vote, but have not been applied to sequential data.We present two methods to infer ground truth labels from sequential annotations where we assume judgments are not independent, based on the observation that an annotator's segments all tend to be several utterances long.The data consists of crowd labels for annotation of discourse segment boundaries.The new methods extend Hidden Markov Models to relax the independence assumption.The two methods are distinct, so positive labels proposed by both are taken to be ground truth.In addition, results of the models are checked using metrics that test whether an annotator's accuracy relative to a given model remains consistent across different conversations.
Ziheng Huang 0010, Jialu Zhong, Rebecca J. Passonneau
EMNLP3
2014 Biber Redux: Reconsidering Dimensions of Variation in American English
Rebecca J. Passonneau, Nancy Ide, Songqiao Su, Jesse Stuart
COLING1
2014 Annotating the MASC Corpus with BabelNet
Andrea Moro 0001, Roberto Navigli, Francesco Maria Tucci, Rebecca J. Passonneau
LREC4
2014 Aspectual Properties of Conversational Activities
abstract
Segmentation of spoken discourse into distinct conversational activities has been applied to broadcast news, meetings, monologs, and two-party dialogs. This paper considers the aspectual properties of discourse segments, meaning how they transpire in time. Classifiers were con-structed to distinguish between segment boundaries and non-boundaries, where the sizes of utterance spans to represent data instances were varied, and the locations of segment boundaries relative to these in-stances. Classifier performance was better for representations that included the end of one discourse segment combined with the beginning of the next. In addition, classi-fication accuracy was better for segments in which speakers accomplish goals with distinctive start and end points. 1
Rebecca J. Passonneau, Boxuan Guan, Cho Ho Yeung, Yuan Du, Emma Conner
SIGDIAL Conference1
2014 The Benefits of a Model of Annotation
abstract
Standard agreement measures for interannotator reliability are neither necessary nor sufficient to ensure a high quality corpus. In a case study of word sense annotation, conventional methods for evaluating labels from trained annotators are contrasted with a probabilistic annotation model applied to crowdsourced data. The annotation model provides far more information, including a certainty measure for each gold standard label; the crowdsourced data was collected at less than half the cost of the conventional approach.
Rebecca J. Passonneau, Bob Carpenter
Trans. Assoc. Comput. Linguistics1
2014 Erratum: "The Benefits of a Model of Annotation"
abstract
Correction for the URL in footnote 7.
Rebecca J. Passonneau, Bob Carpenter
Trans. Assoc. Comput. Linguistics1
2013 Semantic Frames to Predict Stock Price Movement
Boyi Xie, Rebecca J. Passonneau, Leon Wu, Germán Creamer
ACL (1)2
2013 Open Dialogue Management for Relational Databases
Ben Hixon, Rebecca J. Passonneau
HLT-NAACL2
2012 Toward Habitable Assistance from Spoken Dialogue Systems
Susan L. Epstein, Rebecca J. Passonneau, Tiziana Ligorio, Joshua B. Gordon
IAAI2
2012 Multivariate Assessment of a Repair Program for a New York City Electrical Grid
abstract
We assess the impact of an inspection repair program administered to the secondary electrical grid in New York City. The question of interest is whether repairs reduce the incidence of future events that cause service disruptions ranging from minor to serious ones. A key challenge in defining treatment and control groups in the absence of a randomized experiment involved an inherent bias in selection of electrical structures to be inspected in a given year. To compensate for the bias, we construct separate models for each year of the propensity for a structure to have an inspection repair. The propensity models account for differences across years in the structures that get inspected. To model the treatment outcome, we use a statistical approach based on the additive effects of many weak learners. Our results indicate that inspection repairs are more beneficial earlier in the five-year inspection cycle, which accords with the inherent bias to inspect structures in earlier years that are known to have problems.
Rebecca J. Passonneau, Ashish Tomar, Somnath Sarkar, Haimonti Dutta, Axinia Radeva
ICMLA (2)1
2012 Empirical Comparisons of MASC Word Sense Annotations
Gerard de Melo, Collin F. Baker, Nancy Ide, Rebecca J. Passonneau, Christiane Fellbaum
LREC4
2012 The MASC Word Sense Corpus
Rebecca J. Passonneau, Collin F. Baker, Christiane Fellbaum, Nancy Ide
LREC1
2012 Supervised HDP Using Prior Knowledge
Boyi Xie, Rebecca J. Passonneau
NLDB2
2012 Progressive Clustering with Learned Seeds: An Event Categorization System for Power Grid
Boyi Xie, Rebecca J. Passonneau, Haimonti Dutta, Jing-Yeu Miaw, Axinia Radeva, Ashish Tomar, Cynthia Rudin
SEKE2
2012 Semantic Specificity in Spoken Dialogue Requests
Ben Hixon, Rebecca J. Passonneau, Susan L. Epstein
SIGDIAL Conference2
2011 BUGMINER: Software Reliability Analysis Via Data Mining of Bug Reports
Leon Wu, Boyi Xie, Gail E. Kaiser, Rebecca J. Passonneau
SEKE4
2011 Learning to Balance Grounding Rationales for Dialogue Systems
Joshua B. Gordon, Rebecca J. Passonneau, Susan L. Epstein
SIGDIAL Conference2
2011 PARADISE-style Evaluation of a Human-Human Library Corpus
Rebecca J. Passonneau, Irene Alvarado, Phil Crone, Simon Jerome
SIGDIAL Conference1
2011 Embedded Wizardry
Rebecca J. Passonneau, Susan L. Epstein, Tiziana Ligorio, Joshua B. Gordon
SIGDIAL Conference1
2010 An Evaluation Framework for Natural Language Understanding in Spoken Dialogue Systems
Joshua B. Gordon, Rebecca J. Passonneau
LREC2
2010 Word Sense Annotation of Polysemous Words by Multiple Annotators
Rebecca J. Passonneau, Ansaf Salleb-Aouissi, Vikas Bhardwaj, Nancy Ide
LREC1
2010 Learning about Voice Search for Spoken Dialogue Systems
Rebecca J. Passonneau, Susan L. Epstein, Tiziana Ligorio, Joshua B. Gordon, Pravin Bhutada
HLT-NAACL1
2010 Wizards' dialogue strategies to handle noisy speech recognition
abstract
This paper reports on a novel approach to the design and implementation of a spoken dialogue system. A human subject, or wizard, is presented with input of the sort intended for the dialogue system, and selects from among a set of pre-defined actions. The wizard has access to hypotheses generated by noisy automated speech recognition and queries a database with them using partial matching. During the ambitious study reported here, different wizards exhibited different behaviors, elicited different degrees of caller affinity for the system, and achieved different degrees of accuracy on retrieval of the requested items. Our data illustrates that wizards did not trust automated speech recognition hypotheses when they could not lead to a correct database match, and instead asked informed questions. The wealth of data and the richness of the interactions are a valuable resource with which to model expert wizard behavior.
Tiziana Ligorio, Susan L. Epstein, Rebecca J. Passonneau
SLT3
2010 A process for predicting manhole events in Manhattan
Cynthia Rudin, Rebecca J. Passonneau, Axinia Radeva, Haimonti Dutta, Steve Ierome, Delfina Isaac
Mach. Learn.2
2010 Interlingual annotation of parallel text corpora: a new framework for annotation and evaluation
abstract
Abstract This paper focuses on an important step in the creation of a system of meaning representation and the development of semantically annotated parallel corpora, for use in applications such as machine translation, question answering, text summarization, and information retrieval. The work described below constitutes the first effort of any kind to annotate multiple translations of foreign-language texts with interlingual content. Three levels of representation are introduced: deep syntactic dependencies (IL0), intermediate semantic representations (IL1), and a normalized representation that unifies conversives, nonliteral language, and paraphrase (IL2). The resulting annotated, multilingually induced, parallel corpora will be useful as an empirical basis for a wide range of research, including the development and evaluation of interlingual NLP systems and paraphrase-extraction systems as well as a host of other research and development efforts in theoretical and applied linguistics, foreign language pedagogy, translation studies, and other related disciplines.
Bonnie J. Dorr, Rebecca J. Passonneau, David Farwell, Rebecca Green, Nizar Habash, Stephen Helmreich, Eduard H. Hovy, Lori S. Levin, Keith J. Miller, Teruko Mitamura, Owen Rambow, Advaith Siddharthan
Nat. Lang. Eng.2
2010 Formal and functional assessment of the pyramid method for summary content evaluation
abstract
Abstract Pyramid annotation makes it possible to evaluate quantitatively and qualitatively the content of machine-generated (or human) summaries. Evaluation methods must prove themselves against the same measuring stick – evaluation – as other research methods. First, a formal assessment of pyramid data from the 2003 Document Understanding Conference (DUC) is presented; this addresses whether the form of annotation is reliable and whether score results are consistent across annotators. A combination of interannotator reliability measures of the two manual annotation phases (pyramid creation and annotation of system peer summaries against pyramid models), and significance tests of the similarity of system scores from distinct annotations, produces highly reliable results. The most rigorous test consists of a comparison of peer system rankings produced from two independent sets of pyramid and peer annotations, which produce essentially the same rankings. Three years of DUC data (2003, 2005, 2006) are used to assess the reliability of the method across distinct evaluation settings: distinct systems, document sets, summary lengths, and numbers of model summaries. This functional assessment addresses the method's ability to discriminate systems across years. Results indicate that the statistical power of the method is more than sufficient to identify statistically significant differences among systems, and that the statistical power varies little across the 3 years.
Rebecca J. Passonneau
Nat. Lang. Eng.1
2009 Semantic Clustering for a Functional Text Classification Task
Tom Lippincott, Rebecca J. Passonneau
CICLing2
2009 Reducing Noise in Labels and Features for a Real World Dataset: Application of NLP Corpus Annotation Methods
Rebecca J. Passonneau, Cynthia Rudin, Axinia Radeva, Zhi An Liu
CICLing1
2009 Report Cards for Manholes: Eliciting Expert Feedback for a Learning Task
abstract
We present a manhole profiling tool, developed as part of the Columbia/Con Edison machine learning project on manhole event prediction, and discuss its role in evaluating our machine learning model in three important ways: elimination of outliers, elimination of falsely predictive features, and assessment of the quality of the model. The model produces a ranked list of tens of thousands of manholes in Manhattan, where the ranking criterion is vulnerability to serious events such as fires, explosions and smoking manholes. Con Edison set two goals for the model, namely accuracy and intuitiveness, and this tool made it possible for us to address both of these goals. The tool automatically assembles a "report card" or "profile" highlighting data associated with a given manhole. Prior to the processing work that underlies the profiling tool, case studies of a single manhole took several days and resulted in an incomplete study; locating manholes such as those we present in this work would have been extremely difficult. The model is currently assisting Con Edison in determining repair priorities for the secondary electrical grid.
Axinia Radeva, Cynthia Rudin, Rebecca J. Passonneau, Delfina Isaac
ICMLA3
2009 Contrasting the Interaction Structure of an Email and a Telephone Corpus: A Machine Learning Approach to Annotation of Dialogue Function Units
Rebecca J. Passonneau, Owen Rambow
SIGDIAL Conference2
2009 Computational linguistics for metadata building (CLiMB): using text mining for the automatic identification, categorization, and disambiguation of subject terms for image metadata
Judith L. Klavans, Carolyn Sheffield, Eileen G. Abels, Jimmy Lin, Rebecca J. Passonneau, Tandeep Sidhu, Dagobert Soergel
Multim. Tools Appl.5
2008 MASC: the Manually Annotated Sub-Corpus of American English
Nancy Ide, Collin F. Baker, Christiane Fellbaum, Charles J. Fillmore, Rebecca J. Passonneau
LREC5
2008 Relation between Agreement Measures on Human Labeling and Machine Learning Performance: Results from an Art History Domain
Rebecca J. Passonneau, Tom Lippincott, Tae Yano, Judith L. Klavans
LREC1
2006 Measuring Agreement on Set-valued Items (MASI) for Semantic and Pragmatic Annotation
Rebecca J. Passonneau
LREC1
2006 CLiMB ToolKit: A Case Study of Iterative Evaluation in a Multidisciplinary Project
Rebecca J. Passonneau, Roberta Blitz, David K. Elson, Angela Giral, Judith L. Klavans
LREC1
2006 Inter-annotator Agreement on a Multilingual Semantic Annotation Task
Rebecca J. Passonneau, Nizar Habash, Owen Rambow
LREC1
2005 Do summaries help?
abstract
We describe a task-based evaluation to determine whether multi-document summaries measurably improve user performance whe using online news browsing systems for directed research. We evaluated the multi-document summaries generated by Newsblaster, a robust news browsing system that clusters online news articles and summarizes multiple articles on each event. Four groups of subjects were asked to perform the same time-restricted fact-gathering tasks, reading news under different conditions: no summaries at all, single sentence summaries drawn from one of the articles, Newsblaster multi-document summaries, and human summaries. Our results show that, in comparison to source documents only, the quality of reports assembled using Newsblaster summaries was significantly better and user satisfaction was higher with both Newsblaster and human summaries.
Kathy McKeown, Rebecca J. Passonneau, David K. Elson, Ani Nenkova, Julia Hirschberg
SIGIR2
2004 Computing Reliability for Coreference Annotation
Rebecca J. Passonneau
LREC1
2004 Evaluating Content Selection in Summarization: The Pyramid Method
Ani Nenkova, Rebecca J. Passonneau
HLT-NAACL2
2002 DARPA communicator evaluation: progress from 2000 to 2001
abstract
This paper describes the evaluation methodology and results of the DARPA Communicator spoken dialog system evaluation experiments in 2000 and 2001. Nine spoken dialog systems in the travel planning domain participated in the experiments resulting in a total corpus of 1904 dialogs. We describe and compare the experimental design of the 2000 and 2001 DARPA evaluations. We describe how we established a performance baseline in 2001 for complex tasks. We present our overall approach to data collection, the metrics collected, and the application of PARADISE to these data sets. We compare the results we achieved in 2000 for a number of core metrics with those for 2001. These results demonstrate large performance improvements from 2000 to 2001 and show that the Communicator program goal of conversational interaction for complex tasks has been achieved.
Marilyn A. Walker, Alexander I. Rudnicky, John S. Aberdeen, Elizabeth Owen Bratt, John S. Garofolo, Helen Hastie, Audrey N. Le, Bryan L. Pellom, Alexandros Potamianos, Rebecca J. Passonneau, Rashmi Prasad, Salim Roukos, Gregory A. Sanders, Stephanie Seneff, David Stallard
INTERSPEECH10
2002 DARPA communicator: cross-system results for the 2001 evaluation
abstract
This paper describes the evaluation methodology and results of the 2001 DARPA Communicator evaluation. The experiment spanned 6 months of 2001 and involved eight DARPA Communicator systems in the travel planning domain. It resulted in a corpus of 1242 dialogs which include many more dialogues for complex tasks than the 2000 evaluation. We describe the experimental design, the approach to data collection, and the results. We compare the results by the type of travel plan and by system. The results demonstrate some large differences across sites and show that the complex trips are clearly more difficult.
Marilyn A. Walker, Alexander I. Rudnicky, Rashmi Prasad, John S. Aberdeen, Elizabeth Owen Bratt, John S. Garofolo, Helen Hastie, Audrey N. Le, Bryan L. Pellom, Alexandros Potamianos, Rebecca J. Passonneau, Salim Roukos, Gregory A. Sanders, Stephanie Seneff, David Stallard
INTERSPEECH11
2001 Quantitative and Qualitative Evaluation of Darpa Communicator Spoken Dialogue Systems
abstract
This paper describes the application of the PARADISE evaluation framework to the corpus of 662 human-computer dialogues collected in the June 2000 Darpa Communicator data collection. We describe results based on the standard logfile metrics as well as results based on additional qualitative metrics derived using the DATE dialogue act tagging scheme. We show that performance models derived via using the standard metrics can account for 37% of the variance in user satisfaction, and that the addition of DATE metrics improved the models by an absolute 5%.
Marilyn A. Walker, Rebecca J. Passonneau, Julie E. Boland
ACL2
1997 Discourse Segmentation by Human and Automated Means
Rebecca J. Passonneau, Diane J. Litman
Comput. Linguistics1
1995 Combining Multiple Knowledge Sources for Discourse Segmentation
abstract
We predict discourse segment boundaries from linguistic features of utterances, using a corpus of spoken narratives as data. We present two methods for developing segmentation algorithms from training data: hand tuning and machine learning. When multiple types of features are used, results approach human performance on an independent test set (both methods), and using cross-validation (machine learning).
Diane J. Litman, Rebecca J. Passonneau
ACL2
1995 Integrating Gricean and Attentional Constraints
Rebecca J. Passonneau
IJCAI1
1993 Temporal Centering
abstract
We present a semantic and pragmatic account of the anaphoric properties of past and perfect that improves on previous work by integrating discourse structure, aspectual type, surface structure and commonsense knowledge. A novel aspect of our account is that we distinguish between two kinds of temporal intervals in the interpretation of temporal operators --- discourse reference intervals and event intervals. This distinction makes it possible to develop an analogy between centering and temporal centering, which operates on discourse reference intervals. Our temporal property-sharing principle is a defeasible inference rule on the logical form. Along with lexical and causal reasoning, it plays a role in incrementally resolving underspecified aspects of the event structure representation of an utterance against the current context.
Megumi Kameyama, Rebecca J. Passonneau, Massimo Poesio
ACL2
1993 Intention-Based Segmentation: Human Reliability and Correlation with Linguistic Cues
abstract
Certain spans of utterances in a discourse, referred to here as segments, are widely assumed to form coherent units. Further, the segmental structure of discourse has been claimed to constrain and be constrained by many phenomena. However, there is weak consensus on the nature of segments and the criteria for recognizing or generating them. We present quantitative results of a two part study using a corpus of spontaneous, narrative monologues. The first part evaluates the statistical reliability of human segmentation of our corpus, where speaker intention is the segmentation criterion. We then use the subjects' segmentations to evaluate the correlation of discourse segmentation with three linguistic cues (referential noun phrases, cue words, and pauses), using information retrieval metrics.
Rebecca J. Passonneau, Diane J. Litman
ACL1
1993 The KERNEL Text Understanding System
Martha Palmer, Rebecca J. Passonneau, Carl Weir, Tim Finin
Artif. Intell.2
1991 Some Facts about Centers, Indexicals, and Demonstratives
abstract
Certain pronoun contexts are argued to establish a local center (LC), i.e., a conventionalized indexical similar to 1st/2nd pers. pronouns. Demonstrative pronouns, also indexicals, are shown to access entities that are not LCs because they lack discourse relevance or because they are not yet in the universe of discourse.
Rebecca J. Passonneau
ACL1
1990 Integrating Natural Language Processing and Knowledge Based Processing
Rebecca J. Passonneau, Carl Weir, Tim Finin, Martha Palmer
AAAI1
1989 Getting at Discourse Referents
abstract
I examine how discourse anaphoric uses of the definite pronoun it contrast with similar uses of the demonstrative pronoun that. Their distinct contexts of use are characterized in terms of two contextual features---persistence of grammatical subject and persistence of grammatical form---which together demonstrate very clearly the interrelation among lexical choice, grammatical choices and the dimension of time in signalling the dynamic attentional state of a discourse.
Rebecca J. Passonneau
ACL1
1988 Sentence Fragments Regular Structures
abstract
This paper describes an analysis of telegraphic fragments as regular structures (not errors) handled by minimal extensions to a system designed for processing the standard language. The modular approach which has been implemented in the Unisys natural language processing system PUNDIT is based on a division of labor in which syntax regulates the occurrence and distribution of elided elements, and semantics and pragmatics use the system's standard mechanisms to interpret them.
Marcia C. Linebarger, Deborah A. Dahl, Lynette Hirschman, Rebecca J. Passonneau
ACL4
1988 A Computational Model of the Semantics of Tense and Aspect
Rebecca J. Passonneau
Comput. Linguistics1
1987 Nominalizations in PUNDIT
abstract
This paper describes the treatment of nominalizations in the PUNDIT text processing system. A single semantic definition is used for both nominalizations and the verbs to which they are related, with the same semantic roles, decompositions, and selectional restrictions on the semantic roles. However, because syntactically nominalizations are noun phrases, the processing which produces the semantic representation is different in several respects from that used for clauses. (1) The rules relating the syntactic positions of the constituents to the roles that they can fill are different. (2) The fact that nominalizations are untensed while clauses normally are tensed means that an alternative treatment of time is required for nominalizations. (3) Because none of the arguments of a nominalization is syntactically obligatory, some differences in the control of the filling of roles are required, in particular, roles can be filled as part of reference resolution for the nominalization. The differences in processing are captured by allowing the semantic interpreter to operate in two different modes, one for clauses, and one for nominalizations. Because many nominalizations are noun-noun compounds, this approach also addresses this problem, by suggesting a way of dealing with one relatively tractable subset of noun-noun compounds.
Deborah A. Dahl, Martha Palmer, Rebecca J. Passonneau
ACL3
1987 Situations and Intervals
abstract
The PUNDIT system processes natural language descriptions of situations and the intervals over which they hold using an algorithm that integrates aspect and tense logic. It analyses the tense and aspect of the main verb to generate representations of three types of situations---states, processes and events---and to locate the situations with respect to the time at which the text was produced. Each situation type has a distinct temporal structure, represented in terms of one or more intervals. Further, every interval has two features whose different values capture the aspectual differences between the three different situation types. Capturing these differences makes it possible to represent very precisely the times for which predications are asserted to hold.
Rebecca J. Passonneau
ACL1