Jonathan K. Kummerfeld

dblp:84/9011 · DBLP profile ↗
← Back
38ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0001-5030-3016ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 6 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Your Students Don't Use LLMs Like You Wish They Did
abstract
Educational NLP systems are typically evaluated using engagement metrics and satisfaction surveys, which are at best a proxy for meeting pedagogical goals.We introduce six computational metrics for automated evaluation of pedagogical alignment in student-AI dialogue.We validate our metrics through analysis of 12,650 messages across 500 conversations from four courses.Using our metrics, we identify a fundamental misalignment: educators design conversational tutors for sustained learning dialogue, but students mainly use them for answerextraction.Deployment context is the strongest predictor of usage patterns, outweighing student preference or system design: when AI tools are optional, usage concentrates around deadlines; when integrated into course structure, students ask for solutions to verbatim assignment questions.Whole-dialogue evaluation misses these turn-by-turn patterns.Our metrics will enable researchers building educational dialogue systems to measure whether they are achieving their pedagogical goals.
Sebastian Kobler, Matthew Clemson, Angela Sun, Jonathan K. Kummerfeld
ACL (1)4
2025 Less is More: Explainable and Efficient ICD Code Prediction with Clinical Entities
abstract
Clinical coding, assigning standardized codes to medical notes, is critical for epidemiological research, hospital planning, and reimbursement.Neural coding models generally process entire discharge summaries, which are often lengthy and contain information that is not relevant to coding.We propose an approach that combines Named Entity Recognition (NER) and Assertion Classification (AC) to filter for clinically important content before supervised code prediction.On MIMIC-IV, a standard evaluation dataset, our approach achieves near-equivalent performance to a state-of-the-art full-text baseline while using only 22% of the content and reducing training time by over half.Additionally, mapping model attention to complete entity spans yields coherent, clinically meaningful explanations, capturing coding-relevant modifiers such as acuity and laterality.We release a newly annotated NER+AC dataset for MIMIC-IV, designed specifically for ICD coding.Our entitycentric approach lays the foundation for more transparent and cost-effective assisted coding.
James C. Douglas, Yidong Gan, Ben Hachey, Jonathan K. Kummerfeld
ACL (1)4
2025 Aligning AI Research with the Needs of Clinical Coding Workflows: Eight Recommendations Based on US Data Analysis and Critical Review
abstract
Clinical coding is crucial for healthcare billing and data analysis.Manual clinical coding is labour-intensive and error-prone, which has motivated research towards full automation of the process.However, our analysis, based on US English electronic health records and automated coding research using these records, shows that widely used evaluation methods are not aligned with real clinical contexts.For example, evaluations that focus on the top 50 most common codes are an oversimplification, as there are thousands of codes used in practice.This position paper aims to align AI coding research more closely with practical challenges of clinical coding.Based on our analysis, we offer eight specific recommendations, suggesting ways to improve current evaluation methods.Additionally, we propose new AI-based methods beyond automated coding, suggesting alternative approaches to assist clinical coders in their workflows.
Yidong Gan, Maciej Rybinski, Ben Hachey, Jonathan K. Kummerfeld
ACL (1)4
2025 AbstractExplorer: Leveraging Structure-Mapping Theory to Enhance Comparative Close Reading at Scale
Ziwei Gu, Joyce Zhou, Ning-Er (Nina) Lei, Jonathan K. Kummerfeld, Mahmood Jasim, Narges Mahyar, Elena L. Glassman
UIST4
2024 More Victories, Less Cooperation: Assessing Cicero's Diplomacy Play
abstract
Wichayaporn Wongkamjan, Feng Gu, Yanze Wang, Ulf Hermjakob, Jonathan May, Brandon M. Stewart, Jonathan K. Kummerfeld, Denis Peskoff, Jordan Lee Boyd-Graber. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wichayaporn Wongkamjan, Ulf Hermjakob, Jonathan May, Brandon M. Stewart, Jonathan K. Kummerfeld, Denis Peskoff, Jordan L. Boyd-Graber
ACL (1)7
2024 Supporting Sensemaking of Large Language Model Outputs at Scale
abstract
Large language models (LLMs) are capable of generating multiple responses to a single prompt, yet little effort has been expended to help end-users or system designers make use of this capability. In this paper, we explore how to present many LLM responses at once. We design five features, which include both pre-existing and novel methods for computing similarities and differences across textual documents, as well as how to render their outputs. We report on a controlled user study (n=24) and eight case studies evaluating these features and how they support users in different tasks. We find that the features support a wide variety of sensemaking tasks and even make tasks tractable that our participants previously considered to be too difficult to attempt. Finally, we present design guidelines to inform future explorations of new LLM interfaces.
Katy Ilonka Gero, Chelse Swoopes, Ziwei Gu, Jonathan K. Kummerfeld, Elena L. Glassman
CHI4
2024 An AI-Resilient Text Rendering Technique for Reading and Skimming Documents
abstract
Readers find text difficult to consume for many reasons. Summarization can address some of these difficulties, but introduce others, such as omitting, misrepresenting, or hallucinating information, which can be hard for a reader to notice. One approach to addressing this problem is to instead modify how the original text is rendered to make important information more salient. We introduce Grammar-Preserving Text Saliency Modulation (GP-TSM), a text rendering method with a novel means of identifying what to de-emphasize. Specifically, GP-TSM uses a recursive sentence compression method to identify successive levels of detail beyond the core meaning of a passage, which are de-emphasized by rendering words in successively lighter but still legible gray text. In a lab study (n=18), participants preferred GP-TSM over pre-existing word-level text rendering methods and were able to answer GRE reading comprehension questions more efficiently.
Ziwei Gu, Ian Arawjo, Kenneth Li 0002, Jonathan K. Kummerfeld, Elena L. Glassman
CHI4
2024 A Comparative Multidimensional Analysis of Empathetic Systems
abstract
Recently, empathetic dialogue systems have received significant attention.While some researchers have noted limitations, e.g., that these systems tend to generate generic utterances, no study has systematically verified these issues.We survey 21 systems, asking what progress has been made on the task.We observe multiple limitations of current evaluation procedures.Most critically, studies tend to rely on a single non-reproducible empathy score, which inadequately reflects the multidimensional nature of empathy.To better understand the differences between systems, we comprehensively analyze each system with automated methods that are grounded in a variety of aspects of empathy.We find that recent systems lack three important aspects of empathy: specificity, reflection levels, and diversity.Based on our results, we discuss problematic behaviors that may have gone undetected in prior evaluations, and offer guidance for developing future systems. 1
Andrew Lee 0001, Jonathan K. Kummerfeld, Lawrence C. An, Rada Mihalcea
EACL (1)2
2024 Do Text-to-Vis Benchmarks Test Real Use of Visualisations?
abstract
Large language models are able to generate code for visualisations in response to simple user requests.This is a useful application and an appealing one for NLP research because plots of data provide grounding for language.However, there are relatively few benchmarks, and those that exist may not be representative of what users do in practice.This paper investigates whether benchmarks reflect realworld use through an empirical study comparing benchmark datasets with code from public repositories.Our findings reveal a substantial gap, with evaluations not testing the same distribution of chart types, attributes, and actions as real-world examples.One dataset is representative, but requires extensive modification to become a practical end-to-end benchmark.This shows that new benchmarks are needed to support the development of systems that truly address users' visualisation needs.These observations will guide future data creation, highlighting which features hold genuine significance for users.
Hy Nguyen, Xuefei He, Andrew Reeson, Cécile Paris, Josiah Poon, Jonathan K. Kummerfeld
EMNLP6
2024 A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
abstract
While alignment algorithms are commonly used to tune pre-trained language models towards user preferences, we lack explanations for the underlying mechanisms in which models become ``aligned'', thus making it difficult to explain phenomena like jailbreaks. In this work we study a popular algorithm, direct preference optimization (DPO), and the mechanisms by which it reduces toxicity. Namely, we first study how toxicity is represented and elicited in pre-trained language models (GPT2-medium, Llama2-7b). We then apply DPO with a carefully crafted pairwise dataset to reduce toxicity. We examine how the resulting models avert toxic outputs, and find that capabilities learned from pre-training are not removed, but rather bypassed. We use this insight to demonstrate a simple method to un-align the models, reverting them back to their toxic behavior.
Andrew Lee 0001, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, Rada Mihalcea
ICML5
2024 SQLucid: Grounding Natural Language Database Queries with Interactive Explanations
abstract
Though recent advances in machine learning have led to significant improvements in natural language interfaces for databases, the accuracy and reliability of these systems remain limited, especially in high-stakes domains. This paper introduces SQLucid, a novel user interface that bridges the gap between non-expert users and complex database querying processes. SQLucid addresses existing limitations by integrating visual correspondence, intermediate query results, and editable step-by-step SQL explanations in natural language to facilitate user understanding and engagement. This unique blend of features empowers users to understand and refine SQL queries easily and precisely. Two user studies and one quantitative experiment were conducted to validate SQLucid’s effectiveness, showing significant improvement in task completion accuracy and user confidence compared to existing interfaces. Our code is available at https://github.com/magic-YuanTian/SQLucid.
Jonathan K. Kummerfeld, Toby Jia-Jun Li, Tianyi Zhang 0001
UIST2
2023 Empathy Identification Systems are not Accurately Accounting for Context
abstract
Understanding empathy in text dialogue data is a difficult, yet critical, skill for effective human-machine interaction.In this work, we ask whether systems are making meaningful progress on this challenge.We consider a simple model that checks if an input utterance is similar to a small set of empathetic examples.Crucially, the model does not look at what the utterance is a response to, i.e., the dialogue context.This model performs comparably to prior work on standard benchmarks and even outperforms state-of-the-art models for empathetic rationale extraction by 16.7 points on T-F1 and 4.3 on IOU-F1.This indicates that current systems rely on the surface form of the response, rather than whether it is suitable in context.To confirm this, we create examples with dialogue contexts that change the interpretation of the response and show that current systems continue to label utterances as empathetic.We discuss the implications of our findings, including improvements for empathetic benchmarks and how our model can be an informative baseline.
Andrew Lee 0001, Jonathan K. Kummerfeld, Lawrence C. An, Rada Mihalcea
EACL2
2023 Interactive Text-to-SQL Generation via Editable Step-by-Step Explanations
abstract
Relational databases play an important role in business, science, and more.However, many users cannot fully unleash the analytical power of relational databases, because they are not familiar with database languages such as SQL.Many techniques have been proposed to automatically generate SQL from natural language, but they suffer from two issues: (1) they still make many mistakes, particularly for complex queries, and (2) they do not provide a flexible way for non-expert users to validate and refine incorrect queries.To address these issues, we introduce a new interaction mechanism that allows users to directly edit a stepby-step explanation of a query to fix errors.Our experiments on multiple datasets, as well as a user study with 24 participants, demonstrate that our approach can achieve better performance than multiple SOTA approaches.
Zheng Zhang 0043, Zheng Ning, Toby Jia-Jun Li, Jonathan K. Kummerfeld, Tianyi Zhang 0001
EMNLP5
2022 Leveraging Similar Users for Personalized Language Modeling with Limited Data
abstract
Charles Welch, Chenxi Gu, Jonathan Kummerfeld, Veronica Perez-Rosas, Rada Mihalcea. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Charles Welch, Chenxi Gu, Jonathan K. Kummerfeld, Verónica Pérez-Rosas, Rada Mihalcea
ACL (1)3
2022 Using Paraphrases to Study Properties of Contextual Embeddings
abstract
We use paraphrases as a unique source of data to analyze contextualized embeddings, with a particular focus on BERT.Because paraphrases naturally encode consistent word and phrase semantics, they provide a unique lens for investigating properties of embeddings.Using the Paraphrase Database's alignments, we study words within paraphrases as well as phrase representations.We find that contextual embeddings effectively handle polysemous words, but give synonyms surprisingly different representations in many cases.We confirm previous findings that BERT is sensitive to word order, but find slightly different patterns than prior work in terms of the level of contextualization across BERT's layers.
Laura Burdick, Jonathan K. Kummerfeld, Rada Mihalcea
NAACL-HLT2
2021 Analyzing the Surprising Variability in Word Embedding Stability Across Languages
abstract
Word embeddings are powerful representations that form the foundation of many natural language processing architectures, both in English and in other languages.To gain further insight into word embeddings, we explore their stability (e.g., overlap between the nearest neighbors of a word in different embedding spaces) in diverse languages.We discuss linguistic properties that are related to stability, drawing out insights about correlations with affixing, language gender systems, and other features.This has implications for embedding use, particularly in research that uses them to study language trends.
Laura Burdick, Jonathan K. Kummerfeld, Rada Mihalcea
EMNLP (1)2
2021 Overview of the Eighth Dialog System Technology Challenge: DSTC8
abstract
This paper introduces the Eighth Dialog System Technology Challenge. In line with recent challenges, the eighth edition focuses on applying end-to-end dialog technologies in a pragmatic way for multi-domain task-completion, noetic response selection, audio visual scene-aware dialog, and schema-guided dialog state tracking tasks. This paper describes the task definition, provided datasets, baselines and evaluation set-up for each track. We also summarize the results of the submitted systems to highlight the overall trends of the state-of-the-art technologies for the tasks.
Seokhwan Kim, Michel Galley, R. Chulaka Gunasekara, Adam Atkinson, Baolin Peng, Hannes Schulz, Jianfeng Gao 0001, Jinchao Li, Mahmoud Adada, Minlie Huang, Luis A. Lastras, Jonathan K. Kummerfeld, Walter S. Lasecki, Chiori Hori, Anoop Cherian, Tim K. Marks, Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara
IEEE ACM Trans. Audio Speech Lang. Process.13
2020 Crowdsourced Detection of Emotionally Manipulative Language
abstract
Detecting rhetoric that manipulates readers' emotions requires distinguishing intrinsically emotional content (IEC; e.g., a parent losing a child) from emotionally manipulative language (EML; e.g., using fear-inducing language to spread anti-vaccine propaganda). However, this remains an open classification challenge for both automatic and crowdsourcing approaches. Machine Learning approaches only work in narrow domains where labeled training data is available, and non-expert annotators tend to conflate IEC with EML. We introduce an approach, anchor comparison, that leverages workers' ability to identify and remove instances of EML in text to create a paraphrased "anchor text", which is then used as a comparison point to classify EML in the original content. We evaluate our approach with a dataset of news-style text snippets and show that precision and recall can be tuned for system builders' needs. Our contribution is a crowdsourcing approach that enables non-expert disentanglement of social references from content.
Jordan S. Huffaker, Jonathan K. Kummerfeld, Walter S. Lasecki, Mark S. Ackerman
CHI2
2020 Inconsistencies in Crowdsourced Slot-Filling Annotations: A Typology and Identification Methods
abstract
Slot-filling models in task-driven dialog systems rely on carefully annotated training data.However, annotations by crowd workers are often inconsistent or contain errors.Simple solutions like manually checking annotations or having multiple workers label each sample are expensive and waste effort on samples that are correct.If we can identify inconsistencies, we can focus effort where it is needed.Toward this end, we define six inconsistency types in slot-filling annotations.Using three new noisy crowd-annotated datasets, we show that a wide range of inconsistencies occur and can impact system performance if not addressed.We then introduce automatic methods of identifying inconsistencies.Experiments on our new datasets show that these methods effectively reveal inconsistencies in data, though there is further scope for improvement.
Stefan Larson, Adrian Cheung, Anish Mahendran, Kevin Leach, Jonathan K. Kummerfeld
COLING5
2020 Exploring the Value of Personalized Word Embeddings
abstract
In this paper, we introduce personalized word embeddings, and examine their value for language modeling.We compare the performance of our proposed prediction model when using personalized versus generic word representations, and study how these representations can be leveraged for improved performance.We provide insight into what types of words can be more accurately predicted when building personalized models.Our results show that a subset of words belonging to specific psycholinguistic categories tend to vary more in their representations across users and that combining generic and personalized word embeddings yields the best performance, with a 4.7% relative reduction in perplexity.Additionally, we show that a language model using personalized word embeddings can be effectively used for authorship attribution.
Charles Welch, Jonathan K. Kummerfeld, Verónica Pérez-Rosas, Rada Mihalcea
COLING2
2020 Iterative Feature Mining for Constraint-Based Data Collection to Increase Data Diversity and Model Robustness
abstract
Stefan Larson, Anthony Zheng, Anish Mahendran, Rishi Tekriwal, Adrian Cheung, Eric Guldan, Kevin Leach, Jonathan K. Kummerfeld. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Stefan Larson, Anthony Zheng, Anish Mahendran, Rishi Tekriwal, Adrian Cheung, Eric Guldan, Kevin Leach, Jonathan K. Kummerfeld
EMNLP (1)8
2020 Compositional Demographic Word Embeddings
abstract
Word embeddings are usually derived from corpora containing text from many individuals, thus leading to general purpose representations rather than individually personalized representations.While personalized embeddings can be useful to improve language model performance and other language processing tasks, they can only be computed for people with a large amount of longitudinal data, which is not the case for new users.We propose a new form of personalized word embeddings that use demographic-specific word representations derived compositionally from full or partial demographic information for a user (i.e., gender, age, location, religion).We show that the resulting demographic-aware word representations outperform generic word representations on two tasks for English: language modeling and word associations.We further explore the trade-off between the number of available attributes and their relative effectiveness and discuss the ethical implications of using them.
Charles Welch, Jonathan K. Kummerfeld, Verónica Pérez-Rosas, Rada Mihalcea
EMNLP (1)2
2020 Improving Low Compute Language Modeling with In-Domain Embedding Initialisation
abstract
Many NLP applications, such as biomedical data and technical support, have 10-100 million tokens of in-domain data and limited computational resources for learning from it.How should we train a language model in this scenario?Most language modeling research considers either a small dataset with a closed vocabulary (like the standard 1 million token Penn Treebank), or the whole web with bytepair encoding.We show that for our target setting in English, initialising and freezing input embeddings using in-domain data can improve language model performance by providing a useful representation of rare words, and this pattern holds across several different domains.In the process, we show that the standard convention of tying input and output embeddings does not improve perplexity when initializing with embeddings trained on in-domain data.
Charles Welch, Rada Mihalcea, Jonathan K. Kummerfeld
EMNLP (1)3
2020 Overview of the seventh Dialog System Technology Challenge: DSTC7
Luis Fernando D'Haro, Koichiro Yoshino, Chiori Hori, Tim K. Marks, Lazaros Polymenakos, Jonathan K. Kummerfeld, Michel Galley, Xiang Gao 0011
Comput. Speech Lang.6
2019 A Large-Scale Corpus for Conversation Disentanglement
abstract
Jonathan K. Kummerfeld, Sai R. Gouravajhala, Joseph J. Peper, Vignesh Athreya, Chulaka Gunasekara, Jatin Ganhotra, Siva Sankalp Patel, Lazaros C Polymenakos, Walter Lasecki. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Jonathan K. Kummerfeld, Sai R. Gouravajhala, Joseph Peper, Vignesh Athreya, R. Chulaka Gunasekara, Jatin Ganhotra, Siva Sankalp Patel, Lazaros Polymenakos, Walter S. Lasecki
ACL (1)1
2019 Look Who's Talking: Inferring Speaker Attributes from Personal Longitudinal Dialog
Charles Welch, Verónica Pérez-Rosas, Jonathan K. Kummerfeld, Rada Mihalcea
CICLing (2)3
2019 An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction
abstract
Stefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang, Jason Mars. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Stefan Larson, Anish Mahendran, Joseph Peper, Christopher Clarke, Andrew Lee 0001, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael Laurenzano, Lingjia Tang, Jason Mars
EMNLP/IJCNLP (1)7
2019 No-Press Diplomacy: Modeling Multi-Agent Gameplay
abstract
Diplomacy is a seven-player non-stochastic, non-cooperative game, where agents acquire resources through a mix of teamwork and betrayal. Reliance on trust and coordination makes Diplomacy the first non-cooperative multi-agent benchmark for complex sequential social dilemmas in a rich environment. In this work, we focus on training an agent that learns to play the No Press version of Diplomacy where there is no dedicated communication channel between players. We present DipNet, a neural-network-based policy model for No Press Diplomacy. The model was trained on a new dataset of more than 150,000 human games. Our model is trained by supervised learning (SL) from expert trajectories, which is then used to initialize a reinforcement learning (RL) agent trained through self-play. Both the SL and the RL agent demonstrate state-of-the-art No Press performance by beating popular rule-based bots.
Philip Paquette, Steven Bocco, Max O. Smith, Satya Ortiz-Gagne, Jonathan K. Kummerfeld, Joelle Pineau, Satinder Singh 0001, Aaron C. Courville
NeurIPS6
2018 Improving Text-to-SQL Evaluation Methodology
abstract
Catherine Finegan-Dollak, Jonathan K. Kummerfeld, Li Zhang, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang, Dragomir Radev. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Catherine Finegan-Dollak, Jonathan K. Kummerfeld, Li Zhang 0039, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang 0037, Dragomir R. Radev
ACL (1)2
2018 World Knowledge for Abstract Meaning Representation Parsing
Charles Welch, Jonathan K. Kummerfeld, Song Feng 0002, Rada Mihalcea
LREC2
2018 Factors Influencing the Surprising Instability of Word Embeddings
abstract
Laura Wendlandt, Jonathan K. Kummerfeld, Rada Mihalcea. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Laura Burdick, Jonathan K. Kummerfeld, Rada Mihalcea
NAACL-HLT2
2017 Identifying Products in Online Cybercrime Marketplaces: A Dataset for Fine-grained Domain Adaptation
abstract
Greg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Rebecca Portnoff, Sadia Afroz, Damon McCoy, Kirill Levchenko, Vern Paxson. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2017.
Greg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Rebecca S. Portnoff, Sadia Afroz 0001, Damon McCoy, Kirill Levchenko, Vern Paxson
EMNLP2
2017 Tools for Automated Analysis of Cybercriminal Markets
abstract
Underground forums are widely used by criminals to buy and sell a host of stolen items, datasets, resources, and criminal services. These forums contain important resources for understanding cybercrime. However, the number of forums, their size, and the domain expertise required to understand the markets makes manual exploration of these forums unscalable. In this work, we propose an automated, top-down approach for analyzing underground forums. Our approach uses natural language processing and machine learning to automatically generate high-level information about underground forums, first identifying posts related to transactions, and then extracting products and prices. We also demonstrate, via a pair of case studies, how an analyst can use these automated approaches to investigate other categories of products and transactions. We use eight distinct forums to assess our tools: Antichat, Blackhat World, Carders, Darkode, Hack Forums, Hell, L33tCrew and Nulled. Our automated approach is fast and accurate, achieving over 80% accuracy in detecting post category, product, and prices.
Rebecca S. Portnoff, Sadia Afroz 0001, Greg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Damon McCoy, Kirill Levchenko, Vern Paxson
WWW4
2017 Parsing with Traces: An O(n^4) Algorithm and a Structural Representation
abstract
General treebank analyses are graph structured, but parsers are typically restricted to tree structures for efficiency and modeling reasons. We propose a new representation and algorithm for a class of graph structures that is flexible enough to cover almost all treebank structures, while still admitting efficient learning and inference. In particular, we consider directed, acyclic, one-endpoint-crossing graph structures, which cover most long-distance dislocation, shared argumentation, and similar tree-violating linguistic phenomena. We describe how to convert phrase structure parses, including traces, to our new representation in a reversible manner. Our dynamic program uniquely decomposes structures, is sound and complete, and covers 97.3% of the Penn English Treebank. We also implement a proof-of-concept parser that recovers a range of null elements and trace types.
Jonathan K. Kummerfeld, Daniel Klein 0001
Trans. Assoc. Comput. Linguistics1
2015 An Empirical Analysis of Optimization for Max-Margin NLP
abstract
Despite the convexity of structured maxmargin objectives (Taskar et al., 2004;Tsochantaridis et al., 2004), the many ways to optimize them are not equally effective in practice.We compare a range of online optimization methods over a variety of structured NLP tasks (coreference, summarization, parsing, etc) and find several broad trends.First, margin methods do tend to outperform both likelihood and the perceptron.Second, for max-margin objectives, primal optimization methods are often more robust and progress faster than dual methods.This advantage is most pronounced for tasks with dense or continuous-valued features.Overall, we argue for a particularly simple online primal subgradient descent method that, despite being rarely mentioned in the literature, is surprisingly effective in relation to its alternatives.
Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Daniel Klein 0001
EMNLP1
2013 Error-Driven Analysis of Challenges in Coreference Resolution
abstract
Coreference resolution metrics quantify errors but do not analyze them.Here, we consider an automated method of categorizing errors in the output of a coreference system into intuitive underlying error types.Using this tool, we first compare the error distributions across a large set of systems, then analyze common errors across the top ten systems, empirically characterizing the major unsolved challenges of the coreference resolution task.
Jonathan K. Kummerfeld, Daniel Klein 0001
EMNLP1
2012 Parser Showdown at the Wall Street Corral: An Empirical Investigation of Error Types in Parser Output
Jonathan K. Kummerfeld, David Hall 0006, James R. Curran, Daniel Klein 0001
EMNLP-CoNLL1
2010 Faster Parsing by Supertagger Adaptation
Jonathan K. Kummerfeld, Jessika Roesner, Tim Dawborn, James Haggerty, James R. Curran, Stephen Clark
ACL1