Bailey Kuehl

dblp:290/1424 · also Bailey E. Kuehl · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
10since 2021 · last 2024
0009-0005-9704-4791ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Information extraction and text analysis · 51% Language models and text generation · 40% Knowledge representation and reasoning · 9%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Medical and health informatics · 67% Bioinformatics and computational biology · 33%

Topics — the 15 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Empirical software engineering
mining software repositories
0.812024
ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer Reviews · ACL (1) 2024
Natural language and speech › Information extraction and text analysis
entity linking
0.712023
S2abEL: A Dataset for Entity Linking from Scientific Tables · EMNLP 2023
Information retrieval
evaluation
0.712023
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023
Information retrieval › text summarization
multi-document summarization
0.712023
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023
Information retrieval › text summarization
summarization evaluation
0.712023
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023
Information retrieval
text summarization
0.712023
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023
Natural language and speech › Information extraction and text analysis
fact-checking
0.612022
Generating Scientific Claims for Zero-Shot Scientific Fact Checking · ACL (1) 2022
Information retrieval › document retrieval › domain-specific retrieval
scientific literature search
0.612022
A Search Engine for Discovery of Scientific Challenges and Directions · AAAI 2022
Information retrieval
search engines
0.612022
A Search Engine for Discovery of Scientific Challenges and Directions · AAAI 2022
Natural language and speech › Language models and text generation › text summarization
multi-document summarization
0.512021
MS\^2: Multi-Document Summarization of Medical Studies · EMNLP (1) 2021
Natural language and speech › Language models and text generation
text summarization
0.512021
MS\^2: Multi-Document Summarization of Medical Studies · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis
scientific text
0.212024
ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer Reviews · ACL (1) 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
knowledge base construction
0.212023
S2abEL: A Dataset for Entity Linking from Scientific Tables · EMNLP 2023
Medical and health informatics › clinical text processing
clinical text summarization
0.212023
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023
Bioinformatics and computational biology
biomedical knowledge discovery
0.212022
A Search Engine for Discovery of Scientific Challenges and Directions · AAAI 2022

Methods — techniques the papers use, named apart from their topics

text classification · 1.7expert annotation · 1.7corpus construction · 1.5automated metrics · 1.3ROUGE · 1.3summarization evaluation metric · 1.0BART · 1.0neural entity linking · 0.7zero-shot learning · 0.6language model generation · 0.6
YearPublicationVenuePosition
2024 ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer Reviews
abstract
Mike D’Arcy, Alexis Ross, Erin Bransom, Bailey Kuehl, Jonathan Bragg, Tom Hope, Doug Downey. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Mike D'Arcy, Alexis Ross, Erin Bransom, Bailey Kuehl, Jonathan Bragg, Tom Hope, Doug Downey
ACL (1)4
2023 Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations
abstract
-gram similarity metrics such as ROUGE. Better automated evaluation metrics are needed, but few resources exist to assess metrics when they are proposed. Therefore, we introduce a dataset of human-assessed summary quality facets and pairwise preferences to encourage and support the development of better automated evaluation methods for literature review MDS. We take advantage of community submissions to the Multi-document Summarization for Literature Review (MSLR) shared task to compile a diverse and representative sample of generated summaries. We analyze how automated summarization evaluation metrics correlate with lexical features of generated summaries, to other automated metrics including several we propose in this work, and to aspects of human-assessed summary quality. We find that not only do automated metrics fail to capture aspects of quality as assessed by humans, in many cases the system rankings produced by these metrics are anti-correlated with rankings according to human annotators.
Lucy Lu Wang, Yulia Otmakhova 0001, Jay DeYoung, Hung-Thinh Truong, Bailey Kuehl, Erin Bransom, Byron C. Wallace
ACL (1)5
2023 LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization
abstract
Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, Kyle Lo. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, Kyle Lo
EACL3
2023 S2abEL: A Dataset for Entity Linking from Scientific Tables
abstract
Entity linking (EL) is the task of linking a textual mention to its corresponding entry in a knowledge base, and is critical for many knowledge-intensive NLP applications.When applied to tables in scientific papers, EL is a step toward large-scale scientific knowledge bases that could enable advanced scientific question answering and analytics.We present the first dataset for EL in scientific tables.EL for scientific tables is especially challenging because scientific knowledge bases can be very incomplete, and disambiguating table mentions typically requires understanding the paper's text in addition to the table.Our dataset, Scientific Table Entity Linking (S2abEL), focuses on EL in machine learning results tables and includes hand-labeled cell types, attributed sources, and entity links from the PaperswithCode taxonomy for 8,429 cells from 732 tables.We introduce a neural baseline method designed for EL on scientific tables containing many out-of-knowledge-base mentions, and show that it significantly outperforms a state-of-the-art generic table EL method.The best baselines fall below human performance, and our analysis highlights avenues for improvement.
Yuze Lou, Bailey Kuehl, Erin Bransom, Sergey Feldman, Aakanksha Naik, Doug Downey
EMNLP2
2022 A Search Engine for Discovery of Scientific Challenges and Directions
abstract
Keeping track of scientific challenges, advances and emerging directions is a fundamental part of research. However, researchers face a flood of papers that hinders discovery of important knowledge. In biomedicine, this directly impacts human lives. To address this problem, we present a novel task of extraction and search of scientific challenges and directions, to facilitate rapid knowledge discovery. We construct and release an expert-annotated corpus of texts sampled from full-length papers, labeled with novel semantic categories that generalize across many types of challenges and directions. We focus on a large corpus of interdisciplinary work relating to the COVID-19 pandemic, ranging from biomedicine to areas such as AI and economics. We apply a model trained on our data to identify challenges and directions across the corpus and build a dedicated search engine. In experiments with 19 researchers and clinicians using our system, we outperform a popular scientific search engine in assisting knowledge discovery. Finally, we show that models trained on our resource generalize to the wider biomedical domain and to AI papers, highlighting its broad utility. We make our data, model and search engine publicly available.
Dan Lahav, Jon Saad-Falcon, Bailey Kuehl, Sophie Johnson, Sravanthi Parasa, Noam Shomron, Polo Chau, Diyi Yang, Eric Horvitz, Daniel S. Weld, Tom Hope
AAAI3
2022 Generating Scientific Claims for Zero-Shot Scientific Fact Checking
abstract
Dustin Wright, David Wadden, Kyle Lo, Bailey Kuehl, Arman Cohan, Isabelle Augenstein, Lucy Wang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Dustin Wright 0001, Dave Wadden, Kyle Lo, Bailey Kuehl, Arman Cohan, Isabelle Augenstein, Lucy Lu Wang
ACL (1)4
2022 MultiCite: Modeling realistic citations requires moving beyond the single-sentence single-label setting
abstract
Anne Lauscher, Brandon Ko, Bailey Kuehl, Sophie Johnson, Arman Cohan, David Jurgens, Kyle Lo. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Anne Lauscher, Brandon Ko, Bailey Kuehl, Sophie Johnson, Arman Cohan, David Jurgens, Kyle Lo
NAACL-HLT3
2022 VILA: Improving Structured Content Extraction from Scientific PDFs Using Visual Layout Groups
abstract
Abstract Accurately extracting structured content from PDFs is a critical first step for NLP over scientific papers. Recent work has improved extraction accuracy by incorporating elementary layout information, for example, each token’s 2D position on the page, into language model pretraining. We introduce new methods that explicitly model VIsual LAyout (VILA) groups, that is, text lines or text blocks, to further improve performance. In our I-VILA approach, we show that simply inserting special tokens denoting layout group boundaries into model inputs can lead to a 1.9% Macro F1 improvement in token classification. In the H-VILA approach, we show that hierarchical encoding of layout-groups can result in up to 47% inference time reduction with less than 0.8% Macro F1 loss. Unlike prior layout-aware approaches, our methods do not require expensive additional pretraining, only fine-tuning, which we show can reduce training cost by up to 95%. Experiments are conducted on a newly curated evaluation suite, S2-VLUE, that unifies existing automatically labeled datasets and includes a new dataset of manual annotations covering diverse papers from 19 scientific disciplines. Pre-trained weights, benchmark datasets, and source code are available at https://github.com/allenai/VILA.
Shannon Shen 0001, Kyle Lo, Lucy Lu Wang, Bailey Kuehl, Daniel S. Weld, Doug Downey
Trans. Assoc. Comput. Linguistics4
2021 SciA11y: Converting Scientific Papers to Accessible HTML
abstract
We present SciA11y, a system that renders inaccessible scientific paper PDFs into HTML. SciA11y uses machine learning models to extract and understand the content of scientific PDFs, and reorganizes the resulting paper components into a form that better supports skimming and scanning for blind and low vision (BLV) readers. SciA11y adds navigation features such as tagged headings, a table of contents, and bidirectional links between inline citations and references, which allow readers to resolve citations without losing their context. A set of 1.5 million open access papers are processed and available at https://scia11y.org/. This system is a first step in addressing scientific PDF accessibility, and may significantly improve the experience of paper reading for BLV users.
Lucy Lu Wang, Isabel Cachola, Jonathan Bragg, Evie Yu-Yen Cheng, Chelsea Haupt, Matt Latzke, Bailey Kuehl, Madeleine van Zuylen, Linda Wagner, Daniel S. Weld
ASSETS7
2021 MS\^2: Multi-Document Summarization of Medical Studies
abstract
To assess the effectiveness of any medical intervention, researchers must conduct a timeintensive and manual literature review.NLP systems can help to automate or assist in parts of this expensive process.In support of this goal, we release MSˆ2 (Multi-Document Summarization of Medical Studies), a dataset of over 470k documents and 20K summaries derived from the scientific literature.This dataset facilitates the development of systems that can assess and aggregate contradictory evidence across multiple studies, and is the first large-scale, publicly available multi-document summarization dataset in the biomedical domain.We experiment with a summarization system based on BART, with promising early results, though significant work remains to achieve higher summarization quality.We formulate our summarization inputs and targets in both free text and structured forms and modify a recently proposed metric to assess the quality of our system's generated summaries.
Jay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl, Lucy Lu Wang
EMNLP (1)4