Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jay DeYoung

dblp:136/8673 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Artificial intelligence
2 papers
Language models and text generation · 72% Trustworthy machine learning · 28%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Medical and health informatics · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
evaluation
0.712023
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023
Information retrieval › text summarization
multi-document summarization
0.712023
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023
Information retrieval › text summarization
summarization evaluation
0.712023
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023
Information retrieval
text summarization
0.712023
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023
Natural language and speech › Language models and text generation › text summarization
multi-document summarization
0.512021
MS\^2: Multi-Document Summarization of Medical Studies · EMNLP (1) 2021
Natural language and speech › Language models and text generation
text summarization
0.512021
MS\^2: Multi-Document Summarization of Medical Studies · EMNLP (1) 2021
Machine learning › Trustworthy machine learning
interpretability
0.412020
ERASER: A Benchmark to Evaluate Rationalized NLP Models · ACL 2020
Information retrieval
cross-language information retrieval
0.412019
Neural-Network Lexical Translation for Cross-lingual IR from Text and Speech · SIGIR 2019
Medical and health informatics › clinical text processing
clinical text summarization
0.212023
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023
Information retrieval › document retrieval
spoken document retrieval
0.112019
Neural-Network Lexical Translation for Cross-lingual IR from Text and Speech · SIGIR 2019

Methods — techniques the papers use, named apart from their topics

automated metrics · 1.3ROUGE · 1.3summarization evaluation metric · 1.0BART · 1.0word alignment · 0.4neural network · 0.4character sequence encoding · 0.4
YearPublicationVenuePosition
2024 Do Multi-Document Summarization Models Synthesize?
abstract
Abstract Multi-document summarization entails producing concise synopses of collections of inputs. For some applications, the synopsis should accurately synthesize inputs with respect to a key aspect, e.g., a synopsis of film reviews written about a particular movie should reflect the average critic consensus. As a more consequential example, narrative summaries that accompany biomedical systematic reviews of clinical trial results should accurately summarize the potentially conflicting results from individual trials. In this paper we ask: To what extent do modern multi-document summarization models implicitly perform this sort of synthesis? We run experiments over opinion and evidence synthesis datasets using a suite of summarization models, from fine-tuned transformers to GPT-4. We find that existing models partially perform synthesis, but imperfectly: Even the best performing models are over-sensitive to changes in input ordering and under-sensitive to changes in input compositions (e.g., ratio of positive to negative reviews). We propose a simple, general, effective method for improving model synthesis capabilities by generating an explicitly diverse set of candidate outputs, and then selecting from these the string best aligned with the expected aggregate measure for the inputs, or abstaining when the model produces no good candidate.
Jay DeYoung, Stephanie C. Martinez, Iain James Marshall, Byron C. Wallace
Trans. Assoc. Comput. Linguistics1
2023 Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations
abstract
-gram similarity metrics such as ROUGE. Better automated evaluation metrics are needed, but few resources exist to assess metrics when they are proposed. Therefore, we introduce a dataset of human-assessed summary quality facets and pairwise preferences to encourage and support the development of better automated evaluation methods for literature review MDS. We take advantage of community submissions to the Multi-document Summarization for Literature Review (MSLR) shared task to compile a diverse and representative sample of generated summaries. We analyze how automated summarization evaluation metrics correlate with lexical features of generated summaries, to other automated metrics including several we propose in this work, and to aspects of human-assessed summary quality. We find that not only do automated metrics fail to capture aspects of quality as assessed by humans, in many cases the system rankings produced by these metrics are anti-correlated with rankings according to human annotators.
Lucy Lu Wang, Yulia Otmakhova 0001, Jay DeYoung, Hung-Thinh Truong, Bailey Kuehl, Erin Bransom, Byron C. Wallace
ACL (1)3
2022 Entity Anchored ICD Coding
Jay DeYoung, Han-Chin Shing, Luyang Kong, Christopher Winestock, Chaitanya P. Shivade
AMIA1
2021 MS\^2: Multi-Document Summarization of Medical Studies
abstract
To assess the effectiveness of any medical intervention, researchers must conduct a timeintensive and manual literature review.NLP systems can help to automate or assist in parts of this expensive process.In support of this goal, we release MSˆ2 (Multi-Document Summarization of Medical Studies), a dataset of over 470k documents and 20K summaries derived from the scientific literature.This dataset facilitates the development of systems that can assess and aggregate contradictory evidence across multiple studies, and is the first large-scale, publicly available multi-document summarization dataset in the biomedical domain.We experiment with a summarization system based on BART, with promising early results, though significant work remains to achieve higher summarization quality.We formulate our summarization inputs and targets in both free text and structured forms and modify a recently proposed metric to assess the quality of our system's generated summaries.
Jay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl, Lucy Lu Wang
EMNLP (1)1
2020 ERASER: A Benchmark to Evaluate Rationalized NLP Models
abstract
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, Byron C. Wallace. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Jay DeYoung, Nazneen Fatema Rajani, Eric P. Lehman, Caiming Xiong, Richard Socher, Byron C. Wallace
ACL1
2019 Neural-Network Lexical Translation for Cross-lingual IR from Text and Speech
abstract
We propose a neural network model to estimate word translation probabilities for Cross-Lingual Information Retrieval (CLIR). The model estimates better probabilities for word translations than automatic word alignments alone, and generalizes to unseen source-target word pairs. We further improve the lexical neural translation model (and subsequently CLIR), by incorporating source word context, and by encoding the character sequences of input source words to generate translations of out-of-vocabulary words. To be effective, neural network models typically need training on large amounts of data labeled directly on the final task, in this case relevance to queries. In contrast, our approach only requires parallel data to train the translation model, and uses an unsupervised model to compute CLIR relevance scores.
Rabih Zbib, Lingjun Zhao, Damianos Karakos, William Hartmann, Jay DeYoung, Zhongqiang Huang, Zhuolin Jiang, Noah Rivkin, Le Zhang 0002, Richard M. Schwartz, John Makhoul
SIGIR5
2018 Combining rule-based and statistical mechanisms for low-resource named entity recognition
Ryan Gabbard, Jay DeYoung, Constantine Lignos, Marjorie Freedman, Ralph M. Weischedel
Mach. Transl.2
2015 A Concrete Chinese NLP Pipeline
abstract
Nanyun Peng, Francis Ferraro, Mo Yu, Nicholas Andrews, Jay DeYoung, Max Thomas, Matthew R. Gormley, Travis Wolfe, Craig Harman, Benjamin Van Durme, Mark Dredze. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. 2015.
Nanyun Peng 0001, Francis Ferraro, Mo Yu, Nicholas Andrews, Jay DeYoung, Max Thomas, Matthew R. Gormley, Travis Wolfe, Craig Harman, Benjamin Van Durme, Mark Dredze
HLT-NAACL5