VLDB 2026 Research / reviewers in the wild / expert
Jay DeYoung
dblp:136/8673
· DBLP profile ↗
8ranked-venue papers
4as first author
4since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% | |
| Artificial intelligence
2 papers |
Language models and text generation · 72% Trustworthy machine learning · 28% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Medical and health informatics · 100% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
evaluation |
0.7 | 1 | 2023 | Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023 |
Information retrieval › text summarization
multi-document summarization |
0.7 | 1 | 2023 | Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023 |
Information retrieval › text summarization
summarization evaluation |
0.7 | 1 | 2023 | Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023 |
Information retrieval
text summarization |
0.7 | 1 | 2023 | Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023 |
Natural language and speech › Language models and text generation › text summarization
multi-document summarization |
0.5 | 1 | 2021 | MS\^2: Multi-Document Summarization of Medical Studies · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation
text summarization |
0.5 | 1 | 2021 | MS\^2: Multi-Document Summarization of Medical Studies · EMNLP (1) 2021 |
Machine learning › Trustworthy machine learning
interpretability |
0.4 | 1 | 2020 | ERASER: A Benchmark to Evaluate Rationalized NLP Models · ACL 2020 |
Information retrieval
cross-language information retrieval |
0.4 | 1 | 2019 | Neural-Network Lexical Translation for Cross-lingual IR from Text and Speech · SIGIR 2019 |
Medical and health informatics › clinical text processing
clinical text summarization |
0.2 | 1 | 2023 | Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations · ACL (1) 2023 |
Information retrieval › document retrieval
spoken document retrieval |
0.1 | 1 | 2019 | Neural-Network Lexical Translation for Cross-lingual IR from Text and Speech · SIGIR 2019 |
Methods — techniques the papers use, named apart from their topics
automated metrics · 1.3ROUGE · 1.3summarization evaluation metric · 1.0BART · 1.0word alignment · 0.4neural network · 0.4character sequence encoding · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Do Multi-Document Summarization Models Synthesize?abstractAbstract Multi-document summarization entails producing concise synopses of collections of inputs. For some applications, the synopsis should accurately synthesize inputs with respect to a key aspect, e.g., a synopsis of film reviews written about a particular movie should reflect the average critic consensus. As a more consequential example, narrative summaries that accompany biomedical systematic reviews of clinical trial results should accurately summarize the potentially conflicting results from individual trials. In this paper we ask: To what extent do modern multi-document summarization models implicitly perform this sort of synthesis? We run experiments over opinion and evidence synthesis datasets using a suite of summarization models, from fine-tuned transformers to GPT-4. We find that existing models partially perform synthesis, but imperfectly: Even the best performing models are over-sensitive to changes in input ordering and under-sensitive to changes in input compositions (e.g., ratio of positive to negative reviews). We propose a simple, general, effective method for improving model synthesis capabilities by generating an explicitly diverse set of candidate outputs, and then selecting from these the string best aligned with the expected aggregate measure for the inputs, or abstaining when the model produces no good candidate. Jay DeYoung, Stephanie C. Martinez, Iain James Marshall, Byron C. Wallace |
Trans. Assoc. Comput. Linguistics | 1 |
| 2023 | Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluationsabstract-gram similarity metrics such as ROUGE. Better automated evaluation metrics are needed, but few resources exist to assess metrics when they are proposed. Therefore, we introduce a dataset of human-assessed summary quality facets and pairwise preferences to encourage and support the development of better automated evaluation methods for literature review MDS. We take advantage of community submissions to the Multi-document Summarization for Literature Review (MSLR) shared task to compile a diverse and representative sample of generated summaries. We analyze how automated summarization evaluation metrics correlate with lexical features of generated summaries, to other automated metrics including several we propose in this work, and to aspects of human-assessed summary quality. We find that not only do automated metrics fail to capture aspects of quality as assessed by humans, in many cases the system rankings produced by these metrics are anti-correlated with rankings according to human annotators. Lucy Lu Wang, Yulia Otmakhova 0001, Jay DeYoung, Hung-Thinh Truong, Bailey Kuehl, Erin Bransom, Byron C. Wallace |
ACL (1) | 3 |
| 2022 | Entity Anchored ICD Coding
Jay DeYoung, Han-Chin Shing, Luyang Kong, Christopher Winestock, Chaitanya P. Shivade |
AMIA | 1 |
| 2021 | MS\^2: Multi-Document Summarization of Medical StudiesabstractTo assess the effectiveness of any medical intervention, researchers must conduct a timeintensive and manual literature review.NLP systems can help to automate or assist in parts of this expensive process.In support of this goal, we release MSˆ2 (Multi-Document Summarization of Medical Studies), a dataset of over 470k documents and 20K summaries derived from the scientific literature.This dataset facilitates the development of systems that can assess and aggregate contradictory evidence across multiple studies, and is the first large-scale, publicly available multi-document summarization dataset in the biomedical domain.We experiment with a summarization system based on BART, with promising early results, though significant work remains to achieve higher summarization quality.We formulate our summarization inputs and targets in both free text and structured forms and modify a recently proposed metric to assess the quality of our system's generated summaries. Jay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl, Lucy Lu Wang |
EMNLP (1) | 1 |
| 2020 | ERASER: A Benchmark to Evaluate Rationalized NLP ModelsabstractJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, Byron C. Wallace. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Jay DeYoung, Nazneen Fatema Rajani, Eric P. Lehman, Caiming Xiong, Richard Socher, Byron C. Wallace |
ACL | 1 |
| 2019 | Neural-Network Lexical Translation for Cross-lingual IR from Text and SpeechabstractWe propose a neural network model to estimate word translation probabilities for Cross-Lingual Information Retrieval (CLIR). The model estimates better probabilities for word translations than automatic word alignments alone, and generalizes to unseen source-target word pairs. We further improve the lexical neural translation model (and subsequently CLIR), by incorporating source word context, and by encoding the character sequences of input source words to generate translations of out-of-vocabulary words. To be effective, neural network models typically need training on large amounts of data labeled directly on the final task, in this case relevance to queries. In contrast, our approach only requires parallel data to train the translation model, and uses an unsupervised model to compute CLIR relevance scores. Rabih Zbib, Lingjun Zhao, Damianos Karakos, William Hartmann, Jay DeYoung, Zhongqiang Huang, Zhuolin Jiang, Noah Rivkin, Le Zhang 0002, Richard M. Schwartz, John Makhoul |
SIGIR | 5 |
| 2018 | Combining rule-based and statistical mechanisms for low-resource named entity recognition
Ryan Gabbard, Jay DeYoung, Constantine Lignos, Marjorie Freedman, Ralph M. Weischedel |
Mach. Transl. | 2 |
| 2015 | A Concrete Chinese NLP PipelineabstractNanyun Peng, Francis Ferraro, Mo Yu, Nicholas Andrews, Jay DeYoung, Max Thomas, Matthew R. Gormley, Travis Wolfe, Craig Harman, Benjamin Van Durme, Mark Dredze. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. 2015. Nanyun Peng 0001, Francis Ferraro, Mo Yu, Nicholas Andrews, Jay DeYoung, Max Thomas, Matthew R. Gormley, Travis Wolfe, Craig Harman, Benjamin Van Durme, Mark Dredze |
HLT-NAACL | 5 |