VLDB 2026 Research / reviewers in the wild / expert
Robert L. Logan IV
dblp:210/2652
· DBLP profile ↗
12ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Language models and text generation · 63% Information extraction and text analysis · 19% Trustworthy machine learning · 9% | |
| Databases, data mining, and information retrieval
1 paper |
Knowledge graphs · 100% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › evaluation of language models
faithfulness evaluation |
0.7 | 1 | 2023 | BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics · ACL (1) 2023 |
Natural language and speech › Language models and text generation › large language model evaluation
meta-evaluation |
0.7 | 1 | 2023 | BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics · ACL (1) 2023 |
Natural language and speech › Language models and text generation
text generation evaluation |
0.7 | 1 | 2023 | BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics · ACL (1) 2023 |
Natural language and speech › Language models and text generation › in-context learning
few-shot prompting |
0.6 | 1 | 2022 | Continued Pretraining for Better Zero- and Few-Shot Promptability · EMNLP 2022 |
Computer vision › Vision and language › vision-language model
prompt learning |
0.6 | 1 | 2022 | Continued Pretraining for Better Zero- and Few-Shot Promptability · EMNLP 2022 |
Natural language and speech › Language models and text generation
prompt tuning |
0.6 | 1 | 2022 | Continued Pretraining for Better Zero- and Few-Shot Promptability · EMNLP 2022 |
Natural language and speech › Information extraction and text analysis
coreference resolution |
0.5 | 1 | 2021 | Benchmarking Scalable Methods for Streaming Cross Document Entity Coreference · ACL/IJCNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis › coreference resolution
cross-document coreference resolution |
0.5 | 1 | 2021 | Benchmarking Scalable Methods for Streaming Cross Document Entity Coreference · ACL/IJCNLP (1) 2021 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.5 | 1 | 2021 | Active Bayesian Assessment of Black-Box Classifiers · AAAI 2021 |
Natural language and speech › Language models and text generation › prompting › prompt engineering
prompt generation |
0.4 | 1 | 2020 | AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.4 | 1 | 2020 | AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › text representation
contextualized word embeddings |
0.4 | 1 | 2019 | Knowledge Enhanced Contextual Word Representations · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Language models and text generation › large language model › knowledge in language models
knowledge-grounded language model |
0.4 | 1 | 2019 | Barack's Wife Hillary: Using Knowledge Graphs for Fact-Aware Language Modeling · ACL (1) 2019 |
Machine learning › Trustworthy machine learning
robustness |
0.1 | 1 | 2021 | Active Bayesian Assessment of Black-Box Classifiers · AAAI 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph |
0.1 | 1 | 2019 | Knowledge Enhanced Contextual Word Representations · EMNLP/IJCNLP (1) 2019 |
Methods — techniques the papers use, named apart from their topics
neural language model · 0.8knowledge graph fact selection and copying · 0.8minimal pair construction · 0.7multi-task learning · 0.6MAML-style meta-learning · 0.6bayesian inference · 0.5active learning · 0.5masked language model · 0.4importance sampling · 0.4gradient-guided search · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Uchaguzi-2022: A Dataset of Citizen Reports on the 2022 Kenyan ElectionabstractOnline reporting platforms have enabled citizens around the world to collectively share their opinions and report in real time on events impacting their local communities. Systematically organizing (e.g., categorizing by attributes) and geotagging large amounts of crowdsourced information is crucial to ensuring that accurate and meaningful insights can be drawn from this data and used by policy makers to bring about positive change. These tasks, however, typically require extensive manual annotation efforts. In this paper we present Uchaguzi-2022, a dataset of 14k categorized and geotagged citizen reports related to the 2022 Kenyan General Election containing mentions of election-related issues such as official misconduct, vote count irregularities, and acts of violence. We use this dataset to investigate whether language models can assist in scalably categorizing and geotagging reports, thus highlighting its potential application in the AI for Social Good space. Roberto Mondini, Neema Kotonya, Robert L. Logan IV, Elizabeth M. Olson, Angela Oduor Lungati, Daniel Duke Odongo, Tim Ombasa, Hemank Lamba, Aoife Cahill, Joel R. Tetreault, Alejandro Jaimes |
COLING | 3 |
| 2023 | BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness MetricsabstractLiang Ma, Shuyang Cao, Robert L Logan IV, Di Lu, Shihao Ran, Ke Zhang, Joel Tetreault, Alejandro Jaimes. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shuyang Cao, Robert L. Logan IV, Di Lu 0003, Shihao Ran, Ke Zhang 0013, Joel R. Tetreault, Alejandro Jaimes |
ACL (1) | 3 |
| 2023 | Evaluating the generalisability of neural rumour verification modelsabstractResearch on automated social media rumour verification, the task of identifying the veracity of questionable information circulating on social media, has yielded neural models achieving high performance, with accuracy scores that often exceed 90%. However, none of these studies focus on the real-world generalisability of the proposed approaches, that is whether the models perform well on datasets other than those on which they were initially trained and tested. In this work we aim to fill this gap by assessing the generalisability of top performing neural rumour verification models covering a range of different architectures from the perspectives of both topic and temporal robustness. For a more complete evaluation of generalisability, we collect and release COVID-RV, a novel dataset of Twitter conversations revolving around COVID-19 rumours. Unlike other existing COVID-19 datasets, our COVID-RV contains conversations around rumours that follow the format of prominent rumour verification benchmarks, while being different from them in terms of topic and time scale, thus allowing better assessment of the temporal robustness of the models. We evaluate model performance on COVID-RV and three popular rumour verification datasets to understand limitations and advantages of different model architectures, training datasets and evaluation scenarios. We find a dramatic drop in performance when testing models on a different dataset from that used for training. Further, we evaluate the ability of models to generalise in a few-shot learning setup, as well as when word embeddings are updated with the vocabulary of a new, unseen rumour. Drawing upon our experiments we discuss challenges and make recommendations for future research directions in addressing this important problem. Elena Kochkina, Tamanna Hossain, Robert L. Logan IV, Miguel Arana-Catania, Rob Procter, Arkaitz Zubiaga, Sameer Singh 0001, Yulan He 0001, Maria Liakata |
Inf. Process. Manag. | 3 |
| 2022 | Continued Pretraining for Better Zero- and Few-Shot PromptabilityabstractRecently introduced language model prompting methods can achieve high accuracy in zeroand few-shot settings while requiring few to no learned task-specific parameters.Nevertheless, these methods still often trail behind full model finetuning.In this work, we investigate if a dedicated continued pretraining stage could improve "promptability", i.e., zero-shot performance with natural language prompts or few-shot performance with prompt tuning.We reveal settings where existing continued pretraining methods lack promptability.We also identify current methodological gaps, which we fill with thorough large-scale experiments.We demonstrate that a simple recipe, continued pretraining that incorporates a trainable prompt during multi-task learning, leads to improved promptability in both zero-and fewshot settings compared to existing methods, up to 31% relative.On the other hand, we find that continued pretraining using MAML-style metalearning, a method that directly optimizes fewshot promptability, yields subpar performance.We validate our findings with two prompt tuning methods, and, based on our results, we provide concrete recommendations to optimize promptability for different use cases. Zhaofeng Wu, Robert L. Logan IV, Pete Walsh 0001, Akshita Bhagia, Dirk Groeneveld, Sameer Singh 0001, Iz Beltagy |
EMNLP | 2 |
| 2022 | FRUIT: Faithfully Reflecting Updated Information in TextabstractRobert Iv, Alexandre Passos, Sameer Singh, Ming-Wei Chang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Robert L. Logan IV, Alexandre Tachard Passos, Sameer Singh 0001, Ming-Wei Chang |
NAACL-HLT | 1 |
| 2021 | Active Bayesian Assessment of Black-Box Classifiers
Disi Ji, Robert L. Logan IV, Padhraic Smyth, Mark Steyvers |
AAAI | 2 |
| 2021 | Benchmarking Scalable Methods for Streaming Cross Document Entity CoreferenceabstractRobert L Logan IV, Andrew McCallum, Sameer Singh, Dan Bikel. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Robert L. Logan IV, Andrew McCallum, Sameer Singh 0001, Dan Bikel |
ACL/IJCNLP (1) | 1 |
| 2020 | On Importance Sampling-Based Evaluation of Latent Language ModelsabstractLanguage models that use additional latent structures (e.g., syntax trees, coreference chains, and knowledge graph links) provide several advantages over traditional language models.However, likelihood-based evaluation of these models is often intractable as it requires marginalizing over the latent space.Existing methods avoid this issue by using importance sampling.Although this approach has asymptotic guarantees, analysis is rarely conducted on the effect of decisions such as sample size, granularity of sample aggregation, and the proposal distribution on the reported estimates.In this paper, we measure the effect these factors have on perplexity estimates for three different latent language models.In addition, we elucidate subtle differences in how importance sampling is applied, which can have substantial effects on the final estimates, as well as provide theoretical results that reinforce the validity of importance sampling for evaluating latent language models. Robert L. Logan IV, Matt Gardner 0001, Sameer Singh 0001 |
ACL | 1 |
| 2020 | AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsabstractThe remarkable success of pretrained language models has motivated the study of what kinds of knowledge these models learn during pretraining.Reformulating tasks as fillin-the-blanks problems (e.g., cloze tests) is a natural approach for gauging such knowledge, however, its usage is limited by the manual effort and guesswork required to write suitable prompts.To address this, we develop AUTOPROMPT, an automated method to create prompts for a diverse set of tasks, based on a gradient-guided search.Using AUTO-PROMPT, we show that masked language models (MLMs) have an inherent capability to perform sentiment analysis and natural language inference without additional parameters or finetuning, sometimes achieving performance on par with recent state-of-the-art supervised models.We also show that our prompts elicit more accurate factual knowledge from MLMs than the manually created prompts on the LAMA benchmark, and that MLMs can be used as relation extractors more effectively than supervised relation extraction models.These results demonstrate that automatically generated prompts are a viable parameter-free alternative to existing probing methods, and as pretrained LMs become more sophisticated and capable, potentially a replacement for finetuning. Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, Sameer Singh 0001 |
EMNLP (1) | 3 |
| 2019 | Barack's Wife Hillary: Using Knowledge Graphs for Fact-Aware Language ModelingabstractModeling human language requires the ability to not only generate fluent text but also encode factual knowledge.However, traditional language models are only capable of remembering facts seen at training time, and often have difficulty recalling them.To address this, we introduce the knowledge graph language model (KGLM), a neural language model with mechanisms for selecting and copying facts from a knowledge graph that are relevant to the context.These mechanisms enable the model to render information it has never seen before, as well as generate out-of-vocabulary tokens.We also introduce the Linked WikiText-2 dataset, 1 a corpus of annotated text aligned to the Wikidata knowledge graph whose contents (roughly) match the popular WikiText-2 benchmark (Merity et al., 2017).In experiments, we demonstrate that the KGLM achieves significantly better performance than a strong baseline language model.We additionally compare different language models' ability to complete sentences requiring factual knowledge, and show that the KGLM outperforms even very large language models in generating facts. Robert L. Logan IV, Nelson F. Liu, Matthew E. Peters, Matt Gardner 0001, Sameer Singh 0001 |
ACL (1) | 1 |
| 2019 | Knowledge Enhanced Contextual Word RepresentationsabstractMatthew E. Peters, Mark Neumann, Robert Logan, Roy Schwartz, Vidur Joshi, Sameer Singh, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Matthew E. Peters, Mark Neumann, Robert L. Logan IV, Roy Schwartz 0001, Vidur Joshi, Sameer Singh 0001, Noah A. Smith |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Detecting conversation topics in primary care office visits from transcripts of patient-provider interactionsabstractOBJECTIVE: Amid electronic health records, laboratory tests, and other technology, office-based patient and provider communication is still the heart of primary medical care. Patients typically present multiple complaints, requiring physicians to decide how to balance competing demands. How this time is allocated has implications for patient satisfaction, payments, and quality of care. We investigate the effectiveness of machine learning methods for automated annotation of medical topics in patient-provider dialog transcripts. MATERIALS AND METHODS: We used dialog transcripts from 279 primary care visits to predict talk-turn topic labels. Different machine learning models were trained to operate on single or multiple local talk-turns (logistic classifiers, support vector machines, gated recurrent units) as well as sequential models that integrate information across talk-turn sequences (conditional random fields, hidden Markov models, and hierarchical gated recurrent units). RESULTS: Evaluation was performed using cross-validation to measure 1) classification accuracy for talk-turns and 2) precision, recall, and F1 scores at the visit level. Experimental results showed that sequential models had higher classification accuracy at the talk-turn level and higher precision at the visit level. Independent models had higher recall scores at the visit level compared with sequential models. CONCLUSIONS: Incorporating sequential information across talk-turns improves the accuracy of topic prediction in patient-provider dialog by smoothing out noisy information from talk-turns. Although the results are promising, more advanced prediction techniques and larger labeled datasets will likely be required to achieve prediction performance appropriate for real-world clinical applications. Dimitrios Kotzias, Patty Kuo, Robert L. Logan IV, Kritzia Merced, Sameer Singh 0001, Michael Tanana, Efi Karra Taniskidou, Jennifer Elston-Lafata, David C. Atkins, Ming Tai-Seale, Zac E. Imel, Padhraic Smyth |
J. Am. Medical Informatics Assoc. | 4 |