VLDB 2026 Research / reviewers in the wild / expert
Kevin Small
dblp:82/6573
· DBLP profile ↗
24ranked-venue papers
1as first author
12since 2021 · last 2026
0009-0003-8705-1512ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WiNELL: Wikipedia Never-Ending Updating with LLM Agents
Revanth Gangi Reddy, Tanay Dixit, Jiaxin Qin, Cheng Qian 0008, Jiawei Han 0001, Kevin Small, Ruhi Sarikaya, Heng Ji 0001 |
WWW | 7 |
| 2025 | Persona-DB: Efficient Large Language Model Personalization for Response Prediction with Collaborative Data RefinementabstractThe increasing demand for personalized interactions with large language models (LLMs) calls for methodologies capable of accurately and efficiently identifying user opinions and preferences. Retrieval augmentation emerges as an effective strategy, as it can accommodate a vast number of users without the costs from fine-tuning. Existing research, however, has largely focused on enhancing the retrieval stage and devoted limited exploration toward optimizing the representation of the database, a crucial aspect for tasks such as personalization. In this work, we examine the problem from a novel angle, focusing on how data can be better represented for more data-efficient retrieval in the context of LLM customization. To tackle this challenge, we introduce Persona-DB, a simple yet effective framework consisting of a hierarchical construction process to improve generalization across task contexts and collaborative refinement to effectively bridge knowledge gaps among users. In the evaluation of response prediction, Persona-DB demonstrates superior context efficiency in maintaining accuracy with a significantly reduced retrieval size, a critical advantage in scenarios with extensive histories or limited context windows. Our experiments also indicate a marked improvement of over 10% under cold-start scenarios, when users have extremely sparse data. Furthermore, our analysis reveals the increasing importance of collaborative knowledge as the retrieval capacity expands. Chenkai Sun, Ke Yang 0003, Revanth Gangi Reddy, Yi R. Fung 0001, Hou Pong Chan, Kevin Small, ChengXiang Zhai, Heng Ji 0001 |
COLING | 6 |
| 2024 | EVEDIT: Event-based Knowledge Editing for Deterministic Knowledge PropagationabstractThe dynamic nature of real-world information necessitates knowledge editing (KE) in large language models (LLMs).This edited knowledge should propagate and facilitate the deduction of new information based on existing model knowledge.We define the existing related knowledge in a LLM serving as the origination of knowledge propagation as "deduction anchors".However, most of current KE approaches only operate on (subject, relation, object) triples.Both theoretically and empirically, we observe that this simplified setting often leads to uncertainty when determining the deduction anchors, causing low confidence in their responses.To mitigate this issue, we propose a novel task of event-based knowledge editing that pairs facts with event descriptions.This task manifests both as a closer simulation of real-world editing scenarios and a more logically sound setting, implicitly defining the deduction anchor and enabling LLMs to propagate knowledge confidently.We curate a new benchmark dataset EVEDIT derived from the COUNTERFACT dataset and validate its superiority in improving model confidence.Moreover, as we observe that the event-based setting is notably challenging for existing approaches, we propose a novel approach Self-Edit that showcases stronger performance, achieving 55.6% consistency improvement while maintaining the naturalness of generation. 1 Implicitly define the deduction anchor, ensuring model certainty. Previous Simple edits:Messi is a Dutch citizen.In 2024, Lionel Messi made the decision to move to Netherlands and applied for Dutch citizenship.After necessary procedures, he was granted Dutch citizenship and became a citizen of Netherlands.Q: Is Messi a citizen of Argentina in 2023?Q:Where was Messi born ?Q: Did Messi won the World Cup in 2022 ?Ignore the deduction anchor, leading to model uncertainty. Jiateng Liu, Pengfei Yu 0001, Yuji Zhang 0002, Ruhi Sarikaya, Kevin Small, Heng Ji 0001 |
EMNLP | 7 |
| 2023 | SumREN: Summarizing Reported Speech about Events in NewsabstractA primary objective of news articles is to establish the factual record for an event, frequently achieved by conveying both the details of the specified event (i.e., the 5 Ws; Who, What, Where, When and Why regarding the event) and how people reacted to it (i.e., reported statements). However, existing work on news summarization almost exclusively focuses on the event details. In this work, we propose the novel task of summarizing the reactions of different speakers, as expressed by their reported statements, to a given event. To this end, we create a new multi-document summarization benchmark, SumREN, comprising 745 summaries of reported statements from various public figures obtained from 633 news articles discussing 132 events. We propose an automatic silver-training data generation approach for our task, which helps smaller models like BART achieve GPT-3 level performance on this task. Finally, we introduce a pipeline-based framework for summarizing reported speech, which we empirically show to generate summaries that are more abstractive and factual than baseline query-focused summarization approaches. Revanth Gangi Reddy, Heba Elfardy, Hou Pong Chan, Kevin Small, Heng Ji 0001 |
AAAI | 4 |
| 2023 | Enhancing Multi-Document Summarization with Cross-Document Graph-based Information ExtractionabstractInformation extraction (IE) and summarization are closely related, both tasked with presenting a subset of the information contained in a natural language text. However, while IE extracts structural representations, summarization aims to abstract the most salient information into a generated text summary – thus potentially encountering the technical limitations of current text generation methods (e.g., hallucination). To mitigate this risk, this work uses structured IE graphs to enhance the abstractive summarization task. Specifically, we focus on improving Multi-Document Summarization (MDS) performance by using cross-document IE output, incorporating two novel components: (1) the use of auxiliary entity and event recognition systems to focus the summary generation model; (2) incorporating an alignment loss between IE nodes and their text spans to reduce inconsistencies between the IE graphs and text representations. Operationally, both the IE nodes and corresponding text spans are projected into the same embedding space and pairwise distance is minimized. Experimental results on multiple MDS benchmarks show that summaries generated by our model are more factually consistent with the source documents than baseline models while maintaining the same level of abstractiveness. Heba Elfardy, Markus Dreyer, Kevin Small, Heng Ji 0001, Mohit Bansal |
EACL | 4 |
| 2023 | Background Summarization of Event TimelinesabstractGenerating concise summaries of news events is a challenging natural language processing task.While journalists often curate timelines to highlight key sub-events, newcomers to a news event face challenges in catching up on its historical context.In this paper, we address this need by introducing the task of background news summarization, which complements each timeline update with a background summary of relevant preceding events.We construct a dataset by merging existing timeline datasets and asking human annotators to write a background summary for each timestep of each news event.We establish strong baseline performance using state-of-the-art summarization systems and propose a query-focused variant to generate background summaries.To evaluate background summary quality, we present a question-answering-based evaluation metric, Background Utility Score (BUS), which measures the percentage of questions about a current event timestep that a background summary answers.Our experiments show the effectiveness of instruction fine-tuned systems such as Flan-T5, in addition to strong zero-shot performance using GPT-3.5. 1 Adithya Pratapa, Kevin Small, Markus Dreyer |
EMNLP | 2 |
| 2022 | A Zero-Shot Claim Detection Framework Using Question AnsweringabstractIn recent years, there has been an increasing interest in claim detection as an important building block for misinformation detection. This involves detecting more fine-grained attributes relating to the claim, such as the claimer, claim topic, claim object pertaining to the topic, etc. Yet, a notable bottleneck of existing claim detection approaches is their portability to emerging events and low-resource training data settings. In this regard, we propose a fine-grained claim detection framework that leverages zero-shot Question Answering (QA) using directed questions to solve a diverse set of sub-tasks such as topic filtering, claim object detection, and claimer detection. We show that our approach significantly outperforms various zero-shot, few-shot and task-specific baselines on the NewsClaims benchmark (Reddy et al., 2021). Revanth Gangi Reddy, Sai Chetan Chinthakindi, Yi R. Fung 0001, Kevin Small, Heng Ji 0001 |
COLING | 4 |
| 2022 | NewsClaims: A New Benchmark for Claim Detection from News with Attribute KnowledgeabstractRevanth Gangi Reddy, Sai Chetan Chinthakindi, Zhenhailong Wang, Yi Fung, Kathryn Conger, Ahmed ELsayed, Martha Palmer, Preslav Nakov, Eduard Hovy, Kevin Small, Heng Ji. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Revanth Gangi Reddy, Sai Chetan Chinthakindi, Zhenhailong Wang, Yi R. Fung 0001, Kathryn Conger, Ahmed Elsayed, Martha Palmer, Preslav Nakov, Eduard H. Hovy, Kevin Small, Heng Ji 0001 |
EMNLP | 10 |
| 2022 | Building a Dataset for Automatically Learning to Detect Questions Requiring ClarificationabstractQuestion Answering (QA) systems aim to return correct and concise answers in response to user questions. QA research generally assumes all questions are intelligible and unambiguous, which is unrealistic in practice as questions frequently encountered by virtual assistants are ambiguous or noisy. In this work, we propose to make QA systems more robust via the following two-step process: (1) classify if the input question is intelligible and (2) for such questions with contextual ambiguity, return a clarification question. We describe a new open-domain clarification corpus containing user questions sampled from Quora, which is useful for building machine learning approaches to solving these tasks. Ivano Lauriola, Kevin Small, Alessandro Moschitti |
LREC | 2 |
| 2022 | Answer Consolidation: Formulation and BenchmarkingabstractWenxuan Zhou, Qiang Ning, Heba Elfardy, Kevin Small, Muhao Chen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Wenxuan Zhou 0002, Qiang Ning, Heba Elfardy, Kevin Small, Muhao Chen 0001 |
NAACL-HLT | 4 |
| 2021 | Inverse Reinforcement Learning with Natural Language GoalsabstractHumans generally use natural language to communicate task requirements to each other. Ideally, natural language should also be usable for communicating goals to autonomous machines (e.g., robots) to minimize friction in task specification. However, understanding and mapping natural language goals to sequences of states and actions is challenging. Specifically, existing work along these lines has encountered difficulty in generalizing learned policies to new natural language goals and environments. In this paper, we propose a novel adversarial inverse reinforcement learning algorithm to learn a language-conditioned policy and reward function. To improve generalization of the learned policy and reward function, we use a variational goal generator to relabel trajectories and sample diverse goals during training. Our algorithm outperforms multiple baselines by a large margin on a vision-based natural language instruction following dataset (Room-2-Room), demonstrating a promising advance in enabling the use of natural language instructions in specifying agent goals. Li Zhou 0006, Kevin Small |
AAAI | 2 |
| 2021 | Generating Self-Contained and Summary-Centric Question Answer Pairs via Differentiable Reward Imitation LearningabstractMotivated by suggested question generation in conversational news recommendation systems, we propose a model for generating question-answer pairs (QA pairs) with selfcontained, summary-centric questions and length-constrained, article-summarizing answers.We begin by collecting a new dataset of news articles with questions as titles and pairing them with summaries of varying length.This dataset is used to learn a QA pair generation model producing summaries as answers that balance brevity with sufficiency jointly with their corresponding questions.We then reinforce the QA pair generation process with a differentiable reward function to mitigate exposure bias, a common problem in natural language generation.Both automatic metrics and human evaluation demonstrate these QA pairs successfully capture the central gists of the articles and achieve high answer accuracy.1 Li Zhou 0006, Kevin Small, Sandeep Atluri |
EMNLP (1) | 2 |
| 2020 | Fluent Response Generation for Conversational Question AnsweringabstractQuestion answering (QA) is an important aspect of open-domain conversational agents, garnering specific research focus in the conversational QA (ConvQA) subtask.One notable limitation of recent ConvQA efforts is the response being answer span extraction from the target corpus, thus ignoring the natural language generation (NLG) aspect of high-quality conversational agents.In this work, we propose a method for situating QA responses within a SEQ2SEQ NLG approach to generate fluent grammatical answer responses while maintaining correctness.From a technical perspective, we use data augmentation to generate training data for an end-to-end system.Specifically, we develop Syntactic Transformations (STs) to produce question-specific candidate answer responses and rank them using a BERT-based classifier (Devlin et al., 2019).Human evaluation on SQuAD 2.0 data (Rajpurkar et al., 2018) demonstrate that the proposed model outperforms baseline CoQA and QuAC models in generating conversational responses.We further show our model's scalability by conducting tests on the CoQA dataset. 1 Ashutosh Baheti, Alan Ritter, Kevin Small |
ACL | 3 |
| 2011 | Class Imbalance, ReduxabstractClass imbalance (i.e., scenarios in which classes are unequally represented in the training data) occurs in many real-world learning tasks. Yet despite its practical importance, there is no established theory of class imbalance, and existing methods for handling it are therefore not well motivated. In this work, we approach the problem of imbalance from a probabilistic perspective, and from this vantage identify dataset characteristics (such as dimensionality, sparsity, etc.) that exacerbate the problem. Motivated by this theory, we advocate the approach of bagging an ensemble of classifiers induced over balanced bootstrap training samples, arguing that this strategy will often succeed where others fail. Thus in addition to providing a theoretical understanding of class imbalance, corroborated by our experiments on both simulated and real datasets, we provide practical guidance for the data mining practitioner working with imbalanced data. Byron C. Wallace, Kevin Small, Carla E. Brodley, Thomas A. Trikalinos |
ICDM | 2 |
| 2011 | The Constrained Weight Space SVM: Learning with Ranked Features
Kevin Small, Byron C. Wallace, Carla E. Brodley, Thomas A. Trikalinos |
ICML | 1 |
| 2011 | Who Should Label What? Instance Allocation in Multiple Expert Active LearningabstractThe active learning (AL) framework is an increasingly popular strategy for reducing the amount of human labeling effort required to induce a predictive model. Most work in AL has assumed that a single, infallible oracle provides labels requested by the learner at a fixed cost. However, real-world applications suitable for AL often include multiple domain experts who provide labels of varying cost and quality. We explore this multiple expert active learning (MEAL) scenario and develop a novel algorithm for instance allocation that exploits the meta-cognitive abilities of novice (cheap) experts in order to make the best use of the experienced (expensive) annotators. We demonstrate that this strategy outperforms strong baseline approaches to MEAL on both a sentiment analysis dataset and two datasets from our motivating application of biomedical citation screening. Furthermore, we provide evidence that novice labelers are often aware of which instances they are likely to mislabel. Byron C. Wallace, Kevin Small, Carla E. Brodley, Thomas A. Trikalinos |
SDM | 2 |
| 2010 | Active learning for biomedical citation screeningabstractActive learning (AL) is an increasingly popular strategy for mitigating the amount of labeled data required to train classifiers, thereby reducing annotator effort. We describe a real-world, deployed application of AL to the problem of biomedical citation screening for systematic reviews at the Tufts Medical Center's Evidence-based Practice Center. We propose a novel active learning strategy that exploits a priori domain knowledge provided by the expert (specifically, labeled features)and extend this model via a Linear Programming algorithm for situations where the expert can provide ranked labeled features. Our methods outperform existing AL strategies on three real-world systematic review datasets. We argue that evaluation must be specific to the scenario under consideration. To this end, we propose a new evaluation framework for finite-pool scenarios, wherein the primary aim is to label a fixed set of examples rather than to simply induce a good predictive model. We use a method from medical decision theory for eliciting the relative costs of false positives and false negatives from the domain expert, constructing a utility measure of classification performance that integrates the expert preferences. Our findings suggest that the expert can, and should, provide more information than instance labels alone. In addition to achieving strong empirical results on the citation screening problem, this work outlines many important steps for moving away from simulated active learning and toward deploying AL for real-world applications. Byron C. Wallace, Kevin Small, Carla E. Brodley, Thomas A. Trikalinos |
KDD | 2 |
| 2009 | Interactive Feature Space Construction using Semantic Information
Dan Roth 0001, Kevin Small |
CoNLL | 2 |
| 2009 | Unsupervised Rank Aggregation with Domain-Specific Expertise
Alexandre Klementiev, Dan Roth 0001, Kevin Small, Ivan Titov 0001 |
IJCAI | 3 |
| 2008 | Active Learning for Pipeline Models
Dan Roth 0001, Kevin Small |
AAAI | 2 |
| 2008 | Unsupervised rank aggregation with distance-based modelsabstractThe need to meaningfully combine sets of rankings often comes up when one deals with ranked data. Although a number of heuristic and supervised learning approaches to rank aggregation exist, they require domain knowledge or supervised ranked data, both of which are expensive to acquire. In order to address these limitations, we propose a mathematical and algorithmic framework for learning to aggregate (partial) rankings without supervision. We instantiate the framework for the cases of combining permutations and combining top-k lists, and propose a novel metric for the latter. Experiments in both scenarios demonstrate the effectiveness of the proposed formalism. Alexandre Klementiev, Dan Roth 0001, Kevin Small |
ICML | 3 |
| 2007 | An Unsupervised Learning Algorithm for Rank Aggregation
Alexandre Klementiev, Dan Roth 0001, Kevin Small |
ECML | 3 |
| 2007 | All links are not the same: evaluating word alignments for statistical machine translation
Paul C. Davis, Zhuli Xie, Kevin Small |
MTSummit | 3 |
| 2006 | Margin-Based Active Learning for Structured Output Spaces
Dan Roth 0001, Kevin Small |
ECML | 2 |