VLDB 2026 Research / reviewers in the wild / expert
Stefan Larson
dblp:239/4267
· DBLP profile ↗
13ranked-venue papers
9as first author
9since 2021 · last 2025
0000-0003-4907-7888ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 8 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Spurious Cues in RVL-CDIP and Tobacco3482 Document Classification: The Case of ID CodesabstractRVL-CDIP and Tobacco3482 are commonly used document classification benchmarks, but recent work on explainability has revealed that ID codes stamped on the documents in these datasets may be used by machine learning models to learn shortcuts on the classification task. In this paper, we present an in-depth investigation into the influence and impact of these ID codes on model performance. We annotate ID codes in documents from RVL-CDIP and Tobacco3482 and find that shallow learning models can achieve classification accuracy scores of roughly 40% on RVL-CDIP and 60% on Tobacco3482 using only features derived from the ID codes. We also find that a state-of-the-art document classifier sees a performance drop of 11 accuracy points on RVL-CDIP when ID codes are removed from the data. Finally, we train an ID code detection model in order to remove ID codes from RVL-CDIP and Tobacco3482 and make this data publicly available. Stefan Larson, Sharad Duwal, Brian Vilnrotter, Gayatri Chakkithara, Vedant Padwal, Kevin Leach |
DocEng | 1 |
| 2025 | Document Classification using File NamesabstractRapid document classification is critical in several time-sensitive applications like digital forensics and large-scale media classification. Traditional approaches that rely on heavy-duty deep learning models fall short due to high inference times over vast input datasets and computational resources associated with analyzing whole documents. In this paper, we present a method using lightweight supervised learning models, combined with a TF-IDF feature extraction-based tokenization method, to accurately and efficiently classify documents based solely on their file names, which substantially reduces inference time. Experiments on two datasets introduced in this paper show that our file name classifiers correctly predict more than 90% of in-scope documents with 99.63% and 96.57% accuracy while being 442x faster than more complex models such as DiT. Our results demonstrate that incorporating lightweight file name classification as a front-end to document analysis pipelines can efficiently process vast document datasets in critical scenarios, enabling fast and more reliable document classification. Stefan Larson, Kevin Leach |
DocEng | 2 |
| 2024 | Generating Hard-Negative Out-of-Scope Data with ChatGPT for Intent ClassificationabstractIntent classifiers must be able to distinguish when a user’s utterance does not belong to any supported intent to avoid producing incorrect and unrelated system responses. Although out-of-scope (OOS) detection for intent classifiers has been studied, previous work has not yet studied changes in classifier performance against hard-negative out-of-scope utterances (i.e., inputs that share common features with in-scope data, but are actually out-of-scope). We present an automated technique to generate hard-negative OOS data using ChatGPT. We use our technique to build five new hard-negative OOS datasets, and evaluate each against three benchmark intent classifiers. We show that classifiers struggle to correctly identify hard-negative OOS utterances more than general OOS utterances. Finally, we show that incorporating hard-negative OOS data for training improves model robustness when detecting hard-negative OOS data and general OOS data. Our technique, datasets, and evaluation address an important void in the field, offering a straightforward and inexpensive way to collect hard-negative OOS data and improve intent classifiers’ robustness. Stefan Larson, Kevin Leach |
LREC/COLING | 2 |
| 2024 | De-Identification of Sensitive Personal Data in Datasets Derived from IIT-CDIPabstractStefan Larson, Nicole Cornehl Lima, Santiago Pedroza Diaz, Amogh Manoj Joshi, Siddharth Betala, Jamiu Tunde Suleiman, Yash Mathur, Kaushal Kumar Prajapati, Ramla Alakraa, Junjie Shen, Temi Okotore, Kevin Leach. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Stefan Larson, Nicole Lima, Santiago Diaz, Amogh Manoj Joshi, Siddharth Betala, Jamiu Suleiman, Yash Mathur, Kaushal Prajapati, Ramla Alakraa, Junjie Shen 0011, Temi Okotore, Kevin Leach |
EMNLP | 1 |
| 2023 | On Evaluation of Document Classifiers using RVL-CDIPabstractThe RVL-CDIP benchmark is widely used for measuring performance on the task of document classification.Despite its widespread use, we reveal several undesirable characteristics of the RVL-CDIP benchmark.These include (1) substantial amounts of label noise, which we estimate to be 8.1% (ranging between 1.6% to 16.9% per document category); (2) presence of many ambiguous or multi-label documents; (3) a large overlap between test and train splits, which can inflate model performance metrics; and (4) presence of sensitive personally-identifiable information like US Social Security numbers (SSNs).We argue that there is a risk in using RVL-CDIP for benchmarking document classifiers, as its limited scope, presence of errors (state-of-the-art models now achieve accuracy error rates that are within our estimated label error rate), and lack of diversity make it less than ideal for benchmarking.We further advocate for the creation of a new document classification benchmark, and provide recommendations for what characteristics such a resource should include. Model (Reported by) Modality Accuracy Stefan Larson, Gordon Lim, Kevin Leach |
EACL | 1 |
| 2023 | Augraphy: A Data Augmentation Library for Document Images
Alexander Groleau, Kok Wei Chee, Stefan Larson, Samay Maini, Jonathan Boarman |
ICDAR (3) | 3 |
| 2022 | Evaluating Out-of-Distribution Performance on Document Image ClassifiersabstractThe ability of a document classifier to handle inputs that are drawn from a distribution different from the training distribution is crucial for robust deployment and generalizability. The RVL-CDIP corpus is the de facto standard benchmark for document classification, yet to our knowledge all studies that use this corpus do not include evaluation on out-of-distribution documents. In this paper, we curate and release a new out-of-distribution benchmark for evaluating out-of-distribution performance for document classifiers. Our new out-of-distribution benchmark consists of two types of documents: those that are not part of any of the 16 in-domain RVL-CDIP categories (RVL-CDIP-O), and those that are one of the 16 in-domain categories yet are drawn from a distribution different from that of the original RVL-CDIP dataset (RVL-CDIP-N). While prior work on document classification for in-domain RVL-CDIP documents reports high accuracy scores, we find that these models exhibit accuracy drops of between roughly 15-30% on our new out-of-domain RVL-CDIP-N benchmark, and further struggle to distinguish between in-domain RVL-CDIP-N and out-of-domain RVL-CDIP-O inputs. Our new benchmark provides researchers with a valuable new resource for analyzing out-of-distribution performance on document classifiers. Stefan Larson, Gordon Lim, Yutong Ai, David Kuang, Kevin Leach |
NeurIPS | 1 |
| 2022 | Redwood: Using Collision Detection to Grow a Large-Scale Intent Classification DatasetabstractDialog systems must be capable of incorporating new skills via updates over time in order to reflect new use cases or deployment scenarios.Similarly, developers of such ML-driven systems need to be able to add new training data to an already-existing dataset to support these new skills.In intent classification systems, problems can arise if training data for a new skill's intent overlaps semantically with an alreadyexisting intent.We call such cases collisions.This paper introduces the task of intent collision detection between multiple datasets for the purposes of growing a system's skillset.We introduce several methods for detecting collisions, and evaluate our methods on real datasets that exhibit collisions.To highlight the need for intent collision detection, we show that model performance suffers if new data is added in such a way that does not arbitrate colliding intents.Finally, we use collision detection to construct and benchmark a new dataset, Redwood, which is composed of 451 intent categories from 13 original intent classification datasets, making it the largest publicly available intent classification benchmark. Stefan Larson, Kevin Leach |
SIGDIAL | 1 |
| 2021 | LSOIE: A Large-Scale Dataset for Supervised Open Information ExtractionabstractOpen Information Extraction (OIE) systemsseek to compress the factual propositions of a sentence into a series of n-ary tuples.These tuples are useful for downstream tasks in natural language processing like knowledge base creation, textual entailment, and natural language understanding.However, current OIE datasets are limited in both size and diversity.We introduce a new dataset by converting the QA-SRL 2.0 dataset to a large-scale OIE dataset (LSOIE).Our LSOIE dataset is 20 times larger than the next largest human-annotated OIE dataset.We construct and evaluate several benchmark OIE models on LSOIE, providing baselines for future improvements on the task.Our LSOIE data, models, and code are made publicly available.1 Jacob Solawetz, Stefan Larson |
EACL | 2 |
| 2020 | Inconsistencies in Crowdsourced Slot-Filling Annotations: A Typology and Identification MethodsabstractSlot-filling models in task-driven dialog systems rely on carefully annotated training data.However, annotations by crowd workers are often inconsistent or contain errors.Simple solutions like manually checking annotations or having multiple workers label each sample are expensive and waste effort on samples that are correct.If we can identify inconsistencies, we can focus effort where it is needed.Toward this end, we define six inconsistency types in slot-filling annotations.Using three new noisy crowd-annotated datasets, we show that a wide range of inconsistencies occur and can impact system performance if not addressed.We then introduce automatic methods of identifying inconsistencies.Experiments on our new datasets show that these methods effectively reveal inconsistencies in data, though there is further scope for improvement. Stefan Larson, Adrian Cheung, Anish Mahendran, Kevin Leach, Jonathan K. Kummerfeld |
COLING | 1 |
| 2020 | Iterative Feature Mining for Constraint-Based Data Collection to Increase Data Diversity and Model RobustnessabstractStefan Larson, Anthony Zheng, Anish Mahendran, Rishi Tekriwal, Adrian Cheung, Eric Guldan, Kevin Leach, Jonathan K. Kummerfeld. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Stefan Larson, Anthony Zheng, Anish Mahendran, Rishi Tekriwal, Adrian Cheung, Eric Guldan, Kevin Leach, Jonathan K. Kummerfeld |
EMNLP (1) | 1 |
| 2020 | Data Query Language and Corpus Tools for Slot-Filling and Intent Classification DataabstractTypical machine learning approaches to developing task-oriented dialog systems require the collection and management of large amounts of training data, especially for the tasks of intent classification and slot-filling. Managing this data can be cumbersome without dedicated tools to help the dialog system designer understand the nature of the data. This paper presents a toolkit for analyzing slot-filling and intent classification corpora. We present a toolkit that includes (1) a new lightweight and readable data and file format for intent classification and slot-filling corpora, (2) a new query language for searching intent classification and slot-filling corpora, and (3) tools for understanding the structure and makeup for such corpora. We apply our toolkit to several well-known NLU datasets, and demonstrate that our toolkit can be used to uncover interesting and surprising insights. By releasing our toolkit to the research community, we hope to enable others to develop more robust and intelligent slot-filling and intent classification models. Stefan Larson, Eric Guldan, Kevin Leach |
LREC | 1 |
| 2019 | An Evaluation Dataset for Intent Classification and Out-of-Scope PredictionabstractStefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang, Jason Mars. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Stefan Larson, Anish Mahendran, Joseph Peper, Christopher Clarke, Andrew Lee 0001, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael Laurenzano, Lingjia Tang, Jason Mars |
EMNLP/IJCNLP (1) | 1 |