VLDB 2026 Research / reviewers in the wild / expert
Benjamin W. Domingue
dblp:248/9322 · also Ben Domingue
· DBLP profile ↗
8ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-3894-9049ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Trustworthy machine learning · 56% Question answering and dialogue systems · 32% Probabilistic and Bayesian machine learning · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computing education · 100% |
Topics — the 4 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
question generation |
1.0 | 1 | 2026 | Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study · AAAI 2026 |
Machine learning › Trustworthy machine learning
benchmark validity |
0.9 | 1 | 2025 | Fantastic Bugs and Where to Find Them in AI Benchmarks · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.4 | 1 | 2019 | Curve Fitting from Probabilistic Emissions and Applications to Dynamic Item Response Theory · ICDM 2019 |
Computing education › educational assessment
item response theory |
0.4 | 1 | 2019 | Curve Fitting from Probabilistic Emissions and Applications to Dynamic Item Response Theory · ICDM 2019 |
Methods — techniques the papers use, named apart from their topics
iterative refinement · 2.0item response theory · 2.0LLM-generated critique · 2.0statistical analysis · 0.9LLM judge · 0.9variational inference · 0.8grafting · 0.8expectation-maximization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Assessing the Quality of AI-Generated Exams: A Large-Scale Field StudyabstractWhile large language models (LLMs) challenge conventional methods of teaching and learning, they present an exciting opportunity to improve efficiency and scale high-quality instruction. One promising application is the generation of customized exams, tailored to specific course content. There has been significant recent excitement on automatically generating questions using artificial intelligence, but also comparatively little work evaluating the psychometric quality of these items in real-world educational settings. Filling this gap is an important step toward understanding generative AI's role in effective test design. In this study, we introduce and evaluate an iterative refinement strategy for question generation, repeatedly producing, assessing, and improving questions through cycles of LLM-generated critique and revision. We evaluate the quality of these AI-generated questions in a large-scale field study involving 91 classes---covering computer science, mathematics, chemistry, and more---in dozens of colleges across the United States, comprising nearly 1700 students. Our analysis, based on item response theory (IRT), suggests that for students in our sample the AI-generated questions performed comparably to expert-created questions designed for standardized exams. Our results illustrate the power of AI to make high-quality assessments more readily available, benefiting both teachers and students. Calvin Isley, Joshua Gilbert, Evangelos Kassos, Michaela Kocher, Allen Nie, Emma Brunskill, Benjamin W. Domingue, Jake Hofman, Joscha Legewie, Teddy Svoronos, Charlotte Tuminelli, Sharad Goel |
AAAI | 7 |
| 2026 | Understanding Student Effort Using Response-Time Propensities During Problem Solving
Conrad Borchers, Lijin Zhang, Tomohiro Nagashima, Benjamin W. Domingue |
L@S | 5 |
| 2026 | Revisiting the Regularity of Student Learning Rate: Sensitivity to Which Observations Are IncludedabstractMixed-effects models fit to observational practice data are widely used in learning analytics to estimate student-level variation in initial knowledge and learning rate, and the resulting estimates increasingly inform substantive claims about learners. We examine whether such estimates can be read as properties of learners or whether they depend on choices about which observations the model is fit to. As a case study, we revisit the ''astonishing regularity'' reported by Koedinger et al. (2023): that students vary substantially in initial knowledge but much less in learning rate. The finding is based on fits of the individual Additive Factors Model (iAFM) to 27 educational datasets, and rests on a model-derived estimate of student-level learning-rate variation being small in absolute terms. We refit the same model on the same datasets under two specifications, each varying how much of each student's practice on a given skill is used in fitting. The estimate of student-level variation in initial knowledge stays approximately stable across both specifications. The estimate of student-level variation in learning rate does not: it inflates by a median of 118% under one specification and is several times larger under the other. The same model, fit to the same data, returns substantially different estimates of how much students vary in learning rate depending on which observations are included. When estimates from mixed-effects models on observational practice data are used to support substantive claims about learners, sensitivity to such choices deserves a central place in how those estimates are reported and read. Guilherme Lichand, Cristina Barnard, Lucas Klotz, Candace Thille, Yunsung Kim, Benjamin W. Domingue |
L@S | 7 |
| 2025 | Fantastic Bugs and Where to Find Them in AI BenchmarksabstractBenchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of benchmark questions is not only infeasible but also a critical bottleneck for reliable evaluation. In this work, we introduce a framework for systematic benchmark revision that leverages statistical analysis of response patterns to flag potentially invalid questions for further expert review. Our approach builds on a core assumption commonly used in AI evaluations that the mean score sufficiently summarizes model performance. This implies a unidimensional latent construct underlying the measurement experiment, yielding expected ranges for various statistics for each item. When empirically estimated values for these statistics fall outside the expected range for an item, the item is more likely to be problematic. Across nine widely used benchmarks, our method guides expert review to identify problematic questions with up to 84\% precision. In addition, we introduce an LLM‑judge first pass to review questions, further reducing human effort. Together, these components provide an efficient and scalable framework for systematic benchmark revision. Sang T. Truong, Yuheng Tu 0001, Michael Hardy, Anka Reuel, Zeyu Tang 0002, Jirayu Burapacheep, Jonathan Perera, Chibuike Uwakwe, Benjamin W. Domingue, Nick Haber, Oluwasanmi Koyejo |
NeurIPS | 9 |
| 2020 | AI and Holistic Review: Informing Human Reading in College AdmissionsabstractCollege admissions in the United States is carried out by a human-centered method of evaluation known as holistic review, which typically involves reading original narrative essays submitted by each applicant. The legitimacy and fairness of holistic review, which gives human readers significant discretion over determining each applicant's fitness for admission, has been repeatedly challenged in courtrooms and the public sphere. Using a unique corpus of 283,676 application essays submitted to a large, selective, state university system between 2015 and 2016, we assess the extent to which applicant demographic characteristics can be inferred from application essays. We find a relatively interpretable classifier (logistic regression) was able to predict gender and household income with high levels of accuracy. Findings suggest that data auditing might be useful in informing holistic review, and perhaps other evaluative systems, by checking potential bias in human or computational readings. A. J. Alvero, Noah Arthurs, Anthony Lising Antonio, Benjamin W. Domingue, Ben Gebre-Medhin, Sonia Giebel, Mitchell L. Stevens |
AIES | 4 |
| 2020 | Variational Item Response Theory: Fast, Accurate, and Expressive
Mike Wu, Richard Lee Davis, Benjamin W. Domingue, Chris Piech, Noah D. Goodman |
EDM | 3 |
| 2019 | Curve Fitting from Probabilistic Emissions and Applications to Dynamic Item Response TheoryabstractItem response theory (IRT) models are widely used in psychometrics and educational measurement, being deployed in many high stakes tests such as the GRE aptitude test. IRT has largely focused on estimation of a single latent trait (e.g. ability) that remains static through the collection of item responses. However, in contemporary settings where item responses are being continuously collected, such as Massive Open Online Courses (MOOCs), interest will naturally be on the dynamics of ability, thus complicating usage of traditional IRT models. We propose DynAEsti, an augmentation of the traditional IRT Expectation Maximization algorithm that allows ability to be a continuously varying curve over time. In the process, we develop CurvFiFE, a novel non-parametric continuous-time technique that handles the curve-fitting/regression problem extended to address more general probabilistic emissions (as opposed to simply noisy data points). Furthermore, to accomplish this, we develop a novel technique called grafting, which can successfully approximate distributions represented by graphical models when other popular techniques like Loopy Belief Propogation (LBP) and Variational Inference (VI) fail. The performance of DynAEsti is evaluated through simulation, where we achieve results comparable to the optimal of what is observed in the static ability scenario. Finally, DynAEsti is applied to a longitudinal performance dataset (80-years of competitive golf at the 18-hole Masters Tournament) to demonstrate its ability to recover key properties of human performance and the heterogeneous characteristics of the different holes. Python code for CurvFiFE and DynAEsti is publicly available at github.com/chausies/DynAEstiAndCurvFiFE. Ajay Shanker Tripathi, Benjamin W. Domingue |
ICDM | 2 |
| 2017 | Making the Grade: How Learner Engagement Changes After Passing a Course
David Lang, Benjamin W. Domingue, Alex Kindel, Andreas Paepcke |
EDM | 2 |