Benjamin W. Domingue

dblp:248/9322 · also Ben Domingue · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-3894-9049ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 56% Question answering and dialogue systems · 32% Probabilistic and Bayesian machine learning · 12%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computing education · 100%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
question generation
1.012026
Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study · AAAI 2026
Machine learning › Trustworthy machine learning
benchmark validity
0.912025
Fantastic Bugs and Where to Find Them in AI Benchmarks · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.412019
Curve Fitting from Probabilistic Emissions and Applications to Dynamic Item Response Theory · ICDM 2019
Computing education › educational assessment
item response theory
0.412019
Curve Fitting from Probabilistic Emissions and Applications to Dynamic Item Response Theory · ICDM 2019

Methods — techniques the papers use, named apart from their topics

iterative refinement · 2.0item response theory · 2.0LLM-generated critique · 2.0statistical analysis · 0.9LLM judge · 0.9variational inference · 0.8grafting · 0.8expectation-maximization · 0.8
YearPublicationVenuePosition
2026 Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study
abstract
While large language models (LLMs) challenge conventional methods of teaching and learning, they present an exciting opportunity to improve efficiency and scale high-quality instruction. One promising application is the generation of customized exams, tailored to specific course content. There has been significant recent excitement on automatically generating questions using artificial intelligence, but also comparatively little work evaluating the psychometric quality of these items in real-world educational settings. Filling this gap is an important step toward understanding generative AI's role in effective test design. In this study, we introduce and evaluate an iterative refinement strategy for question generation, repeatedly producing, assessing, and improving questions through cycles of LLM-generated critique and revision. We evaluate the quality of these AI-generated questions in a large-scale field study involving 91 classes---covering computer science, mathematics, chemistry, and more---in dozens of colleges across the United States, comprising nearly 1700 students. Our analysis, based on item response theory (IRT), suggests that for students in our sample the AI-generated questions performed comparably to expert-created questions designed for standardized exams. Our results illustrate the power of AI to make high-quality assessments more readily available, benefiting both teachers and students.
Calvin Isley, Joshua Gilbert, Evangelos Kassos, Michaela Kocher, Allen Nie, Emma Brunskill, Benjamin W. Domingue, Jake Hofman, Joscha Legewie, Teddy Svoronos, Charlotte Tuminelli, Sharad Goel
AAAI7
2026 Understanding Student Effort Using Response-Time Propensities During Problem Solving
Conrad Borchers, Lijin Zhang, Tomohiro Nagashima, Benjamin W. Domingue
L@S5
2026 Revisiting the Regularity of Student Learning Rate: Sensitivity to Which Observations Are Included
abstract
Mixed-effects models fit to observational practice data are widely used in learning analytics to estimate student-level variation in initial knowledge and learning rate, and the resulting estimates increasingly inform substantive claims about learners. We examine whether such estimates can be read as properties of learners or whether they depend on choices about which observations the model is fit to. As a case study, we revisit the ''astonishing regularity'' reported by Koedinger et al. (2023): that students vary substantially in initial knowledge but much less in learning rate. The finding is based on fits of the individual Additive Factors Model (iAFM) to 27 educational datasets, and rests on a model-derived estimate of student-level learning-rate variation being small in absolute terms. We refit the same model on the same datasets under two specifications, each varying how much of each student's practice on a given skill is used in fitting. The estimate of student-level variation in initial knowledge stays approximately stable across both specifications. The estimate of student-level variation in learning rate does not: it inflates by a median of 118% under one specification and is several times larger under the other. The same model, fit to the same data, returns substantially different estimates of how much students vary in learning rate depending on which observations are included. When estimates from mixed-effects models on observational practice data are used to support substantive claims about learners, sensitivity to such choices deserves a central place in how those estimates are reported and read.
Guilherme Lichand, Cristina Barnard, Lucas Klotz, Candace Thille, Yunsung Kim, Benjamin W. Domingue
L@S7
2025 Fantastic Bugs and Where to Find Them in AI Benchmarks
abstract
Benchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of benchmark questions is not only infeasible but also a critical bottleneck for reliable evaluation. In this work, we introduce a framework for systematic benchmark revision that leverages statistical analysis of response patterns to flag potentially invalid questions for further expert review. Our approach builds on a core assumption commonly used in AI evaluations that the mean score sufficiently summarizes model performance. This implies a unidimensional latent construct underlying the measurement experiment, yielding expected ranges for various statistics for each item. When empirically estimated values for these statistics fall outside the expected range for an item, the item is more likely to be problematic. Across nine widely used benchmarks, our method guides expert review to identify problematic questions with up to 84\% precision. In addition, we introduce an LLM‑judge first pass to review questions, further reducing human effort. Together, these components provide an efficient and scalable framework for systematic benchmark revision.
Sang T. Truong, Yuheng Tu 0001, Michael Hardy, Anka Reuel, Zeyu Tang 0002, Jirayu Burapacheep, Jonathan Perera, Chibuike Uwakwe, Benjamin W. Domingue, Nick Haber, Oluwasanmi Koyejo
NeurIPS9
2020 AI and Holistic Review: Informing Human Reading in College Admissions
abstract
College admissions in the United States is carried out by a human-centered method of evaluation known as holistic review, which typically involves reading original narrative essays submitted by each applicant. The legitimacy and fairness of holistic review, which gives human readers significant discretion over determining each applicant's fitness for admission, has been repeatedly challenged in courtrooms and the public sphere. Using a unique corpus of 283,676 application essays submitted to a large, selective, state university system between 2015 and 2016, we assess the extent to which applicant demographic characteristics can be inferred from application essays. We find a relatively interpretable classifier (logistic regression) was able to predict gender and household income with high levels of accuracy. Findings suggest that data auditing might be useful in informing holistic review, and perhaps other evaluative systems, by checking potential bias in human or computational readings.
A. J. Alvero, Noah Arthurs, Anthony Lising Antonio, Benjamin W. Domingue, Ben Gebre-Medhin, Sonia Giebel, Mitchell L. Stevens
AIES4
2020 Variational Item Response Theory: Fast, Accurate, and Expressive
Mike Wu, Richard Lee Davis, Benjamin W. Domingue, Chris Piech, Noah D. Goodman
EDM3
2019 Curve Fitting from Probabilistic Emissions and Applications to Dynamic Item Response Theory
abstract
Item response theory (IRT) models are widely used in psychometrics and educational measurement, being deployed in many high stakes tests such as the GRE aptitude test. IRT has largely focused on estimation of a single latent trait (e.g. ability) that remains static through the collection of item responses. However, in contemporary settings where item responses are being continuously collected, such as Massive Open Online Courses (MOOCs), interest will naturally be on the dynamics of ability, thus complicating usage of traditional IRT models. We propose DynAEsti, an augmentation of the traditional IRT Expectation Maximization algorithm that allows ability to be a continuously varying curve over time. In the process, we develop CurvFiFE, a novel non-parametric continuous-time technique that handles the curve-fitting/regression problem extended to address more general probabilistic emissions (as opposed to simply noisy data points). Furthermore, to accomplish this, we develop a novel technique called grafting, which can successfully approximate distributions represented by graphical models when other popular techniques like Loopy Belief Propogation (LBP) and Variational Inference (VI) fail. The performance of DynAEsti is evaluated through simulation, where we achieve results comparable to the optimal of what is observed in the static ability scenario. Finally, DynAEsti is applied to a longitudinal performance dataset (80-years of competitive golf at the 18-hole Masters Tournament) to demonstrate its ability to recover key properties of human performance and the heterogeneous characteristics of the different holes. Python code for CurvFiFE and DynAEsti is publicly available at github.com/chausies/DynAEstiAndCurvFiFE.
Ajay Shanker Tripathi, Benjamin W. Domingue
ICDM2
2017 Making the Grade: How Learner Engagement Changes After Passing a Course
David Lang, Benjamin W. Domingue, Alex Kindel, Andreas Paepcke
EDM2