VLDB 2026 Research / reviewers in the wild / expert
Jason D. Yeatman
dblp:24/10396
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-2686-1293ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 74% Vision and language · 26% | |
| Human-computer interaction and pervasive computing
1 paper |
Learning and educational technologies · 100% |
Topics — the 3 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › large language model evaluation
human-model comparison |
0.8 | 1 | 2024 | DevBench: A multimodal developmental benchmark for language learning · NeurIPS 2024 |
Computer vision › Vision and language
vision-language model |
0.8 | 1 | 2024 | DevBench: A multimodal developmental benchmark for language learning · NeurIPS 2024 |
Natural language and speech › Language models and text generation › LLM agents
LLM-based simulation |
0.7 | 1 | 2023 | Generating and Evaluating Tests for K-12 Students with Language Model Simulations: A Case Study on Sentence Reading Efficiency · EMNLP 2023 |
Methods — techniques the papers use, named apart from their topics
optimal transport · 1.3fine-tuning · 1.3behavioral data analysis · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A software ecosystem for brain tractometry processing, analysis, and insightabstractTractometry uses diffusion-weighted magnetic resonance imaging (dMRI) to assess physical properties of brain connections. Here, we present an integrative ecosystem of software that performs all steps of tractometry: post-processing of dMRI data, delineation of major white matter pathways, and modeling of the tissue properties within them. This ecosystem also provides a set of interoperable and extensible tools for visualization and interpretation of the results that extract insights from these measurements. These include novel machine learning and statistical analysis methods adapted to the characteristic structure of tract-based data. We benchmark the performance of these statistical analysis methods in different datasets and analysis tasks, including hypothesis testing on group differences and predictive analysis of subject age. We also demonstrate that computational advances implemented in the software offer orders of magnitude of acceleration. Taken together, these open-source software tools-freely available at https://tractometry.org-provide a transformative environment for the analysis of dMRI data. John Kruper, Adam C. Richie-Halford, Joanna Qiao, Asa Gilmore, Kelly Chang, Mareike Grotheer, Ethan Roy, Sendy Caffarra, Teresa Gómez, Sam Chou, Matthew Cieslak, Serge Koudoro, Eleftherios Garyfallidis, Theodore D. Satterthwaite, Jason D. Yeatman, Ariel Rokem |
PLoS Comput. Biol. | 15 |
| 2024 | DevBench: A multimodal developmental benchmark for language learningabstractHow (dis)similar are the learning trajectories of vision–language models and children? Recent modeling work has attempted to understand the gap between models’ and humans’ data efficiency by constructing models trained on less data, especially multimodal naturalistic data. However, such models are often evaluated on adult-level benchmarks, with limited breadth in language abilities tested, and without direct comparison to behavioral data. We introduce DevBench, a multimodal benchmark comprising seven language evaluation tasks spanning the domains of lexical, syntactic, and semantic ability, with behavioral data from both children and adults. We evaluate a set of vision–language models on these tasks, comparing models and humans on their response patterns, not their absolute performance. Across tasks, models exhibit variation in their closeness to human response patterns, and models that perform better on a task also more closely resemble human behavioral responses. We also examine the developmental trajectory of OpenCLIP over training, finding that greater training results in closer approximations to adult response patterns. DevBench thus provides a benchmark for comparing models to human language development. These comparisons highlight ways in which model and human language learning processes diverge, providing insight into entry points for improving language models. Alvin Wei Ming Tan, Chunhua Yu, Bria Long, Wanjing Ma, Tonya Murray, Rebecca D. Silverman, Jason D. Yeatman, Michael C. Frank |
NeurIPS | 7 |
| 2023 | Generating and Evaluating Tests for K-12 Students with Language Model Simulations: A Case Study on Sentence Reading EfficiencyabstractDeveloping an educational test can be expensive and time-consuming, as each item must be written by experts and then evaluated by collecting hundreds of student responses.Moreover, many tests require multiple distinct sets of questions administered throughout the school year to closely monitor students' progress, known as parallel tests.In this study, we focus on tests of silent sentence reading efficiency, used to assess students' reading ability over time.To generate high-quality parallel tests, we propose to fine-tune large language models (LLMs) to simulate how previous students would have responded to unseen items.With these simulated responses, we can estimate each item's difficulty and ambiguity.We first use GPT-4 to generate new test items following a list of expert-developed rules and then apply a fine-tuned LLM to filter the items based on criteria from psychological measurements.We also propose an optimal-transport-inspired technique for generating parallel tests and show the generated tests closely correspond to the original test's difficulty and reliability based on crowdworker responses.Our evaluation of a generated test with 234 students from grades 2 to 8 produces test scores highly correlated (r=0.93) to those of a standard test form written by human experts and evaluated across thousands of K-12 students. Eric Zelikman, Wanjing Anya Ma, Jasmine E. Tran, Diyi Yang, Jason D. Yeatman, Nick Haber |
EMNLP | 5 |
| 2022 | Automated generation of sentence reading fluency test items
Julia White 0001, Amy Burkhardt, Jason D. Yeatman, Noah D. Goodman |
CogSci | 3 |
| 2021 | Multidimensional analysis and detection of informative features in human brain white matterabstractThe white matter contains long-range connections between different brain regions and the organization of these connections holds important implications for brain function in health and disease. Tractometry uses diffusion-weighted magnetic resonance imaging (dMRI) to quantify tissue properties along the trajectories of these connections. Statistical inference from tractometry usually either averages these quantities along the length of each fiber bundle or computes regression models separately for each point along every one of the bundles. These approaches are limited in their sensitivity, in the former case, or in their statistical power, in the latter. We developed a method based on the sparse group lasso (SGL) that takes into account tissue properties along all of the bundles and selects informative features by enforcing both global and bundle-level sparsity. We demonstrate the performance of the method in two settings: i) in a classification setting, patients with amyotrophic lateral sclerosis (ALS) are accurately distinguished from matched controls. Furthermore, SGL identifies the corticospinal tract as important for this classification, correctly finding the parts of the white matter known to be affected by the disease. ii) In a regression setting, SGL accurately predicts "brain age." In this case, the weights are distributed throughout the white matter indicating that many different regions of the white matter change over the lifespan. Thus, SGL leverages the multivariate relationships between diffusion properties in multiple bundles to make accurate phenotypic predictions while simultaneously discovering the most relevant features of the white matter. Adam C. Richie-Halford, Jason D. Yeatman, Noah Simon, Ariel Rokem |
PLoS Comput. Biol. | 2 |