VLDB 2026 Research / reviewers in the wild / expert
David S. Hayden
dblp:34/46
· DBLP profile ↗
10ranked-venue papers
7as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 38% Autonomous driving · 25% Trustworthy machine learning · 10% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 38% Algorithmic game theory and mechanism design · 38% Information theory · 23% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.7 | 2 | 2025 | Generative Data Mining with Longtail-Guided Diffusion · ICML 2025 Causal Composition Diffusion Model for Closed-loop Traffic Generation · CVPR 2025 |
Machine learning › Generative modeling
autoregressive model |
0.9 | 1 | 2025 | DriveGPT: Scaling Autoregressive Behavior Models for Driving · ICML 2025 |
Natural language and speech › Language models and text generation › neural language model
autoregressive transformer |
0.9 | 1 | 2025 | DriveGPT: Scaling Autoregressive Behavior Models for Driving · ICML 2025 |
Robotics › Autonomous driving
behavior modeling |
0.9 | 1 | 2025 | DriveGPT: Scaling Autoregressive Behavior Models for Driving · ICML 2025 |
Machine learning › Deep learning architectures and training
data augmentation |
0.9 | 1 | 2025 | Generative Data Mining with Longtail-Guided Diffusion · ICML 2025 |
Machine learning › Trustworthy machine learning › uncertainty estimation
epistemic uncertainty |
0.9 | 1 | 2025 | Generative Data Mining with Longtail-Guided Diffusion · ICML 2025 |
Machine learning › Generative modeling › diffusion model
latent diffusion model |
0.9 | 1 | 2025 | Generative Data Mining with Longtail-Guided Diffusion · ICML 2025 |
Machine learning › Efficient and distributed learning › large-scale learning
model scaling |
0.9 | 1 | 2025 | DriveGPT: Scaling Autoregressive Behavior Models for Driving · ICML 2025 |
Machine learning › Generative modeling › diffusion model
structure-guided diffusion |
0.9 | 1 | 2025 | Causal Composition Diffusion Model for Closed-loop Traffic Generation · CVPR 2025 |
Robotics › Autonomous driving › scenario generation
traffic scenario generation |
0.9 | 1 | 2025 | Causal Composition Diffusion Model for Closed-loop Traffic Generation · CVPR 2025 |
Robotics › Autonomous driving
trajectory prediction |
0.9 | 1 | 2025 | DriveGPT: Scaling Autoregressive Behavior Models for Driving · ICML 2025 |
Computer vision › 3D vision › 3d motion analysis
articulated motion analysis |
0.4 | 1 | 2020 | Nonparametric Object and Parts Modeling With Lie Group Dynamics · CVPR 2020 |
Algorithmic game theory and mechanism design
resource allocation |
0.4 | 1 | 2020 | Sequential Bayesian Experimental Design with Variable Cost Structure · NeurIPS 2020 |
Machine learning › Trustworthy machine learning
safety evaluation |
0.3 | 1 | 2025 | Causal Composition Diffusion Model for Closed-loop Traffic Generation · CVPR 2025 |
Robotics › Autonomous driving › autonomous vehicle testing
simulation-based testing |
0.3 | 1 | 2025 | Causal Composition Diffusion Model for Closed-loop Traffic Generation · CVPR 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model |
0.1 | 1 | 2020 | Nonparametric Object and Parts Modeling With Lie Group Dynamics · CVPR 2020 |
Information theory › information measures
mutual information |
0.1 | 1 | 2020 | Sequential Bayesian Experimental Design with Variable Cost Structure · NeurIPS 2020 |
Information theory › information measures › mutual information
mutual information bound |
0.1 | 1 | 2020 | Sequential Bayesian Experimental Design with Variable Cost Structure · NeurIPS 2020 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 0.9transformer · 0.9pre-training · 0.9longtail guidance · 0.9latent diffusion · 0.9diffusion model · 0.9constrained optimization · 0.9causal structure learning · 0.9autoregressive modeling · 0.9two-sided bounds · 0.4knapsack optimization · 0.4gibbs sampling · 0.4best-arm identification · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Causal Composition Diffusion Model for Closed-loop Traffic GenerationabstractSimulation is critical for safety evaluation in autonomous driving, particularly in capturing complex interactive behaviors. However, generating realistic and controllable traffic scenarios in long-tail situations remains a significant challenge. Existing generative models suffer from the conflicting objective between user-defined controllability and realism constraints, which is amplified in safety-critical contexts. In this work, we introduce the Causal Compositional Diffusion Model (CCDiff), a structure-guided diffusion framework to address these challenges. We first formulate the learning of controllable and realistic closed-loop simulation as a constrained optimization problem. Then, CCDiff maximizes controllability while adhering to realism by automatically identifying and injecting causal structures directly into the diffusion process, providing structured guidance to enhance both realism and controllability. Through rigorous evaluations on benchmark datasets and in a closed-loop simulator, CCDiff demonstrates substantial gains over state-of-the-art approaches in generating realistic and user-preferred trajectories. Our results show CCDiff’s effectiveness in extracting and leveraging causal structures, showing improved closed-loop performance based on key metrics such as collision rate, off-road rate, FDE, and comfort. For more details, welcome to check our project website. Haohong Lin, Tung Phan, David S. Hayden, Huan Zhang 0001, Ding Zhao, Siddhartha S. Srinivasa, Eric M. Wolff, Hongge Chen |
CVPR | 4 |
| 2025 | Generative Data Mining with Longtail-Guided DiffusionabstractIt is difficult to anticipate the myriad challenges that a predictive model will encounter once deployed. Common practice entails a reactive, cyclical approach: model deployment, data mining, and retraining. We instead develop a proactive longtail discovery process by imagining additional data during training. In particular, we develop general model-based longtail signals, including a differentiable, single forward pass formulation of epistemic uncertainty that does not impact model parameters or predictive performance but can flag rare or hard inputs. We leverage these signals as guidance to generate additional training data from a latent diffusion model in a process we call Longtail Guidance (LTG). Crucially, we can perform LTG without retraining the diffusion model or the predictive model, and we do not need to expose the predictive model to intermediate diffusion states. Data generated by LTG exhibit semantically meaningful variation, yield significant generalization improvements on numerous image classification benchmarks, and can be analyzed by a VLM to proactively discover, textually explain, and address conceptual gaps in a deployed predictive model. David S. Hayden, Mao Ye 0006, Timur Garipov, Gregory P. Meyer, Carl Vondrick, Yuning Chai, Eric M. Wolff, Siddhartha S. Srinivasa |
ICML | 1 |
| 2025 | DriveGPT: Scaling Autoregressive Behavior Models for DrivingabstractWe present DriveGPT, a scalable behavior model for autonomous driving. We model driving as a sequential decision-making task, and learn a transformer model to predict future agent states as tokens in an autoregressive fashion. We scale up our model parameters and training data by multiple orders of magnitude, enabling us to explore the scaling properties in terms of dataset size, model parameters, and compute. We evaluate DriveGPT across different scales in a planning task, through both quantitative metrics and qualitative examples, including closed-loop driving in complex real-world scenarios. In a separate prediction task, DriveGPT outperforms state-of-the-art baselines and exhibits improved performance by pretraining on a large-scale dataset, further validating the benefits of data scaling. Eric M. Wolff, Paul Vernaza, Tung Phan-Minh, Hongge Chen, David S. Hayden, Mark Edmonds, Brian Pierce, Xinxin Chen, Pratik Elias Jacob, Xiaobai Chen, Chingiz Tairbekov, Pratik Agarwal, Tianshi Gao, Yuning Chai, Siddhartha S. Srinivasa |
ICML | 6 |
| 2020 | Nonparametric Object and Parts Modeling With Lie Group DynamicsabstractArticulated motion analysis often utilizes strong prior knowledge such as a known or trained parts model for humans. Yet, the world contains a variety of articulating objects--mammals, insects, mechanized structures--where the number and configuration of parts for a particular object is unknown in advance. Here, we relax such strong assumptions via an unsupervised, Bayesian nonparametric parts model that infers an unknown number of parts with motions coupled by a body dynamic and parameterized by SE(D), the Lie group of rigid transformations. We derive an inference procedure that utilizes short observation sequences (image, depth, point cloud or mesh) of an object in motion without need for markers or learned body models. Efficient Gibbs decompositions for inference over distributions on SE(D) demonstrate robust part decompositions of moving objects under both 3D and 2D observation models. The inferred representation permits novel analysis, such as object segmentation by relative part motion, and transfers to new observations of the same object type. David S. Hayden, Jason L. Pacheco, John W. Fisher III |
CVPR | 1 |
| 2020 | Sequential Bayesian Experimental Design with Variable Cost StructureabstractMutual information (MI) is a commonly adopted utility function in Bayesian optimal experimental design (BOED). While theoretically appealing, MI evaluation poses a significant computational burden for most real world applications. As a result, many algorithms utilize MI bounds as proxies that lack regret-style guarantees. Here, we utilize two-sided bounds to provide such guarantees. Bounds are successively refined/tightened through additional computation until a desired guarantee is achieved. We consider the problem of adaptively allocating computational resources in BOED. Our approach achieves the same guarantee as existing methods, but with fewer evaluations of the costly MI reward. We adapt knapsack optimization of best arm identification problems, with important differences that impact overall algorithm design and performance. First, observations of MI rewards are biased. Second, evaluating experiments incurs shared costs amongst all experiments (posterior sampling) in addition to per experiment costs that may vary with increasing evaluation. We propose and demonstrate an algorithm that accounts for these variable costs in the refinement decision. Sue Zheng, David S. Hayden, Jason L. Pacheco, John W. Fisher III |
NeurIPS | 2 |
| 2012 | Using Clustering and Metric Learning to Improve Science Return of Remote Sensed ImageryabstractCurrent and proposed remote space missions, such as the proposed aerial exploration of Titan by an aerobot, often can collect more data than can be communicated back to Earth. Autonomous selective downlink algorithms can choose informative subsets of data to improve the science value of these bandwidth-limited transmissions. This requires statistical descriptors of the data that reflect very abstract and subtle distinctions in science content. We propose a metric learning strategy that teaches algorithms how best to cluster new data based on training examples supplied by domain scientists. We demonstrate that clustering informed by metric learning produces results that more closely match multiple scientists’ labelings of aerial data than do clusterings based on random or periodic sampling. A new metric-learning strategy accommodates training sets produced by multiple scientists with different and potentially inconsistent mission objectives. Our methods are fit for current spacecraft processors (e.g., RAD750) and would further benefit from more advanced spacecraft processor architectures, such as OPERA. David S. Hayden, Steve A. Chien, David R. Thompson 0001, Rebecca Castaño |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2011 | Note-taker 3.0: an assistive technology enabling students who are legally blind to take notes in classabstractWhile the Americans with Disabilities Act mandates that universities provide visually disabled students with human note-takers, studies have shown that it is vital that students take their own notes during classroom lectures. Because students are cognitively engaged while taking notes, their retention is better, even if they never review their notes after class. However, students with visual disabilities are at a disadvantage compared to their sighted peers when taking notes - especially in fast-pace class presentations. They find it more difficult to rapidly switch back and forth between viewing the front of the room and viewing the notes they are taking. Currently available assistive technologies do not adequately address this need to rapidly switch back and forth. This paper presents the results of a 3-year study aimed at the development of an assistive technology that is specifically aimed at allowing students with visual disabilities to take handwritten and/or typed notes in the classroom, without relying on any classroom infrastructure, or any special accommodations by the lecturer or the institution. David S. Hayden, Michael J. Astrauskas, Liqing Zhou, John A. Black Jr. |
ASSETS | 1 |
| 2010 | Note-taker 2.0: the next step toward enabling students who are legally blind to take notes in classabstractIn-class note-taking is a vital learning activity in secondary and post-secondary classrooms. The process of note-taking helps students stay focused on the instruction, forces them to cognitively process what is being presented, and better retain what has been taught, even if they never refer to their notes after the class. However, note-taking is difficult for students with low vision, or who are legally blind for two reasons. First, they are less able to see what is being presented at the front of them room, and second, they must repeatedly switch between the far-sight task of viewing the front of the room, and the near-sight task of taking notes. This paper describes ongoing research aimed at developing a portable assistive device (called the Note-Taker) that a student can take to class, to assist in the process of taking notes. It describes the principles that have guided the development of the proof-of-concept Note-Taker prototype and the Note-Taker 2.0 prototype. Initial testing of those prototypes has been encouraging, but some significant problems remain to be solved. Proposed solutions are currently being implemented, and appear to be effective. If ongoing usability testing confirms their effectiveness, they will be implemented on the planned Note-Taker 3.0 prototype. David S. Hayden, Liqing Zhou, Michael J. Astrauskas, John A. Black Jr. |
ASSETS | 1 |
| 2008 | Note-taker: enabling students who are legally blind to take notes in classabstractThe act of note-taking is a key component of learning in secondary and post-secondary classrooms. Students who take notes retain information from classroom lectures better, even if they never refer to those notes afterward. However, students who are legally blind, and who wish to take notes in their classrooms are at a disadvantage. Simply equipping classrooms with lecture recording systems does not substitute for note taking, since it does not actively engage the student in note-taking during the lecture. In this paper we detail the problems encountered by one math and computer science student who is legally blind, and we present our proposed solution: the CUbiC Note-Taker, which is a highly portable device that requires no prior classroom setup, and does not require lecturers to adapt their presentations. We also present results from two case studies of the Note-Taker, totaling more than 200 hours of in-class use. David S. Hayden, Dirk Colbry, John A. Black Jr., Sethuraman Panchanathan |
ASSETS | 1 |
| 2008 | Enabling the legally blind in classroom note-takingabstractClassroom note-taking has been shown to be beneficial, even if the student never reviews his/her own notes. Students that are legally blind are thus at a disadvantage because they face significant barriers to note-taking in the classroom. This presentation demonstrates a working prototype of the CUbiC Note Taker, which is a highly portable device that allows students who are legally blind to take their own notes in class without any special in-classroom accommodations, and without requiring lecturers to adapt their presentation in any way. David S. Hayden, Dirk Colbry, John A. Black Jr., Sethuraman Panchanathan |
ASSETS | 1 |