VLDB 2026 Research / reviewers in the wild / expert
Lynn Vonder Haar
dblp:336/4772 · also Lynn Vonderhaar
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2026
0009-0003-0555-3640ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Survey of Machine Learning Lifecycle Provenance: Models, Approaches, and Tools
Lynn Vonder Haar, Tyler Procko, Omar Ochoa |
ENASE (2) | 1 |
| 2026 | Verifying Machine Learning Testability Requirements with Provenance
Lynn Vonder Haar, Tyler Procko, Omar Ochoa |
ICSOFT | 1 |
| 2025 | Generating and Verifying Synthetic Datasets with Requirements EngineeringabstractWith the rise of generative Artificial Intelligence (AI), Machine Learning (ML) developers are becoming less reliant on real data to train their models. Data insufficiency can be resolved by using synthetic data generated by a diffusion model. However, beyond ad hoc interpretation of a generative model's outputs, there is little assurance of the synthetic data's adherence to the data requirement specifications. Adherence of synthetic data to these specifications is critical given that they describe desired downstream model behavior. Therefore, without proper verification methods for this synthetic data, ML developers cannot be confident in the behavior of the downstream model. This paper presents a verification method for generating synthetic data to train downstream ML models by prompting the generative model using requirement specifications and tracing elements of the output back to the prompt. The purpose of this research is to embed requirements engineering into the data augmentation process to increase the rigor and acceptance of these generative AI models to train downstream ML models. This improves the transparency of the data augmentation process, potentially increasing the trust of stakeholders in the generated data, and the use of generative models for data augmentation in a wider range of applications. This also provides a more traditional approach to synthetic data generation to guide ML developers in augmenting their datasets, thus incorporating a more rigorous engineering process into the ML development, i.e., ML Engineering. Lynn Vonder Haar, Timothy Elvira, Omar Ochoa |
CAIN | 1 |
| 2024 | Exploring Testing Methods for Large Language ModelsabstractLarge Language Models (LLMs) are extensive aggregations of human language, designed to understand and generate sophisticated text. LLMs are becoming ubiquitous in a range of applications, from social media to code generation. With their immense size, LLMs face scalability challenges, making testing methods particularly difficult to implement effectively. Traditional machine learning and software testing methods, derived and adapted for LLMs, test these models to a point; however, they still struggle to accurately capture the full complexity of model behavior. This paper aims to capture the current efforts and techniques in testing LLMs, specifically focusing on stress testing, mutation testing, regression testing, metamorphic testing, and adversarial testing. This survey focuses on how traditional testing methods must be adapted to fit the needs of LLMs. Furthermore, while this area is fairly novel, there are still gaps in the literature that have been identified for future research. Timothy Elvira, Tyler Procko, Lynn Vonder Haar, Omar Ochoa |
ICMLA | 3 |
| 2023 | An analysis of explainability methods for convolutional neural networks
Lynn Vonder Haar, Timothy Elvira, Omar Ochoa |
Eng. Appl. Artif. Intell. | 1 |