VLDB 2026 Research / reviewers in the wild / expert
Kevin Wu
dblp:31/10853
· DBLP profile ↗
12ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Identifying Students' Code Quality Defects while Contributing to Large Code BasesabstractLow-quality code can cost a company significant time and effort. As a result, code quality has been consistently studied in computing education research, especially in the context of CS1 students. However, less research has examined students' code quality while working on existing code bases (i.e., in tasks they are expected to do in industry). In this paper, we identify 1) common code quality defects introduced by upper-division students while contributing to an existing code base, 2) the severity, tool support, and language independence of those defects, and 3) programming experiences that may be associated with students' frequencies of defects, such as internship experience and use of Python (which was the language used in the programming tasks). In an upper division software engineering course, 48 students worked individually to 1) modify an existing feature and 2) implement a new feature in an open-source code base. Using an existing framework of code quality defects by Řechtáčková et al., we conducted a manual code review of all student submissions and found that students created defects related to Poor Design, Poor Documentation, Poor Formatting, and Unused Code at a high frequency. Students also seemed to copy-paste code from other files, which introduced defects related to Unused Code and Poor Design to their submission. Though our regression analysis did not reveal statistically significant predictors, students with prior internships, on average, introduced more code quality defects than those without any internship experience. Anshul Shah 0002, Thomas Rexin, Gonzalo Allen-Perez, Kevin Wu, William G. Griswold, Adalbert Gerald Soosai Raj |
ITiCSE (1) | 4 |
| 2025 | AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack IntegrationabstractAs large language models (LLMs) become increasingly capable, security and safety evaluation are crucial. While current red teaming approaches have made strides in assessing LLM vulnerabilities, they often rely heavily on human input and lack comprehensive coverage of emerging attack vectors. This paper introduces AutoRedTeamer, a novel framework for fully automated, end-to-end red teaming against LLMs. AutoRedTeamer combines a multi-agent architecture with a memory-guided attack selection mechanism to enable continuous discovery and integration of new attack vectors. The dual-agent framework consists of a red teaming agent that can operate from high-level risk categories alone to generate and execute test cases, and a strategy proposer agent that autonomously discovers and implements new attacks by analyzing recent research. This modular design allows AutoRedTeamer to adapt to emerging threats while maintaining strong performance on existing attack vectors. We demonstrate AutoRedTeamer’s effectiveness across diverse evaluation settings, achieving 20% higher attack success rates on HarmBench against Llama-3.1-70B while reducing computational costs by 46% compared to existing approaches. AutoRedTeamer also matches the diversity of human-curated benchmarks in generating test cases, providing a comprehensive, scalable, and continuously evolving framework for evaluating the security of AI systems. Andy Zhou, Kevin Wu, Francesco Pinto, Zhaorun Chen, Yi Zeng 0005, Yu Yang 0007, Oluwasanmi Koyejo, James Zou 0001, Bo Li 0026 |
NeurIPS | 2 |
| 2025 | An Analysis of Students' Testing Processes in CS1abstractUnderstanding students' testing processes in a CS1 course is crucial in helping instructors of introductory courses determine the necessary content to teach. Prior work highlights the importance of teaching testing practices to students, as there is concern for students' testing abilities upon graduation of an university CS program. Given that testing is an implicit programming process, we aim to examine how students in CS1 go about testing their code in programming assignments. Because of the consistent research showing the achievement gap between students with and without prior experience in introductory classes, our analysis also aims to understand specific differences in testing processes between the two groups. Leveraging a dataset of over 300 students with over 50,000 snapshots of student code during their development process, we applied metrics related to incremental testing and determined the usage of diagnostic print statements and the usage of designing test cases beyond the given tests (in which we refer to as ' custom test cases '). A large majority of the students used neither diagnostic print statements nor custom test cases in their programming assignments. Additionally, the three testing practices we examined do not seem to significantly contribute to the achievement gap due to prior experience to students' success, suggesting a need for further investigation into which practices do account for that success. Gonzalo Allen-Perez, Luis Millan, Brandon Nghiem, Kevin Wu, Anshul Shah 0002, Adalbert Gerald Soosai Raj |
SIGCSE (1) | 4 |
| 2024 | Towards Fine-Grained Sidewalk Accessibility Assessment with Deep Learning: Initial Benchmarks and an Open DatasetabstractWe examine the feasibility of using deep learning to infer 33 classes of sidewalk accessibility conditions in pre-cropped streetscape images, including bumpy, brick/cobblestone, cracks, height difference (uplifts), narrow, uneven/slanted, pole, and sign. We present two experiments: first, a comparison between two state-of-the-art computer vision models, Meta’s DINOv2 and OpenAI’s CLIP-ViT, on a cleaned dataset of ∼ 24k images; second, an examination of a larger but noisier crowdsourced dataset (∼ 87k images) on the best performing model from Experiment 1. Though preliminary, Experiment 1 shows that certain sidewalk conditions can be identified with high precision and recall, such as missing tactile warnings on curb ramps and grass grown on sidewalks, while Experiment 2 demonstrates that larger but noisier training data can have a detrimental effect on performance. We contribute an open dataset and classification benchmarks to advance this important area. Kevin Wu, Minchu Kulkarni, Michael Saugstad, Peyton Anton Rapo, Jeremy Freiburger, Chu Li 0001, Jon Froehlich |
ASSETS | 2 |
| 2024 | DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion ModelsabstractQuantifying the impact of training data points is crucial for understanding the outputs of machine learning models and for improving the transparency of the AI pipeline. The influence function is a principled and popular data attribution method, but its computational cost often makes it challenging to use. This issue becomes more pronounced in the setting of large language models and text-to-image models. In this work, we propose DataInf, an efficient influence approximation method that is practical for large-scale generative AI models. Leveraging an easy-to-compute closed-form expression, DataInf outperforms existing influence computation algorithms in terms of computational and memory efficiency. Our theoretical analysis shows that DataInf is particularly well-suited for parameter-efficient fine-tuning techniques such as LoRA. Through systematic empirical evaluations, we show that DataInf accurately approximates influence scores and is orders of magnitude faster than existing methods. In applications to RoBERTa-large, Llama-2-13B-chat, and stable-diffusion-v1.5 models, DataInf effectively identifies the most influential fine-tuning examples better than other approximate influence scores. Moreover, it can help to identify which data points are mislabeled. Yongchan Kwon, Eric Wu, Kevin Wu, James Zou 0001 |
ICLR | 3 |
| 2024 | ClashEval: Quantifying the tug-of-war between an LLM's internal prior and external evidenceabstractRetrieval augmented generation (RAG) is frequently used to mitigate hallucinations and provide up-to-date knowledge for large language models (LLMs). However, given that document retrieval is an imprecise task and sometimes results in erroneous or even harmful content being presented in context, this raises the question of how LLMs handle retrieved information: If the provided content is incorrect, does the model know to ignore it, or does it recapitulate the error? Conversely, when the model's initial response is incorrect, does it always know to use the retrieved information to correct itself, or does it insist on its wrong prior response? To answer this, we curate a dataset of over 1200 questions across six domains (e.g., drug dosages, Olympic records, locations) along with content relevant to answering each question. We further apply precise perturbations to the answers in the content that range from subtle to blatant errors.We benchmark six top-performing LLMs, including GPT-4o, on this dataset and find that LLMs are susceptible to adopting incorrect retrieved content, overriding their own correct prior knowledge over 60\% of the time. However, the more unrealistic the retrieved content is (i.e. more deviated from truth), the less likely the model is to adopt it. Also, the less confident a model is in its initial response (via measuring token probabilities), the more likely it is to adopt the information in the retrieved content. We exploit this finding and demonstrate simple methods for improving model accuracy where there is conflicting retrieved content. Our results highlight a difficult task and benchmark for LLMs -- namely, their ability to correctly discern when it is wrong in light of correct retrieved content and to reject cases when the provided content is incorrect. Our dataset, called ClashEval, and evaluations are open-sourced to allow for future benchmarking on top-performing models at https://github.com/kevinwu23/StanfordClashEval. Kevin Wu, Eric Wu, James Zou 0001 |
NeurIPS | 1 |
| 2023 | BusStopCV: A Real-time AI Assistant for Labeling Bus Stop Accessibility Features in Streetscape ImageryabstractPublic transportation provides vital connectivity to people with disabilities, facilitating access to work, education, and health services. While modern navigation applications provide a suite of information about transit options—including real-time updates about bus or train arrivals—they lack data about the accessibility of the transit stops themselves. Bus stop features such as seatings, shelters, and landing areas are critical, but few cities provide this information. In this demo paper, we introduce BusStopCV, a Human+AI web prototype for scalably collecting data on bus stop features using real-time computer vision and human labeling. We describe BusStopCV’s design, custom training with the YOLOv8 model, and an evaluation of 100 randomly selected bus stops in Seattle, WA. Our findings demonstrate the potential of BusStopCV and highlight opportunities for future work. Minchu Kulkarni, Chu Li 0001, Jaye Jungmin Ahn, Katrina Oi Yau Ma, Zhihan Zhang 0002, Michael Saugstad, Kevin Wu, Yochai Eisenberg, Valerie Novack, Brent C. Chamberlain, Jon Froehlich |
ASSETS | 7 |
| 2020 | Bilingual Multi-word Expressions, Multiple-correspondence, and their cultivation from parallel patents: The Chinese-English case
Benjamin Ka-Yin T'sou, Ka-Po Chow, John Lee 0001, Ka-Fai Yip, Yaxuan Ji, Kevin Wu |
PACLIC | 6 |
| 2020 | Detecting hidden webcams with delay-tolerant similarity of simultaneous observation
Kevin Wu, Brent Lagesse |
Pervasive Mob. Comput. | 1 |
| 2019 | Do You See What I See?Detecting Hidden Streaming Cameras Through Similarity of Simultaneous ObservationabstractSmall, low-cost, wireless cameras are becoming increasingly commonplace making surreptitious observation of people more difficult to detect. Previous work in detecting hidden cameras has only addressed limited environments in small spaces where the user has significant control of the environment. To address this problem in a less constrained scope of environments, we introduce the concept of similarity of simultaneous observation where the user utilizes a camera (Wi-Fi camera, camera on a mobile phone or laptop) to compare timing patterns of data transmitted by potentially hidden cameras and the timing patterns that are expected from the scene that the known camera is recording. To analyze the patterns, we applied several similarity measures and demonstrated an accuracy of over 87% and and F1 score of 0.88 using an efficient threshold-based classification. Furthermore, we used our data set to train a neural network and saw improved results with accuracy as high as 97% and an F1 score over 0.95 for both indoors and outdoors settings. From these results, we conclude that similarity of simultaneous observation is a feasible method for detecting hidden wireless cameras that are streaming video of a user. Our work removes significant limitations that have been put on previous detection methods. Kevin Wu, Brent Lagesse |
PerCom | 1 |
| 2015 | Helium: lifting high-performance stencil kernels from stripped x86 binaries to halide DSL codeabstractHighly optimized programs are prone to bit rot, where performance quickly becomes suboptimal in the face of new hardware and compiler techniques. In this paper we show how to automatically lift performance-critical stencil kernels from a stripped x86 binary and generate the corresponding code in the high-level domain-specific language Halide. Using Halide’s state-of-the-art optimizations targeting current hardware, we show that new optimized versions of these kernels can replace the originals to rejuvenate the application for newer hardware. The original optimized code for kernels in stripped binaries is nearly impossible to analyze statically. Instead, we rely on dynamic traces to regenerate the kernels. We perform buffer structure reconstruction to identify input, intermediate and output buffer shapes. We abstract from a forest of concrete dependency trees which contain absolute memory addresses to symbolic trees suitable for high-level code generation. This is done by canonicalizing trees, clustering them based on structure, inferring higher-dimensional buffer accesses and finally by solving a set of linear equations based on buffer accesses to lift them up to simple, high-level expressions. Helium can handle highly optimized, complex stencil kernels with input-dependent conditionals. We lift seven kernels from Adobe Photoshop giving a 75% performance improvement, four kernels from IrfanView, leading to 4.97× performance, and one stencil from the miniGMG multigrid benchmark netting a 4.25× improvement in performance. We manually rejuvenated Photoshop by replacing eleven of Photoshop’s filters with our lifted implementations, giving 1.12× speedup without affecting the user experience. Charith Mendis, Jeffrey Bosboom, Kevin Wu, Shoaib Kamil 0001, Jonathan Ragan-Kelley, Sylvain Paris, Saman P. Amarasinghe |
PLDI | 3 |
| 2009 | Using transition test to understand timing behavior of logic circuits on UltraSPARCTM T2 familyabstractDelay test is crucial for finding slow paths and slow ICs, both during bringup and during speed binning. Path delay test has traditionally been considered to be superior in finding slow paths. This paper describes our experiments indicating that this is not always the case. For the UltraSPARC T2 microprocessor series we found that transition delay test often ran slower, was more effective in finding the root cause of the slow path, and correlated well with functional diags also used for speed binning. Transition test does a better job finding delay issues related to the impact of simultaneous switching and coupling noise on chip speed. We used transition test to measure the impact on chip timing of voltage, temperature, and we also used it to confirm the results of improving slow paths. Liang-Chi Chen, Paul Dickinson, Peter Dahlgren, Scott Davidson 0001, Olivier Caty, Kevin Wu |
ITC | 6 |