VLDB 2026 Research / reviewers in the wild / expert
Boyue Caroline Hu
dblp:245/0715
· DBLP profile ↗
7ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0002-7090-6276ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Debugging and Runtime Analysis of Neural Networks with VLMs (A Case Study)abstractDebugging of Deep Neural Networks (DNNs), particularly vision models, is very challenging due to the complex and opaque decision-making processes in these networks. In this paper, we explore multi-modal Vision-Language Models (VLMs), such as CLIP, to automatically interpret the opaque representation space of vision models using natural language. This in turn, enables a semantic analysis of model behavior using human-understandable concepts, without requiring costly human annotations. Key to our approach is the notion of semantic heatmap, that succinctly captures the statistical properties of DNNs in terms of the concepts discovered with the VLM and that are computed off-line using a held-out data set. We show the utility of semantic heatmaps for fault localization - an essential step in debugging - in vision models. Our proposed technique helps localize the fault in the network (encoder vs head) and also highlights the responsible high-level concepts, by leveraging novel differential heatmaps, which summarize the semantic differences between the correct and incorrect behaviour of the analyzed DNN. We further propose a lightweight runtime analysis to detect and filter-out defects at runtime, thus improving the reliability of the analyzed DNNs. The runtime analysis works by measuring and comparing the similarity between the heatmap computed for a new (unseen) input and the heatmaps computed a-priori for correct vs incorrect DNN behavior. We consider two types of defects: misclassifications and vulnerabilities to adversarial attacks. We demonstrate the debugging and runtime analysis on a case study involving a complex ResNet-based classifier trained on the RIVAL10 dataset. Boyue Caroline Hu, Divya Gopinath, Corina Pasareanu, Nina Narodytska, Ravi Mangal, Susmit Jha |
CAIN | 1 |
| 2025 | Assessing Visually-Continuous Corruption Robustness of Neural Networks Relative to Human PerformanceabstractNeural Networks (NNs) have surpassed human accuracy in image classification on ImageNet, yet they often lack robustness against image corruption, i.e., corruption robustness, with such robustness being seemingly effortless for human perception. In this paper, we propose visually-continuous corruption robustness (VCR) - an extension of corruption robustness to allow assessing it over the wide and continuous range of changes that correspond to the human perceptive quality (i.e., from the original image to the full distortion of all perceived visual information), along with two novel human-aware metrics for NN evaluation. To compare VCR of NNs with human perception, we conducted extensive experiments on 14 commonly used image corruptions with 7,718 human participants and state-of-the-art robust NN models with different training objectives (e.g., standard, adversarial, corruption robustness), different architectures (e.g., convolution NNs, vision transformers), and different amounts of training data augmentation. Our study showed that: 1) assessing robustness against continuous corruption can reveal insufficient robustness undetected by existing benchmarks; as a result, 2) the gap between NN and human robustness is larger than previously known; and finally, 3) some image corruptions have a similar impact on human perception, offering opportunities for more cost-effective robustness assessments. Huakun Shen, Boyue Caroline Hu, Krzysztof Czarnecki 0001, Lina Marsso, Marsha Chechik |
WACV | 2 |
| 2023 | Towards Feature-Based Analysis of the Machine Learning Development LifecycleabstractThe safety and trustworthiness of systems with components that are based on Machine Learning (ML) require an in-depth understanding and analysis of all stages in its Development Lifecycle (MLDL). High-level abstractions of desired functionalities, model behaviour, and data are called features, and they have been studied by different communities across all MLDL stages. In this paper, we propose to support Software Engineering analysis of the MLDL through features, calling it feature-based analysis of the MLDL. First, to achieve a shared understanding of features among different experts, we establish a taxonomy of existing feature definitions currently used in various MLDL stages. Through this taxonomy, we map features from different stages to each other, discover gaps and future research directions and identify areas of collaboration between Software Engineering and other MLDL experts. Boyue Caroline Hu, Marsha Chechik |
ESEC/SIGSOFT FSE | 1 |
| 2023 | DecompoVision: Reliability Analysis of Machine Vision Components through Decomposition and ReuseabstractAnalyzing reliability of Machine Vision Components (MVC) against scene changes (such as rain or fog) in their operational environment is crucial for safety-critical applications. Safety analysis relies on the availability of precisely specified and, ideally, machine-verifiable requirements. The state-of-the-art reliability framework ICRAF developed machine-verifiable requirements obtained using human performance data. However, ICRAF is limited to analyzing reliability of MVCs solving simple vision tasks, such as image classification. Yet, many real-world safety-critical systems require solving more complex vision tasks, such as object detection and instance segmentation. Fortunately, many complex vision tasks (which we call “c-tasks”) can be represented as a sequence of simple vision subtasks. For instance, object detection can be decomposed as object localization followed by classification. Based on this fact, in this paper, we show that the analysis of c-tasks can also be decomposed as a sequential analysis of their simple subtasks, which allows us to apply existing techniques for analyzing simple vision tasks. Specifically, we propose a modular reliability framework, DecompoVision, that decomposes: (1) the problem of solving a c-task, (2) the reliability requirements, and (3) the reliability analysis, and, as a result, provides deeper insights into MVC reliability. DecompoVision extends ICRAF to handle complex vision tasks and enables reuse of existing artifacts across different c-tasks. We capture new reliability gaps by checking our requirements on 13 widely used object detection MVCs, and, for the first time, benchmark segmentation MVCs. Boyue Caroline Hu, Lina Marsso, Nikita Dvornik, Huakun Shen, Marsha Chechik |
ESEC/SIGSOFT FSE | 1 |
| 2022 | If a Human Can See It, So Should Your System: Reliability Requirements for Machine Vision ComponentsabstractMachine Vision Components (MVC) are becoming safety-critical. Assuring their quality, including safety, is essential for their successful deployment. Assurance relies on the availability of precisely specified and, ideally, machine-verifiable requirements. MVCs with state-of-the-art performance rely on machine learning (ML) and training data, but largely lack such requirements. Boyue Caroline Hu, Lina Marsso, Krzysztof Czarnecki 0001, Rick Salay, Huakun Shen, Marsha Chechik |
ICSE | 1 |
| 2022 | What to Check: Systematic Selection of Transformations for Analyzing Reliability of Machine Vision ComponentsabstractMachine Vision Components (MVCs) are deployed in safety-critical systems, such as autonomous driving, and their reliability must be checked against scene changes, e.g., rain, that may lead to hazardous situations in the deployment environment. Many scene changes leading to hazardous situations may be hard to reproduce on demand, so existing approaches for MVC reliability analysis use synthetic image transformations to simulate such changes. Therefore, the question of how to select the image transformations to simulate specific hazardous situations is essential to MVC reliability analysis. Yet, this problem has not been addressed by the scientific community so far. In this paper, we propose a framework for mapping between hazardous situations and relevant image transformations using their descriptions. Our framework includes a systematic description mapping process DMaP, a method autoDMaP for automating this process, and coverage metrics measuring how well a list of transformations can simulate a list of hazardous situations. We show the applicability of our framework by mapping hazardous situations from an existing checklist, i.e., CV-HAZOP, to a list of synthetic image transformations from a state-of-the-art transformation library, i.e., Albumentation. As part of evaluation, we conducted an experiment and showed that, compared with the manual, ad-hoc mapping produced by image processing experts, DMaP and autoDMaP resulted in better precision and recall. Additionally, using our new coverage metrics, we found that image transformations considered by state-of-the-art libraries and reliability benchmarks are far from fully simulating the CV-HAZOP hazardous situations, and the MVCs that perform best on these benchmarks have significant reliability gaps against these situations. Boyue Caroline Hu, Lina Marsso, Krzysztof Czarnecki 0001, Marsha Chechik |
ISSRE | 1 |
| 2019 | Support for user generated evolutions of goal modelsabstractGoal models are used in early phase requirements engineering to elicit stakeholders' intentions, analyze dependencies, and help stakeholders make trade-off decisions about the project and its interaction with the environment. The Evolving Intentions framework extended goal model analysis to evaluate how models change over time, by creating simulation paths showing possible evolutions of the model. More recently, we extended this analysis to allow users to explore states along the path and generate their own simulation paths. However, this approach is limited by users' ability to comprehend the state space, which grows exponentially with the size of the model. In this paper, we explore using filters to reduce the number of viewable solutions enabling users to create their own simulation results. We present our approach and initial validation, including an analysis of prior models and a review of expert feedback. Boyue Caroline Hu, Alicia M. Grubb |
MiSE@ICSE | 1 |