EDBT 2026 Demo / reviewers in the wild / expert
Boqi Chen
dblp:172/9949
· DBLP profile ↗
20ranked-venue papers
11as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Grounding Generative AI in Software Engineering: Are We There Yet?
Mootez Saad, José Antonio Hernández López, Boqi Chen, Neil A. Ernst, Dániel Varró, Tushar Sharma 0001 |
SANER | 3 |
| 2026 | Certifying robustness of graph convolutional networks for node perturbation with polyhedra abstract interpretation
Boqi Chen, Kristóf Marussy, Oszkár Semeráth, Gunter Mussbacher, Dániel Varró |
Data Min. Knowl. Discov. | 1 |
| 2025 | The Power of Types: Exploring the Impact of Type Checking on Neural Bug Detection in Dynamically Typed Languagesabstract[Motivation] Automated bug detection in dynamically typed languages such as Python is essential for maintaining code quality. The lack of mandatory type annotations in such languages can lead to errors that are challenging to identify early with traditional static analysis tools. Recent progress in deep neural networks has led to increased use of neural bug detectors. In statically typed languages, a type checker is integrated into the compiler and thus taken into consideration when the neural bug detector is designed for these languages. [Problem] However, prior studies overlook this aspect during the training and testing of neural bug detectors for dynamically typed languages. When an optional type checker is used, assessing existing neural bug detectors on bugs easily detectable by type checkers may impact their performance estimation. Moreover, including these bugs in the training set of neural bug detectors can shift their detection focus toward the wrong type of bugs. [Contribution] We explore the impact of type checking on various neural bug detectors for variable misuse bugs, a common type targeted by neural bug detectors. Existing synthetic and real-world datasets are type-checked to evaluate the prevalence of type-related bugs. Then, we investigate how type-related bugs influence the training and testing of the neural bug detectors. [Findings] Our findings indicate that existing bug detection datasets contain a significant proportion of type-related bugs. Building on this insight, we discover integrating the neural bug detector with a type checker can be beneficial, especially when the code is annotated with types. Further investigation reveals neural bug detectors perform better on type-related bugs than other bugs. Moreover, removing type-related bugs from the training data helps improve neural bug detectors' ability to identify bugs beyond the scope of type checkers. Boqi Chen, José Antonio Hernández López, Gunter Mussbacher, Dániel Varró |
ICSE | 1 |
| 2025 | MIXPINN: Mixed-Material Simulations by Physics-Informed Neural NetworkabstractSimulating the complex interactions between soft tissues and rigid anatomy is critical for applications in surgical training, planning, and robotic-assisted interventions. Traditional Finite Element Method (FEM)-based simulations, while accurate, are computationally expensive and impractical for real-time scenarios. Learning-based approaches have shown promise in accelerating predictions but have fallen short in modeling soft-rigid interactions effectively. We introduce MIXPINN, a physics-informed Graph Neural Network (GNN) framework for mixed-material simulations, explicitly capturing soft-rigid interactions using graph-based augmentations. Our approach integrates Virtual Nodes (VNs) and Virtual Edges (VEs) to enhance rigid body constraint satisfaction while preserving computational efficiency. By leveraging a graph-based representation of biomechanical structures, MIXPINN learns high-fidelity deformations from FEM-generated data and achieves real-time inference with sub-millimeter accuracy. We validate our method in a realistic clinical scenario, demonstrating superior performance compared to baseline GNN models and traditional FEM methods. Our results show that MIXPINN reduces computational cost by an order of magnitude while maintaining high physical accuracy, making it a viable solution for real-time surgical simulation and robotic-assisted procedures. Xintian Yuan, Yunke Ao, Boqi Chen, Philipp Fürnstahl |
IROS | 3 |
| 2025 | Revisiting Automatic Data Curation for Vision Foundation Models in Digital Pathology
Boqi Chen, Cédric Vincent-Cuaz, Lydia A. Schoenpflug, Manuel Madeira, Lisa Fournier, Vaishnavi Subramanian, Sonali Andani, Samuel Ruipérez-Campillo, Julia E. Vogt, Raphaëlle Luisier, Dorina Thanou, Viktor H. Koelzer, Pascal Frossard, Gabriele Campanella, Gunnar Rätsch |
MICCAI (6) | 1 |
| 2025 | MCeT: Behavioral Model Correctness Evaluation using Large Language ModelsabstractBehavioral model diagrams, e.g., sequence diagrams, are an essential form of documentation that are typically designed by system engineers from requirements documentation, either fully manually or assisted by design tools. With the growing use of Large Language Models (LLM) as AI modeling assistants, more automation will be involved in generating diagrams. This necessitates the advancement of automatic model correctness evaluation tools. Such a tool can be used to evaluate both manually and AI automatically generated models; to provide feedback to system engineers, and enable AI assistants to self-evaluate and self-enhance their generated models. In this paper, we propose MCeT, the first fully automated tool to evaluate the correctness of a behavioral model, sequence diagrams in particular, against its corresponding requirements text and produce a list of issues that the model has. We utilize LLMs for the correctness evaluation tasks as they have shown outstanding natural language understanding ability. However, we show that directly asking an LLM to compare a diagram to requirements finds less than 35% of issues that experienced engineers can find. We propose to supplement the direct check with a fine-grained, multi-perspective approach; we split the diagram into atomic, non-divisible interactions, and split the requirements text into atomic, self-contained items. We compare the diagram with atomic requirements and each diagramatom with the requirements. We also propose a self-consistency checking approach that combines perspectives to mitigate LLM hallucinated issues. Our combined approach improves upon the precision of the direct approach from 0.58 to 0.81 in a dataset of real requirements. Moreover, the approach finds 90% more issues that the experienced engineers found than the direct approach, and reports an average of 6 new issues per diagram. Khaled E. Ahmed, Jialing Song, Boqi Chen, Ou Wei, Bingzhou Zheng |
MODELS | 3 |
| 2025 | SHERPA: A Model-Driven Framework for Large Language Model ExecutionabstractRecently, large language models (LLMs) have achieved widespread application across various fields. Despite their impressive capabilities, LLMs suffer from a lack of structured reasoning ability, particularly for complex tasks requiring domain-specific best practices, which are often unavailable in the training data. Although multi-step prompting methods incorporating human best practices, such as chain-of-thought and tree-of-thought, have gained popularity, they lack a general mechanism to control LLM behavior. In this paper, we propose SHERPA, a model-driven framework to improve the LLM performance on complex tasks by explicitly incorporating domain-specific best practices into hierarchical state machines. By structuring the LLM execution processes using state machines, SHERPA enables more fine-grained control over their behavior via rules or decisions driven by machine learning-based approaches, including LLMs. We show that SHERPA is applicable to a wide variety of tasks-specifically, code generation, class name generation, and question answering-replicating previously proposed approaches while further improving the performance. We demonstrate the effectiveness of SHERPA for the aforementioned tasks using various LLMs. Our systematic evaluation compares different state machine configurations against baseline approaches without state machines. Results show that integrating well-designed state machines significantly improves the quality of LLM outputs, and is particularly beneficial for complex tasks with well-established human best practices but lacking data used for training LLMs. Boqi Chen, Kua Chene, José Antonio Hernández López, Gunter Mussbacher, Dániel Varró, Amir Feizpour |
MODELS | 1 |
| 2025 | Accurate and Consistent Graph Model Generation from Text with Large Language ModelsabstractGraph model generation from natural language description is an important task with many applications in software engineering. With the rise of large language models (LLMs), there is a growing interest in using LLMs for graph model generation. Nevertheless, LLM-based graph model generation typically produces partially correct models that suffer from three main issues: (1) syntax violations: the generated model may not adhere to the syntax defined by its metamodel, (2) constraint inconsistencies: the structure of the model might not conform to some domain-specific constraints, and (3) inaccuracy: due to the inherent uncertainty in LLMs, the models can include inaccurate, hallucinated elements. While the first issue is often addressed through techniques such as constraint decoding or filtering, the latter two remain largely unaddressed. Motivated by recent self-consistency approaches in LLMs, we propose a novel abstraction-concretization framework that enhances the consistency and quality of generated graph models by considering multiple outputs from an LLM. Our approach first constructs a probabilistic partial model that aggregates all candidate outputs and then refines this partial model into the most appropriate concrete model that satisfies all constraints. We evaluate our framework on several popular open-source and closed-source LLMs using diverse datasets for model generation tasks. The results demonstrate that our approach significantly improves both the consistency and quality of the generated graph models. Boqi Chen, Ou Wei, Bingzhou Zheng, Gunter Mussbacher |
MODELS | 1 |
| 2025 | LLM-based Satisfiability Checking of String Requirements by Consistent Data and Checker GenerationabstractRequirements over strings, commonly represented using natural language (NL), are particularly relevant for software systems due to their heavy reliance on string data manipulation. While individual requirements can usually be analyzed manually, verifying properties (e.g., satisfiability) over sets of NL requirements is particularly challenging. Formal approaches (e.g., SMT solvers) may efficiently verify such properties, but are known to have theoretical limitations. Additionally, the translation of NL requirements into formal constraints typically requires significant manual effort. Recently, large language models (LLMs) have emerged as an alternative approach for formal reasoning tasks, but their effectiveness in verifying requirements over strings is less studied. In this paper, we introduce a hybrid approach that verifies the satisfiability of NL requirements over strings by using LLMs (1) to derive a satisfiability outcome (and a consistent string, if possible), and (2) to generate declarative (i.e., SMT) and imperative (i.e., Python) checkers, used to validate the correctness of (1). In our experiments, we assess the performance of four LLMs. Results show that LLMs effectively translate natural language into checkers, even achieving perfect testing accuracy for Python-based checkers. These checkers substantially help LLMs in generating a consistent string and accurately identifying unsatisfiable requirements, leading to more than doubled generation success rate and F1-score in certain cases compared to baselines without generated checkers. Boqi Chen, Aren A. Babikian, Shuzhao Feng, Dániel Varró, Gunter Mussbacher |
RE | 1 |
| 2025 | Generalizable Single-Source Cross-Modality Medical Image Segmentation via Invariant Causal MechanismsabstractSingle-source domain generalization (SDG) aims to learn a model from a single source domain that can generalize well on unseen target domains. This is an important task in computer vision, particularly relevant to medical imaging where domain shifts are common. In this work, we consider a challenging yet practical setting: SDG for cross-modality medical image segmentation. We combine causality-inspired theoretical insights on learning domain-invariant representations with recent advancements in diffusion-based augmentation to improve generalization across diverse imaging modalities. Guided by the “intervention-augmentation equivariant” principle, we use controlled diffusion models (DMs) to simulate diverse imaging styles while preserving the content, leveraging rich generative priors in large-scale pretrained DMs to comprehensively perturb the multidimensional style variable. Extensive experiments on challenging cross-modality segmentation tasks demonstrate that our approach consistently outperforms state-of-the-art SDG methods across three distinct anatomies and imaging modalities. The source code is available at https://github.com/ratschlab/ICMSeg. Boqi Chen, Yuanzhi Zhu 0001, Yunke Ao, Sebastiano Caprara, Reto Sutter, Gunnar Rätsch, Ender Konukoglu, Anna Susmelj |
WACV | 1 |
| 2025 | On Inter-Dataset Code Duplication and Data Leakage in Large Language ModelsabstractMotivation.Large language models (LLMs) have exhibited remarkable proficiency in diverse software engineering (SE) tasks, such as code summarization, code translation, and code search. Handling such tasks typically involves acquiring foundational coding knowledge on large, general-purpose datasets during a pre-training phase, and subsequently refining on smaller, task-specific datasets as part of a fine-tuning phase.Problem statement.Data leakagei.e.,using information of the test set to perform the model training, is a well-known issue in training of machine learning models. A manifestation of this issue is the intersection of the training and testing splits. Whileintra-datasetcode duplication examines this intersection within a given dataset and has been addressed in prior research,inter-dataset code duplication, which gauges the overlap between different datasets, remains largely unexplored. If this phenomenon exists, it could compromise the integrity ofLLMevaluations because of the inclusion of fine-tuning test samples that were already encountered during pre-training, resulting in inflated performance metrics.Contribution.This paper explores the phenomenon of inter-dataset code duplication and its impact on evaluatingLLMs across diverseSEtasks.Study design.We conduct an empirical study using theCodeSearchNetdataset (csn), a widely adopted pre-training dataset, and five fine-tuning datasets used for variousSEtasks. We first identify the intersection between the pre-training and fine-tuning datasets using a deduplication process. Next, we pre-train two versions ofLLMs using a subset ofcsn: one leakyLLM, which includes the identified intersection in its pre-training set, and one non-leakyLLMthat excludes these samples. Finally, we fine-tune both models and compare their performances using fine-tuning test samples that are part of the intersection.Results.Our findings reveal a potential threat to the evaluation ofLLMs across multipleSEtasks, stemming from the inter-dataset code duplication phenomenon. We also demonstrate that this threat is accentuated by the chosen fine-tuning technique. Furthermore, we provide evidence that open-source models such asCodeBERT,GraphCodeBERT, andUnixCodercould be affected by inter-dataset duplication. Based on our findings, we delve into prior research that may be susceptible to this threat. Additionally, we offer guidance toSEresearchers on strategies to prevent inter-dataset code duplication. José Antonio Hernández López, Boqi Chen, Mootez Saad, Tushar Sharma 0001, Dániel Varró |
IEEE Trans. Software Eng. | 2 |
| 2024 | A Unified Model for Longitudinal Multi-Modal Multi-View Prediction with Missingness
Boqi Chen, Junier B. Oliva, Marc Niethammer |
MICCAI (12) | 1 |
| 2023 | MRIS: A Multi-modal Retrieval Approach for Image Synthesis on Diverse Modalities
Boqi Chen, Marc Niethammer |
MICCAI (10) | 1 |
| 2023 | Generative appearance replay for continual unsupervised domain adaptationabstractDeep learning models can achieve high accuracy when trained on large amounts of labeled data. However, real-world scenarios often involve several challenges: Training data may become available in installments, may originate from multiple different domains, and may not contain labels for training. Certain settings, for instance medical applications, often involve further restrictions that prohibit retention of previously seen data due to privacy regulations. In this work, to address such challenges, we study unsupervised segmentation in continual learning scenarios that involve domain shift. To that end, we introduce GarDA (Generative Appearance Replay for continual Domain Adaptation), a generative-replay based approach that can adapt a segmentation model sequentially to new domains with unlabeled data. In contrast to single-step unsupervised domain adaptation (UDA), continual adaptation to a sequence of domains enables leveraging and consolidation of information from multiple domains. Unlike previous approaches in incremental UDA, our method does not require access to previously seen data, making it applicable in many practical scenarios. We evaluate GarDA on three datasets with different organs and modalities, where it substantially outperforms existing techniques. Our code is available at: https://github.com/histocartography/generative-appearance-replay. Boqi Chen, Kevin Thandiackal, Pushpak Pati, Orcun Goksel |
Medical Image Anal. | 1 |
| 2022 | Differentiable Zooming for Multiple Instance Learning on Whole-Slide Images
Kevin Thandiackal, Boqi Chen, Pushpak Pati, Guillaume Jaume, Drew F. K. Williamson, Maria Gabrani, Orcun Goksel |
ECCV (21) | 2 |
| 2022 | Consistent Scene Graph Generation by Constraint OptimizationabstractScene graph generation takes an image and derives a graph representation of key objects in the image and their relations. This core computer vision task is often used in autonomous driving, where traditional software and machine learning (ML) components are used in tandem. However, in such a safety-critical context, valid scene graphs can be further restricted by consistency constraints captured by domain or safety experts. Existing ML approaches for scene graph generation focus exclusively on relation-level accuracy but provide little to no guarantee that consistency constraints are satisfied in the generated scene graphs. In this paper, we aim to complement existing ML-based approaches by a post-processing step using constraint optimization over probabilistic scene graphs that can (1) guarantee that no consistency constraints are violated and (2) improve the overall accuracy of scene graph generation by fixing constraint violations. We evaluate the effectiveness of our approach using well-known, and novel metrics in the context of two popular ML datasets augmented with consistency constraints and two ML-based scene graph generation approaches as baselines. Boqi Chen, Kristóf Marussy, Sebastian Pilarski, Oszkár Semeráth, Dániel Varró |
ASE | 1 |
| 2022 | An Empirical Study of Type-Related Defects in Python ProjectsabstractIn recent years,Pythonhas experienced an explosive growth in adoption, particularly among open source projects. WhilePython's dynamically-typed nature provides developers with powerful programming abstractions, that same dynamic type system allows for type-related defects to accumulate in code bases. To aid in the early detection of type-related defects, type annotations were introduced into thePythonecosystem (i.e., PEP-484) and static type checkers likemypyhave appeared on the market. While applying a type checker likemypycan in theory help to catch type-related defects before they impact users, little is known about the real impact of adopting a type checker to reveal defects inPythonprojects. In this paper, we study the extent to whichPythonprojects benefit from such type checking features. For this purpose, we mine the issue tracking and version control repositories of 210Pythonprojects on GitHub. Inspired by the work of Gaoet al.on type-related defects in JavaScript, we add type annotations to test whethermypydetects an error that would have helped developers to avoid real defects. We observe that 15 percent of the defects could have been prevented bymypy. Moreover, we find that there is no significant difference between the experience level of developers committing type-related defects and the experience of developers committing defects that are not type-related. In addition, a manual analysis of the anti-patterns that most commonly lead to type-checking faults reveals that the redefinition ofPythonreferences, dynamic attribute initialization and incorrectly handled Null objects are the most common causes of type-related faults. Since our study is conducted on fixed public defects that have gone through code reviews and multiple test cycles, these results represent a lower bound on the benefits of adopting a type checker. Therefore, we recommend incorporating a static type checker likemypyinto the development workflow, as not only will it prevent type-related defects but also mitigate certain anti-patterns during development. Faizan Khan, Boqi Chen, Dániel Varró, Shane McIntosh |
IEEE Trans. Software Eng. | 2 |
| 2021 | Automated generation of consistent, diverse and structurally realistic graph modelsabstractAbstract In this paper, we present a novel technique to automatically synthesize consistent, diverse and structurally realistic domain-specific graph models. A graph model is (1) consistent if it is metamodel-compliant and it satisfies the well-formedness constraints of the domain; (2) it is diverse if local neighborhoods of nodes are highly different; and (1) it is structurally realistic if a synthetic graph is at a close distance to a representative real model according to various graph metrics used in network science, databases or software engineering. Our approach grows models by model extension operators using a hill-climbing strategy in a way that (A) ensures that there are no constraint violation in the models (for consistency reasons), while (B) more realistic candidates are selected to minimize a target metric value (wrt. the representative real model). We evaluate the effectiveness of the approach for generating realistic models using multiple metrics for guidance heuristics and compared to other model generators in the context of three case studies with a large set of real human models. We also highlight that our technique is able to generate a diverse set of models, which is a requirement in many testing scenarios. Oszkár Semeráth, Aren A. Babikian, Boqi Chen, Chuning Li, Kristóf Marussy, Gábor Szárnyas, Dániel Varró |
Softw. Syst. Model. | 3 |
| 2019 | Focused Context Balancing for Robust Offline Policy EvaluationabstractPrecisely evaluating the effect of new policies (e.g. ad-placement models, recommendation functions, ranking functions) is one of the most important problems for improving interactive systems. The conventional policy evaluation methods rely on online A/B tests, but they are usually extremely expensive and may have undesirable impacts. Recently, Inverse Propensity Score (IPS) estimators are proposed as alternatives to evaluate the effect of new policy with offline logged data that was collected from a different policy in the past. They tend to remove the distribution shift induced by past policy. However, they ignore the distribution shift that would be induced by the new policy, which results in imprecise evaluation. Moreover, their performances rely on accurate estimation of propensity score, which can not be guaranteed or validated in practice. In this paper, we propose a non-parametric method, named Focused Context Balancing (FCB) algorithm, to learn sample weights for context balancing, so that the distribution shift induced by the past policy and new policy can be eliminated respectively. To validate the effectiveness of our FCB algorithm, we conduct extensive experiments on both synthetic and real world datasets. The experimental results clearly demonstrate that our FCB algorithm outperforms existing estimators by achieving more precise and robust results for offline policy evaluation. Hao Zou 0001, Kun Kuang 0001, Boqi Chen, Peixuan Chen, Peng Cui 0001 |
KDD | 3 |
| 2015 | Reference image based method of region of interest enhancement for haze imageabstractDifferent from general algorithms of haze removal and low lighting image enhancement, which only use the information of image to process, this paper adds a reference image to get more information for the algorithm and focuses on enhancing region of interest of an image based on the reference one. With the reference image, the haze one can be divided into Region of Interest (RoI) and Region of no Interest (non-RoI). Furthermore, the reference image can provide more useful information for computing the transmission map and atmospheric light. For the non-RoI region, a more robust transmission map and minimizing reconstruction error cost function based method to estimate atmospheric light has been proposed. Because the atmospheric light is a global variable, the optimized one is also suitable for the RoI region. With the global optimized atmospheric light, an optimized transmission map can be got for the RoI region. The RoI region can be enhanced via the optimal transmission map and atmosphere light. Theoretical analysis gives eloquent proof proving that the proposed method is definitely better than the traditional dark-channel-prior-based methods due to our better transmission map and atmosphere light. Extensive experiments also show the expected results. Wuzhen Shi, Xinwei Gao, Boqi Chen, Feng Jiang 0001, Debin Zhao |
ICIP | 3 |