VLDB 2026 Research / reviewers in the wild / expert
Nan Yang 0009
dblp:51/1629-9
· DBLP profile ↗
7ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0002-3071-0244ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AskGraph: A Dependency-Aware Code Assistant Powered by Code Graphs and LLM-Generated Cypher QueriesabstractLarge Language Models (LLMs) have transformed code assistants by enabling personalization, interactivity, and higher abstraction. However, these assistants often struggle with a common limitation; they generate responses based on a limited set of relevant code snippets retrieved from the codebase using semantic similarity search. This mechanism prevents them from viewing the code structure holistically, making it difficult to give accurate and complete answers to questions on code dependencies and structure. This paper introduces a dependency-aware code assistant that answers structural questions developers cannot easily pose to general-purpose assistants like GitHub Copilot. We achieve this by enriching the LLM with dependency facts obtained from a code graph generated by a static-analysis pipeline customized specifically for industry-scale codebases. The dependency information is queried from a Neo4j database, which stores the code graph, via Text-to-Cypher translation powered by LLMs. Cypher is a query language, designed specifically for querying graph-structured data. We evaluated our solution at Philips Healthcare. Specifically, we performed a benchmark with 420 collected questions and a user study with seven industrial software engineers. By analyzing the results, we identified common mistakes made by GPT40 in the Text-to-Cypher translation to query code graphs, including syntax, schema and semantic errors. This work lays the foundation for advancing Cypher query generation on industryscale code graphs and for augmenting graph-based code analysis with LLMs. Nan Yang 0009, Joseph Reynolds, Laurens Prast, Rosilde Corvino |
ICSME | 1 |
| 2025 | Towards Synthesis-Based Engineering for Cyber-Physical Production SystemsabstractContains fulltext : 318879.pdf (Publisher’s version ) (Open Access) Wytse Oortwijn, Yuri Blankenstein, Jos Hegge, Dennis Hendriks, Piërre van de Laar, Bram van der Sanden, Laura van Veen, Nan Yang 0009 |
MODELSWARD | 8 |
| 2023 | An interview study about the use of logs in embedded software engineering
Nan Yang 0009, Pieter J. L. Cuijpers, Dennis Hendriks, Ramon R. H. Schiffelers, Johan J. Lukkien, Alexander Serebrenik |
Empir. Softw. Eng. | 1 |
| 2021 | Logs and models in engineering complex embedded systemsabstractComplex embedded systems, such as robotics, automotive and high-tech manufacturing, are hard to maintain due to their complex nature. To advance our understanding of the software engineering practice for complex embedded systems, we conducted a series of empirical studies at ASML, a leading manufacturer of lithography machines for semi-conductor industry. We started with an interview study exploring how developers use execution logs, essential artifacts that capture the runtime behavior of software systems. The empirical insights obtained from this study led us to explore subtopics about model inference from logs, modeling practice and log comparison. Motivated by the observation that developers often manually sketch behavioral models based on logs, we propose a model inference technique that can extract models by combining log analysis, and analysis of a running system under stimuli. As observed in this model inference study, the transition from code to models requires developers to work with a hybrid system which consists of handwritten code and models. We then study modeling practices and the roles of model in such hybrid systems. Particularly, we study why developers violate modeling guidelines, providing implications for researchers and tool builders to support developers in modeling complex embedded systems. Another interesting observation from the interview study is that developers face challenges in comparing multiple logs generated from such systems. We therefore conduct a literature study to provide an overview of the existing techniques and identify the limitations of the existing techniques. In this project, we study logs and models in complex embedded systems, providing tool builders, researchers and practitioners with implications to facilitate log analysis, model inference, modeling practice and log comparison. Nan Yang 0009, Pieter J. L. Cuijpers, Ramon R. H. Schiffelers, Johan J. Lukkien, Alexander Serebrenik |
ICSME | 1 |
| 2021 | Single-state state machines in model-driven software engineering: an exploratory studyabstractAbstract Context Models, as the main artifact in model-driven engineering, have been extensively used in the area of embedded systems for code generation and verification. One of the most popular behavioral modeling techniques is the state machine. Many state machine modeling guidelines recommend that a state machine should have more than one state in order to be meaningful. However, single-state state machines (SSSMs) violating this recommendation have been used in modeling cases reported in the literature. Objective We aim for understanding the phenomenon of using SSSMs in practice as understanding why developers violate the modeling guidelines is the first step towards improvement of modeling tools and practice. Method To study the phenomenon, we conducted an exploratory study which consists of two complementary studies. The first study investigated the prevalence and role of SSSMs in the domain of embedded systems, as well as the reasons why developers use them and their perceived advantages and disadvantages. We employed the sequential explanatory strategy, including repository mining and interview, to study 1500 state machines from 26 components at ASML, a leading company in manufacturing lithography machines from the semiconductor industry. In the second study, we investigated the evolutionary aspects of SSSMs, exploring when SSSMs are introduced to the systems and how developers modify them by mining the largest state-machine-based component from the company. Results We observe that 25 out of 26 components contain SSSMs. Our interviews suggest that SSSMs are used to interface with the existing code, to deal with tool limitations, to facilitate maintenance and to ease verification. Our study on the evolutionary aspects of SSSMs reveals that the need for SSSMs to deal with tool limitations grew continuously over the years. Moreover, only a minority of SSSMs have been changed between SSSM and multiple-state state machine (MSSM) during their evolution. The most frequent modifications developers made to SSSMs is inserting events with constraints on the execution of the events. Conclusions Based on our results, we provide implications for developers and tool builders. Furthermore, we formulate hypotheses about the effectiveness of SSSMs, the impacts of SSSMs on development, maintenance and verification as well as the evolution of SSSMs. Nan Yang 0009, Pieter J. L. Cuijpers, Ramon R. H. Schiffelers, Johan J. Lukkien, Alexander Serebrenik |
Empir. Softw. Eng. | 1 |
| 2020 | Painting Flowers: Reasons for Using Single-State State Machines in Model-Driven EngineeringabstractModels, as the main artifact in model-driven engineering, have been extensively used in the area of embedded systems for code generation and verification. One of the most popular behavioral modeling techniques is state machine. Many state machine modeling guidelines recommend that a state machine should have more than one state in order to be meaningful. However, single-state state machines (SSSMs) violating this recommendation have been used in modeling cases reported in the literature. Nan Yang 0009, Pieter J. L. Cuijpers, Ramon R. H. Schiffelers, Johan J. Lukkien, Alexander Serebrenik |
MSR | 1 |
| 2019 | Improving Model Inference in Industry by Combining Active and Passive LearningabstractInferring behavioral models (e.g., state machines) of software systems is an important element of re-engineering activities. Model inference techniques can be categorized as active or passive learning, constructing models by (dynamically) interacting with systems or (statically) analyzing traces, respectively. Application of those techniques in the industry is, however, hindered by the trade-off between learning time and completeness achieved (active learning) or by incomplete input logs (passive learning). We investigate the learning time/completeness achieved trade-off of active learning with a pilot study at ASML, provider of lithography systems for the semiconductor industry. To resolve the trade-off we advocate extending active learning with execution logs and passive learning results. We apply the extended approach to eighteen components used in ASML TWINSCAN lithography machines. Compared to traditional active learning, our approach significantly reduces the active learning time. Moreover, it is capable of learning the behavior missed by the traditional active learning approach. Nan Yang 0009, Kousar Aslam, Ramon R. H. Schiffelers, Leonard Lensink, Dennis Hendriks, Loek Cleophas, Alexander Serebrenik |
SANER | 1 |