VLDB 2026 Research / reviewers in the wild / expert
Zishuo Ding
dblp:276/3298
· DBLP profile ↗
19ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0002-0803-5609ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 17 · 6 first-author · 16 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A multi-language perspective on the robustness of LLM code generation
Zishuo Ding, Jinqiu Yang 0001 |
Empir. Softw. Eng. | 2 |
| 2026 | Enhancing Log Sentiments: An Exploratory Study of Sentiments and Emotions with Software LogsabstractSoftware logs serve as valuable resources for understanding system running and are extensively used in diverse software maintenance tasks. Logs are generated by logging statements in the code, which are written by developers. Therefore, logs may reflect developers’ sentiments about the described situations. Consequently, when developers and system administrators read logs, the sentiments embedded in logs may influence their understanding. Although the sentiments associated with logs can convey valuable information, such information is not leveraged in research and practice. Previous research has primarily relied on verbosity levels of logs to gauge sentiments, which does not really capture the sentiments and emotions perceived by humans. To bridge this gap, in this article, we first conduct an exploratory study to investigate sentiments and emotions that are communicated within logs. Our study encompasses five anomaly log datasets from LogHub and a dataset involving eight open-source Apache Java projects. We find that 8% of the logs express sentiments and emotions though developers are suggested to write them in an objective way. While most log messages might not explicitly express sentiments and emotions, they can still implicitly evoke sentiments and emotions in those who read them. Therefore, we exploit issue reports referencing logs to capture such sentiments and emotions. In these issue reports, 47.5% exhibit emotions, with 54.7% of those emotions being related to logs and 8.1% directly addressing logs. Furthermore, we demonstrate the potential of leveraging sentiment analysis to complement verbosity levels in logs, showcasing how sentiment information can offer novel insights and enhance log analysis. Specifically, by applying automatic tools, we identify 41 issue reports (9.8% on average) with negative sentiment and 55 reports (13.2% on average) with negative emotions, all referencing INFO or DEBUG logs (i.e., low severity). After manually verifying and filtering exception logs, we uncover three main concerns from 22 critical instances. Youshuai Tan, Zishuo Ding, Jinfu Chen 0002, Jifeng Xuan, Weiyi Shang |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | Defects4Log: Benchmarking LLMs for Logging Code Defect Detection and ReasoningabstractLogging code is written by developers to capture system runtime behavior and plays a vital role in debugging, performance analysis, and system monitoring. However, defects in logging code can undermine the usefulness of logs and lead to misinterpretations. Although prior work has identified several logging defect patterns and provided valuable insights into logging practices, these studies often focus on a narrow range of defect patterns derived from limited sources (e.g., commit histories) and lack a systematic and comprehensive analysis. Moreover, large language models (LLMs) have demonstrated promising generalization and reasoning capabilities across a variety of code-related tasks, yet their potential for detecting logging code defects remains largely unexploredIn this paper, we derive a comprehensive taxonomy of logging code defects, which encompasses seven logging code defect patterns with 14 detailed scenarios. We further construct a benchmark dataset, Defects4Log, consisting of 164 developer-verified real-world logging defects. Then we propose an automated framework that leverages various prompting strategies and contextual information to evaluate LLMs’ capability in detecting and reasoning logging code defects. Experimental results reveal that LLMs generally struggle to accurately detect and reason logging code defects based on the source code only. However, incorporating proper knowledge (e.g., detailed scenarios of defect patterns) can lead to 10.9% improvement in detection accuracy. Overall, our findings provide actionable guidance for practitioners to avoid common defect patterns and establish a foundation for improving LLM-based reasoning in logging code defect detection. Zhenhao Li 0002, Zishuo Ding |
ASE | 3 |
| 2025 | Who's to Blame? Rethinking the Brittleness of Automated Web GUI Testing from a Pragmatic PerspectiveabstractAutomated web GUI testing is important for software quality, however, its effectiveness is often undermined by test case brittleness, especially in continuously evolving real-world applications. In this experience paper, we pragmatically investigate the root causes of brittleness. We first analyze why legacy test cases, derived from the Mind2Web dataset, fail when executed on current web application versions. Our findings reveal that brittleness stems from multifaceted factors, including test script design, web application complexity, and automation framework limitations. A longitudinal study further shows that 81.7% of repaired tests break again within six months, primarily due to similar recurring issues, highlighting the persistent nature of brittleness. We further demonstrate that Large Language Models, when provided with human-like diagnostic context, can successfully repair a substantial portion of these brittle tests, though human expertise remains important for more complex scenarios. Our findings emphasize that brittleness is a multifaceted problem requiring collaboration between different parts involved in the automation testing. Haonan Zhang 0006, Kundi Yao, Zishuo Ding, Lizhi Liao, Weiyi Shang |
ASE | 3 |
| 2025 | Improving Qa System Testing Efficiency Through White-Box Test PrioritizationabstractEffective testing of sequence-to-sequence (seq2seq) models, such as those used in question answering (QA) systems, is essential for ensuring their reliability. While recent efforts have introduced metamorphic testing strategies to detect bugs without requiring ground-truth labels, the efficiency of these methods remains limited by their lack of test case prioritization. Executing all test cases uniformly can lead to wasted resources and slower fault discovery. In this paper, we propose a white-box prioritization framework that ranks test cases based on internal signals extracted from the underlying model. Building upon a prior work that introduced two whitebox techniques (i.e., GRI and WALI) for identifying vulnerable tokens, we adapt these techniques to the task of test prioritization. Instead of generating new test inputs, our methods analyze test cases produced by QAQA and prioritize those most likely to uncover faults. We evaluate our approaches on three widely-used QA datasets: BoolQ, NarrativeQA, and SQuAD2. Experimental results show that GRI significantly improves the rate of bug detection under constrained testing budgets, while WALI achieves comparable performance to baseline methods. Our findings demonstrate the value of incorporating white-box insights into the prioritization process, offering a more efficient and effective way to test QA systems. Hanying Shao, Zishuo Ding, Kundi Yao, Haonan Zhang 0006, Weiyi Shang |
QRS | 2 |
| 2025 | Towards effectively testing machine translation systems from white-box perspectives
Hanying Shao, Zishuo Ding, Weiyi Shang, Jinqiu Yang 0001, Nikolaos Tsantalis |
Empir. Softw. Eng. | 2 |
| 2024 | Towards a Robust Waiting Strategy for Web GUI Testing for an Industrial Software SystemabstractAutomated web GUI testing has been widely adopted since manual testing is time-consuming and tedious. Waiting strategy plays a vital role in automated web GUI testing since it significantly impacts the testing performance. Though important, little focus has been set on the waiting strategies in web GUI testing. Existing waiting strategies either wait for a predetermined time, which is not reliable in a dynamic environment, or only wait for a specific condition to be verified, which is often not robust enough to handle the complicated testing scenarios. In this work, we introduce a robust waiting strategy. Instead of waiting for a predetermined time or waiting for the availability of a particular element, our approach waits for a desired state to reach. This is achieved by capturing the Document Object Models (DOM) at the desired point, followed by an offline analysis to identify the differences between the DOMs associated with every two consecutive test actions. Such differences are used to determine the appropriate waiting time when automatically generating tests. Evaluation results with an industrial web application indicate that our approach produces more robust tests than the conventional waiting strategies used in web GUI testing. Furthermore, our generated tests are more representative of the recorded usage scenarios and are efficient with low overhead in test execution time. Haonan Zhang 0006, Lizhi Liao, Zishuo Ding, Weiyi Shang, Nidhi Narula, Catalin Sporea, Andrei Toma, Sarah Sajedi |
ASE | 3 |
| 2024 | GreenStableYolo: Optimizing Inference Time and Image Quality of Text-to-Image Generation
Jingzhi Gong, Giordano d'Aloisio, Zishuo Ding, Yulong Ye, William B. Langdon, Federica Sarro |
SSBSE | 4 |
| 2024 | Temporal knowledge graph reasoning based on evolutional representation and contrastive learning
Qiuying Ma, Xuan Zhang 0002, Zishuo Ding, Chen Gao 0006, Weiyi Shang, Qiong Nong, Yubin Ma, Zhi Jin 0001 |
Appl. Intell. | 3 |
| 2024 | Few-shot relational triple extraction with hierarchical prototype optimization
Chen Gao 0006, Xuan Zhang 0002, Zhi Jin 0001, Weiyi Shang, Yubing Ma, LinYu Li 0001, Zishuo Ding, Yuqin Liang |
Pattern Recognit. | 7 |
| 2024 | LoGenText-Plus: Improving Neural Machine Translation Based Logging Texts Generation with Syntactic TemplatesabstractDevelopers insert logging statements in the source code to collect important runtime information about software systems. The textual descriptions in logging statements (i.e., logging texts) are printed during system executions and exposed to multiple stakeholders including developers, operators, users, and regulatory authorities. Writing proper logging texts is an important but often challenging task for developers. Prior studies find that developers spend significant efforts modifying their logging texts. However, despite extensive research on automated logging suggestions, research on suggesting logging texts rarely exists. To fill this knowledge gap, we first propose LoGenText (initially reported in our conference paper), an automated approach that uses neural machine translation (NMT) models to generate logging texts by translating the related source code into short textual descriptions. LoGenText takes the preceding source code of a logging text as the input and considers other context information, such as the location of the logging statement, to automatically generate the logging text. LoGenText ’s evaluation on 10 open source projects indicates that the approach is promising for automatic logging text generation and significantly outperforms the state-of-the-art approach. Furthermore, we extend LoGenText to LoGenText-Plus by incorporating the syntactic templates of the logging texts. Different from LoGenText , LoGenText-Plus decomposes the logging text generation process into two stages. LoGenText-Plus first adopts an NMT model to generate the syntactic template of the target logging text. Then LoGenText-Plus feeds the source code and the generated template as the input to another NMT model for logging text generation. We also evaluate LoGenText-Plus on the same 10 projects and observe that it outperforms LoGenText on 9 of them. According to a human evaluation from developers’ perspectives, the logging texts generated by LoGenText-Plus have a higher quality than those generated by LoGenText and the prior baseline approach. By manually examining the generated logging texts, we then identify five aspects that can serve as guidance for writing or generating good logging texts. Our work is an important step toward the automated generation of logging statements, which can potentially save developers’ efforts and improve the quality of software logging. Our findings shed light on research opportunities that leverage advances in NMT techniques for automated generation and suggestion of logging statements. Zishuo Ding, Yiming Tang 0002, Heng Li 0007, Weiyi Shang |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2023 | On the Temporal Relations between Logging and CodeabstractPrior work shows that misleading logging texts (i.e., the textual descriptions in logging statements) can be counterproductive for developers during their use of logs. One of the most important types of information provided by logs is the temporal information of the recorded system behavior. For example, a logging text may use a perfective aspect to describe a fact that an important system event has finished. Although prior work has performed extensive studies on automated logging suggestions, few of these studies investigate the temporal relations between logging and code. In this work, we make the first attempt to comprehensively study the temporal relations between logging and its corresponding source code. In particular, we focus on two types of temporal relations: (1) logical temporal relations, which can be inferred from the execution order between the logging statement and the corresponding source code; and (2) semantic temporal relations, which can be inferred based on the semantic meaning of the logging text. We first perform qualitative analyses to study these two types of logging-code temporal relations and the inconsistency between them. As a result, we derive rules to detect these two types of temporal relations and their inconsistencies. Based on these rules, we propose a tool named TempoLo to automatically detect the issues of temporal inconsistencies between logging and code. Through an evaluation of four projects, we find that TempoLo can effectively detect temporal inconsistencies with a small number of false positives. To gather developers' feedback on whether such inconsistencies are worth fixing, we report 15 detected instances from these projects to developers. 13 instances from three projects are confirmed and fixed, while two instances of the remaining project are pending at the time of this writing. Our work lays the foundation for describing temporal relations between logging and code and demonstrates the potential for a deeper understanding of the relationship between logging and code. Zishuo Ding, Yiming Tang 0002, Heng Li 0007, Weiyi Shang |
ICSE | 1 |
| 2023 | CoMSA: A Modeling-Driven Sampling Approach for Configuration Performance TestingabstractHighly configurable systems enable customers to flexibly configure the systems in diverse deployment environments. The flexibility of configurations also poses challenges for performance testing. On one hand, there exist a massive number of possible configurations; while on the other hand, the time and resources are limited for performance testing, which is already a costly process during software development. Modeling the performance of configurations is one of the solutions to reduce the cost of configuration performance testing. Although prior research proposes various modeling and sampling techniques to build configuration performance models, the sampling approaches used in the model typically do not consider the accuracy of the performance models, leading to potential sub-optimal performance modeling results in practice. In this paper, we present a modeling-driven sampling approach (CoMSA) to improve the performance modeling of highly configurable systems. The intuition of CoMSA is to select samples based on their uncertainties to the performance models. In other words, the configurations that have the more uncertain performance prediction results by the performance models are more likely to be selected as further training samples to improve the model. CoMSA is designed by considering both scenarios where 1) the software projects do not have historical performance testing results (cold start) and 2) there exist historical performance testing results (warm start). We evaluate the performance of our approach in four subjects, namely LRZIP, LLVM, x264, and SQLite. Through the evaluation result, we can conclude that our sampling approaches could highly enhance the accuracy of the prediction models and the efficiency of configuration performance testing compared to other baseline sampling approaches. Yuanjie Xia 0002, Zishuo Ding, Weiyi Shang |
ASE | 2 |
| 2023 | IoPV: On Inconsistent Option Performance VariationsabstractMaintaining a good performance of a software system is a primordial task when evolving a software system. The performance regression issues are among the dominant problems that large software systems face. In addition, these large systems tend to be highly configurable, which allows users to change the behaviour of these systems by simply altering the values of certain configuration options. However, such flexibility comes with a cost. Such software systems suffer throughout their evolution from what we refer to as “Inconsistent Option Performance Variation” (IoPV ). An IoPV indicates, for a given commit, that the performance regression or improvement of different values of the same configuration option is inconsistent compared to the prior commit. For instance, a new change might not suffer from any performance regression under the default configuration (i.e., when all the options are set to their default values), while altering one option’s value manifests a regression, which we refer to as a hidden regression as it is not manifested under the default configuration. Similarly, when developers improve the performance of their systems, performance regression might be manifested under a subset of the existing configurations. Unfortunately, such hidden regressions are harmful as they can go unseen to the production environment. In this paper, we first quantify how prevalent (in)consistent performance regression or improvement is among the values of an option. In particular, we study over 803 Hadoop and 502 Cassandra commits, for which we execute a total of 4,902 and 4,197 tests, respectively, amounting to 12,536 machine hours of testing. We observe that IoPV is a common problem that is difficult to manually predict. 69% and 93% of the Hadoop and Cassandra commits have at least one configuration that hides a performance regression. Worse, most of the commits have different options or tests leading to IoPV and hiding performance regressions. Therefore, we propose a prediction model that identifies whether a given combination of commit, test, and option (CTO) manifests an IoPV. Our evaluation for different models shows that random forest is the best performing classifier, with a median AUC of 0.91 and 0.82 for Hadoop and Cassandra, respectively. Our paper defines and provides scientific evidence about the IoPV problem and its prevalence, which can be explored by future work. In addition, we provide an initial machine learning model for predicting IoPV. Jinfu Chen 0002, Zishuo Ding, Yiming Tang 0002, Mohammed Sayagh, Heng Li 0007, Bram Adams, Weiyi Shang |
ESEC/SIGSOFT FSE | 2 |
| 2023 | StableYolo: Optimizing Image Generation for Large Language Models
Harel Berger, Aidan Dakhama, Zishuo Ding, Karine Even-Mendoza, David A. Kelly, Héctor D. Menéndez 0001, Rebecca Moussa, Federica Sarro |
SSBSE | 3 |
| 2023 | Towards Learning Generalizable Code Embeddings Using Task-agnostic Graph Convolutional NetworksabstractCode embeddings have seen increasing applications in software engineering (SE) research and practice recently. Despite the advances in embedding techniques applied in SE research, one of the main challenges is their generalizability. A recent study finds that code embeddings may not be readily leveraged for the downstream tasks that the embeddings are not particularly trained for. Therefore, in this article, we propose GraphCodeVec , which represents the source code as graphs and leverages the Graph Convolutional Networks to learn more generalizable code embeddings in a task-agnostic manner. The edges in the graph representation are automatically constructed from the paths in the abstract syntax trees, and the nodes from the tokens in the source code. To evaluate the effectiveness of GraphCodeVec , we consider three downstream benchmark tasks (i.e., code comment generation, code authorship identification, and code clones detection) that are used in a prior benchmarking of code embeddings and add three new downstream tasks (i.e., source code classification, logging statements prediction, and software defect prediction), resulting in a total of six downstream tasks that are considered in our evaluation. For each downstream task, we apply the embeddings learned by GraphCodeVec and the embeddings learned from four baseline approaches and compare their respective performance. We find that GraphCodeVec outperforms all the baselines in five out of the six downstream tasks, and its performance is relatively stable across different tasks and datasets. In addition, we perform ablation experiments to understand the impacts of the training context (i.e., the graph context extracted from the abstract syntax trees) and the training model (i.e., the Graph Convolutional Networks) on the effectiveness of the generated embeddings. The results show that both the graph context and the Graph Convolutional Networks can benefit GraphCodeVec in producing high-quality embeddings for the downstream tasks, while the improvement by Graph Convolutional Networks is more robust across different downstream tasks and datasets. Our findings suggest that future research and practice may consider using graph-based deep learning methods to capture the structural information of the source code for SE tasks. Zishuo Ding, Heng Li 0007, Weiyi Shang, Tse-Hsun (Peter) Chen |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2022 | LoGenText: Automatically Generating Logging Texts Using Neural Machine TranslationabstractThe textual descriptions in logging statements (i.e., logging texts) are printed during system executions and exposed to multiple stakeholders including developers, operators, users, and regulatory authorities. Writing proper logging texts is an important but often challenging task for developers. However, despite extensive research on automated logging suggestions, research on suggesting logging texts rarely exists. In this paper, we present LoGenText, an automated approach that generates logging texts by translating the related source code into short textual descriptions. LoGenText takes the preceding source code of a logging text as the input and considers other context information such as the location of the logging statement, to automatically generate the logging text using neural machine translation models. We evaluate LoGenText on 10 open-source projects, and compare the automatically generated logging texts with the developer-inserted logging texts in the source code. We find that LoGenText generates logging texts that achieve BLEU scores of 23.3 to 41.8 and ROUGE-L scores of 42.1 to 53.9, which outperforms the state-of-the-art approach by a large margin. In addition, we perform a human evaluation involving 42 participants, which further demonstrates the quality of the logging texts generated by LoGenText. Our work is an important step towards automated generation of logging statements, which can potentially save developers' efforts and improve the quality of software logging. Zishuo Ding, Heng Li 0007, Weiyi Shang |
SANER | 1 |
| 2022 | Can pre-trained code embeddings improve model performance? Revisiting the use of code embeddings in software engineering tasks
Zishuo Ding, Heng Li 0007, Weiyi Shang, Tse-Hsun (Peter) Chen |
Empir. Softw. Eng. | 1 |
| 2020 | Towards the use of the readily available tests from the release pipeline as performance tests: are we there yet?abstractPerformance is one of the important aspects of software quality. Performance issues exist widely in software systems, and the process of fixing the performance issues is an essential step in the release cycle of software systems. Although performance testing is widely adopted in practice, it is still expensive and time-consuming. In particular, the performance testing is usually conducted after the system is built in a dedicated testing environment. The challenges of performance testing make it difficult to fit into the common DevOps process in software development. On the other hand, there exist a large number of tests readily available, that are executed regularly within the release pipeline during software development. In this paper, we perform an exploratory study to determine whether such readily available tests are capable of serving as performance tests. In particular, we would like to see whether the performance of these tests can demonstrate performance improvements obtained from fixing real-life performance issues. We collect 127 performance issues from Hadoop and Cassandra, and evaluate the performance of the readily available tests from the commits before and after the performance issue fixes. We find that most of the improvements from the fixes to performance issues can be demonstrated using the readily available tests in the release pipeline. However, only a very small portion of the tests can be used for demonstrating the improvements. By manually examining the tests, we identify eight reasons that a test cannot demonstrate performance improvements even though it covers the changed source code of the issue fix. Finally, we build random forest classifiers determining the important metrics influencing the readily available tests (not) being able to demonstrate performance improvements from issue fixes. We find that the test code itself and the source code covered by the test are important factors, while the factors related to the code changes in the performance issues fixes have a low importance. Practitioners may focus on designing and improving the tests, instead of fine-tuning tests for different performance issues fixes. Our findings can be used as a guideline for practitioners to reduce the amount of effort spent on leveraging and designing tests that run in the release pipeline for performance assurance activities. Zishuo Ding, Jinfu Chen 0002, Weiyi Shang |
ICSE | 1 |