VLDB 2026 Research / reviewers in the wild / expert
Dong Liu 0025
dblp:98/1737-25
· DBLP profile ↗
10ranked-venue papers
4as first author
6since 2021 · last 2025
0009-0001-6940-0083ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TRACED: A Temporal Graph Neural Networks-based Model for Data PrefetchingabstractIn modern microarchitectures, machine-learning-based prefetchers use past memory requests to learn access patterns and predict memory addresses, thereby prefetching data into the cache to mitigate the processor-memory speed gap. However, they face two key challenges in capturing irregular access patterns generated by complex data structures and algorithms. One is data dispersion: the disorderliness of memory addresses makes it difficult for prefetchers to extract meaningful data features. The other is temporal and spatial complexity: existing prefetchers fail to effectively learn temporal and spatial characteristics, and thus are unable to explore more complex access patterns. To resolve these challenges, we propose TRACED, a novel temporal graph neural network-based prefetcher aimed at learning access patterns of memory addresses. TRACED consists of two key components: a dynamic clustering component and a temporal graph neural network component. In the dynamic clustering component, we introduce a similarity function to quantify the similarity of memory addresses. Based on the quantified similarity, we dynamically group unordered memory addresses into different clusters. This ensures that the memory addresses in each cluster are ordered and change smoothly, thus resolving the first challenge. The temporal graph neural network component constructs a spatiotemporal graph to represent relationships among memory addresses. This helps capture temporal and spatial characteristics both across and within clusters, thus resolving the second challenge. This article demonstrates the effectiveness of the proposed prefetcher through experiments. Specifically, in terms of accuracy, TRACED outperforms BO, SPP, DOMINO, Delta-LSTM, and VOYAGER by 2.29%–40.83% on average. Furthermore, TRACED attains remarkable coverage of 55.67% and IPC of 43.75%, outperforming all competing approaches in both metrics. He Jiang 0001, Liuwei Fu, Dong Liu 0025, Zhilei Ren, Yuting Chen 0001, Lei Qiao 0002 |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | What's Wrong With Low-Code Development Platforms? An Empirical Study of Low-Code Development Platform BugsabstractLow-code development platforms (LCDPs) are increasingly being introduced and leveraged by major IT enterprises to lower the threshold and promote the efficiency of software development. Like other software systems, LCDPs are also inevitable to have bugs. The bugs in LCDPs may cause unpredictable consequences as they pose risks to all the downstream software products. However, to the best of our knowledge, there exist no studies that ever consider the bugs caused by LCDPs. To handle the LCDP bugs better, in this article, we conduct an empirical study of the characteristics of LCDP bugs by examining 974 confirmed bugs of four dominant LCDPs (i.e., OutSystems, Mendix, Appsmith, and Budibase) from both commercial and open-source domains. These bugs are analyzed from three perspectives, including bug root causes, bug symptoms, and the affected stages of LCDPs. Based on the analysis, we obtain a series of valuable findings. For example, around 60% of the bugs reside in the stage of designing and specifying the developed applications. Over 37% of the bugs lead LCDPs to behave unexpectedly but without showing explicit signs. Moreover, the bugs relevant to the incorrect graphics of user interfaces are significant due to the characteristics of LCDPs. These findings point out the guidelines, challenges, and future directions to address LCDP bugs. Dong Liu 0025, He Jiang 0001, Shikai Guo, Yuting Chen 0001, Lei Qiao 0002 |
IEEE Trans. Reliab. | 1 |
| 2023 | A Comprehensive Study of WebAssembly Runtime BugsabstractWebAssembly runtime is the infrastructure for executing WebAssembly, which is widely used as an execution engine by web browsers or blockchain platforms. Bugs in the WebAssembly runtime can lead to unexpected behavior and even security vulnerabilities in any application that relies on it. Therefore, to aid developers in understanding the WebAssembly runtime, a thorough investigation of bugs in the WebAssembly runtime should be conducted. To accomplish this, we carry out the first empirical analysis of 867 real bugs across four popular WebAssembly runtimes (V8, SpiderMonkey, Wasmer, and Wasmtime). We analyze the WebAssembly runtime bug characteristics based on their root causes, symptoms, bug-fixing time, and the number of files and lines of code involved in the bug fixes. Here are a few major research findings: 1) Incorrect Algorithm Implementation accounts for 25.49% of WebAssembly runtime bugs, the most prevalent of all root causes; 2) The most prevalent symptom is Crash, which accounts for 56.86% of WebAssembly runtime bugs; 3) At the median, the bug-fixing time are 13, 4, 5, and 6 days for V8, SpiderMonkey, Wasmer, and Wasmtime respectively; 4) Over 50% of bug fixes in the four WebAssembly runtimes involve only one file, while more than 90% of bug fixes involve no more than 8 files; 5) The median source code lines for bug fixes for V8, SpiderMonkey, Wasmer, and Wasmtime are 18.5, 14, 26, and 36 lines, respectively. Overall, our research summarizes 18 findings and discusses the broad implications for WebAssembly runtime bug detection, localization, debugging, and repair based on the key findings. Zhide Zhou, Zhilei Ren, Dong Liu 0025, He Jiang 0001 |
SANER | 4 |
| 2022 | Surrogate-Assisted Multi-objective Optimization for Compiler Optimization Sequence Selection
Guojun Gao, Lei Qiao 0002, Dong Liu 0025, Shifei Chen, He Jiang 0001 |
PPSN (2) | 3 |
| 2022 | Detecting Compiler Bugs Via a Deep Learning-Based FrameworkabstractCompiler testing is the most widely used way to assure compiler quality. However, since compilers require a large number of sophisticated test programs as inputs, the existing approaches in compiler testing still have a limited capability in generating both syntactically valid and diverse test programs. In this paper, we propose DeepGen, a deep learning-based approach to support compiler testing through the inference of a generative model for compiler inputs. First, DeepGen trains a Transformer-XL model based on a large corpus of seed programs, and uses the trained model to generate syntactically valid programs. Then, DeepGen adopts a sampling strategy in the inference phase to generate diverse test programs. Finally, DeepGen leverages differential testing on the generated programs to discover compiler bugs. We have evaluated DeepGen over two popular C++ compilers GCC and LLVM, and the results confirm the effectiveness of our approach. DeepGen detects 35.29%, 53.33%, and 187.50% more bugs than three existing approaches, i.e. DeepSmith, DeepFuzz, and Csmith, respectively. In addition, 30.43% bugs detected by DeepGen are not detected by other approaches. Furthermore, DeepGen has successfully detected 38 bugs in the latest development versions of GCC and LLVM; 21 of them have been confirmed/fixed by the developers. Zhilei Ren, He Jiang 0001, Lei Qiao 0002, Dong Liu 0025, Zhide Zhou, Weiqiang Kong |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2022 | DPWord2Vec: Better Representation of Design Patterns in SemanticsabstractWith the plain text descriptions of design patterns, developers could better learn and understand the definitions and usage scenarios of design patterns. To facilitate the automatic usage of these descriptions, e.g., recommending design patterns by free-text queries, design patterns and natural languages should be adequately associated. Existing studies usually use texts in design pattern books as the representations of design patterns to calculate similarities with the queries. However, this way is problematic. Lots of information of design patterns may be absent from design pattern books and many words would be out of vocabulary due to the content limitation of these books. To overcome these issues, a more comprehensive method should be constructed to estimate the relatedness between design patterns and natural language words. Motivated by Word2Vec, in this study, we propose DPWord2Vec that embeds design patterns and natural language words into vectors simultaneously. We first build a corpus containing more than 400 thousand documents extracted from design pattern books, Wikipedia, and Stack Overflow. Next, we redefine the concept of context window to associate design patterns with words. Then, the design pattern and word vector representations are learnt by leveraging an advanced word embedding method. The learnt design pattern and word vectors can be universally used in textual description based design pattern tasks. An evaluation shows that DPWord2Vec outperforms the baseline algorithms by 24.2-120.9 percent in measuring the similarities between design patterns and words in terms of Spearman’s rank correlation coefficient. Moreover, we adopt DPWord2Vec on two typical design pattern tasks. In the design pattern tag recommendation task, the DPWord2Vec-based method outperforms two state-of-the-art algorithms by 6.6 and 32.7 percent respectively when considering$Recall@10$. In the design pattern selection task, DPWord2Vec improves the existing methods by 6.5-70.7 percent in terms of MRR. Dong Liu 0025, He Jiang 0001, Zhilei Ren, Lei Qiao 0002, Zuohua Ding |
IEEE Trans. Software Eng. | 1 |
| 2020 | Mining Design Pattern Use Scenarios and Related Design Pattern Pairs: A Case Study on Online Posts
Dong Liu 0025, Zhilei Ren, Zhongtian Long, Guojun Gao, He Jiang 0001 |
J. Comput. Sci. Technol. | 1 |
| 2018 | Unsupervised deep bug report summarizationabstractBug report summarization is an effective way to reduce the considerable time in wading through numerous bug reports. Although some supervised and unsupervised algorithms have been proposed for this task, their performance is still limited, due to the particular characteristics of bug reports, including the evaluation behaviours in bug reports, the diverse sentences in software language and natural language, and the domain-specific predefined fields. In this study, we conduct the first exploration of the deep learning network on bug report summarization. Our approach, called DeepSum, is a novel stepped auto-encoder network with evaluation enhancement and predefined fields enhancement modules, which successfully integrates the bug report characteristics into a deep neural network. DeepSum is unsupervised. It significantly reduces the efforts on labeling huge training sets. Extensive experiments show that DeepSum outperforms the comparative algorithms by up to 13.2% and 9.2% in terms of F-score and Rouge-n metrics respectively over the public datasets, and achieves the state-of-the-art performance. Our work shows promising prospects for deep learning to summarize millions of bug reports. He Jiang 0001, Dong Liu 0025, Zhilei Ren, Ge Li 0001 |
ICPC | 3 |
| 2017 | Length-Changeable Incremental Extreme Learning Machine
Youxi Wu, Dong Liu 0025, He Jiang 0001 |
J. Comput. Sci. Technol. | 2 |
| 2016 | FP-ELM: An online sequential learning algorithm for dealing with concept drift
Dong Liu 0025, Youxi Wu, He Jiang 0001 |
Neurocomputing | 1 |