VLDB 2026 Research / reviewers in the wild / expert
Liang Dou 0001
dblp:10/10585-1
· DBLP profile ↗
13ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0003-3044-3841ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RetroProphet: Think One More Edit in Single-Step Retrosynthesis
Yu Zhang 0112, Luojian Xie, Chunyun Xiao, Liang Dou 0001 |
ICIC (28) | 5 |
| 2025 | Copy-Augmented Representation for Structure Invariant Template-Free RetrosynthesisabstractRetrosynthesis prediction is fundamental to drug discovery and chemical synthesis, requiring the identification of reactants that can produce a target molecule. Current templatefree methods struggle to capture the structural invariance inherent in chemical reactions, where substantial molecular scaffolds remain unchanged, leading to unnecessarily large search spaces and reduced prediction accuracy. We introduce C-SMILES, a novel molecular representation that decomposes traditional SMILES into element-token pairs with five special tokens, effectively minimizing editing distance between reactants and products. Building upon this representation, we incorporate a copy-augmented mechanism that dynamically determines whether to generate new tokens or preserve unchanged molecular fragments from the product. Our approach integrates SMILES alignment guidance to enhance attention consistency with ground-truth atom mappings, enabling more chemically coherent predictions. Comprehensive evaluation on USPTO-50K and large-scale USPTO-FULL datasets demonstrates significant improvements: 67.2 % top-1 accuracy on USPTO-50K and 50.8% on USPTO-FULL, with 99.9 % validity in generated molecules. This work establishes a new paradigm for structure-aware molecular generation with direct applications in computational drug discovery. Jiaxi Zhuang, Yu Zhang 0112, Liang Dou 0001, Aimin Zhou |
BIBM | 3 |
| 2023 | Heterogeneous Directed Hypergraph Neural Network over abstract syntax tree (AST) for Code ClassificationabstractCode classification is a difficult issue in program understanding and automatic coding.Due to the elusive syntax and complicated semantics in programs, most existing studies use techniques based on abstract syntax tree (AST) and graph neural network (GNN) to create code representations for code classification.These techniques utilize the structure and semantic information of the code, but they only take into account pairwise associations and neglect the high-order correlations that already exist between nodes in the AST, which may result in the loss of code structural information.On the other hand, while a general hypergraph can encode high-order data correlations, it is homogeneous and undirected which will result in a lack of semantic and structural information such as node types, edge types, and directions between child nodes and parent nodes when modeling AST.In this study, we propose to represent AST as a heterogeneous directed hypergraph (HDHG) and process the graph by heterogeneous directed hypergraph neural network (HDHGN) for code classification.Our method improves code understanding and can represent high-order data correlations beyond paired interactions.We assess heterogeneous directed hypergraph neural network (HDHGN) on public datasets of Python and Java programs.Our method outperforms previous AST-based and GNN-based methods, which demonstrates the capability of our model. Guang Yang 0068, Tiancheng Jin, Liang Dou 0001 |
SEKE | 3 |
| 2022 | TBPM-DDIE: Transformer Based Pretrained Method for predicting Drug-Drug Interactions EventsabstractMillions of patients die from Drug-Drug Interactions (DDIs) each year and therefore DDIs have attracted widespread attention. A large number of deep learning-based methods have been used to testify whether there is an interaction between drugs in previous studies. However, few researchers pay attention to the specific interaction events between drugs. Recently, some new models for predicting DDIs events emerged, based on labeled drug pairs and leaving the unlabeled drug data not fully utilized. To make full use of unlabeled data to help predict the specific interaction events between drugs, we propose Transformer Based Pretrained Methods for improving the prediction of Drug-Drug Interactions Events (TBPM-DDIE) to extract a latent vector with drug structure and semantic information and then concatenate the vector and the diverse features of drugs as input to the DDIs events classifier. The TBPM-DDIE model can be divided into three parts, including the Transformer pretrained part, the Diverse features of drugs part, and the DDIs events classifier part. We have done experiments on the real-world dataset and compared it with the latest models. The results show that TBPM-DDIE can achieve state-of-the-art effects on predicting DDIs events. Zhenjie Shao, Liang Dou 0001 |
COMPSAC | 3 |
| 2021 | Mixture Density Networks for Tropical Cyclone Tracks Prediction in South China SeaabstractThe forecasting of Tropical Cyclone (TC) Tracks in the South China Sea is important to cope with the associated disasters. But increasing its forecasting accuracy is a hard thing due to many factors. The main objective in the presented study is to develop models to deliver more accurate forecasts of TC Tracks over the South China Sea. The model proposed in this study is the TC Tracks Probability Forecasting Framework based on the Mixture Density Network (MDN). The TC Tracks Probability Forecasting Framework calculates the joint probability distribution of latitude and longitude and consists of latitude MDN and longitude MDN. MDN is a method that models the conditional probability distribution of the target data by the conventional neural network and mixture model. Forecast error is measured by calculating the distance between the real position and forecast position of TC. A decrease of 19.49 km in mean forecast error is obtained by our proposed model compared to the stepwise regression model, which is widely used in TC Tracks forecast. What's more, the model gives the probability distribution of each track prediction. This message can be used well in the TC track forecast cone. Fengyun Hao, Liang Dou 0001 |
IJCNN | 2 |
| 2021 | FinFuzzer: One Step Further in Fuzzing Fintech SystemsabstractComprehensive testing is of high importance to ensure the reliability of software systems, especially for systems with high stakes such as FinTech systems. In this paper, we share our observations of the Ant Group’s status quo in testing their financial services, specifically on the importance of properly transforming relevant external environment settings and prioritizing input object fields for mutation during automated fuzzing. Based on these observations, we propose FinFuzzer, an automated fuzz testing framework that detects and transforms relevant environmental settings into system inputs, prioritizes input object fields, and mutates system inputs on both environment settings and high-priority object fields. Our evaluation of FinFuzzer against four FinTech systems developed by the Ant Group shows that FinFuzzer can outperform a state-of-the-art approach in terms of line coverage in much shorter time. Qingshun Wang, Lihua Xu, Haotian Zhang 0026, Liang Dou 0001, Liang He 0001, Tao Xie 0001 |
ASE | 6 |
| 2019 | Enhancing the Healthcare Retrieval with a Self-adaptive Saturated Density Function
Yang Song 0010, Wenxin Hu, Liang He 0001, Liang Dou 0001 |
PAKDD (1) | 4 |
| 2019 | FinExpert: domain-specific test generation for FinTech systemsabstractTo assure high quality of software systems, the comprehensiveness of the created test suite and efficiency of the adopted testing process are highly crucial, especially in the FinTech industry, due to a FinTech system’s complicated system logic, mission-critical nature, and large test suite. However, the state of the testing practice in the FinTech industry still heavily relies on manual efforts. Our recent research efforts contributed our previous approach as the first attempt to automate the testing process in China Foreign Exchange Trade System (CFETS) Information Technology Co. Ltd., a subsidiary of China’s Central Bank that provides China’s foreign exchange transactions, and revealed that automating test generation for such complex trading platform could help alleviate some of these manual efforts. In this paper, we investigate further the dilemmas faced in testing the CFETS trading platform, identify the importance of domain knowledge in its testing process, and propose a new approach of domain-specific test generation to further improve the effectiveness and efficiency of our previous approach in industrial settings. We also present findings of our empirical studies of conducting domain-specific testing on subsystems of the CFETS Trading Platform. Tiancheng Jin, Qingshun Wang, Lihua Xu, Chunmei Pan, Liang Dou 0001, Haifeng Qian, Liang He 0001, Tao Xie 0001 |
ESEC/SIGSOFT FSE | 5 |
| 2018 | Mitigating SDN Flow Table OverflowabstractThe flow table in OpenFlow switches plays a critical role in OpenFlow-based Software Defined Networking (SDN), which stores the rules populated by the controllers for controlling and directing the packet flows in SDN. The limited capacity of flow table becomes a performance bottleneck of SDN and new target for malicious attacks, as well. This paper analyzes the timeout impact on the flow table performance and proposes the Dynamic LRU flow entry rule eviction algorithm to mitigate the SDN flow table overflow and improve the SDN performance. Hanwu Luo, Wenzhen Li, Liang Dou 0001 |
COMPSAC (1) | 4 |
| 2018 | FACTS: automated black-box testing of FinTech systemsabstractFinTech, short for ``financial technology,'' has advanced the process of transforming financial business from a traditional manual-process-driven to an automation-driven model by providing various software platforms. However, the current FinTech-industry still heavily depends on manual testing, which becomes the bottleneck of FinTech industry development. To automate the testing process, we propose an approach of black-box testing for a FinTech system with effective tool support for both test generation and test oracles. For test generation, we first extract input categories from business-logic specifications, and then mutate real data collected from system logs with values randomly picked from each extracted input category. For test oracles, we propose a new technique of priority differential testing where we evaluate execution results of system-test inputs on the system's head (i.e., latest) version in the version repository (1) against the last legacy version in the version repository (only when the executed test inputs are on new, not-yet-deployed services) and (2) against both the currently-deployed version and the last legacy version (only when the test inputs are on existing, deployed services). When we rank the behavior-inconsistency results for developers to inspect, for the latter case, we give the currently-deployed version as a higher-priority source of behavior to check. We apply our approach to the CSTP subsystem, one of the largest data processing and forwarding modules of the China Foreign Exchange Trade System (CFETS) platform, whose annual total transaction volume reaches 150 trillion US dollars. Extensive experimental results show that our approach can substantially boost the branch coverage by approximately 40%, and is also efficient to identify common faults in the FinTech system. Qingshun Wang, Lintao Gu, Minhui Xue 0001, Lihua Xu, Wenyu Niu, Liang Dou 0001, Liang He 0001, Tao Xie 0001 |
ESEC/SIGSOFT FSE | 6 |
| 2017 | Mechanized semantics and refinement of UML-StatechartsabstractThe Unified Modeling Language (UML) is an industry standard for modeling analysis and design. However, the semantics of UML is not precisely defined and the correctness of refinement relations cannot be verified. In this study, we use the theorem proof assistant Coq to formalize and mechanize the semantics of UML-Statecharts and the refinement relations between models. Based on the mechanized semantics, the desired properties of both the semantics and the refinement relations can be described and proven as predicates and lemmas. This approach provides a promising way to obtain certified fault-free modeling and refinement. Feng Sheng, Liang Dou 0001, Zongyuan Yang |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2015 | PerfBlower: Quickly Detecting Memory-Related Performance Problems via AmplificationabstractPerformance problems in managed languages are extremely difficult to find. Despite many efforts to find those problems, most existing work focuses on how to debug a user-provided test execution in which performance problems already manifest. It remains largely unknown how to effectively find performance bugs before software release. As a result, performance bugs often escape to production runs, hurting software reliability and user experience. This paper describes PerfBlower, a general performance testing framework that allows developers to quickly test Java programs to find memory-related performance problems. PerfBlower provides (1) a novel specification language ISL to describe a general class of performance problems that have observable symptoms; (2) an automated test oracle via \emph{virtual amplification}; and (3) precise reference-path-based diagnostic information via object mirroring. Using this framework, we have amplified three different types of problems. Our experimental results demonstrate that (1) ISL is expressive enough to describe various memory-related performance problems; (2) PerfBlower successfully distinguishes executions with and without problems; 8 unknown problems are quickly discovered under small workloads; and (3) PerfBlower outperforms existing detectors and does not miss any bugs studied before in the literature. Lu Fang 0003, Liang Dou 0001, Guoqing Harry Xu |
ECOOP | 2 |
| 2013 | A metamodeling approach for pattern specification and managementabstractThe formal specification of design patterns is central to pattern research and is the foundation of solving various pattern-related problems. In this paper, we propose a metamodeling approach for pattern specification, in which a pattern is modeled as a meta-level class and its participants are meta-level references. Instead of defining a new metamodel, we reuse the Unified Modeling Language (UML) metamodel and incorporate the concepts of Variable and Set into our approach, which are unavailable in the UML but essential for pattern specification. Our approach provides straightforward solutions for pattern-related problems, such as pattern instantiation, evolution, and implementation. By integrating the solutions into a single framework, we can construct a pattern management system, in which patterns can be instantiated, evolved, and implemented in a correct and manageable way. Liang Dou 0001, Zongyuan Yang |
J. Zhejiang Univ. Sci. C | 1 |