VLDB 2026 Research / reviewers in the wild / expert
Dalin Zhang 0003
dblp:98/9560-3
· DBLP profile ↗
24ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0003-0346-7020ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Non-stationary Spatiotemporal Hawkes Process for Railway Delay Causality Learning
Jubao Cheng, Dalin Zhang 0003, Shunjie Yang, Yunjuan Peng, Rong-Hua Li 0001 |
DASFAA (4) | 2 |
| 2026 | Adaptive Multiprototype Gaussian Prototypical Network for Few-Shot UAV Spectrogram Classification
Shouyue Fang, Xiaoqiang Zhu, Yingying Yao, Zhenyan Ji, Dalin Zhang 0003 |
IEEE Internet Things J. | 7 |
| 2025 | Software Trustworthiness Assessment Based on Causal Evidence GraphabstractIn safety critical domains such as aerospace and defense, software must perform with high reliability, making trustworthiness assessment indispensable. The trustworthiness assessment integrates key attributes such as reliability and maintainability, drawing evidence from across the entire software life cycle and fusing it to determine overall trustworthiness. However, these indicators are not independent, since causal dependencies among them increase the complexity of evaluation. To address this, we build a life cycle Causal Evidence Graph to model cross stage links, causal directions, confounders, mediators, and calibrate evidence weights by causal consistency. To resolve conflicts, we use falsity, credibility, and uncertainty measures, apply game theoretic weighted aggregation, and then perform Dempster Shafer fusion. Preliminary experimental results show that our approach satisfies monotonicity, acceleration, sensitivity, and substitutability, and yields greater interpretability than methods that ignore causal relations. Dalin Zhang 0003, Zhenguo Ding |
APSEC | 2 |
| 2025 | Towards Adaptive Network Defense: A Self-evolving Threat Detection Framework
Chaoqun Guo, Dalin Zhang 0003 |
Inscrypt (2) | 2 |
| 2025 | Poster: SCL-IDS - A Semi-Supervised Continual Learning Framework for Adaptive Intrusion DetectionabstractAs modern network infrastructures grow in complexity, intrusion detection systems (IDS) face increasing challenges in detecting emerging threats under non-stationary conditions. Real-world traffic is characterized by continual attack evolution, severe label scarcity, and imbalanced distributions. Traditional IDS approaches often fail to maintain accuracy over time due to catastrophic forgetting, while supervised methods demand costly manual annotations. Chaoqun Guo, Dalin Zhang 0003 |
ICNP | 2 |
| 2025 | Advanced persistent threat detection via mining long-term features in provenance graphs
Fan Xu 0009, Qinxin Zhao, Nan Wang 0015, Meiqi Gao, Xuezhi Wen, Dalin Zhang 0003 |
Frontiers Comput. Sci. | 7 |
| 2024 | Path Exploration Strategy for Symbolic Execution based on Multi-strategy Active LearningabstractThis paper proposes a novel symbolic execution path exploration strategy named MS-ALS (Multi-strategy Active Learning Search). MS-ALS integrates multiple heuristic methods and introduces a machine learning model to learn symbolic states of program paths from the training set, aiming to predict the reward of symbolic states of program paths for selecting the optimal states. To obtain an accurate predictive model, this paper employs an active learning approach based on multiple query strategies. It selects symbolic states of program paths with high uncertainty, representativeness, and low redundancy from the pool of states to annotate, feeding back the state samples to the model to guide it towards more accurate predictions. This enables symbolic execution tools to explore input programs more efficiently. Experiments show that MS-ALS achieves higher code coverage and identifies more security violations compared to baseline methods. Additionally, test cases generated by MS-ALS also exhibit higher quality, improving AFL’s path discovery in fuzz testing when used as initial seeds. Lianying He, Dalin Zhang 0003, Dongqing Zhu, Junwen Zhang 0004, Jiqiang Liu |
Internetware | 2 |
| 2024 | GLADformer: A Mixed Perspective for Graph-Level Anomaly Detection
Fan Xu 0009, Nan Wang 0015, Hao Wu 0094, Xuezhi Wen, Dalin Zhang 0003, Siyang Lu, Binyong Li, Wei Gong 0001, Hai Wan, Xibin Zhao |
ECML/PKDD (6) | 5 |
| 2024 | Fairness based on anomaly score and adaptive weight in network attack detection
Xuezhi Wen, Meiqi Gao, Nan Wang 0015, Jiahui Ma, Dalin Zhang 0003, Xibin Zhao, Jiqiang Liu |
Inf. Sci. | 5 |
| 2024 | Enhanced evolutionary automated program repair by finer-granularity ingredients and better search algorithmsabstractSummary Bug repair is time consuming and tedious, which hampers software maintenance. To alleviate the burden, automated program repair (APR) is proposed and has been fruitful in the last decade. Evolutionary repair is the seminal work of this field and proliferated a family of approaches. The performance of evolutionary repair approaches is affected by two main factors: (1) search space, which defines all possible patches, and (2) search algorithms, which navigate the space. Although recent approaches have achieved remarkable progress, the main challenges of the two factors still remain. On one hand, the different kinds of search space are very coarse for containing correct patches. On the other hand, the search process guided by genetic algorithms is inefficient in finding the correct patches in an appropriate time budget. In this paper, we propose MicroRepair, a new evolutionary repair approach to address the two challenges. Rather than finding statement‐level patches like existing genetic repair approaches, MicroRepair enlarges the search space by breaking the statements into finer‐granularity ingredients that consist of AST leaves. As the search space grows exponentially, the former search algorithms may become inefficient in navigating the larger space. We utilize the best multiobjective search algorithm selected from our empirical comparison of a set of search algorithms. At last, we find redundancies search in the existing genetic process, and we further design a history‐aware search strategy to accelerate the process. We evaluated MicroRepair on 224 bugs of real‐world from the benchmark Defects4J and compared it with several state‐of‐the‐art repair approaches. The evaluation results show that MicroRepair correctly repaired 26 bugs with a precision of 62%, which significantly outperforms the state‐of‐the‐art evolutionary APR approaches in terms of precision. Moreover, the history‐aware search boosts the repair execution speed by 4% on average. Bo Wang 0050, Guizhuang Liu, Youfang Lin, Shuang Ren, Dalin Zhang 0003 |
J. Softw. Evol. Process. | 6 |
| 2024 | A Multi-Source Dynamic Temporal Point Process Model for Train Delay PredictionabstractTrain delay prediction is a key technology for intelligent train scheduling and passenger services. We propose a train delay prediction model that takes into account the asynchrony of train events, the dynamics of train operations, and the diversity of influencing factors. Firstly, we consider train operations as discrete sequences of train events and propose a train arrival neural temporal point process (TANTPP) framework focused on predicting train delays that explicitly models the asynchrony of train events. Secondly, we introduce a multi-source dynamic spatiotemporal embedding method for the feature encoder in the TANTPP framework, which enhances the capability to capture the features of train operation networks. Thirdly, to better capture the distribution of train events in the TANTPP framework, we utilize a log-normal mixture hybrid method to learn the probability density distribution of train arrival events. Finally, the experimental result on real-world datasets demonstrates that the TANTPP model outperforms current state-of-the-art models, reducing the MAE by 10.85%, the RMSE by 9.8%, the RRSE by 3.78% and the MAPE by 10.11% on average. To the best of our knowledge, this is the first study to utilize neural temporal point processes to enhance train delay prediction. Dalin Zhang 0003, Chenyue Du, Yunjuan Peng, Jiqiang Liu, Sabah Mohammed, Alessandro Calvi |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | MASTER: Multi-Source Transfer Weighted Ensemble Learning for Multiple Sources Cross-Project Defect PredictionabstractBackground:Multi-source cross-project defect prediction (MSCPDP) attempts to transfer defect knowledge learned from multiple source projects to the target project. MSCPDP has drawn increasing attention from academic and industry communities owing to its advantages compared with single-source cross-project defect prediction (SSCPDP). However, two main problems, which are how to effectively extract the transferable knowledge from each source dataset and how to measure the amount of knowledge transferred from each source dataset to the target dataset, seriously restrict the performance of existing MSCPDP models.Objective:In this paper, we propose a novel multi-source transfer weighted ensemble learning (MASTER) method for MSCPDP.Method:MASTER measures the weight of each source dataset based on feature importance and distribution difference and then extracts the transferable knowledge based on the proposed feature-weighted transfer learning algorithm. Experiments are performed on 30 software projects. We compare MASTER with the latest state-of-the-art MSCPDP methods with statistical test in terms of famous effort-unaware measures (i.e., PD, PF, AUC, and MCC) and two widely used effort-aware measures (Popt20% and IFA).Result:The experiment results show that: 1) MASTER can substantially improve the prediction performance compared with the baselines, e.g., an improvement of at least 49.1% in MCC, 48.1% in IFA; 2) MASTER significantly outperforms each baseline on most datasets in terms of AUC, MCC,Popt20% and IFA; 3) MSCPDP model significantly performs better than the mean case of SSCPDP model on most datasets and even outperforms the best case of SSCPDP on some datasets.Conclusion:It can be concluded that 1) it is very necessary to conduct MSCPDP, and 2) the proposed MASTER is a more promising alternative for MSCPDP. Haonan Tong, Dalin Zhang 0003, Jiqiang Liu, Weiwei Xing, Lingyun Lu, Wei Lu 0010, Yumei Wu |
IEEE Trans. Software Eng. | 2 |
| 2023 | DTC: Addressing the long-tailed problem in intrusion detection through the divide-then-conquer paradigmabstractIntrusion detection systems (IDS) analyze the monitored data to detect patterns or signatures that correspond to known cyberattack techniques, vulnerabilities, or deviations from established baselines. They employ different algorithms and techniques to identify potential threats. When building "Known Patterns and Signatures", it always faces the long-tailed problem, which refers to the imbalanced distribution of different types of network traffic or events in a dataset, which means that there are very few instances of certain types of intrusions compared to more common types of network traffic or benign events. Such a situation poses a great challenge to deep learning-based or machine learning-based detection models on how to handle it. Models may struggle to learn from long-tailed distributions because they tend to bias their predictions toward the majority, and perform well on common events but poorly on rare intrusions. To address this problem, different from previous methods concerning training the model on the whole samples to obtain a balanced data distribution, we first focus on dividing the whole samples into a balanced group and an imbalanced one, then, we train a detection model on the balanced group. Cycle over and over again. Specifically, We propose to use a Gaussian mixture flow filter to progressively perform sample aggregation, continuously transforming the long-tail distribution into a more balanced which allows us to train the classifier on the obtained balanced group. The separated training samples with high distribution balance make it easier to train subsequent classifiers and mitigate the head-to-tail bias. Through extensive experiments, we have achieved new state-of-the-art performance on common intrusion detection datasets such as UNSW-NB15, CIC-IDS2017, and NSL-KDD. These results demonstrate that it is possible to surpass carefully constructed balanced datasets by progressively distinguishing the head class and the tail class. Chaoqun Guo, Nan Wang 0015, Yuanlin Sun, Dalin Zhang 0003 |
ICPADS | 4 |
| 2023 | Fairness with adaptive weight in network attack detectionabstractNetwork attacks aim to exploit vulnerabilities inherent in network protocols, which is widely used in many real-world applications. In the process of network anomaly detection, most methods train the model by minimizing the average empirical risk of all samples. However, due to the uneven distribution of samples from different protocols, detection models tend to be biased against minority protocols groups. To address this issue, we propose an adaptive weight assignment method for network attack detection, which emphasizes more on error-prone samples in prediction and enhances adequate representation of minority groups for fairness. We conduct experiments on two widely used datasets, KDD and NSLKDD. According to the results, our method achieves better performance than state-of-the-art methods for classification and regression tasks, and is robust to label noise in the test dataset. Xuezhi Wen, Nan Wang 0015, Yuanlin Sun, Fan Xu 0009, Dalin Zhang 0003, Xibin Zhao |
ICPADS | 5 |
| 2023 | Deep Reinforcement Learning Guided Decision Tree Learning For Program SynthesisabstractSyntax-Guided Synthesis (SyGuS) is a general solving framework for program synthesis. Previous researchers have proposed a divide-and-conquer strategy to divide the grammar search space in order to search separately. This framework converts the program synthesis problem into a decision tree classification problem. It unifies the obtained partial results into a complete program by learning decision trees according to conditional expression statements. However, due to the unknowns and uncertainties of sample label categories and sample attribute features in program synthesis, as well as the limitation of traditional decision tree learning methods to immediate gains, it results in long solving times and the large size of the candidate program. To address the above problems, we propose a deep reinforcement learning-based synthesis strategy for syntax-guided programs: a) the decision tree learning process is modeled as a Markov chain decision process, addressing the learning of knowledge patterns under conceptual drift and the weighting of the importance of sample attribute features; b) a graph neural network is used to extract the feature information of the sample attributes as action embedding feature values; (c) a reward function is proposed to evaluate the classification accuracy of the sample attributes as a feedback mechanism. We have implemented our approach in a tool called RLSolver. According to the experiments, the solving times ofRLSolver are comparable to EUSolver on small-scale task datasets. RLSolver substantially reduces the solving times on large-scale task datasets, solving two more SyGus tasks than EUSolver in the same solving time limit. In addition, RLSolver reduces the number of decision tree learning rounds by 60% and the size of decision tree size by 40% after using the multi-round model training strategy and early-stop strategy. Mingrui Yang, Dalin Zhang 0003 |
SANER | 2 |
| 2023 | An Interpretable Station Delay Prediction Model Based on Graph Community Neural Network and Time-Series Fuzzy Decision TreeabstractHigh-speed train delay prediction has always been one of the important research issues in the railway dispatching. Accurate and interpretable delay prediction can enable staff to implement preventive measures and scheduling decisions in advance, and guide relevant departments to cooperate in completing complex transportation tasks, so as to improve rail transit operations, service quality, and the efficiency of train operation. This article proposes a new interpretable model based on graph community neural network and time-series fuzzy decision tree. This model can well capture the influence of spatiotemporal characteristics, train community structure, and multifactor in high-speed train station delay prediction. Besides, the time series fuzzy decision tree based on multiobjective optimization and reduced error pruning can mine potential decision rules to improve the model's interpretability, transparency, and high reliability. Finally, we prove that the prediction effect of the proposed model is superior than the other seven state-of-the-art models and our model is interpretable. Dalin Zhang 0003, Yunjuan Peng, Chenyue Du, Nan Wang 0015, Mincong Tang, Lingyun Lu, Jiqiang Liu |
IEEE Trans. Fuzzy Syst. | 1 |
| 2022 | Enhanced Evolutionary Automated Program Repair by Finer-Granularity Ingredients and Better Search AlgorithmsabstractBug repair is time-consuming and tedious, which hampers software maintenance. To alleviate the burden, automated program repair (APR) is proposed and has been fruitful in the last decade. Evolutionary repair is the seminal work of this field and proliferated a family of approaches. The performance of evolutionary repair approaches is affected by two main factors: (1) search space, which defines all possible patches, and (2) search algorithms, which navigates the space. Although recent approaches have achieved remarkable progress, the main challenges of the two factors still remain. On one hand, the different kinds of search space are very coarse for containing correct patches. On the other hand, the search process guided by genetic algorithms is inefficient to find the correct patches in an appropriate time budget. Bo Wang 0050, Guizhuang Liu, Youfang Lin, Shuang Ren, Dalin Zhang 0003 |
Internetware | 6 |
| 2022 | Deep-Reinforcement-Learning-based User-Preference-Aware Rate Adaptation for Video StreamingabstractOnline video is the most popular Internet application. As the throughput would frequently change under different network conditions, it is important to adaptively select the proper bitrate and improve user’s quality of experience. In this paper, we propose a new DRL-based rate adaption algorithm for video streaming, which holistically captures user’s preference of video contents, network throughput and buffer occupancy, and select the proper bitrate for video to improve the QoE. Specifically, we use 3D Convolutional neural (C3D) network to learn the spatio-temporal features, and implement the semantic analysis of videos. We also apply the Term Frequency-Inverse Document Frequency (TF-IDF) method to analyze the user’s preference of different scene types, according to its viewing history. The dynamic adaptive streaming is formulated as a Markov Decision Process (MDP) problem, and use the Actor-Critic (A3C) algorithm to dynamically choose the optimal bitrate. As corroborated by simulations, our algorithm can accurately obtain the user’s preference, keep the bitrate allocation consistent with the user’s preference, and maintain video quality. Compared with the state-of-the-art Pensieve algorithm, our algorithm improves the average QoE by at least 12.5%. It also has a significant improvement over other baseline methods. Lingyun Lu, Wei Ni 0001, Haifeng Du, Dalin Zhang 0003 |
WoWMoM | 5 |
| 2022 | Eliminating the high false-positive rate in defect prediction through BayesNet with adjustable weightabstractAbstract In defect prediction, a high false‐positive rate (FPR) caused by class imbalance not only increases the workload of testing and development but also consumes unnecessary costs. Many defect models against class imbalance have been proposed to improve the accuracy of defect prediction, but their ability to reduce FPR is unclear. To solve these problems, we first proposed a BayesNet with adjustable weights, called WBN, to reduce the FPR in software defect prediction, which is an algorithm independent of data preprocessing techniques. The mechanism of our WBN is to change the sampling probability of the misclassified instances when training the defect model, making the BayesNet model focus more on false alarm instances. And then, we investigate the FPR of five mainstream defect models for solving class imbalance and select them as comparison models to test the validity of our methods. The experimental result on eight open‐source projects shows that a) our WBN, in in‐version defect prediction (IVDP) and cross‐version defect prediction (CVDP), effectively reduces FPR with means of 0.384 and 0.322, respectively; b) compared with improved subclass discriminant analysis (ISDA) that is the lowest FPR in all control models, our WBN not only reduced the FPR but maintained recall whose mean value was 0.797, whereas ISDA did not, with an average recall of only 0.397; c) our WBN, in CVDP, not only reduces FPR, but also has significant superiority over five control defect models and baseline. Besides, we also found that the class imbalance difference between the test set and the training set has an impact on CVDP performance, recommending that practitioners choose the best dataset for CVDP from the defect data of the historical version through special technology. Yanyang Zhao, Dalin Zhang 0003, Yunzhan Gong |
Expert Syst. J. Knowl. Eng. | 3 |
| 2022 | ST-TLF: Cross-version defect prediction framework based transfer learningabstractCross-version defect prediction (CVDP) is a practical scenario in which defect prediction models are derived from defect data of historical versions to predict potential defects in the current version. Prior research employed defect data of the latest historical version as the training set using the empirical recommended method, ignoring the concept drift between versions, which undermines the accuracy of CVDP. We customized a Selected Training set and Transfer Learning Framework (ST-TLF) with two objectives: a) to obtain the best training set for the version at hand, proposing an approach to select the training set from the historical data; b) to eliminate the concept drift, designing a transfer strategy for CVDP. To evaluate the performance of ST-TLF, we investigated three research problems, covering the generalization of ST-TLF for multiple classifiers, the accuracy of our training set matching methods, and the performance of ST-TLF in CVDP compared against state-of-the-art approaches. The results reflect that (a) the eight classifiers we examined are all boosted under our ST-TLF, where SVM improves 49.74% considering MCC, as is similar to others; (b) when performing the best training set matching, the accuracy of the method proposed by us is 82.4%, while the experience recommended method is only 41.2%; (c) comparing the 12 control methods, our ST-TLF (with BayesNet), against the best contrast method P15-NB, improves the average MCC by 18.84%. Our framework ST-TLF with various classifiers can work well in CVDP. The training set selection method we proposed can effectively match the best training set for the current version, breaking through the limitation of relying on experience recommendation, which has been ignored in other studies. Also, ST-TLF can efficiently elevate the CVDP performance compared with random forest and 12 control methods. Yanyang Zhao, Yuwei Zhang 0003, Dalin Zhang 0003, Yunzhan Gong, Dahai Jin |
Inf. Softw. Technol. | 4 |
| 2022 | Train Time Delay Prediction for High-Speed Train Dispatching Based on Spatio-Temporal Graph Convolutional NetworkabstractTrain delay prediction can improve the quality of train dispatching, which helps the dispatcher to estimate the running state of the train more accurately and make reasonable dispatching decision. The delay of one train is affected by many factors, such as passenger flow, fault, extreme weather, dispatching strategy. The departure time of one train is generally determined by dispatchers, which is limited by their strategy and knowledge. The existing train delay prediction methods cannot comprehensively consider the temporal and spatial dependence between the multiple trains and routes. In this paper, we don’t try to predict the specific delay time of one train, but predict the collective cumulative effect of train delay over a certain period, which is represented by the total number of arrival delays in one station. We propose a deep learning framework, train spatio-temporal graph convolutional network (TSTGCN), to predict the collective cumulative effect of train delay in one station for train dispatching and emergency plans. The proposed model is mainly composed of the recent, daily and weekly components. Each component contains two parts: spatio-temporal attention mechanism and spatio-temporal convolution, which can effectively capture spatio-temporal characteristics. The weighted fusion of the three components produces the final prediction result. The experiments on the train operation data from China Railway Passenger Ticket System demonstrate that TSTGCN clearly outperforms the existing advanced baselines in train delay prediction. Dalin Zhang 0003, Yunjuan Peng, Daohua Wu, Hongwei Wang 0008, Hailong Zhang 0006 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Reasoning about recursive tree traversalsabstractTraversals are commonly seen in tree data structures, and performance-enhancing transformations between tree traversals are critical for many applications. Existing approaches to reasoning about tree traversals and their transformations are ad hoc, with various limitations on the classes of traversals they can handle, the granularity of dependence analysis, and the types of possible transformations. We propose Retreet, a framework in which one can describe general recursive tree traversals, precisely represent iterations, schedules and dependences, and automatically check data-race-freeness and transformation correctness. The crux of the framework is a stack-based representation for iterations and an encoding to Monadic Second-Order (MSO) logic over trees. Experiments show that Retreet can automatically verify optimizations for complex traversals on real-world data structures, such as CSS and cycletrees, which are not possible before. Our framework is also integrated with other MSO-based analysis techniques to verify even more challenging program transformations. Yanjun Wang 0010, Dalin Zhang 0003, Xiaokang Qiu |
PPoPP | 3 |
| 2015 | A hybrid static analysis refinement approach within internetware environmentabstractIn this paper, we propose a hybrid refinement approach to improve the accuracy of static analysis. It keeps condition constraints information during forward dataflow analysis and gets the satisfiability of a warning by a constraint solver taking as input such information and path conditions; data regression analysis can remedy the capability of handling loops and library calls of abstract interpretation technique. It has been implemented in our static analysis tool, Defect Testing System (DTS) and deployed on a internetware environment TRUSTIE. Experiment on a large number of C open source projects shows the great improvement this strategy makes. Dalin Zhang 0003, Gang Yin, Dahai Jin, Yunzhan Gong, Tianshuang Wu, Hailong Zhang 0006 |
Internetware | 1 |
| 2013 | Diagnosis-Oriented Alarm CorrelationsabstractDefect detection generally includes two stages: static analysis and alarm inspection. Helping the user in the alarm inspection task is a major challenge for current static analyzers. A large number of independent alarms are against the understanding and may lead developers and managers to reject the use of static analysis tools due to the overhead of alarm inspection. To help with the inspection tasks, we formally introduce alarm correlations. If the occurrence of one alarm causes another alarm to occur, we say they are correlated. We propose a framework for the investigation of the alarms, so as to help classifying them by their correlations. The underlying algorithms were implemented inside our static analysis tool. We choose one common semantic alarm as case study and proved that our method has the effect of reducing 33.1% of alarm identification. Using correlation information, we are able to automate alarm identification that previously had to be done manually. Dalin Zhang 0003, Dahai Jin, Yunzhan Gong, Hailong Zhang 0006 |
APSEC (1) | 1 |