VLDB 2026 Research / reviewers in the wild / expert
Dan Du
dblp:163/7753
· DBLP profile ↗
24ranked-venue papers
5as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 7 since 2021Security and privacy · 7 · 6 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CMD-EPD: A Graph Contrastive Learning Framework with Multi-Dimensional Fusion for Ethereum Phishing DetectionabstractThe burgeoning prevalence of Ethereum phishing behavior has iCSUR-2025-0155mposed substantial constraints on the advancement of blockchain finance, resulting in losses of more than $7.7 billion to date, so it is urgent to detect it in time. Currently, available detection methods usually focus on the spatial features within transaction graphs. These methods often employ shallow mining techniques on small samples. As a result, they may overlook certain aspects of interaction patterns, such as temporal behavior. Additionally, their data mining capability is limited due to the small sample sizes. In this study, we propose a graph contrastive learning framework to enrich features of accounts behavior patterns with restricted samples to overcome these limitations. Firstly, we construct an Ethereum interaction graph with the multi-graph involving more temporal information centered with labeled nodes and lighten it with our strategy. Secondly, to comprehensively characterize the accounts pattern, we design the encoder part with the GAT-LSTM model based on attention mechanism fusing statistical features , fine-grained temporal behavioral features and graph structural semantic features . Thirdly, to moderate the sparsity of phishing nodes, we employ data augmentation and contrastive learning to fully mine sparse node information. Moreover, we carried out an in-depth experimental evaluation. The CMD-EPD approach, boasting an F 1 -score of 0.87, outperformed all comparison methods. We also executed a thorough case study to analyze phishing accounts phenomenological indicators which back up the superiority of our framework. Chuyi Yan, Yinhao Qi, Xueying Han, Dan Du, Zhigang Lu 0002, Meng Shen 0001 |
ACM Trans. Priv. Secur. | 5 |
| 2025 | MPKAN: APT Attack Detection on Audit Logs via Graph Semantic EnhancementabstractAs cloud computing and mobile work blur traditional network boundaries, security measures like static firewalls and signature-based systems are becoming inadequate. Audit logs contain fine-grained OS-level information, but due to their vast volume and the complex relationships between entities, processing and analyzing them remains a significant challenge. In this paper, we introduce MPKAN, a new method for detecting APT attacks, which enhances the information input of nodes and edges, integrates node-level and edge-level information, and improves graph-level semantics. It uses meta-path random walks to enhance semantic connections between nodes, merges multiple edges in the provenance graph into a single edge while retaining the original edge information and operation sequence relationships, and by associating heterogeneous graph neighbors, utilizing the message passing mechanism to iteratively update states based on neighbor node information, and using a knowledge association network to integrate node-level and edge-level information, we can effectively capture local and global structural information in the graph. MPKAN's evaluations on the ATLAS and Darpa datasets demonstrate its excellent performance in complex attack scenarios, achieving an average accuracy of 0.9899 and an F1 score of 0.9853, confirming its effectiveness and efficiency. Dan Du, Yinhao Qi, Bo Jiang 0013, Zhigang Lu 0002 |
CSCWD | 2 |
| 2025 | An Approach for Attack Chain Context Inference and Completion Based on Large Language ModelsabstractAlert underreporting presents a significant challenge to the reconstruction of attack chains, as it often leads to the absence of critical information necessary for fully presenting the entire attack. To address this issue, this paper proposes an approach for attack chain context inference and completion based on Large Language Models. By integrating an attack knowledge base, this approach leverages LLM-driven inference to identify missing attack stages and uncover potential attack behaviors. Experimental results demonstrate that this approach can effectively detect omitted alerts and complete the attack chain, thereby enhancing the integrity of attack detection. Dan Du, Changzhi Zhao, Yunpeng Li 0006, Dongxu Han, Bo Jiang 0013, Zhigang Lu 0002 |
SMC | 1 |
| 2025 | Elucidating the interactions between Kinesin-5/BimC and the microtubule: insights from TIRF microscopy and molecular dynamics simulationsabstractKinesin-5 s are bipolar motor proteins that contribute to cell division by crosslinking and sliding apart antiparallel microtubules inside the mitotic spindle. However, the mechanism underlying the interactions between kinesin-5 and the microtubule remains poorly understood. In this study, we investigated the binding of BimC, a kinesin-5 motor from Aspergillus nidulans, to the microtubule using a combination of total internal reflection fluorescence (TIRF) microscopy and molecular dynamics (MD) simulations. TIRF microscopy experiments revealed that increasing the concentration of KCl in the motility buffer from 0 mM to 150 mM completely abolishes the ability of BimC to bind to the microtubule. Consistent with this experimental finding, MD simulations demonstrated a significant reduction in the strength of electrostatic interactions between BimC and microtubules at 150 mM KCl compared to 0 mM KCl. Furthermore, we identified several salt bridges at the BimC-microtubule interface, with positively charged residues on BimC interacting with negatively charged residues on the tubulin heterodimer. These results provide mechanistic insights into the role of electrostatic interactions in kinesin-5-microtubule binding, advancing our understanding of the molecular underpinnings of kinesin-5 motility. Wenhan Guo, Dan Du, Jason E. Sanchez, Weihong Qiu |
Briefings Bioinform. | 3 |
| 2024 | Multi-language Webshell Detection based on Abstract Syntax Tree and TreeLSTMabstractWebshell is a command execution environment existing in web containers, which is used by attackers to remotely control servers and illegally access website resources. Accurately detecting Webshells is of great significance for maintaining web security. Current research faces several challenges. On the one hand, in order to evade detection, Webshells use a large amount of obfuscation, and existing research methods often use source code or opcode, which cannot fully utilize the semantic and syntactic information of Webshell code. On the other hand, Webshells can be constructed using any web application programming language, while most existing methods only detect one or a few types of Webshells. This paper proposes a novel approach called WS-Tree, which effectively utilizes the semantics and syntax of Webshells by using abstract syntax tree as input features. The TreeLSTM model is used as an encoder to handle node relationships in the syntax tree, thereby achieving the detection of obfuscated and multi-language Webshells. We also propose a new dataset of Webshells containing obfuscated and non-obfuscated to prevent dataset leakage. Extensive experiments demonstrate that our proposed model performs better than the state-of-theart baselines under different webshell programming languages and improves model generalizability. Mengchuan Shang, Xueying Han, Changzhi Zhao, Zelin Cui, Dan Du, Bo Jiang 0013 |
CSCWD | 5 |
| 2024 | Deep Dive into Insider Threats: Malicious Activity Detection within EnterpriseabstractWith the digital transformation of enterprises, the increasing complexity of their internal information systems poses a growing challenge in terms of insider threats. Most existing research focuses on user-level and session-level insider threat detection, neglecting activity-level detection, leading to a lack of fine-grained insider threat detection. To tackle the aforementioned issue, we propose MADE, a novel method for detecting malicious activities within enterprise environments. MADE first encodes user multi-source activity logs into activity sequences and learns the semantic representations of activities within the sequences through embedding. Following this, we design an activity detection network based on Bidirectional Long Short-Term Memory (BiLSTM), Convolutional Neural Network (CNN), and Conditional Random Field (CRF). Combining adversarial training, our activity detection network learns the embedded activity sequences and identifies malicious activities. Extensive experimental results on the CERT R4.2 and R5.2 datasets demonstrate the effectiveness of our proposed MADE method. Haitao Xiao, Dan Du, Zhigang Lu 0002 |
CSCWD | 2 |
| 2024 | MAD-LLM: A Novel Approach for Alert-Based Multi-stage Attack Detection via LLMabstractIn the realm of cybersecurity, detecting multi-stage attacks is vital for uncovering the actual intentions and strategies of attackers. However, the detection of multi-stage attacks is fraught with challenges due to the proliferation of alerts from diverse sources, the heterogeneity of their formats, and the weak correlations among them. This study explores MAD-LLM, a novel approach for alert-based multi-stage attack detection using Large Language Model (LLM). Leveraging their advanced natural language processing capabilities, LLM demonstrates significant advantages in text comprehension and pattern recognition. This research attempts to aggregate and correlate security alerts using the prompt engineering capabilities of LLM to reconstruct multi-stage attack chain. Experimental results indicate that LLM exhibit excellent performance in the task of multi-stage attack detection, providing an innovative solution for cybersecurity defense. This study also offers insights and implications for the application of LLM in other fields. Dan Du, Xingmao Guan, Bo Jiang 0013, Huamin Feng |
ISPA | 1 |
| 2024 | A novel approach to study multi-domain motions in JAK1's activation mechanism based on energy landscapeabstractThe family of Janus Kinases (JAKs) associated with the JAK-signal transducers and activators of transcription signaling pathway plays a vital role in the regulation of various cellular processes. The conformational change of JAKs is the fundamental steps for activation, affecting multiple intracellular signaling pathways. However, the transitional process from inactive to active kinase is still a mystery. This study is aimed at investigating the electrostatic properties and transitional states of JAK1 to a fully activation to a catalytically active enzyme. To achieve this goal, structures of the inhibited/activated full-length JAK1 were modelled and the energies of JAK1 with Tyrosine Kinase (TK) domain at different positions were calculated, and Dijkstra's method was applied to find the energetically smoothest path. Through a comparison of the energetically smoothest paths of kinase inactivating P733L and S703I mutations, an evaluation of the reasons why these mutations lead to negative or positive regulation of JAK1 are provided. Our energy analysis suggests that activation of JAK1 is thermodynamically spontaneous, with the inhibition resulting from an energy barrier at the initial steps of activation, specifically the release of the TK domain from the inhibited Four-point-one, Ezrin, Radixin, Moesin-PK cavity. Overall, this work provides insights into the potential pathway for TK translocation and the activation mechanism of JAK1. Georgialina Rodriguez, Gaoshu Zhao, Jason E. Sanchez, Wenhan Guo, Dan Du, Omar J. Rodriguez Moncivais, Dehua Hu, Robert Arthur Kirken |
Briefings Bioinform. | 6 |
| 2024 | Unveiling shadows: A comprehensive framework for insider threat detection based on statistical and sequential analysis
Haitao Xiao, Zhigang Lu 0002, Dan Du |
Comput. Secur. | 5 |
| 2024 | Phishing behavior detection on different blockchains via adversarial domain adaptationabstractAbstract Despite the growing attention on blockchain, phishing activities have surged, particularly on newly established chains. Acknowledging the challenge of limited intelligence in the early stages of new chains, we propose ADA-Spear-an automatic phishing detection model utilizing a dversarial d omain a daptive learning which symbolizes the method’s ability to penetrate various heterogeneous blockchains for phishing detection. The model effectively identifies phishing behavior in new chains with limited reliable labels, addressing challenges such as significant distribution drift, low attribute overlap, and limited inter-chain connections. Our approach includes a subgraph construction strategy to align heterogeneous chains, a layered deep learning encoder capturing both temporal and spatial information, and integrated adversarial domain adaptive learning in end-to-end model training. Validation in Ethereum, Bitcoin, and EOSIO environments demonstrates ADA-Spear’s effectiveness, achieving an average F1 score of 77.41 on new chains after knowledge transfer, surpassing existing detection methods. Chuyi Yan, Xueying Han, Dan Du, Zhigang Lu 0002 |
Cybersecur. | 4 |
| 2024 | AKGNN: Attribute Knowledge Graph Neural Networks Recommendation for Corporate Volunteer ActivitiesabstractDue to the collective decision-making nature of enterprises, the process of accepting recommendations is predominantly characterized by an analytical synthesis of objective requirements and cost-effectiveness, rather than being rooted in individual interests. This distinguishes enterprise recommendation scenarios from those tailored for individuals or groups formed by similar individuals, rendering traditional recommendation algorithms less applicable in the corporate context. To overcome the challenges, by taking the corporate volunteer as an example, which aims to recommend volunteer activities to enterprises, we propose a novel recommendation model calledAttributeKnowledgeGraphNeuralNetworks (AKGNN). Specifically, a novel comprehensive attribute knowledge graph is constructed for enterprises and volunteer activities, based on which we obtain the feature representation. Then we utilize anextendedVariationalAuto-Encoder (eVAE) model to learn the preferences representation and then we utilize a GNN model to learn the comprehensive representation with representation of the similar nodes. Finally, all the comprehensive representations are input to the prediction layer. Extensive experiments have been conducted on real datasets, confirming the advantages of the AKGNN model. We delineate the challenges faced by recommendation algorithms in Business-to-Business (B2B) platforms and introduces a novel research approach utilizing attribute knowledge graphs. Dan Du, Pei-Yuan Lai, Yan-Fei Wang, De-Zhang Liao, Min Chen 0003 |
IEEE Trans. Big Data | 1 |
| 2023 | TAElog: A Novel Transformer AutoEncoder-Based Log Anomaly Detection Method
Changzhi Zhao, Kezhen Huang, Xueying Han, Dan Du, Yutian Zhou, Zhigang Lu 0002 |
Inscrypt (2) | 5 |
| 2023 | HANDOM: Heterogeneous Attention Network Model for Malicious Domain Detection
Qing Wang 0041, Cong Dong, Shijie Jian, Dan Du, Zhigang Lu 0002, Yinhao Qi, Dongxu Han, Xiaobo Ma 0001, Fei Wang 0014 |
Comput. Secur. | 4 |
| 2023 | SA-RPN: A Spacial Aware Region Proposal Network for Acne DetectionabstractAutomated detection of skin lesions offers excellent potential for interpretative diagnosis and precise treatment of acne vulgar. However, the blurry boundary and small size of lesions make it challenging to detect acne lesions with traditional object detection methods. To better understand the acne detection task, we construct a new benchmark dataset named AcneSCU, consisting of 276 facial images with 31777 instance-level annotations from clinical dermatology. To the best of our knowledge, AcneSCU is the first acne dataset with high-resolution imageries, precise annotations, and fine-grained lesion categories, which enables the comprehensive study of acne detection. More importantly, we propose a novel method called Spatial Aware Region Proposal Network (SA-RPN) to improve the proposal quality of two-stage detection methods. Specifically, the representation learning for the classification and localization task is disentangled with a double head component to promote the proposals for hard samples. Then, Normalized Wasserstein Distance of each proposal is predicted to improve the correlation between the classification scores and the proposals' intersection-over-unions (IoUs). SA-RPN can serve as a plug-and-play module to enhance standard two-stage detectors. Extensive experiments are conducted on both AcneSCU and the public dataset ACNE04, and the results show that the proposed method can consistently outperform state-of-the-art methods. Jianwei Zhang 0016, Lei Zhang 0005, Junyou Wang, Jiaqi Li 0019, Xian Jiang, Dan Du |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | An Approach for Predicting the Costs of Forwarding Contracts using Gradient BoostingabstractPredicting the cost of forwarding contract is a severe challenge to road transport management system.The transportation cost of a forwarding contract often depends on many factors.It is hard for humans to evaluate the various factors in transportation and calculate the cost of forwarding contract.In this paper, we propose an approach to address such a problem by following the sequence of machine learning steps which consist of data analysis, feature engineering and model construction.First, we conduct a detailed analysis of the given data.Then, we generate effective features to characterize the cost of forwarding contract and eliminate redundant features.Finally, in the model construction phase, we propose a gradient boosting decision tree based method to train and predict the cost of forwarding contract.The proposed approach achieves RMSE scores of 0.1391 on the test set, which is the 2 nd final score in the competition. Haitao Xiao, Dan Du, Zhigang Lu 0002 |
FedCSIS | 3 |
| 2022 | DGGCN: Dictionary based DGA detection method based on DomainGraph and GCNabstractNowadays, malware uses Algorithmically Generated Domains (AGDs) to establish communication with Command and Control (C&C) servers. Dictionary based Domain Generation Algorithm (DGA) selects words from the frequently changed dictionaries to generate AGDs similar to benign domains, which degrades the accuracy of string based detection method. To combat this, we propose a DGA detection method based on DomainGraph and GCN (Graph Convolutional Network) which detects cross-dictionary AGDs based on the association relation between domains instead of lexical features. Starting from the association relation between domains rather than the lexical features of the domain itself, we can detect the unknown AGDs from a known AGD, regardless of the DGA dictionary they use. The proposed method exploits the fact that string association of benign domains is weak, while AGDs' association is strong. DGGCN composes a domain segmentation method, constructs a graph composed of domains (DomainGraph) based on segmentations and adopts GCN to detect AGDs. We conduct the experiments on public datasets under three settings: detecting AGDs generated by familiar dictionaries, unfamiliar dictionaries and confusing dictionaries. The results reveal that DGGCN can detect cross-dictionary AGDs similar to benign domains more accurately and robustly. Haoran Jiao, Qing Wang 0041, Zhaoshan Fan, Dan Du |
ICCCN | 5 |
| 2022 | DCC-Find: DNS Covert Channel Detection by Features Concatenation-Based LSTMabstractDNS (Domain Name System) plays an important role in network communication and it is rarely blocked by firewalls and intrusion detection systems (IDS). It is a suitable way for attackers to build DCC (DNS Covert Channel), which is used for data exfiltration. In recent years, some DCC detection methods have been proposed based on deep learning and there is no need for manual feature extraction. However, some expert knowledge is helpful to express the DNS characteristic. In this paper, we propose a FC-LSTM (Features Concatenation-based LSTM) model to detect DCC. The statistical features are concatenated with the output features of the LSTM model. This method makes the expression of DNS domain names more abundant. The experimental results have shown that the DCC traffic can be identified from normal traffic via this model, and the recognition rate is significantly improved compared with the traditional LSTM model and CNN model. In addition, we implement multi-classification in terms of the DCC tools (some of them are used in APT32). We also add generalization DNS packets (simulating APT34 traffic using DCC for stealing and attacking) to verify the robustness of our model. The FC-LSTM model has a good detection performance as well. Dongxu Han, Pu Dong, Xiang Cui, Jiawen Diao, Qing Wang 0041, Dan Du |
TrustCom | 7 |
| 2022 | Only Header: a reliable encrypted traffic classification framework without privacy risk
Susu Cui, Cong Dong, Zhigang Lu 0002, Dan Du |
Soft Comput. | 5 |
| 2021 | WP-GBDT: An Approach for Winner Prediction using Gradient Boosting Decision TreeabstractPredicting victories in video games from rich history of gameplay logs is a severe challenge to game developers. It is hard for humans to evaluate the real-time game situation and predict who will win the video game. In this paper, we propose an approach to this problem by following the sequence of machine learning steps which consist of feature engineering, feature selection, and model construction. We conduct a detailed analysis of the game logs and generate effective features from different granularity gameplay logs in the feature engineering phase. Then, we design a group based recursive feature elimination method for feature selection. In model construction, we present an ensemble approach that combines stacking and averaging for prediction to improve the generalization performance of models. The proposed approach achieves AUC scores of 0.8997 on the test set, which is the highest final score in the competition. Haitao Xiao, Dan Du, Zhigang Lu 0002 |
IEEE BigData | 3 |
| 2020 | Perception and Production of Mandarin Initial Stops by Native Urdu Speakers
Dan Du, Xianjin Zhu, Jinsong Zhang 0001 |
INTERSPEECH | 1 |
| 2019 | Long-Term Span Traffic Prediction Model Based on STL Decomposition and LSTMabstractWith the increasing complexity of the network, the current network traffic has strong nonlinearity and burstiness. Therefore, the traditional traffic prediction model is no longer applicable. The neural network model, especially the LSTM, can well fit the nonlinearity of time-series data and preserve the information memory of the past. However, as for the periodicity of long-term span network traffic data, the neural network model does not perform well. Based on this, this paper proposes LTS-TP (Long-Term Span Traffic Prediction model), a network traffic prediction model, to solve the problem. First, the model decomposes the collected network traffic data using the improved STL decomposition algorithm to preserve the seasonal component. Then, the trend component and the remainder component are input into the Seq2Seq model based on the LSTM added with the improved attention mechanism for prediction. Finally, the predicted value of the output is added to the seasonal component, and the final network traffic prediction value is obtained. In the simulation part, this paper uses the MAWI public data set to test the proposed network traffic prediction model and compared performance with other models. The results show that the network traffic prediction model proposed in this paper has a good predictive effect on long-term span network traffic data. Yonghua Huo, Dan Du, Yang Yang 0006 |
APNOMS | 3 |
| 2019 | Identifying Truly Suspicious Events and False Alarms Based on Alert GraphabstractAs a cyber security protection technology, Intrusion Detection System (IDS), through real-time monitoring, issues alerts when detecting malicious events. It is one of the most widely used network security products, yet still has high false positive rates. False positive alerts will not only waste a lot of resources and time to process, but also have bad effects on the correlation analysis and attack path detection. Therefore, reducing the false positives rate is one of the important means to improve the performance of IDS. In this paper, we propose an effective model for false positives identification using gradient boosting tree models based on the analysis of security features of the IDS alerts. Firstly, we analyze alarms from aggregation and correlation by constructing a correlated alert graph based on IP addresses. Secondly, we design a novel bidirectional recursive feature elimination method combining with random forest for feature selection. Finally, the ensemble methods are employed from boosting tree models in our approach for better improvement. Zhigang Lu 0002, Dan Du, Yaopeng Han |
IEEE BigData | 4 |
| 2019 | The Production of Chinese Affricates /ts/ and /tsh/ by Native Urdu Speakers
Dan Du, Jinsong Zhang 0001 |
INTERSPEECH | 1 |
| 2017 | Stealthy Domain Generation AlgorithmsabstractBotnets are groups of compromised computers that botmasters (botherders) use to launch attacks over the Internet. To avoid detection, botnets use DNS fast flux to change the mapping between IP addresses and domain names periodically. Domain generation algorithms (DGAs) are employed to generate a large number of domain names. Detection techniques have been proposed to identify malicious domain names generated by DGAs. Three metrics, Kullback-Leibler (KL) distance, Edit distance (ED), and Jaccard index (JI), are used to detect botnet domains with up to 100% detection rate and 2.5% false-positive rate. In this paper, we propose two DGAs that use hidden Markov models (HMMs) and probabilistic context-free grammars (PCFGs), respectively. Experiment results show that DGA detection metrics (KL, JI, and ED) and detection systems (BotDigger and Pleiades) have difficulty detecting domain names generated using the proposed approaches. Game theory is used to optimize strategies for both botmasters and security personnel. Results show that, to optimize DGA detection, security personnel should use the ED detection technique with probability 0.78 and JI detection with probability 0.22, and botmasters should choose the HMM-based DGA with probability 0.67 and PCFG-based DGA with probability 0.33. Yu Fu 0005, Lu Yu 0001, Oluwakemi Hambolu, Ilker Özçelik, Benafsh Husain, Jingxuan Sun, Karan Sapra, Dan Du, Christopher Tate Beasley, Richard R. Brooks |
IEEE Trans. Inf. Forensics Secur. | 8 |