Kun Zhu 0024

dblp:230/9968 · DBLP profile ↗
← Back
24ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0002-5773-5089ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Targeting Borderline Fraudsters: Multi-View Hypergraph Fraud Detection with LLM-Guided Contrastive Learning
abstract
Graph fraud detection (GFD) on transaction networks is crucial for safeguarding financial systems. However, due to the limited perspective of existing graph neural networks (GNNs) in the single transaction view, sophisticated fraudsters can disguise themselves to exhibit weak fraud signals, appearing as borderline fraudsters. To address this challenge, we propose MH-LGC, a multi-view hypergraph fraud detection model with large language model (LLM) guided contrastive learning. MH-LGC tackles two key limitations of existing GNN-based GFD methods: (1) Due to the local aggregation mechanism, existing methods struggle to capture high-order trading patterns among distant fraudsters. MH-LGC introduces two temporal hyper-views as complements to the transaction view and employs a Temporal Hypergraph Attention Network (THAN) to integrate the three views. (2) Most GFD methods overlook the rich semantic cues embedded in transaction data. Although some general graph learning studies have explored LLM integration, the high computational overhead and task-specific fine-tuning make them impractical for GFD tasks. MH-LGC introduces a semantic view through a fine-tuning-free LLM-Guided Contrastive learning (LGC), adopting a novel paradigm for integrating GNN and LLM to reduce the computational overhead of LLM. Extensive experiments on three real-world datasets demonstrate that MH-LGC outperforms twelve state-of-the-art baselines, with AUC improvements ranging from 1.10% to 5.70%.
Rui Ou, Kun Zhu 0024, Jiangtong Li, Chaochao Chen 0001, Yuhua Xu 0011, Changjun Jiang 0002
AAAI2
2026 STG-GNN: Semi-supervised Temporal Graph Learning Based on Group Strategy Against Credit Card Fraud
Rongkun Cui, Kun Zhu 0024
ICIC (7)2
2026 Bridging Cognitive Neuroscience and Graph Intelligence: Hippocampus-Inspired Multi-View Hypergraph Learning for Web Finance Fraud
abstract
Online financial services constitute an essential component of contemporary web ecosystems, yet their openness introduces substantial exposure to fraud that harms vulnerable users and weakens trust in digital finance. Such threats have become a significant web harm that erodes societal fairness and affects the well-being of online communities. However, existing detection methods based on graph neural networks (GNNs) struggle with two persistent challenges: (1) long-tailed data distributions, which obscure rare but critical fraudulent cases, and (2) fraud camouflage, where malicious transactions mimic benign behaviors to evade detection. To fill these gaps, we propose HIMVH, a Hippocampus-Inspired Multi-View Hypergraph learning model for web finance fraud detection. Specifically, drawing inspiration from the scene conflict monitoring role of the hippocampus, we design a cross-view inconsistency perception module that captures subtle discrepancies and behavioral heterogeneity across multiple transaction views. This module enables the model to identify subtle cross-view conflicts for detecting online camouflaged fraudulent behaviors. Furthermore, inspired by the match-mismatch novelty detection mechanism of the CA1 region, we introduce a novelty-aware hypergraph learning module that measures feature deviations from neighborhood expectations and adaptively reweights messages, thereby enhancing sensitivity to online rare fraud patterns in the long-tailed settings. Extensive experiments on six web-based financial fraud datasets demonstrate that HIMVH achieves 6.42% improvement in AUC, 9.74% in F1 and 39.14% in AP on average over 15 SOTA models.
Rongkun Cui, Kun Zhu 0024, Qi Zhang 0020
WWW3
2026 STG-DGR: Fraud Detection on Streaming Transaction Graphs with Diffusion-based Generative Replay
abstract
Fraud detection on streaming transaction graphs (STGs) faces challenges on the catastrophic forgetting of previously learned fraud patterns when adapting to evolving patterns. Although some Graph Continual Learning (GCL) approaches mitigate this issue by storing and revisiting historical samples, practical storage constraints prevent them from fully preserving previous patterns. In this work, we propose STG-DGR, a streaming GNN model with diffusion-based generative replay that generates synthetic samples to retain previously learned patterns without storing real samples. The generation of replay samples for STGs faces two key challenges: (1) Heterogeneity challenge of generating STG samples with discrete adjacency table, user features, transaction features, and transaction timestamps. (2) Dependency challenge of capturing bottom-up dependencies across layers in STG samples. To address these challenges, STG-DGR integrates two novel components: (1) a Computational Subgraph Processor (CSP) that transforms heterogeneous STG samples into well-organized hierarchical subgraphs, and (2) a Diffusion-based Subgraph Generator (DSG) that captures the bottom-up dependencies using a novel Transformer-based Hierarchical Denoising Network (THDN), and generates synthetic replay samples that preserve these dependencies. Extensive experiments on four streaming fraud detection datasets demonstrate STG-DGR's superiority in reducing forgetting and improving accuracy over nineteen state-of-the-art baselines.
Rui Ou, Kun Zhu 0024, Jiangtong Li, Chaochao Chen 0001, Changjun Jiang 0002
WWW2
2026 Dynamic Min-Max Multi-Dimensional Reinforcement Backdoor Attacks and Orchestrated Closed-Loop Defense in Fairness-Aware Web Federated Finance
abstract
In the rapidly evolving web-based financial ecosystem where digital banking services become critical infrastructure for underserved communities, credit card fraud disproportionately affects vulnerable populations relying on financial platforms. However, previous studies overlook extreme data scarcity conditions, particularly at small-to-medium web banks that serve as crucial gateways for vulnerable communities. This paper addresses the fundamental challenge of building inclusive and secure financial systems operable at true web scale. To overcome this deficiency, we propose a novel web-based fairness-aware federated fraud detection model, CLARF, which utilizes the designed privacy-enhanced representation fusion and fraud-aware contrastive learning modules to enhance detection performance under conditions of data scarcity and label imbalance. Furthermore, current federated fraud detection systems critically neglect vulnerability to backdoor attacks, where malicious actors can implant hidden triggers during model aggregation, compromising system integrity. We propose a novel dynamic web Min-Max adversarial game framework where attackers employ hybrid multi-stage reinforcement learning with multi-dimensional reward mechanisms to dynamically evolve triggers that achieve excellent tradeoff between stealthiness and effectiveness. Defender adapts a closed-loop Selection-Evaluation-Suppression framework where high-reliability clients are selected via Fisher information to carry out reverse trigger engineering. Then clients' confidence scores are calculated as weights to minimize Attack Success Rate (ASR) during aggregation. Extensive experiments on six financial fraud datasets demonstrate the superiority of CLARF model and Min-Max adversarial game paradigm compared with multiple SOTA models.
Ruixiao Zhu, Kun Zhu 0024, Qi Zhang 0020, Changjun Jiang 0002
WWW2
2026 Beyond catastrophic forgetting: A continual learning-driven multi-modal fusion model for saliency prediction in dynamic scenes
Jiaqi Wang 0003, Dandan Zhu 0001, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai
Expert Syst. Appl.4
2026 Developing Evolving Adaptability in Biological Intelligence: A Novel Biologically-Inspired Continual Learning Model for Video Saliency Prediction
abstract
In the era of deep learning, video saliency prediction task still remains major challenge due to the issue of catastrophic forgetting during feature learning. Most prior works commonly employ generative replay strategies to generate pseudo-samples from previous tasks, enabling them to recall the data distribution. However, scaling up generative replay to accommodate class-incremental and task-incremental settings poses challenges, as generated data with low quality can severely deteriorate performance. Additionally, existing advances mainly focus on preserving memory stability to alleviate catastrophic forgetting, but they remain difficult to flexibly adapt to incremental changes in dynamic scenes. To achieve a better balance between memory stability and learning plasticity, we propose a novel biologically-inspired continual learning (BICL) model tailored to effectively predict human attention in dynamic scenes while mitigate catastrophic forgetting. In particular, inspired by the function of the hippocampus in the human neural system, we elaborately design a visual saliency memory bank module to explicitly store and retrieve representative features from previous tasks. Furthermore, drawing inspiration from the Drosophila $\gamma$γMB system, we propose an active forgetting strategy equipped with multiple parallel adaptive learner modules, which can appropriately attenuate old memories in parameter distribution to enhance learning plasticity to adapt to new tasks, and accordingly to ensure compatibility among multiple learners. Notably, without compromising the performance of old tasks, our proposed model can achieve a better trade-off between memory stability and learning plasticity. Through extensive experiments on several benchmark datasets, our model not only enhances performance in task-incremental settings, but also potentially provides deep insights into neurological adaptive mechanisms.
Dandan Zhu 0001, Kaiwei Zhang, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 HFDFM: A Heterogeneous Credit Card Fraud Detection Model Based on Federated Learning With Membership Privacy
abstract
Credit card fraud brings serious losses to both cardholders and card issuers. To reduce losses caused by fraudulent behaviors, banking institutions establish credit card fraud detection (CCFD) models to identify potential fraudulent behaviors. To develop more effective fraud detection models, banking institutions need to collaborate on model training. Federated learning (FL) enables collaboration on fraud detection model training without exchanging data between banking institutions. Nevertheless, the distribution of transaction data in the real-world is heterogeneous among banking institutions, which may lead to convergence issues in global fraud detection models. Furthermore, the behavior and weights of the model may implicitly contain the cardholders’ personal information, which makes existing federated models prone to transaction data leakage. In this article, we propose aheterogeneous credit cardfrauddetection model based onfederated learning withmembership privacy, called HFDFM. To protect the sensitive information of cardholders, we design a novel mechanism that ensures training-data confidentiality by minimizing the accuracy of the best black-box membership inference attack (MIA) against the model. Unlike previous FL frameworks that either optimize for client drift or for membership privacy, HFDFM simultaneously mitigates drift and provides certified membership privacy through a single min–max game. Additionally, we use the control variable to rectify the client drift in its local update in the scenario of heterogeneous data. Extensive experimental results on three mainstream real-world transaction datasets demonstrate that the proposed HFDFM has advantages in utility compared to 11 SOTA baselines, and the proposed HFDFM can mitigate the risks of MIAs (near random guess).
Jun Wu 0006, Kun Zhu 0024, Rongkun Cui, Zhe Liu 0001, Changjun Jiang 0002
IEEE Trans. Comput. Soc. Syst.2
2026 MEScan360: A Memory-Enhanced Scanpath Prediction Model for Omnidirectional Images
Dandan Zhu 0001, Kaiwei Zhang, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai
ACM Trans. Multim. Comput. Commun. Appl.4
2026 CamFD: Semi-Supervised Camouflage-Aware Fraud Detection Based on Dynamic Graphs
abstract
Fraud detection on dynamic graph (FDDG) is in high demand across many real-world applications, such as financial transaction networks and social networks. A graph neural network (GNN) is an advanced methodology for learning graph data and has been widely adopted in fraud detection tasks. Despite the promising results, research on GNN-based fraud detection faces two critical challenges. First, due to the lack of temporal consistency mining in existing methods, camouflaged fraudsters can easily evade detection by imitating the behavioral patterns of benign users. Second, with constantly evolving fraud patterns and a scarcity of fraud samples, existing methods can easily overfit to current fraud patterns and struggle to identify the evolved and new ones. To address these challenges, we propose CamFD, a semi-supervised camouflage-aware fraud detection model on dynamic graphs. To detect the temporal inconsistency in the behavioral patterns of camouflaged fraudsters, CamFD is equipped with a novel consistency-sensitive contrastive learning (CSCL) module. CSCL discriminatively learns the consistency and inconsistency in users’ behavioral patterns by constructing temporal and structural contrastive pairs. Since modeling the evolving fraud patterns with sparse fraud samples is almost impractical, CamFD concentrates on modeling the relatively stable patterns of extensive benign users with a multivariate Gaussian distribution modeling (MGDM) module. We conduct extensive experiments on a private credit card fraud dataset as well as three public datasets. The experimental results demonstrate that CamFD outperforms ten state-of-the-art (SOTA) baselines across all datasets, with the area under ROC curve (AUC) improvements ranging from 1.03% to 5.32%.
Rui Ou, Kun Zhu 0024, MengChu Zhou, Changjun Jiang 0002
IEEE Trans. Syst. Man Cybern. Syst.2
2025 FDFRL: Credit Card Fraud Detection Based on Federated Reinforcement Learning
Kun Zhu 0024, Dandan Zhu 0001
ICANN (4)3
2025 LD2Scan: A Lightweight Dual-Temporal Constrained Scanpath Prediction Model for Omnidirectional Images
abstract
Predicting scanpaths in omnidirectional images (ODIs) is essential for simulating human gaze behaviors. However, current methods often struggle with long-term dependencies and exhibit high complexity, which limits their efficiency and scalability. To tackle these challenges, we propose LD2Scan, a lightweight diffusion-based model specifically designed for scanpath prediction in ODIs. It employs Efficient Equivariant (E4) convolution to enhance feature extraction from distorted ODIs while improving computational performance, thereby reducing resource demands. LD2Scan utilizes a dual-graph convolutional network (GCN) to enforce internal time constraints between fixations, integrating semantic-level GCN for sequential fixation modeling and image-level GCN to capture relationships across different images, enriching contextual information. We formulate the scanpath prediction issue as a conditional generation task, refining noisy scanpaths using features encoded by the dual-GCN and robust E4-processed features. Experimental results on several benchmark datasets demonstrate that LD2Scan outperforms existing methods in terms of both accuracy and efficiency.
Dandan Zhu 0001, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai
ICME4
2025 Elevating Mesh Saliency in VR: Introducing a Novel Prediction Network and Dataset
abstract
In computer graphics, polygon meshes stand out as a popular representation providing effective delineation of delicate textures and complex geometries. When dealing with geometric processing tasks for critical regions of the mesh, it is necessary to consider the human visual perception related to saliency. Therefore, we establish a novel mesh saliency dataset, facilitated by a more comprehensive gathering pipeline of eye-tracking from subjects observing mesh models at arbitrary viewpoints in a virtual reality space with six degrees of freedom. Additionally, we propose a mesh saliency prediction model that accurately infers visual attention density maps for complex and irregular mesh surfaces. This model integrates surface curvature and triangular face shape information from multi-scale neighboring ranges as local geometric features, while also leveraging surface spatial positioning as a global feature. Our work aims to preserve critical areas and minimize visual loss in saliency-driven tasks such as mesh simplification, rendering, and texturing. We believe that our research can offer valuable insights for human-centered mesh computation applications.
Kaiwei Zhang, Mohan He, Dandan Zhu 0001, Kun Zhu 0024, Xiongkuo Min, Guangtao Zhai
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Audio-Visual Saliency Prediction Model with Implicit Neural Representation
abstract
With the remarkable advancement of deep learning techniques and the wide availability of large-scale datasets, the performance of audio-visual saliency prediction has been drastically improved. Actually, audio-visual saliency prediction is still at an early exploration stage due to the spatial-temporal signal complexity and dynamic continuity of video content. To our knowledge, most existing audio-visual saliency prediction approaches usually represent videos as 3D grid of RGB values using discrete convolutional neural networks (CNNs), which inevitably incurs video content-agnostic and ignores the dynamic continuity issues. This article proposes a novel parametric audio-visual saliency (PAVS) model with implicit neural representation (INR) to address the aforementioned problems. Specifically, by using the proposed parametric neural network, we can effectively encode the space-time coordinates of video frames into corresponding saliency values, which can significantly enhance the compact feature representation ability. Meanwhile, a parametric feature fusion method is developed to achieve intrinsic interactions between audio and visual information streams, which can adaptively fuse audio and visual features to obtain competitive performance. Notably, without resorting to any specific audio-visual feature fusion strategy, the proposed PAVS model outperforms other state-of-the-art saliency methods by a large margin.
Dandan Zhu 0001, Kun Zhu 0024, Guangtao Zhai, Xiaokang Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 E2DAS: An Efficient Equivariant Dynamic Aggregation Saliency Model for Omnidirectional Images
Dandan Zhu 0001, Kun Zhu 0024, Guangtao Zhai, Xiaokang Yang 0001
ICPR (3)4
2024 TDiffSal: Text-Guided Diffusion Saliency Prediction Model for Images
Dandan Zhu 0001, Kun Zhu 0024, Guangtao Zhai
ICPR (8)4
2024 IAPCP: An Effective Cross-Project Defect Prediction Model via Intra-Domain Alignment and Programming-Based Distribution Adaptation
abstract
Cross‐project defect prediction (CPDP) aims to identify defect‐prone software instances in one project (target) using historical data collected from other software projects (source), which can help maintainers allocate limited testing resources reasonably. Unfortunately, the feature distribution discrepancy between the source and target projects makes it challenging to transfer the matching feature representation and severely hinders CPDP performance. Besides, existing CPDP models require an intensively expensive and time‐consuming process to tune a lot of parameters. To address the above limitations, we propose an effective CPDP model named IAPCP based on distribution adaptation in this study, which consists of two stages: correlation alignment and intra‐domain programming. Correlation alignment first calculates the covariance matrices of the source and target projects and then erases some features of the source project (i.e., whitening operation) and employs the features of the target project (i.e., target covariance) to fill the source project, thereby well aligning the source and target feature distributions and reducing the distribution discrepancy across projects. Intra‐domain programming can directly learn a nonparametric linear transfer defect predictor with strong discriminative capacity by solving a probabilistic annotation matrix (PAM) based on the adjusted features of the source project. The model does not require model selection and parameter tuning. Extensive experiments on a total of 82 cross‐project pairs from 16 software projects demonstrate that IAPCP can achieve competitive CPDP effectiveness and efficiency compared with multiple state‐of‐the‐art baseline models.
Kun Zhu 0024, Dandan Zhu 0001
IET Softw.2
2024 WSBCV: A data-driven cross-version defect model via multi-objective optimization and incremental representation learning
Kun Zhu 0024, Weiping Ding 0001, Dandan Zhu 0001
Inf. Sci.2
2024 IMDAC: A robust intelligent software defect prediction model via multi-objective optimization and end-to-end hybrid deep learning networks
abstract
Abstract Software defect prediction (SDP) aims to build an effective prediction model for historical defect data from software repositories by some specialized techniques or algorithms, and predict the defect proneness of new software modules. Nevertheless, the complex internal intrinsic structure hidden behind the defect data makes it challenging for the built prediction model to capture the most expressive defect feature representations, and largely limits the SDP performance. Fortunately, artificial intelligence is interacting closely with humans and provides powerful intelligent technical support for addressing these SDP issues. In this article, we propose a robust intelligent SDP model called IMDAC based on deep learning and soft computing techniques. This model has three main advantages: (1) an effective deep generative network—InfoGAN (information maximizing GANs) is employed to conduct data augmentation, namely generating sufficient defect instances and achieving defect class balance simultaneously. (2) Select the fewest representative feature subset for the minimum error via an advanced multi‐objective optimization approach—MSEA (multi‐stage evolutionary algorithm). (3) Build a powerful end‐to‐end deep defect predictor by hybrid deep learning techniques—DAE (Denoising AutoEncoder) and CNN (convolutional neural network), which can not only reconstruct a clean “repaired” input with strong robustness and generalization capabilities via DAE, but also learn the abstract deep semantic features with strong discriminating capability via CNN. Experimental results verify the superiority and robustness of the IMDAC model across 15 software projects.
Kun Zhu 0024, Changjun Jiang 0002, Dandan Zhu 0001
Softw. Pract. Exp.1
2022 Software defect prediction based on stacked sparse denoising autoencoders and enhanced extreme learning machine
abstract
Abstract Software defect prediction is an important software quality assurance technique. Nevertheless, the prediction performance of the constructed model is easily susceptible to irrelevant or redundant features in the software projects and is not predominant enough. To address these two issues, a novel defect prediction model called SSEPG based on Stacked Sparse Denoising AutoEncoders (SSDAE) and Extreme Learning Maching (ELM) optimised by Particle Swarm Optimisation (PSO) and another complementary Gravitational Search Algorithm (GSA) are proposed in this paper, which has two main merits: (1) employ a novel deep neural network – SSDAE to extract new combined features, which can effectively learn the robust deep semantic feature representation. (2) integrate strong exploitation capacity of PSO with strong exploration capability of GSA to optimise the input weights and hidden layer biases of ELM, and utilise the superior discriminability of the enhanced ELM to predict the defective modules. The SSDAE is compared with eleven state‐of‐the‐art feature extraction methods in effect and efficiency, and the SSEPG model is compared with multiple baseline models that contain five classic defect predictors and three variants across 24 software defect projects. The experimental results exhibit the superiority of the SSDAE and the SSEPG on six evaluation metrics.
Shi Ying 0002, Kun Zhu 0024, Dandan Zhu 0001
IET Softw.3
2022 IVKMP: A robust data-driven heterogeneous defect model based on deep representation optimization learning
Kun Zhu 0024, Shi Ying 0002, Weiping Ding 0001, Dandan Zhu 0001
Inf. Sci.1
2021 WGNCS: A robust hybrid cross-version defect model via multi-objective optimization and deep enhanced feature representation
Shi Ying 0002, Weiping Ding 0001, Kun Zhu 0024, Dandan Zhu 0001
Inf. Sci.4
2021 Software defect prediction based on enhanced metaheuristic feature selection optimization and a hybrid deep neural network
Kun Zhu 0024, Shi Ying 0002, Dandan Zhu 0001
J. Syst. Softw.1
2020 Within-project and cross-project just-in-time defect prediction based on denoising autoencoder and convolutional neural network
abstract
Just‐in‐time defect prediction is an important and useful branch in software defect prediction. At present, deep learning is a research hotspot in the field of artificial intelligence, which can combine basic defect features into deep semantic features and make up for the shortcomings of machine learning algorithms. However, the mainstream deep learning techniques have not been applied yet in just‐in‐time defect prediction. Therefore, the authors propose a novel just‐in‐time defect prediction model named DAECNN‐JDP based on denoising autoencoder and convolutional neural network in this study, which has three main advantages: (i) Different weights for the position vector of each dimension feature are set, which can be automatically trained by adaptive trainable vector. (ii) Through the training of denoising autoencoder, the input features that are not contaminated by noise can be obtained, thus learning more robust feature representation. (iii) The authors leverage a powerful representation‐learning technique, convolution neural network, to construct the basic change features into the abstract deep semantic features. To evaluate the performance of the DAECNN‐JDP model, they conduct extensive within‐project and cross‐project defect prediction experiments on six large open source projects. The experimental results demonstrate that the superiority of DAECNN‐JDP on five evaluation metrics.
Kun Zhu 0024, Shi Ying 0002, Dandan Zhu 0001
IET Softw.1