Jieren Cheng

dblp:46/4396 · DBLP profile ↗
← Back
44ranked-venue papers
11as first author
35since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Security and privacy · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Federated Graph-level Clustering Network with Attribute Inference
abstract
With the rise of vertical segmentation in real-world data, federated graph-level clustering has gained significant attention in recent years. However, the inherent missing attributes in graph datasets held by certain clients lead to suboptimal local parameter updates and misaligned global parameter consensus. This results in knowledge shifts during negotiation to ultimately impair overall clustering performance. This issue remains largely underexplored in the current advanced research. To bridge this gap, we propose a novel deep learning network called Federated Graph-level Clustering Network with Attribute Inference (FedAI), which utilizes high-confidence prior knowledge from each domain and multi-party collaborative optimization to achieve efficient reasoning of unknown features. Specifically, on the client, high-confidence graph samples are projected into a latent space. We then extract and upload irreversible path digest information and attribute-oriented inference signals from them. On the server, we first identify affinity relationships hierarchically via the improved graph kernel method. We then infer the features of clients lacking node attributes through a prior structure-guide recovery operator, facilitating inter-client knowledge transfer for better clustering. Experimental results on 15 cross-dataset and cross-domain non-IID graph datasets demonstrate that FedAI consistently outperforms existing methods.
Renda Han, Wenxuan Tu, Jingxin Liu 0006, Jieren Cheng
AAAI6
2026 Cross-View Progressive Feature Filtering for Multi-View Graph Clustering in Remote Sensing
abstract
Multi-view clustering of remote sensing data plays a vital role in Earth observation analysis. Recently, deep graph clustering methods based on contrastive learning have significantly improved feature representation capabilities. However, most existing approaches treat all views equally, neglecting the inherent uniqueness and heterogeneity across views, which often results in two major issues: 1) discriminative features from clustering-friendly views are underexplored; and 2) redundant or noisy information from less informative views can degrade the shared representation. To address these challenges, we propose a novel multi-view graph clustering framework termed CF-MVGC for remote sensing data, which dynamically preserves discriminative features and suppresses redundancy by assessing view affinity. Specifically, we employ a dual-stage representation learning strategy to extract both view-specific discriminative features and cross-view consistent representations. To further exploit and adaptively integrate complementary information across views, we design a progressive feature filtering model that dynamically evaluates view affinity using two novel metrics, i.e., view fidelity index (VFI) and view criticality index (VCI). Based on these assessments, the module adaptively modulates feature update and reset signals, reinforcing informative views while suppressing noisy or redundant ones. Views with high affinity receive strengthened update signals to retain valuable features, while those with low affinity are subjected to enhanced reset operations to eliminate noise and redundancy. The resulting high-quality, discriminative representations lead to improved clustering performance, establishing a positive feedback loop. Experimental results on four benchmark datasets demonstrate the effectiveness and superiority of CF-MVGC against its competitors.
Bowen Liu 0020, Xin Peng 0010, Wenxuan Tu, Chengyao Wei, Xiangyan Tang, Jieren Cheng
AAAI6
2026 Anchor-Driven Nyström for Deep Graph-Level Clustering
abstract
Graph-level clustering (GLC), which aims to group entire graphs according to their structural and attribute-based similarities, represents a fundamental yet challenging task in various practical applications. Existing GLC methods primarily fall into two main paradigms: 1) deep graph clustering approaches based on Graph Neural Networks (GNNs), and 2) kernel-based methods that utilize predefined kernels to perform fine-grained structural comparison for clustering. However, GNN-based methods typically learn graph-level representations by aggregating node embeddings through pooling operations, which inevitably leads to substantial information loss and suboptimal clustering performance. In contrast, kernel methods, despite their theoretical expressiveness, suffer from prohibitive computational costs that hinder their scalability to large-scale settings. To solve these issues, we propose a novel graph learning framework named Anchor-driven Nyström for Deep Graph-Level Clustering (ANGC), which computes graph similarity via kernel methods while retaining the scalability of GNNs. Specifically, we first employ GNNs to encode individual graphs into sets of node embeddings. Rather than relying on pooling operations, we compute graph similarities in a kernel space constructed from these embeddings. To enhance both scalability and representational power, we introduce learnable graph Nyström anchors, which support end-to-end optimization and significantly accelerate kernel computations. To further improve the discriminative capability of these anchors, we propose the concept of anchor response discrepancy, that is, the variation in a given anchor’s responses across different samples. By maximizing this discrepancy, the anchors are encouraged to strengthen inter-graph distinctions for better clustering. Extensive experiments demonstrate the effectiveness and superiority of ANGC over existing state-of-the-art methods.
Wenxuan Tu, Lingren Wang, Jieren Cheng
AAAI4
2026 FedPKDA: Personalized Federated Learning with Privacy-Preserving Knowledge Dynamic Alignment
abstract
Personalized Federated Learning (PFL), which aims to customize models for each client while preserving data privacy, has become an important research topic in addressing the challenges of data heterogeneity. Existing studies usually enhance the localization of global parameters by injecting local information into the globally shared model. However, these methods focus excessively on the personalized characteristics of individual clients and fail to fully exploit distinctive information across clients, limiting the quality of local models to represent unseen samples well. To address this issue, we propose a novel personalized Federated Privacy-preserving Knowledge Dynamic Alignment (FedPKDA) framework, which ensures data privacy during both the collection of client-side key information and its incorporation into federated model training. Specifically, to ensure data privacy during the cross-client information collection phase, we first conduct feature clipping and add Laplacian noise to the local prototypes extracted from each client. Further, we compute the centroid of the uploaded local prototypes in a latent space and leverage Mahalanobis distance to guide the generation of global prototypes, thereby preserving the semantic contributions from participating clients. Moreover, to boost the personalization of the local model, we dynamically align representations learned by the shared model with both a set of local prototypes and privacy-preserving global prototypes, facilitating effective cross-client knowledge sharing under heterogeneous settings while preserving client-specific characteristics. Extensive experiments on benchmark datasets have verified the superiority of FedPKDA against its competitors.
Moxuan Zeng, Wenxuan Tu, Yiying Wang, Xiangyan Tang, Jieren Cheng
AAAI7
2026 FedCND: Federated Graph-Level Clustering under Inter-Client Cluster Number Discrepancy
abstract
Federated graph-level clustering (FGC) provides an effective solution for analyzing decentralized graph data with privacy protection. Existing methods typically assume that all clients have the same number of clusters. This assumption simplifies the learning task and has achieved preliminary success. However, this assumption rarely holds in practice, as clients often exhibit substantial heterogeneity in both data distributions and semantic granularity. As a result, cluster-specific knowledge becomes misaligned during server-side aggregation, which ultimately degrades the overall clustering performance. To address this challenge, we propose a novel Federated Graph Clustering under Inter-Client Cluster Number Discrepancy (FedCND) framework, which aligns inter-client heterogeneous distributions by decoupling graph data into public and private patterns. Specifically, after initial local training and clustering on each client, we design a public learner and a private learner to model public and private graph data, respectively. Only anonymized, cluster-level public information is uploaded to the server, while private information remains local. On the server, cluster-level public prototypes are aggregated based on affinities between reconstructed cluster-level graphs, enabling privacy-preserving prototype alignment across clients with heterogeneous cluster numbers and mitigating interference from misaligned information during global aggregation. Finally, private subgraphs derive client-specific prototypes through local relearning, which are subsequently fused with globally oriented public prototypes for better clustering. Extensive experiments demonstrate that the proposed FedCND achieves an average of 4.9% accuracy improvement against current state-of-the-art methods.
Renda Han, Wenxuan Tu, Jingxin Liu 0006, Jieren Cheng
WWW6
2026 Pontis: A decentralized framework for unifying remote attestation and enabling interoperability between heterogeneous TEEs
abstract
In modern decentralized information systems, establishing verifiable trust in remote computing environments has emerged as a critical challenge for secure cross-domain collaboration. Hardware-based Trusted Execution Environments (TEEs) such as Intel SGX and ARM TrustZone offer a promising foundation for addressing this challenge through cryptographically verifiable execution guarantees, but their incompatible Remote Attestation (RA) mechanisms create fundamental barriers to cross-platform trust establishment. Current solutions either focus on single-vendor ecosystems or introduce prohibitive architectural complexity, failing to address the critical need for lightweight, decentralized interoperability. This paper presents Pontis, a decentralized blockchain-based framework that solves unified RA and cross-platform communication challenges for heterogeneous TEEs through three key innovations: (1) a distributed off-chain Coordinator that normalizes vendor-specific attestation protocols, combined with blockchain-anchored decentralized identifiers that provide immutable distributed identities for TEE instances; (2) blockchain-based smart contracts implementing Registration, Attestation, and Management functions for immutable trust status propagation; and (3) secure cross-TEE communication channels built on the Noise protocol framework. Pontis reduces attestation complexity from O ( n 2 ) → O ( n ), achieving up to 99% reduction in required attestations for large-scale deployments. Comprehensive evaluations on heterogeneous TEE platforms, including Intel SGX and ARM TrustZone, demonstrate that Pontis maintains an attestation latency of 60 ms with over 10,000 TEE instances, while establishing a trusted channel between heterogeneous TEEs requires only 5 ms. These results demonstrate the robustness and feasibility of Pontis, establishing a robust foundation for secure, scalable, and flexible cross-platform collaboration in demanding heterogeneous computing environments.
Hong Lei 0001, Xinman Luo, Pengxu Shen, Jieren Cheng
Inf. Process. Manag.7
2026 Bio-inspired night vision: A cat-eye mechanism for low-light image enhancement
Chengchao Wang 0003, Jihao Guo, Jieren Cheng, Wenxuan Tu
Knowl. Based Syst.6
2026 LiteRoute: Lightweight Latent Routing for Low-Light Enhancement
Jihao Guo, Jieren Cheng, Zeshi Chu, Ruilin Gao
IEEE Signal Process. Lett.2
2026 AGCN: A Service Computing-Oriented Anchor-Guided Graph Clustering Network
abstract
Deep graph clustering aims to uncover latent structural patterns within user behavior graphs to effectively group users based on behavioral similarity, thereby enhancing the efficiency of personalized service recommendation. However, most existing generative clustering methods heavily rely on encoder–decoder architectures and their pretraining procedures, which not only increase training complexity but also result in insufficient robustness and unstable optimization. To address these limitations, this paper proposes an Anchor-Guided Graph Clustering Network (AGCN) that eliminates the need for pretraining. The proposed method introduces an anchor-guided learning strategy that automatically selects high-confidence anchor samples farthest from the distribution boundary and utilizes the k-means++ algorithm to generate pseudo-labels in the diffusion feature space. Furthermore, AGCN integrates an anchor expansion strategy and a commonality enhancement mechanism to significantly improve clustering discriminability and recommendation accuracy. Theoretical analysis and extensive experiments on multiple real-world datasets demonstrate that the proposed method outperforms state-of-the-art approaches in terms of both clustering accuracy and stability.
Jieren Cheng, Faqiang Zeng, Jimei Li, Ling Liu 0001, Naixue Xiong, Xin Peng 0010
IEEE Trans. Serv. Comput.1
2025 Federated Graph-Level Clustering Network
abstract
Federated graph learning (FGL), which excels in analyzing non-IID graphs as well as protecting data privacy, has recently emerged as a hot topic. Existing FGL methods usually train the client model using labeled data and then collaboratively learn a global model without sharing their local graph data. However, in real-world scenarios, the lack of data annotations impedes the negotiation of multi-source information at the server, leading to sub-optimal feedback to the clients. To address this issue, we propose a novel unsupervised learning framework called Federated Graph-level Clustering Network (FedGCN), which collects the topology-oriented features of non-IID graphs from clients to generate global consensus representations through multi-source clustering structure sharing. Specifically, in the client, we first preserve the prototype features of each cluster from the structure-oriented embedding through clustering and then upload the learned multiple prototypes that are hard to be reconstructed into the raw graph data. In the server, we generate consensus prototypes from multiple condensed structure-oriented signals through Gaussian estimation, which are subsequently transferred to each client to promote the great encoding capacity of the local model for better clustering. Extensive experiments across multiple non-IID graph datasets have demonstrated the effectiveness and superiority of FedGCN against its competitors.
Jingxin Liu 0006, Jieren Cheng, Renda Han, Wenxuan Tu, Xin Peng 0010
AAAI2
2025 FedPKA: Federated Graph-Level Clustering Network with Personalized Knowledge Aggregation
Jingxin Liu 0006, Wenxuan Tu, Renda Han, Jieren Cheng, Xiangyan Tang
ICIC (16)6
2025 Federated Node-Level Clustering Network with Cross-Subgraph Link Mending
abstract
Subgraphs of a complete graph are usually distributed across multiple devices and can only be accessed locally because the raw data cannot be directly shared. However, existing node-level federated graph learning suffers from at least one of the following issues: 1) heavily relying on labeled graph samples that are difficult to obtain in real-world applications, and 2) partitioning a complete graph into several subgraphs inevitably causes missing links, leading to sub-optimal sample representations. To solve these issues, we propose a novel $\underline{\text{Fed}}$erated $\underline{\text{N}}$ode-level $\underline{\text{C}}$lustering $\underline{\text{N}}$etwork (FedNCN), which mends the destroyed cross-subgraph links using clustering prior knowledge. Specifically, within each client, we first design an MLP-based projector to implicitly preserve key clustering properties of a subgraph in a denoising learning-like manner, and then upload the resultant clustering signals that are hard to reconstruct for subsequent cross-subgraph links restoration. In the server, we maximize the potential affinity between subgraphs stemming from clustering signals by graph similarity estimation and minimize redundant links via the N-Cut criterion. Moreover, we employ a GNN-based generator to learn consensus prototypes from this mended graph, enabling the MLP-GNN joint-optimized learner to enhance data privacy during data transmission and further promote the local model for better clustering. Extensive experiments demonstrate the superiority of FedNCN.
Jingxin Liu 0006, Renda Han, Wenxuan Tu, Jieren Cheng
ICML6
2025 Dual Boost-Driven Graph-Level Clustering Network
abstract
Graph-level clustering remains a pivotal yet formidable challenge in graph learning. Recently, the integration of deep learning with representation learning has demonstrated notable advancements, yielding performance enhancements to a certain degree. However, existing methods suffer from at least one of the following issues: 1) the original graph structure has noise, and 2) during feature propagation and pooling processes, noise is gradually aggregated into the graph-level embeddings through information propagation. Consequently, these two limitations mask clustering-friendly information, leading to suboptimal graph-level clustering performance. To this end, we propose a novel Dual Boost-Driven Graph-Level Clustering Network (DBGCN) to alternately promote graph-level clustering and filtering out interference information in a unified framework. Specifically, in the pooling step, we evaluate the contribution of features at the global and optimize them using a learnable transformation matrix to obtain high-quality graph-level representation, such that the model’s reasoning capability can be improved. Moreover, to enable reliable graph-level clustering, we first identify and suppress information detrimental to clustering by evaluating similarities between graph-level representations, providing more accurate guidance for multi-view fusion. Extensive experiments demonstrated that DBGCN outperforms the state-of-the-art graph-level clustering methods on six benchmark datasets.
Renda Han, Wenxuan Tu, Wenxin Zhang 0005, Jingxin Liu 0006, Jieren Cheng, Huajie Lei, Guangzhen Yao, Lingren Wang, Yu Li 0047
IJCNN7
2025 Discovering Maximum Frequency Consensus: Lightweight Federated Learning for Medical Image Segmentation
Lingren Wang, Wenxuan Tu, Jieren Cheng, Xiangyan Tang
ACM Multimedia3
2025 Hierarchical Shortest-Path Graph Kernel Network
abstract
Graph kernels have emerged as a fundamental and widely adopted technique in graph machine learning. However, most existing graph kernel methods rely on fixed graph similarity estimation that cannot be directly optimized for task-specific objectives, leading to sub-optimal performance. To address this limitation, we propose a kernel-based learning framework called Hierarchical Shortest-Path Graph Kernel Network HSP-GKN, which seamlessly integrates graph similarity estimation with downstream tasks within a unified optimization framework. Specifically, we design a hierarchical shortest-path graph kernel that efficiently preserves both the semantic and structural information of a given graph by transforming it into hierarchical features used for subsequent neural network learning. Building upon this kernel, we develop a novel end-to-end learning framework that matches hierarchical graph features with learnable $hidden$ graph features to produce a similarity vector. This similarity vector subsequently serves as the graph embedding for end-to-end training, enabling the neural network to learn task-specific representations. Extensive experimental results demonstrate the effectiveness and superiority of the designed kernel and its corresponding learning framework compared to current competitors.
Wenxuan Tu, Jieren Cheng
NeurIPS3
2025 FedIGL: Federated Invariant Graph Learning for Non-IID Graphs
abstract
Federated Graph Learning (FGL) effectively facilitates cross-domain graph model training by enabling decentralized learning across multiple domains, while ensuring data privacy through local data storage and communication of model updates instead of raw data. Existing approaches usually assume shared generic knowledge (e.g., prototypes, spectral features) via aggregating local structures statistically to alleviate structural heterogeneity. However, imposing overly strict assumptions about the presumed correlation between structural features and the global objective often fails in generalizing to local tasks, leading to suboptimal performance. To tackle this issue, we propose a **Fed**erated **I**nvariant **G**raph **L**earning (**FedIGL**) framework based on invariant learning, which effectively disrupts spurious correlations and further mines the invariant factors across different distributions. Specifically, a server-side global model is trained to capture client-agnostic subgraph patterns shared across clients, whereas client-side models specialize in client-specific subgraph patterns. Subsequently, without compromising privacy, we propose a novel Bi-Gradient Regularization strategy that introduces gradient constraints to guide the model in identifying client-agnostic and client-specific subgraph patterns for better graph representations. Extensive experiments on graph-level clustering and classification tasks demonstrate the superiority of FedIGL against its competitors.
Lingren Wang, Wenxuan Tu, Jieren Cheng, Jingxin Liu 0006
NeurIPS5
2025 TGDCLNet: Teacher-Guided Denoising Contrastive Learning Network-Based IoT Network Intrusion Detection
abstract
With the rapid development of Internet of Things (IoT) technology, many devices are connecting to networks, making the security risks of IoT devices a growing focus of concern. The demand for enhancing the accuracy of IoT intrusion detection is becoming increasingly urgent. In the face of uneven traffic distribution and the problem of unknown attack types in the network flow data collected by IoT devices, traditional supervised learning models are not robust enough in practical applications. To address these issues, we propose an innovative Teacher-Guided Denoising Contrastive Learning Network (TGDCLNet), which employs a dual autoencoder structure for teacher and student models to perform network intrusion detection. In preprocessing, we innovatively design hybrid rebalancing strategies to adapt to different data distribution scenarios and select the best strategy combination. In the denoising learning module, we propose a denoising multi-dimensional balanced mean squared error loss function to reduce the impact of noise in the collected traffic on the training of both teacher and student models. Furthermore, we design a teacher guidance module, which takes label information as a feature input into the teacher autoencoder, and propose an equilibrium-weighted binary cross-entropy loss function to guide the learning process of the student model. The experimental results on six representative datasets show that our method outperforms other supervised contrastive learning algorithms. We also verify the robustness of our method in the complex and dynamic real-world IoT environment. Our method performs better in challenging environments with unknown attack types and scarce traffic data than other contrastive learning methods, providing a stable and effective solution for IoT intrusion detection.
Yue Yang 0047, Jieren Cheng, Renjie Wu 0011, Xiaoxv Tan, Zhaowu Liu, Haolan Yu, Xiangyan Su
IEEE Internet Things J.2
2025 Revisiting Initializing Then Refining: An Incomplete and Missing Graph Imputation Network
abstract
With the development of various applications, such as recommendation systems and social network analysis, graph data have been ubiquitous in the real world. However, graphs usually suffer from being absent during data collection due to copyright restrictions or privacy-protecting policies. The graph absence could be roughly grouped into attribute-incomplete and attribute-missing cases. Specifically, attribute-incomplete indicates that a portion of the attribute vectors of all nodes are incomplete, while attribute-missing indicates that all attribute vectors of partial nodes are missing. Although various graph imputation methods have been proposed, none of them is custom-designed for a common situation where both types of graph absence exist simultaneously. To fill this gap, we develop a novel graph imputation network termed revisiting initializing then refining (RITR), where both attribute-incomplete and attribute-missing samples are completed under the guidance of a novel initializing-then-refining imputation criterion. Specifically, to complete attribute-incomplete samples, we first initialize the incomplete attributes using Gaussian noise before network learning, and then introduce a structure-attribute consistency constraint to refine incomplete values by approximating a structure-attribute correlation matrix to a high-order structure matrix. To complete attribute-missing samples, we first adopt structure embeddings of attribute-missing samples as the embedding initialization, and then refine these initial values by adaptively aggregating the reliable information of attribute-incomplete samples according to a dynamic affinity structure. To the best of our knowledge, this newly designed method is the first end-to-end unsupervised framework dedicated to handling hybrid-absent graphs. Extensive experiments on six datasets have verified that our methods consistently outperform the existing state-of-the-art competitors. Our source code is available at https://github.com/WxTu/RITR.
Wenxuan Tu, Bin Xiao 0002, Xinwang Liu 0002, Sihang Zhou 0001, Zhiping Cai, Jieren Cheng
IEEE Trans. Neural Networks Learn. Syst.6
2025 Keep It Simple: Self-Adaptive Code Graph Simplification for Accurate Vulnerability Detection
abstract
Software vulnerability detection is crucial for high-quality software development. Recently, some studies utilizing Graph Neural Networks (GNNs) to learn the graph representation of code in vulnerability detection tasks have achieved remarkable success. However, existing graph-based approaches mainly face two limitations that prevent them from generalizing well to large code graphs: (1) the interference of noise information in the code graph; (2) the difficulty in capturing long-distance dependencies within the graph. To mitigate these problems, we propose a novel vulnerability detection method,ANGEL, whose novelty mainly embodies the hierarchical graph refinement and context-aware graph representation learning. The former hierarchically filters redundant information in the code graph, thereby reducing the size of the graph, while the latter collaboratively employs the Graph Transformer and GNN to learn code graph representations from both the global and local perspectives, thus capturing long-distance dependencies. Extensive experiments demonstrate promising results on three widely used benchmark datasets: our method significantly outperforms several other baselines in terms of the accuracy and F1 score. Particularly, in large code graphs,ANGELachieves an improvement in accuracy of 34.27%-161.93% compared to the state-of-the-art method, AMPLE. Such results demonstrate the effectiveness ofANGELin vulnerability detection tasks.
Xin Peng 0010, Shangwen Wang, Yihao Qin, Bo Lin 0011, Liqian Chen, Jieren Cheng, Xiaoguang Mao
IEEE Trans. Software Eng.6
2024 Attribute-Missing Graph Clustering Network
abstract
Deep clustering with attribute-missing graphs, where only a subset of nodes possesses complete attributes while those of others are missing, is an important yet challenging topic in various practical applications. It has become a prevalent learning paradigm in existing studies to perform data imputation first and subsequently conduct clustering using the imputed information. However, these ``two-stage" methods disconnect the clustering and imputation processes, preventing the model from effectively learning clustering-friendly graph embedding. Furthermore, they are not tailored for clustering tasks, leading to inferior clustering results. To solve these issues, we propose a novel Attribute-Missing Graph Clustering (AMGC) method to alternately promote clustering and imputation in a unified framework, where we iteratively produce the clustering-enhanced nearest neighbor information to conduct the data imputation process and utilize the imputed information to implicitly refine the clustering distribution through model optimization. Specifically, in the imputation step, we take the learned clustering information as imputation prompts to help each attribute-missing sample gather highly correlated features within its clusters for data completion, such that the intra-class compactness can be improved. Moreover, to support reliable clustering, we maximize inter-class separability by conducting cost-efficient dual non-contrastive learning over the imputed latent features, which in turn promotes greater graph encoding capability for clustering sub-network. Extensive experiments on five datasets have verified the superiority of AMGC against competitors.
Wenxuan Tu, Renxiang Guan, Sihang Zhou 0001, Chuan Ma 0001, Xin Peng 0010, Zhiping Cai, Zhe Liu 0001, Jieren Cheng, Xinwang Liu 0002
AAAI8
2024 TabSec: A Collaborative Framework for Novel Insider Threat Detection
abstract
In the era of the Internet of Things (IoT) and data sharing, users frequently upload their personal information to enterprise databases to enjoy enhanced service experiences provided by various online services. However, the widespread presence of system vulnerabilities, remote network intrusions, and insider threats significantly increases the exposure of private enterprise data on the internet. If such data is stolen or leaked by attackers, it can result in severe asset losses and business operation disruptions. To address these challenges, this paper proposes a novel threat detection framework, TabITD. This framework integrates Intrusion Detection Systems (IDS) with User and Entity Behavior Analytics (UEBA) strategies to form a collaborative detection system that bridges the gaps in existing systems’ capabilities. It effectively addresses the blurred boundaries between external and insider threats caused by the diversification of attack methods, thereby enhancing the model’s learning ability and overall detection performance. Moreover, the proposed method leverages the TabNet architecture, which employs a sparse attention feature selection mechanism that allows TabNet to select the most relevant features at each decision step, thereby improving the detection of rare-class attacks. We evaluated our proposed solution on two different datasets, achieving average accuracies of 96.71% and 97.25%, respectively. The results demonstrate that this approach can effectively detect malicious behaviors such as masquerade attacks and external threats, significantly enhancing network security defenses and the efficiency of network attack detection.
Xiangyan Tang, Xinyi Cao, Jieren Cheng, Wenxuan Tu, Logan Bo-Yee Liu
ISPA5
2024 Dual Contrastive Learning Network for Graph Clustering
abstract
Graph representation is an important part of graph clustering. Recently, contrastive learning, which maximizes the mutual information between augmented graph views that share the same semantics, has become a popular and powerful paradigm for graph representation. However, in the process of patch contrasting, existing literature tends to learn all features into similar variables, i.e., representation collapse, leading to less discriminative graph representations. To tackle this problem, we propose a novel self-supervised learning method called dual contrastive learning network (DCLN), which aims to reduce the redundant information of learned latent variables in a dual manner. Specifically, the dual curriculum contrastive module (DCCM) is proposed, which approximates the node similarity matrix and feature similarity matrix to a high-order adjacency matrix and an identity matrix, respectively. By doing this, the informative information in high-order neighbors could be well collected and preserved while the irrelevant redundant features among representations could be eliminated, hence improving the discriminative capacity of the graph representation. Moreover, to alleviate the problem of sample imbalance during the contrastive process, we design a curriculum learning strategy, which enables the network to simultaneously learn reliable information from two levels. Extensive experiments on six benchmark datasets have demonstrated the effectiveness and superiority of the proposed algorithm compared with state-of-the-art methods.
Xin Peng 0010, Jieren Cheng, Xiangyan Tang, Jingxin Liu 0006
IEEE Trans. Neural Networks Learn. Syst.2
2024 FedVeca: Federated Vectorized Averaging on Non-IID Data With Adaptive Bi-Directional Global Objective
abstract
Federated Learning (FL) is a distributed machine learning framework in parallel and distributed systems. However, the systems’ Non-Independent and Identically Distributed (Non-IID) data negatively affect the communication efficiency, since clients with different datasets may cause significant gaps to the local gradients in each communication round. In this article, we propose a Federated Vectorized Averaging (FedVeca) method to optimize the FL communication system on Non-IID data. Specifically, we set a novel objective for the global model which is related to the local gradients. The local gradient is defined as a bi-directional vector with step size and direction, where the step size is the number of local updates and the direction is divided into positive and negative according to our definition. In FedVeca, the direction is influenced by the step size, thus we average the bi-directional vectors to reduce the effect of different step sizes. Then, we theoretically analyze the relationship between the step sizes and the global objective, and obtain upper bounds on the step sizes per communication round. Based on the upper bounds, we design an algorithm for the server and the client to adaptively adjusts the step sizes that make the objective close to the optimum. Finally, we conduct experiments on different datasets, models and scenarios by building a prototype system, and the experimental results demonstrate the effectiveness and efficiency of the FedVeca method.
Ping Luo 0007, Jieren Cheng, Naixue Xiong, Zhenhao Liu, Jie Wu 0001
IEEE Trans. Parallel Distributed Syst.2
2024 HSNet: An Intelligent Hierarchical Semantic-Aware Network System for Real-Time Semantic Segmentation
abstract
Semantic segmentation, which aims to accurately identify each pixel, is a meaningful and challenging task. Recently, we witness a strong tendency to improve model efficiency in low-computing applications. However, most real-time methods ignore hierarchical features and context information to improve efficiency, leading to a decrease in the accuracy of semantic segmentation. To this end, we propose a novel system named hierarchical semantic-aware network (HSNet) to refine multilevel context information. HSNet mainly has the following two core modules: 1) hierarchical feature refinement module (HFRM) and 2) cross-scale pyramid fusion module (CPFM). By aggregating hierarchical feature maps, the proposed HFRM learns multilevel feature representation to recover spatial details. Afterward, the dual attention mechanism is developed to refine features from both channel and spatial levels, thereby alleviating the multilevel semantic gap. Meanwhile, the CPFM, which fuses local and global context information in a cross-scale manner, is proposed to enrich semantic information to improve accuracy. Furthermore, HSNet is carefully designed to improve the efficiency of the model by reusing shallow features and reducing channel capacity. Extensive experiments show that our method is effective and superior in segmentation accuracy and inference speed compared with state-of-the-art methods.
Xin Peng 0010, Jieren Cheng, Xiangyan Tang, Ziqi Deng, Wenxuan Tu, Naixue Xiong
IEEE Trans. Syst. Man Cybern. Syst.2
2023 H-MIS: A Hierarchical Multi-Identifier System Based on Blockchain
abstract
With its wide range of applications, the Internet shows a future trend towards abundant and diverse data resources with multiple types of identifiers (multi-identifiers). However, the legacy Domain Name System (DNS) in the current TCP/IP network architecture has failed to manage these identifiers due to the centralized security issue. While some decentralized DNS alternatives have been proposed, they also face scalability issues. In this paper, we propose a blockchain-based Hierarchical Multi-Identifier System, named H-MIS, as a DNS alternative. Specially, it realizes optimal decentralization and scalability by introducing the Zero-Knowledge rollup (ZK-rollup) solution to synchronize the upper and lower on-chain identifier data, as well as off-chain associated resource data. Finally, we implement H-MIS on Ethereum and evaluate its performance. The experimental results indicate that compared to the original MIS and Ethereum Name Service (ENS), H-MIS has advantages in such aspects as efficiency, data consumption, and Gas fees.
Qi Lyu, Hui Li 0022, Xinnan Lin, Han Wang 0022, Hanxu Hou, Yuguo Yin, Qianbin Chen, Selwyn Deng, Jieren Cheng
IEEE Big Data14
2022 Initializing Then Refining: A Simple Graph Attribute Imputation Network
abstract
Representation learning on the attribute-missing graphs, whose connection information is complete while the attribute information of some nodes is missing, is an important yet challenging task. To impute the missing attributes, existing methods isolate the learning processes of attribute and structure information embeddings, and force both resultant representations to align with a common in-discriminative normal distribution, leading to inaccurate imputation. To tackle these issues, we propose a novel graph-oriented imputation framework called initializing then refining (ITR), where we first employ the structure information for initial imputation, and then leverage observed attribute and structure information to adaptively refine the imputed latent variables. Specifically, we first adopt the structure embeddings of attribute-missing samples as the embedding initialization, and then refine these initial values by aggregating the reliable and informative embeddings of attribute-observed samples according to the affinity structure. Specially, in our refining process, the affinity structure is adaptively updated through iterations by calculating the sample-wise correlations upon the recomposed embeddings. Extensive experiments on four benchmark datasets verify the superiority of ITR against state-of-the-art methods.
Wenxuan Tu, Sihang Zhou 0001, Xinwang Liu 0002, Yue Liu 0008, Zhiping Cai, En Zhu, Changwang Zhang, Jieren Cheng
IJCAI8
2022 Foreground object structure transfer for unsupervised domain adaptation
abstract
Unsupervised domain adaptation aims to train a classification model from the labeled source domain for the unlabeled target domain. Since the data distribution of the two domains are different, the model often performs poorly on the target domain. The existing methods align the global features of the source domain and the target domain, and learn the domain invariant features to improve the performance of the model, which ignores the difference between the foreground features and the background features, and does not consider the structural information in the image foreground object. Therefore we proposed a method called foreground object structure transfer (FOST), it avoids the problem of ignoring differences in the structure information of foreground features and background features, exploits foreground feature enhancement from source-to-target transfer during adaptation and structural contrast loss to drive the domain alignment process. FOST relies on prior knowledge to distinguish foreground and background features, and considers the structural information of the object, which makes the intra-class spatial distribution more compact, the interclass spatial distribution more separated, improves the transferability and improves the classification efficiency. Extensive experimental results on various benchmarks under different domain adaptation settings illustrated that our FOST compares favorably against the state-of-the-art domain adaptation methods, we achieved the accuracies of 95.3%, 91.3%, 76.6%, and 87.55% on the ImageCLEF-DA, Office-31, Office-Home, and Visda-2017 data sets, respectively.
Jieren Cheng, Qiaobo Da
Int. J. Intell. Syst.1
2022 MIFNet: A lightweight multiscale information fusion network
abstract
Semantic segmentation technique plays a crucial role in Internet of Things applications, such as industrial robotics and self-driving. Recently deep learning approaches have boosted semantic segmentation accuracy greatly. However, their comprehensive performance in terms of accuracy and efficiency is still far from satisfactory. We observe that (1) accuracy-oriented methods rely on numerous convolution layers and sophisticated architectures, which result in heavy computational complexity and usually take a long time for inference; (2) efficiency-oriented methods fail to capture the multiscale context information for discriminative representations during the feature fusion process, thus leading to suboptimal performance. Previous semantic segmentation approaches fail to address these two challenges simultaneously. To tackle the dilemma of precise segmentation and efficient inference, we propose a novel lightweight Multiscale Information Fusion Network (MIFNet). Specifically, the proposed MIFNet mainly consists of two core components, that is, Pyramid Refinement Connection Module (PRCM) and Lightweight Information Fusion Module (LIFM). The PRCM exploits skip learning to establish dependency between different stages. Meanwhile, the pyramid attention mechanism (PAM) in PRCM, which adjusts the weight of hybrid pyramid attention vector to refine spatial features of low-level, is developed to alleviate the semantic gap. Moreover, the LIFM is designed to detect objects at multiple scales from the global-local perspective. In LIFM, the proposed multiscale dense concatenation (MDC) adopts various dilated convolution to extract multiscale local context information. Extensive experimental results on benchmarks data sets demonstrate the significantly better performance of the proposed MIFNet compared with most existing state-of-the-art methods.
Jieren Cheng, Xin Peng 0010, Xiangyan Tang, Wenxuan Tu, Wenhang Xu
Int. J. Intell. Syst.1
2022 An ensemble framework for interpretable malicious code detection
abstract
Malicious code is an ever-growing security threats to computer systems and networks, while malware detection provides effective defense against malicious codes. In this paper, a brief overview is presented on currently prevalent methods to detect malicious codes, including signature-based methods, behavioral-based detection and machine learning (ML) based ones. More specifically, the potentially effective malicious features are summarized and the novel methods using ML are deeply discussed. Furthermore, an ensemble interpretable framework is explored for automatic and efficient malicious code detection. Based on the knowledge graph of malware, the novel framework inclines to achieve robust malware detection even confronted with unseen malicious codes. Finally, both advantages and disadvantages are discussed and experimental results are outlined to verify the effectiveness of the novel methods.
Jieren Cheng, Jiachen Zheng, Xiaomei Yu
Int. J. Intell. Syst.1
2022 AAFL: Asynchronous-Adaptive Federated Learning in Edge-Based Wireless Communication Systems for Countering Communicable Infectious Diseasess
abstract
With the rapid growth of the coronavirus disease of 2019 (COVID-19) cases, massive amounts of relevant data are being trained on machine learning models for countering communicable infectious diseases. Federated Learning (FL) is a paradigm of distributed machine learning to deal with the individual COVID-19 data, and enable the protection of data privacy. However, FL has low efficiency in Edge-Based wireless communication systems with system heterogeneity. In this paper, we propose an “Asynchronous-Adaptive FL” (AAFL) scheme. Specifically, we allow that medical devices with different performances have a heterogeneous number of local SGD iterations in each communication round, called asynchronous iteration strategy which is balanced under adaptive control. We theoretically analyze the convergence of the AAFL scheme under a given time budget and obtain a mathematical relationship between the heterogeneous number of local SGD iterations and the optimal model parameters. Based on the mathematical relationship, we design an algorithm for parameter server and work nodes to adaptively control the heterogeneous number of local SGD iterations. Subsequently, we build a prototype heterogeneous system and conduct experiments on various scenarios for analyzing the general properties of our algorithm, and then apply our algorithm to public COVID-19 databases. The experimental results and application performance demonstrate the effectiveness and efficiency of our AAFL scheme.
Jieren Cheng, Ping Luo 0007, Naixue Xiong, Jie Wu 0001
IEEE J. Sel. Areas Commun.1
2022 Robust zero-watermarking algorithm for medical images based on SIFT and Bandelet-DCT
Yangxiu Fang, Jing Liu 0041, Jingbing Li, Jieren Cheng, Jiabin Hu, Dan Yi, Xiliang Xiao, Uzair Aslam Bhatti
Multim. Tools Appl.4
2021 Deep Fusion Clustering Network
abstract
Deep clustering is a fundamental yet challenging task for data analysis. Recently we witness a strong tendency of combining autoencoder and graph neural networks to exploit structure information for clustering performance enhancement. However, we observe that existing literature 1) lacks a dynamic fusion mechanism to selectively integrate and refine the information of graph structure and node attributes for consensus representation learning; 2) fails to extract information from both sides for robust target distribution (i.e., “groundtruth” soft labels) generation. To tackle the above issues, we propose a Deep Fusion Clustering Network (DFCN). Specifically, in our network, an interdependency learning-based Structure and Attribute Information Fusion (SAIF) module is proposed to explicitly merge the representations learned by an autoencoder and a graph autoencoder for consensus representation learning. Also, a reliable target distribution generation measure and a triplet self-supervision strategy, which facilitate cross-modality information exploitation, are designed for network training. Extensive experiments on six benchmark datasets have demonstrated that the proposed DFCN consistently outperforms the state-of-the-art deep clustering methods.
Wenxuan Tu, Sihang Zhou 0001, Xinwang Liu 0002, Xifeng Guo 0001, Zhiping Cai, En Zhu, Jieren Cheng
AAAI7
2021 Adaptive XACML access policies for heterogeneous distributed IoT environments
Khaled Riad, Jieren Cheng
Inf. Sci.2
2021 DFFNet: An IoT-perceptive dual feature fusion network for general real-time semantic segmentation
Xiangyan Tang, Wenxuan Tu, Keqiu Li, Jieren Cheng
Inf. Sci.4
2021 A survey of security threats and defense on Blockchain
Jieren Cheng, Luyi Xie, Xiangyan Tang, Naixue Xiong
Multim. Tools Appl.1
2020 ProbInfer: Probability-based AS path inference from multigraph perspective
Xionglve Li, Zhiping Cai, Bingnan Hou, Ning Liu 0015, Fang Liu 0002, Jieren Cheng
Comput. Networks6
2019 Contourlet-DCT based multiple robust watermarkings for medical images
Xiaoqi Wu, Jingbing Li, Rong Tu, Jieren Cheng, Uzair Aslam Bhatti, Jixin Ma 0001
Multim. Tools Appl.4
2019 Corrigendum to "Flow Correlation Degree Optimization Driven Random Forest for Detecting DDoS Attacks in Cloud Computing"
abstract
Correlation Degree Optimization Driven Random Forest for Detecting DDoS Attacks in Cloud Computing" [1], there was an error in the expression below formula (1) in Section 3.2, where theta "" symbol should be replaced with alpha "".Therefore the expression "() = () + (1 -)(), (0 < < 1)" should be corrected to be "() = () + (1 -)(), (0 < < 1)".
Jieren Cheng, Xiangyan Tang, Victor S. Sheng, Wei Guo 0011
Secur. Commun. Networks1
2018 A DDoS Detection Method for Socially Aware Networking Based on Forecasting Fusion Feature Sequence
abstract
Distributed Denial-of-Service (DDoS) is one of the most destructive network attacks. In Socially Aware Networking (SAN), there are many problems in current detection methods, such as low flexibility in detecting different attacks, high false-negative and false-positive rates. In this paper, we propose a DDoS detection method for SAN based on fusion feature series forecasting. Specifically, we define a multi-protocol-fusion feature (MPFF) to characterize normal network flows. Moreover, we utilize the time-series Autoregressive Integrated Moving Average Model (ARIMA) to formally describe the MPFF sequence, which is subsequently used in network flow forecasting and error calculation. Finally, we present the ARIMA detection model with error correction based on MPFF time series to identify DDoS in SAN. The experimental results show that the proposed method can effectively distinguish attacking flows from normal ones. Compared with previous DDoS detection methods for SAN, the proposed method can achieve better performance of detecting DDoS in terms of detection rate, false-positive rate and time delay.
Jieren Cheng, Jinghe Zhou, Qiang Liu 0004, Xiangyan Tang, Yanxiang Guo
Comput. J.1
2018 Flow Correlation Degree Optimization Driven Random Forest for Detecting DDoS Attacks in Cloud Computing
abstract
Distributed denial-of-service (DDoS) has caused major damage to cloud computing, and the false- and missing-alarm rates of existing DDoS attack-detection methods are relatively high in cloud environment. In this paper, we propose a DDoS attack-detection method with enhanced random forest (RF) optimized by genetic algorithm based on flow correlation degree (FCD) feature. We define the FCD feature according to the asymmetric and semidirectivity interaction characteristics and use the two-tuples FCD feature consisting of packet-statistical degree (PSD) and semidirectivity interaction abnormality (SDIA) to describe the features of attack flow and normal flow. Then we use a genetic algorithm based on the FCD feature sequences to optimize two key parameters of the decision tree in the RF: the maximum number of decision trees and the maximum depth of every single decision tree. We apply the trained RF model with optimized parameters to generate the classifier to be used for DDoS attack-detection. The experiment shows that the proposed method can effectively detect DDoS attacks in cloud environment with a higher accuracy rate and lower false- and missing-alarm rates compared to existing DDoS attack-detection methods.
Jieren Cheng, Xiangyan Tang, Victor S. Sheng, Wei Guo 0011
Secur. Commun. Networks1
2018 Adaptive DDoS Attack Detection Method Based on Multiple-Kernel Learning
abstract
Distributed denial of service (DDoS) attacks has caused huge economic losses to society. They have become one of the main threats to Internet security. Most of the current detection methods based on a single feature and fixed model parameters cannot effectively detect early DDoS attacks in cloud and big data environment. In this paper, an adaptive DDoS attack detection method (ADADM) based on multiple-kernel learning (MKL) is proposed. Based on the burstiness of DDoS attack flow, the distribution of addresses, and the interactivity of communication, we define five features to describe the network flow characteristic. Based on the ensemble learning framework, the weight of each dimension is adaptively adjusted by increasing the interclass mean with a gradient ascent and reducing the intraclass variance with a gradient descent, and the classifier is established to identify an early DDoS attack by training simple multiple-kernel learning (SMKL) models with two characteristics including interclass mean squared difference growth (M-SMKL) and intraclass variance descent (S-SMKL). The sliding window mechanism is used to coordinate the S-SMKL and M-SMKL to detect the early DDoS attack. The experimental results indicate that this method can detect DDoS attacks early and accurately.
Jieren Cheng, Chen Zhang 0009, Xiangyan Tang, Victor S. Sheng, Junqi Li
Secur. Commun. Networks1
2009 A Hybrid Parallel Signature Matching Model for Network Security Applications Using SIMD GPU
Chengkun Wu, Jianping Yin, Zhiping Cai, En Zhu, Jieren Cheng
APPT5
2009 DDoS Attack Detection Method Based on Linear Prediction Model
Jieren Cheng, Jianping Yin, Chengkun Wu, Boyun Zhang
ICIC (1)1
2006 A Novel Fairness Property of Electronic Commerce Protocols and Its Game-based Formalization
Jianping Yin, Jieren Cheng
SEKE4