VLDB 2026 Research / reviewers in the wild / expert
Wenxuan Tu
dblp:240/2589
· DBLP profile ↗
70ranked-venue papers
7as first author
68since 2021 · last 2026
0000-0002-1353-2968ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 6 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 3 first-author · 30 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Federated Graph-level Clustering Network with Attribute InferenceabstractWith the rise of vertical segmentation in real-world data, federated graph-level clustering has gained significant attention in recent years. However, the inherent missing attributes in graph datasets held by certain clients lead to suboptimal local parameter updates and misaligned global parameter consensus. This results in knowledge shifts during negotiation to ultimately impair overall clustering performance. This issue remains largely underexplored in the current advanced research. To bridge this gap, we propose a novel deep learning network called Federated Graph-level Clustering Network with Attribute Inference (FedAI), which utilizes high-confidence prior knowledge from each domain and multi-party collaborative optimization to achieve efficient reasoning of unknown features. Specifically, on the client, high-confidence graph samples are projected into a latent space. We then extract and upload irreversible path digest information and attribute-oriented inference signals from them. On the server, we first identify affinity relationships hierarchically via the improved graph kernel method. We then infer the features of clients lacking node attributes through a prior structure-guide recovery operator, facilitating inter-client knowledge transfer for better clustering. Experimental results on 15 cross-dataset and cross-domain non-IID graph datasets demonstrate that FedAI consistently outperforms existing methods. Renda Han, Wenxuan Tu, Jingxin Liu 0006, Jieren Cheng |
AAAI | 3 |
| 2026 | Cross-View Progressive Feature Filtering for Multi-View Graph Clustering in Remote SensingabstractMulti-view clustering of remote sensing data plays a vital role in Earth observation analysis. Recently, deep graph clustering methods based on contrastive learning have significantly improved feature representation capabilities. However, most existing approaches treat all views equally, neglecting the inherent uniqueness and heterogeneity across views, which often results in two major issues: 1) discriminative features from clustering-friendly views are underexplored; and 2) redundant or noisy information from less informative views can degrade the shared representation. To address these challenges, we propose a novel multi-view graph clustering framework termed CF-MVGC for remote sensing data, which dynamically preserves discriminative features and suppresses redundancy by assessing view affinity. Specifically, we employ a dual-stage representation learning strategy to extract both view-specific discriminative features and cross-view consistent representations. To further exploit and adaptively integrate complementary information across views, we design a progressive feature filtering model that dynamically evaluates view affinity using two novel metrics, i.e., view fidelity index (VFI) and view criticality index (VCI). Based on these assessments, the module adaptively modulates feature update and reset signals, reinforcing informative views while suppressing noisy or redundant ones. Views with high affinity receive strengthened update signals to retain valuable features, while those with low affinity are subjected to enhanced reset operations to eliminate noise and redundancy. The resulting high-quality, discriminative representations lead to improved clustering performance, establishing a positive feedback loop. Experimental results on four benchmark datasets demonstrate the effectiveness and superiority of CF-MVGC against its competitors. Bowen Liu 0020, Xin Peng 0010, Wenxuan Tu, Chengyao Wei, Xiangyan Tang, Jieren Cheng |
AAAI | 3 |
| 2026 | Personalized Federated Graph-Level Clustering NetworkabstractIn the federated clustering task, structural heterogeneity across clients inevitably impedes effective multi-source information sharing. To solve this issue, Personalized Federated Learning (PFL) has emerged as a potentially effective solution for image and text clustering. Unlike Euclidean data, graph-structured data exhibits diverse and fragile local patterns, which widely exist in real-world scenarios. Multi-graph data analysis in the federated learning setting is challenging and important, yet remains underexplored. This motivates us to propose a novel PERsonalized Federated graph-lEvel Clustering neTwork (PERFECT), which generates a specialized aggregation strategy for each client by uploading key model parameters and representative samples without sharing private information. Specifically, for each client, we first reconstruct privacy-preserving representative samples in a min-max optimization manner and then upload these samples to the server for subsequent personalized parameter aggregation. On the server, we first extract graph-level embeddings from the uploaded data, and then estimate affinities among multiple learned embeddings to formulate a personalized aggregation strategy for each client. Subsequently, to help each local model better identify the cluster boundaries, we utilize clustering-wise gradient to update the key components in the personalized model parameters from the server. Extensive experimental results have demonstrated the effectiveness and superiority of PERFECT over its competitors. Jingxin Liu 0006, Wenxuan Tu, Renda Han, Guohui Liu, Xiangyan Tang |
AAAI | 2 |
| 2026 | Causally-Aware Attribute Completion for Incomplete Federated Graph ClusteringabstractNode-level federated graph clustering allows multiple unlabeled subgraph holders to collaboratively train on node-level tasks without sharing private information. Existing methods usually assume that the node attributes are complete and have achieved promising progress. However, in the Federated Graph Learning (FGL) scenarios, this assumption is overly strict due to failures in data collection devices. Consequently, most existing FGL frameworks struggle to extract useful features from attribute-incomplete graphs for clustering, yet the issue remains underexplored. To bridge this gap, we propose a causally-aware attribute completion for Incomplete Federated Graph Clustering (IFedGC), which constructs a reliable global causal structure that incorporates clustering-friendly information to guide attribute completion for each subgraph. Specifically, in the attribute completion step, we first construct the causal structure to extract the causal relationships between initialized features, and then upload them to the server. Subsequently, we integrate multiple uploaded causal structures into a global causal one to achieve cross-client attribute completion. Moreover, to support reliable clustering, we first collect the high-confidence cluster centroids from each subgraph using a Graph Neural Network (GNN) model and subsequently aggregate these centroids on the server. The above two steps are seamlessly integrated into a unified FGL framework to obtain a clustering-oriented causal structure, which is sent back to the client to promote high-quality attribute completion for better clustering. Extensive results on five benchmark datasets demonstrate the effectiveness and superiority of IFedGC against its competitors. Jingxin Liu 0006, Wenxuan Tu, Renda Han, Haoyi Li, Xiangyan Tang |
AAAI | 2 |
| 2026 | Anchor-Driven Nyström for Deep Graph-Level ClusteringabstractGraph-level clustering (GLC), which aims to group entire graphs according to their structural and attribute-based similarities, represents a fundamental yet challenging task in various practical applications. Existing GLC methods primarily fall into two main paradigms: 1) deep graph clustering approaches based on Graph Neural Networks (GNNs), and 2) kernel-based methods that utilize predefined kernels to perform fine-grained structural comparison for clustering. However, GNN-based methods typically learn graph-level representations by aggregating node embeddings through pooling operations, which inevitably leads to substantial information loss and suboptimal clustering performance. In contrast, kernel methods, despite their theoretical expressiveness, suffer from prohibitive computational costs that hinder their scalability to large-scale settings. To solve these issues, we propose a novel graph learning framework named Anchor-driven Nyström for Deep Graph-Level Clustering (ANGC), which computes graph similarity via kernel methods while retaining the scalability of GNNs. Specifically, we first employ GNNs to encode individual graphs into sets of node embeddings. Rather than relying on pooling operations, we compute graph similarities in a kernel space constructed from these embeddings. To enhance both scalability and representational power, we introduce learnable graph Nyström anchors, which support end-to-end optimization and significantly accelerate kernel computations. To further improve the discriminative capability of these anchors, we propose the concept of anchor response discrepancy, that is, the variation in a given anchor’s responses across different samples. By maximizing this discrepancy, the anchors are encouraged to strengthen inter-graph distinctions for better clustering. Extensive experiments demonstrate the effectiveness and superiority of ANGC over existing state-of-the-art methods. Wenxuan Tu, Lingren Wang, Jieren Cheng |
AAAI | 2 |
| 2026 | FedPKDA: Personalized Federated Learning with Privacy-Preserving Knowledge Dynamic AlignmentabstractPersonalized Federated Learning (PFL), which aims to customize models for each client while preserving data privacy, has become an important research topic in addressing the challenges of data heterogeneity. Existing studies usually enhance the localization of global parameters by injecting local information into the globally shared model. However, these methods focus excessively on the personalized characteristics of individual clients and fail to fully exploit distinctive information across clients, limiting the quality of local models to represent unseen samples well. To address this issue, we propose a novel personalized Federated Privacy-preserving Knowledge Dynamic Alignment (FedPKDA) framework, which ensures data privacy during both the collection of client-side key information and its incorporation into federated model training. Specifically, to ensure data privacy during the cross-client information collection phase, we first conduct feature clipping and add Laplacian noise to the local prototypes extracted from each client. Further, we compute the centroid of the uploaded local prototypes in a latent space and leverage Mahalanobis distance to guide the generation of global prototypes, thereby preserving the semantic contributions from participating clients. Moreover, to boost the personalization of the local model, we dynamically align representations learned by the shared model with both a set of local prototypes and privacy-preserving global prototypes, facilitating effective cross-client knowledge sharing under heterogeneous settings while preserving client-specific characteristics. Extensive experiments on benchmark datasets have verified the superiority of FedPKDA against its competitors. Moxuan Zeng, Wenxuan Tu, Yiying Wang, Xiangyan Tang, Jieren Cheng |
AAAI | 2 |
| 2026 | Safe-FedLLM: Delving into the Safety of Federated Large Language ModelsabstractFederated learning (FL) addresses privacy and data-silo issues in the training of large language models (LLMs).Most prior work focuses on improving the efficiency of federated learning for LLMs (FedLLM).However, security in open federated environments, particularly defenses against malicious clients, remains underexplored.To investigate the security of FedLLM, we conduct a preliminary study to analyze potential attack surfaces and defensive characteristics from the perspective of LoRA updates.We find two key properties of FedLLM: 1) LLMs are vulnerable to attacks from malicious clients in FL, and 2) LoRA updates exhibit distinct behavioral patterns that can be effectively distinguished by lightweight classifiers.Based on these properties, we propose Safe-FedLLM, a probebased defense framework for FedLLM, which constructs defenses across three levels: Step-Level, Client-Level, and Shadow-Level.The core concept of Safe-FedLLM is to perform probe-based discrimination on each client's local LoRA updates, treating them as highdimensional behavioral features and using a lightweight classifier to determine whether they are malicious.Extensive experiments demonstrate that Safe-FedLLM effectively improves FedLLM's robustness against malicious clients while maintaining competitive performance on benign data.Notably, our method effectively suppresses the impact of malicious data without significantly affecting training speed, and remains effective even under high malicious client ratios. Mingxiang Tao, Wenxuan Tu, Xiangyan Tang |
ACL (1) | 3 |
| 2026 | FedCND: Federated Graph-Level Clustering under Inter-Client Cluster Number DiscrepancyabstractFederated graph-level clustering (FGC) provides an effective solution for analyzing decentralized graph data with privacy protection. Existing methods typically assume that all clients have the same number of clusters. This assumption simplifies the learning task and has achieved preliminary success. However, this assumption rarely holds in practice, as clients often exhibit substantial heterogeneity in both data distributions and semantic granularity. As a result, cluster-specific knowledge becomes misaligned during server-side aggregation, which ultimately degrades the overall clustering performance. To address this challenge, we propose a novel Federated Graph Clustering under Inter-Client Cluster Number Discrepancy (FedCND) framework, which aligns inter-client heterogeneous distributions by decoupling graph data into public and private patterns. Specifically, after initial local training and clustering on each client, we design a public learner and a private learner to model public and private graph data, respectively. Only anonymized, cluster-level public information is uploaded to the server, while private information remains local. On the server, cluster-level public prototypes are aggregated based on affinities between reconstructed cluster-level graphs, enabling privacy-preserving prototype alignment across clients with heterogeneous cluster numbers and mitigating interference from misaligned information during global aggregation. Finally, private subgraphs derive client-specific prototypes through local relearning, which are subsequently fused with globally oriented public prototypes for better clustering. Extensive experiments demonstrate that the proposed FedCND achieves an average of 4.9% accuracy improvement against current state-of-the-art methods. Renda Han, Wenxuan Tu, Jingxin Liu 0006, Jieren Cheng |
WWW | 3 |
| 2026 | Bio-inspired night vision: A cat-eye mechanism for low-light image enhancement
Chengchao Wang 0003, Jihao Guo, Jieren Cheng, Wenxuan Tu |
Knowl. Based Syst. | 7 |
| 2026 | Adaptive feature boosting and distribution refinement for graph clustering
Jingxin Liu 0006, Xiangyan Tang, Renda Han, Wenxuan Tu, Ruili Wang 0001 |
Pattern Recognit. | 4 |
| 2026 | MORSE: Molecular representation learning via structured semantic extraction across hierarchical and asymmetric biological modalities
Mengran Li 0001, Wenbin Xing, Bo Li 0128, Wenxuan Tu, Yongfu Li 0001, Ruxin Wang 0001 |
Pattern Recognit. | 6 |
| 2025 | Structure-Adaptive Multi-View Graph Clustering for Remote Sensing DataabstractMulti-view clustering (MVC) for remote sensing data is a critical and challenging task in Earth observation. Although recent advances in graph neural network (GNN)-based MVC have shown remarkable success, the most prevalent approaches have two major limitations: 1) heavily relying on a predefined yet fixed graph, which limits the performance of clustering because the large number of indistinguishable background samples contained in remote sensing data would introduce noise information and increase structure heterogeneity; 2) ignoring the effect of confusing samples on cluster structure compactness, which leads to fluffy cluster structure and decrease feature discriminability. To address these issues, we propose a Structure-Adaptive Multi-View Graph Clustering method named SAMVGC on remote sensing data which boosts the structure homogeneity and cluster compactness by adaptively learning the graph and cluster structures, respectively. Concretely, we use the geometric structure within the feature embedding space to refine adjacency matrices. The adjacency matrices are dynamically fused with the previous ones to improve the homogeneity and stability of structure information. Additionally, the samples are separated into two categories, including the central (intra-cluster center samples) and the confusing (inter-cluster boundary samples). On the basis, we deploy the contrastive learning paradigm on the central samples within views and the consistent learning paradigm on the confusing samples between views, improving the cluster compactness and consistency. Finally, we conduct extensive experiments on four benchmarks and achieve promising results, well demonstrating the effectiveness and superiority of the proposed method. Renxiang Guan, Wenxuan Tu, Siwei Wang 0001, Jiyuan Liu 0003, Dayu Hu, Chang Tang, Baili Xiao, Xinwang Liu 0002 |
AAAI | 2 |
| 2025 | Federated Graph-Level Clustering NetworkabstractFederated graph learning (FGL), which excels in analyzing non-IID graphs as well as protecting data privacy, has recently emerged as a hot topic. Existing FGL methods usually train the client model using labeled data and then collaboratively learn a global model without sharing their local graph data. However, in real-world scenarios, the lack of data annotations impedes the negotiation of multi-source information at the server, leading to sub-optimal feedback to the clients. To address this issue, we propose a novel unsupervised learning framework called Federated Graph-level Clustering Network (FedGCN), which collects the topology-oriented features of non-IID graphs from clients to generate global consensus representations through multi-source clustering structure sharing. Specifically, in the client, we first preserve the prototype features of each cluster from the structure-oriented embedding through clustering and then upload the learned multiple prototypes that are hard to be reconstructed into the raw graph data. In the server, we generate consensus prototypes from multiple condensed structure-oriented signals through Gaussian estimation, which are subsequently transferred to each client to promote the great encoding capacity of the local model for better clustering. Extensive experiments across multiple non-IID graph datasets have demonstrated the effectiveness and superiority of FedGCN against its competitors. Jingxin Liu 0006, Jieren Cheng, Renda Han, Wenxuan Tu, Xin Peng 0010 |
AAAI | 4 |
| 2025 | FedPKA: Federated Graph-Level Clustering Network with Personalized Knowledge Aggregation
Jingxin Liu 0006, Wenxuan Tu, Renda Han, Jieren Cheng, Xiangyan Tang |
ICIC (16) | 4 |
| 2025 | Federated Node-Level Clustering Network with Cross-Subgraph Link MendingabstractSubgraphs of a complete graph are usually distributed across multiple devices and can only be accessed locally because the raw data cannot be directly shared. However, existing node-level federated graph learning suffers from at least one of the following issues: 1) heavily relying on labeled graph samples that are difficult to obtain in real-world applications, and 2) partitioning a complete graph into several subgraphs inevitably causes missing links, leading to sub-optimal sample representations. To solve these issues, we propose a novel $\underline{\text{Fed}}$erated $\underline{\text{N}}$ode-level $\underline{\text{C}}$lustering $\underline{\text{N}}$etwork (FedNCN), which mends the destroyed cross-subgraph links using clustering prior knowledge. Specifically, within each client, we first design an MLP-based projector to implicitly preserve key clustering properties of a subgraph in a denoising learning-like manner, and then upload the resultant clustering signals that are hard to reconstruct for subsequent cross-subgraph links restoration. In the server, we maximize the potential affinity between subgraphs stemming from clustering signals by graph similarity estimation and minimize redundant links via the N-Cut criterion. Moreover, we employ a GNN-based generator to learn consensus prototypes from this mended graph, enabling the MLP-GNN joint-optimized learner to enhance data privacy during data transmission and further promote the local model for better clustering. Extensive experiments demonstrate the superiority of FedNCN. Jingxin Liu 0006, Renda Han, Wenxuan Tu, Jieren Cheng |
ICML | 3 |
| 2025 | Scalable Attribute-Missing Graph Clustering via Neighborhood DifferentiationabstractDeep graph clustering (DGC), which aims to unsupervisedly separate the nodes in an attribute graph into different clusters, has seen substantial potential in various industrial scenarios like community detection and recommendation. However, the real-world attribute graphs, e.g., social networks interactions, are usually large-scale and attribute-missing. To solve these two problems, we propose a novel DGC method termed **C**omplementary **M**ulti-**V**iew **N**eighborhood **D**ifferentiation ($\textit{CMV-ND}$), which preprocesses graph structural information into multiple views in a complete but non-redundant manner. First, to ensure completeness of the structural information, we propose a recursive neighborhood search that recursively explores the local structure of the graph by completely expanding node neighborhoods across different hop distances. Second, to eliminate the redundancy between neighborhoods at different hops, we introduce a neighborhood differential strategy that ensures no overlapping nodes between the differential hop representations. Then, we construct $K+1$ complementary views from the $K$ differential hop representations and the features of the target node. Last, we apply existing multi-view clustering or DGC methods to the views. Experimental results on six widely used graph datasets demonstrate that CMV-ND significantly improves the performance of various methods. Yaowen Hu, Wenxuan Tu, Yue Liu 0008, Xinhang Wan, Junyi Yan, Taichun Zhou, Xinwang Liu 0002 |
ICML | 2 |
| 2025 | Dual Boost-Driven Graph-Level Clustering NetworkabstractGraph-level clustering remains a pivotal yet formidable challenge in graph learning. Recently, the integration of deep learning with representation learning has demonstrated notable advancements, yielding performance enhancements to a certain degree. However, existing methods suffer from at least one of the following issues: 1) the original graph structure has noise, and 2) during feature propagation and pooling processes, noise is gradually aggregated into the graph-level embeddings through information propagation. Consequently, these two limitations mask clustering-friendly information, leading to suboptimal graph-level clustering performance. To this end, we propose a novel Dual Boost-Driven Graph-Level Clustering Network (DBGCN) to alternately promote graph-level clustering and filtering out interference information in a unified framework. Specifically, in the pooling step, we evaluate the contribution of features at the global and optimize them using a learnable transformation matrix to obtain high-quality graph-level representation, such that the model’s reasoning capability can be improved. Moreover, to enable reliable graph-level clustering, we first identify and suppress information detrimental to clustering by evaluating similarities between graph-level representations, providing more accurate guidance for multi-view fusion. Extensive experiments demonstrated that DBGCN outperforms the state-of-the-art graph-level clustering methods on six benchmark datasets. Renda Han, Wenxuan Tu, Wenxin Zhang 0005, Jingxin Liu 0006, Jieren Cheng, Huajie Lei, Guangzhen Yao, Lingren Wang, Yu Li 0047 |
IJCNN | 2 |
| 2025 | Multi-view Graph Clustering with Dual Structure Awareness for Remote Sensing DataabstractMulti-view clustering plays a pivotal role in remote sensing image analysis, where graph neural network-based methods have demonstrated remarkable potential by modeling data as graphs. However, existing efforts, which construct remote sensing graphs using fixed rules (e.g., K-nearest neighbors), inevitably introduce noisy edges and increase the risk of heterogeneous information diffusion, leading to inferior clustering performance. Although recent works attempt to address this issue by refining the structure, they are designed for single-view data and struggle to extend to multi-view scenarios. To bridge this gap, we propose a dual structure awareness multi-view graph clustering method named DSMVGC, which generates two distinct structures for each view through explicit and implicit perspectives. Specifically, in our method, the learning processes of structure refinement and clustering are alternately optimized to mutually enhance each other. On one hand, the explicit structure updates the topology based on inter-cluster relationships, while the implicit structure captures latent relationships not covered by the explicit structure through adversarial learning. On the other hand, the refined structures not only facilitate homogeneous message passing but also serve as prior knowledge to guide the contrastive loss, thereby enhancing the discriminability of representations for accurate clustering. Extensive experiments on five multi-view remote sensing datasets validate the effectiveness of DSMVGC. Xin Peng 0010, Bowen Liu 0020, Renxiang Guan, Wenxuan Tu |
ACM Multimedia | 4 |
| 2025 | Multi-view Graph Clustering with Dual Relation Optimization for Remote Sensing DataabstractMulti-view clustering (MVC) for remote sensing data has attracted increasing attention due to its ability to exploit complementary information from multiple modalities without requiring labels. Recent graph-based deep clustering methods have shown strong potential in modeling spatial structures inherent in remote sensing data. However, existing approaches often emphasize capturing rich node relations while overlooking the optimization of these relations, leading to noisy connections and weak inter-cluster discrimination. To address this issue, we propose a novel Multi-view Graph Clustering with dual Relation Optimization (MDRO) framework tailored for remote sensing data. Specifically, we first segment the remote sensing image into irregular superpixels to reduce computational complexity and use superpixels as graph nodes. Then, MDRO constructs high-order similarity matrices guided by clustering distribution matrices and performs dual relation optimization to suppress noise relations and strengthen similarity relations. Furthermore, an optimal transportation-based constraint is introduced to guide the formation of robust and balanced cluster assignments, mitigating over-smoothing and trivial solutions in graph learning. Comprehensive experiments on four benchmark remote sensing datasets demonstrate that MDRO consistently outperforms existing single-view and multi-view clustering methods, achieving superior accuracy and robustness. Renxiang Guan, Siwei Wang 0001, Wenxuan Tu, Miaomiao Li 0001, En Zhu, Xinwang Liu 0002, Ping Chen 0004 |
ACM Multimedia | 4 |
| 2025 | Divide-Then-Rule: A Cluster-Driven Hierarchical Interpolator for Attribute-Missing GraphsabstractDeep graph clustering (DGC) for attribute-missing graphs is an unsupervised task aimed at partitioning nodes with incomplete attributes into distinct clusters. Existing imputation methods for attribute-missing graphs often fail to account for the varying amounts of information available across node neighborhoods, leading to unreliable results. To address this issue, we propose a novel method named Divide-Then-Rule Graph Completion (DTRGC). This method first addresses nodes with sufficient known neighborhood information and treats the imputed results as new knowledge to iteratively impute more challenging nodes, while leveraging clustering information to correct imputation errors. Specifically, Dynamic Cluster-Aware Feature Propagation initializes missing node attributes by adjusting propagation weights based on the clustering structure. Subsequently, Hierarchical Neighborhood-Aware Imputation categorizes attribute-missing nodes into three groups based on the completeness of their neighborhood attributes. The imputation is performed hierarchically, prioritizing the groups with nodes that have the most available neighborhood information. The cluster structure is then used to refine the imputation and correct potential errors. Finally, Hop-wise Representation Enhancement integrates information across multiple hops, thereby enriching the expressiveness of node representations. Experimental results on 6 widely used graph datasets show that DTRGC significantly improves the clustering performance of various DGC methods under attribute-missing graphs. Yaowen Hu, Wenxuan Tu, Yue Liu 0008, Miaomiao Li 0001, Wenpeng Lu, Zhigang Luo, Xinwang Liu 0002, Ping Chen 0004 |
ACM Multimedia | 2 |
| 2025 | Discovering Maximum Frequency Consensus: Lightweight Federated Learning for Medical Image Segmentation
Lingren Wang, Wenxuan Tu, Jieren Cheng, Xiangyan Tang |
ACM Multimedia | 2 |
| 2025 | Hierarchical Shortest-Path Graph Kernel NetworkabstractGraph kernels have emerged as a fundamental and widely adopted technique in graph machine learning. However, most existing graph kernel methods rely on fixed graph similarity estimation that cannot be directly optimized for task-specific objectives, leading to sub-optimal performance. To address this limitation, we propose a kernel-based learning framework called Hierarchical Shortest-Path Graph Kernel Network HSP-GKN, which seamlessly integrates graph similarity estimation with downstream tasks within a unified optimization framework. Specifically, we design a hierarchical shortest-path graph kernel that efficiently preserves both the semantic and structural information of a given graph by transforming it into hierarchical features used for subsequent neural network learning. Building upon this kernel, we develop a novel end-to-end learning framework that matches hierarchical graph features with learnable $hidden$ graph features to produce a similarity vector. This similarity vector subsequently serves as the graph embedding for end-to-end training, enabling the neural network to learn task-specific representations. Extensive experimental results demonstrate the effectiveness and superiority of the designed kernel and its corresponding learning framework compared to current competitors. Wenxuan Tu, Jieren Cheng |
NeurIPS | 2 |
| 2025 | FedIGL: Federated Invariant Graph Learning for Non-IID GraphsabstractFederated Graph Learning (FGL) effectively facilitates cross-domain graph model training by enabling decentralized learning across multiple domains, while ensuring data privacy through local data storage and communication of model updates instead of raw data. Existing approaches usually assume shared generic knowledge (e.g., prototypes, spectral features) via aggregating local structures statistically to alleviate structural heterogeneity. However, imposing overly strict assumptions about the presumed correlation between structural features and the global objective often fails in generalizing to local tasks, leading to suboptimal performance. To tackle this issue, we propose a **Fed**erated **I**nvariant **G**raph **L**earning (**FedIGL**) framework based on invariant learning, which effectively disrupts spurious correlations and further mines the invariant factors across different distributions. Specifically, a server-side global model is trained to capture client-agnostic subgraph patterns shared across clients, whereas client-side models specialize in client-specific subgraph patterns. Subsequently, without compromising privacy, we propose a novel Bi-Gradient Regularization strategy that introduces gradient constraints to guide the model in identifying client-agnostic and client-specific subgraph patterns for better graph representations. Extensive experiments on graph-level clustering and classification tasks demonstrate the superiority of FedIGL against its competitors. Lingren Wang, Wenxuan Tu, Jieren Cheng, Jingxin Liu 0006 |
NeurIPS | 2 |
| 2025 | IIM-ARE: An Effective Interactive Incentive Mechanism Based on Adaptive Reputation Evaluation for Mobile Crowd SensingabstractMobile crowd sensing (MCS), as an innovative data acquisition model in the Internet of Things (IoT), employs an incentive mechanism based on users’ reputation evaluation, which is a mainstream reward allocation method. However, in the existing incentive mechanisms based on reputation evaluation, unidirectional incentive strategies and nonadaptive reputation models result in unequal reward allocation. To tackle this issue, we propose an effective interactive incentive mechanism based on adaptive reputation evaluation. Specifically, we generate user status thresholds to classify, rate, and weight user behaviors, based on the average quality thresholds of tasks released or data submitted by different users in each interaction round. Meanwhile, we achieve multiparty consensus by incorporating the obtained user reputation values and combining them with the cumulative reputation values from multiple rounds to obtain adaptive reputation evaluation results. Moreover, we design an interactive incentive strategy that measures users’ incentive values based on their reputation evaluation results in each round, mutually punishing malicious behaviors from both the publisher’s and the worker’s perspectives. Extensive experiments have demonstrated that our method consistently outperforms existing advanced incentive mechanisms. Xiangyan Tang, Jingxin Liu 0006, Keqiu Li, Wenxuan Tu, Xinbin Xu, Naixue Xiong |
IEEE Internet Things J. | 4 |
| 2025 | WAGE: Weight-Sharing Attribute-Missing Graph AutoencoderabstractAttribute-missing graph learning, a common yet challenging problem, has recently attracted considerable attention. Existing efforts have at least one of the following limitations: 1) lack a noise filtering and information enhancing scheme, resulting in less comprehensive data completion; 2) isolate the node attribute and graph structure encoding processes, introducing more parameters and failing to take full advantage of the two types of information; and 3) impose overly strict distribution assumptions on the latent variables, leading to biased or less discriminative node representations. To tackle the issues, based on the idea of introducing intimate information interaction between the two information sources, we propose Weight-sharing Attribute-missing Graph autoEncoder (WAGE) to boost the expressive capacity of node representations for high-quality missing attribute reconstruction. Specifically, three strategies have been conducted. Firstly, we entangle the attribute embedding and structure embedding by introducing a weight-sharing architecture to share the parameters learned by both processes, which allows the network training to benefit from more abundant and diverse information. Secondly, we introduce a $K$K-nearest neighbor-based dual non-local learning mechanism to improve the quality of data imputation by revealing unobserved high-confidence connections while filtering unreliable ones. Thirdly, we manually mask the connections on multiple adjacency matrices and force the structure-oriented embedding sub-network to recover the actual adjacency matrix, thus enforcing the resulting network to be able to selectively exploit more high-order discriminative features for data completion. Extensive experiments on six benchmark datasets demonstrate the effectiveness and superiority of WAGE against state-of-the-art competitors. Wenxuan Tu, Sihang Zhou 0001, Xinwang Liu 0002, Zhiping Cai, Yue Liu 0008, Kunlun He |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Sampling Enhanced Contrastive Multi-View Remote Sensing Data Clustering With Long-Short Range Information MiningabstractMulti-view clustering (MVC) for remote sensing data has demonstrated significant potential in Earth observation, given its ability to aggregate multi-source information without relying on labels. Despite achieving compelling results through the combination of deep encoders and contrastive learning, existing algorithms still face two limitations: inadequate exploration of diverse spatial relationships and inability to guide the selection of sample pairs leads to blind sampling, both of which lead to suboptimal clustering performance. To tackle these challenges, we propose a sampling enhanced contrastive multi-view clustering method for remote sensing data, namely SEC-LSRM. The proposed method incorporates long- and short-range information mining to enhance clustering performance. By aggregating shortrange information extracted through autoencoders and longrange information obtained via graph autoencoders, our method improves the sampling quality of positive and negative sample pairs. To render the extracted features more compact, a multiview correlation reduction strategy is devised to filter out irrelevant information. With the extracted comprehensive features, an adaptive sampling strategy is designed to obtain high-quality positive and negative samples. Subsequently, we select positive and negative sample pairs based on these affinity matrices with idempotence and block diagonal constraints. Moreover, we integrate the optimization of these sample pairs and contrastive learning within the same framework to achieve iterative updates of both. Experiments conducted on multiple multi-view remote sensing datasets illustrate that our proposed SEC-LSRM method achieves excellent and reliable clustering performance. Renxiang Guan, Tianrui Liu 0001, Wenxuan Tu, Chang Tang, Wenhan Luo, Xinwang Liu 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | GZOO: Black-Box Node Injection Attack on Graph Neural Networks via Zeroth-Order OptimizationabstractThe ubiquity of Graph Neural Networks (GNNs) emphasizes the imperative to assess their resilience against node injection attacks, a type of evasion attacks that impact victim models by injecting nodes with fabricated attributes and structures. However, prevailing attacks face two primary limitations: (1) Sequential construction of attributes and structures results in suboptimal outcomes as structure information is overlooked during attribute construction and vice versa. (2) In black-box scenarios, where attackers lack access to victim model architecture and parameters, reliance on surrogate models degrades performance due to architectural discrepancies. To overcome these limitations, we introduce GZOO, a black-box node injection attack that leverages an adversarial graph generator, compromising both attribute and structure sub-generators. This integration crafts optimal attributes and structures by considering their mutual information, enhancing their influence when aggregating information from injected nodes. Furthermore, GZOO proposes a zeroth-order optimization algorithm leveraging prediction results from victim models to estimate gradients for updating generator parameters, eliminating the necessity to train surrogate models. Across sixteen datasets, GZOO significantly outperforms state-of-the-art attacks, achieving remarkable effectiveness and robustness. Notably, on the Cora dataset with the GCN model, GZOO achieves an impressive 95.69% success rate, surpassing the maximum 66.01% achieved by baselines. Hao Yu 0017, Ke Liang 0006, Dayu Hu, Wenxuan Tu, Chuan Ma 0001, Sihang Zhou 0001, Xinwang Liu 0002 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Prototype-Driven Multi-View Attribute-Missing Graph ClusteringabstractAttribute-missing deep graph clustering, which aims to categorize the graph nodes with partial attribute-missing samples into distinct categories in an unsupervised manner, has gained significant popularity. However, most existing researches have at least one of the following issues: 1) seldom exploit diverse clustering structural information to facilitate non-Euclidean data imputation and refine the clustering pattern and 2) ignoring the positive effect of diverse information on feature imputation and representation extraction, resulting in sub-optimal missing feature estimation and inferior clustering performance. To solve these issues, we propose a novelPrototype-drivenMulti-viewAttribute-missingGraphClustering (PMAGC) model that leverages rich structural and diverse information to assist the processes of imputing missing attributes and learning clustering-friendly features. Specifically, we design a multi-view augmentation module that extracts attribute-complete samples as node view and constructs feature and edge views using feature pre-imputation and edge masking techniques. Then, guided by clustering pseudo-labels, we promote the proximity between the prototypes of attribute-missing samples and those of attribute-complete samples within the feature space. Thus, PMAGC cleverly employs both clustering structural information and reliably attribute-complete sample data to assist feature imputation. In addition, we design a prototype-wise contrastive loss, which considers prototypes from different views within the same cluster as positive samples, while treating others as negative samples. Hence, the optimized features could more accurately guide the attribute learning process. Extensive experiments on six graph datasets with missing attributes are conducted to demonstrate the effectiveness of the proposed PMAGE. Renxiang Guan, Wenxuan Tu, Dayu Hu, Weixuan Liang, Ke Liang 0006, Yaowen Hu, Yue Liu 0008, Xinwang Liu 0002 |
IEEE Trans. Multim. | 2 |
| 2025 | Self-Supervised Temporal Graph Learning With Temporal and Structural Intensity AlignmentabstractTemporal graph learning aims to generate high-quality representations for graph-based tasks with dynamic information, which has recently garnered increasing attention. In contrast to static graphs, temporal graphs are typically organized as node interaction sequences over continuous time rather than an adjacency matrix. Most temporal graph learning methods model current interactions by incorporating historical neighborhood. However, such methods only consider first-order temporal information while disregarding crucial high-order structural information, resulting in suboptimal performance. To address this issue, we propose a self-supervised method called S2T for temporal graph learning, which extracts both temporal and structural information to learn more informative node representations. Notably, the initial node representations combine first-order temporal and high-order structural information differently to calculate two conditional intensities. An alignment loss is then introduced to optimize the node representations, narrowing the gap between the two intensities and making them more informative. Concretely, in addition to modeling temporal information using historical neighbor sequences, we further consider structural knowledge at both local and global levels. At the local level, we generate structural intensity by aggregating features from high-order neighbor sequences. At the global level, a global representation is generated based on all nodes to adjust the structural intensity according to the active statuses on different nodes. Extensive experiments demonstrate that the proposed model S2T achieves at most 10.13% performance improvement compared with the state-of-the-art competitors on several datasets. Meng Liu 0014, Ke Liang 0006, Wenxuan Tu, Sihang Zhou 0001, Xinbiao Gan, Xinwang Liu 0002, Kunlun He |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Improved Dual Correlation Reduction Network With Affinity RecoveryabstractDeep graph clustering, which aims to reveal the underlying graph structure and divide the nodes into different clusters without human annotations, is a fundamental yet challenging task. However, we observe that the existing methods suffer from the representation collapse problem and tend to encode samples with different classes into the same latent embedding. Consequently, the discriminative capability of nodes is limited, resulting in suboptimal clustering performance. To address this problem, we propose a novel deep graph clustering algorithm termed improved dual correlation reduction network (IDCRN) through improving the discriminative capability of samples. Specifically, by approximating the cross-view feature correlation matrix to an identity matrix, we reduce the redundancy between different dimensions of features, thus improving the discriminative capability of the latent space explicitly. Meanwhile, the cross-view sample correlation matrix is forced to approximate the designed clustering-refined adjacency matrix to guide the learned latent representation to recover the affinity matrix even across views, thus enhancing the discriminative capability of features implicitly. Moreover, we avoid the collapsed representation caused by the oversmoothing issue in graph convolutional networks (GCNs) through an introduced propagation regularization term, enabling IDCRN to capture the long-range information with the shallow network structure. Extensive experimental results on six benchmarks have demonstrated the effectiveness and efficiency of IDCRN compared with the existing state-of-the-art deep graph clustering algorithms. The code of IDCRN is released at IDCRN. Besides, we share a collection of deep graph clustering, including papers, codes, and datasets at ADGC. Yue Liu 0008, Sihang Zhou 0001, Xihong Yang, Xinwang Liu 0002, Wenxuan Tu, Liang Li 0041, Xin Xu 0001, Fuchun Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Revisiting Initializing Then Refining: An Incomplete and Missing Graph Imputation NetworkabstractWith the development of various applications, such as recommendation systems and social network analysis, graph data have been ubiquitous in the real world. However, graphs usually suffer from being absent during data collection due to copyright restrictions or privacy-protecting policies. The graph absence could be roughly grouped into attribute-incomplete and attribute-missing cases. Specifically, attribute-incomplete indicates that a portion of the attribute vectors of all nodes are incomplete, while attribute-missing indicates that all attribute vectors of partial nodes are missing. Although various graph imputation methods have been proposed, none of them is custom-designed for a common situation where both types of graph absence exist simultaneously. To fill this gap, we develop a novel graph imputation network termed revisiting initializing then refining (RITR), where both attribute-incomplete and attribute-missing samples are completed under the guidance of a novel initializing-then-refining imputation criterion. Specifically, to complete attribute-incomplete samples, we first initialize the incomplete attributes using Gaussian noise before network learning, and then introduce a structure-attribute consistency constraint to refine incomplete values by approximating a structure-attribute correlation matrix to a high-order structure matrix. To complete attribute-missing samples, we first adopt structure embeddings of attribute-missing samples as the embedding initialization, and then refine these initial values by adaptively aggregating the reliable information of attribute-incomplete samples according to a dynamic affinity structure. To the best of our knowledge, this newly designed method is the first end-to-end unsupervised framework dedicated to handling hybrid-absent graphs. Extensive experiments on six datasets have verified that our methods consistently outperform the existing state-of-the-art competitors. Our source code is available at https://github.com/WxTu/RITR. Wenxuan Tu, Bin Xiao 0002, Xinwang Liu 0002, Sihang Zhou 0001, Zhiping Cai, Jieren Cheng |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Hawkes-Enhanced Spatial-Temporal Hypergraph Contrastive Learning Based on Criminal CorrelationsabstractCrime prediction is a crucial yet challenging task within urban computing, which benefits public safety and resource optimization. Over the years, various models have been proposed, and spatial-temporal hypergraph learning models have recently shown outstanding performances. However, three correlations underlying crime are ignored, thus hindering the performance of previous models. Specifically, there are two spatial correlations and one temporal correlation, i.e., (1) co-occurrence of different types of crimes (type spatial correlation), (2) the closer to the crime center, the more dangerous it is around the neighborhood area (neighbor spatial correlation), and (3) the closer between two timestamps, the more relevant events are (hawkes temporal correlation). To this end, we propose Hawkes-enhanced Spatial-Temporal Hypergraph Contrastive Learning framework (HCL), which mines the aforementioned correlations via two specific strategies. Concretely, contrastive learning strategies are designed for two spatial correlations, and hawkes process modeling is adopted for temporal correlations. Extensive experiments demonstrate the promising capacities of HCL from four aspects, i.e., superiority, transferability, effectiveness, and sensitivity. Ke Liang 0006, Sihang Zhou 0001, Meng Liu 0014, Yue Liu 0008, Wenxuan Tu, Yi Zhang 0104, Liming Fang 0001, Zhe Liu 0001, Xinwang Liu 0002 |
AAAI | 5 |
| 2024 | MINES: Message Intercommunication for Inductive Relation Reasoning over Neighbor-Enhanced SubgraphsabstractGraIL and its variants have shown their promising capacities for inductive relation reasoning on knowledge graphs. However, the uni-directional message-passing mechanism hinders such models from exploiting hidden mutual relations between entities in directed graphs. Besides, the enclosing subgraph extraction in most GraIL-based models restricts the model from extracting enough discriminative information for reasoning. Consequently, the expressive ability of these models is limited. To address the problems, we propose a novel GraIL-based framework, termed MINES, by introducing a Message Intercommunication mechanism on the Neighbor-Enhanced Subgraph. Concretely, the message intercommunication mechanism is designed to capture the omitted hidden mutual information. It introduces bi-directed information interactions between connected entities by inserting an undirected/bi-directed GCN layer between uni-directed RGCN layers. Moreover, inspired by the success of involving more neighbors in other graph-based tasks, we extend the neighborhood area beyond the enclosing subgraph to enhance the information collection for inductive relation reasoning. Extensive experiments prove the promising capacity of the proposed MINES from various aspects, especially for the superiority, effectiveness, and transfer ability. Ke Liang 0006, Lingyuan Meng, Sihang Zhou 0001, Wenxuan Tu, Siwei Wang 0001, Yue Liu 0008, Meng Liu 0014, Long Zhao 0002, Xiangjun Dong 0001, Xinwang Liu 0002 |
AAAI | 4 |
| 2024 | Attribute-Missing Graph Clustering NetworkabstractDeep clustering with attribute-missing graphs, where only a subset of nodes possesses complete attributes while those of others are missing, is an important yet challenging topic in various practical applications. It has become a prevalent learning paradigm in existing studies to perform data imputation first and subsequently conduct clustering using the imputed information. However, these ``two-stage" methods disconnect the clustering and imputation processes, preventing the model from effectively learning clustering-friendly graph embedding. Furthermore, they are not tailored for clustering tasks, leading to inferior clustering results. To solve these issues, we propose a novel Attribute-Missing Graph Clustering (AMGC) method to alternately promote clustering and imputation in a unified framework, where we iteratively produce the clustering-enhanced nearest neighbor information to conduct the data imputation process and utilize the imputed information to implicitly refine the clustering distribution through model optimization. Specifically, in the imputation step, we take the learned clustering information as imputation prompts to help each attribute-missing sample gather highly correlated features within its clusters for data completion, such that the intra-class compactness can be improved. Moreover, to support reliable clustering, we maximize inter-class separability by conducting cost-efficient dual non-contrastive learning over the imputed latent features, which in turn promotes greater graph encoding capability for clustering sub-network. Extensive experiments on five datasets have verified the superiority of AMGC against competitors. Wenxuan Tu, Renxiang Guan, Sihang Zhou 0001, Chuan Ma 0001, Xin Peng 0010, Zhiping Cai, Zhe Liu 0001, Jieren Cheng, Xinwang Liu 0002 |
AAAI | 1 |
| 2024 | A Non-parametric Graph Clustering Framework for Multi-View DataabstractMulti-view graph clustering (MVGC) derives encouraging grouping results by seamlessly integrating abundant information inside heterogeneous data, and has captured surging focus recently. Nevertheless, the majority of current MVGC works involve at least one hyper-parameter, which not only requires additional efforts for tuning, but also leads to a complicated solving procedure, largely harming the flexibility and scalability of corresponding algorithms. To this end, in the article we are devoted to getting rid of hyper-parameters, and devise a non-parametric graph clustering (NpGC) framework to more practically partition multi-view data. To be specific, we hold that hyper-parameters play a role in balancing error item and regularization item so as to form high-quality clustering representations. Therefore, under without the assistance of hyper-parameters, how to acquire high-quality representations becomes the key. Inspired by this, we adopt two types of anchors, view-related and view-unrelated, to concurrently mine exclusive characteristics and common characteristics among views. Then, all anchors' information is gathered together via a consensus bipartite graph. By such ways, NpGC extracts both complementary and consistent multi-view features, thereby obtaining superior clustering results. Also, linear complexities enable it to handle datasets with over 120000 samples. Numerous experiments reveal NpGC's strong points compared to lots of classical approaches. Shengju Yu, Siwei Wang 0001, Zhibin Dong, Wenxuan Tu, Suyuan Liu, Zhao Lv, En Zhu |
AAAI | 4 |
| 2024 | Deep Temporal Graph ClusteringabstractDeep graph clustering has recently received significant attention due to its ability to enhance the representation learning capabilities of models in unsupervised scenarios. Nevertheless, deep clustering for temporal graphs, which could capture crucial dynamic interaction information, has not been fully explored. It means that in many clustering-oriented real-world scenarios, temporal graphs can only be processed as static graphs. This not only causes the loss of dynamic information but also triggers huge computational consumption. To solve the problem, we propose a general framework for deep Temporal Graph Clustering called TGC, which introduces deep clustering techniques to suit the interaction sequence-based batch-processing pattern of temporal graphs. In addition, we discuss differences between temporal graph clustering and static graph clustering from several levels. To verify the superiority of the proposed framework TGC, we conduct extensive experiments. The experimental results show that temporal graph clustering enables more flexibility in finding a balance between time and space requirements, and our framework can effectively improve the performance of existing temporal graph learning methods. The code is released: https://github.com/MGitHubL/Deep-Temporal-Graph-Clustering. Meng Liu 0014, Yue Liu 0008, Ke Liang 0006, Wenxuan Tu, Siwei Wang 0001, Sihang Zhou 0001, Xinwang Liu 0002 |
ICLR | 4 |
| 2024 | Towards Resource-friendly, Extensible and Stable Incomplete Multi-view ClusteringabstractIncomplete multi-view clustering (IMVC) methods typically encounter three drawbacks: (1) intense time and/or space overheads; (2) intractable hyper-parameters; (3) non-zero variance results. With these concerns in mind, we give a simple yet effective IMVC scheme, termed as ToRES. Concretely, instead of self-expression affinity, we manage to construct prototype-sample affinity for incomplete data so as to decrease the memory requirements. To eliminate hyper-parameters, besides mining complementary features among views by view-wise prototypes, we also attempt to devise cross-view prototypes to capture consensus features for jointly forming high-quality clustering representation. To avoid the variance, we successfully unify representation learning and clustering operation, and directly optimize the discrete cluster indicators from incomplete data. Then, for the resulting objective function, we provide two equivalent solutions from perspectives of feasible region partitioning and objective transformation. Many results suggest that ToRES exhibits advantages against 20 SOTA algorithms, even in scenarios with a higher ratio of incomplete data. Shengju Yu, Zhibin Dong, Siwei Wang 0001, Xinhang Wan, Yue Liu 0008, Weixuan Liang, Pei Zhang 0008, Wenxuan Tu, Xinwang Liu 0002 |
ICML | 8 |
| 2024 | TabSec: A Collaborative Framework for Novel Insider Threat DetectionabstractIn the era of the Internet of Things (IoT) and data sharing, users frequently upload their personal information to enterprise databases to enjoy enhanced service experiences provided by various online services. However, the widespread presence of system vulnerabilities, remote network intrusions, and insider threats significantly increases the exposure of private enterprise data on the internet. If such data is stolen or leaked by attackers, it can result in severe asset losses and business operation disruptions. To address these challenges, this paper proposes a novel threat detection framework, TabITD. This framework integrates Intrusion Detection Systems (IDS) with User and Entity Behavior Analytics (UEBA) strategies to form a collaborative detection system that bridges the gaps in existing systems’ capabilities. It effectively addresses the blurred boundaries between external and insider threats caused by the diversification of attack methods, thereby enhancing the model’s learning ability and overall detection performance. Moreover, the proposed method leverages the TabNet architecture, which employs a sparse attention feature selection mechanism that allows TabNet to select the most relevant features at each decision step, thereby improving the detection of rare-class attacks. We evaluated our proposed solution on two different datasets, achieving average accuracies of 96.71% and 97.25%, respectively. The results demonstrate that this approach can effectively detect malicious behaviors such as masquerade attacks and external threats, significantly enhancing network security defenses and the efficiency of network attack detection. Xiangyan Tang, Xinyi Cao, Jieren Cheng, Wenxuan Tu, Logan Bo-Yee Liu |
ISPA | 6 |
| 2024 | Dynamic Position Transformation and Boundary Refinement Network for Left Atrial Segmentation
Fangqiang Xu, Wenxuan Tu, Malitha Gunawardhana, Jiayuan Yang, Yun Gu, Jichao Zhao |
MICCAI (8) | 2 |
| 2024 | Simple Yet Effective: Structure Guided Pre-trained Transformer for Multi-modal Knowledge Graph ReasoningabstractVarious information in different modalities in an intuitive way in multi-modal knowledge graphs (MKGs), which are utilized in different downstream tasks, like recommendation. However, most MKGs are still far from complete, which motivates the flourishing of MKG reasoning models. Recently, with the development of general artificial intelligence, pre-trained transformers have drawn increasing attention, especially in multi-modal scenarios. However, the research of multi-modal pre-trained transformers (MPT) for knowledge graph reasoning (KGR) is still at an early stage. As the biggest difference between MKG and other multi-modal data, the rich structural information underlying the MKG is still not fully utilized in previous MPT. Most of them only use the graph structure as a retrieval map for matching images and texts connected with the same entity, which hinders their reasoning performances. To this end, the graph Structure Guided Multi-modal Pre-trained Transformer is proposed for knowledge graph reasoning (SGMPT). Specifically, the graph structure encoder is adopted for structural feature encoding. Then, a structure-guided fusion module with two simple yet effective strategies, i.e., weighted summation and alignment constraint, is designed to inject the structural information into both the textual and visual features. To the best of our knowledge, SGMPT is the first MPT for multi-modal KGR, which mines structural information underlying MKGs. Extensive experiments on FB15k-237-IMG and WN18-IMG, demonstrate that our SGMPT outperforms existing state-of-the-art models, and proves the effectiveness of the designed strategies. Ke Liang 0006, Lingyuan Meng, Yue Liu 0008, Meng Liu 0014, Suyuan Liu, Wenxuan Tu, Siwei Wang 0001, Sihang Zhou 0001, Xinwang Liu 0002 |
ACM Multimedia | 7 |
| 2024 | A Survey of Knowledge Graph Reasoning on Graph Types: Static, Dynamic, and Multi-ModalabstractKnowledge graph reasoning (KGR), aiming to deduce new facts from existing facts based on mined logic rules underlying knowledge graphs (KGs), has become a fast-growing research direction. It has been proven to significantly benefit the usage of KGs in many AI applications, such as question answering, recommendation systems, and etc. According to the graph types, existing KGR models can be roughly divided into three categories, i.e., static models, temporal models, and multi-modal models. Early works in this domain mainly focus on static KGR, and recent works try to leverage the temporal and multi-modal information, which are more practical and closer to real-world. However, no survey papers and open-source repositories comprehensively summarize and discuss models in this important direction. To fill the gap, we conduct a first survey for knowledge graph reasoning tracing from static to temporal and then to multi-modal KGs. Concretely, the models are reviewed based on bi-level taxonomy, i.e., top-level (graph types) and base-level (techniques and scenarios). Besides, the performances, as well as datasets, are summarized and presented. Moreover, we point out the challenges and potential opportunities to enlighten the readers. Ke Liang 0006, Lingyuan Meng, Meng Liu 0014, Yue Liu 0008, Wenxuan Tu, Siwei Wang 0001, Sihang Zhou 0001, Xinwang Liu 0002, Fuchun Sun 0001, Kunlun He |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | GANN: Graph Alignment Neural Network for semi-supervised learning
Linxuan Song, Wenxuan Tu, Sihang Zhou 0001, En Zhu |
Pattern Recognit. | 2 |
| 2024 | Contrastive Multiview Subspace Clustering of Hyperspectral Images Based on Graph Convolutional NetworksabstractHigh-dimensional and complex spectral structures make the clustering of hyperspectral images (HSI) a challenging task. Subspace clustering is an effective approach for addressing this problem. However, current subspace clustering algorithms are primarily designed for a single view and do not fully exploit the spatial or textural feature information in HSI. In this study, contrastive multi-view subspace clustering of HSI was proposed based on graph convolutional networks. Pixel neighbor textural and spatial-spectral information were sent to construct two graph convolutional subspaces to learn their affinity matrices. To maximize the interaction between different views, a contrastive learning algorithm was introduced to promote the consistency of positive samples and assist the model in extracting robust features. An attention-based fusion module was used to adaptively integrate these affinity matrices, constructing a more discriminative affinity matrix. The model was evaluated using four popular HSI datasets: Indian Pines, Pavia University, Houston, and Xu Zhou. It achieved overall accuracies of 97.61%, 96.69%, 87.21%, and 97.65%, respectively, and significantly outperformed state-of-the-art clustering methods. In conclusion, the proposed model effectively improves the clustering accuracy of HSI. Our implementation is available at https://github.com/GuanRX/CMSCGC. Renxiang Guan, Wenxuan Tu, Jun Wang 0118, Yue Liu 0008, Xianju Li, Chang Tang, Ruyi Feng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Spatial-Spectral Graph Contrastive Clustering With Hard Sample Mining for Hyperspectral ImagesabstractHyperspectral image (HSI) clustering is a fundamental yet challenging task that groups image pixels with similar features into distinct clusters. Among various approaches, contrastive learning methods, which employ the concept of encouraging semantically similar samples to move closer together while pushing semantically inconsistent samples apart, have garnered significant attention due to their promising performance. However, the most prevalent approaches face two major limitations: 1) treating all samples indiscriminately during optimization, where the abundance of well-categorized samples overwhelms the feature learning process and 2) tending to introduce noise when constructing positive sample pairs through view augmentation or searching the nearest neighbors, which would cause semantic drift of sample features. To solve these issues, we propose a graph autoencoder-based deep clustering framework named spatial–spectral graph contrastive clustering with hard sample mining (SSGCC) that constructs spatial–spectral dual views without data augmentation and focuses more on hard samples rather than treating all samples equally with the aid of spatial–spectral features. Concretely, we extract the spectral features and the neighborhood spatial features of the samples as dual branches to avoid the noise caused by data augmentation and develop the cluster-oriented consistency learning to facilitate the exchange of knowledge between the two spectral–spatial perspectives. In addition, we propose a hard sample mining-based contrastive learning scheme with the aid of spatial–spectral features. To better measure the importance of the samples, we combine spatial features and spectral features to calculate the similarity between sample pairs. The weights of hard sample pairs are dynamically up-weight while the easy ones are down-weighting to improve the discriminative capability. Extensive experiments on four benchmark HSI datasets demonstrate the effectiveness and superiority of the proposed methods against state-of-the-art ones. Renxiang Guan, Wenxuan Tu, Hao Yu 0017, Dayu Hu, Yuzeng Chen, Chang Tang, Qiangqiang Yuan, Xinwang Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Scalable and Structural Multi-View Graph Clustering With Adaptive Anchor FusionabstractAnchor graph has been recently proposed to accelerate multi-view graph clustering and widely applied in various large-scale applications. Different from capturing full instance relationships, these methods choose small portion anchors among each view, construct single-view anchor graphs and combine them into the unified graph. Despite its efficiency, we observe that: (i) Existing mechanism adopts a separable two-step procedure-anchor graph construction and individual graph fusion, which may degrade the clustering performance. (ii)These methods determine the number of selected anchors to be equal among all the views, which may destruct the data distribution diversity. A more flexible multi-view anchor graph fusion framework with diverse magnitudes is desired to enhance the representation ability. (iii) During the latter fusion process, current anchor graph fusion framework follows simple linearly-combined style while the intrinsic clustering structures are ignored. To address these issues, we propose a novel scalable and flexible anchor graph fusion framework for multi-view graph clustering method in this paper. Specially, the anchor graph construction and graph alignment are jointly optimized in our unified framework to boost clustering quality. Moreover, we present a novel structural alignment regularization to adaptively fuse multiple anchor graphs with different magnitudes. In addition, our proposed method inherits the linear complexity of existing anchor strategies respecting to the sample number, which is time-economical for large-scale data. Experiments conducted on various benchmark datasets demonstrate the superiority and effectiveness of the newly proposed anchor graph fusion framework against the existing state-of-the-arts over the clustering performance promotion and time expenditure. Our code is publicly available at https://github.com/wangsiwei2010/SMVAGC-SF. Siwei Wang 0001, Xinwang Liu 0002, Suyuan Liu, Wenxuan Tu, En Zhu |
IEEE Trans. Image Process. | 4 |
| 2024 | Knowledge Graph Contrastive Learning Based on Relation-Symmetrical StructureabstractKnowledge graph embedding (KGE) aims at learning powerful representations to benefit various artificial intelligence applications. Meanwhile, contrastive learning has been widely leveraged in graph learning as an effective mechanism to enhance the discriminative capacity of the learned representations. However, the complex structures of KG make it hard to construct appropriate contrastive pairs. Only a few attempts have integrated contrastive learning strategies with KGE. But, most of them rely on language models (e.g.,Bert) for contrastive pair construction instead of fully mining information underlying the graph structure, hindering expressive ability. Surprisingly, we find that the entities within a relational symmetrical structure are usually similar and correlated. To this end, we propose a knowledge graph contrastive learning framework based on relation-symmetrical structure, KGE-SymCL, which mines symmetrical structure information in KGs to enhance the discriminative ability of KGE models. Concretely, a plug-and-play approach is proposed by taking entities in the relation-symmetrical positions as positive pairs. Besides, a self-supervised alignment loss is designed to pull together positive pairs. Experimental results on link prediction and entity classification datasets demonstrate that our KGE-SymCL can be easily adopted to various KGE models for performance improvements. Moreover, extensive experiments show that our model could outperform other state-of-the-art baselines. Ke Liang 0006, Yue Liu 0008, Sihang Zhou 0001, Wenxuan Tu, Yi Wen 0001, Xihong Yang, Xiangjun Dong 0001, Xinwang Liu 0002 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | RARE: Robust Masked Graph AutoencoderabstractMasked graph autoencoder (MGAE) has emerged as a promising self-supervised graph pre-training (SGP) paradigm due to its simplicity and effectiveness. However, existing efforts perform the mask-then-reconstruct operation in the raw data space as is done in computer vision (CV) and natural language processing (NLP) areas, while neglecting the important non-Euclidean property of graph data. As a result, the highly unstable local structures largely increase the uncertainty in inferring masked data and decrease the reliability of the exploited self-supervision signals, leading to inferior representations for downstream evaluations. To address this issue, we propose a novel SGP method termed Robust mAsked gRaph autoEncoder (RARE) to improve the certainty in inferring masked data and the reliability of the self-supervision mechanism by further masking and reconstructing node samples in the high-order latent feature space. Through both theoretical and empirical analyses, we have discovered that performing a joint mask-then-reconstruct strategy in both latent feature and raw data spaces could yield improved stability and performance. To this end, we elaborately design a masked latent feature completion scheme, which predicts latent features of masked nodes under the guidance of high-order sample correlations that are hard to be observed from the raw data perspective. Specifically, we first adopt a latent feature predictor to predict the masked latent features from the visible ones. Next, we encode the raw data of masked samples with a momentum graph encoder and subsequently employ the resulting representations to improve the predicted results through latent feature matching. Extensive experiments on seventeen datasets have demonstrated the effectiveness and robustness of RARE against state-of-the-art (SOTA) competitors across three downstream tasks. Our source code is available athttps://github.com/WxTu/RARE. Wenxuan Tu, Qing Liao 0001, Sihang Zhou 0001, Xin Peng 0010, Chuan Ma 0001, Zhe Liu 0001, Xinwang Liu 0002, Zhiping Cai, Kunlun He |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Deep Fusion Clustering Network With Reliable Structure PreservationabstractDeep clustering, which can elegantly exploit data representation to seek a partition of the samples, has attracted intensive attention. Recently, combining auto-encoder (AE) with graph neural networks (GNNs) has accomplished excellent performance by introducing structural information implied among data in clustering tasks. However, we observe that there are some limitations of most existing works: 1) in practical graph datasets, there exist some noisy or inaccurate connections among nodes, which would confuse network learning and cause biased representations, thus leading to unsatisfied clustering performance; 2) lacking dynamic information fusion module to carefully combine and refine the node attributes and the graph structural information to learn more consistent representations; and 3) failing to exploit the two separated views' information for generating a more robust target distribution. To solve these problems, we propose a novel method termed deep fusion clustering network with reliable structure preservation (DFCN-RSP). Specifically, the random walk mechanism is introduced to boost the reliability of the original graph structure by measuring localized structure similarities among nodes. It can simultaneously filter out noisy connections and supplement reliable connections in the original graph. Moreover, we provide a transformer-based graph auto-encoder (TGAE) that can use a self-attention mechanism with the localized structure similarity information to fine-tune the fused topology structure among nodes layer by layer. Furthermore, we provide a dynamic cross-modality fusion strategy to combine the representations learned from both TGAE and AE. Also, we design a triplet self-supervision strategy and a target distribution generation measure to explore the cross-modality information. The experimental results on five public benchmark datasets reflect that DFCN-RSP is more competitive than the state-of-the-art deep clustering algorithms. The corresponding code is available at https://github.com/gongleii/DFCN-RSP. Lei Gong 0008, Wenxuan Tu, Sihang Zhou 0001, Long Zhao 0002, Zhe Liu 0001, Xinwang Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Simple Contrastive Graph ClusteringabstractContrastive learning has recently attracted plenty of attention in deep graph clustering due to its promising performance. However, complicated data augmentations and time-consuming graph convolutional operations undermine the efficiency of these methods. To solve this problem, we propose a simple contrastive graph clustering (SCGC) algorithm to improve the existing methods from the perspectives of network architecture, data augmentation, and objective function. As to the architecture, our network includes two main parts, that is, preprocessing and network backbone. A simple low-pass denoising operation conducts neighbor information aggregation as an independent preprocessing, and only two multilayer perceptrons (MLPs) are included as the backbone. For data augmentation, instead of introducing complex operations over graphs, we construct two augmented views of the same vertex by designing parameter unshared Siamese encoders and perturbing the node embeddings directly. Finally, as to the objective function, to further improve the clustering performance, a novel cross-view structural consistency objective function is designed to enhance the discriminative capability of the learned network. Extensive experimental results on seven benchmark datasets validate our proposed algorithm's effectiveness and superiority. Significantly, our algorithm outperforms the recent contrastive deep clustering competitors with at least seven times speedup on average. The code of SCGC is released at SCGC. Besides, we share a collection of deep graph clustering, including papers, codes, and datasets at ADGC. Yue Liu 0008, Xihong Yang, Sihang Zhou 0001, Xinwang Liu 0002, Siwei Wang 0001, Ke Liang 0006, Wenxuan Tu, Liang Li 0041 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Hierarchically Contrastive Hard Sample Mining for Graph Self-Supervised PretrainingabstractContrastive learning has recently emerged as a powerful technique for graph self-supervised pretraining (GSP). By maximizing the mutual information (MI) between a positive sample pair, the network is forced to extract discriminative information from graphs to generate high-quality sample representations. However, we observe that, in the process of MI maximization (Infomax), the existing contrastive GSP algorithms suffer from at least one of the following problems: 1) treat all samples equally during optimization and 2) fall into a single contrasting pattern within the graph. Consequently, the vast number of well-categorized samples overwhelms the representation learning process, and limited information is accumulated, thus deteriorating the learning capability of the network. To solve these issues, in this article, by fusing the information from different views and conducting hard sample mining in a hierarchically contrastive manner, we propose a novel GSP algorithm called hierarchically contrastive hard sample mining (HCHSM). The hierarchical property of this algorithm is manifested in two aspects. First, according to the results of multilevel MI estimation in different views, the MI-based hard sample selection (MHSS) module keeps filtering the easy nodes and drives the network to focus more on hard nodes. Second, to collect more comprehensive information for hard sample learning, we introduce a hierarchically contrastive scheme to sequentially force the learned node representations to involve multilevel intrinsic graph features. In this way, as the contrastive granularity goes finer, the complementary information from different levels can be uniformly encoded to boost the discrimination of hard samples and enhance the quality of the learned graph embedding. Extensive experiments on seven benchmark datasets indicate that the HCHSM performs better than other competitors on node classification and node clustering tasks. The source code of HCHSM is available at https://github.com/WxTu/HCHSM. Wenxuan Tu, Shaohua Kevin Zhou, Xinwang Liu 0002, Chunpeng Ge 0001, Zhiping Cai, Yue Liu 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | HSNet: An Intelligent Hierarchical Semantic-Aware Network System for Real-Time Semantic SegmentationabstractSemantic segmentation, which aims to accurately identify each pixel, is a meaningful and challenging task. Recently, we witness a strong tendency to improve model efficiency in low-computing applications. However, most real-time methods ignore hierarchical features and context information to improve efficiency, leading to a decrease in the accuracy of semantic segmentation. To this end, we propose a novel system named hierarchical semantic-aware network (HSNet) to refine multilevel context information. HSNet mainly has the following two core modules: 1) hierarchical feature refinement module (HFRM) and 2) cross-scale pyramid fusion module (CPFM). By aggregating hierarchical feature maps, the proposed HFRM learns multilevel feature representation to recover spatial details. Afterward, the dual attention mechanism is developed to refine features from both channel and spatial levels, thereby alleviating the multilevel semantic gap. Meanwhile, the CPFM, which fuses local and global context information in a cross-scale manner, is proposed to enrich semantic information to improve accuracy. Furthermore, HSNet is carefully designed to improve the efficiency of the model by reusing shallow features and reducing channel capacity. Extensive experiments show that our method is effective and superior in segmentation accuracy and inference speed compared with state-of-the-art methods. Xin Peng 0010, Jieren Cheng, Xiangyan Tang, Ziqi Deng, Wenxuan Tu, Naixue Xiong |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2023 | Hard Sample Aware Network for Contrastive Deep Graph ClusteringabstractContrastive deep graph clustering, which aims to divide nodes into disjoint groups via contrastive mechanisms, is a challenging research spot. Among the recent works, hard sample mining-based algorithms have achieved great attention for their promising performance. However, we find that the existing hard sample mining methods have two problems as follows. 1) In the hardness measurement, the important structural information is overlooked for similarity calculation, degrading the representativeness of the selected hard negative samples. 2) Previous works merely focus on the hard negative sample pairs while neglecting the hard positive sample pairs. Nevertheless, samples within the same cluster but with low similarity should also be carefully learned. To solve the problems, we propose a novel contrastive deep graph clustering method dubbed Hard Sample Aware Network (HSAN) by introducing a comprehensive similarity measure criterion and a general dynamic sample weighing strategy. Concretely, in our algorithm, the similarities between samples are calculated by considering both the attribute embeddings and the structure embeddings, better revealing sample relationships and assisting hardness measurement. Moreover, under the guidance of the carefully collected high-confidence clustering information, our proposed weight modulating function will first recognize the positive and negative samples and then dynamically up-weight the hard sample pairs while down-weighting the easy ones. In this way, our method can mine not only the hard negative samples but also the hard positive sample, thus improving the discriminative capability of the samples further. Extensive experiments and analyses demonstrate the superiority and effectiveness of our proposed method. The source code of HSAN is shared at https://github.com/yueliu1999/HSAN and a collection (papers, codes and, datasets) of deep graph clustering is shared at https://github.com/yueliu1999/Awesome-Deep-Graph-Clustering on Github. Yue Liu 0008, Xihong Yang, Sihang Zhou 0001, Xinwang Liu 0002, Ke Liang 0006, Wenxuan Tu, Liang Li 0041, Jingcan Duan, Cancan Chen |
AAAI | 7 |
| 2023 | Cluster-Guided Contrastive Graph Clustering NetworkabstractBenefiting from the intrinsic supervision information exploitation capability, contrastive learning has achieved promising performance in the field of deep graph clustering recently. However, we observe that two drawbacks of the positive and negative sample construction mechanisms limit the performance of existing algorithms from further improvement. 1) The quality of positive samples heavily depends on the carefully designed data augmentations, while inappropriate data augmentations would easily lead to the semantic drift and indiscriminative positive samples. 2) The constructed negative samples are not reliable for ignoring important clustering information. To solve these problems, we propose a Cluster-guided Contrastive deep Graph Clustering network (CCGC) by mining the intrinsic supervision information in the high-confidence clustering results. Specifically, instead of conducting complex node or edge perturbation, we construct two views of the graph by designing special Siamese encoders whose weights are not shared between the sibling sub-networks. Then, guided by the high-confidence clustering information, we carefully select and construct the positive samples from the same high-confidence cluster in two views. Moreover, to construct semantic meaningful negative sample pairs, we regard the centers of different high-confidence clusters as negative samples, thus improving the discriminative capability and reliability of the constructed sample pairs. Lastly, we design an objective function to pull close the samples from the same cluster while pushing away those from other clusters by maximizing and minimizing the cross-view cosine similarity between positive and negative samples. Extensive experimental results on six datasets demonstrate the effectiveness of CCGC compared with the existing state-of-the-art algorithms. The code of CCGC is available at https://github.com/xihongyang1999/CCGC on Github. Xihong Yang, Yue Liu 0008, Sihang Zhou 0001, Siwei Wang 0001, Wenxuan Tu, Qun Zheng, Xinwang Liu 0002, Liming Fang 0001, En Zhu |
AAAI | 5 |
| 2023 | TMac: Temporal Multi-Modal Graph Learning for Acoustic Event ClassificationabstractAudiovisual data is everywhere in this digital age, which raises higher requirements for the deep learning models developed on them. To well handle the information of the multi-modal data is the key to a better audiovisual modal. We observe that these audiovisual data naturally have temporal attributes, such as the time information for each frame in the video. More concretely, such data is inherently multi-modal according to both audio and visual cues, which proceed in a strict chronological order. It indicates that temporal information is important in multi-modal acoustic event modeling for both intra- and inter-modal. However, existing methods deal with each modal feature independently and simply fuse them together, which neglects the mining of temporal relation and thus leads to sub-optimal performance. With this motivation, we propose a Temporal Multi-modal graph learning method for Acoustic event Classification, called TMac, by modeling such temporal information via graph learning techniques. In particular, we construct a temporal graph for each acoustic event, dividing its audio data and video data into multiple segments. Each segment can be considered as a node, and the temporal relationships between nodes can be considered as timestamps on their edges. In this case, we can smoothly capture the dynamic information in intra-modal and inter-modal. Several experiments are conducted to demonstrate TMac outperforms other SOTA models in performance. Our code is available at https://github.com/MGitHubL/TMac. Meng Liu 0014, Ke Liang 0006, Dayu Hu, Hao Yu 0017, Yue Liu 0008, Lingyuan Meng, Wenxuan Tu, Sihang Zhou 0001, Xinwang Liu 0002 |
ACM Multimedia | 7 |
| 2023 | Learn from Relational Correlations and Periodic Events for Temporal Knowledge Graph ReasoningabstractReasoning on temporal knowledge graphs (TKGR), aiming to infer missing events along the timeline, has been widely studied to alleviate incompleteness issues in TKG, which is composed of a series of KG snapshots at different timestamps. Two types of information, i.e., intra-snapshot structural information and inter-snapshot temporal interactions, mainly contribute to the learned representations for reasoning in previous models. However, these models fail to leverage (1) semantic correlations between relationships for the former information and (2) the periodic temporal patterns along the timeline for the latter one. Thus, such insufficient mining manners hinder expressive ability, leading to sub-optimal performances. To address these limitations, we propose a novel reasoning model, termed RPC, which sufficiently mines the information underlying the Relational correlations and Periodic patterns via two novel Correspondence units, i.e., relational correspondence unit (RCU) and periodic correspondence unit (PCU). Concretely, relational graph convolutional network (RGCN) and RCU are used to encode the intra-snapshot graph structural information for entities and relations, respectively. Besides, the gated recurrent units (GRU) and PCU are designed for sequential and periodic inter-snapshot temporal interactions, separately. Moreover, the model-agnostic time vectors are generated by time2vector encoders to guide the time-dependent decoder for fact scoring. Extensive experiments on six benchmark datasets show that RPC outperforms the state-of-the-art TKGR models, and also demonstrate the effectiveness of two novel strategies in our model. Ke Liang 0006, Lingyuan Meng, Meng Liu 0014, Yue Liu 0008, Wenxuan Tu, Siwei Wang 0001, Sihang Zhou 0001, Xinwang Liu 0002 |
SIGIR | 5 |
| 2023 | scDFC: A deep fusion clustering method for single-cell RNA-seq dataabstractClustering methods have been widely used in single-cell RNA-seq data for investigating tumor heterogeneity. Since traditional clustering methods fail to capture the high-dimension methods, deep clustering methods have drawn increasing attention these years due to their promising strengths on the task. However, existing methods consider either the attribute information of each cell or the structure information between different cells. In other words, they cannot sufficiently make use of all of this information simultaneously. To this end, we propose a novel single-cell deep fusion clustering model, which contains two modules, i.e. an attributed feature clustering module and a structure-attention feature clustering module. More concretely, two elegantly designed autoencoders are built to handle both features regardless of their data types. Experiments have demonstrated the validity of the proposed approach, showing that it is efficient to fuse attributes, structure, and attention information on single-cell RNA-seq data. This work will be further beneficial for investigating cell subpopulations and tumor microenvironment. The Python implementation of our work is now freely available at https://github.com/DayuHuu/scDFC. Dayu Hu, Ke Liang 0006, Sihang Zhou 0001, Wenxuan Tu, Meng Liu 0014, Xinwang Liu 0002 |
Briefings Bioinform. | 4 |
| 2022 | Deep Graph Clustering via Dual Correlation ReductionabstractDeep graph clustering, which aims to reveal the underlying graph structure and divide the nodes into different groups, has attracted intensive attention in recent years. However, we observe that, in the process of node encoding, existing methods suffer from representation collapse which tends to map all data into the same representation. Consequently, the discriminative capability of the node representation is limited, leading to unsatisfied clustering performance. To address this issue, we propose a novel self-supervised deep graph clustering method termed Dual Correlation Reduction Network (DCRN) by reducing information correlation in a dual manner. Specifically, in our method, we first design a siamese network to encode samples. Then by forcing the cross-view sample correlation matrix and cross-view feature correlation matrix to approximate two identity matrices, respectively, we reduce the information correlation in the dual-level, thus improving the discriminative capability of the resulting features. Moreover, in order to alleviate representation collapse caused by over-smoothing in GCN, we introduce a propagation regularization term to enable the network to gain long-distance information with the shallow network structure. Extensive experimental results on six benchmark datasets demonstrate the effectiveness of the proposed DCRN against the existing state-of-the-art methods. The code of DCRN is available at https://github.com/yueliu1999/DCRN and a collection (papers, codes and, datasets) of deep graph clustering is shared at https://github.com/yueliu1999/Awesome-Deep-Graph-Clustering on Github. Yue Liu 0008, Wenxuan Tu, Sihang Zhou 0001, Xinwang Liu 0002, Linxuan Song, Xihong Yang, En Zhu |
AAAI | 2 |
| 2022 | Highly-efficient Incomplete Largescale Multiview Clustering with Consensus Bipartite GraphabstractMultiview clustering has received increasing attention due to its effectiveness in fusing complementary information without manual annotations. Most previous methods hold the assumption that each instance appears in all views. However, it is not uncommon to see that some views may contain some missing instances, which gives rise to incomplete multi-view clustering (IMVC) in literature. Although many IMVC methods have been recently proposed, they always encounter high complexity and expensive time expenditure from being applied into large-scale tasks. In this paper, we present a flexible highly-efficient incomplete large-scale multi-view clustering approach based on bipartite graph framework to solve these issues. Specifically, we formalize multi-view anchor learning and incomplete bipartite graph into a unified framework, which coordinates with each other to boost cluster performance. By introducing the flexible bipartite graph framework to handle IMVC for the first practice, our proposed method enjoys linear complexity respecting to instance numbers, which is more applicable for large-scale IMVC tasks. Comprehensive experimental results on various benchmark datasets demonstrate the effectiveness and efficiency of our proposed algorithm against other IMVC competitors. The code is available at11https://github.com/wangsiwei2010/CVPR22-IMVC-CBG. Siwei Wang 0001, Xinwang Liu 0002, Li Liu 0002, Wenxuan Tu, Xinzhong Zhu, Jiyuan Liu 0003, Sihang Zhou 0001, En Zhu |
CVPR | 4 |
| 2022 | Attributed Graph Clustering with Dual Redundancy ReductionabstractAttributed graph clustering is a basic yet essential method for graph data exploration. Recent efforts over graph contrastive learning have achieved impressive clustering performance. However, we observe that the commonly adopted InfoMax operation tends to capture redundant information, limiting the downstream clustering performance. To this end, we develop a novel method termed attributed graph clustering with dual redundancy reduction (AGC-DRR) to reduce the information redundancy in both input space and latent feature space. Specifically, for the input space redundancy reduction, we introduce an adversarial learning mechanism to adaptively learn a redundant edge-dropping matrix to ensure the diversity of the compared sample pairs. To reduce the redundancy in the latent space, we force the correlation matrix of the cross-augmentation sample embedding to approximate an identity matrix. Consequently, the learned network is forced to be robust against perturbation while discriminative against different samples. Extensive experiments have demonstrated that AGC-DRR outperforms the state-of-the-art clustering methods on most of our benchmarks. The corresponding code is available at https://github.com/gongleii/AGC-DRR. Lei Gong 0008, Sihang Zhou 0001, Wenxuan Tu, Xinwang Liu 0002 |
IJCAI | 3 |
| 2022 | Initializing Then Refining: A Simple Graph Attribute Imputation NetworkabstractRepresentation learning on the attribute-missing graphs, whose connection information is complete while the attribute information of some nodes is missing, is an important yet challenging task. To impute the missing attributes, existing methods isolate the learning processes of attribute and structure information embeddings, and force both resultant representations to align with a common in-discriminative normal distribution, leading to inaccurate imputation. To tackle these issues, we propose a novel graph-oriented imputation framework called initializing then refining (ITR), where we first employ the structure information for initial imputation, and then leverage observed attribute and structure information to adaptively refine the imputed latent variables. Specifically, we first adopt the structure embeddings of attribute-missing samples as the embedding initialization, and then refine these initial values by aggregating the reliable and informative embeddings of attribute-observed samples according to the affinity structure. Specially, in our refining process, the affinity structure is adaptively updated through iterations by calculating the sample-wise correlations upon the recomposed embeddings. Extensive experiments on four benchmark datasets verify the superiority of ITR against state-of-the-art methods. Wenxuan Tu, Sihang Zhou 0001, Xinwang Liu 0002, Yue Liu 0008, Zhiping Cai, En Zhu, Changwang Zhang, Jieren Cheng |
IJCAI | 1 |
| 2022 | Align then Fusion: Generalized Large-scale Multi-view Clustering with Anchor Matching CorrespondencesabstractMulti-view anchor graph clustering selects representative anchors to avoid full pair-wise similarities and therefore reduce the complexity of graph methods. Although widely applied in large-scale applications, existing approaches do not pay sufficient attention to establishing correct correspondences between the anchor sets across views. To be specific, anchor graphs obtained from different views are not aligned column-wisely. Such an Anchor-Unaligned Problem (AUP) would cause inaccurate graph fusion and degrade the clustering performance. Under multi-view scenarios, generating correct correspondences could be extremely difficult since anchors are not consistent in feature dimensions. To solve this challenging issue, we propose the first study of the generalized and flexible anchor graph fusion framework termed Fast Multi-View Anchor-Correspondence Clustering (FMVACC). Specifically, we show how to find anchor correspondence with both feature and structure information, after which anchor graph fusion is performed column-wisely. Moreover, we theoretically show the connection between FMVACC and existing multi-view late fusion and partial view-aligned clustering, which further demonstrates our generality. Extensive experiments on seven benchmark datasets demonstrate the effectiveness and efficiency of our proposed method. Moreover, the proposed alignment module also shows significant performance improvement applying to existing multi-view anchor graph competitors indicating the importance of anchor alignment. Our code is available at \url{https://github.com/wangsiwei2010/NeurIPS22-FMVACC}. Siwei Wang 0001, Xinwang Liu 0002, Suyuan Liu, Jiaqi Jin, Wenxuan Tu, Xinzhong Zhu, En Zhu |
NeurIPS | 5 |
| 2022 | MIFNet: A lightweight multiscale information fusion networkabstractSemantic segmentation technique plays a crucial role in Internet of Things applications, such as industrial robotics and self-driving. Recently deep learning approaches have boosted semantic segmentation accuracy greatly. However, their comprehensive performance in terms of accuracy and efficiency is still far from satisfactory. We observe that (1) accuracy-oriented methods rely on numerous convolution layers and sophisticated architectures, which result in heavy computational complexity and usually take a long time for inference; (2) efficiency-oriented methods fail to capture the multiscale context information for discriminative representations during the feature fusion process, thus leading to suboptimal performance. Previous semantic segmentation approaches fail to address these two challenges simultaneously. To tackle the dilemma of precise segmentation and efficient inference, we propose a novel lightweight Multiscale Information Fusion Network (MIFNet). Specifically, the proposed MIFNet mainly consists of two core components, that is, Pyramid Refinement Connection Module (PRCM) and Lightweight Information Fusion Module (LIFM). The PRCM exploits skip learning to establish dependency between different stages. Meanwhile, the pyramid attention mechanism (PAM) in PRCM, which adjusts the weight of hybrid pyramid attention vector to refine spatial features of low-level, is developed to alleviate the semantic gap. Moreover, the LIFM is designed to detect objects at multiple scales from the global-local perspective. In LIFM, the proposed multiscale dense concatenation (MDC) adopts various dilated convolution to extract multiscale local context information. Extensive experimental results on benchmarks data sets demonstrate the significantly better performance of the proposed MIFNet compared with most existing state-of-the-art methods. Jieren Cheng, Xin Peng 0010, Xiangyan Tang, Wenxuan Tu, Wenhang Xu |
Int. J. Intell. Syst. | 4 |
| 2022 | Spare simple MKKM with semi-infinite linear program optimizationabstractMultiple kernel clustering (MKC) optimally combines a group of predefined kernel matrices to improve clustering performance. Although demonstrating promising performance in various applications, most of existing approaches adopt the min–min formulation, which could be sensitive to perturbation with adversarial samples. Moreover, existing MKC algorithms often involve several hypermeters preventing them into further real applications. To address these issues, we propose a parameter-free effective sparse simple multiple kernel k-means algorithm with max–min optimization formulation in this paper. To be specific, we propose to optimize the widely used unsupervised kernel alignment criterion by minimizing the kernel coefficient and maximizing the clustering partition matrix. Unlike traditional min–min formulation, the max–min kernel alignment is robust to adversarial sample perturbation and free of hyper-parameters. An optimization method based on semi-infinite linear program is designed to solve the complicated optimization problem. Extensive experiments on six multiple kernel benchmark data sets demonstrate the effectiveness of the proposed method. Miaomiao Li 0001, Wenxuan Tu, Jiyuan Liu 0003, Jiahao Ying |
Int. J. Intell. Syst. | 3 |
| 2021 | Deep Fusion Clustering NetworkabstractDeep clustering is a fundamental yet challenging task for data analysis. Recently we witness a strong tendency of combining autoencoder and graph neural networks to exploit structure information for clustering performance enhancement. However, we observe that existing literature 1) lacks a dynamic fusion mechanism to selectively integrate and refine the information of graph structure and node attributes for consensus representation learning; 2) fails to extract information from both sides for robust target distribution (i.e., “groundtruth” soft labels) generation. To tackle the above issues, we propose a Deep Fusion Clustering Network (DFCN). Specifically, in our network, an interdependency learning-based Structure and Attribute Information Fusion (SAIF) module is proposed to explicitly merge the representations learned by an autoencoder and a graph autoencoder for consensus representation learning. Also, a reliable target distribution generation measure and a triplet self-supervision strategy, which facilitate cross-modality information exploitation, are designed for network training. Extensive experiments on six benchmark datasets have demonstrated that the proposed DFCN consistently outperforms the state-of-the-art deep clustering methods. Wenxuan Tu, Sihang Zhou 0001, Xinwang Liu 0002, Xifeng Guo 0001, Zhiping Cai, En Zhu, Jieren Cheng |
AAAI | 1 |
| 2021 | One Pass Late Fusion Multi-view ClusteringabstractExisting late fusion multi-view clustering (LFMVC) optimally integrates a group of pre-specified base partition matrices to learn a consensus one. It is then taken as the input of the widely used k-means to generate the cluster labels. As observed, the learning of the consensus partition matrix and the generation of cluster labels are separately done. These two procedures lack necessary negotiation and can not best serve for each other, which may adversely affect the clustering performance. To address this issue, we propose to unify the aforementioned two learning procedures into a single optimization, in which the consensus partition matrix can better serve for the generation of cluster labels, and the latter is able to guide the learning of the former. To optimize the resultant optimization problem, we develop a four-step alternate algorithm with proved convergence. We theoretically analyze the clustering generalization error of the proposed algorithm on unseen data. Comprehensive experiments on multiple benchmark datasets demonstrate the superiority of our algorithm in terms of both clustering accuracy and computational efficiency. It is expected that the simplicity and effectiveness of our algorithm will make it a good option to be considered for practical multi-view clustering applications. Xinwang Liu 0002, Li Liu 0002, Qing Liao 0001, Siwei Wang 0001, Yi Zhang 0104, Wenxuan Tu, Chang Tang, Jiyuan Liu 0003, En Zhu |
ICML | 6 |
| 2021 | Self-Representation Subspace Clustering for Incomplete Multi-view DataabstractIncomplete multi-view clustering is an important research topic in multimedia where partial data entries of one or more views are missing. Current subspace clustering approaches mostly employ matrix factorization on the observed feature matrices to address this issue. Meanwhile, self-representation technique is left unexplored, since it explicitly relies on full data entries to construct the coefficient matrix, which is contradictory to the incomplete data setting. However, it is widely observed that self-representation subspace method enjoys a better clustering performance over the factorization based one. Therefore, we adapt it to incomplete data by jointly performing data imputation and self-representation learning. To the best of our knowledge, this is the first attempt in incomplete multi-view clustering literature. Besides, the proposed method is carefully compared with current advances in experiment with respect to different missing ratios, verifying its effectiveness. Jiyuan Liu 0003, Xinwang Liu 0002, Yi Zhang 0104, Pei Zhang 0008, Wenxuan Tu, Siwei Wang 0001, Sihang Zhou 0001, Weixuan Liang, Siqi Wang 0001, Yuexiang Yang |
ACM Multimedia | 5 |
| 2021 | Scalable Multi-view Subspace Clustering with Unified AnchorsabstractMulti-view subspace clustering has received widespread attention to effectively fuse multi-view information among multimedia applications. Considering that most existing approaches' cubic time complexity makes it challenging to apply to realistic large-scale scenarios, some researchers have addressed this challenge by sampling anchor points to capture distributions in different views. However, the separation of the heuristic sampling and clustering process leads to weak discriminate anchor points. Moreover, the complementary multi-view information has not been well utilized since the graphs are constructed independently by the anchors from the corresponding views. To address these issues, we propose a Scalable Multi-view Subspace Clustering with Unified Anchors (SMVSC). To be specific, we combine anchor learning and graph construction into a unified optimization framework. Therefore, the learned anchors can represent the actual latent data distribution more accurately, leading to a more discriminative clustering structure. Most importantly, the linear time complexity of our proposed algorithm allows the multi-view subspace clustering approach to be applied to large-scale data. Then, we design a four-step alternative optimization algorithm with proven convergence. Compared with state-of-the-art multi-view subspace clustering methods and large-scale oriented methods, the experimental results on several datasets demonstrate that our SMVSC method achieves comparable or better clustering performance much more efficiently. The code of SMVSC is available at https://github.com/Jeaninezpp/SMVSC. Mengjing Sun, Pei Zhang 0008, Siwei Wang 0001, Sihang Zhou 0001, Wenxuan Tu, Xinwang Liu 0002, En Zhu |
ACM Multimedia | 5 |
| 2021 | DFFNet: An IoT-perceptive dual feature fusion network for general real-time semantic segmentation
Xiangyan Tang, Wenxuan Tu, Keqiu Li, Jieren Cheng |
Inf. Sci. | 2 |
| 2020 | Hybrid Dilated Convolution Network Using Attentive Kernels for Real-Time Semantic Segmentation
Jiankai He, Bin Jiang 0006, Chao Yang 0015, Wenxuan Tu |
PRCV (1) | 4 |
| 2020 | Context-Integrated and Feature-Refined Network for Lightweight Object ParsingabstractSemantic segmentation for lightweight object parsing is a very challenging task, because both accuracy and efficiency (e.g., execution speed, memory footprint or computational complexity) should all be taken into account. However, most previous works pay too much attention to one-sided perspective, either accuracy or speed, and ignore others, which poses a great limitation to actual demands of intelligent devices. To tackle this dilemma, we propose a novel lightweight architecture named Context-Integrated and Feature-Refined Network (CIFReNet). The core components of CIFReNet are the Long-skip Refinement Module (LRM) and the Multi-scale Context Integration Module (MCIM). The LRM is designed to ease the propagation of spatial information between low-level and high-level stages. Furthermore, channel attention mechanism is introduced into the process of long-skip learning to boost the quality of low-level feature refinement. Meanwhile, the MCIM consists of three cascaded Dense Semantic Pyramid (DSP) blocks with image-level features, which is presented to encode multiple context information and enlarge the field of view. Specifically, the proposed DSP block exploits a dense feature sampling strategy to enhance the information representations without significantly increasing the computation cost. Comprehensive experiments are conducted on three benchmark datasets for object parsing including Cityscapes, CamVid, and Helen. As indicated, the proposed method reaches a better trade-off between accuracy and efficiency compared with the other state-of-the-art methods. Bin Jiang 0006, Wenxuan Tu, Chao Yang 0015, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 2 |