VLDB 2026 Research / reviewers in the wild / expert
Xinzhong Zhu
dblp:25/2001
· DBLP profile ↗
81ranked-venue papers
4as first author
58since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 3 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 1 first-author · 21 since 2021Databases, data management, data science and information retrieval · 11 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Prototype region calibration guided federated domain generalization
Wenjie Yao, Suxia Zhu, Libao Zhang, Guanglu Sun, Xinzhong Zhu |
Inf. Process. Manag. | 7 |
| 2026 | RAC-DMVC: Reliability-Aware Contrastive Deep Multi-View Clustering Under Multi-Source NoiseabstractMulti-view clustering (MVC), which aims to separate the multi-view data into distinct clusters in an unsupervised manner, is a fundamental yet challenging task. To enhance its applicability in real-world scenarios, this paper addresses a more challenging task: MVC under multi-source noises, including missing noise and observation noise. To this end, we propose a novel framework, Reliability-Aware Contrastive Deep Multi-View Clustering (RAC-DMVC), which constructs a reliability graph to guide robust representation learning under noisy environments. Specifically, to address observation noise, we introduce a cross-view reconstruction to enhances robustness at the data level, and a reliability-aware noise contrastive learning to mitigates bias in positive and negative pairs selection caused by noisy representations. To handle missing noise, we design a dual-attention imputation to capture shared information across views while preserving view-specific features. In addition, a self-supervised cluster distillation module further refines the learned representations and improves the clustering performance. Extensive experiments on five benchmark datasets demonstrate that RAC-DMVC outperforms SOTA methods on multiple evaluation metrics and maintains excellent performance under varying ratios of noise. Shihao Dong, Yue Liu 0008, Xiaotong Zhou, Yuhui Zheng, Xinzhong Zhu |
AAAI | 6 |
| 2026 | Point Cloud Quantization Through Multimodal Prompting for 3D UnderstandingabstractVector quantization has emerged as a powerful tool in large-scale multimodal models, unifying heterogeneous representations through discrete token encoding. However, its effectiveness hinges on robust codebook design. Current prototype-based approaches relying on trainable vectors or clustered centroids fall short in representativeness and interpretability, even as multimodal alignment demonstrates its promise in vision-language models. To address these limitations, we propose a simple multimodal prompting-driven quantization framework for point cloud analysis. Our methodology is built upon two core insights: 1) Text embeddings from pre-trained models inherently encode visual semantics through many-to-one contrastive alignment, naturally serving as robust prototype priors; and 2) Multimodal prompts enable adaptive refinement of these prototypes, effectively mitigating vision-language semantic gaps. The framework introduces a dual-constrained quantization space, enforced by compactness and separation regularization, which seamlessly integrates visual and prototype features, resulting in hybrid representations that jointly encode geometric and semantic information. Furthermore, we employ Gumbel-Softmax relaxation to achieve differentiable discretization while maintaining quantization sparsity. Extensive experiments on the ModelNet40 and ScanObjectNN datasets clearly demonstrate the superior effectiveness of the proposed method. Wencheng Zhu, Xinzhong Zhu, Pengfei Zhu 0001 |
AAAI | 4 |
| 2026 | ExtendAttack: Attacking Servers of LRMs via Extending ReasoningabstractLarge Reasoning Models (LRMs) have demonstrated promising performance in complex tasks. However, the resource-consuming reasoning processes may be exploited by attackers to maliciously occupy the resources of the servers, leading to a crash, like the DDoS attack in cyber. To this end, we propose a novel attack method on LRMs termed ExtendAttack to maliciously occupy the resources of servers by stealthily extending the reasoning processes of LRMs. Concretely, we systematically obfuscate characters within a benign prompt, transforming them into a complex, poly-base ASCII representation. This compels the model to perform a series of computationally intensive decoding sub-tasks that are deeply embedded within the semantic structure of the query itself. Extensive experiments demonstrate the effectiveness of our proposed ExtendAttack. Remarkably, it significantly increases response length and latency, with the former increasing by over 2.7 times for the o3 model on the HumanEval benchmark. Besides, it preserves the original meaning of the query and achieves comparable answer accuracy, showing the stealthiness. Zhenhao Zhu, Yue Liu 0008, Yingwei Ma, Hongcheng Gao, Nuo Chen 0002, Yanpei Guo, Wenjie Qu 0001, Zifeng Kang, Xinzhong Zhu, Jiaheng Zhang |
AAAI | 11 |
| 2026 | IPeDet: An end-to-end fine-grained feature aggregation network for UAV infrared pedestrian detection
Yi Li 0068, Xinzhong Zhu |
Adv. Eng. Informatics | 3 |
| 2026 | KSCNet: Exploring KAN and state space model collaboration network for small object detection from UAV imagery
Yiming Sun 0003, Pengfei Zhu 0001, Xinzhong Zhu |
Expert Syst. Appl. | 6 |
| 2026 | Revocable Policy Hiding Bilateral Access Control Scheme With Equality Test for IIoT EnvironmentabstractWith the development of the Industrial Internet of Things (IIoT), the amount of data generated by industrial manufacturing equipment is growing rapidly, creating a critical need for secure and efficient cloud-based data sharing. While bilateral access control enables both data senders and receivers to define their own fine-grained access policies, existing schemes transmit these policies in plaintext, exposing sensitive operational metadata and creating security vulnerabilities. Furthermore, they lack essential functionalities for practical IIoT deployment, including efficient duplicate data detection, a robust user revocation mechanism, and computationally lightweight operations suitable for resource-constrained devices. To address these challenges, this paper proposes a novel policy-hiding, lightweight, and revocable bilateral access control scheme with equality test (RBAC-ET) for cloud-enabled IIoT systems. RBAC-ET is the first scheme to integrally combine five critical capabilities: first, fine-grained bilateral access control using Linear Secret Sharing Schemes (LSSS) supporting AND/OR logical operations. Second, policy confidentiality is achieved by encrypting access policies within the ciphertext to prevent the leakage of sensitive relationships. Third, the equality test functionality enables cloud servers to identify duplicate ciphertexts for efficient storage and processing without requiring decryption. Fourth, a proposed lightweight design is based on elliptic curve scalar multiplication operations instead of the more resource-intensive bilinear pairing operations. Fifth, a robust user revocation mechanism that ensures both forward and backward secrecy in dynamic industrial environments. Theoretical analysis and experimental results demonstrate that RBAC-ET offers superior security and functionality while maintaining computational efficiency comparable to that of state-of-the-art schemes, making it a scalable and practical solution for modern IIoT data-sharing applications. Junaid Hassan, Zhen Qin 0002, Muhammad Arslan Rauf, Muhammad Umar Aftab, Negalign Wake Hundera, Xinzhong Zhu |
IEEE Internet Things J. | 6 |
| 2026 | Structure adversarial augmented graph anomaly detection via multi-view contrastive learning
Ruidong Wang 0001, Yue Liu 0008, Xinzhong Zhu |
Knowl. Based Syst. | 5 |
| 2026 | Semidefinite program-inspired continuous relaxation robust multi-view clustering for large-scale data
Ziying Wang, Hechuan Lin, Xinzhong Zhu |
Knowl. Based Syst. | 6 |
| 2026 | Balanced Multi-View ClusteringabstractMulti-view clustering (MvC) aims to integrate information from different views to enhance the capability of the model in capturing the underlying data structures. The widely used joint training paradigm in MvC potentially does not fully leverage the multi-view information, due to the imbalanced and under-optimized view-specific features caused by the uniform learning objective for all views. For instance, particular views with more discriminative information could dominate the learning process in the joint training paradigm, leading to other views being under-optimized. To alleviate this issue, we first analyze the imbalanced phenomenon in the joint-training paradigm of multi-view clustering from the perspective of gradient descent for each view-specific feature extractor. Then, we propose a novel balanced multi-view clustering (BMvC) method, which introduces a view-specific contrastive regularization (VCR) to modulate the optimization of each view. Concretely, VCR preserves the sample similarities captured from the joint features and view-specific ones into the clustering distributions corresponding to view-specific features to enhance the learning process of view-specific feature extractors. Additionally, an analysis is provided to illustrate that VCR adaptively modulates the magnitudes of gradients for updating the parameters of view-specific feature extractors to achieve a balanced multi-view learning procedure. In such a manner, BMvC achieves a better trade-off between the exploitation of view-specific patterns and the exploration of view-invariance patterns to fully learn the multi-view information for the clustering task. Finally, a set of experiments are conducted to verify the superiority of the proposed method compared with state-of-the-art approaches both on eight benchmark MvC datasets and two spatially resolved transcriptomics datasets. Zhenglai Li, Jun Wang 0118, Chang Tang, Xinzhong Zhu, Wei Zhang 0049, Xinwang Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Efficient one-pass incomplete multi-view clustering via anchor alignment and high-order correlation exploitation
Tingzhou Yan, Xinzhong Zhu, Chang Tang |
Pattern Recognit. | 4 |
| 2026 | PyMIF-Net: Pyramid Mutual Information Fusion Network for multimodal skin lesion classification
Xinzhong Zhu |
Pattern Recognit. | 9 |
| 2026 | HCA-Net: Hierarchical Contextual Attention Network for Lightweight and Accurate Polyp SegmentationabstractEarly detection of colorectal polyps is crucial for clinical screening and cancer prevention, where accurate and efficient automatic segmentation plays a pivotal role. However, colonoscopy images often suffer from low contrast, blurred boundaries, and scale variations, making segmentation challenging. Existing encoder-decoder networks (e.g., U-Net) suffer from asymmetric supervision and feature redundancy, which in turn lead to semantic inconsistency and loss of fine details. While deeper or hybrid designs alleviate these issues, their high complexity and computational burden limit feasibility in real-time clinical practice. To address these challenges, we propose a lightweight segmentation framework, Hierarchical Contextual Attention Network (HCA-Net), consisting of the Redundancy-Suppressed Dual-Path Downsampling (RS-DPD) module and the Boundary-Aware Semantic Alignment Upsampling (BA-SAU) module, applied to the encoder and decoder, respectively. RS-DPD suppresses redundancy while preserving fine-grained details through a dual-path design, whereas BA-SAU leverages cross-layer contextual attention to enforce semantic consistency and enhance boundary sensitivity. Both modules are built upon our proposed Hierarchical Contextual Attention (HCA) mechanism, which combines convolutional projection with pooling-based compression to achieve efficient global modeling and accurate local boundary restoration. In addition, a composite boundary-aware loss function is designed to improve pixel-level accuracy, structural consistency, and robustness in low-contrast and boundary-ambiguous regions. Extensive experiments on public colorectal polyp datasets demonstrate that HCA-Net achieves state-of-the-art (SOTA) segmentation accuracy with significantly improved efficiency, while maintaining robustness under low-contrast and blurred-boundary conditions. Xinzhong Zhu, Huiling Chen 0001, Yun Liu 0011, Chang Tang, Miaomiao Li 0001, Shanfu Lu |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | RemDet: Rethinking Efficient Model Design for UAV Object DetectionabstractObject detection in Unmanned Aerial Vehicle (UAV) images has emerged as a focal area of research, which presents two significant challenges: i) objects are typically small and dense within vast images; ii) computational resource constraints render most models unsuitable for real-time deployment. Current real-time object detectors are not optimized for UAV images, and complex methods designed for small object detection often lack real-time capabilities. To address these challenges, we propose a novel detector, RemDet (Reparameter efficient multiplication Detector). Our contributions are as follows: 1) Rethinking the challenges of existing detectors for small and dense UAV images, and proposing information loss as a design guideline for efficient models. 2) We introduce the ChannelC2f module to enhance small object detection performance, demonstrating that high-dimensional representations can effectively mitigate information loss. 3) We design the GatedFFN module to provide not only strong performance but also low latency, effectively addressing the challenges of real-time detection. Our research reveals that GatedFFN, through the use of multiplication, is more cost-effective than feed-forward networks for high-dimensional representation. 4) We propose the CED module, which combines the advantages of ViT and CNN downsampling to effectively reduce information loss. It specifically enhances context information for small and dense objects. Extensive experiments on large UAV datasets, Visdrone and UAVDT, validate the real-time efficiency and superior performance of our methods. On the challenging UAV dataset VisDrone, our methods not only provided state-of-the-art results, improving detection by more than 3.4%, but also achieve 110 FPS on a single 4090. Xinzhong Zhu |
AAAI | 5 |
| 2025 | Mamba YOLO: A Simple Baseline for Object Detection with State Space ModelabstractDriven by the rapid development of deep learning technology, the YOLO series has set a new benchmark for real-time object detectors. Additionally, transformer-based structures have emerged as the most powerful solution in the field, greatly extending the model's receptive field and achieving significant performance improvements. However, this improvement comes at a cost, as the quadratic complexity of the self-attentive mechanism increases the computational burden of the model. To address this problem, we introduce a simple yet effective baseline approach called Mamba YOLO. Our contributions are as follows: 1) We propose that the ODMamba backbone introduce a State Space Model (SSM) with linear complexity to address the quadratic complexity of self-attention. Unlike the other Transformer-base and SSM-base method, ODMamba is simple to train without pretraining. 2) For real-time requirement, we designed the macro structure of ODMamba, determined the optimal stage ratio and scaling size. 3) We design the RG Block that employs a multi-branch structure to model the channel dimensions, which addresses the possible limitations of SSM in sequence modeling, such as insufficient receptive fields and weak image localization. This design captures localized image dependencies more accurately and significantly. Extensive experiments on the publicly available COCO benchmark dataset show that Mamba YOLO achieves state-of-the-art performance compared to previous methods. Specifically, a tiny version of Mamba YOLO achieves a 7.5% improvement in mAP on a single 4090 GPU with an inference time of 1.5 ms. Chen Li 0025, Xinzhong Zhu |
AAAI | 4 |
| 2025 | Multi-DAT: Dynamic Job Task Scheduling Method Based on Multi-Agent Reinforcement LearningabstractThe scheduling of aircraft support tasks requires efficient planning based on the available resources in each sup-port position. The many-to-many characteristics of tasks and the dynamic nature of scheduling environments place high demands on the real-time responsiveness of algorithms. Additionally, the complexity inherent in task scheduling and the need for flexible sequential processing further complicate decision-making. Existing methods often struggle to achieve both fast response times and effective task message capture. To address these challenges, we propose a novel scheduling method called Multi-DAT, which is designed to optimize dynamic task allocation using a multi-agent reinforcement learning algorithm. We improve on the traditional QMIX algorithm to select the shortest duration tasks based on priority, combined with a newly designed long-term reward function, which integrates long-term historical actions into the scheduling algorithm. Experimental results demonstrate that our proposed method outperforms the traditional rule-based method and seven other multi-agent reinforcement learning-based scheduling algorithms in terms of scheduling performance. Linwei Yao, Kuan Li, Xinzhong Zhu |
CSCWD | 4 |
| 2025 | One-step Incomplete Multi-view Clustering based on Bipartite Graph LearningabstractAlthough previous graph-based multi-view clustering algorithms have made remarkable progress, most of them still face the following two limitations: 1. Many existing methods rely on k-means for the discretization of spectral embeddings, which cannot directly learn graphs with discrete cluster structures and require two steps for clustering results. 2. Practical applications may contain some missing instances, which require Incomplete Multi-View Clustering (IMVC) methods to hold them. In this paper, we propose a novel method named One-step Incomplete Multi-View Clustering based on Bipartite Graph Learning (OIMVC-BGL) which aims to solve the above problems. OIMVC-BGL first constructs bipartite graphs from all views with an anchor-based subspace learning method. Then, OIMVC-BGL fuses these graphs to obtain a consensus bipartite graph with an adaptive weight manner. Finally, OIMVC-BGL imposes a Laplacian rank constraint on the consensus bipartite graph to obtain the results directly. Experiments conducted on benchmark datasets verify the effectiveness of OIMVC-BGL. Hechuan Lin, Ziying Wang, Xinzhong Zhu |
ICASSP | 5 |
| 2025 | RestorMamba: An Enhanced Synergistic State Space Model for Image RestorationabstractIn this paper, we introduce an image inpainting method based on the State Space Model (SSM), named Restoration Mamba (RestorMamba). This approach incorporates effi-cient long-range dependency modeling within the network, which is particularly suited for the complexities of high-texture and high-resolution image restoration scenarios. To benefit from a broader context while maintaining global receptive fields, we have designed two pivotal modules: Skip Scan and Enhanced Synergistic Mamba (ESM) Block. Our experimental results demonstrate that RestorMamba achieves state-of-the-art performance in tasks such as image deraining and image denoising, encompassing Gaussian grayscale / color denoising and real image denoising. Chen Li 0025, Xinzhong Zhu |
ICASSP | 4 |
| 2025 | MambaInst: Lightweight State Space Model for Real-Time Instance SegmentationabstractIn this paper, we propose a lightweight and efficient state-space model-based instance segmentation network named MambaInst, which extracts deep semantic features through a LightSSM Block consisting of gating mechanisms and residual connectivity to model long-distance spatial dependencies with linear computational complexity. We design a novel downsampling method called FRDown to efficiently capture contextual information, thereby improving the network’s local information perception. With its excellent model architecture and simple training method, MambaInst-B achieves 40.8% in Mask mAP on a single 4090 GPU with an inference time of 2.28 ms on the COCO-seg. Our proposal demonstrates first proof of SSM’s effectiveness in real-time instance segmentation, setting a new performance benchmark for Mamba-based techniques in this particular application. Xinzhong Zhu |
ICASSP | 4 |
| 2025 | Pose-Enhanced 3D Rotary Embedding for Multi-View 3D Object Detection
Ke Sheng, Xinzhong Zhu |
ICIC (1) | 3 |
| 2025 | Multi-Q3IM: Job Dynamic Task Scheduling Based on Multi-agent Deep Reinforcement Learning in Resource-Constrained Environments
Linwei Yao, Kuan Li, Xinzhong Zhu |
ICIC (13) | 4 |
| 2025 | Center-Oriented Prototype Contrastive ClusteringabstractContrastive learning is widely used in clustering tasks due to its discriminative representation. However, the conflict problem between classes is difficult to solve effectively. Existing methods try to solve this problem through prototype contrast, but there is a deviation between the calculation of hard prototypes and the true cluster center. To address this problem, we propose a center-oriented prototype contrastive clustering framework, which consists of a soft prototype contrastive module and a dual consistency learning module. In short, the soft prototype contrastive module uses the probability that the sample belongs to the cluster center as a weight to calculate the prototype of each category, while avoiding inter-class conflicts and reducing prototype drift. The dual consistency learning module aligns different transformations of the same sample and the neighborhoods of different samples respectively, ensuring that the features have transformation-invariant semantic information and compact intra-cluster distribution, while providing reliable guarantees for the calculation of prototypes. Extensive experiments on five datasets show that the proposed method is effective compared to the SOTA. Our code is published on https://github.com/LouisDong95/CPCC. Shihao Dong, Xiaotong Zhou, Yuhui Zheng, Xinzhong Zhu |
ICME | 5 |
| 2025 | Rethinking Cross-Modality Fusion Mamba from a Frequency Domain PerspectiveabstractThis paper introduces a Cross-Modality Fusion method for multispectral object detection (MSOD), which is based on State Space Model (SSM) and frequency modeling. The approach benefits from efficient long-range dependency modeling by processing the perceptually significant frequency information within the complex coupling of RGB and infrared (IR) images, integrating global context from different modalities. Additionally, we propose a novel down-sampling method based on wavelet transform (WT), designed to reduce the spatial resolution of feature maps while filtering out as much redundant information as possible. Experimental results demonstrate that our method achieves state-of-the-art (SOTA) performance and faster inference in MSOD. Yun Liu 0011, Xinzhong Zhu |
ICME | 5 |
| 2025 | Task-Gated Multi-Expert Collaboration Network for Degraded Multi-Modal Image FusionabstractMulti-modal image fusion aims to integrate complementary information from different modalities to enhance perceptual capabilities in applications such as rescue and security. However, real-world imaging often suffers from degradation issues, such as noise, blur, and haze in visible imaging, as well as stripe noise in infrared imaging, which significantly degrades model performance. To address these challenges, we propose a task-gated multi-expert collaboration network (TG-ECNet) for degraded multi-modal image fusion. The core of our model lies in the task-aware gating and multi-expert collaborative framework, where the task-aware gating operates in two stages: degradation-aware gating dynamically allocates expert groups for restoration based on degradation types, and fusion-aware gating guides feature integration across modalities to balance information retention between fusion and restoration tasks. To achieve this, we design a two-stage training strategy that unifies the learning of restoration and fusion tasks. This strategy resolves the inherent conflict in information processing between the two tasks, enabling all-in-one multi-modal image restoration and fusion. Experimental results demonstrate that TG-ECNet significantly enhances fusion performance under diverse complex degradation conditions and improves robustness in downstream applications. The code is available at https://github.com/LeeX54946/TG-ECNet. Yiming Sun 0003, Pengfei Zhu 0001, Qinghua Hu, Dongwei Ren, Xinzhong Zhu |
ICML | 7 |
| 2025 | Multi-view Clustering via Bi-level Decoupling and Consistency LearningabstractMulti-view clustering has shown to be an effective method for analyzing underlying patterns in multi-view data. The performance of clustering can be improved by learning the consistency and complementarity between multi-view features, however, cluster-oriented representation learning is often over-looked. In this paper, we propose a novel Bi-level Decoupling and Consistency Learning framework (BDCL) to further explore the effective representation for multi-view data to enhance inter-cluster discriminability and intra-cluster compactness of features in multi-view clustering. Our framework comprises three modules: 1) The multi-view instance learning module aligns the consistent information while preserving the private features between views through reconstruction autoencoder and contrastive learning. 2) The bi-level decoupling of features and clusters enhances the discriminability of feature space and cluster space. 3) The consistency learning module treats the different views of the sample and their neighbors as positive pairs, learns the consistency of their clustering assignments, and further compresses the intra-cluster space. Experimental results on five benchmark datasets demonstrate the superiority of the proposed method compared with the SOTA methods. Our code is published on https://github.com/LouisDong95/BDCL. Shihao Dong, Yuhui Zheng, Xinzhong Zhu |
IJCNN | 4 |
| 2025 | High-Performance Discriminative Tracking with Spatio-Temporal Template Fusion
Xuedong He, Xinzhong Zhu, Hongbo Li 0001 |
ACM Multimedia | 3 |
| 2025 | Towards Single-Source Domain Generalized Object Detection via Causal Visual PromptsabstractSingle-source Domain Generalized Object Detection (SDGOD), as a cutting-edge research topic in computer vision, aims to enhance model generalization capability in unseen target domains through single-source domain training. Current mainstream approaches attempt to mitigate domain discrepancies via data augmentation techniques. However, due to domain shift and limited domain‑specific knowledge, models tend to fall into the pitfall of spurious correlations. This manifests as the model's over-reliance on simplistic classification features (e.g., color) rather than essential domain-invariant representations like object contours. To address this critical challenge, we propose the Cauvis (Causal Visual Prompts) method. First, we introduce a Cross-Attention Prompts module that mitigates bias from spurious features by integrating visual prompts with cross-attention. To address the inadequate domain knowledge coverage and spurious feature entanglement in visual prompts for single-domain generalization, we propose a dual-branch adapter that disentangles causal-spurious features while achieving domain adaptation via high-frequency feature extraction. Cauvis achieves state-of-the-art performance with 15.9–31.4\% gains over existing domain generalization methods on SDGOD datasets, while exhibiting significant robustness advantages in complex interference environments. Chen Li 0025, Changxin Gao, Xinzhong Zhu |
NeurIPS | 6 |
| 2025 | SOD-SRB: A General Super-Resolution Auxiliary Branch for Small Object Detection
Xinzhong Zhu, Yunzhong Si, Hongbo Li 0001 |
PRCV (18) | 3 |
| 2025 | Learning color prompt and position constraint for visual tracking
Xuedong He, Xinzhong Zhu, Yunliang Jiang |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Attention-modulated frequency-aware pooling via spatial guidance
Yunzhong Si, Xinzhong Zhu, Rihao Liu |
Neurocomputing | 3 |
| 2025 | SCSA: Exploring the synergistic effects between spatial and channel attention
Yunzhong Si, Xinzhong Zhu, Hongbo Li 0001 |
Neurocomputing | 3 |
| 2025 | Diversity Learning Guided Dual Graph Autoencoder for Unsupervised Hyperspectral Band SelectionabstractHyperspectral band selection, aimed at identifying key spectral bands from the original image, is crucial for reducing dimensionality and enhancing computational efficiency in hyperspectral image (HSI) analysis. Graph learning-based methods have attracted considerable attention due to their efficiency in representing structural correlations between bands and their powerful capability to extract features. However, existing methods have limitations in utilizing spatial relationships among bands and learning their discriminative characteristics. To address these limitations, we propose a Diversity Learning Guided Dual Graph Autoencoder (DLG-DGAE) for unsupervised hyperspectral band selection. In our framework, we integrate a Dual Graph Autoencoder (DGAE) module designed to extract information from both the spatial and spectral relationships among bands, thus fully capturing the structural similarity of the bands. Additionally, we introduce a Spectral Diversity Learning (SDL) strategy to reduce redundant information in the latent representation and enhance the discriminative properties of each band. In the final step, we proceed to cluster the fused latent embeddings. Within each cluster, we select the band exhibiting the highest information entropy as the representative band. Through extensive experimentation on three publicly available datasets, our results consistently indicate that the proposed method surpasses other state-of-the-art techniques. The code is available athttps://github.com/fengwe1/DLG-DGAE. Chang Tang, Xinwang Liu 0002, Junjun Jiang, Xianju Li, Xinzhong Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | FedRDA: Representation Deviation Alignment in Heterogeneous Federated LearningabstractFederatedlearning has garnered significant attention in the Internet of Things and healthcare applications due to its ability to train a shared global model across distributed clients. However, imbalanced data distribution leads to model discrepancies among clients. Most existing methods adopt implicit alignment strategies while overlooking explicit modeling of geometric and directional discrepancies in feature representations, which undermines local model optimization. To address this issue, we propose a method of representation deviation alignment in federated learning, which projects features onto the principal feature space to measure deviations between local and global feature representations explicitly. Specifically, Federated learning with Representation Deviation Alignment (FedRDA) employs a feature encoder to extract compact features and construct unbiased principal feature spaces for global and local models. Then, the residual projection in the feature space serves as a quantitative measure of the representation deviation, effectively capturing the latent direction differences between models. Besides, we introduce a representation consistency alignment strategy, which ensures that the distribution of local client features becomes more uniform within the global feature space. Extensive experiments on SVHN, CIFAR-10, CIFAR-100, Tiny-ImageNet, and GC10 demonstrate that FedRDA effectively reduces the classifier bias caused by representational differences. Wenjie Yao, Guanglu Sun, Suxia Zhu, Ruidong Wang 0001, Xinzhong Zhu, Xiguang Wei |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Dynamic Ensemble Framework for Imbalanced Data ClassificationabstractDynamic ensemble has significantly greater potential space to improve the classification of imbalanced data compared to static ensemble. However, dynamic ensemble schemes are far less successful than static ensemble methods in the imbalanced learning field. Through an in-depth analysis on the behavior characteristics of dynamic ensemble, we find that there are some important problems that need to be addressed to release the full potential of dynamic ensemble, including but not limited to, correcting the component classifiers’ bias towards the majority classes, increasing the proportions of the positive classifiers (i.e., the component classifiers making correct prediction) for difficult samples, and providing the accurate competence estimations on the hard-to-classify samples w.r.t the classifier pool. Inspired by these, we propose a Dynamic Ensemble Framework for imbalanced data classification (imDEF). imDEF first uses the data generation method OREM$\mathrm{_{G}}$to generate multiple artificial synthetic datasets, which have diverse class distributions by rebalancing the original imbalanced data. Based on each of such synthetic datasets, imDEF then utilizes a Classification Error-aware Self-Paced Sampling Ensemble (SPSE$\mathrm{_{CE}}$) method to gradually focus more on difficult samples, to create a low-biased classifier pool and increase the proportions of the positive classifiers for the difficult samples. Finally, imDEF constructs a referee system to achieve the competence estimations by leveraging an Ensemble Margin-aware Self-Paced Sampling Ensemble (SPSE$\mathrm{_{EM}}$) method. SPSE$\mathrm{_{EM}}$incrementally strengthens the learning of the hard-to-classify samples, so that the competent levels of component classifiers could be estimated accurately. Extensive experiments demonstrate the effectiveness of imDEF. The source codes have been made publicly available on GitHub. Tuanfei Zhu, Xingchen Hu 0001, Xinwang Liu 0002, En Zhu, Xinzhong Zhu |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | V-DDPM: MRI Rician Noise Removal Model Based on VST and DDPMabstractMagnetic resonance imaging (MRI) often contains Rician noise. Unlike additive Gaussian noise, the distribution of Rician noise is related to the data of the image, making it more difficult to remove. It has been shown that the variance stabilizing transformation (VST) can transform the Rician noise distribution into a variance stabilized Gaussian distribution. Utilizing this property, combined with the diffusion model, we proposed an algorithm for the removal of Rician noise from MRI. The algorithm first performs VST on the MRI containing Rician noise and then carries out diffusion denoising. A Gaussian noise sequence is added to the diffusion process; afterward, the diffusion process is reversed to provide different denoising levels through Markov chain modeling. Finally, the denoised image is obtained through the inverse of VST. Experimental results demonstrate the performance of our method in removing Rician noise from magnetic resonance images compared to DDPM denoising alone, while preserving detailed information better. Noise removal was also performed for MRI with a simple structure. Xinzhong Zhu, Negalign Wake Hundera |
ICASSP | 3 |
| 2024 | Scalable Multiple Kernel Clustering: Learning Clustering Structure from ExpectationabstractIn this paper, we derive an upper bound of the difference between a kernel matrix and its expectation under a mild assumption. Specifically, we assume that the true distribution of the training data is an unknown isotropic Gaussian distribution. When the kernel function is a Gaussian kernel, and the mean of each cluster is sufficiently separated, we find that the expectation of a kernel matrix can be close to a rank-$k$ matrix, where $k$ is the cluster number. Moreover, we prove that the normalized kernel matrix of the training set deviates (w.r.t. Frobenius norm) from its expectation in the order of $\widetilde{\mathcal{O}}(1/\sqrt{d})$, where $d$ is the dimension of samples. Based on the above theoretical results, we propose a novel multiple kernel clustering framework which attempts to learn the information of the expectation kernel matrices. First, we aim to minimize the distance between each base kernel and a rank-$k$ matrix, which is a proxy of the expectation kernel. Then, we fuse these rank-$k$ matrices into a consensus rank-$k$ matrix to find the clustering structure. Using an anchor-based method, the proposed framework is flexible with the sizes of input kernel matrices and able to handle large-scale datasets. We also provide the approximation guarantee by deriving two non-asymptotic bounds for the consensus kernel and clustering indicator matrices. Finally, we conduct extensive experiments to verify the clustering performance of the proposed method and the correctness of the proposed theoretical results. Weixuan Liang, En Zhu, Shengju Yu, Xinzhong Zhu, Xinwang Liu 0002 |
ICML | 5 |
| 2024 | Reliable Attribute-missing Multi-view Clustering with Instance-level and feature-level Cooperative ImputationabstractMulti-view clustering (MVC) constitutes a distinct approach to data mining within the field of machine learning. Due to limitations in the data collection process, missing attributes are frequently encountered. However, existing MVC methods primarily focus on missing instances, showing limited attention to missing attributes. A small number of studies employ the reconstruction of missing instances to address missing attributes, potentially overlooking the synergistic effects between the instance and feature spaces, which could lead to distorted imputation outcomes. Furthermore, current methods uniformly treat all missing attributes as zero values, thus failing to differentiate between real and technical zeroes, potentially resulting in data over-imputation. To mitigate these challenges, we introduce a novel Reliable Attribute-Missing Multi-View Clustering method (RAM-MVC). Specifically, feature reconstruction is utilized to address missing attributes, while similarity graphs are simultaneously constructed within the instance and feature spaces. By leveraging structural information from both spaces, RAM-MVC learns a high-quality feature reconstruction matrix during the joint optimization process. Additionally, we introduce a reliable imputation guidance module that distinguishes between real and technical attribute-missing events, enabling discriminative imputation. The proposed RAM-MVC method outperforms nine baseline methods, as evidenced by real-world experiments using single-cell multi-view data. Dayu Hu, Suyuan Liu, Jun Wang 0118, Junpu Zhang, Siwei Wang 0001, Xingchen Hu 0001, Xinzhong Zhu, Chang Tang, Xinwang Liu 0002 |
ACM Multimedia | 7 |
| 2024 | View Gap Matters: Cross-view Topology and Information Decoupling for Multi-view ClusteringabstractMulti-view clustering, a pivotal technology in multimedia research, aims to leverage complementary information from diverse perspectives to enhance clustering performance. The current multi-view clustering methods normally enforce the reduction of distances between any pair of views, overlooking the heterogeneity between views, thereby sacrificing the diverse and valuable insights inherent in multi-view data. In this paper, we propose a Tree-Based View-Gap Maintaining Multi-View Clustering (TGM-MVC) method. Our approach introduces a novel conceptualization of multiple views as a graph structure. In this structure, each view corresponds to a node, with the view gap, calculated by the cosine distance between views, acting as the edge. Through graph pruning, we derive the minimum spanning tree of the views, reflecting the neighbouring relationships among them. Specifically, we applied a share-specific learning framework, and generate view trees for both view-shared and view-specific information. Concerning shared information, we only narrow the distance between adjacent views, while for specific information, we maintain the view gap between neighboring views. Theoretical analysis highlights the risks of eliminating the view gap, and comprehensive experiments validate the efficacy of our proposed TGM-MVC method. Fangdi Wang, Jiaqi Jin, Zhibin Dong, Xihong Yang, Xinwang Liu 0002, Xinzhong Zhu, Siwei Wang 0001, Tianrui Liu 0001, En Zhu |
ACM Multimedia | 7 |
| 2024 | Deep Contrastive Clustering via Hard positive sample Debiased
Xinzhong Zhu |
Neurocomputing | 3 |
| 2024 | RFSC-net: Re-parameterization forward semantic compensation network in low-light environments
Xinzhong Zhu, Yunzhong Si |
Image Vis. Comput. | 3 |
| 2024 | Graph Contrastive Multi-view Learning: A Pre-training Framework for Graph ClassificationabstractRecent advancements in node and graph classification tasks can be attributed to the implementation of contrastive learning and similarity search. Despite considerable progress, these approaches present challenges. The integration of similarity search introduces an additional layer of complexity to the model. At the same time, applying contrastive learning to non-transferable domains or out-of-domain datasets results in less competitive outcomes. In this work, we propose maintaining domain specificity for these tasks, which has demonstrated the potential to improve performance by eliminating the need for additional similarity searches. We adopt a fraction of domain-specific datasets for pre-training purposes, generating augmented pairs that retain structural similarity to the original graph, thereby broadening the number of views. This strategy involves a comprehensive exploration of optimal augmentations to devise multi-view embeddings. An evaluation protocol, which focuses on error minimization, accuracy enhancement, and overfitting prevention, guides this process to learn inherent, transferable structural representations that span diverse datasets. We combine pre-trained embeddings and the source graph as a beneficial input, leveraging local and global graph information to enrich downstream tasks. Furthermore, to maximize the utility of negative samples in contrastive learning, we extend the training mechanism during the pre-training stage. Our method consistently outperforms comparative baseline approaches in comprehensive experiments conducted on benchmark graph datasets of varying sizes and characteristics, establishing new state-of-the-art results. Michael Adjeisah, Xinzhong Zhu |
Knowl. Based Syst. | 2 |
| 2024 | Fast Approximated Multiple Kernel K-MeansabstractMultiple Kernel Clustering (MKC) has emerged as a prominent research domain in recent decades due to its capacity to exploit diverse information from multiple views by learning an optimal kernel. Despite the successes achieved by various MKC methods, a significant challenge lies in the computational complexity associated with generating a consensus partition from the optimal kernel matrix, typically of size$n \times n$, where$n$represents the number of samples. This computational bottleneck restricts the practical applicability of these methods when confronted with large-scale datasets. Furthermore, certain existing MKC algorithms derive the consensus partition matrix by fusing all base partitions. However, this fusion process may inadvertently overlook critical information embedded in individual base kernels, potentially leading to inferior clustering performance. In light of these challenges, we introduce an innovative and efficient multiple kernel$k$-means approach, denoted as FAMKKM. Notably, FAMKKM incorporates two approximated partition matrices instead of the original individual partition matric for each base kernel. This strategic substitution significantly reduces computational complexity. Additionally, FAMKKM leverages the original kernel information to guide the fusion of all base partitions, thereby enhancing the quality of the resulting consensus partition matrix. Finally, we substantiate the efficacy and efficiency of the proposed FAMKKM through extensive experiments conducted on six benchmark datasets. Our results demonstrate its superiority over state-of-the-art methods. The demo code of this work is publicly available athttps://github.com/WangJun2023/FAMKKM Jun Wang 0118, Chang Tang, Xinwang Liu 0002, Wei Zhang 0049, En Zhu, Xinzhong Zhu |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2024 | Eigenvalue Ratio Inspired Partition Learning and Fusion for Multiple Kernel ClusteringabstractMultiple kernel clustering (MKC) aims to extract and integrate the clustering information from a set of pre-defined kernels for handling data which cannot be linearly separated well. More precisely, existing MKC methods generally devote to learn the complementary information from a set of kernel partitions, whose feature dimensions are commonly fixed as the upper bound$n$or lower bound$c$, where$n$and$c$represents the number of samples and clusters, respectively. However, the adopting of the lower bound or upper bound generally leads to poor clustering performance caused by the lack or redundancy of clustering information carried by kernel partitions. To tackle this issue, we propose a novel late fusion multiple kernel clustering method, termed as Eigenvalue Ratio Inspired Partition Learning and Fusion for Multiple Kernel Clustering (ERMKC), in this paper. Specifically, we propose an eigenvalue ratio based criterion to guide the kernel partition learning for each single kernel matrix, which ensures more suitable feature dimensions for the learnt kernel partitions. In addition, we also propose a novel late fusion model for fusing the learnt kernel partitions optimally. Furthermore, we conduct extensive experiments on numerous benchmark datasets to evaluate the proposed ERMKC method, whose results verify the effectiveness and advantage of the proposed method compared to the other state-of-the-art methods. Wenqi Yang, Chang Tang, Xinzhong Zhu, Xinwang Liu 0002 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Iterative Deep Structural Graph Contrast Clustering for Multiview Raw DataabstractMultiview clustering has attracted increasing attention to automatically divide instances into various groups without manual annotations. Traditional shadow methods discover the internal structure of data, while deep multiview clustering (DMVC) utilizes neural networks with clustering-friendly data embeddings. Although both of them achieve impressive performance in practical applications, we find that the former heavily relies on the quality of raw features, while the latter ignores the structure information of data. To address the above issue, we propose a novel method termed iterative deep structural graph contrast clustering (IDSGCC) for multiview raw data consisting of topology learning (TL), representation learning (RL), and graph structure contrastive learning to achieve better performance. The TL module aims to obtain a structured global graph with constraint structural information and then guides the RL to preserve the structural information. In the RL module, graph convolutional network (GCN) takes the global structural graph and raw features as inputs to aggregate the samples of the same cluster and keep the samples of different clusters away. Unlike previous methods performing contrastive learning at the representation level of the samples, in the graph contrastive learning module, we conduct contrastive learning at the graph structure level by imposing a regularization term on the similarity matrix. The credible neighbors of the samples are constructed as positive pairs through the credible graph, and other samples are constructed as negative pairs. The three modules promote each other and finally obtain clustering-friendly embedding. Also, we set up an iterative update mechanism to update the topology to obtain a more credible topology. Impressive clustering results are obtained through the iterative mechanism. Comparative experiments on eight multiview datasets show that our model outperforms the state-of-the-art traditional and deep clustering competitors. Zhibin Dong, Jiaqi Jin, Yuyang Xiao, Siwei Wang 0001, Xinzhong Zhu, Xinwang Liu 0002, En Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Joint Optimization of System Bandwidth and Transmitting Power in Space-Air-Ground Integrated Mobile Edge Computing
Yuan Qiu 0006, Jianwei Niu 0002, Tao Ren 0001, Xinzhong Zhu, Kuntuo Zhu |
ICA3PP (6) | 6 |
| 2023 | Unsupervised feature selection via multiple graph fusion and feature weight learning
Chang Tang, Wei Zhang 0049, Xinwang Liu 0002, Xinzhong Zhu, En Zhu |
Sci. China Inf. Sci. | 5 |
| 2023 | Mutual structure learning for multiple kernel clustering
Zhenglai Li, Chang Tang, Zhiguo Wan, Kun Sun 0002, Wei Zhang 0049, Xinzhong Zhu |
Inf. Sci. | 7 |
| 2023 | Region-Aware Hierarchical Latent Feature Representation Learning-Guided Clustering for Hyperspectral Band SelectionabstractHyperspectral band selection aims to identify an optimal subset of bands for hyperspectral images (HSIs). For most existing clustering-based band selection methods, they directly stretch each band into a single feature vector and employ the pixelwise features to address band redundancy. In this way, they do not take full consideration of the spatial information and deal with the importance of different regions in HSIs, which leads to a nonoptimal selection. To address these issues, a region-aware hierarchical latent feature representation learning-guided clustering (HLFC) method is proposed. Specifically, in order to fully preserve the spatial information of HSIs, the superpixel segmentation algorithm is adopted to segment HSIs into multiple regions first. For each segmented region, the similarity graph is constructed to reflect the bands-wise similarity, and its corresponding Laplacian matrix is generated for learning low-dimensional latent features in a hierarchical way. All latent features are then fused to form a unified feature representation of HSIs. Finally, k -means clustering is utilized on the unified feature representation matrix to generate multiple clusters from which the band with maximum information entropy is selected to form the final subset of bands. Extensive experimental results demonstrate that the proposed clustering method can achieve superior performance than the state-of-the-art representative methods on the band selection. The demo code of this work is publicly available at https://github.com/WangJun2023/HLFC. Jun Wang 0118, Chang Tang, Xinwang Liu 0002, Wei Zhang 0049, Wanqing Li 0001, Xinzhong Zhu, Lizhe Wang 0001, Albert Y. Zomaya |
IEEE Trans. Cybern. | 6 |
| 2023 | Localized Incomplete Multiple Kernel k-Means With Matrix-Induced RegularizationabstractLocalized incomplete multiple kernel k -means (LI-MKKM) is recently put forward to boost the clustering accuracy via optimally utilizing a quantity of prespecified incomplete base kernel matrices. Despite achieving significant achievement in a variety of applications, we find out that LI-MKKM does not sufficiently consider the diversity and the complementary of the base kernels. This could make the imputation of incomplete kernels less effective, and vice versa degrades on the subsequent clustering. To tackle these problems, an improved LI-MKKM, called LI-MKKM with matrix-induced regularization (LI-MKKM-MR), is proposed by incorporating a matrix-induced regularization term to handle the correlation among base kernels. The incorporated regularization term is beneficial to decrease the probability of simultaneously selecting two similar kernels and increase the probability of selecting two kernels with moderate differences. After that, we establish a three-step iterative algorithm to solve the corresponding optimization objective and analyze its convergence. Moreover, we theoretically show that the local kernel alignment is a special case of its global one with normalizing each base kernel matrices. Based on the above observation, the generalization error bound of the proposed algorithm is derived to theoretically justify its effectiveness. Finally, extensive experiments on several public datasets have been conducted to evaluate the clustering performance of the LI-MKKM-MR. As indicated, the experimental results have demonstrated that our algorithm consistently outperforms the state-of-the-art ones, verifying the superior performance of the proposed algorithm. Miaomiao Li 0001, Jingyuan Xia, Qing Liao 0001, Xinzhong Zhu, Xinwang Liu 0002 |
IEEE Trans. Cybern. | 5 |
| 2023 | Spatial and Spectral Structure Preserved Self-Representation for Unsupervised Hyperspectral Band SelectionabstractAs an effective manner to reduce data redundancy and processing inconvenience, hyperspectral band selection aims to select a subset of informative and discriminative bands from the original data cube. Although a large number of approaches have been proposed and obtained great success, they still face at least two issues. Firstly, most of the previous methods only consider the redundancy between neighbor bands, while the global information has been ignored. Secondly, each band is often treated as a whole and reshaped to a feature vector without considering the spatial structure of different regions. In this paper, in order to address these issues, we propose a spatial and spectral structure preserved self-representation model for unsupervised hyperspectral band selection without using any label information, referred to as S4P briefly. Different from previous methods that stretch each band into a feature vector, the first principal component of the original hyperspectral cube is segmented into different superpixels, which can reflect the spatial structure of homogeneous regions. Then each band can be represented by a superpixel level feature vector and the self-representation model is utilized to learn the spectral correlation of different bands. In addition, an adaptive and weighted multiple graph fusion term is designed to generate a unified similarity graph between different superpixels, which is used to capture the spatial structure in the self-representation space. Finally, anl2,1-norm is imposed on the self-representation coefficient matrix to measure the band importance. We design an alternative update scheme to optimize the resultant problem, the self-representation coefficient matrix and the superpixel-wise similarity graph can boost each other during the updating process to obtain optimal results. Extensive experiments with detailed analysis of three public datasets are conducted to validate the superiority of the proposed S4P when compared with other state-of-the-art competitors. Chang Tang, Jun Wang 0118, Xinwang Liu 0002, Weiying Xie, Xianju Li, Xinzhong Zhu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Highly-efficient Incomplete Largescale Multiview Clustering with Consensus Bipartite GraphabstractMultiview clustering has received increasing attention due to its effectiveness in fusing complementary information without manual annotations. Most previous methods hold the assumption that each instance appears in all views. However, it is not uncommon to see that some views may contain some missing instances, which gives rise to incomplete multi-view clustering (IMVC) in literature. Although many IMVC methods have been recently proposed, they always encounter high complexity and expensive time expenditure from being applied into large-scale tasks. In this paper, we present a flexible highly-efficient incomplete large-scale multi-view clustering approach based on bipartite graph framework to solve these issues. Specifically, we formalize multi-view anchor learning and incomplete bipartite graph into a unified framework, which coordinates with each other to boost cluster performance. By introducing the flexible bipartite graph framework to handle IMVC for the first practice, our proposed method enjoys linear complexity respecting to instance numbers, which is more applicable for large-scale IMVC tasks. Comprehensive experimental results on various benchmark datasets demonstrate the effectiveness and efficiency of our proposed algorithm against other IMVC competitors. The code is available at11https://github.com/wangsiwei2010/CVPR22-IMVC-CBG. Siwei Wang 0001, Xinwang Liu 0002, Li Liu 0002, Wenxuan Tu, Xinzhong Zhu, Jiyuan Liu 0003, Sihang Zhou 0001, En Zhu |
CVPR | 5 |
| 2022 | Detecting Anomalous Events from Unlabeled Videos via Temporal Masked Auto-EncodingabstractUnsupervised video anomaly detection (UVAD) intends to discern anomalous events from fully unlabeled videos. However, existing UVAD methods suffer from poor performance. Inspired by recent masked autoencoder (MAE) [1], we propose Temporal Masked Auto-Encoding (TMAE) as an effective end-to-end UVAD method. Specifically, we first denote video events by spatial-temporal cubes (STCs), which are built by temporally consecutive foreground patches from unlabeled videos. Then, half of patches in an STC are masked along the temporal dimension, while a vision transformer (ViT) is trained to exploit unmasked patches to predict masked patches. The rare and unusual nature of anomaly will result in a poorer prediction for anomalous events, which enables us to discriminate anomalies from unlabeled videos and compute the anomaly scores. Furthermore, to utilize motion clues in videos, we also propose to apply TMAE on optical flow, which can further boost performance. Experiments show that TMAE significantly outperforms existing UVAD methods by a notable margin (3.9%–6.6% AUC). Jingtao Hu, Siqi Wang 0001, En Zhu, Zhiping Cai, Xinzhong Zhu |
ICME | 6 |
| 2022 | Align then Fusion: Generalized Large-scale Multi-view Clustering with Anchor Matching CorrespondencesabstractMulti-view anchor graph clustering selects representative anchors to avoid full pair-wise similarities and therefore reduce the complexity of graph methods. Although widely applied in large-scale applications, existing approaches do not pay sufficient attention to establishing correct correspondences between the anchor sets across views. To be specific, anchor graphs obtained from different views are not aligned column-wisely. Such an Anchor-Unaligned Problem (AUP) would cause inaccurate graph fusion and degrade the clustering performance. Under multi-view scenarios, generating correct correspondences could be extremely difficult since anchors are not consistent in feature dimensions. To solve this challenging issue, we propose the first study of the generalized and flexible anchor graph fusion framework termed Fast Multi-View Anchor-Correspondence Clustering (FMVACC). Specifically, we show how to find anchor correspondence with both feature and structure information, after which anchor graph fusion is performed column-wisely. Moreover, we theoretically show the connection between FMVACC and existing multi-view late fusion and partial view-aligned clustering, which further demonstrates our generality. Extensive experiments on seven benchmark datasets demonstrate the effectiveness and efficiency of our proposed method. Moreover, the proposed alignment module also shows significant performance improvement applying to existing multi-view anchor graph competitors indicating the importance of anchor alignment. Our code is available at \url{https://github.com/wangsiwei2010/NeurIPS22-FMVACC}. Siwei Wang 0001, Xinwang Liu 0002, Suyuan Liu, Jiaqi Jin, Wenxuan Tu, Xinzhong Zhu, En Zhu |
NeurIPS | 6 |
| 2022 | 3-D Auxiliary Classifier GAN for Hyperspectral Anomaly Detection via Weakly Supervised LearningabstractHyperspectral anomaly detection (AD) is important in Earth observation and remote sensing. However, the low spatial resolution of hyperspectral images, insufficient samples and lack of prior information limit the detection accuracy. To solve these problems, in this paper, we propose an auxiliary classifier generative adversarial network model based on a three-dimensional (3D) convolutional neural network named 3D AC-GAN. Firstly, the model is based on a 3D convolutional neural network design, with 3D tensors as samples. The network maintains valuable image spatial spectrum joint features to achieve good detection results. It can also generate sufficient samples to achieve dataset augmentation, solving the overfitting problem in GAN training. Secondly, we train the model with a weakly supervised method. The label of the samples is obtained through the coarse scanning method. Then, the AC-GAN is trained with the bootstrapping method to mitigate the impact of noise labels. The experimental results show that our proposed algorithm outperforms state-of-the-art AD algorithms. Huanlin Luo, Haowen Zhu, Shengyang Liu, Xinzhong Zhu, Jinmei Lai 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Fast Parameter-Free Multi-View Subspace Clustering With Consensus Anchor GuidanceabstractMulti-view subspace clustering has attracted intensive attention to effectively fuse multi-view information by exploring appropriate graph structures. Although existing works have made impressive progress in clustering performance, most of them suffer from the cubic time complexity which could prevent them from being efficiently applied into large-scale applications. To improve the efficiency, anchor sampling mechanism has been proposed to select vital landmarks to represent the whole data. However, existing anchor selecting usually follows the heuristic sampling strategy, e.g. k -means or uniform sampling. As a result, the procedures of anchor selecting and subsequent subspace graph construction are separated from each other which may adversely affect clustering performance. Moreover, the involved hyper-parameters further limit the application of traditional algorithms. To address these issues, we propose a novel subspace clustering method termed Fast Parameter-free Multi-view Subspace Clustering with Consensus Anchor Guidance (FPMVS-CAG). Firstly, we jointly conduct anchor selection and subspace graph construction into a unified optimization formulation. By this way, the two processes can be negotiated with each other to promote clustering quality. Moreover, our proposed FPMVS-CAG is proved to have linear time complexity with respect to the sample number. In addition, FPMVS-CAG can automatically learn an optimal anchor subspace graph without any extra hyper-parameters. Extensive experimental results on various benchmark datasets demonstrate the effectiveness and efficiency of the proposed method against the existing state-of-the-art multi-view subspace clustering competitors. These merits make FPMVS-CAG more suitable for large-scale subspace clustering. The code of FPMVS-CAG is publicly available at https://github.com/wangsiwei2010/FPMVS-CAG. Siwei Wang 0001, Xinwang Liu 0002, Xinzhong Zhu, Pei Zhang 0008, Yi Zhang 0104, En Zhu |
IEEE Trans. Image Process. | 3 |
| 2022 | Adaptive Semantic-Spatio-Temporal Graph Convolutional Network for Lip ReadingabstractThe goal of this work is to recognize words, phrases, and sentences being spoken by a talking face without given the audio. Current deep learning approaches for lip reading focus on exploring the appearance and optical flow information of videos. However, these methods do not fully exploit the characteristics of lip motion. In addition to appearance and optical flow, the mouth contour deformation usually conveys significant information that is complementary to others. However, the modeling of dynamic mouth contour has received little attention than that of appearance and optical flow. In this work, we propose a novel model of dynamic mouth contours called Adaptive Semantic-Spatio-Temporal Graph Convolution Network (ASST-GCN), to go beyond previous methods by automatically learning both the spatial and temporal information from videos. To combine the complementary information from appearance and mouth contour, a two-stream visual front-end network is proposed. Experimental results demonstrate that the proposed method significantly outperforms the state-of-the-art lip reading methods on several large-scale lip reading benchmarks. Changchong Sheng, Xinzhong Zhu, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Multim. | 2 |
| 2021 | Partial multiview clustering with locality graph regularizationabstractMultiview clustering (MVC) collects complementary and abundant information, which draws much attention in machine learning and data mining community. Existing MVC methods usually hold the assumption that all the views are complete. However, multiple source data are often incomplete in real-world applications, and so on sensor failure or unfinished collection process, which gives rise to incomplete multiview clustering (IMVC). Although enormous efforts have been devoted in IMVC, there still are some urgent issues that need to be solved: (i) The locality among multiple views has not been utilized in the existing mechanism; (ii) Existing methods inappropriately force all the views to share consensus representation while ignoring specific structures. In this paper, we propose a novel method termed partial MVC with locality graph regularization to address these issues. First, followed the traditional IMVC approaches, we construct weighted semi-nonnegative matrix factorization models to handle incomplete multiview data. Then, upon the consensus representation matrix, the locality graph is constructed for regularizing the shared feature matrix. Moreover, we add the coefficient regression term to constraint the various base matrices among views. We incorporate the three aforementioned processes into a unified framework, whereas they can negotiate with each other serving for learning tasks. An effective iterative algorithm is proposed to solve the resultant optimization problem with theoretically guaranteed convergence. The comprehensive experiment results on several benchmarks demonstrate the effectiveness of the proposed method. Huiqiang Lian, Siwei Wang 0001, Miaomiao Li 0001, Xinzhong Zhu, Xinwang Liu 0002 |
Int. J. Intell. Syst. | 5 |
| 2021 | Gaussian Mixture Model Clustering with Incomplete DataabstractGaussian mixture model (GMM) clustering has been extensively studied due to its effectiveness and efficiency. Though demonstrating promising performance in various applications, it cannot effectively address the absent features among data, which is not uncommon in practical applications. In this article, different from existing approaches that first impute the absence and then perform GMM clustering tasks on the imputed data, we propose to integrate the imputation and GMM clustering into a unified learning procedure. Specifically, the missing data is filled by the result of GMM clustering, and the imputed data is then taken for GMM clustering. These two steps alternatively negotiate with each other to achieve optimum. By this way, the imputed data can best serve for GMM clustering. A two-step alternative algorithm with proved convergence is carefully designed to solve the resultant optimization problem. Extensive experiments have been conducted on eight UCI benchmark datasets, and the results have validated the effectiveness of the proposed algorithm. Yi Zhang 0104, Miaomiao Li 0001, Siwei Wang 0001, Sisi Dai, Lei Luo 0002, En Zhu, Xinzhong Zhu, Chaoyun Yao |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2020 | CGD: Multi-View Clustering via Cross-View Graph DiffusionabstractGraph based multi-view clustering has been paid great attention by exploring the neighborhood relationship among data points from multiple views. Though achieving great success in various applications, we observe that most of previous methods learn a consensus graph by building certain data representation models, which at least bears the following drawbacks. First, their clustering performance highly depends on the data representation capability of the model. Second, solving these resultant optimization models usually results in high computational complexity. Third, there are often some hyper-parameters in these models need to tune for obtaining the optimal results. In this work, we propose a general, effective and parameter-free method with convergence guarantee to learn a unified graph for multi-view data clustering via cross-view graph diffusion (CGD), which is the first attempt to employ diffusion process for multi-view clustering. The proposed CGD takes the traditional predefined graph matrices of different views as input, and learns an improved graph for each single view via an iterative cross diffusion process by 1) capturing the underlying manifold geometry structure of original data points, and 2) leveraging the complementary information among multiple graphs. The final unified graph used for clustering is obtained by averaging the improved view associated graphs. Extensive experiments on several benchmark datasets are conducted to demonstrate the effectiveness of the proposed method in terms of seven clustering evaluation metrics. Chang Tang, Xinwang Liu 0002, Xinzhong Zhu, En Zhu, Zhigang Luo, Lizhe Wang 0001, Wen Gao 0001 |
AAAI | 3 |
| 2020 | R²MRF: Defocus Blur Detection via Recurrently Refining Multi-Scale Residual FeaturesabstractDefocus blur detection aims to separate the in-focus and out-of-focus regions in an image. Although attracting more and more attention due to its remarkable potential applications, there are still several challenges for accurate defocus blur detection, such as the interference of background clutter, sensitivity to scales and missing boundary details of defocus blur regions. In order to address these issues, we propose a deep neural network which Recurrently Refines Multi-scale Residual Features (R2MRF) for defocus blur detection. We firstly extract multi-scale deep features by utilizing a fully convolutional network. For each layer, we design a novel recurrent residual refinement branch embedded with multiple residual refinement modules (RRMs) to more accurately detect blur regions from the input image. Considering that the features from bottom layers are able to capture rich low-level features for details preservation while the features from top layers are capable of characterizing the semantic information for locating blur regions, we aggregate the deep features from different layers to learn the residual between the intermediate prediction and the ground truth for each recurrent step in each residual refinement branch. Since the defocus degree is sensitive to image scales, we finally fuse the side output of each branch to obtain the final blur detection map. We evaluate the proposed network on two commonly used defocus blur detection benchmark datasets by comparing it with other 11 state-of-the-art methods. Extensive experimental results with ablation studies demonstrate that R2MRF consistently and significantly outperforms the competitors in terms of both efficiency and accuracy. Chang Tang, Xinwang Liu 0002, Xinzhong Zhu, En Zhu, Kun Sun 0002, Pichao Wang, Lizhe Wang 0001, Albert Y. Zomaya |
AAAI | 3 |
| 2020 | Absent Multiple Kernel Learning AlgorithmsabstractMultiple kernel learning (MKL) has been intensively studied during the past decade. It optimally combines the multiple channels of each sample to improve classification performance. However, existing MKL algorithms cannot effectively handle the situation where some channels of the samples are missing, which is not uncommon in practical applications. This paper proposes three absent MKL (AMKL) algorithms to address this issue. Different from existing approaches where missing channels are first imputed and then a standard MKL algorithm is deployed on the imputed data, our algorithms directly classify each sample based on its observed channels, without performing imputation. Specifically, we define a margin for each sample in its own relevant space, a space corresponding to the observed channels of that sample. The proposed AMKL algorithms then maximize the minimum of all sample-based margins, and this leads to a difficult optimization problem. We first provide two two-step iterative algorithms to approximately solve this problem. After that, we show that this problem can be reformulated as a convex one by applying the representer theorem. This makes it readily be solved via existing convex optimization packages. In addition, we provide a generalization error bound to justify the proposed AMKL algorithms from a theoretical perspective. Extensive experiments are conducted on nine UCI and six MKL benchmark datasets to compare the proposed algorithms with existing imputation-based methods. As demonstrated, our algorithms achieve superior performance and the improvement is more significant with the increase of missing ratio. Xinwang Liu 0002, Lei Wang 0001, Xinzhong Zhu, Miaomiao Li 0001, En Zhu, Tongliang Liu, Li Liu 0002, Yong Dou, Jianping Yin |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Multiple Kernel $k$k-Means with Incomplete KernelsabstractMultiple kernel clustering (MKC) algorithms optimally combine a group of pre-specified base kernel matrices to improve clustering performance. However, existing MKC algorithms cannot efficiently address the situation where some rows and columns of base kernel matrices are absent. This paper proposes two simple yet effective algorithms to address this issue. Different from existing approaches where incomplete kernel matrices are first imputed and a standard MKC algorithm is applied to the imputed kernel matrices, our first algorithm integrates imputation and clustering into a unified learning procedure. Specifically, we perform multiple kernel clustering directly with the presence of incomplete kernel matrices, which are treated as auxiliary variables to be jointly optimized. Our algorithm does not require that there be at least one complete base kernel matrix over all the samples. Also, it adaptively imputes incomplete kernel matrices and combines them to best serve clustering. Moreover, we further improve this algorithm by encouraging these incomplete kernel matrices to mutually complete each other. The three-step iterative algorithm is designed to solve the resultant optimization problems. After that, we theoretically study the generalization bound of the proposed algorithms. Extensive experiments are conducted on 13 benchmark data sets to compare the proposed algorithms with existing imputation-based methods. Our algorithms consistently achieve superior performance and the improvement becomes more significant with increasing missing ratio, verifying the effectiveness and advantages of the proposed joint imputation and clustering. Xinwang Liu 0002, Xinzhong Zhu, Miaomiao Li 0001, Lei Wang 0001, En Zhu, Tongliang Liu, Marius Kloft, Dinggang Shen, Jianping Yin, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Adaptive Self-Paced Deep Clustering with Data AugmentationabstractDeep clustering gains superior performance than conventional clustering by jointly performing feature learning and cluster assignment. Although numerous deep clustering algorithms have emerged in various applications, most of them fail to learn robust cluster-oriented features which in turn hurts the final clustering performance. To solve this problem, we propose a two-stage deep clustering algorithm by incorporating data augmentation and self-paced learning. Specifically, in the first stage, we learn robust features by training an autoencoder with examples that are augmented by random shifting and rotating the given clean examples. Then, in the second stage, we encourage the learned features to be cluster-oriented by alternatively finetuning the encoder with the augmented examples and updating the cluster assignments of the clean examples. During finetuning the encoder, the target of each augmented example in the loss function is the center of the cluster to which the clean example is assigned. The targets may be computed incorrectly, and the examples with incorrect targets could mislead the encoder network. To stabilize the network training, we select most confident examples in each iteration by utilizing the adaptive self-paced learning. Extensive experiments validate that our algorithm outperforms the state of the arts on four image datasets. Xifeng Guo 0001, Xinwang Liu 0002, En Zhu, Xinzhong Zhu, Miaomiao Li 0001, Xin Xu 0001, Jianping Yin |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Feature Selective Projection with Low-Rank Embedding and Dual Laplacian RegularizationabstractFeature extraction and feature selection have been regarded as two independent dimensionality reduction methods in most of the existing literature. In this paper, we propose to integrate both approaches into a unified framework and design an unsupervised linear feature selective projection (FSP) for feature extraction with low-rank embedding and dual Laplacian regularization, with the aim to exploit the intrinsic relationship among data and suppress the impact of noise. Specifically, a projection matrix with an l2,1-norm regularization is introduced to project original high dimensional data points into a new subspace with lower dimension, where the l2,1-norm regularization can endow the projection with good interpretability. We deploy a coefficient matrix with low rank constraint to reconstruct the data points and the l2,1-norm is imposed to regularize the data reconstruction errors in the low-dimensional subspace and make FSP robust to noise. Furthermore, a dual graph Laplacian regularization term is imposed on the low dimensional data and data reconstruction matrix for preserving the local manifold geometrical structure of data. Finally, an alternatively iterative algorithm is carefully designed for solving the proposed optimization model. Theoretical convergence and computational complexity analysis of the algorithm are also provided. Comprehensive experiments on various benchmark datasets have been carried out to evaluate the performance of the proposed FSP. As indicated, our algorithm significantly outperforms other state-of-the-art methods for feature extraction. Chang Tang, Xinwang Liu 0002, Xinzhong Zhu, Jian Xiong 0002, Miaomiao Li 0001, Jingyuan Xia, Xiangke Wang, Lizhe Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2019 | Efficient and Effective Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering (IMVC) optimally fuses multiple pre-specified incomplete views to improve clustering performance. Among various excellent solutions, the recently proposed multiple kernel k-means with incomplete kernels (MKKM-IK) forms a benchmark, which redefines IMVC as a joint optimization problem where the clustering and kernel matrix imputation tasks are alternately performed until convergence. Though demonstrating promising performance in various applications, we observe that the manner of kernel matrix imputation in MKKM-IK would incur intensive computational and storage complexities, overcomplicated optimization and limitedly improved clustering performance. In this paper, we propose an Efficient and Effective Incomplete Multi-view Clustering (EE-IMVC) algorithm to address these issues. Instead of completing the incomplete kernel matrices, EE-IMVC proposes to impute each incomplete base matrix generated by incomplete views with a learned consensus clustering matrix. We carefully develop a three-step iterative algorithm to solve the resultant optimization problem with linear computational complexity and theoretically prove its convergence. Further, we conduct comprehensive experiments to study the proposed EE-IMVC in terms of clustering accuracy, running time, evolution of the learned consensus clustering matrix and the convergence. As indicated, our algorithm significantly and consistently outperforms some state-of-the-art algorithms with much less running time and memory. Xinwang Liu 0002, Xinzhong Zhu, Miaomiao Li 0001, Chang Tang, En Zhu, Jianping Yin, Wen Gao 0001 |
AAAI | 2 |
| 2019 | Cross-View Local Structure Preserved Diversity and Consensus Learning for Multi-View Unsupervised Feature SelectionabstractMulti-view unsupervised feature selection (MV-UFS) aims to select a feature subset from multi-view data without using the labels of samples. However, we observe that existing MV-UFS algorithms do not well consider the local structure of cross views and the diversity of different views, which could adversely affect the performance of subsequent learning tasks. In this paper, we propose a cross-view local structure preserved diversity and consensus semantic learning model for MV-UFS, termed CRV-DCL briefly, to address these issues. Specifically, we project each view of data into a common semantic label space which is composed of a consensus part and a diversity part, with the aim to capture both the common information and distinguishing knowledge across different views. Further, an inter-view similarity graph between each pairwise view and an intra-view similarity graph of each view are respectively constructed to preserve the local structure of data in different views and different samples in the same view. An l2,1-norm constraint is imposed on the feature projection matrix to select discriminative features. We carefully design an efficient algorithm with convergence guarantee to solve the resultant optimization problem. Extensive experimental study is conducted on six publicly real multi-view datasets and the experimental results well demonstrate the effectiveness of CRV-DCL. Chang Tang, Xinzhong Zhu, Xinwang Liu 0002, Lizhe Wang 0001 |
AAAI | 2 |
| 2019 | DeFusionNET: Defocus Blur Detection via Recurrently Fusing and Refining Multi-Scale Deep FeaturesabstractDefocus blur detection aims to detect out-of-focus regions from an image. Although attracting more and more attention due to its widespread applications, defocus blur detection still confronts several challenges such as the interference of background clutter, sensitivity to scales and missing boundary details of defocus blur regions. To deal with these issues, we propose a deep neural network which recurrently fuses and refines multi-scale deep features (DeFusionNet) for defocus blur detection. We firstly utilize a fully convolutional network to extract multi-scale deep features. The features from bottom layers are able to capture rich low-level features for details preservation, while the features from top layers can characterize the semantic information to locate blur regions. These features from different layers are fused as shallow features and semantic features, respectively. After that, the fused shallow features are propagated to top layers for refining the fine details of detected defocus blur regions, and the fused semantic features are propagated to bottom layers to assist in better locating the defocus regions. The feature fusing and refining are carried out in a recurrent manner. Also, we finally fuse the output of each layer at the last recurrent step to obtain the final defocus blur map by considering the sensitivity to scales of the defocus degree. Experiments on two commonly used defocus blur detection benchmark datasets are conducted to demonstrate the superority of DeFusionNet when compared with other 10 competitors. Code and more results can be found at: http://tangchang.net. Chang Tang, Xinzhong Zhu, Xinwang Liu 0002, Lizhe Wang 0001, Albert Y. Zomaya |
CVPR | 2 |
| 2019 | Salient Object Detection via Recurrently Aggregating Spatial Attention Weighted Cross-Level Deep FeaturesabstractThis paper proposes a novel deep saliency detection network by recurrently aggregating and refining features in a cross-level and spatial attention-aware manner. In this way, the features integrated from multiple layers can be used to refine layer-wise features step by step and the complementary information in different layers can be fully captured for detecting salient objects with different scales, i.e., the features integrated from low-level layers can serve to refine the details of detected salient objects while the features integrated from high-level layers with semantic information can benefit the locating of salient objects. In addition, by considering that only partial regions of an image are salient, we embed a spatial attention-aware module to suppress the non-salient regions and highlight salient objects. Finally, different saliency detection results from different layers are fused to generate the final saliency map. Experimental results on five benchmark datasets demonstrate that our proposed method outperforms other 14 state-of-the-art competitors. Chang Tang, Xinzhong Zhu, Xinwang Liu 0002, Pichao Wang |
ICME | 2 |
| 2019 | CoinNet: Copy Initialization Network for Multispectral Imagery Semantic SegmentationabstractRemote sensing imagery semantic segmentation refers to assigning a label to every pixel. Recently, deep convolutional neural networks (CNNs)-based methods have presented an impressive performance in this task. Due to the lack of sufficient labeled remote sensing images, researchers usually utilized transfer learning (TL) strategies to fine tune networks which were pretrained in huge RGB-scene data sets. Unfortunately, this manner may not work if the target images are multispectral/hyperspectral. The basic assumption of TL is that the low-level features extracted by the former layers are similar in most data sets, hence users only require to train the parameters in the last layers that are specific to different tasks. However, if one should use a pretrained deep model in RGB data for multispectral /hyperspectral imagery semantic segmentation, the structure of the input layer has to be adjusted. In this case, the first convolutional layer has to be trained using the multispectral /hyperspectral data sets which are much smaller. Apparently, the feature representation ability of the first convolutional layer will decrease and it may further harm the following layers. In this letter, we propose a new deep learning model, COpy INitialization Network (CoinNet), for multispectral imagery semantic segmentation. The major advantage of CoinNet is that it can make full use of the initial parameters in the pretrained network's first convolutional layer. Comparison experiments on a challenging multispectral data set have demonstrated the effectiveness of the proposed improvement. The demo and a trained network will be published in our homepage. Bin Pan, Zhenwei Shi 0001, Tianyang Shi, Xinzhong Zhu |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2019 | Late Fusion Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering optimally integrates a group of pre-specified incomplete views to improve clustering performance. Among various excellent solutions, multiple kernel $k$k-means with incomplete kernels forms a benchmark, which redefines the incomplete multi-view clustering as a joint optimization problem where the imputation and clustering are alternatively performed until convergence. However, the comparatively intensive computational and storage complexities preclude it from practical applications. To address these issues, we propose Late Fusion Incomplete Multi-view Clustering (LF-IMVC) which effectively and efficiently integrates the incomplete clustering matrices generated by incomplete views. Specifically, our algorithm jointly learns a consensus clustering matrix, imputes each incomplete base matrix, and optimizes the corresponding permutation matrices. We develop a three-step iterative algorithm to solve the resultant optimization problem with linear computational complexity and theoretically prove its convergence. Further, we conduct comprehensive experiments to study the proposed LF-IMVC in terms of clustering accuracy, running time, advantages of late fusion multi-view clustering, evolution of the learned consensus clustering matrix, parameter sensitivity and convergence. As indicated, our algorithm significantly and consistently outperforms some state-of-the-art algorithms with much less running time and memory. Xinwang Liu 0002, Xinzhong Zhu, Miaomiao Li 0001, Lei Wang 0001, Chang Tang, Jianping Yin, Dinggang Shen, Huaimin Wang 0001, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Triangle Lasso for Simultaneous Clustering and Optimization in Graph DatasetsabstractRecently, network lasso has dawn much attention due to its remarkable performance on simultaneous clustering and optimization. However, it usually suffers from the imperfect data (noise, missing values, etc.), and yields sub-optimal solutions. The reason is that it finds the similar instances according to their features directly, which is usually impacted by the imperfect data, and thus returns sub-optimal results. In this paper, we propose triangle lasso to avoid its disadvantage for graph datasets. In a graph dataset, each instance is represented by a vertex. If two instances have many common adjacent vertices, they tend to become similar. Although some instances are profiled by the imperfect data, it is still able to find the similar counterparts. Furthermore, we develop an efficient algorithm based on Alternating Direction Method of Multipliers (ADMM) to obtain a moderately accurate solution. In addition, we present a dual method to obtain the accurate solution with the low additional time consumption. We demonstrate through extensive numerical experiments that triangle lasso is robust to the imperfect data. It usually yields a better performance than the state-of-the-art method when performing data analysis tasks in practical scenarios. Kai Xu 0004, En Zhu, Xinwang Liu 0002, Xinzhong Zhu, Jianping Yin |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2019 | Learning a Joint Affinity Graph for Multiview Subspace ClusteringabstractWith the ability to exploit the internal structure of data, graph-based models have received a lot of attention and have achieved great success in multiview subspace clustering for multimedia data. Most of the existing methods individually construct an affinity graph for each single view and fuse the result obtained from each single graph. However, the common representation shared by different views and the complementary diversity across these views are not efficiently exploited. In addition, noise and outliers are often mixed in original data, which adversely degenerate the clustering performance of many existing methods. In this paper, we propose addressing these issues by learning a joint affinity graph for multiview subspace clustering based on a low-rank representation with diversity regularization and a rank constraint. Specifically, a low-rank representation model is employed to learn a shared sample representation coefficient matrix to generate the affinity graph. At the same time, we use diversity regularization to learn the optimal weights for each view, which can suppress the redundancy and enhance the diversity among different feature views. In addition, the cluster number is used to promote affinity graph learning by using a rank constraint. The final clustering result is obtained by using normalized cuts on the learned affinity graph. An efficient algorithm based on an augmented Lagrangian multiplier with alternating direction minimization is carefully designed to solve the resulting optimization problem. Extensive experiments on various real-world datasets are conducted, and the results demonstrate well the effectiveness of the proposed algorithm. Chang Tang, Xinzhong Zhu, Xinwang Liu 0002, Miaomiao Li 0001, Pichao Wang, Changqing Zhang 0002, Lizhe Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | Localized Incomplete Multiple Kernel k-meansabstractThe recently proposed multiple kernel k-means with incomplete kernels (MKKM-IK) optimally integrates a group of pre-specified incomplete kernel matrices to improve clustering performance. Though it demonstrates promising performance in various applications, we observe that it does not \emph{sufficiently consider the local structure among data and indiscriminately forces all pairwise sample similarity to equally align with their ideal similarity values}. This could make the incomplete kernels less effectively imputed, and in turn adversely affect the clustering performance. In this paper, we propose a novel localized incomplete multiple kernel k-means (LI-MKKM) algorithm to address this issue. Different from existing MKKM-IK, LI-MKKM only requires the similarity of a sample to its k-nearest neighbors to align with their ideal similarity values. This helps the clustering algorithm to focus on closer sample pairs that shall stay together and avoids involving unreliable similarity evaluation for farther sample pairs. We carefully design a three-step iterative algorithm to solve the resultant optimization problem and theoretically prove its convergence. Comprehensive experiments on eight benchmark datasets demonstrate that our algorithm significantly outperforms the state-of-the-art comparable algorithms proposed in the recent literature, verifying the advantage of considering local structure. Xinzhong Zhu, Xinwang Liu 0002, Miaomiao Li 0001, En Zhu, Li Liu 0002, Zhiping Cai, Jianping Yin, Wen Gao 0001 |
IJCAI | 1 |
| 2018 | Robust graph regularized unsupervised feature selectionabstractRecent research indicates the critical importance of preserving local geometric structure of data in unsupervised feature selection (UFS), and the well studied graph Laplacian is usually deployed to capture this property. By using a squared l 2 -norm, we observe that conventional graph Laplacian is sensitive to noisy data, leading to unsatisfying data processing performance. To address this issue, we propose a unified UFS framework via feature self-representation and robust graph regularization , with the aim at reducing the sensitivity to outliers from the following two aspects: i) an l 2, 1 -norm is used to characterize the feature representation residual matrix; and ii) an l 1 -norm based graph Laplacian regularization term is adopted to preserve the local geometric structure of data. By this way, the proposed framework is able to reduce the effect of noisy data on feature selection. Furthermore, the proposed l 1 -norm based graph Laplacian is readily extendible, which can be easily integrated into other UFS methods and machine learning tasks with local geometrical structure of data being preserved. As demonstrated on ten challenging benchmark data sets, our algorithm significantly and consistently outperforms state-of-the-art UFS methods in the literature, suggesting the effectiveness of the proposed UFS framework. Chang Tang, Xinzhong Zhu, Jiajia Chen 0010, Pichao Wang, Xinwang Liu 0002, Jie Tian 0001 |
Expert Syst. Appl. | 2 |
| 2018 | Saliency detection via affinity graph learning and weighted manifold ranking
Xinzhong Zhu, Chang Tang, Pichao Wang, Minhui Wang, Jiajia Chen 0010, Jie Tian 0001 |
Neurocomputing | 1 |
| 2014 | Multi-scale local binary pattern with filters for spoof fingerprint detection
Xiaofei Jia, Xin Yang 0001, Kai Cao 0001, Yali Zang, Ning Zhang 0015, Ruwei Dai, Xinzhong Zhu, Jie Tian 0001 |
Inf. Sci. | 7 |
| 2013 | Super-class Discriminant Analysis: A novel solution for heteroscedasticity
Xinzhong Zhu |
Pattern Recognit. Lett. | 1 |
| 2008 | Multimedia Data Modeling Through a Semantic View Mechanism
Jianmin Zhao, Xinzhong Zhu |
World Wide Web | 3 |
| 2006 | Web Services Based On Grid TechnologyabstractCreating an integrated virtual computing environment based on grid will provide a technical platform and supporting environment for realizing a seamless access and full share of entire resources on the Internet. This text has introduced a kind of Web service (grid-based Web service, GRWS) based on grid mainly. This paper discusses the concept, architecture, working mechanism and security of the grid based on Web services, and finally gives a brief summary Jianmin Zhao, Xinzhong Zhu |
CSCWD | 3 |
| 2006 | A New Communication Mechanism Based on Virtual Agent for Mobile AgentabstractMobile agent technology can be used as one of the key technologies for new componentware frameworks. Communication mechanism of mobile agent plays a very important role in mobile agent systems. Communication mechanism could satisfy some requirements, such as location-transparency, reliability and high-efficiency. This paper introduces the current various communication algorithms and mechanisms for mobile agent, and summarizes their advantages and disadvantages. Subsequently, a new communication mechanism for mobile agent with the conception of virtual agent is put forward in this paper. This communication mechanism could implement the synchronism of communication and migration to a great extent, which could greatly improve the communication efficiency and reliability of the whole net Jianmin Zhao, Xinzhong Zhu, Xinpeng Zhuang |
CSCWD | 2 |
| 2006 | A Research of Collaborative CAD System Based on Multi-AgentabstractTo improve the features of cooperative design of products in distributed heterogeneous cooperation, methods of building cooperation interface and integration of heterogeneous design environment are presented with multi-agent technology. In products development, designs at various levels, stages and different places can cooperate and communicate with each other. So it takes a shorter time to develop a new product and high-quality of the product is improved. A system framework was put forward and a software environment supporting the remote collaborative design of product was developed and implemented to use Xinzhong Zhu, Jianmin Zhao, Rong Long |
CSCWD | 1 |