Zihua Zhao

dblp:299/4676 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Privileged information assisted learning from noisy correspondence
Zihua Zhao, Tianjie Dai, Mengxi Chen, Jiangchao Yao, Bo Han 0003, Ya Zhang 0002, Yanfeng Wang 0001
Neurocomputing1
2026 Dual-granularity Sinkhorn Distillation for Enhanced Learning from Long-Tailed Noisy Data
Feng Hong 0004, Zihua Zhao, Zhihan Zhou 0002, Jiangchao Yao, Dongsheng Li 0002, Ya Zhang 0002, Yanfeng Wang 0001
Mach. Learn.3
2026 Self-paced and structured graph-based ensemble clustering
Haoliang Tang, Zihua Zhao, Rong Wang 0001, Feiping Nie 0001
Signal Process.4
2026 Multi-view clustering via contrastive approach
Haonan Xin, Zihua Zhao, Jiacong Xiao, Rong Wang 0001
Signal Process.3
2026 SimMTC: Simple Multi-View Tensor Clustering
abstract
Tensor-based multi-view clustering algorithms have attracted considerable attention due to their superior clustering performance. However, these algorithms typically treat each view independently, failing to utilize the complementary information across all views, thus lacking globality. Additionally, employing low-rank tensor constraints to extract consistent information among views may result in the loss of important information due to weak consistency constraints. These limitations significantly hinder the clustering performance. To address these issues, we propose Simple Multi-view Tensor Clustering (SimMTC), which achieves globality and strong consistency. SimMTC first applies Fast Fourier Transform (FFT) to the anchor graphs to obtain high-frequency and low-frequency information, which encode similarities between samples and anchors from all views, thereby capturing global information. Orthogonal tensor factorization is then conducted in the frequency domain. Moreover, a novel strong consistency constraint based on FFT is introduced, which enhances the extraction of consistent information in the frequency domain. What's more, an efficient alternating optimization algorithm is designed to solve the optimization problem in SimMTC. Finally, extensive experiments on real-world datasets demonstrate that SimMTC achieves state-of-the-art clustering performance. The code has been made publicly available on GitHub at: https://github.com/haonanxin/SimMTC_code.
Haonan Xin, Zhezheng Hao, Zihua Zhao, Rong Wang 0001, Feiping Nie 0001
IEEE Trans. Image Process.4
2025 High-Order Anchor Graph-Based Clustering for Efficient Structured Proximity Matrix Learning
Zihua Zhao, Yuyu Jia, Fangyuan Xie, Rong Wang 0001
ADMA (4)1
2025 Multi-modal Medical Diagnosis via Large-small Model Collaboration
abstract
Recent advances in medical AI have shown a clear trend towards large models in healthcare. However, developing large models for multi-modal medical diagnosis remains challenging due to a lack of sufficient modal-complete medical data. Most existing multi-modal diagnostic models are relatively small and struggle with limited feature extraction capabilities. To bridge this gap, we propose AdaCoMed, an adaptive collaborative-learning framework that synergistically integrates the off-the-shelf medical single-modal large models with multi-modal small models. Our framework first employs a mixture-of-modality-experts (MoME) architecture to combine features extracted from multiple single-modal medical large models, and then introduces a novel adaptive co-learning mechanism to collaborate with a multi-modal small model. This co-learning mechanism, guided by an adaptive weighting strategy, dynamically balances the complementary strengths between the MoMEfused large model features and the cross-modal reasoning capabilities of the small model. Extensive experiments on two representative multi-modal medical datasets (MIMICIV-MM and MMIST ccRCC) across six modalities and four diagnostic tasks demonstrate consistent improvements over state-of-the-art baselines, making it a promising solution for real-world medical diagnosis applications. The code is available at https://github.com/Zoew420/AdaCoMed.
Zihua Zhao, Jiangchao Yao, Ya Zhang 0002, Jiajun Bu, Haishuai Wang
CVPR2
2025 Consensus Graph-Based Spectral Ensemble Clustering via Low-Rank Tensor Learning
abstract
Ensemble clustering using co-association matrices integrates multiple base clusterings but often overlooks interactions between crucial samples and base clusterings. This neglect can introduce noise and lead to information loss and instability. To address these issues, we propose the Consensus Graph-Based Spectral Ensemble Clustering via Low-Rank Tensor Learning (SECGTL) model. SECGTL organizes base clusterings into a third-order tensor and applies the Fast Fourier Transform (FFT) to capture inter-relations in the frequency domain. By rotating the tensor and minimizing the Tensor Schatten p-norm, SECGTL extracts shared information in a low-rank space, reducing noise and enhancing the learned common graph. With Laplacian rank constraints, SECGTL directly learns a graph with c-connected components, representing the clustering structure without post-processing. Extensive experiments on real-world datasets demonstrate SECGTL’s superior performance and robustness to noise.
Haonan Xin, Zihua Zhao, Jie Wang 0164, Rong Wang 0001
ICASSP3
2025 Differential-Informed Sample Selection Accelerates Multimodal Contrastive Learning
abstract
The remarkable success of contrastive-learning-based multimodal models has been greatly driven by training on ever-larger datasets with expensive compute consumption. Sample selection as an alternative efficient paradigm plays an important direction to accelerate the training process. However, recent advances on sample selection either mostly rely on an oracle model to offline select a high-quality coreset, which is limited in the cold-start scenarios, or focus on online selection based on real-time model predictions, which has not sufficiently or efficiently considered the noisy correspondence. To address this dilemma, we propose a novel Differential-Informed Sample Selection (DISSect) method, which accurately and efficiently discriminates the noisy correspondence for training acceleration. Specifically, we rethink the impact of noisy correspondence on contrastive learning and propose that the differential between the predicted correlation of the current model and that of a historical model is more informative to characterize sample quality. Based on this, we construct a robust differential-based sample selection and analyze its theoretical insights. Extensive experiments on three benchmark datasets and various downstream tasks demonstrate the consistent superiority of DISSect over current state-of-the-art methods. Source code is available at: https://github.com/MediaBrain-SJTU/DISSect.
Zihua Zhao, Feng Hong 0004, Mengxi Chen, Pengyi Chen, Benyuan Liu, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
ICCV1
2025 Bidirectional fusion for deep contrastive multi-view clustering
Jie Wang 0164, Weizhong Yu, Zihua Zhao, Zongcheng Miao, Feiping Nie 0001
Expert Syst. Appl.4
2025 Enhancing Clustering Performance With Tensorized High-Order Bipartite Graphs: A Structured Graph Learning Approach
abstract
Clustering based on structured graph learning involves acquiring a proximity matrix with an explicit clustering structure from the original one. However, the original proximity matrix often lacks some must-links compared to the groundtruth, constraining the upper bound of clustering performance. High-order proximity information can mitigate this limitation, yet traditional high-order proximity matrix-based methods are time-intensive. To tackle this, we propose the Tensorized High-order Bipartite Graphs-based structured proximity matrix learning method (THBG). Firstly, we introduce a high-order bipartite graph proximity matrix with a swift computation method, incorporating high-order information and significantly reducing computational overhead. Secondly, we apply tensor nuclear norm minimization to the tensor composed of high-order bipartite graphs, learning a low-rank tensor representation that effectively harnesses the consistency of high-order information. Concurrently, a structured bipartite graph proximity matrix with an explicit clustering structure is adaptively learned based on the low-rank tensor representation and Laplace rank constraint. Experimental results demonstrate the superiority and great potential of this method. Code available:https://anonymous.4open.science/r/THBG-D10D.
Zihua Zhao, Haonan Xin, Rong Wang 0001, Danyang Wu, Zheng Wang 0037, Feiping Nie 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Fuzzy Clustering via Orthogonal Tensor Decomposition on High-Order Anchor Graphs
abstract
Clusteringis an important unsupervised learning technique widely applied in data analysis and pattern recognition. Graph-based clustering methods have gained attention for their ability to effectively model complex data structures. However, traditional methods mainly rely on first-order proximity information, which struggles to capture high-order structural relationships between data points. Such limitation can significantly degrade clustering performance. To address this issue, we propose a novel fuzzy clustering approach leveraging orthogonal tensor decomposition on high-order anchor graphs (OTDHAG). Unlike conventional high-order graph-based methods that rely on self-multiplication of proximity matrices, which are computationally expensive, our method introduces the high-order anchor graph with low computational complexity, and exploits high-order proximity information via orthogonal tensor factorization on high-order anchor graphs. Meanwhile, tensor nuclear norm regularization is adopted to enhance clustering consistency, thereby improving clustering quality. Experiments on 10 benchmark datasets (e.g., MNIST, Wine, PIE) demonstrate that OTDHAG outperforms 14 state-of-the-art methods, achieving 92.3% accuracy on Wine (vs. 86.5% for DenoHOG) and 88.7% NMI on MNIST (vs. 82.1% for AOPL-Root). In conclusion, OTDHAG not only significantly improves clustering accuracy and computational efficiency but also provides a scalable solution for clustering tasks. Code available:https://anonymous.4open.science/r/HOAGTD-CED0.
Zihua Zhao, Xinyi Hui, Rong Wang 0001, Feiping Nie 0001
IEEE Trans. Fuzzy Syst.1
2025 Graph-Based Clustering: High-Order Bipartite Graph for Proximity Learning
abstract
Structured proximity matrix learning, one of the mainstream directions in clustering research, refers to learning a proximity matrix with an explicit clustering structure from the original first-order proximity matrix. Due to the complexity of the data structure, the original first-order proximity matrix always lacks some must-links compared to the groundtruth proximity matrix. It is worth noting that high-order proximity matrices can provide missed must-link information. However, the computation of high-order proximity matrices and clustering based on them are expensive. To solve the above problem, inspired by the anchor bipartite graph, we present a novel high-order bipartite graph proximity matrix and a fast method to compute it. This proposed high-order bipartite graph proximity matrix contains high-order proximity information and can significantly reduce the computational complexity of the whole clustering process. Furthermore, we introduce an efficient and simple high-order bipartite graph fusion framework that can adaptively assign weights to each order of the high-order bipartite graph matrices. Finally, under the Laplace rank constraint, a consensus structured bipartite graph proximity matrix is obtained. At the same time, an efficient solution algorithm is proposed for this model. The model's efficacy is underscored through rigorous experiments, highlighting its superior clustering performance and time efficiency. Code available:https://anonymous.4open.science/r/HBGC-F6C4.
Zihua Zhao, Danyang Wu, Rong Wang 0001, Zheng Wang 0037, Feiping Nie 0001, Xuelong Li 0001
IEEE Trans. Knowl. Data Eng.1
2024 Mitigating Noisy Correspondence by Geometrical Structure Consistency Learning
abstract
Noisy correspondence that refers to mismatches in cross-modal data pairs, is prevalent on human-annotated or web-crawled datasets. Prior approaches to leverage such data mainly consider the application of uni-modal noisy label learning without amending the impact on both cross-modal and intra-modal geometrical structures in multimodal learning. Actually, we find that both structures are effective to discriminate noisy correspondence through structural differences when being wellestablished. Inspired by this observation, we introduce a Geometrical Structure Consistency (GSC) method to infer the true correspon-dence. Specifically, GSC ensures the preservation of geometrical structures within and between modalities, allowing for the accurate discrimination of noisy samples based on structural differences. Utilizing these inferred true correspondence labels, GSC refines the learning of geometrical structures by filtering out the noisy samples. Experiments across four cross-modal datasets confirm that GSC effectively identifies noisy samples and significantly outperforms the current leading methods. Source code is available at: https://github.com/MediaBrain-SJTU/GSC.
Zihua Zhao, Mengxi Chen, Tianjie Dai, Jiangchao Yao, Bo Han 0003, Ya Zhang 0002, Yanfeng Wang 0001
CVPR1
2024 Exploring Training on Heterogeneous Data with Mixture of Low-rank Adapters
abstract
Training a unified model to take multiple targets into account is a trend towards artificial general intelligence. However, how to efficiently mitigate the training conflicts among heterogeneous data collected from different domains or tasks remains under-explored. In this study, we explore to leverage Mixture of Low-rank Adapters (MoLA) to mitigate conflicts in heterogeneous data training, which requires to jointly train the multiple low-rank adapters and their shared backbone. Specifically, we introduce two variants of MoLA, namely, MoLA-Grad and MoLA-Router, to respectively handle the target-aware and target-agnostic scenarios during inference. The former uses task identifiers to assign personalized low-rank adapters to each task, disentangling task-specific knowledge towards their adapters, thereby mitigating heterogeneity conflicts. The latter uses a novel Task-wise Decorrelation (TwD) loss to intervene the router to learn oriented weight combinations of adapters to homogeneous tasks, achieving similar effects. We conduct comprehensive experiments to verify the superiority of MoLA over previous state-of-the-art methods and present in-depth analysis on its working mechanism. Source code is available at: https://github.com/MediaBrain-SJTU/MoLA
Zihua Zhao, Haolin Li 0001, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
ICML2
2024 Probabilistic Conformal Distillation for Enhancing Missing Modality Robustness
abstract
Multimodal models trained on modality-complete data are plagued with severe performance degradation when encountering modality-missing data. Prevalent cross-modal knowledge distillation-based methods precisely align the representation of modality-missing data and that of its modality-complete counterpart to enhance robustness. However, due to the irreparable information asymmetry, this determinate alignment is too stringent, easily inducing modality-missing features to capture spurious factors erroneously. In this paper, a novel multimodal Probabilistic Conformal Distillation (PCD) method is proposed, which considers the inherent indeterminacy in this alignment. Given a modality-missing input, our goal is to learn the unknown Probability Density Function (PDF) of the mapped variables in the modality-complete space, rather than relying on the brute-force point alignment. Specifically, PCD models the modality-missing feature as a probabilistic distribution, enabling it to satisfy two characteristics of the PDF. One is the extremes of probabilities of modality-complete feature points on the PDF, and the other is the geometric consistency between the modeled distributions and the peak points of different PDFs. Extensive experiments on a range of benchmark datasets demonstrate the superiority of PCD over state-of-the-art methods. Code is available at: https://github.com/mxchen-mc/PCD.
Mengxi Chen, Fei Zhang 0016, Zihua Zhao, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001
NeurIPS3
2024 An Balanced, and Scalable Graph-Based Multiview Clustering Method
abstract
In recent years, graph-based multiview clustering methods have become a research hotspot in the clustering field. However, most existing methods lack consideration of cluster balance in their results. In fact, cluster balance is crucial in many real-world scenarios. Additionally, graph-based multiview clustering methods often suffer from high time consumption and cannot handle large-scale datasets. To address these issues, this paper proposes a novel graph-based multiview clustering method. The method is built upon the bipartite graph. Specifically, it employs a label propagation mechanism to update the smaller anchor label matrix rather than the sample label matrix, significantly reducing the computational cost. The introduced balance constraint in the proposed model contributes to achieving balanced clustering results. The entire clustering model combines information from multiple views through graph fusion. The joint graph and view weight parameters in the model are obtained through task-driven self-supervised learning. Moreover, the model can directly obtain clustering results without the need for the two-stage processing typically used in general spectral clustering. Finally, extensive experiments on toy datasets and real-world datasets are conducted to validate the superiority of the proposed method in terms of clustering performance, clustering balance, and time expenditure.
Zihua Zhao, Feiping Nie 0001, Rong Wang 0001, Zheng Wang 0037, Xuelong Li 0001
IEEE Trans. Knowl. Data Eng.1
2024 Graph Joint Representation Clustering via Penalized Graph Contrastive Learning
abstract
Graph clustering based on graph contrastive learning (GCL) is one of the dominant paradigms in the current graph clustering research field. However, those GCL-based methods often yield false negative samples, which can distort the learned representations and limit clustering performance. In order to alleviate this issue, we propose the idea of maintaining mutual information (MI) between the representations and the inputs to mitigate the loss of semantic information of false negative samples. We demonstrate the validity of this proposal through relevant experiments. Since maximizing MI can be approximately replaced by minimizing reconstruction error, we further propose a graph clustering method based on GCL penalized by reconstruction error, in which our carefully designed reconstruction decoder, as well as reconstruction error term, improve the clustering performance. In addition, we use a pseudo-label-guided strategy to improve the GCL process and further alleviate the problem of false negative samples. Our experiment results demonstrate the superiority and great potential of our proposed graph clustering method compared with state-of-the-art algorithms.
Zihua Zhao, Rong Wang 0001, Zheng Wang 0037, Feiping Nie 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 Reinforcing Pretrained Models for Generating Attractive Text Advertisements
abstract
We study how pretrained language models can be enhanced by using deep reinforcement learning to generate attractive text advertisements that reach the high quality standard of real-world advertiser mediums. To improve ad attractiveness without hampering user experience, we propose a model-based reinforcement learning framework for text ad generation, which constructs a model for the environment dynamics and avoids large sample complexity. Based on the framework, we develop Masked-Sequence Policy Gradient, a reinforcement learning algorithm that integrates efficiently with pretrained models and explores the action space effectively. Our method has been deployed to production in Microsoft Bing. Automatic offline experiments, human evaluation, and online experiments demonstrate the superior performance of our method.
Xiting Wang, Xinwei Gu, Zihua Zhao, Yulan Yan, Bhuvan Middha, Xing Xie 0001
KDD4