Zihan Fang 0002

dblp:273/3457-2 · DBLP profile ↗
← Back
28ranked-venue papers
8as first author
28since 2021 · last 2026
0000-0003-2009-9039ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DIN: Dual Impulse Network for Multi-view Representation Learning
abstract
Multi-view representation learning, which utilizes multiple channels to improve perceptual accuracy, is recognized for its effectiveness in the analysis of multi-view data. However, deploying these methods in real-world scenarios presents two primary challenges. 1) Lack of Variegation: Multi-view representation techniques commonly observe along a singular axis, i.e., the attribute axis; 2) Insufficient Relationship: Most multi-view models lack mechanisms for exploring potential relationships between attribute axis and channel axis. To mitigate these obstacles, we design a Dual Impulse Network framework for multi-view representation learning (DIN) to train a feature representation. In this framework, a strategy observed along the channel axis and attribute axis simultaneously is introduced, and two different representations are generated by two analogous impulse networks, which are capable of extracting information corresponding to different axes. Furthermore, we incorporate an integration network that analyzes the potential relationship between attribute axis and channel axis to generate two attention matrices. The final two feature representations derived from these attention matrices are aggregated to amplify the expression of internal information. Comprehensive experimental results support the efficacy and superiority of the proposed framework, demonstrating improvements in classification performance compared to state-of-the-art methods.
Yilin Wu 0001, Weihong Lin, Renjie Lin, Zihan Fang 0002, Shide Du, Shiping Wang
AAAI4
2026 Revisiting multi-view semi-supervised classification: a reinforcement learning perspective
Zhicheng Wei, Zhiling Cai, Zhibin Shi, Zihan Fang 0002, Mingjian Fu 0001, Shiping Wang
Appl. Intell.4
2026 Prior-constrained multi-view clustering with evidence theory
Zhengnan Chen, Zihan Fang 0002, Shiping Wang
Neurocomputing3
2026 Attention meets convolution: Dual-channel multi-view fusion network
Fu Zhao, Zihan Fang 0002, Shide Du, Zhiling Cai, Hongju Cheng, Shiping Wang
Inf. Sci.2
2026 Beyond local aggregation: Global graph contrastive learning for multi-view fusion
Xueyang Min, Zihan Fang 0002, Weihong Lin, Shiping Wang
Neural Networks3
2026 Harnessing noisy LLM annotations: Confidence-calibrated node selection on text-attributed graphs
Zihan Fang 0002, Shide Du, Zhihao Wu 0003, Zhiling Cai, Yanchao Tan, Shiping Wang, Zhouchen Lin
Pattern Recognit.1
2026 SOI-Net: Structural Optimization-Inspired Interpretable Network for Incomplete Multi-View Clustering
abstract
Data missing is a common issue in real-world applications, posing significant challenges for incomplete data processing. Traditional incomplete multi-view clustering methods rely on manually-designed optimization problems based on prior interpretable knowledge, considering the full utilization of available data. However, their limited feature extraction capability may become a bottleneck. In contrast, deep optimization methods leverage learning-based nonlinear transformations for clustering. They primarily achieve data imputation through the generalization ability of deep models, but their model interpretability may be limited by the black-box nature. Moreover, most existing methods only explore the structure of each view independently, where these structures are fixed and cannot form a complete unified structure. To address these issues, we propose a Structural Optimization-inspired Interpretable Network (SOI-Net) for incomplete multi-view clustering. Specifically, we project the features of all views into a unified representation space with the un-missing information of the views as constraints. By optimizing consistent structural information, we preserve the structures of missing modalities in the unified representation space, thereby mitigating the impact of missing data. Meanwhile, we derive network components based on the optimization problem to guide the learning of structure and representation. The practical significance of these network components provides model design-level interpretability. Extensive experiments on six datasets validate the effectiveness of SOI-Net in handling incomplete multi-view clustering task.
Zihan Fang 0002, Zhiling Cai, Shide Du, Wei Huang 0037, Shiping Wang
IEEE Trans. Multim.2
2026 Interpretable Multi-View Representation Learning Towards Complex Scenes: From Homogeneity to Heterogeneity
abstract
Multi-view representation learning is recognized for its effectiveness in multi-source data analysis, yet it faces significant challenges: 1) Deep model structures remain opaque, lacking interpretability; 2) Research on compatibility models toward multi-feature and multi-relationship data is insufficient. In this paper, we introduce an interpretable multi-view representation learning framework specifically designed for the complex multi view scene. The barrier to achieving compatibility stems from the need to simultaneously process the homogeneous information inherent in multiple features and the heterogeneous characteristic of multi-relation data. To address this, we design an objective function solved by iterative methods to learn comprehensive relationships and the consistent representation. The introduction of comprehensive relationships aims to mitigate mutual interference among different data types while combining information abstracted from original features and relationships into a unified representation. We then convert iterative solutions into feed-forward network layers with embedded learnable modules, resulting in a deep network architecture that is interpretable at the design level. Extensive experimental results demonstrate the superior performance of the proposed method over state-of-the art approaches.
Ying Zou 0024, Zihan Fang 0002, Shide Du, Yilin Wu 0001, Hong Zhao 0002, Shiping Wang
IEEE Trans. Multim.2
2026 Be Reliable: An Interpretable Attribute-Oriented Representation Learning Framework
abstract
Representation learning techniques effectively unveil latent patterns within raw data. However, the learning process is often marred by uncertainties, such as variations in data quality and heterogeneous scenarios, which greatly affect the reliability of representation learning. In this article, we introduce a reliable representation learning framework to establish a connection between data attributes and modeling strategies, namely the interpretable attribute-oriented representation learning framework. First, by focusing on the inherent knowledge embedded in the data, we decouple it into four principal attributes: fidelity, topology, invariance, and discriminability. To explicitly address these attributes, we incorporate them into an optimization-derived framework using corresponding general loss terms. Furthermore, by treating the iterative solution process as a bridge, each derived network module possesses traceable interpretability, thus laying a reliable foundation. Ultimately, we extend the proposed framework to multisource heterogeneous scenarios, enabling it to adapt to complex environments while maintaining reliability. In essence, our work aims to seamlessly integrate deep representations with prior knowledge during the learning process, thereby creating a solid basis for dependable modeling. Networks derived from the proposed framework achieve promising results, particularly in complex multisource heterogeneous environments, demonstrating both their effectiveness and reliability. The code is available at https://github.com/ZihanFang11/2025_AORLNet_TNNLS.
Zihan Fang 0002, Shide Du, Ying Zou 0024, Yanchao Tan, Na Song, Shiping Wang
IEEE Trans. Neural Networks Learn. Syst.1
2025 OpenViewer: Openness-Aware Multi-View Learning
abstract
Multi-view learning methods leverage multiple data sources to enhance perception by mining correlations across views, typically relying on predefined categories. However, deploying these models in real-world scenarios presents two primary openness challenges. 1) Lack of Interpretability: The integration mechanisms of multi-view data in existing black-box models remain poorly explained; 2) Insufficient Generalization: Most models are not adapted to multi-view scenarios involving unknown categories. To address these challenges, we propose OpenViewer, an openness-aware multi-view learning framework with theoretical support. This framework begins with a Pseudo-Unknown Sample Generation Mechanism to efficiently simulate open multi-view environments and previously adapt to potential unknown samples. Subsequently, we introduce an Expression-Enhanced Deep Unfolding Network to intuitively promote interpretability by systematically constructing functional prior-mapping modules and effectively providing a more transparent integration mechanism for multi-view data. Additionally, we establish a Perception-Augmented Open-Set Training Regime to significantly enhance generalization by precisely boosting confidences for known categories and carefully suppressing inappropriate confidences for unknown ones. Experimental results demonstrate that OpenViewer effectively addresses openness challenges while ensuring recognition performance for both known and unknown samples.
Shide Du, Zihan Fang 0002, Yanchao Tan, Changwei Wang 0001, Shiping Wang, Wenzhong Guo
AAAI2
2025 Constraint-Aware Multi-View Clustering via Graph Contrastive Learning
Zhengnan Chen, Zihan Fang 0002, Shide Du, Shiping Wang
ADMA (2)3
2025 HiTuner: Hierarchical Semantic Fusion Model Fine-Tuning on Text-Attributed Graphs
abstract
Text-Attributed Graphs (TAGs) are vital for modeling entity relationships across various domains. Graph Neural Networks have become cornerstone for processing graph structures, while the integration of text attributes remains a prominent research. The development of Large Language Models (LLMs) provides new opportunities for advancing textual encoding in TAGs. However, LLMs face challenges in specialized domains due to their limited task-specific knowledge, and fine-tuning them for specific tasks demands significant resources. To cope with the above challenges, we propose HiTuner, a novel framework that leverages fine-tuned Pre-trained Language Models (PLMs) with domain expertise as tuner to enhance the hierarchical LLM contextualized representations for modeling TAGs. Specifically, we first strategically select hierarchical hidden states of LLM to form a set of diverse and complementary descriptions as input for the sparse projection operator. Concurrently, a hybrid representation learning is developed to amalgamate the broad linguistic comprehension of LLMs with task-specific insights of the fine-tuned PLMs. Finally, HiTuner employs a confidence network to adaptively fuse the semantically-augmented representations. Empirical results across benchmark datasets spanning various domains validate the effectiveness of the proposed framework. Our codes are available at: https://github.com/ZihanFang11/HiTuner
Zihan Fang 0002, Zhiling Cai, Yuxuan Zheng, Shide Du, Yanchao Tan, Shiping Wang
IJCAI1
2025 Strategy-Architecture Synergy: A Multi-View Graph Contrastive Paradigm for Consistent Representations
abstract
Facing the growing diversity of multi-view data, multi-view graph-based models have made encouraging progress in handling multi-view data modeled as graphs. Graph Contrastive Learning (GCL) naturally fits multi-view graph data by treating their inherent views as augmentations. However, the development of GCL on multi-view graph data is still in the infant stage. Challenges remain in designing strategies that coordinate preprocessing and contrastive learning, and in developing model architectures that automatically meet the needs of diverse views. To tackle these, we propose a framework named CAMEL, which refines consistency learning by introducing a tailored contrastive paradigm for multi-view graphs. Initially, we theoretically analyze the positive effect of edge-dropping preprocessing on the consistency and quantify the factors that influence it. Paired with a learnable model architecture, the proposed adaptive edge-dropping preprocessing strategy is guided by dynamic topology, making the heterogeneity of views more controllable and better aligned with contrastive learning. Finally, we design a neighborhood consistency multi-view contrastive objective that enhances consistency information interaction by extending positive samples. Extensive experiments on downstream tasks, including node classification and clustering, validate the superiority of our proposed model.
Shuman Zhuang, Zhihao Wu 0003, Zihan Fang 0002, Jiali Yin, Ximeng Liu
IJCAI4
2025 LargeMvC-Net: Anchor-based Deep Unfolding Network for Large-scale Multi-view Clustering
abstract
Deep anchor-based multi-view clustering methods enhance the scalability of neural networks by utilizing representative anchors to reduce the computational complexity of large-scale clustering. Despite their scalability advantages, existing approaches often incorporate anchor structures in a heuristic or task-agnostic manner, either through post-hoc graph construction or as auxiliary components for message passing. Such designs overlook the core structural demands of anchor-based clustering, neglecting key optimization principles. To bridge this gap, we revisit the underlying optimization problem of large-scale anchor-based multi-view clustering and unfold its iterative solution into a novel deep network architecture, termed LargeMvC-Net. The proposed model decomposes the anchor-based clustering process into three modules: RepresentModule, NoiseModule, and AnchorModule, corresponding to representation learning, noise suppression, and anchor estimation. Each module is derived by unfolding a step of the original optimization procedure into a dedicated network component, providing structural clarity and optimization traceability. In addition, an unsupervised reconstruction loss aligns each view with the anchor-induced latent space, encouraging consistent clustering structures across views. Extensive experiments on several large-scale multi-view benchmarks show that LargeMvC-Net consistently outperforms state-of-the-art methods in terms of both effectiveness and scalability. The source data, code. https://github.com/dushide/LargeMvC-Net_ACMMM_2025, and extended version http://arxiv.org/abs/2507.20980 are available.
Shide Du, Zihan Fang 0002, Wendi Zhao, Yilin Wu 0001, Changwei Wang 0001, Shiping Wang
ACM Multimedia3
2025 Enhancing Multi-view Open-set Learning via Ambiguity Uncertainty Calibration and View-wise Debiasing
abstract
Existing multi-view learning models struggle in open-set scenarios due to their implicit assumption of class completeness. Moreover, static view-induced biases, which arise from spurious view-label associations formed during training, further degrade their ability to recognize unknown categories. In this paper, we propose a multi-view open-set learning framework via ambiguity uncertainty calibration and view-wise debiasing. To simulate ambiguous samples, we design O-Mix, a novel synthesis strategy to generate virtual samples with calibrated open-set ambiguity uncertainty. These samples are further processed by an auxiliary ambiguity perception network that captures atypical patterns for improved open-set adaptation. Furthermore, we incorporate an HSIC-based contrastive debiasing module that enforces independence between view-specific ambiguous and view-consistent representations, encouraging the model to learn generalizable features. Extensive experiments on diverse multi-view benchmarks demonstrate that the proposed framework consistently enhances unknown-class recognition while preserving strong closed-set performance. The source code are available https://github.com/ZihanFang11/2025_MOCD_ACMMM
Zihan Fang 0002, Lan Du 0002, Shide Du, Zhiling Cai, Shiping Wang
ACM Multimedia1
2025 Optimization-oriented multi-view representation learning in implicit bi-topological spaces
Shiyang Lan, Shide Du, Zihan Fang 0002, Zhiling Cai, Wei Huang 0037, Shiping Wang
Inf. Sci.3
2025 Order-flexible graph attention network for multi-view subspace clustering
Gangshuo Bao, Yongquan Shi, Yueyang Pi, Zihan Fang 0002, Shiping Wang
Knowl. Based Syst.4
2025 MOAL: Multi-view Out-of-distribution Awareness Learning
Xuzheng Wang, Zihan Fang 0002, Shide Du, Wenzhong Guo, Shiping Wang
Neural Networks2
2025 A Universal Interpretable Multiview Clustering Framework: From Homogeneity to Heterogeneity
abstract
Traditional multiview clustering relies on manually-designed optimization problems based on prior interpretable knowledge to cluster objects with similar attributes, but it may be hindered by limited feature extraction capability. In contrast, deep multiview clustering overcomes this limitation by utilizing learning-based nonlinear transformations for clustering, but it may be restricted by model interpretability caused by blackbox networks. Besides, previous multiview clustering methods only consider homogeneous or heterogeneous situations, resulting in restricted scalability and generalizability. To address the aforementioned issues, we design a universal interpretable clustering framework that accommodates both homogeneous and heterogeneous multiview scenarios. To realize this purpose: 1) we revisit the interpretable knowledge-driven design architecture of traditional multiview methods and formulate clustering optimization problems on multiview data attributes in homogeneous scenarios; 2) the optimization problem is leveraged to derive network modules that learn shared and self-expressive representations for clustering, with the practical meaning of each network component providing model design-level interpretability; 3) the proposed method is extended from homogeneous to heterogeneous scenarios, enhancing its universality for a broader spectrum of multiview clustering tasks; and 4) tailored training loss for the clustering task is employed to inversely enhance the affinity between objects with similar attributes. Extensive experimental results on both homogeneous and heterogeneous multiview datasets demonstrate the superior effectiveness and adaptability of the proposed framework compared to state-of-the-art clustering methods.
Zihan Fang 0002, Shide Du, Shiping Wang
IEEE Trans. Comput. Soc. Syst.2
2024 Beyond the Known: Ambiguity-Aware Multi-view Learning
abstract
The inherent variability and unpredictability in open multi-view learning scenarios infuse considerable ambiguity into the learning and decision-making processes of predictors. This demands that predictors not only recognize familiar patterns but also adaptively interpret unknown ones out of training scope. To address this challenge, we propose an Ambiguity-Aware Multi-view Learning Framework, which integrates four synergistic modules into an end-to-end framework to achieve generalizability and reliability beyond the known. By introducing the mixed samples to broaden the learning sample space, accompanied by corresponding soft labels to encapsulate their inherent uncertainty, the proposed method adapts to the distribution of potentially unknown samples in advance. Furthermore, an instance-level sparse inference is implemented to learn sparse approximated points in the multiple view embedding space, and individual view representations are gated by view-level confidence mappings. Finally, a multi-view consistent representation is obtained by dynamically assigning weights based on the degree of cluster-level dispersion. Extensive experiments demonstrate that our approach is effective and stable compared with other state-of-the-art methods in open-world recognition situations.
Zihan Fang 0002, Shide Du, Shiping Wang
ACM Multimedia1
2024 IMPRL-Net: interpretable multi-view proximity representation learning network
Shiyang Lan, Zihan Fang 0002, Shide Du, Zhiling Cai, Shiping Wang
Neural Comput. Appl.2
2024 Multi-view heterogeneous graph learning with compressed hypergraph neural networks
Aiping Huang, Zihan Fang 0002, Zhihao Wu 0003, Yanchao Tan, Peng Han 0005, Shiping Wang, Le Zhang 0001
Neural Networks2
2024 Attention-based stackable graph convolutional network for multi-view learning
Ying Zou 0024, Zihan Fang 0002, Shiping Wang
Neural Networks4
2024 Revisiting multi-view learning: A perspective of implicitly heterogeneous Graph Convolutional Network
Ying Zou 0024, Zihan Fang 0002, Zhihao Wu 0003, Chenghui Zheng, Shiping Wang
Neural Networks2
2024 Representation Learning Meets Optimization-Derived Networks: From Single-View to Multi-View
abstract
Existing representation learning approaches lie predominantly in designing models empirically without rigorous mathematical guidelines, neglecting interpretation in terms of modeling. In this work, we propose an optimization-derived representation learning network that embraces both interpretation and extensibility. To ensure interpretability at the design level, we adopt a transparent approach in customizing the representation learning network from an optimization perspective. This involves modularly stitching together components to meet specific requirements, enhancing flexibility and generality. Then, we convert the iterative solution of the convex optimization objective into the corresponding feed-forward network layers by embedding learnable modules. These above optimization-derived layers are seamlessly integrated into a deep neural network architecture, allowing for training in an end-to-end fashion. Furthermore, extra view-wise weights are introduced for multi-view learning to discriminate the contributions of representations from different views. The proposed method outperforms several advanced approaches on semi-supervised classification tasks, demonstrating its feasibility and effectiveness.
Zihan Fang 0002, Shide Du, Zhiling Cai, Shiyang Lan, Yanchao Tan, Shiping Wang
IEEE Trans. Multim.1
2023 DMRL-Net: Differentiable Multi-view Representation Learning Network
abstract
In recent years, deep multi-view representation learning has made considerable achievements due to its excellent nonlinear mapping capability. Yet its development is limited by the challenge of interpreting the underlying structure. For this reason, we propose a differentiable multi-view representation learning network to address the aforementioned issue, which is equipped with the interpretable working mechanism of sparse low-rank decomposition and outstanding representation ability of neural networks. The network is constructed by stacking multiple differentiable blocks that are explicitly reformulated from the optimization objective exhibiting interpretability. Benefiting from end-to-end optimization of deep networks, it can efficiently learn an interpretable deep representation of high-dimensional features from multi-view data. Extensive experimental results on several benchmark multi-view datasets demonstrate the effectiveness of the learned representation in comparison to several state-of-the-art algorithms.
Zihan Fang 0002, Shide Du, Shiping Wang
ICME1
2023 Bridging Trustworthiness and Open-World Learning: An Exploratory Neural Approach for Enhancing Interpretability, Generalization, and Robustness
abstract
As researchers strive to narrow the gap between machine intelligence and human through the development of artificial intelligence multimedia technologies, it is imperative that we recognize the critical importance of trustworthiness in open-world, which has become ubiquitous in all aspects of daily life for everyone. However, several challenges may create a crisis of trust in current open-world artificial multimedia systems that need to be bridged: 1) Insufficient explanation of predictive results; 2) Inadequate generalization for learning models; 3) Poor adaptability to uncertain environments. Consequently, we explore a neural program to bridge trustworthiness and open-world learning, extending from single-modal to multi-modal scenarios for readers.1) To enhance design-level interpretability, we first customize trustworthy networks with specific physical meanings; 2) We then design environmental well-being task-interfaces via flexible learning regularizers for improving the generalization of trustworthy learning; 3) We propose to increase the robustness of trustworthy learning by integrating open-world recognition losses with agent mechanisms. Eventually, we enhance various trustworthy properties through the establishment of design-level explainability, environmental well-being task-interfaces and open-world recognition programs. As a result, these designed open-world protocols are applicable across a wide range of surroundings, under open-world multimedia recognition scenarios with significant performance improvements observed.
Shide Du, Zihan Fang 0002, Shiyang Lan, Yanchao Tan, Manuel Günther, Shiping Wang, Wenzhong Guo
ACM Multimedia2
2023 DBO-Net: Differentiable bi-level optimization network for multi-view clustering
Zihan Fang 0002, Shide Du, Xincan Lin, Jinbin Yang, Shiping Wang, Yiqing Shi
Inf. Sci.1