Mingze Yao

dblp:300/2082 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Conditional Prompt Learning via Degradation Perception for Underwater Image Enhancement
abstract
Underwater Image Enhancement (UIE) focuses on improving visual quality from various underwater scenes. Existing methods simplistically treat various degradations as homogeneous, disregarding their intrinsic connections and causing models to blindly learn, resulting in conflicting optimization goals and visual distortions. To address above limitations, we propose a Conditional Prompt Learning via Degradation Perception (CPLDP) model, which employs conditional prompt as degradation perception priors and guides underwater image enhancement. Specifically, we show that the natural language prompts not only promote distinguishing different degraded images, but also aid in exploring more details with semantic information. Therefore, our method generates five key degradation prompts (green/blue/green-blue color casts, uneven illumination and haze) with conditional prompt learning. Subsequently, considering the intrinsic relationships among different degradations, we employ degradation perceptions as priors and fine-tune the learning strategy to enhance underwater images. During training, an adaptive loss function with multi-degradations is designed, allowing it to effectively handle the task conflicts among multiple underwater degradations. Additionally, we conduct a human visual-based underwater dataset with various degradation types by subjective statistics. Extensive experiments on both full-reference and non-reference datasets demonstrate that our CPLDP can achieve better visual results and outperforms state-of-the-art UIE methods across various degradation scenarios.
Mingze Yao, Zhiying Jiang, Xianping Fu, Huibing Wang
AAAI1
2026 Tensorized Fine-Grained Incomplete Multiview Clustering via Intrinsic Structure Recovery
abstract
Incomplete multiview clustering (IMVC) focuses on exploiting the complementary and consistent information from multiple incomplete views for dividing unlabeled multiview data into corresponding clusters. Most existing methods seek to recover the missing samples of the views while inevitably having an influence on the intrinsic structure of the original space. Moreover, previous IMVC algorithms treat the samples of each view equally, which learns the consensus representation in a view-level manner and thus neglects that different views contribute to each individual sample diversely. These situations are a limitation to effective recovery of absent samples and fuse the heterogeneous information, which leads to suboptimal clustering performance. To tackle the above issues, this article proposes a novel approach called tensorized fine-grained IMVC via intrinsic structure recovery (TFIR), which flexibly recovers the intrinsic structure of the original data and attains the unified representation based on the fine-grained fusion strategy. Specifically, TFIR infers the incomplete data via available instances’ relations to improve the accuracy of learned intersample-specific representations. Afterward, TFIR stacks the specific representations of multiple views into a tensor to preserve the intrasample consistency. Finally, TFIR explores the complementarity among various views from the fine-grained sample perspective to obtain a rich underlying structure for clustering. Experimental results on the eight benchmark datasets clearly verify the remarkable superiority of TFIR compared with the most state-of-the-art methods.
Huibing Wang, Luyan Cui, Mingze Yao, Xianping Fu
IEEE Trans. Syst. Man Cybern. Syst.3
2025 Consensus-Guided Incomplete Multi-view Clustering via Cross-view Affinities Learning
abstract
Incomplete multi-view clustering (IMC) has garnered substantial attention due to its capacity to handle unlabeled data. Existing methods predominantly explore pairwise consistency between every two views. However, such consistency is highly susceptible to missing samples and outliers within a certain view and thus deviates from the true clustering distribution. Moreover, dual-view interaction neglects the collaboration effects of multiple views, making it challenging to capture the holistic characteristics across views. In response to these issues, we propose a novel Consensus-Guided Incomplete Multi-view Clustering via Cross-view Affinities Learning (CAL). Specifically, CAL reconstructs views with available instances to mine sample-wise affinities and harness comprehensive content information within views. Subsequently, to extract clean structural information, CAL imposes a structured sparse constraint on the representation tensor to eliminate biased errors. Furthermore, by integrating the consensus representation into a representation tensor, CAL can employ high-order interaction of multiple views to depict the semantic correlation between views while acquiring a unified structural graph across multiple views. Extensive experiments on seven benchmark datasets demonstrate that CAL outperforms some state-of-the-art methods in clustering performance. The code is available at https://github.com/whbdmu/CAL.
Huibing Wang, Jinjia Peng, Yawei Chen, Mingze Yao, Xianping Fu, Yang Wang 0023
IJCAI5
2025 Scalable Multi-view Clustering based on Tight Anchor Distribution
Yawei Chen, Huibing Wang, Mingze Yao, Jinjia Peng, Guangqi Jiang, Jiqing Zhang
ACM Multimedia3
2025 Dual-Constraint Multi-view Fuzzy Clustering with Scalable Anchor Graph Learning
Luyan Cui, Huibing Wang, Yawei Chen, Mingze Yao, Xianping Fu, Jiqing Zhang
ACM Multimedia4
2025 Consensus guided incomplete multi-view clustering via geometric consistency learning
Huibing Wang, Mingze Yao, Yawei Chen, Jinjia Peng, Guangqi Jiang, Xianping Fu
Appl. Intell.3
2025 Detail-focused and polarization-guided multi-modality fusion for underwater image clarity enhancing
Mingze Yao, Huibing Wang, Xianping Fu
Eng. Appl. Artif. Intell.1
2025 Focus More on What? Guiding Multi-Task Training for End-to-End Person Search
abstract
End-to-end person search is a research domain that executes the task of locating and identifying a target individual from a large number of scene images via a multi-task framework. However, a major challenge for the learning of end-to-end methods is the inherent conflicts between the two sub-tasks: person detection focuses on identifying generic features of persons, while person re-identification (Re-ID) strives to find unique, distinguishing features for matching the target person. Unlike previous research focusing on model architectures, this paper delves into the end-to-end person search training process. We find that the unbalanced and conflicting training issues significantly impair the learning efficiency of the Re-ID sub-task, which directly influences person search accuracy. To address this, we propose a novel Guiding Multi-Task Training (GMT) framework that facilitates end-to-end balanced learning for person search. We introduce a Guiding Multi-Task Harmonious Learning (GMHL) module, which decouple the features and then performs intra- and cross-task feature interaction to enhance the learning of each sub-task. Moreover, GMT employs a Balancing Multi-Task Oriented Fusing (BMOF) method to explicitly enhance Re-ID sub-task learning through additional Re-ID training and target-guided multi-model parameters fusion. Extensive experiments on 2 benchmark datasets, CUHK-SYSU and PRW, show that GMT achieves leading performance with 96.0% mAP and 61.3% mAP, respectively.
Boyu Cai, Huibing Wang, Mingze Yao, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.3
2025 Tensor Completion Framework by Graph Refinement for Incomplete Multi-View Clustering
abstract
Incomplete Multi-view Clustering (IMVC) endeavors to harness information from multiple incomplete views to partition multi-view data into their respective clusters. How to recover missing information with lossless fidelity is the core of IMVC, which is of vital importance but challenging. Most of the existing methods include a feature recovery step to mitigate the negative impact of missing samples on the feature graph, however, these IMVC algorithms simply utilize the correlation between samples to recover the relationship between the unmissing instances and the missing instances while ignoring the consistency between views, which leads to often unsatisfactory recovery results. In addition, previous IMVC algorithms focus more on the recovery of incomplete data, ignoring the effect of the error term on incomplete graphs. This can mislead the recovery process of IMVC algorithm and the feature graph can be affected by anomalous information, which leads to degradation of clustering performance. To address this gap, this paper introduces the Tensor Completion Framework by Graph Refinement for Incomplete Multi-view Clustering (IMVC-TGR). IMVC-TGR separates the redundant information in each affine graph by graph refinement operation, aiming to mitigate the negative impact of error terms and redundant information on the feature graph during the recovery process. Meanwhile, IMVC-TGR stacks the feature graphs into tensors to explore intra-view correlation and inter-view consistency, so as to recover the relationship between missing samples and non-missing samples, and improve the quality of the feature graphs. Finally, IMVC-TGR introduces semantic consistency constraints and self-weighted fusion strategies into the high-quality feature graphs, aiming at preserving the complementary information between different views while balancing the contributions of the refined representation matrices of different views. The experimental results on multiple different datasets indicate that IMVC-TGR can achieve state-of-the-art performance.
Huibing Wang, Yawei Chen, Mingze Yao, Jinjia Peng, Xianping Fu
IEEE Trans. Multim.3
2025 Between/Within View Information Completing for Tensorial Incomplete Multi-View Clustering
abstract
Incomplete Multi-view Clustering (IMvC) receives increasing attention due to its effectiveness in solving data-missing problems. With the information loss in incomplete situations, the core of IMvC needs to consider effectively overcoming the challenge of missing views, that is, exploring the underlying correlations from available data and recovering the missing information. However, most existing IMvC methods overemphasize the recovery-first principle with integrating the existing data from different views while neglecting the influence of view consistency in IMvC task together with valuable within view information. In this paper, a novel Between/Within View Information Completing for Tensorial Incomplete Multi-view Clustering (BWIC-TIMC) has been proposed, in which between/within view information is jointly exploited for effectively completing the missing views. Specifically, the proposed method designs a dual tensor constraint module, which focuses on simultaneously exploring the view-specific correlations of incomplete views and enforcing the between view consistency across different views. With the dual tensor constraint, between/within view information can be effectively integrated for completing missing views for IMvC task. Furthermore, in order to balance different contributions of multiple views and alleviate the problem of feature degeneration, BWIC-TIMC implements an adaptive fusion graph learning strategy for consensus representation learning. Extensive comparative experiments with the-state-of-art baselines can demonstrate the effectiveness of BWIC-TIMC.
Mingze Yao, Huibing Wang, Yawei Chen, Xianping Fu
IEEE Trans. Multim.1
2024 Cascade transformers with dynamic attention for video question answering
abstract
Visual question answering (VQA) has become a hot study topic with challenging motivation of correctly answering the videos or images questions in recent years. However, the existing VQA model mostly aimed at answering questions about images and performed poorly in the video question answering (VideoQA) domain. VideoQA needs to simultaneously consider the correlations between video frames and the dynamic information of multiple objects in video. Therefore, we propose a novel Cascade Transformers with Dynamic Attention for Video Question Answering (CTDA-QA), which aims to simultaneously solve the above considerations. Specifically, the proposed CTDA-QA model utilizes multiple transformers structure to encode videos for reasoning complex spatial and temporal information, which is different from the previous recurrent neural network methods. Besides, in order to effectively capture the dynamic information from various scenarios in videos, a flexible attention module has been proposed to explore the essential relations between objects in a dynamic timeline. Finally, to avoid spurious answers and fully explore the cross-modal relationships, a mixed-supervised learning strategy is designed for optimizing the reasoning tasks. The experiments on several benchmark video question–answer datasets clearly verify the performance and effectiveness of CTDA-QA, which contains the results in contrast to the state-of-the-art methods. Besides, the provided ablation study and visualization results further reveal the potential of CTDA-QA.
Tingfei Yan, Mingze Yao, Huibing Wang
Comput. Vis. Image Underst.3
2024 Manifold-Based Incomplete Multi-View Clustering via Bi-Consistency Guidance
abstract
Incomplete multi-view clustering primarily focuses on dividing unlabeled data into corresponding categories with missing instances, and has received intensive attention due to its superiority in real applications. Considering the influence of incomplete data, the existing methods mostly attempt to recover data by adding extra terms. However, for the unsupervised methods, a simple recovery strategy will cause errors and outlying value accumulations, which will affect the performance of the methods. Broadly, the previous methods have not taken the effectiveness of recovered instances into consideration, or cannot flexibly balance the discrepancies between recovered data and original data. To address these problems, we propose a novel method termed Manifold-based Incomplete Multi-view clustering via Bi-consistency guidance (MIMB), which flexibly recovers incomplete data among various views, and attempts to achieve biconsistency guidance via reverse regularization. In particular, MIMB adds reconstruction terms to representation learning by recovering missing instances, which dynamically examines the latent consensus representation. Moreover, to preserve the consistency information among multiple views, MIMB implements a biconsistency guidance strategy with reverse regularization of the consensus representation and proposes a manifold embedding measure for exploring the hidden structure of the recovered data. Notably, MIMB aims to balance the importance of different views, and introduces an adaptive weight term for each view. Finally, an optimization algorithm with an alternating iteration optimization strategy is designed for final clustering. Extensive experimental results on 6 benchmark datasets are provided to confirm that MIMB can significantly obtain superior results as compared with several state-of-the-art baselines.
Huibing Wang, Mingze Yao, Yawei Chen, Yunqiu Xu, Haipeng Liu 0004, Wei Jia 0001, Xianping Fu, Yang Wang 0023
IEEE Trans. Multim.2
2024 Graph-Collaborated Auto-Encoder Hashing for Multiview Binary Clustering
abstract
Unsupervised hashing methods have attracted widespread attention with the explosive growth of large-scale data, which can greatly reduce storage and computation by learning compact binary codes. Existing unsupervised hashing methods attempt to exploit the valuable information from samples, which fails to take the local geometric structure of unlabeled samples into consideration. Moreover, hashing based on auto-encoders aims to minimize the reconstruction loss between the input data and binary codes, which ignores the potential consistency and complementarity of multiple sources data. To address the above issues, we propose a hashing algorithm based on auto-encoders for multiview binary clustering, which dynamically learns affinity graphs with low-rank constraints and adopts collaboratively learning between auto-encoders and affinity graphs to learn a unified binary code, called graph-collaborated auto-encoder (GCAE) hashing for multiview binary clustering. Specifically, we propose a multiview affinity graphs' learning model with low-rank constraint, which can mine the underlying geometric information from multiview data. Then, we design an encoder-decoder paradigm to collaborate the multiple affinity graphs, which can learn a unified binary code effectively. Notably, we impose the decorrelation and code balance constraints on binary codes to reduce the quantization errors. Finally, we use an alternating iterative optimization scheme to obtain the multiview clustering results. Extensive experimental results on five public datasets are provided to reveal the effectiveness of the algorithm and its superior performance over other state-of-the-art alternatives.
Huibing Wang, Mingze Yao, Guangqi Jiang, Zetian Mi, Xianping Fu
IEEE Trans. Neural Networks Learn. Syst.2
2023 Domain adaptive person search via GAN-based scene synthesis for cross-scene videos
Huibing Wang, Tianxiang Cui, Mingze Yao, Huijuan Pang, Yushan Du
Image Vis. Comput.3
2021 A Hierarchical Approach to Multi-Agent Path Finding
abstract
Solving Multi-Agent Path Finding (MAPF) instances optimally is NP-hard, and existing optimal and bounded suboptimal MAPF solvers thus usually do not scale to large MAPF instances. Greedy MAPF solvers scale to large MAPF instances, but their solution qualities are often bad. In this paper, we therefore propose a novel MAPF solver, Hierarchical Multi-Agent Path Planner (HMAPP), which creates a spatial hierarchy by partitioning the environment into multiple regions and decomposes a MAPF instance into smaller MAPF sub-instances for each region. For each sub-instance, it uses a bounded-suboptimal MAPF solver to solve it with good solution quality. Our experimental results show that HMAPP is able to solve as large MAPF instances as greedy MAPF solvers while achieving better solution qualities on various maps.
Han Zhang 0018, Mingze Yao, Ziang Liu 0002, Jiaoyang Li 0001, Lucas Terr, Shao-Hung Chan, T. K. Satish Kumar, Sven Koenig
SOCS2