Wenbiao Yan

dblp:341/1579 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0001-9884-9923ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
abstract
Dongxu Zhang, Yiding Sun, Cheng Tan, Wenbiao Yan, Ning Yang, Jihua Zhu, Haijun Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Wenbiao Yan, Ning Yang 0005, Jihua Zhu, Haijun Zhang 0001
ACL (1)4
2026 IGASA: Integrated Geometry-Aware and Skip-Attention Modules for Enhanced Point Cloud Registration
Jihua Zhu, Wenbiao Yan, Peilin Fan, Huimin Lu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Point-RMAE: Reinforcement Masked Autoencoder for 3D Representation Learning
abstract
The Mainstream 3D masked point modeling representation learning community typically employs predefined, fixed-ratio random or block masking strategies, aiming to obtain optimal representations and achieve high downstream performance. However, these empirical designs overlook the significant geometric information and structural importance differences that are inherent among different 3D points, leading to a suboptimal trade-off between the representation capture capabilities and reconstruction difficulty of such masking strategies. To address this issue, we are the first to present this decision-making problem to a reinforcement learning agent and propose a Reinforcement Masked Autoencoder for 3D representation learning, named Point-RMAE. Guided by geometric features as state factor, this method leverages the Masking Strategy Analyzer and the Dynamic Masking Generator to adaptively decide and apply the masking strategy during pretraining. The Masking Ratio Scheduling module dynamically adjusts the masking ratio based on the optimal strategy. Subsequently, the analyzer is updated by multiscale rewards derived from reconstruction quality level, distribution-aware feedback, and policy exploration. Notably, to enrich the Reward Function with distribution-aware signals and avoid decision collapse issue, we propose a Flow Matching Point Cloud Fast Generator that guides the selected masking decisions. Our method achieves outstanding performance across downstream tasks such as shape classification, medical diagnosis, object detection, action recognition, denoising and multiscale scene segmentation on ten popular 3D and 4D datasets. More importantly, Point-RMAE pioneers the application of reinforcement learning in 3D self-supervised representation learning.
Haozhe Cheng, Lintong Wei, Wenbiao Yan, Jinqian Chen, Kun Yue, Jihua Zhu
IEEE Trans. Image Process.4
2025 Relationship completion for incomplete multi-view clustering
Minghong Wu, Jihua Zhu, Wenbiao Yan, Qinghai Zheng
Neural Networks3
2025 Partially multi-view clustering via re-alignment
Wenbiao Yan, Jihua Zhu, Jinqian Chen, Haozhe Cheng, Shunshun Bai, Liang Duan, Qinghai Zheng
Neural Networks1
2025 Graph Variational Multi-View Clustering
abstract
Multi-view clustering (MVC) aims to extract consensus information from multi-source data and has developed rapidly. While generative model-based methods perform well by leveraging predefined priors, they often overlook inter-instance relationships, which are essential for high-quality clustering. To address this issue, we propose Graph Variational Multi-view Clustering (GVMVC), which integrates graph information into the generative process. Specifically, we treat both the original multi-view features and graph information from each view as observed data, guiding the learning of latent representations. The key principles of this approach are: 1) enhancing discriminative feature learning through graph integration, and 2) ensuring consistent multi-view learning via graph-based constraints. Extensive experiments show that GVMVC outperforms state-of-the-art methods across various datasets and metrics. Code is available at https://github.com/WenB777/GVMVC.git.
Wenbiao Yan, Jihua Zhu, Jinqian Chen, Haozhe Cheng, Qinghai Zheng
IEEE Trans. Circuits Syst. Video Technol.1
2025 Neighbor-Based Completion for Addressing Incomplete Multiview Clustering
abstract
Driven by the complementarity and consistency inherent in multiview data, multiview clustering (MVC) has garnered widespread attention in various domains. Real-world data often encounters the issue of missing information, leading to a surge of interest in the domain of incomplete MVC (IMVC). Despite existing approaches having made significant progress in addressing IMVC, two significant challenges persist: 1) many alignment-based methodologies tend to overlook the topological relationships among instances and 2) the view representations based on completion lack reconstructive properties, casting doubt on their alignment with the actual view representations. In response, we present a novel approach termed neighbor-based completion for addressing IMVC (NBIMVC), which capitalizes on the topological information among instances and the consistent information across views. Specifically, our method uses autoencoders to learn feature representations for each view and leverages nearest-neighbor relationships between unique and complete instances to complete missing features in missing views. Subsequently, we enforce hard negative alignment constraints on complete paired instances in the feature space. Finally, we ensure the consistency of views in the semantic space by employing cluster information and a shared clustering network, which facilitates the final multiview categories output and effectively resolves the IMVC problem. Extensive experimental evaluations validate the efficacy of our proposed method, showcasing comparable or superior performance to existing approaches.
Wenbiao Yan, Jihua Zhu, Yiyang Zhou, Jinqian Chen, Haozhe Cheng, Kun Yue, Qinghai Zheng
IEEE Trans. Neural Networks Learn. Syst.1
2024 HumanRefiner: Benchmarking Abnormal Human Generation and Refining with Coarse-to-Fine Pose-Reversible Guidance
Guian Fang, Wenbiao Yan, Yuanfan Guo, Jianhua Han, Zutao Jiang, Hang Xu 0004, Shengcai Liao, Xiaodan Liang
ECCV (32)2
2024 MCoCo: Multi-level Consistency Collaborative multi-view clustering
Yiyang Zhou, Qinghai Zheng, Wenbiao Yan, Jihua Zhu
Expert Syst. Appl.4
2024 Multi-view Semantic Consistency based Information Bottleneck for Clustering
Wenbiao Yan, Yiyang Zhou, Qinghai Zheng, Jihua Zhu
Knowl. Based Syst.1
2024 PTM: Torus Masking for 3D Representation Learning Guided by Robust and Trusted Teachers
abstract
3D Masked Point Modeling (MPM) typically involves randomly or blockly discarding points or patches and then reconstructing them, offering a promising avenue for exploring geometric representation. By surveying current masking strategies, we have found that random-masked regions are provided with excessive context, reducing modeling difficulty but impeding knowledge transfer. While, block-masked regions lack sufficient guidance, resulting in significant generated noise. To address these issues, we propose PTM, a novel Transformer-style 3D MPM method employing a torus masking strategy. Specifically, a high-density area is chosen as the masked region, forming a torus by retaining small-radius neighborhoods around the center point. To mitigate torus modeling noise, the designed robust teacher model captures density scale to construct noise embedding, utilizing a reverse fit function for reconstruction assistance. Furthermore, the proposed trusted teacher model defines the multi-modal global descriptor as subjective evidence. On a semantic level, we form semi-subjective trusted evidence to guide reconstruction by evaluating the contribution of each subjective evidence to 3D representation. Downstream fine-tuning tasks validate the state-of-the-art performance of PTM in multi-scale point cloud classification and segmentation.
Haozhe Cheng, Jihua Zhu, Naiwen Hu, Jinqian Chen, Wenbiao Yan
IEEE Trans. Circuits Syst. Video Technol.5
2023 Contrastive Label Enhancement
abstract
Label distribution learning (LDL) is a new machine learning paradigm for solving label ambiguity. Since it is difficult to directly obtain label distributions, many studies are focusing on how to recover label distributions from logical labels, dubbed label enhancement (LE). Existing LE methods estimate label distributions by simply building a mapping relationship between features and label distributions under the supervision of logical labels. They typically overlook the fact that both features and logical labels are descriptions of the instance from different views. Therefore, we propose a novel method called Contrastive Label Enhancement (ConLE) which integrates features and logical labels into the unified projection space to generate high-level features by contrastive learning strategy. In this approach, features and logical labels belonging to the same sample are pulled closer, while those of different samples are projected farther away from each other in the projection space. Subsequently, we leverage the obtained high-level features to gain label distributions through a well-designed training strategy that considers the consistency of label attributes. Extensive experiments on LDL benchmark datasets demonstrate the effectiveness and superiority of our method.
Yiyang Zhou, Jihua Zhu, Xinyuan Liu 0001, Wenbiao Yan
IJCAI5