EDBT 2026 Demo / reviewers in the wild / expert
Haozhe Cheng
dblp:295/6942
· DBLP profile ↗
17ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-8723-9924ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Point-SRA: Self-Representation Alignment for 3D Representation LearningabstractMasked autoencoders (MAE) have become a dominant paradigm in 3D representation learning, setting new performance benchmarks across various downstream tasks. Existing methods with fixed mask ratios neglect multi-level representational correlations and intrinsic geometric structures, while relying on point-wise reconstruction assumptions that conflict with the diversity of point cloud. To address these issues, we propose a 3D representation learning method, termed Point-SRA, which aligns representations through self-distillation and probabilistic modeling. Specifically, we assign different masking ratios to the MAE to capture complementary geometric and semantic information, while the MeanFlow Transformer (MFT) leverages cross-modal conditional embeddings to enable diverse probabilistic reconstruction. Our analysis further reveals that representations at different time steps in MFT also exhibit complementarity. Therefore, a Dual Self-Representation Alignment mechanism is proposed at both the MAE and MFT levels. Finally, we design a Flow-Conditioned Fine-Tuning Architecture to fully exploit the point cloud distribution learned via MeanFlow. Point-SRA outperforms Point-MAE by 5.37% on ScanObjectNN. On intracranial aneurysm segmentation, it reaches 96.07% mean IoU for arteries and 86.87% for aneurysms. For 3D object detection, Point-SRA achieves 47.3% AP@50, surpassing MaskPoint by 5.12%. Lintong Wei, Haozhe Cheng, Jihua Zhu, Kaibing Zhang |
AAAI | 3 |
| 2026 | HyperPoint: Multimodal 3D foundation model in hyperbolic space
Haozhe Cheng, Chaoyi Lu, Zhengqiao Li, Minghong Wu, Huimin Lu 0001, Jihua Zhu |
Pattern Recognit. | 2 |
| 2026 | Curve3D: Curvature-aware masked autoencoder for self-supervised point cloud understanding
Chaoyi Lu, Haozhe Cheng, Huimin Lu 0001, Jihua Zhu |
Pattern Recognit. | 3 |
| 2026 | Point-RMAE: Reinforcement Masked Autoencoder for 3D Representation LearningabstractThe Mainstream 3D masked point modeling representation learning community typically employs predefined, fixed-ratio random or block masking strategies, aiming to obtain optimal representations and achieve high downstream performance. However, these empirical designs overlook the significant geometric information and structural importance differences that are inherent among different 3D points, leading to a suboptimal trade-off between the representation capture capabilities and reconstruction difficulty of such masking strategies. To address this issue, we are the first to present this decision-making problem to a reinforcement learning agent and propose a Reinforcement Masked Autoencoder for 3D representation learning, named Point-RMAE. Guided by geometric features as state factor, this method leverages the Masking Strategy Analyzer and the Dynamic Masking Generator to adaptively decide and apply the masking strategy during pretraining. The Masking Ratio Scheduling module dynamically adjusts the masking ratio based on the optimal strategy. Subsequently, the analyzer is updated by multiscale rewards derived from reconstruction quality level, distribution-aware feedback, and policy exploration. Notably, to enrich the Reward Function with distribution-aware signals and avoid decision collapse issue, we propose a Flow Matching Point Cloud Fast Generator that guides the selected masking decisions. Our method achieves outstanding performance across downstream tasks such as shape classification, medical diagnosis, object detection, action recognition, denoising and multiscale scene segmentation on ten popular 3D and 4D datasets. More importantly, Point-RMAE pioneers the application of reinforcement learning in 3D self-supervised representation learning. Haozhe Cheng, Lintong Wei, Wenbiao Yan, Jinqian Chen, Kun Yue, Jihua Zhu |
IEEE Trans. Image Process. | 1 |
| 2025 | PointDico: Contrastive 3D Representation Learning Guided by Diffusion ModelsabstractSelf-supervised representation learning has shown significant improvement in Natural Language Processing and 2D Computer Vision. However, existing methods face difficulties in representing 3D data because of its unordered and uneven density. Through an in-depth analysis of mainstream contrastive and generative approaches, we find that contrastive models tend to suffer from overfitting, while 3D Mask Autoencoders struggle to handle unordered point clouds. This motivates us to learn 3D representations by sharing the merits of diffusion and contrast models, which is non-trivial due to the pattern difference between the two paradigms. In this paper, we propose PointDico, a novel model that seamlessly integrates these methods. PointDico learns from both denoising generative modeling and cross-modal contrastive learning through knowledge distillation, where the diffusion model serves as a guide for the contrastive model. We introduce a hierarchical pyramid conditional generator for multi-scale geometric feature extraction and employ a dual-channel design to effectively integrate local and global contextual information. PointDico achieves a new state-of-the-art in 3D representation learning, e.g., 94.32% accuracy on ScanObjectNN, 86.5% Inst. mIoU on ShapeNetPart. Haozhe Cheng |
IJCNN | 3 |
| 2025 | BeyondPoints: Curve Fusion and Attention-Driven Local Feature Learning for 3-D Semantic SegmentationabstractEfficient semantic segmentation of large-scale point cloud scenes is regarded as a fundamental and essential task for perceiving and understanding 3-D environments. It is also recognized as a key technology for environmental perception and intelligent decision-making in Internet of Things (IoT) applications. However, the diversity of objects and occlusion issues within scenes often hinder the ability of existing networks to effectively represent varying object shapes, leading to point information ambiguity and loss caused by pose variations. To address these challenges, a novel curve fusion and attention-driven local feature learning network (BeyondPoints) is proposed for point cloud segmentation. The proposed network consists of three key modules: 1) a local feature enhancement (LFE) module; 2) a dual-axis attention (DAA) module; and 3) a hybrid curve fusion (HCF) module. Specifically, the LFE module explicitly models spatial relationships and decouples local aggregation, effectively integrating additional geometric information into local features to compensate for point information loss and enhance environmental perception. To further alleviate local spatial perception ambiguity, the DAA module is designed to extract critical information from different spatial positions, thereby enhancing the representation of significant regions in point clouds and achieving more precise semantic segmentation. Finally, the HCF module serializes point cloud data to reduce model computational complexity while enabling cross-domain feature fusion, effectively integrating contextual information and suppressing noise interference. Extensive experiments on benchmark datasets, such as S3DIS, ScanNetV2, and SemanticKITTI, validate the exceptional segmentation performance of the proposed BeyondPoints network. Particularly, its superior performance in large-scale point cloud scenes underscores its potential to enhance environmental information perception in IoT scenarios. Liguo Luo, Kaibing Zhang, Haozhe Cheng, Xiaogai Chen |
IEEE Internet Things J. | 4 |
| 2025 | Partially multi-view clustering via re-alignment
Wenbiao Yan, Jihua Zhu, Jinqian Chen, Haozhe Cheng, Shunshun Bai, Liang Duan, Qinghai Zheng |
Neural Networks | 4 |
| 2025 | Graph Variational Multi-View ClusteringabstractMulti-view clustering (MVC) aims to extract consensus information from multi-source data and has developed rapidly. While generative model-based methods perform well by leveraging predefined priors, they often overlook inter-instance relationships, which are essential for high-quality clustering. To address this issue, we propose Graph Variational Multi-view Clustering (GVMVC), which integrates graph information into the generative process. Specifically, we treat both the original multi-view features and graph information from each view as observed data, guiding the learning of latent representations. The key principles of this approach are: 1) enhancing discriminative feature learning through graph integration, and 2) ensuring consistent multi-view learning via graph-based constraints. Extensive experiments show that GVMVC outperforms state-of-the-art methods across various datasets and metrics. Code is available at https://github.com/WenB777/GVMVC.git. Wenbiao Yan, Jihua Zhu, Jinqian Chen, Haozhe Cheng, Qinghai Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Neighbor-Based Completion for Addressing Incomplete Multiview ClusteringabstractDriven by the complementarity and consistency inherent in multiview data, multiview clustering (MVC) has garnered widespread attention in various domains. Real-world data often encounters the issue of missing information, leading to a surge of interest in the domain of incomplete MVC (IMVC). Despite existing approaches having made significant progress in addressing IMVC, two significant challenges persist: 1) many alignment-based methodologies tend to overlook the topological relationships among instances and 2) the view representations based on completion lack reconstructive properties, casting doubt on their alignment with the actual view representations. In response, we present a novel approach termed neighbor-based completion for addressing IMVC (NBIMVC), which capitalizes on the topological information among instances and the consistent information across views. Specifically, our method uses autoencoders to learn feature representations for each view and leverages nearest-neighbor relationships between unique and complete instances to complete missing features in missing views. Subsequently, we enforce hard negative alignment constraints on complete paired instances in the feature space. Finally, we ensure the consistency of views in the semantic space by employing cluster information and a shared clustering network, which facilitates the final multiview categories output and effectively resolves the IMVC problem. Extensive experimental evaluations validate the efficacy of our proposed method, showcasing comparable or superior performance to existing approaches. Wenbiao Yan, Jihua Zhu, Yiyang Zhou, Jinqian Chen, Haozhe Cheng, Kun Yue, Qinghai Zheng |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Turbo: Informativity-Driven Acceleration Plug-In for Vision-Language Large Models
Chen Ju, Haicheng Wang, Haozhe Cheng, Xu Chen 0026, Zhonghua Zhai, Jinsong Lan, Shuai Xiao 0002, Bo Zheng 0007 |
ECCV (46) | 3 |
| 2024 | Hyperbolic Image-and-Pointcloud Contrastive Learning for 3D Classificationabstract3D contrastive representation learning has exhibited remarkable efficacy across various downstream tasks. However, existing contrastive learning paradigms based on cosine similarity fail to deeply explore the potential intra-modal hierarchical and cross-modal semantic correlations about multi-modal data in Euclidean space. In response, we seek solutions in hyperbolic space and propose a hyperbolic image-and-pointcloud contrastive learning method (HyperIPC). For the intra-modal branch, we rely on the intrinsic geometric structure to explore the hyperbolic embedding representation of point cloud to capture invariant features. For the cross-modal branch, we leverage images to guide the point cloud in establishing strong semantic hierarchical correlations. Empirical experiments underscore the outstanding classification performance of HyperIPC. Notably, HyperIPC enhances object classification results by 2.8% and few-shot classification outcomes by 5.9% on ScanObjectNN compared to the baseline. Furthermore, ablation studies and confirmatory testing validate the rationality of HyperIPC’s parameter settings and the effectiveness of its submodules. Naiwen Hu, Haozhe Cheng, Jihua Zhu |
IROS | 2 |
| 2024 | EDGCNet: Joint dynamic hyperbolic graph convolution and dual squeeze-and-attention for 3D point cloud segmentation
Haozhe Cheng, Jihua Zhu |
Expert Syst. Appl. | 1 |
| 2024 | Multi-Trusted Cross-Modal Information Bottleneck for 3D self-supervised representation learning
Haozhe Cheng, Jihua Zhu, Zhongyu Li 0002 |
Knowl. Based Syst. | 1 |
| 2024 | Trusted 3D self-supervised representation learning with cross-modal settings
Haozhe Cheng, Jihua Zhu |
Mach. Vis. Appl. | 2 |
| 2024 | PTM: Torus Masking for 3D Representation Learning Guided by Robust and Trusted Teachersabstract3D Masked Point Modeling (MPM) typically involves randomly or blockly discarding points or patches and then reconstructing them, offering a promising avenue for exploring geometric representation. By surveying current masking strategies, we have found that random-masked regions are provided with excessive context, reducing modeling difficulty but impeding knowledge transfer. While, block-masked regions lack sufficient guidance, resulting in significant generated noise. To address these issues, we propose PTM, a novel Transformer-style 3D MPM method employing a torus masking strategy. Specifically, a high-density area is chosen as the masked region, forming a torus by retaining small-radius neighborhoods around the center point. To mitigate torus modeling noise, the designed robust teacher model captures density scale to construct noise embedding, utilizing a reverse fit function for reconstruction assistance. Furthermore, the proposed trusted teacher model defines the multi-modal global descriptor as subjective evidence. On a semantic level, we form semi-subjective trusted evidence to guide reconstruction by evaluating the contribution of each subjective evidence to 3D representation. Downstream fine-tuning tasks validate the state-of-the-art performance of PTM in multi-scale point cloud classification and segmentation. Haozhe Cheng, Jihua Zhu, Naiwen Hu, Jinqian Chen, Wenbiao Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Transition Information Enhanced Disentangled Graph Neural Networks for session-based recommendation
Ansong Li, Jihua Zhu, Zhongyu Li 0002, Haozhe Cheng |
Expert Syst. Appl. | 4 |
| 2021 | PTANet: Triple Attention Network for point cloud semantic segmentation
Haozhe Cheng, Mao-Xin Luo, Kaibing Zhang |
Eng. Appl. Artif. Intell. | 1 |