EDBT 2026 Demo / reviewers in the wild / expert
Yaohua Zha
dblp:344/5717
· DBLP profile ↗
17ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0001-9789-452XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 13 · 6 first-author · 13 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CASL: Curvature-Augmented Self-supervised Learning for 3D Anomaly DetectionabstractDeep learning-based 3D anomaly detection methods have demonstrated significant potential in industrial manufacturing. However, many approaches are specifically designed for anomaly detection tasks, which limits their generalizability to other 3D tasks. In contrast, self-supervised point cloud models aim for general representation learning, yet our investigation reveals that these classical models are suboptimal at anomaly detection under the unified fine-tuning paradigm. This motivates us to develop a more generalizable 3D model that can effectively detect anomalies without relying on task-specific designs. Interestingly, we find that using only the curvature of each point as its anomaly score already outperforms several classical self-supervised and dedicated anomaly detection models, highlighting the critical role of curvature in 3D anomaly detection. In this paper, we propose a Curvature-Augmented Self-supervised Learning (CASL) framework based on a reconstruction paradigm. Built upon the classical U-Net architecture, our approach introduces multi-scale curvature prompts to guide the decoder in predicting the coordinates of each point. Without relying on any dedicated anomaly detection mechanisms, it achieves leading detection performance through straightforward anomaly classification fine-tuning. Moreover, the learned representations generalize well to standard 3D understanding tasks such as point cloud classification. Yaohua Zha, Xue Yuerong, Chunlin Fan, Yuansong Wang, Tao Dai 0001, Ke Chen 0004, Shutao Xia |
AAAI | 1 |
| 2025 | GCD-Sampling: A General Cross-scale Decoupled Sampling for Point CloudabstractSampling strategy (e.g., fixed farthest point sampling) of point cloud has been an essential step for developing practical solutions in 3D computer vision tasks. Previous fixed sampling is simple, but suffer from suboptimal performance for downstream tasks. To adapt to target networks properly, adaptive sampling methods with trainable parameters have been recently developed to enhance the performance. However, existing adaptive sampling methods still suffer from the over-coupling problem of target network, and thus become model-specific, which limits their practical applications. To address this issue, we propose a novel general cross-scale decoupled sampling method (GCD-sampling) for point cloud, which consists of original feature cache, cross-scale feature fusion and convex combination learning for better feature extraction. To reduce the coupling relationship with the target task network, our method only utilizes the point cloud coordinates as the input and output of itself. Besides, we introduce an arbitrary scale structure to enable parameter sharing across multi-scale sampling in point cloud networks. Extensive experiments on different architectures demonstrate the effectiveness of our method over other existing adaptive sampling methods. Tao Dai 0001, Yanzi Wang, Jianyu Xiong, Yaohua Zha, Shutao Xia, Zexuan Zhu 0001 |
AAAI | 4 |
| 2025 | Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context LearningabstractVisual In-Context Learning (VICL) enables adaptively solving vision tasks by leveraging pixel demonstrations, mimicking human-like task completion through analogy. Prompt selection is critical in VICL, but current methods assume the existence of a single "ideal" prompt in a pool of candidates, which in practice may not hold true. Multiple suitable prompts may exist, but individually they often fall short, leading to difficulties in selection and the exclusion of useful context. To address this, we propose a new perspective: prompt condensation. Rather than relying on a single prompt, candidate prompts collaborate to efficiently integrate informative contexts without sacrificing resolution. We devise Condenser, a lightweight external plugin that compresses relevant fine-grained context across multiple prompts. Optimized end-to-end with the backbone, Condenser ensures accurate integration of contextual cues. Experiments demonstrate Condenser outperforms state-of-the-arts across benchmark tasks, showing superior context compression, scalability with more prompts, and enhanced computational efficiency compared to ensemble methods, positioning it as a highly competitive solution for VICL. Code is open-sourced at https://github.com/gimpong/CVPR25-Condenser. Jinpeng Wang 0002, Tianci Luo, Yaohua Zha, Ruisheng Luo, Bin Chen 0011, Tao Dai 0001, Long Chen 0016, Yaowei Wang 0001, Shutao Xia |
CVPR | 3 |
| 2025 | MambaIRv2: Attentive State Space RestorationabstractThe Mamba-based image restoration backbones have recently demonstrated significant potential in balancing global reception and computational efficiency. However, the inherent causal modeling limitation of Mamba, where each token depends solely on its predecessors in the scanned sequence, restricts the full utilization of pixels across the image and thus presents new challenges in image restoration. In this work, we propose MambaIRv2, which equips Mamba with the non-causal modeling ability similar to ViTs to reach the attentive state space restoration model. Specifically, the proposed attentive state-space equation allows to attend beyond the scanned sequence and facilitate image unfolding with just one single scan. Moreover, we further introduce a semantic-guided neighboring mechanism to encourage interaction between distant but similar pixels. Extensive experiments show our MambaIRv2 outperforms SRFormer by even 0.35dB PSNR for lightweight SR even with 9.3% less parameters and suppresses HAT on classic SR by up to 0.29dB. Code is available at https://github.com/csguoh/MambaIR. Hang Guo 0002, Yaohua Zha, Yulun Zhang 0001, Wenbo Li 0002, Tao Dai 0001, Shutao Xia, Yawei Li 0001 |
CVPR | 3 |
| 2025 | Adapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal PerceptionabstractPoint cloud video understanding is becoming increasingly important in fields such as robotics, autonomous driving, and augmented reality, as they can accurately represent object motion and environmental changes. Despite the progress made in self-supervised learning methods for point cloud video understanding, the limited availability of 4D data and the high computational cost of training 4D-specific models remain significant obstacles. In this paper, we investigate the potential of transferring pre-trained static 3D point cloud models to the 4D domain, pointing out the limitations of static models that capture only spatial information while neglecting temporal dynamics. To address this, we propose a novel Cross-frame Spatio-temporal Adaptation (CSA) strategy by introducing the Point Tube Adapter as the embedding layer and the Geometric Constraint Temporal Adapter (GCTA) to enforce temporal consistency across frames. This strategy extracts both short-term and long-term temporal dynamics, effectively integrating them with spatial features and enriching the model’s understanding of temporal changes in point cloud videos. Extensive experiments on 3D action and gesture recognition tasks demonstrate that our method achieves state-of-the-art performance, establishing its effectiveness for point cloud video understanding. Code is available at: https://github.com/LvBaixuan/Point-CSA. Baixuan Lv, Yaohua Zha, Tao Dai 0001, Xue Yuerong, Ke Chen 0004, Shutao Xia |
CVPR | 2 |
| 2025 | PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba AdapterabstractApplying pre-trained models to assist point cloud understanding has recently become a mainstream paradigm in 3D perception. However, existing application strategies are straightforward, utilizing only the final output of the pre-trained model for various task heads. It neglects the rich complementary information in the intermediate layer, thereby failing to fully unlock the potential of pre-trained models. To overcome this limitation, we propose an orthogonal solution: Point Mamba Adapter (PMA), which constructs an ordered feature sequence from all layers of the pre-trained model and leverages Mamba to fuse all complementary semantics, thereby promoting comprehensive point cloud understanding. Constructing this ordered sequence is non-trivial due to the inherent isotropy of 3D space. Therefore, we further propose a geometry-constrained gate prompt generator (G2PG) shared across different layers, which applies shared geometric constraints to the output gates of the Mamba and dynamically optimizes the spatial order, thus enabling more effective integration of multi-layer information. Extensive experiments conducted on challenging point cloud datasets across various tasks demonstrate that our PMA elevates the capability for point cloud understanding to a new level by fusing diverse complementary intermediate features. Code is available at https://github.com/zyh16143998882/PMA. Yaohua Zha, Yanzi Wang, Hang Guo 0002, Jinpeng Wang 0002, Tao Dai 0001, Bin Chen 0011, Zhihao Ouyang, Xue Yuerong, Ke Chen 0004, Shutao Xia |
CVPR | 1 |
| 2025 | Expert-Enhanced Masked Point Modeling for Point Cloud Self-Supervised LearningabstractRecently, learning-based point cloud analysis has played a crucial role in robotic perception. Masked Point Modeling (MPM), owing to its powerful representational capabilities, has become the mainstream point cloud self-supervised learning method. However, existing MPM-based methods often suffer from the problem of negative transfer, due to the disparity in semantic distribution between upstream data and downstream data. To address this issue, we propose an expert enhancement strategy for existing MPM-based methods. Specifically, we insert a Sparse Mixture of Experts (SMoE) layer after each block of the backbone network, which utilizes a multi-branch expert architecture with routers that allocate data of different semantics to the appropriate experts for analysis. During the pre-training phase, our expert-enhanced model not only learns universal 3D representations for the backbone network but also acquires powerful semantic routing capabilities for all expert layers. In the fine-tuning phase, we freeze all backbones and conduct end-to-end fine-tuning solely on our expert layers to adaptively select multiple experts most relevant to the semantics of each downstream data for analysis. Extensive downstream experiments demonstrate the superiority of our method, especially outperforming baseline (Point-MAE) by 5.16%, 5.86%, and 4.62% in three variants of ScanObjectNN while utilizing only 12% of its trainable parameters. Our code is released at https://github.com/chenchen1104/point_e2mae. Yaohua Zha, Naiqi Li, Tao Dai 0001, Bin Chen 0011, Shutao Xia |
ICRA | 2 |
| 2025 | Point Cloud Mixture-of-Domain-Experts Model for 3D Self-supervised LearningabstractPoint clouds, as a primary representation of 3D data, can be categorized into scene domain point clouds and object domain point clouds. Point cloud self-supervised learning (SSL) has become a mainstream paradigm for learning 3D representations. However, existing point cloud SSL primarily focuses on learning domain-specific 3D representations within a single domain, neglecting the complementary nature of cross-domain knowledge, which limits the learning of 3D representations. In this paper, we propose to learn a comprehensive Point cloud Mixture-of-Domain-Experts model (Point-MoDE) via a block-to-scene pre-training strategy. Specifically, We first propose a mixture-of-domain-expert model consisting of scene domain experts and multiple shared object domain experts. Furthermore, we propose a block-to-scene pretraining strategy, which leverages the features of point blocks in the object domain to regress their initial positions in the scene domain through object-level block mask reconstruction and scene-level block position regression. By integrating the complementary knowledge between object and scene, this strategy simultaneously facilitates the learning of both object-domain and scene-domain representations, leading to a more comprehensive 3D representation. Extensive experiments in downstream tasks demonstrate the superiority of our model. Yaohua Zha, Tao Dai 0001, Hang Guo 0002, Yanzi Wang, Bin Chen 0011, Ke Chen 0004, Shutao Xia |
IJCAI | 1 |
| 2024 | Towards Compact 3D Representations via Point Feature Enhancement Masked AutoencodersabstractLearning 3D representation plays a critical role in masked autoencoder (MAE) based pre-training methods for point cloud, including single-modal and cross-modal based MAE. Specifically, although cross-modal MAE methods learn strong 3D representations via the auxiliary of other modal knowledge, they often suffer from heavy computational burdens and heavily rely on massive cross-modal data pairs that are often unavailable, which hinders their applications in practice. Instead, single-modal methods with solely point clouds as input are preferred in real applications due to their simplicity and efficiency. However, such methods easily suffer from limited 3D representations with global random mask input. To learn compact 3D representations, we propose a simple yet effective Point Feature Enhancement Masked Autoencoders (Point-FEMAE), which mainly consists of a global branch and a local branch to capture latent semantic features. Specifically, to learn more compact features, a share-parameter Transformer encoder is introduced to extract point features from the global and local unmasked patches obtained by global random and local block mask strategies, followed by a specific decoder to reconstruct. Meanwhile, to further enhance features in the local branch, we propose a Local Enhancement Module with local patch convolution to perceive fine-grained local context at larger scales. Our method significantly improves the pre-training efficiency compared to cross-modal alternatives, and extensive downstream experiments underscore the state-of-the-art effectiveness, particularly outperforming our baseline (Point-MAE) by 5.16%, 5.00%, and 5.04% in three variants of ScanObjectNN, respectively. Code is available at https://github.com/zyh16143998882/AAAI24-PointFEMAE. Yaohua Zha, Huizhen Ji, Jinmin Li, Rongsheng Li, Tao Dai 0001, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
AAAI | 1 |
| 2024 | LR-MAE: Locate while Reconstructing with Masked Autoencoders for Point Cloud Self-supervised LearningabstractAs an efficient self-supervised pre-training approach, Masked autoencoder (MAE) has shown promising improvement across various 3D point cloud understanding tasks. However, the pretext task of existing point-based MAE is to reconstruct the geometry of masked points only, hence it learns features at lower semantic levels which is not appropriate for high-level downstream tasks. To address this challenge, we propose a novel self-supervised approach named Locate while Reconstructing with Masked Autoencoders (LR-MAE). Specifically, a multi-head decoder is designed to simultaneously localize the global position of masked patches while reconstructing masked points, aimed at learning better semantic features that align with downstream tasks. Moreover, we design a random query patch detection strategy for 3D object detection tasks in the pre-training stage, which significantly boosts the model performance with faster convergence speed. Extensive experiments show that our LR-MAE achieves superior performance on various point cloud understanding tasks. By fine-tuning on downstream datasets, LR-MAE outperforms the Point-MAE baseline by 3.65% classification accuracy on the ScanObjectNN dataset, and significantly exceeds the 3DETR baseline by 6.1% AP50on the ScanNetV2 dataset. Code is available at https://github.com/cathy-ji/LR-MAE. Huizhen Ji, Yaohua Zha, Qingmin Liao |
ICME | 2 |
| 2024 | Invertible Residual Rescaling Models
Jinmin Li, Tao Dai 0001, Yaohua Zha, Yilu Luo, Longfei Lu, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
IJCAI | 3 |
| 2024 | Large Point-to-Gaussian Model for Image-to-3D Generation
Longfei Lu, Huachen Gao, Tao Dai 0001, Yaohua Zha, Zhi Hou, Junta Wu, Shutao Xia |
ACM Multimedia | 4 |
| 2024 | ReFIR: Grounding Large Restoration Models with Retrieval AugmentationabstractRecent advances in diffusion-based Large Restoration Models (LRMs) have significantly improved photo-realistic image restoration by leveraging the internal knowledge embedded within model weights. However, existing LRMs often suffer from the hallucination dilemma, i.e., producing incorrect contents or textures when dealing with severe degradations, due to their heavy reliance on limited internal knowledge. In this paper, we propose an orthogonal solution called the Retrieval-augmented Framework for Image Restoration (ReFIR), which incorporates retrieved images as external knowledge to extend the knowledge boundary of existing LRMs in generating details faithful to the original scene. Specifically, we first introduce the nearest neighbor lookup to retrieve content-relevant high-quality images as reference, after which we propose the cross-image injection to modify existing LRMs to utilize high-quality textures from retrieved images. Thanks to the additional external knowledge, our ReFIR can well handle the hallucination challenge and facilitate faithfully results. Extensive experiments demonstrate that ReFIR can achieve not only high-fidelity but also realistic restoration results. Importantly, our ReFIR requires no training and is adaptable to various LRMs. Hang Guo 0002, Tao Dai 0001, Zhihao Ouyang, Taolin Zhang 0003, Yaohua Zha, Bin Chen 0011, Shutao Xia |
NeurIPS | 5 |
| 2024 | LCM: Locally Constrained Compact Point Cloud Model for Masked Point ModelingabstractThe pre-trained point cloud model based on Masked Point Modeling (MPM) has exhibited substantial improvements across various tasks. However, these models heavily rely on the Transformer, leading to quadratic complexity and limited decoder, hindering their practice application. To address this limitation, we first conduct a comprehensive analysis of existing Transformer-based MPM, emphasizing the idea that redundancy reduction is crucial for point cloud analysis. To this end, we propose a Locally constrained Compact point cloud Model (LCM) consisting of a locally constrained compact encoder and a locally constrained Mamba-based decoder. Our encoder replaces self-attention with our local aggregation layers to achieve an elegant balance between performance and efficiency. Considering the varying information density between masked and unmasked patches in the decoder inputs of MPM, we introduce a locally constrained Mamba-based decoder. This decoder ensures linear complexity while maximizing the perception of point cloud geometry information from unmasked patches with higher information density. Extensive experimental results show that our compact model significantly surpasses existing Transformer-based models in both performance and efficiency, especially our LCM-based Point-MAE model, compared to the Transformer-based model, achieved an improvement of 1.84%, 0.67%, and 0.60% in performance on the three variants of ScanObjectNN while reducing parameters by 88% and computation by 73%. The code is available at https://github.com/zyh16143998882/LCM. Yaohua Zha, Naiqi Li, Yanzi Wang, Tao Dai 0001, Hang Guo 0002, Bin Chen 0011, Zhi Wang 0001, Zhihao Ouyang, Shutao Xia |
NeurIPS | 1 |
| 2023 | Semantic Preserving Learning for Task-Oriented Point Cloud DownsamplingabstractRecent years have witnessed a tremendous growth in the scale and resolution of point clouds. To facilitate the applications of point cloud in downsampling tasks (e.g., point cloud classification), several task-oriented downsampling works have been developed by training with the task-specific loss with one-hot encoded label. However, these methods still suffer from performance degradation at high downsampling scales. In this paper, we propose a general semantic-preserved downsampling framework (SPDF) for point clouds by exploiting the rich knowledge inherent in the task network. Specifically, we firstly refine the previous pipeline to generate richer semantic supervised information. Then, the semantic feature learning is subdivided into label-level and feature-level to guide the training of downsampling network, which can better limit the semantic loss during downsampling. Extensive experiments on the benchmark dataset show that SPDF outperforms state-of-the-art downsampling methods. Jianyu Xiong, Tao Dai 0001, Yaohua Zha, Xin Wang 0001, Shutao Xia |
ICASSP | 3 |
| 2023 | SFR: Semantic-Aware Feature Rendering of Point CloudabstractMulti-view projection methods have demonstrated their ability to reach state-of-the-art performance in point cloud downstream tasks(e.g., classification and retrieval). These methods first require rendering the point cloud into 2D multi-view images. However, conventional methods only project the geometry of the point cloud, and such projections inevitably suffer from a loss of point cloud semantic information due to dimensionality reduction. We propose a semantic-aware and task-oriented differentiable feature rendering (SFR), which reduces the information loss during projection by generating rendered images with more point cloud semantic information for downstream tasks. Our SFR method can be applied as a plug-and-play module added to any multi-view-based backbone network for end-to-end training. Extensive experiments on benchmark datasets show that our SFR method reaches state-of-the-art performance and brings general improvements to point cloud classification and retrieval tasks. Yaohua Zha, Rongsheng Li, Tao Dai 0001, Jianyu Xiong, Xin Wang 0001, Shutao Xia |
ICASSP | 1 |
| 2023 | Instance-aware Dynamic Prompt Tuning for Pre-trained Point Cloud ModelsabstractPre-trained point cloud models have found extensive applications in 3D understanding tasks like object classification and part segmentation. However, the prevailing strategy of full fine-tuning in downstream tasks leads to large per-task storage overhead for model parameters, which limits the efficiency when applying large-scale pre-trained models. Inspired by the recent success of visual prompt tuning (VPT), this paper attempts to explore prompt tuning on pre-trained point cloud models, to pursue an elegant balance between performance and parameter efficiency. We find while instance-agnostic static prompting, e.g. VPT, shows some efficacy in downstream transfer, it is vulnerable to the distribution diversity caused by various types of noises in real-world point cloud data. To conquer this limitation, we propose a novel Instance-aware Dynamic Prompt Tuning (IDPT) strategy for pre-trained point cloud models. The essence of IDPT is to develop a dynamic prompt generation module to perceive semantic prior features of each point cloud instance and generate adaptive prompt tokens to enhance the model's robustness. Notably, extensive experiments demonstrate that IDPT outperforms full finetuning in most tasks with a mere 7% of the trainable parameters, providing a promising solution to parameter-efficient learning for pre-trained point cloud models. Code is available at https://github.com/zyh16143998882/ICCV23-IDPT. Yaohua Zha, Jinpeng Wang 0002, Tao Dai 0001, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
ICCV | 1 |