EDBT 2026 Demo / reviewers in the wild / expert
Zhirui Gao
dblp:342/7837
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-7108-7962ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Curve-Aware Gaussian Splatting for 3D Parametric Curve ReconstructionabstractThis paper presents an end-to-end framework for reconstructing 3D parametric curves directly from multi-view edge maps. Contrasting with existing two-stage methods that follow a sequential ``edge point cloud reconstruction and parametric curve fitting'' pipeline, our one-stage approach optimizes 3D parametric curves directly from 2D edge maps, eliminating error accumulation caused by the inherent optimization gap between disconnected stages. However, parametric curves inherently lack suitability for rendering-based multi-view optimization, necessitating a complementary representation that preserves their geometric properties while enabling differentiable rendering. We propose a novel bi-directional coupling mechanism between parametric curves and edge-oriented Gaussian components. This tight correspondence formulates a curve-aware Gaussian representation, \textbf{CurveGaussian}, that enables differentiable rendering of 3D curves, allowing direct optimization guided by multi-view evidence. Furthermore, we introduce a dynamically adaptive topology optimization framework during training to refine curve structures through linearization, merging, splitting, and pruning operations. Comprehensive evaluations on the ABC dataset and real-world benchmarks demonstrate our one-stage method's superiority over two-stage alternatives, particularly in producing cleaner and more robust reconstructions. Additionally, by directly optimizing parametric curves, our method significantly reduces the parameter count during training, achieving both higher efficiency and superior performance compared to existing approaches. Zhirui Gao, Renjiao Yi, Yaqiao Dai, Xuening Zhu, Wei Chen 0009, Chenyang Zhu 0002, Kai Xu 0004 |
ICCV | 1 |
| 2025 | Self-Supervised Learning of Hybrid Part-Aware 3D Representations of 2D Gaussians and Superquadrics
Zhirui Gao, Renjiao Yi, Yuhang Huang 0006, Wei Chen 0009, Chenyang Zhu 0002, Kai Xu 0004 |
ICCV | 1 |
| 2025 | BoxFusion: Reconstruction-Free Open-Vocabulary 3D Object Detection via Real-Time Multi-View Box FusionabstractAbstract Open‐vocabulary 3D object detection has gained significant interest due to its critical applications in autonomous driving and embodied AI. Existing detection methods, whether offline or online, typically rely on dense point cloud reconstruction, which imposes substantial computational overhead and memory constraints, hindering real‐time deployment in downstream tasks. To address this, we propose a novel reconstruction‐free online framework tailored for memory‐efficient and real‐time 3D detection. Specifically, given streaming posed RGB‐D video input, we leverage Cubify Anything as a pre‐trained visual foundation model (VFM) for single‐view 3D object detection, coupled with CLIP to capture open‐vocabulary semantics of detected objects. To fuse all detected bounding boxes across different views into a unified one, we employ an association module for correspondences of multi‐views and an optimization module to fuse the 3D bounding boxes of the same instance. The association module utilizes 3D Non‐Maximum Suppression (NMS) and a box correspondence matching module. The optimization module uses an IoU‐guided efficient random optimization technique based on particle filtering to enforce multi‐view consistency of the 3D bounding boxes while minimizing computational complexity. Extensive experiments on CA‐1M and ScanNetV2 datasets demonstrate that our method achieves state‐of‐the‐art performance among online methods. Benefiting from this novel reconstruction‐free paradigm for 3D object detection, our method exhibits great generalization abilities in various scenarios, enabling real‐time perception even in environments exceeding 1000 square meters. Yuqing Lan, Chenyang Zhu 0002, Zhirui Gao, Jiazhao Zhang, Renjiao Yi, Yijie Wang 0001, Kai Xu 0004 |
Comput. Graph. Forum | 3 |
| 2025 | Generic Objects as Pose Probes for Few-Shot View SynthesisabstractRadiance fields, including NeRFs and 3D Gaussians, demonstrate great potential in high-fidelity rendering and scene reconstruction, while they require a substantial number of posed images as input. COLMAP is frequently employed for preprocessing to estimate poses. However, COLMAP necessitates a large number of feature matches to operate effectively, and struggles with scenes characterized by sparse features, large baselines, or few-view images. We aim to tackle few-view NeRF reconstruction using only 3 to 6 unposed scene images, freeing from COLMAP initializations. Inspired by the idea of calibration boards in traditional pose calibration, we propose a novel approach of utilizing everyday objects, commonly found in both images and real life, as “pose probes”. By initializing the probe object as a cube shape, we apply a dual-branch volume rendering optimization (object NeRF and scene NeRF) to constrain the pose optimization and jointly refine the geometry. PnP matching is used to initialize poses between images incrementally, where only a few feature matches are enough. PoseProbe achieves state-of-the-art performance in pose estimation and novel view synthesis across multiple datasets in experiments. We demonstrate its effectiveness, particularly in few-view and large-baseline scenes where COLMAP struggles. In ablations, using different objects in a scene yields comparable performance, showing that PoseProbe is robust to the choice of probe objects. Our project page is available at:https://zhirui-gao.github.io/PoseProbe.github.io/ Zhirui Gao, Renjiao Yi, Chenyang Zhu 0002, Ke Zhuang, Wei Chen 0009, Kai Xu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | FDC-NeRF: Learning Pose-Free Neural Radiance Fields with Flow-Depth ConsistencyabstractLearning neural radiance fields (NeRF) without camera poses has been widely studied. However, recent methods lack explicit and effective supervision for pose estimation, resulting in ambiguous optimization of camera pose and NeRF geometry during joint training, particularly in scenarios involving large camera movements. In this paper, we propose FDCNeRF that leverages the direction information contained in the RGB-based optical flow and depth-based virtual flow as a direct guidance for camera pose optimization to reduce pose-geometry ambiguity. Additionally, we introduce Adaptive Pose-Aware Sampling (APAS) to replace the previous random ray sampling strategy, which reduces the difficulty of pose learning in early stages and preserves the diversity of rays in later stages. Experiments on the challenging Tanks and Temples dataset demonstrate that our method achieves state-of-the-art results in both novel view synthesis quality and pose estimation accuracy. Huachen Gao, Shihe Shen, Kaiqiang Xiong, Zhirui Gao, Yugui Xie, Ronggang Wang |
ICASSP | 6 |
| 2024 | Learning accurate template matching with differentiable coarse-to-fine correspondence refinementabstractTemplate matching is a fundamental task in computer vision and has been studied for decades. It plays an essential role in manufacturing industry for estimating the poses of different parts, facilitating downstream tasks such as robotic grasping. Existing methods fail when the template and source images have different modalities, cluttered backgrounds, or weak textures. They also rarely consider geometric transformations via homographies, which commonly exist even for planar industrial parts. To tackle the challenges, we propose an accurate template matching method based on differentiable coarse-to-fine correspondence refinement. We use an edge-aware module to overcome the domain gap between the mask template and the grayscale image, allowing robust matching. An initial warp is estimated using coarse correspondences based on novel structure-aware information provided by transformers. This initial alignment is passed to a refinement network using references and aligned images to obtain sub-pixel level correspondences which are used to give the final geometric transformation. Extensive evaluation shows that our method to be significantly better than state-of-the-art methods and baselines, providing good generalization ability and visually plausible results even on unseen real data. Zhirui Gao, Renjiao Yi, Zheng Qin 0002, Yunfan Ye, Chenyang Zhu 0002, Kai Xu 0004 |
Comput. Vis. Media | 1 |
| 2023 | NEF: Neural Edge Fields for 3D Parametric Curve Reconstruction from Multi-View ImagesabstractWe study the problem of reconstructing 3D feature curves of an object from a set of calibrated multi-view images. To do so, we learn a neural implicit field representing the density distribution of 3D edges which we refer to as Neural Edge Field (NEF). Inspired by NeRF [20], NEF is optimized with a view-based rendering loss where a 2D edge map is rendered at a given view and is compared to the ground-truth edge map extracted from the image of that view. The rendering-based differentiable optimization of NEF fully exploits 2D edge detection, without needing a supervision of 3D edges, a 3D geometric operator or cross-view edge correspondence. Several technical designs are devised to ensure learning a range-limited and view-independent NEF for robust edge extraction. The final parametric 3D curves are extracted from NEF with an iterative optimization method. On our benchmark with synthetic data, we demonstrate that NEF outperforms existing state-of-the-art methods on all metrics. Project page: https://yunfan1202.github.io/NEF/. Yunfan Ye, Renjiao Yi, Zhirui Gao, Chenyang Zhu 0002, Zhiping Cai, Kai Xu 0004 |
CVPR | 3 |
| 2023 | 2D3D-MATR: 2D-3D Matching Transformer for Detection-free Registration between Images and Point CloudsabstractThe commonly adopted detect-then-match approach to registration finds difficulties in the cross-modality cases due to the incompatible keypoint detection and inconsistent feature description. We propose, 2D3D-MATR, a detection-free method for accurate and robust registration between images and point clouds. Our method adopts a coarse-to-fine pipeline where it first computes coarse correspondences between downsampled patches of the input image and the point cloud and then extends them to form dense correspondences between pixels and points within the patch region. The coarse-level patch matching is based on transformer which jointly learns global contextual constraints with self-attention and cross-modality correlations with cross-attention. To resolve the scale ambiguity in patch matching, we construct a multi-scale pyramid for each image patch and learn to find for each point patch the best matching image patch at a proper resolution level. Extensive experiments on two public benchmarks demonstrate that 2D3D-MATR outperforms the previous state-of-the-art P2-Net by around 20 percentage points on inlier ratio and over 10 points on registration recall. Our code and models are available at https://github.com/minhaolee/2D3DMATR. Minhao Li, Zheng Qin 0002, Zhirui Gao, Renjiao Yi, Chenyang Zhu 0002, Yulan Guo, Kai Xu 0004 |
ICCV | 3 |
| 2023 | Delving Into Crispness: Guided Label Refinement for Crisp Edge DetectionabstractLearning-based edge detection usually suffers from predicting thick edges. Through extensive quantitative study with a new edge crispness measure, we find that noisy human-labeled edges are the main cause of thick predictions. Based on this observation, we advocate that more attention should be paid on label quality than on model design to achieve crisp edge detection. To this end, we propose an effective Canny-guided refinement of human-labeled edges whose result can be used to train crisp edge detectors. Essentially, it seeks for a subset of over-detected Canny edges that best align human labels. We show that several existing edge detectors can be turned into a crisp edge detector through training on our refined edge maps. Experiments demonstrate that deep models trained with refined edges achieve significant performance boost of crispness from 17.4% to 30.6%. With the PiDiNet backbone, our method improves ODS and OIS by 12.2% and 12.6% on the Multicue dataset, respectively, without relying on non-maximal suppression. We further conduct experiments and show the superiority of our crisp edge detection for optical flow estimation and image segmentation. Yunfan Ye, Renjiao Yi, Zhirui Gao, Zhiping Cai, Kai Xu 0004 |
IEEE Trans. Image Process. | 3 |