VLDB 2026 Research / reviewers in the wild / expert
Renjiao Yi
dblp:147/1338
· DBLP profile ↗
31ranked-venue papers
4as first author
26since 2021 · last 2026
0000-0002-6057-1089ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 22 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From dark flash images to relightable 3D scenes with photometric stereo priors
Xuening Zhu, Renjiao Yi, Xin Wen 0005, Xuesong Xu, Hailiang Hou, Kai Xu 0004, Chenyang Zhu 0002 |
Frontiers Comput. Sci. | 2 |
| 2026 | A generalizable neural representation framework for arbitrary-scale volumetric CT super-resolution
Xin Wen 0005, Renjiao Yi, Xuening Zhu, Chenyang Zhu 0002, Kai Xu 0004, Kunlun He |
Neurocomputing | 2 |
| 2026 | RemixFusion: Residual-based Mixed Representation for Large-scale Online RGB-D ReconstructionabstractThe introduction of the neural implicit representation has notably propelled the advancement of online dense reconstruction techniques. Compared to traditional explicit representations, such as TSDF, it substantially improves the mapping completeness and memory efficiency. However, the lack of reconstruction details and the time-consuming learning of neural representations hinder the widespread application of neural-based methods to large-scale online reconstruction. We introduce RemixFusion, a novel residual-based mixed representation for scene reconstruction and camera pose estimation dedicated to high-quality and large-scale online RGB-D reconstruction. In particular, we propose a residual-based map representation comprised of an explicit coarse TSDF grid and an implicit neural module that produces residuals representing fine-grained details to be added to the coarse grid. Such mixed representation allows for detail-rich reconstruction with bounded time and memory budget, contrasting with the overly-smoothed results by the purely implicit representations, thus paving the way for high-quality camera tracking. Furthermore, we extend the residual-based representation to handle multi-frame joint pose optimization via bundle adjustment (BA). In contrast to the existing methods, which optimize poses directly, we opt to optimize pose changes. Combined with a novel technique for adaptive gradient amplification, our method attains better optimization convergence and global optimality. Furthermore, we adopt a local moving volume to factorize the whole mixed scene representation with a divide-and-conquer design to facilitate efficient online learning in our residual-based framework. Extensive experiments demonstrate that our method surpasses all state-of-the-art ones, including those based either on explicit or implicit representations, in terms of the accuracy of both mapping and tracking on large-scale scenes. Project page can be found at https://lanlan96.github.io/RemixFusion/ . Yuqing Lan, Chenyang Zhu 0002, Shuaifeng Zhi, Jiazhao Zhang, Zhoufeng Wang, Renjiao Yi, Yijie Wang 0001, Kai Xu 0004 |
ACM Trans. Graph. | 6 |
| 2025 | VasTSD: Learning 3D Vascular Tree-state Space Diffusion Model for Angiography SynthesisabstractAngiography imaging is a medical imaging technique that enhances the visibility of blood vessels within the body by using contrast agents. Angiographic images can effectively assist in the diagnosis of vascular diseases. However, contrast agents may bring extra radiation exposure which is harmful to patients with health risks. To mitigate these concerns, in this paper, we aim to automatically generate angiography from non-angiographic inputs, by leveraging and enhancing the inherent physical properties of vascular structures. Previous methods relying on 2D slice-based angiography synthesis struggle with maintaining continuity in 3D vascular structures and exhibit limited effectiveness across different imaging modalities. We propose VasTSD, a 3D vascular tree-state space diffusion model to synthesize angiography from 3D non-angiographic volumes, with a novel state space serialization approach that dynamically constructs vascular tree topologies, integrating these with a diffusion-based generative model to ensure the generation of anatomically continuous vasculature in 3D volumes. A pre-trained vision embedder is employed to construct vascular state space representations, enabling consistent modeling of vascular structures across multiple modalities. Extensive experiments on various angiographic datasets demonstrate the superiority of VasTSD over prior works, achieving enhanced continuity of blood vessels in synthesized angiographic synthesis for multiple modalities and anatomical regions. Project page: https://jefferyzhifeng.github.io/projects/VasTSD/ Renjiao Yi, Xin Wen 0005, Chenyang Zhu 0002, Kai Xu 0004 |
CVPR | 2 |
| 2025 | Curve-Aware Gaussian Splatting for 3D Parametric Curve ReconstructionabstractThis paper presents an end-to-end framework for reconstructing 3D parametric curves directly from multi-view edge maps. Contrasting with existing two-stage methods that follow a sequential ``edge point cloud reconstruction and parametric curve fitting'' pipeline, our one-stage approach optimizes 3D parametric curves directly from 2D edge maps, eliminating error accumulation caused by the inherent optimization gap between disconnected stages. However, parametric curves inherently lack suitability for rendering-based multi-view optimization, necessitating a complementary representation that preserves their geometric properties while enabling differentiable rendering. We propose a novel bi-directional coupling mechanism between parametric curves and edge-oriented Gaussian components. This tight correspondence formulates a curve-aware Gaussian representation, \textbf{CurveGaussian}, that enables differentiable rendering of 3D curves, allowing direct optimization guided by multi-view evidence. Furthermore, we introduce a dynamically adaptive topology optimization framework during training to refine curve structures through linearization, merging, splitting, and pruning operations. Comprehensive evaluations on the ABC dataset and real-world benchmarks demonstrate our one-stage method's superiority over two-stage alternatives, particularly in producing cleaner and more robust reconstructions. Additionally, by directly optimizing parametric curves, our method significantly reduces the parameter count during training, achieving both higher efficiency and superior performance compared to existing approaches. Zhirui Gao, Renjiao Yi, Yaqiao Dai, Xuening Zhu, Wei Chen 0009, Chenyang Zhu 0002, Kai Xu 0004 |
ICCV | 2 |
| 2025 | Self-Supervised Learning of Hybrid Part-Aware 3D Representations of 2D Gaussians and Superquadrics
Zhirui Gao, Renjiao Yi, Yuhang Huang 0006, Wei Chen 0009, Chenyang Zhu 0002, Kai Xu 0004 |
ICCV | 2 |
| 2025 | BoxFusion: Reconstruction-Free Open-Vocabulary 3D Object Detection via Real-Time Multi-View Box FusionabstractAbstract Open‐vocabulary 3D object detection has gained significant interest due to its critical applications in autonomous driving and embodied AI. Existing detection methods, whether offline or online, typically rely on dense point cloud reconstruction, which imposes substantial computational overhead and memory constraints, hindering real‐time deployment in downstream tasks. To address this, we propose a novel reconstruction‐free online framework tailored for memory‐efficient and real‐time 3D detection. Specifically, given streaming posed RGB‐D video input, we leverage Cubify Anything as a pre‐trained visual foundation model (VFM) for single‐view 3D object detection, coupled with CLIP to capture open‐vocabulary semantics of detected objects. To fuse all detected bounding boxes across different views into a unified one, we employ an association module for correspondences of multi‐views and an optimization module to fuse the 3D bounding boxes of the same instance. The association module utilizes 3D Non‐Maximum Suppression (NMS) and a box correspondence matching module. The optimization module uses an IoU‐guided efficient random optimization technique based on particle filtering to enforce multi‐view consistency of the 3D bounding boxes while minimizing computational complexity. Extensive experiments on CA‐1M and ScanNetV2 datasets demonstrate that our method achieves state‐of‐the‐art performance among online methods. Benefiting from this novel reconstruction‐free paradigm for 3D object detection, our method exhibits great generalization abilities in various scenarios, enabling real‐time perception even in environments exceeding 1000 square meters. Yuqing Lan, Chenyang Zhu 0002, Zhirui Gao, Jiazhao Zhang, Renjiao Yi, Yijie Wang 0001, Kai Xu 0004 |
Comput. Graph. Forum | 6 |
| 2025 | DISCO: Efficient Diffusion Solver for large-scale Combinatorial Optimization problemsabstractCombinatorial Optimization (CO) problems are fundamentally important in numerous real-world applications across diverse industries, notably computer graphics, characterized by entailing enormous solution space and demanding time-sensitive response. Despite recent advancements in neural solvers, their limited expressiveness struggles to capture the multi-modal nature of CO landscapes. While some research has adopted diffusion models, these methods sample solutions indiscriminately from the entire NP-complete solution space with time-consuming denoising processes, limiting scalability for large-scale problems. We propose DISCO , an efficient DI ffusion S olver for large-scale C ombinatorial O ptimization problems that excels in both solution quality and inference speed. DISCO’s efficacy is twofold: First, it enhances solution quality by constraining the sampling space to a more meaningful domain guided by solution residues, while preserving the multi-modal properties of the output distributions. Second, it accelerates the denoising process through an analytically solvable approach, enabling solution sampling with very few reverse-time steps and significantly reducing inference time. This inference-speed advantage is further amplified by Jittor, a high-performance learning framework based on just-in-time compiling and meta-operators. DISCO delivers strong performance on large-scale Traveling Salesman Problems and challenging Maximal Independent Set benchmarks, with inference duration up to 5.38 times faster than existing diffusion solver alternatives. We apply DISCO to design 2D/3D TSP Art, enabling the generation of fluid stroke sequences at reduced path costs. By incorporating DISCO’s multi-modal property into a divide-and-conquer strategy, it can further generalize to solve unseen-scale instances out of the box. Hang Zhao 0018, Kexiong Yu, Yuhang Huang 0006, Renjiao Yi, Chenyang Zhu 0002, Kai Xu 0004 |
Graph. Model. | 4 |
| 2025 | CAD-NeRF: learning NeRFs from uncalibrated few-view images by CAD model retrieval
Xin Wen 0005, Xuening Zhu, Renjiao Yi, Chenyang Zhu 0002, Kai Xu 0004 |
Frontiers Comput. Sci. | 3 |
| 2025 | Generic Objects as Pose Probes for Few-Shot View SynthesisabstractRadiance fields, including NeRFs and 3D Gaussians, demonstrate great potential in high-fidelity rendering and scene reconstruction, while they require a substantial number of posed images as input. COLMAP is frequently employed for preprocessing to estimate poses. However, COLMAP necessitates a large number of feature matches to operate effectively, and struggles with scenes characterized by sparse features, large baselines, or few-view images. We aim to tackle few-view NeRF reconstruction using only 3 to 6 unposed scene images, freeing from COLMAP initializations. Inspired by the idea of calibration boards in traditional pose calibration, we propose a novel approach of utilizing everyday objects, commonly found in both images and real life, as “pose probes”. By initializing the probe object as a cube shape, we apply a dual-branch volume rendering optimization (object NeRF and scene NeRF) to constrain the pose optimization and jointly refine the geometry. PnP matching is used to initialize poses between images incrementally, where only a few feature matches are enough. PoseProbe achieves state-of-the-art performance in pose estimation and novel view synthesis across multiple datasets in experiments. We demonstrate its effectiveness, particularly in few-view and large-baseline scenes where COLMAP struggles. In ablations, using different objects in a scene yields comparable performance, showing that PoseProbe is robust to the choice of probe objects. Our project page is available at:https://zhirui-gao.github.io/PoseProbe.github.io/ Zhirui Gao, Renjiao Yi, Chenyang Zhu 0002, Ke Zhuang, Wei Chen 0009, Kai Xu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Relighting Scenes With Object Insertions in Neural Radiance FieldsabstractInserting objects into scenes and performing realistic relighting are common applications in augmented reality (AR). Previous methods focused on inserting virtual objects using CAD models or real objects from single-view images, resulting in highly limited AR application scenarios. We introduce a novel pipeline based on Neural Radiance Fields (NeRFs) for seamlessly integrating objects into scenes, from two sets of images depicting the object and scene. This approach enables novel view synthesis, realistic relighting, and supports physical interactions such as shadow casting between objects. The lighting environment is in a hybrid representation of Spherical Harmonics and Spherical Gaussians, representing both high- and low-frequency lighting components very well, and supporting non-Lambertian surfaces. Specifically, we leverage the benefits of volume rendering and introduce an innovative approach for efficient shadow rendering by comparing the depth maps between the camera view and the light source view and generating vivid soft shadows. The proposed method achieves realistic relighting effects in extensive experimental evaluations. Xuening Zhu, Renjiao Yi, Xin Wen 0005, Chenyang Zhu 0002, Kai Xu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Tensorformer: Normalized Matrix Attention Transformer for High-Quality Point Cloud ReconstructionabstractSurface reconstruction from raw point clouds has been studied for decades in the computer graphics community, which is highly demanded by modeling and rendering applications nowadays. Classic solutions, such as Poisson surface reconstruction, require point normals as extra input to perform reasonable results. Modern transformer-based methods can work without normals, while the results are less fine-grained due to limited encoding performance in local fusion from discrete points. We introduce a novel normalized matrix attention transformer (Tensorformer) to perform high-quality reconstruction. The proposedmatrix attentionallows for simultaneous point-wise and channel-wise message passing, while the previous vector attention loses neighbor point information across different channels. It brings more degree of freedom in feature learning and thus facilitates better modeling of local geometries. Our method achieves state-of-the-art on two commonly used datasets, ShapeNetCore and ABC, and attains 4% improvements on IOU on ShapeNet. Our implementation will be released upon acceptance. Hui Tian 0005, Zheng Qin 0002, Renjiao Yi, Chenyang Zhu 0002, Kai Xu 0004 |
IEEE Trans. Multim. | 3 |
| 2025 | Angio-Diff: learning a self-supervised adversarial diffusion model for angiographic geometry generation
Renjiao Yi, Xin Wen 0005, Chenyang Zhu 0002, Kai Xu 0004, Kunlun He |
Vis. Comput. | 2 |
| 2024 | DiffusionEdge: Diffusion Probabilistic Model for Crisp Edge DetectionabstractLimited by the encoder-decoder architecture, learning-based edge detectors usually have difficulty predicting edge maps that satisfy both correctness and crispness. With the recent success of the diffusion probabilistic model (DPM), we found it is especially suitable for accurate and crisp edge detection since the denoising process is directly applied to the original image size. Therefore, we propose the first diffusion model for the task of general edge detection, which we call DiffusionEdge. To avoid expensive computational resources while retaining the final performance, we apply DPM in the latent space and enable the classic cross-entropy loss which is uncertainty-aware in pixel level to directly optimize the parameters in latent space in a distillation manner. We also adopt a decoupled architecture to speed up the denoising process and propose a corresponding adaptive Fourier filter to adjust the latent features of specific frequencies. With all the technical designs, DiffusionEdge can be stably trained with limited resources, predicting crisp and accurate edge maps with much fewer augmentation strategies. Extensive experiments on four edge detection benchmarks demonstrate the superiority of DiffusionEdge both in correctness and crispness. On the NYUDv2 dataset, compared to the second best, we increase the ODS, OIS (without post-processing) and AC by 30.2%, 28.1% and 65.1%, respectively. Code: https://github.com/GuHuangAI/DiffusionEdge. Yunfan Ye, Kai Xu 0004, Yuhang Huang 0006, Renjiao Yi, Zhiping Cai |
AAAI | 4 |
| 2024 | MaskEditor: Instruct 3D Object Editing with Learned Masks
Xinyao Liu, Kai Xu 0004, Yuhang Huang 0006, Renjiao Yi, Chenyang Zhu 0002 |
PRCV (6) | 4 |
| 2024 | Learning accurate template matching with differentiable coarse-to-fine correspondence refinementabstractTemplate matching is a fundamental task in computer vision and has been studied for decades. It plays an essential role in manufacturing industry for estimating the poses of different parts, facilitating downstream tasks such as robotic grasping. Existing methods fail when the template and source images have different modalities, cluttered backgrounds, or weak textures. They also rarely consider geometric transformations via homographies, which commonly exist even for planar industrial parts. To tackle the challenges, we propose an accurate template matching method based on differentiable coarse-to-fine correspondence refinement. We use an edge-aware module to overcome the domain gap between the mask template and the grayscale image, allowing robust matching. An initial warp is estimated using coarse correspondences based on novel structure-aware information provided by transformers. This initial alignment is passed to a refinement network using references and aligned images to obtain sub-pixel level correspondences which are used to give the final geometric transformation. Extensive evaluation shows that our method to be significantly better than state-of-the-art methods and baselines, providing good generalization ability and visually plausible results even on unseen real data. Zhirui Gao, Renjiao Yi, Zheng Qin 0002, Yunfan Ye, Chenyang Zhu 0002, Kai Xu 0004 |
Comput. Vis. Media | 2 |
| 2024 | THP: Tensor-field-driven hierarchical path planning for autonomous scene exploration with depth sensorsabstractIt is challenging to automatically explore an unknown 3D environment with a robot only equipped with depth sensors due to the limited field of view. We introduce THP, a tensor field-based framework for efficient environment exploration which can better utilize the encoded depth information through the geometric characteristics of tensor fields. Specifically, a corresponding tensor field is constructed incrementally and guides the robot to formulate optimal global exploration paths and a collision-free local movement strategy. Degenerate points generated during the exploration are adopted as anchors to formulate a hierarchical TSP for global path optimization. This novel strategy can help the robot avoid long-distance round trips more effectively while maintaining scanning completeness. Furthermore, the tensor field also enables a local movement strategy to avoid collision based on particle advection. As a result, the framework can eliminate massive, time-consuming recalculations of local movement paths. We have experimentally evaluate our method with a ground robot in 8 complex indoor scenes. Our method can on average achieve 14% better exploration efficiency and 21% better exploration completeness than state-of-the-art alternatives using LiDAR scans. Moreover, compared to similar methods, our method makes path decisions 39% faster due to our hierarchical exploration strategy. Yuefeng Xi, Chenyang Zhu 0002, Yao Duan, Renjiao Yi, Hongjun He, Kai Xu 0004 |
Comput. Vis. Media | 4 |
| 2024 | STEdge: Self-Training Edge Detection With Multilayer Teaching and RegularizationabstractLearning-based edge detection has hereunto been strongly supervised with pixel-wise annotations which are tedious to obtain manually. We study the problem of self-training edge detection, leveraging the untapped wealth of large-scale unlabeled image datasets. We design a self-supervised framework with multilayer regularization and self-teaching. In particular, we impose a consistency regularization which enforces the outputs from each of the multiple layers to be consistent for the input image and its perturbed counterpart. We adopt L0-smoothing as the "perturbation" to encourage edge prediction lying on salient boundaries following the cluster assumption in self-supervised learning. Meanwhile, the network is trained with multilayer supervision by pseudo labels which are initialized with Canny edges and then iteratively refined by the network as the training proceeds. The regularization and self-teaching together attain a good balance of precision and recall, leading to a significant performance boost over supervised methods, with lightweight refinement on the target dataset. Through extensive experiments, our method demonstrates strong cross-dataset generality and can improve the original performance of edge detectors after self-training and fine-tuning. Yunfan Ye, Renjiao Yi, Zhiping Cai, Kai Xu 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Multi-Resolution Monocular Depth Map Fusion by Self-Supervised Gradient-Based CompositionabstractMonocular depth estimation is a challenging problem on which deep neural networks have demonstrated great potential. However, depth maps predicted by existing deep models usually lack fine-grained details due to convolution operations and down-samplings in networks. We find that increasing input resolution is helpful to preserve more local details while the estimation at low resolution is more accurate globally. Therefore, we propose a novel depth map fusion module to combine the advantages of estimations with multi-resolution inputs. Instead of merging the low- and high-resolution estimations equally, we adopt the core idea of Poisson fusion, trying to implant the gradient domain of high-resolution depth into the low-resolution depth. While classic Poisson fusion requires a fusion mask as supervision, we propose a self-supervised framework based on guided image filtering. We demonstrate that this gradient-based composition performs much better at noisy immunity, compared with the state-of-the-art depth map fusion method. Our lightweight depth fusion is one-shot and runs in real-time, making it 80X faster than a state-of-the-art depth fusion method. Quantitative evaluations demonstrate that the proposed method can be integrated into many fully convolutional monocular depth estimation backbones with a significant performance boost, leading to state-of-the-art results of detail enhancement on depth maps. Codes are released at https://github.com/yuinsky/gradient-based-depth-map-fusion. Yaqiao Dai, Renjiao Yi, Chenyang Zhu 0002, Hongjun He, Kai Xu 0004 |
AAAI | 2 |
| 2023 | NEF: Neural Edge Fields for 3D Parametric Curve Reconstruction from Multi-View ImagesabstractWe study the problem of reconstructing 3D feature curves of an object from a set of calibrated multi-view images. To do so, we learn a neural implicit field representing the density distribution of 3D edges which we refer to as Neural Edge Field (NEF). Inspired by NeRF [20], NEF is optimized with a view-based rendering loss where a 2D edge map is rendered at a given view and is compared to the ground-truth edge map extracted from the image of that view. The rendering-based differentiable optimization of NEF fully exploits 2D edge detection, without needing a supervision of 3D edges, a 3D geometric operator or cross-view edge correspondence. Several technical designs are devised to ensure learning a range-limited and view-independent NEF for robust edge extraction. The final parametric 3D curves are extracted from NEF with an iterative optimization method. On our benchmark with synthetic data, we demonstrate that NEF outperforms existing state-of-the-art methods on all metrics. Project page: https://yunfan1202.github.io/NEF/. Yunfan Ye, Renjiao Yi, Zhirui Gao, Chenyang Zhu 0002, Zhiping Cai, Kai Xu 0004 |
CVPR | 2 |
| 2023 | Weakly-supervised Single-view Image RelightingabstractWe present a learning-based approach to relight a single image of Lambertian and low-frequency specular objects. Our method enables inserting objects from photographs into new scenes and relighting them under the new environment lighting, which is essential for AR applications. To relight the object, we solve both inverse rendering and re-rendering. To resolve the ill-posed inverse rendering, we propose a weakly-supervised method by a low-rank constraint. To facilitate the weakly-supervised training, we contribute Relit, a large-scale (750K images) dataset of videos with aligned objects under changing illuminations. For re-rendering, we propose a differentiable specular rendering layer to render low-frequency non-Lambertian materials under various illuminations of spherical harmonics. The whole pipeline is end-to-end and efficient, allowing for a mobile app implementation of AR object insertion. Extensive evaluations demonstrate that our method achieves state-of-the-art performance. Project page: https://renjiaoyi.github.io/relighting/. Renjiao Yi, Chenyang Zhu 0002, Kai Xu 0004 |
CVPR | 1 |
| 2023 | 2D3D-MATR: 2D-3D Matching Transformer for Detection-free Registration between Images and Point CloudsabstractThe commonly adopted detect-then-match approach to registration finds difficulties in the cross-modality cases due to the incompatible keypoint detection and inconsistent feature description. We propose, 2D3D-MATR, a detection-free method for accurate and robust registration between images and point clouds. Our method adopts a coarse-to-fine pipeline where it first computes coarse correspondences between downsampled patches of the input image and the point cloud and then extends them to form dense correspondences between pixels and points within the patch region. The coarse-level patch matching is based on transformer which jointly learns global contextual constraints with self-attention and cross-modality correlations with cross-attention. To resolve the scale ambiguity in patch matching, we construct a multi-scale pyramid for each image patch and learn to find for each point patch the best matching image patch at a proper resolution level. Extensive experiments on two public benchmarks demonstrate that 2D3D-MATR outperforms the previous state-of-the-art P2-Net by around 20 percentage points on inlier ratio and over 10 points on registration recall. Our code and models are available at https://github.com/minhaolee/2D3DMATR. Minhao Li, Zheng Qin 0002, Zhirui Gao, Renjiao Yi, Chenyang Zhu 0002, Yulan Guo, Kai Xu 0004 |
ICCV | 4 |
| 2023 | EFECL: Feature encoding enhancement with contrastive learning for indoor 3D object detectionabstractGood proposal initials are critical for 3D object detection applications. However, due to the significant geometry variation of indoor scenes, incomplete and noisy proposals are inevitable in most cases. Mining feature information among these “bad” proposals may mislead the detection. Contrastive learning provides a feasible way for representing proposals, which can align complete and incomplete/noisy proposals in feature space. The aligned feature space can help us build robust 3D representation even if bad proposals are given. Therefore, we devise a new contrast learning framework for indoor 3D object detection, called EFECL, that learns robust 3D representations by contrastive learning of proposals on two different levels. Specifically, we optimize both instance-level and category-level contrasts to align features by capturing instance-specific characteristics and semantic-aware common patterns. Furthermore, we propose an enhanced feature aggregation module to extract more general and informative features for contrastive learning. Evaluations on ScanNet V2 and SUN RGB-D benchmarks demonstrate the generalizability and effectiveness of our method, and our method can achieve 12.3% and 7.3% improvements on both datasets over the benchmark alternatives. The code and models are publicly available at https://github.com/YaraDuan/EFECL . Yao Duan, Renjiao Yi, Yuanming Gao, Kai Xu 0004, Chenyang Zhu 0002 |
Comput. Vis. Media | 2 |
| 2023 | 6DOF pose estimation of a 3D rigid object based on edge-enhanced point pair featuresabstractThe point pair feature (PPF) is widely used for 6D pose estimation. In this paper, we propose an efficient 6D pose estimation method based on the PPF framework. We introduce a well-targeted down-sampling strategy that focuses on edge areas for efficient feature extraction for complex geometry. A pose hypothesis validation approach is proposed to resolve ambiguity due to symmetry by calculating the edge matching degree. We perform evaluations on two challenging datasets and one real-world collected dataset, demonstrating the superiority of our method for pose estimation for geometrically complex, occluded, symmetrical objects. We further validate our method by applying it to simulated punctures. Chenyi Liu, Renjiao Yi, Chenyang Zhu 0002, Kai Xu 0004 |
Comput. Vis. Media | 4 |
| 2023 | Delving Into Crispness: Guided Label Refinement for Crisp Edge DetectionabstractLearning-based edge detection usually suffers from predicting thick edges. Through extensive quantitative study with a new edge crispness measure, we find that noisy human-labeled edges are the main cause of thick predictions. Based on this observation, we advocate that more attention should be paid on label quality than on model design to achieve crisp edge detection. To this end, we propose an effective Canny-guided refinement of human-labeled edges whose result can be used to train crisp edge detectors. Essentially, it seeks for a subset of over-detected Canny edges that best align human labels. We show that several existing edge detectors can be turned into a crisp edge detector through training on our refined edge maps. Experiments demonstrate that deep models trained with refined edges achieve significant performance boost of crispness from 17.4% to 30.6%. With the PiDiNet backbone, our method improves ODS and OIS by 12.2% and 12.6% on the Multicue dataset, respectively, without relying on non-maximal suppression. We further conduct experiments and show the superiority of our crisp edge detection for optical flow estimation and image segmentation. Yunfan Ye, Renjiao Yi, Zhirui Gao, Zhiping Cai, Kai Xu 0004 |
IEEE Trans. Image Process. | 2 |
| 2022 | DisARM: Displacement Aware Relation Module for 3D DetectionabstractWe introduce Displacement Aware Relation Module (DisARM), a novel neural network module for enhancing the performance of 3D object detection in point cloud scenes. The core idea is extracting the most principal contextual information is critical for detection while the target is incomplete or featureless. We find that relations between proposals provide a good representation to describe the context. However, adopting relations between all the object or patch proposals for detection is inefficient, and an imbalanced combination of local and global relations brings extra noise that could mislead the training. Rather than working with all relations, we find that training with relations only between the most representative ones, or an-chors, can significantly boost the detection performance. Good anchors should be semantic-aware with no ambiguity and able to describe the whole layout of a scene with no redundancy. To find the anchors, we first perform a preliminary relation anchor module with an objectness-aware sampling approach and then devise a displacement based module for weighing the relation importance for better utilization of contextual information. This lightweight relation module leads to significantly higher accuracy of object instance detection when being plugged into the state-of-the-art detectors. Evaluations on the public benchmarks of real-world scenes show that our method achieves the state-of-the-art performance on both SUN RGB-D and Scan-Net V2. The code and models are publicly available at https://github.com/YaraDuan/DisARM. Yao Duan, Chenyang Zhu 0002, Yuqing Lan, Renjiao Yi, Xinwang Liu 0002, Kai Xu 0004 |
CVPR | 4 |
| 2020 | Leveraging Multi-View Image Sets for Unsupervised Intrinsic Image Decomposition and Highlight SeparationabstractWe present an unsupervised approach for factorizing object appearance into highlight, shading, and albedo layers, trained by multi-view real images. To do so, we construct a multi-view dataset by collecting numerous customer product photos online, which exhibit large illumination variations that make them suitable for training of reflectance separation and can facilitate object-level decomposition. The main contribution of our approach is a proposed image representation based on local color distributions that allows training to be insensitive to the local misalignments of multi-view images. In addition, we present a new guidance cue for unsupervised training that exploits synergy between highlight separation and intrinsic image decomposition. Over a broad range of objects, our technique is shown to yield state-of-the-art results for both of these tasks. Renjiao Yi, Ping Tan 0002, Stephen Lin 0001 |
AAAI | 1 |
| 2018 | Faces as Lighting Probes via Unsupervised Deep Highlight Extraction
Renjiao Yi, Chenyang Zhu 0002, Ping Tan 0002, Stephen Lin 0001 |
ECCV (9) | 1 |
| 2018 | SCORES: shape composition with recursive substructure priorsabstractWe introduce SCORES, a recursive neural network for shape composition. Our network takes as input sets of parts from two or more source 3D shapes and a rough initial placement of the parts. It outputs an optimized part structure for the composed shape, leading to high-quality geometry construction. A unique feature of our composition network is that it is not merely learning how to connect parts. Our goal is to produce a coherent and plausible 3D shape, despite large incompatibilities among the input parts. The network may significantly alter the geometry and structure of the input parts and synthesize a novel shape structure based on the inputs, while adding or removing parts to minimize a structure plausibility loss. We design SCORES as a recursive autoencoder network. During encoding, the input parts are recursively grouped to generate a root code. During synthesis, the root code is decoded, recursively, to produce a new, coherent part assembly. Assembled shape structures may be novel, with little global resemblance to training exemplars, yet have plausible substructures. SCORES therefore learns a hierarchical substructure shape prior based on per-node losses. It is trained on structured shapes from ShapeNet, and is applied iteratively to reduce the plausibility loss. We show results of shape composition from multiple sources over different categories of man-made shapes and compare with state-of-the-art alternatives, demonstrating that our network can significantly expand the range of composable shapes for assembly-based modeling. Chenyang Zhu 0002, Kai Xu 0004, Siddhartha Chaudhuri, Renjiao Yi, Hao (Richard) Zhang |
ACM Trans. Graph. | 4 |
| 2017 | Deformation-driven shape correspondence via shape recognitionabstractMany approaches to shape comparison and recognition start by establishing a shape correspondence. We "turn the table" and show that quality shape correspondences can be obtained by performing many shape recognition tasks. What is more, the method we develop computes a fine-grained, topology-varying part correspondence between two 3D shapes where the core evaluation mechanism only recognizes shapes globally. This is made possible by casting the part correspondence problem in a deformation-driven framework and relying on a data-driven "deformation energy" which rates visual similarity between deformed shapes and models from a shape repository. Our basic premise is that if a correspondence between two chairs (or airplanes, bicycles, etc.) is correct, then a reasonable deformation between the two chairs anchored on the correspondence ought to produce plausible , "chair-like" in-between shapes. Given two 3D shapes belonging to the same category, we perform a top-down, hierarchical search for part correspondences. For a candidate correspondence at each level of the search hierarchy, we deform one input shape into the other, while respecting the correspondence, and rate the correspondence based on how well the resulting deformed shapes resemble other shapes from ShapeNet belonging to the same category as the inputs. The resemblance, i.e., plausibility, is measured by comparing multi-view depth images over category-specific features learned for the various shape categories. We demonstrate clear improvements over state-of-the-art approaches through tests covering extensive sets of man-made models with rich geometric and topological variations. Chenyang Zhu 0002, Renjiao Yi, Wallace P. Lira, Ibraheem Alhashim, Kai Xu 0004, Hao (Richard) Zhang |
ACM Trans. Graph. | 2 |
| 2016 | Automatic Fence Segmentation in Videos of Dynamic ScenesabstractWe present a fully automatic approach to detect and segment fence-like occluders from a video clip. Unlike previous approaches that usually assume either static scenes or cameras, our method is capable of handling both dynamic scenes and moving cameras. Under a bottom-up framework, it first clusters pixels into coherent groups using color and motion features. These pixel groups are then analyzed in a fully connected graph, and labeled as either fence or non-fence using graph-cut optimization. Finally, we solve a dense Conditional Random Filed (CRF) constructed from multiple frames to enhance both spatial accuracy and temporal coherence of the segmentation. Once segmented, one can use existing hole-filling methods to generate a fencefree output. Extensive evaluation suggests that our method outperforms previous automatic and interactive approaches on complex examples captured by mobile devices. Renjiao Yi, Jue Wang 0001, Ping Tan 0002 |
CVPR | 1 |