Shen Cai

dblp:127/2781 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 9 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Deep-Augmented Point Pair Features for 6D Pose Estimation of Flat Objects
Sipei Fang, Shibin Xie, Yubin Hu, Shen Cai
ICIC (18)6
2026 A Concise P3P Method Based on Homography Decomposition
Shibin Xie, Shen Cai
ICIC (1)6
2026 Geometry-Guided Depth Correction for Metric Relative Pose Estimation
abstract
In recent years, Monocular Depth Estimation (MDE) has evolved from predicting affine-invariant relative depth to estimating metric-scale (absolute) depth. However, local geometric inconsistencies in single-view depth maps and scale inconsistencies across different views still severely hinder their practical application in 3D matching and relative pose estimation. To address these challenges, we propose a geometry-guided depth correction framework for metric-scale relative pose estimation. Our approach first leverages pre-trained foundation models to extract initial metric depth, semi-dense correspondences, and high-dimensional semantic features from dual-view images. We then introduce a local depth refinement module to correct geometric deviations. Finally, the corrected depth of stereo-matched pairs is integrated into a differentiable RANSAC framework to jointly optimize the relative pose with consistent scale. Experiments on ScanNet and 7-Scenes demonstrate that our method achieves superior performance and robustness across various challenging scenarios.
Shibin Xie, Xiaokang Fang, Yanting Zhang 0001, Shen Cai
ICMR8
2025 Three-view Focal Length Recovery From Homographies
abstract
In this paper, we propose a novel approach for recovering focal lengths from three-view homographies. By examining the consistency of normal vectors between two homographies, we derive new explicit constraints between the focal lengths and homographies using an elimination technique. We demonstrate that three-view homographies provide two additional constraints, enabling the recovery of one or two focal lengths. We discuss four possible cases, including three cameras having an unknown equal focal length, three cameras having two different unknown focal lengths, three cameras where one focal length is known, and the other two cameras have equal or different unknown focal lengths. All the problems can be converted into solving polynomials in one or two unknowns, which can be efficiently solved using Sturm sequence or hidden variable technique. Evaluation using both synthetic and real data shows that the proposed solvers are both faster and more accurate than methods relying on existing two-view solvers. The code and data are available on https://github.com/kocurvik/hf.
Yaqing Ding 0001, Viktor Kocur, Zuzana Berger Haladová, Qianliang Wu, Shen Cai, Jian Yang 0003, Zuzana Kukelova
CVPR5
2025 Neural Implicit Reconstruction and Fast Rendering Based on Dual Spherical Shell
abstract
In recent years, neural implicit representation has gained widespread attention for 3D model reconstruction and rendering. In particular, Signed Distance Field (SDF) has become a prevalent technique in these applications. However, existing SDF-based reconstruction and rendering algorithms often suffer from high storage requirements and slow rendering speeds. To overcome these limitations, we propose a novel dual spherical shell (DSS) representation that enables accurate reconstruction while maintaining low storage demands. By leveraging a Multi-Layer Perceptron (MLP) to overfit the SDF within the interlayer of concentric spheres, our method facilitates model compression and reconstruction using neural SDFs. Furthermore, our approach improves rendering efficiency by exploiting the intrinsic properties of spherical shells, allowing us to achieve early termination of ray tracing and parallel sphere tracing along each ray. Additionally, our technique is applicable to neural radiance fields (NeRFs), optimizing the sampling process by reducing the number of points and concentrating on regions near the surface. Experimental results across various models validate the effectiveness of our proposed methodology. Source codes are available at https://github.com/cscvlab/Dual-Spherical-Shell.
Binghao Wang, Shen Cai
ICME5
2025 Inverse Farthest Point Sampling (IFPS): A Universal and Hierarchical Shell Representation for Discrete Data
Nayu Ding, Long Wan, Zhijun Fang 0001, Shen Cai, Lin Gao 0004
ICMR7
2025 Fast and Interpretable 2D Homography Decomposition: Similarity-Kernel-Similarity and Affine-Core-Affine Transformations
abstract
In this article, we present two fast and interpretable decomposition methods for 2D homography, which are named Similarity-Kernel-Similarity (SKS) and Affine-Core-Affine (ACA) transformations respectively. Under the minimal 4-point configuration, two similarity transformations in SKS are computed by two anchor points on source and target planes, respectively. Then, the other two point correspondences can be exploited to compute the middle kernel transformation with only four parameters. Furthermore, ACA uses three anchor points to compute the source and the target affine transformations, followed by computation of the middle core transformation utilizing the other one point correspondence. ACA can compute a homography up to a scale with only 85 floating-point operations (FLOPs), without even any division operations. Therefore, as a plug-in module, ACA facilitates various traditional feature-based Random Sample Consensus (RANSAC) pipelines, as well as deep homography pipelines estimating 4-point offsets. In addition to the advantages of geometric parameterization and computational efficiency, SKS and ACA can express each element of homography by a polynomial of input coordinates (7th degree to 9th degree), extend the existing essential Similarity-Affine-Projective (SAP) decomposition and calculate 2D affine transformations in a unified way.
Shen Cai, Zhanhao Wu, Lingxi Guo, Jiachun Wang, Junchi Yan, Shuhan Shen
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
abstract
Point-pixel registration between LiDAR point clouds and camera images is a fundamental yet challenging task in autonomous driving and robotic perception. A key difficulty lies in the modality gap between unstructured point clouds and structured images, especially under sparse single-frame LiDAR settings. Existing methods typically extract features separately from point clouds and images, then rely on hand-crafted or learned matching strategies. This separate encoding fails to bridge the modality gap effectively, and more critically, these methods struggle with the sparsity and noise of single-frame LiDAR, often requiring point cloud accumulation or additional priors to improve reliability. Inspired by recent progress in detector-free matching paradigms, we revisit the projection-based approach and introduce the detector-free framework for direct point-pixel matching between LiDAR and camera views. To further enhance matching reliability, we introduce a repeatability scoring mechanism that acts as a soft visibility prior. This guides the network to suppress unreliable matches in regions with low intensity variation, improving robustness under sparse input. Extensive experiments on KITTI, nuScenes, and MIAS-LCEC-TF70 benchmarks demonstrate that our method achieves state-of-the-art performance, outperforming prior approaches on nuScenes (even those relying on accumulated point clouds), despite using only single-frame LiDAR.
Yanting Zhang 0001, Fangjun Ding, Shen Cai, Yanchao Dong, Rui Fan 0001
IEEE Trans Autom. Sci. Eng.5
2024 Unsigned Orthogonal Distance Fields: An Accurate Neural Implicit Representation for Diverse 3D Shapes
abstract
Neural implicit representation of geometric shapes has witnessed considerable advancements in recent years. However, common distance field based implicit represen-tations, specifically signed distance field (SDF) for water-tight shapes or unsigned distance field (UDF) for arbitrary shapes, routinely suffer from degradation of reconstruction accuracy when converting to explicit surface points and meshes. In this paper, we introduce a novel neural implicit representation based on unsigned orthogonal distance fields (UODFs). In UODFs, the minimal unsigned distance from any spatial point to the shape surface is de-fined solely in one orthogonal direction, contrasting with the multi-directional determination made by SDF and UDF. Consequently, every point in the 3D UODFs can directly access its closest surface points along three orthogonal di-rections. This distinctive feature leverages the accurate re-construction of surface points without interpolation errors. We verify the effectiveness of UODFs through a range of re-construction examples, extending from simple watertight or non-watertight shapes to complex shapes that include hol-lows, internal or assembling structures.
Long Wan, Nayu Ding, Shuhan Shen, Shen Cai, Lin Gao 0004
CVPR6
2024 Taming Serverless Cold Start of Cloud Model Inference With Edge Computing
abstract
Serverless computing is envisioned as the de-facto standard for next-generation cloud computing. However, the cold start dilemma has impeded its adoption by delay-sensitive and burst applications. In this paper, we propose to tame serverless cold start in a cloud inference system with edge computing. Specifically, the proposed solution smooths the serverless cloud workload with user-owned edge computing, reducing the number of cold starts. Leveraging the configurability of requests and serverless functions, the proposed solution further reduces the transmission latency and serverless cost by adapting request configuration (e.g., image resolution) and function configuration (e.g., memory). To alleviate the potential inference accuracy degradation incurred by configuration adaption, we aim to strike a nice balance between inference latency, cost, and accuracy. However, achieving this goal is non-trivial since the underlying optimization is non-convex and involves future uncertain information. To simultaneously address dual challenges, the presented cold-start-aware online algorithms apply the regularization technique to decompose the problem into separate convex subproblems. Then, it applies lazy switching to smooth the number of provisioned functions and thus reduces the cold start. Through rigorous theoretical analysis, realistic prototype evaluations on AWS Lambda, and trace-driven simulations, we comprehensively validate the theoretical and empirical performance of our proposed solution.
Kongyange Zhao, Zhi Zhou 0006, Lei Jiao 0002, Shen Cai, Fei Xu 0009, Xu Chen 0004
IEEE Trans. Mob. Comput.4
2022 High-fidelity 3D Model Compression based on Key Spheres
abstract
In recent years, neural signed distance function (SDF) has become one of the most effective representation methods for 3D models. By learning continuous SDFs in 3D space, neural networks can predict the distance from a given query space point to its closest object surface, whose positive and negative signs denote inside and outside of the object, respectively. Training a specific network for each 3D model, which individually embeds its shape, can realize compressed representation of objects by storing fewer network (and possibly latent) parameters. Consequently, reconstruction through network inference and surface recovery can be achieved. In this paper, we propose an SDF prediction network using explicit key spheres as input. Key spheres are extracted from the internal space of objects, whose centers either have relatively larger SDF values (sphere radii), or are located at essential positions. By inputting the spatial information of multiple spheres which imply different local shapes, the proposed method can significantly improve the reconstruction accuracy with a negligible storage cost. Compared to previous works, our method achieves the high-fidelity and high-compression 3D object coding and reconstruction. Experiments conducted on three datasets verify the superior performance of our method.
Yuanzhan Li, Yuqi Liu 0001, Shen Cai, Yanting Zhang 0001
DCC5
2022 An Efficient End-to-End 3D Voxel Reconstruction based on Neural Architecture Search
abstract
Using neural networks to represent 3D objects has become popular. However, many previous works employ neural networks with fixed architecture and size to represent different 3D objects, which lead to excessive network parameters for simple objects and limited reconstruction accuracy for complex objects. For each 3D model, it is desirable to have an end-to-end neural network with as few parameters as possible to achieve high-fidelity reconstruction. In this paper, we propose an efficient voxel reconstruction method utilizing neural architecture search (NAS) and binary classification. Taking the number of layers, the number of nodes in each layer, and the activation function of each layer as the search space, a specific network architecture can be obtained based on reinforcement learning technology. Furthermore, to get rid of the traditional surface reconstruction algorithms (e.g., marching cube) used after network inference, we complete the end-to-end network by classifying binary voxels. Compared to other signed distance field (SDF) prediction or binary classification networks, our method achieves significantly higher reconstruction accuracy using fewer network parameters.
Yongdong Huang, Yuanzhan Li, Xulong Cao, Shen Cai, Yuqi Liu 0001
ICPR5
2022 Spherical Transformer: Adapting Spherical Signal to Convolutional Networks
Yuqi Liu 0001, Haikuan Du, Shen Cai
PRCV (3)4
2021 SN-Graph: A Minimalist 3D Object Representation for Classification
abstract
Using deep learning techniques to process 3D objects has achieved many successes. However, few methods focus on the representation of 3D objects, which could be more effective for specific tasks than traditional representations, such as point clouds, voxels, and multi-view images. In this paper, we propose a Sphere Node Graph (SN-Graph) to represent 3D objects. Specifically, we extract a certain number of internal spheres (as nodes) from the signed distance field (SDF), and then establish connections (as edges) among the sphere nodes to construct a graph, which is seamlessly suitable for 3D analysis using graph neural network (GNN). Experiments conducted on the ModelNet40 dataset show that when there are fewer nodes in the graph or the tested objects are rotated arbitrarily, the classification accuracy of SN-Graph is significantly higher than the state-of-the-art methods.
Yuqi Liu 0001, Shen Cai, Yanting Zhang 0001, Yuanzhan Li, Xiaoyu Chi
ICME4
2020 InSphereNet: A Concise Representation and Classification Method for 3D Object
Haikuan Du, Shen Cai
MMM (2)4
2017 A versatile homography computation method based on two real points
Shen Cai, Zhanhao Wu, Yuncai Liu
Image Vis. Comput.2
2016 A rapid method for detecting objects with rectangular structures based on line correspondences
Jiachun Wang, Shen Cai
Multim. Tools Appl.3
2013 Camera calibration with enclosing ellipses by an extended application of generalized eigenvalue decomposition
Shen Cai, Longxiang Huang, Yuncai Liu
Mach. Vis. Appl.1