VLDB 2026 Research / reviewers in the wild / expert
Shuwei Shao
dblp:304/4196
· DBLP profile ↗
17ranked-venue papers
10as first author
17since 2021 · last 2026
0000-0001-8057-1599ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Foundation Models: Distilling Geometric Priors for Lightweight Monocular Depth Estimation in EndoscopyabstractIn recent times, geometric foundation models have demonstrated remarkable performance in depth estimation tasks, benefiting from exposure to large-scale data that enables the learning of intricate geometric structures and spatial dependencies. However, their large parameter sizes and high computational complexity pose significant challenges in meeting the efficiency requirements of downstream surgical applications. Consequently, the design of a high-performance yet lightweight monocular depth estimator has become a focal point of research. To this end, we harness the rich geometric priors encoded in geometric foundation models and introduce a novel trinity distillation scheme that transfers geometric knowledge across three complementary dimensions, namely spatial, spectral and gradient, into a compact depth estimator. To further enhance prediction quality, we develop a semantic distribution alignment strategy to effectively suppress pseudo-texture artifacts arising from the limited semantic representation capability of the lightweight estimator. Extensive experiments on the SCARED, SERV-CT, Hamlyn, and C3VD datasets demonstrate that the proposed method either surpasses or achieves comparable performance to previous state-of-the-art competitors, with a smaller model size and reduced computational overhead. Code will be available at: https://github.com/ShuweiShao/LiteNet. Kejin Zhu, Shuwei Shao, Yongming Yang, Zhongyu Tian, Baochang Zhang 0001, Zhe Min |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure PriorsabstractSelf-supervised depth estimation from monocular cameras in diverse outdoor conditions, such as daytime, rain, and nighttime, is challenging due to the difficulty of learning universal representations and the severe lack of labeled real-world adverse data. Previous methods either rely on synthetic inputs and pseudo-depth labels or directly apply daytime strategies to adverse conditions, resulting in suboptimal results. In this paper, we present the first synthetic-to-real robust depth estimation framework, incorporating motion and structure priors to capture real-world knowledge effectively. In the synthetic adaptation, we transfer motion-structure knowledge inside cost volumes for better robust representation, using a frozen daytime model to train a depth estimator in synthetic adverse conditions. In the innovative real adaptation, which targets to fix synthetic-real gaps, models trained earlier identify the weather-insensitive regions with a designed consistency-reweighting strategy to emphasize valid pseudo-labels. We introduce a new regularization by gathering explicit depth distribution to constrain the model facing real-world data. Experiments show that our method outperforms the state-of-the-art across diverse conditions in multi-frame and single-frame evaluations. We achieve improvements of 7.5% and 4.3% in Ab-sRel and RMSE on average for nuScenes and Robotcar datasets (daytime, nighttime, rain). In zero-shot evaluation of DrivingStereo (rain, fog), our method generalizes better than previous ones. The code is at Syn2Real-Depth. Weilong Yan, Shuwei Shao, Robby T. Tan |
CVPR | 4 |
| 2025 | Learnable patchmatch and self-teaching for multi-frame depth estimation in monocular endoscopy
Shuwei Shao, Zhongcai Pei, Weihai Chen, Xingming Wu, Zhong Liu 0005 |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | IEBins: Iterative Elastic Bins for Monocular Depth Estimation and Completion
Shuwei Shao, Zhongcai Pei, Weihai Chen, Peter C. Y. Chen, Zhengguo Li |
Int. J. Comput. Vis. | 1 |
| 2025 | MonoDiffusion: Self-Supervised Monocular Depth Estimation Using Diffusion ModelabstractOver the past few years, self-supervised monocular depth estimation has received widespread attention. Most efforts focus on designing different types of network architectures and loss functions or handling edge cases, for example, occlusion and dynamic objects. In this work, we take another path and propose a novel conditional diffusion-based generative framework for self-supervised monocular depth estimation, dubbed MonoDiffusion. Because the depth ground-truth is unavailable in a self-supervised setting, we develop a new pseudo ground-truth diffusion process to assist the diffusion for training. Instead of diffusing at a fixed high resolution, we perform diffusion in a coarse-to-fine manner that allows for faster inference time without sacrificing accuracy or even better accuracy. Furthermore, we develop a simple yet effective contrastive depth reconstruction mechanism to enhance the denoising ability of model. It is worth noting that the proposed MonoDiffusion has the property of naturally acquiring the depth uncertainty that is essential to be implemented in safety-critical cases. Extensive experiments on the KITTI, Make3D and DIML datasets indicate that our MonoDiffusion outperforms prior state-of-the-art self-supervised competitors. The source code will be publicly available upon the acceptance. Shuwei Shao, Zhongcai Pei, Weihai Chen, Dingchi Sun, Peter C. Y. Chen, Zhengguo Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Digging into Contrastive Learning for Robust Depth Estimation with Diffusion ModelsabstractRecently, diffusion-based depth estimation methods have drawn widespread attention due to their elegant denoising patterns and promising performance. However, they are typically unreliable under adverse conditions prevalent in real-world scenarios, such as rainy, snowy, etc. In this paper, we propose a novel robust depth estimation method called D4RD, featuring a custom contrastive learning mode tailored for diffusion models to mitigate performance degradation in complex environments. Concretely, we integrate the strength of knowledge distillation into contrastive learning, building the `trinity' contrastive scheme. This scheme utilizes the sampled noise of the forward diffusion process as a natural reference, guiding the predicted noise in diverse scenes toward a more stable and precise optimum. Moreover, we extend noise-level trinity to encompass more generic feature and image levels, establishing a multi-level contrast to distribute the burden of robust perception across the overall network. Before addressing complex scenarios, we enhance the stability of the baseline diffusion model with three straightforward yet effective improvements, which facilitate convergence and remove depth outliers. Extensive experiments demonstrate that D4RD surpasses existing state-of-the-art solutions on synthetic corruption datasets and real-world weather conditions. Source code and data are available at \url{https://github.com/wangjiyuan9/D4RD}. Jiyuan Wang 0001, Chunyu Lin, Lang Nie, Kang Liao, Shuwei Shao, Yao Zhao 0001 |
ACM Multimedia | 5 |
| 2024 | F2Depth: Self-supervised indoor monocular depth estimation via optical flow consistency and feature map synthesis
Huijie Zhao, Shuwei Shao, Baochang Zhang 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | NDDepth: Normal-Distance Assisted Monocular Depth Estimation and CompletionabstractOver the past few years, monocular depth estimation and completion have been paid more and more attention from the computer vision community because of their widespread applications. In this paper, we introduce novel physics (geometry)-driven deep learning frameworks for these two tasks by assuming that 3D scenes are constituted with piece-wise planes. Instead of directly estimating the depth map or completing the sparse depth map, we propose to estimate the surface normal and plane-to-origin distance maps or complete the sparse surface normal and distance maps as intermediate outputs. To this end, we develop a normal-distance head that outputs pixel-level surface normal and distance. Afterthat, the surface normal and distance maps are regularized by a developed plane-aware consistency constraint, which are then transformed into depth maps. Furthermore, we integrate an additional depth head to strengthen the robustness of the proposed frameworks. Extensive experiments on the NYU-Depth-v2, KITTI and SUN RGB-D datasets demonstrate that our method exceeds in performance prior state-of-the-art monocular depth estimation and completion competitors. Shuwei Shao, Zhongcai Pei, Weihai Chen, Peter C. Y. Chen, Zhengguo Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | URCDC-Depth: Uncertainty Rectified Cross-Distillation With CutFlip for Monocular Depth EstimationabstractThis work aims to estimate a high-quality depth map from a single RGB image. Due to the lack of depth clues, making full use of the long-range correlation and local information is critical for accurate depth estimation. To this end, we introduce an uncertainty rectified cross-distillation between the Transformer and convolutional neural network (CNN) to achieve a comprehensive depth estimator. Specifically, we utilize the depth estimates from the Transformer branch and CNN branch as pseudo labels to teach each other. At the same time, the pixel-wise depth uncertainty is modeled to mitigate the negative impact of noisy pseudo labels. To avoid the large capacity gap induced by the strong Transformer branch deteriorating the cross-distillation, we transfer the feature maps from the Transformer to the CNN and develop coupling units to assist the weak CNN branch in leveraging the transferred features. Furthermore, we introduce CutFlip, a surprisingly simple yet highly effective data augmentation technique, which forces the model to focus on more valuable depth reasoning clues apart from the vertical image position. Extensive experiments demonstrate that our model, termedURCDC-Depth, exceeds in performance previous state-of-the-art approaches on the KITTI, NYU-Depth-v2 and SUN RGB-D datasets, with no additional computational burden in the evaluation phase. The source code will be publicly available upon acceptance. The source code is available athttps://github.com/ShuweiShao/URCDC-Depth. Shuwei Shao, Zhongcai Pei, Weihai Chen, Zhong Liu 0005, Zhengguo Li |
IEEE Trans. Multim. | 1 |
| 2023 | NDDepth: Normal-Distance Assisted Monocular Depth EstimationabstractMonocular depth estimation has drawn widespread attention from the vision community due to its broad applications. In this paper, we propose a novel physics (geometry)-driven deep learning framework for monocular depth estimation by assuming that 3D scenes are constituted by piece-wise planes. Particularly, we introduce a new normal-distance head that outputs pixel-level surface normal and plane-to-origin distance for deriving depth at each position. Meanwhile, the normal and distance are regularized by a developed plane-aware consistency constraint. We further integrate an additional depth head to improve the robustness of the proposed framework. To fully exploit the strengths of these two heads, we develop an effective contrastive iterative refinement module that refines depth in a complementary manner according to the depth uncertainty. Extensive experiments indicate that the proposed method exceeds previous state-of-the-art competitors on the NYU-Depth-v2, KITTI and SUN RGB-D datasets. Notably, it ranks 1st among all submissions on the KITTI depth prediction online benchmark at the submission time. The source code is available at https://github.com/ShuweiShao/NDDepth. Shuwei Shao, Zhongcai Pei, Weihai Chen, Xingming Wu, Zhengguo Li |
ICCV | 1 |
| 2023 | Monocular Depth Estimation: A SurveyabstractMonocular depth estimation is an ill-posed task in computer vision, which holds great significance in the fields such as artificial intelligence, virtual reality, augmented reality, path planning, unmanned driving, and navigation guidance. The primary objective of monocular depth estimation is to predict the depth value of each pixel or infer depth information, given just a single red-green-blue (RGB) image as input. Traditional monocular depth estimation methods rely on limited depth cues, such as strict scene conditions. With the significant advancements in computer vision and artificial intelligence, monocular depth estimation using deep learning has been extensively researched and has yielded substantial results. This paper presents a comprehensive survey of monocular depth estimation. Firstly, we give an overall introduction to monocular depth estimation and explain it from traditional and deep learning-based methods, respectively. To specify, supervised, self-supervised and semi-supervised models are described in detail in deep learning-based methods. Additionally, we introduce publicly available benchmark datasets and evaluation metrics commonly used in this field. Finally, we discuss the current challenges and promising prospects for the development of monocular depth estimation. Dong Wang 0051, Zhong Liu 0005, Shuwei Shao, Xingming Wu, Weihai Chen, Zhengguo Li |
IECON | 3 |
| 2023 | IEBins: Iterative Elastic Bins for Monocular Depth EstimationabstractMonocular depth estimation (MDE) is a fundamental topic of geometric computer vision and a core technique for many downstream applications. Recently, several methods reframe the MDE as a classification-regression problem where a linear combination of probabilistic distribution and bin centers is used to predict depth. In this paper, we propose a novel concept of iterative elastic bins (IEBins) for the classification-regression-based MDE. The proposed IEBins aims to search for high-quality depth by progressively optimizing the search range, which involves multiple stages and each stage performs a finer-grained depth search in the target bin on top of its previous stage. To alleviate the possible error accumulation during the iterative process, we utilize a novel elastic target bin to replace the original target bin, the width of which is adjusted elastically based on the depth uncertainty. Furthermore, we develop a dedicated framework composed of a feature extractor and an iterative optimizer that has powerful temporal context modeling capabilities benefiting from the GRU-based architecture. Extensive experiments on the KITTI, NYU-Depth-v2 and SUN RGB-D datasets demonstrate that the proposed method surpasses prior state-of-the-art competitors. The source code is publicly available at https://github.com/ShuweiShao/IEBins. Shuwei Shao, Zhongcai Pei, Xingming Wu, Zhong Liu 0005, Weihai Chen, Zhengguo Li |
NeurIPS | 1 |
| 2023 | A geometry-aware deep network for depth estimation in monocular endoscopy
Yongming Yang, Shuwei Shao, Chengdong Wu 0001, Hao Liu 0008 |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Self-Supervised Monocular Depth Estimation With Self-Reference Distillation and Disparity Offset RefinementabstractMonocular depth estimation plays a fundamental role in computer vision. Due to the costly acquisition of depth ground truth, self-supervised methods that leverage adjacent frames to establish a supervision signal have emerged as the most promising paradigms. In this work, we propose two novel ideas to improve self-supervised monocular depth estimation: 1) self-reference distillation and 2) disparity offset refinement. Specifically, we use a parameter-optimized model as the teacher updated as the training epochs to provide additional supervision during the training process. The teacher model has the same structure as the student model, with weights inherited from the historical student model. In addition, a multiview check is introduced to filter out the outliers produced by the teacher model. Furthermore, we leverage the contextual consistency between high-level and low-level features to obtain multiscale disparity offsets, which are used to refine the disparity output incrementally by aligning disparity information at different scales. The experimental results on the KITTI and Make3D datasets show that our method outperforms previous state-of-the-art competitors. Zhong Liu 0005, Shuwei Shao, Xingming Wu, Weihai Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Towards Comprehensive Monocular Depth Estimation: Multiple Heads are Better Than OneabstractDepth estimation attracts widespread attention in the computer vision community. However, it is still quite difficult to recover an accurate depth map using only one RGB image. We observe a phenomenon that existing methods tend to fail in different cases, caused by differences in network architecture, loss function and so on. In this work, we investigate into the phenomenon and propose to integrate the strengths of multiple weak depth predictor to build a comprehensive and accurate depth predictor, which is critical for many real-world applications, e.g., 3D reconstruction. Specifically, we construct multiple base (weak) depth predictors by utilizing different Transformer-based and convolutional neural network (CNN)-based architectures. Transformer establishes long-range correlation while CNN preserves local information ignored by Transformer due to the spatial inductive bias. Therefore, the coupling of Transformer and CNN contributes to the generation of complementary depth estimates, which are essential to achieve a comprehensive depth predictor. Then, we design mixers to learn from multiple weak predictions and adaptively fuse them into a strong depth estimate. The resultant model, which we refer to as Transformer-assisted depth ensembles (TEDepth). On the standard NYU-Depth-v2 and KITTI datasets, we thoroughly explore how the neural ensembles affect the depth estimation and demonstrate that our TEDepth achieves better results than previous state-of-the-art approaches. To validate the generalizability across cameras, we directly apply the models trained on NYU-Depth-v2 to the SUN RGB-D dataset without any fine-tuning, and the superior results emphasize its strong generalizability. Shuwei Shao, Zhongcai Pei, Zhong Liu 0005, Weihai Chen, Wentao Zhu 0001, Xingming Wu, Baochang Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Self-Supervised monocular depth and ego-Motion estimation in endoscopy: Appearance flow to the rescue
Shuwei Shao, Zhongcai Pei, Weihai Chen, Wentao Zhu 0001, Xingming Wu, Dianmin Sun, Baochang Zhang 0001 |
Medical Image Anal. | 1 |
| 2021 | Self-Supervised Learning for Monocular Depth Estimation on Minimally Invasive Surgery ScenesabstractSelf-supervised learning algorithms that compute depth map from monocular videos have achieved remarkable performance on urban scenes and have been applied extensively. These techniques still face significant challenges, however, when applied directly to endoscopic videos because of the brightness variations from frame to frame and inadequate representation learning during the training phase. Inspired by the optical flow for motion alignment between adjacent frames, we design a AFNet with structural stability loss and residual-based smoothness loss to learn the appearance flow across adjacent frames, which handles the brightness inconsistency issue efficaciously. In addition, we propose a novel self-attention mechanism named feature scaling module to alleviate the inadequate representation learning problem. In a comparison study to the current state-of-the-art self-supervised methods explored for urban videos on the SCARED dataset, the developed model surpasses existing methods by a large margin. Shuwei Shao, Zhongcai Pei, Weihai Chen, Baochang Zhang 0001, Xingming Wu, Dianmin Sun, David S. Doermann |
ICRA | 1 |