VLDB 2026 Research / reviewers in the wild / expert
Cheng Han 0002
dblp:53/6096-2
· DBLP profile ↗
15ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0003-3735-0162ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GADRNet: Geometry-Prior-Guided Adaptive Distortion Rectification Network for Panoramic Depth EstimationabstractPanoramic depth estimation is fundamental to comprehensive 3D scene understanding and serves as a core component of many panoramic vision tasks. However, panoramic images are typically represented in 2D equirectangular projection, which introduces severe spatial distortions from the equator to the poles, posing a central challenge for accurate depth estimation. Existing methods address this issue but fail to fully exploit the intrinsic geometric priors of panoramic imagery. In this paper, we propose a Geometry-Prior-Guided Adaptive Distortion Rectification Network (GADRNet) for panoramic depth estimation. By explicitly modeling panoramic geometric priors and introducing an adaptive receptive-field selection mechanism, GADRNet effectively improves monocular panoramic depth estimation. Specifically, we introduce a Distortion-aware Weight Map that adaptively modulates feature responses according to regional distortion levels, guiding the model to emphasize less-distorted equatorial regions. Moreover, we propose an Adaptive Geometric Distortion Rectification Module that injects geometric priors into dual-scale deformable convolutions, enabling adaptive receptive-field selection to address spatially varying distortions and extract multi-scale features. To further enhance representation learning, we design a Global-Local Scene Understanding Module to jointly capture global context and fine-grained local details. In addition, a knowledge distillation strategy is incorporated to further improve performance and generalization. Extensive experiments on two public real-world panoramic benchmarks demonstrate that GADRNet outperforms existing methods and achieves superior performance. Guoxu Li, Cheng Han 0002, Chao Zhang 0105, Tongzhou Zhang 0001 |
ICMR | 2 |
| 2026 | SketchFormer3D: Generating 3D shapes from sketches with implicit SDF priors via diffusion models
Cheng Han 0002, Chao Zhang 0105, Yu Liu 0072 |
Expert Syst. Appl. | 3 |
| 2026 | Mamba - CNN collaborative learning for panoramic semantic segmentation via online knowledge distillation
Chao Xu 0021, Jiayue Xu, Cheng Han 0002, Hua Li 0024 |
Expert Syst. Appl. | 3 |
| 2026 | Enhanced Multi-Scale PoseNet for Self-Supervised Monocular Depth EstimationabstractMonocular depth estimation is essential for 3D perception in applications such as autonomous driving and robotics. Self-supervised methods avoid depth labels but often rely on shallow pose networks with weak temporal modeling, leading to unstable predictions. We propose EMSP-Net, an Enhanced Multi-Scale PoseNet for self-supervised monocular depth estimation. It introduces a hierarchical feature fusion encoder, a temporal attention-context decoder, and a pose consistency loss to jointly improve feature extraction, temporal stability, and geometric constraints. On the KITTI dataset, EMSP-Net achieved an absolute relative error of 0.105 and a squared relative error of 0.708. In the Make3D cross-domain test, its strong robustness was further demonstrated. Chao Zhang 0105, Cheng Han 0002, Tiancheng Shao, Shichao Zhao |
IEEE Signal Process. Lett. | 3 |
| 2025 | Adaptive Occupancy Map Guided Neural Representation for Video Based Point Cloud Attribute CompressionabstractThe advancement of video-based point cloud(V-PCC) is essential for the widespread adoption of point cloud technology. However, it faces significant challenges, including the occupancy map (OM) containing numerous empty areas due to improper patch packaging, and serious compression time-consuming. Excessive empty areas can burden the subsequent video compression. Prolonged processing times may hinder the widespread adoption of V-PCC. Thus we propose data-adaptive patch packing (DAPP) to reduce empty areas during the projection process. In addition, considering the progressive nature of neural representations for videos, we propose OM-guided neural representations for videos (OM-NeRV), which utilizes advanced neural representation techniques to compress video effectively. First, it utilizes the OM to enable the network to prioritize processing valuable data when representing point clouds. Next, by utilizing context features, the residual motion features are combined with the current frame's features, thereby enhancing the network's representational capacity. Experiments have shown that our method outperforms the latest proposed advanced methods in terms of attribute rate-distortion performance and time efficiency. Yirong Chi, Cheng Han 0002, Shiyu Lu, Hao Luo 0002 |
DCC | 2 |
| 2024 | Objective Evaluation of VR Sickness and Analysis of Its Relationship with VR Presence
Cheng Han 0002, Yuechen Zhang, Yongqing Cai |
ICIC (10) | 3 |
| 2024 | Sttcnerf: Style Transfer of Neural Radiance Fields for 3d Scene Based on Texture Consistency ConstraintabstractNeural radiance fields (NeRF) have been successfully applied to many visual tasks. Due to the limitations of traditional style transfer methods in the 3D domain, style transfer methods based on NeRF have emerged. However, existing methods fail to generate stylized images with clear scene textures and high cross-view consistency. Therefore, we design STTCNeRF, a novel method for 3D scene style transfer. We propose a texture consistency loss based on our framework's Merge Net to capture the source textures and inherent consistency, achieving clear and consistent stylized effects. Meanwhile, we propose comprehensive latent style codes as conditional features, remapping them through our proposed Style Net to obtain style vectors conforming to the distributions of 2D stylized images. We utilize the Merge Net to render stylized images that are both reasonable and natural. Extensive experimental results show that STTCNeRF outperforms existing methods in terms of visual perception, image quality and cross-view consistency. Wudi Chen, Chao Zhang 0105, Cheng Han 0002, Yanjie Ma, Yongqing Cai |
ICME | 3 |
| 2024 | Effective fusion module with dilation convolution for monocular panoramic depth estimateabstractAbstract Depth estimation from monocular panoramic image is a crucial step in 3D reconstruction, which is a close relationship with virtual reality and metaverse technologies. In recent years, some methods, such as HRDFuse, BiFuse++, and UniFuse, have employed a two‐branch neural network leveraging two common projections: equirectangular and cubemap projections (CMPs). The equirectangular projection (ERP) provides a complete field of view but introduces distortion, while the CMP avoids distortion but introduces discontinuity at the boundary of the cube. In order to address the issue of distortion and discontinuity, the authors propose an efficient depth estimation fusion module to balance the feature mapping of the two projections. Moreover, for the ERP, the authors propose a novel inflated network architecture to extend the receptive field and effectively harness visual information. Extensive experiments show that the authors’ method predicts more clear boundaries and accurate depth results while outperforming mainstream panoramic depth estimation algorithms. Cheng Han 0002, Yongqing Cai, Xinpeng Pan |
IET Image Process. | 1 |
| 2023 | MsF-HigherHRNet: Multi-scale Feature Fusion for Human Pose Estimation in Crowded Scenes
Cuihong Yu, Cheng Han 0002, Chao Zhang 0105 |
CAD/Graphics | 2 |
| 2023 | PCformer: A parallel convolutional transformer network for 360° depth estimationabstractAbstract 360° depth estimation has been extensively studied because 360° images provide a full field of view of the surrounding environment as well as a detailed description of the entire scene. However, most well‐studied convolutional neural networks (CNNs) for 360° depth estimation can extract local features well, but fail to capture rich global features from the panorama due to a fixed receptive field in CNNs. PCformer, a parallel convolutional transformer network that combines the benefits of CNNs and transformers, is proposed for 360° depth estimation. The transformer has the nature to model long‐range dependency and extract global features. With PCformer, both global dependency and local spatial features can be efficiently captured. To fully incorporate global and local features, a dual attention fusion module is designed. Besides, a distortion‐weighted loss function is designed to reduce the distortion in panoramas. Extensive experiments demonstrate that the proposed method achieves competitive results against the state‐of‐the‐art methods on three benchmark datasets. Additional experiments also demonstrate that the proposed model has benefits in terms of model complexity and generalisation capability. Chao Xu 0021, Cheng Han 0002, Chao Zhang 0105 |
IET Comput. Vis. | 3 |
| 2022 | Learning high-quality depth map from 360° multi-exposure imagery
Chao Xu 0021, Cheng Han 0002, Chao Zhang 0105 |
Multim. Tools Appl. | 3 |
| 2020 | Upper body motion recognition based on key frame and random forest regression
Bo Li 0076, Baoxing Bai, Cheng Han 0002 |
Multim. Tools Appl. | 3 |
| 2020 | Hand Gesture Recognition Enhancement Based on Spatial Fuzzy Matching in Leap MotionabstractGesture recognition is an important human-computer interaction interface. This article introduces a novel hand gesture recognition system based on Leap Motion gen.2. In this system, a spatial fuzzy matching (SFM) algorithm is first presented by matching and fusing spatial information to construct a fused gesture dataset. For dynamic hand recognition, an initial frame correction strategy based on SFM is proposed to fast initialize the trajectory of test gesture with respect to the gesture dataset. A notable feature of this system is that it can run on ordinary laptops due to the small size of the fused dataset, which accelerates the calculation of recognition rate. Experimental results show that the system recognizes static hand gestures at recognition rates of 94%-100% and over 90% of dynamic gestures using our collected dataset. This can greatly enhance the usability of Leap Motion. Hua Li 0024, Huan Wang 0006, Cheng Han 0002, Jianping Zhao 0002 |
IEEE Trans. Ind. Informatics | 4 |
| 2019 | Gesture Recognition Based on Kinect v2 and Leap Motion Data FusionabstractThis study proposed a method for multiple motion-sensitive devices (i.e. one Kinect v2 and two Leap Motions) to integrate gesture data in Unity. Other depth cameras could replace the Kinect. The general steps in integrating gesture data for motion-sensitive devices were introduced as follows. (1) A method was proposed to recognize the fingertip from depth images for the Kinect v2. (2) Coordinates observed by three motion-sensitive devices were aligned in space in three steps. First, preliminary coordinate conversion parameters were obtained through joint calibration of the three devices. Second, two types of devices were approached to the observed value of the standard Leap Motion by the least squares method twice (i.e. one Kinect and one Leap Motion on the first round, then two Leap Motions on the second round). (3) Data of the three devices were aligned with time by using Unity while applying the data plan. On this basis, a human hand interacted with a virtual object in Unity. Experimental results demonstrated that the proposed method had a small recognition error of hand joints and realized the natural interaction between the human hand and virtual objects. Bo Li 0076, Chao Zhang 0105, Cheng Han 0002, Baoxing Bai |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2017 | Autonomous Perceptual Projection Correction Technique of Deep Heterogeneous Surface
Baoxing Bai, Cheng Han 0002, Chao Zhang 0105, Yuying Du |
ICONIP (3) | 3 |