EDBT 2026 Demo / reviewers in the wild / expert
Kunhong Li 0001
dblp:245/4906-1
· DBLP profile ↗
13ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0003-3297-1824ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EA-GPnP: Efficient and Accurate Generalized-Perspective-n-Point Solution via Optimized Null Space Analysis
Yi Zhang 0130, Baoqiong Wang, Kunhong Li 0001, Xiuqi Wang, Ye Zhang 0037, Yueqiang Zhang, Yulan Guo |
IEEE Trans. Robotics | 3 |
| 2025 | DualNet: Robust Self-Supervised Stereo Matching with Pseudo-Label SupervisionabstractSelf-supervised stereo matching has drawn attention due to its ability to estimate disparity without needing ground-truth data. However, existing self-supervised stereo matching methods heavily rely on the photo-metric consistency assumption, which is vulnerable to natural disturbances, resulting in ambiguous supervision and inferior performance compared to the supervised ones. To relax the limitation of the photo-metric consistency assumption and even bypass this assumption, we propose a novel self-supervised framework named DualNet, which consists of two key steps: robust self-supervised teacher learning and pseudo-label supervised student training. Specifically, the teacher model is first trained in a self-supervised manner with a focus on feature-metric consistency and data augmentation consistency. Then, the output of the teacher model is geometrically constrained to obtain high-quality pseudo labels. Benefiting from these high-quality pseudo labels, the student model can outperform its teacher model by a large margin. With the two well-designed steps, the proposed framework DualNet ranks 1st among all self-supervised methods on multiple benchmarks, surprisingly even outperforming several supervised counterparts. Yun Wang 0053, Jiahao Zheng 0001, Chenghao Zhang 0003, Zhanjie Zhang, Kunhong Li 0001, Junjie Hu 0003 |
AAAI | 5 |
| 2025 | Self-Distilled Stereo Matching: Real-Time Domain Generalization for Robotic Depth PerceptionabstractWhile human vision inherently achieves robust cross-domain depth estimation through binocular coordination, robotic systems employing stereo matching still confront significant challenges in maintaining robustness across domains when performing real-time environmental depth perception. Furthermore, most stereo matching methods struggle with challenging regions such as object boundaries and non-overlapping areas on the left side of the left image, resulting in disparity maps that are relatively indistinct and lacking fine details. In this paper, we propose Learning More in Challenging Areas (LMC) to alleviate this problem, which enhances the domain generalization of the model through targeted training on challenging regions. LMC is a simple yet effective data-driven training framework primarily based on self-distillation. Specifically, 1) We pre-train models on a high-frequency dataset to improve perception ability on object boundaries; 2) We develop a self-distillation training strategy to benefit learning in non-overlapping areas on the left side of the left image; 3) We design an adaptive difficult area mask to balance the loss weight on other undefined challenging regions. Under our proposed training framework, GwcNet achieves 33% and 23% performance improvements in autonomous driving benchmarks KITTI 2012 and KITTI 2015 respectively, while preserving real-time inference efficiency without computational overhead. Xuxin Zhang, Kunhong Li 0001, Runqing Jiang, Ye Zhang 0037, Yulan Guo |
IROS | 2 |
| 2025 | ADStereo: Efficient Stereo Matching With Adaptive Downsampling and Disparity AlignmentabstractThe balance between accuracy and computational efficiency is crucial for the applications of deep learning-based stereo matching algorithms in real-world scenarios. Since matching cost aggregation is usually the most computationally expensive component, a common practice is to construct cost volumes at a low resolution for aggregation and then directly regress a high-resolution disparity map. However, current solutions often suffer from limitations such as the loss of discriminative features caused by downsampling operations that treat all pixels equally, and spatial misalignment resulting from repeated downsampling and upsampling. To overcome these challenges, this paper presents two sampling strategies: the Adaptive Downsampling Module (ADM) and the Disparity Alignment Module (DAM), to prioritize real-time inference while ensuring accuracy. The ADM leverages local features to learn adaptive weights, enabling more effective downsampling while preserving crucial structure information. On the other hand, the DAM employs a learnable interpolation strategy to predict transformation offsets of pixels, thereby mitigating the spatial misalignment issue. Building upon these modules, we introduce ADStereo, a real-time yet accurate network that achieves highly competitive performance on multiple public benchmarks. Specifically, our ADStereo runs over faster than the current state-of-the-art CREStereo (0.054s vs. ) under the same hardware while achieving comparable accuracy (1.82% vs. 1.69%) on the KITTI stereo 2015 benchmark. The codes are available at: https://github.com/cocowy1/ADStereo. Yun Wang 0053, Kunhong Li 0001, Longguang Wang, Junjie Hu 0003, Dapeng Oliver Wu, Yulan Guo |
IEEE Trans. Image Process. | 2 |
| 2024 | LoS: Local Structure-Guided Stereo MatchingabstractEstimating disparities in challenging areas is difficult and limits the performance of stereo matching models. In this paper, we exploit local structure information (LSI) to better handle these areas. Specifically, our LSI comprises a series of key elements, including the slant plane (parameterised by disparity gradients), disparity offset details and neighbouring relations. This LSI empowers our method to effectively handle intricate structures, including object boundaries and curved surfaces. We bootstrap the LSI from monocular depth and subsequently refine it to bet-ter capture the underlying scene geometry constraints in an iterative manner. Building upon the LSI, we introduce the Local Structure-Guided Propagation (LSGP), which enhances the disparity initialization, optimization, and refinement processes. By combining LSGP with a Gated Re-current Unit (GRU), we present our novel stereo matching method, referred to as Local Structure-guided stereo matching (LoS). Remarkably, LoS achieves top-ranking results on four widely recognized public benchmark datasets (ETH3D, Middlebury, KITTI 15 & 12) and robust vision challenge, demonstrating the superior capabilities of our model. Kunhong Li 0001, Longguang Wang, Ye Zhang 0037, Shunbo Zhou, Yulan Guo |
CVPR | 1 |
| 2024 | Distractor-Free Novel View Synthesis via Exploiting Memorization Effect in Optimization
Kunhong Li 0001, Minglin Chen, Longguang Wang, Shunbo Zhou, Yulan Guo |
ECCV (54) | 2 |
| 2024 | Learning Representations from Foundation Models for Domain Generalized Stereo Matching
Longguang Wang, Kunhong Li 0001, Yun Wang 0053, Yulan Guo |
ECCV (42) | 3 |
| 2024 | Self-Supervised Multi-Frame Monocular Depth Estimation for Dynamic ScenesabstractSelf-supervised multi-frame depth estimation outperforms single-frame approaches by utilizing not only appearance information, but also geometric information. A common practice for multi-frame methods is to employ feature-metric bundle adjustment (FBA) to refine depth map initialized from the single-frame prior. However, FBA cannot always provide effective residual updates due to unreliable matching costs, which are corrupted by thin texture, occlusion, and especially object motion. To tackle this problem, we propose a context-aware transformer (CAT) to refine the corrupted matching costs by leveraging the spatial context information. Specifically, the CAT adaptively aggregates matching costs according to the spatial affinity inferred from local appearance context, and produces reliable contextual costs for FBA. Moreover, we design a motion-aware regularization loss to provide supervision for regions with moving objects, making CAT competent for dynamic scenes. Extensive experiments and analyses on the KITTI and Cityscapes datasets demonstrate the effectiveness and superior generalization capability of our approach. Guanghui Wu, Hao Liu 0061, Longguang Wang, Kunhong Li 0001, Yulan Guo, Zengping Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Cost Volume Aggregation in Stereo Matching Revisited: A Disparity Classification PerspectiveabstractCost aggregation plays a critical role in existing stereo matching methods. In this paper, we revisit cost aggregation in stereo matching from disparity classification and propose a generic yet efficient Disparity Context Aggregation (DCA) module to improve the performance of CNN-based methods. Our approach is based on an insight that a coarse disparity class prior is beneficial to disparity regression. To obtain such a prior, we first classify pixels in an image into several disparity classes and treat pixels within the same class as homogeneous regions. We then generate homogeneous region representations and incorporate these representations into the cost volume to suppress irrelevant information while enhancing the matching ability for cost aggregation. With the help of homogeneous region representations, efficient and informative cost aggregation can be achieved with only a shallow 3D CNN. Our DCA module is fully-differentiable and well-compatible with different network architectures, which can be seamlessly plugged into existing networks to improve performance with small additional overheads. It is demonstrated that our DCA module can effectively exploit disparity class priors to improve the performance of cost aggregation. Based on our DCA, we design a highly accurate network named DCANet, which achieves state-of-the-art performance on several benchmarks. Yun Wang 0053, Longguang Wang, Kunhong Li 0001, Dapeng Oliver Wu, Yulan Guo |
IEEE Trans. Image Process. | 3 |
| 2022 | Decoupling Makes Weakly Supervised Local Feature BetterabstractWeakly supervised learning can help local feature methods to overcome the obstacle of acquiring a large-scale dataset with densely labeled correspondences. However, since weak supervision cannot distinguish the losses caused by the detection and description steps, directly conducting weakly supervised learning within a joint training describe-then-detect pipeline suffers limited performance. In this paper, we propose a decoupled training describe-then-detect pipeline tailored for weakly supervised local feature learning. Within our pipeline, the detection step is decoupled from the description step and postponed until discriminative and robust descriptors are learned. In addition, we introduce a line-to-window search strategy to explicitly use the camera pose information for better descriptor learning. Extensive experiments show that our method, namely PoSFeat (Camera Pose Supervised Feature), outperforms previous fully and weakly supervised methods and achieves state-of-the-art performance on a wide range of downstream task. Kunhong Li 0001, Longguang Wang, Li Liu 0002, Qing Ran, Kai Xu 0004, Yulan Guo |
CVPR | 1 |
| 2022 | Soft Exemplar Highlighting for Cross-View Image-Based Geo-LocalizationabstractThe goal of ground-to-aerial image geo-localization is to determine the location of a ground query image by matching it against a reference database consisting of aerial/satellite images. This task is highly challenging due to the large appearance difference caused by extreme changes in viewpoint and orientation. In this work, we show that the training difficulty is an important cue that can be leveraged to improve metric learning on cross-view images. More specifically, we propose a new Soft Exemplar Highlighting (SEH) loss to achieve online soft selection of exemplars. Adaptive weights are generated for exemplars by measuring their associated training difficulty using distance rectified logistic regression. These weights are then constrained to remove simple exemplars from training and truncate the large weights of extremely hard exemplars to escape from the trap with a local optimal solution. We further use the proposed SEH loss to train two mainstream convolutional neural networks for ground-to-aerial image-based geo-localization. Experimental results on two benchmark cross-view image datasets demonstrate that the proposed method achieves significant improvements in feature discriminativeness and outperforms the state-of-the-art image-based geo-localization methods. Yulan Guo, Kunhong Li 0001, Farid Boussaïd, Mohammed Bennamoun |
IEEE Trans. Image Process. | 3 |
| 2021 | Distortion-Aware Monocular Depth Estimation for Omnidirectional ImagesabstractImage distortion is a main challenge for tasks on panoramas. In this work, we propose a Distortion-Aware Monocular Omnidirectional (DAMO) network to estimate dense depth maps from indoor panoramas. First, we introduce a distortion-aware module to extract semantic features from omnidirectional images. Specifically, we exploit deformable convolution to adjust its sampling grids to geometric distortions on panoramas. We also utilize a strip pooling module to sample against horizontal distortion introduced by inverse gnomonic projection. Second, we introduce a plug-and-play spherical-aware weight matrix for our loss function to handle the uneven distribution of areas projected from a sphere. Experiments on the 360D dataset show that the proposed method can effectively extract semantic features from distorted panoramas and alleviate the supervision bias caused by distortion. It achieves the state-of-the-art performance on the 360D dataset with high efficiency. Hong-Xiang Chen, Kunhong Li 0001, Zhiheng Fu, Zonghao Chen, Yulan Guo |
IEEE Signal Process. Lett. | 2 |
| 2021 | Adv-Depth: Self-Supervised Monocular Depth Estimation With an Adversarial LossabstractLoss function plays a key role in self-supervised monocular depth estimation methods. Current reprojection loss functions are hand-designed and mainly focus on local patch similarity but overlook the global distribution differences between a synthetic image and a target image. In this paper, we leverage global distribution differences by introducing an adversarial loss into the training stage of self-supervised depth estimation. Specifically, we formulate this task as a novel view synthesis problem. We use a depth estimation module and a pose estimation module to form a generator, and then design a discriminator to learn the global distribution differences between real and synthetic images. With the learned global distribution differences, the adversarial loss can be back-propagated to the depth estimation module to improve its performance. Experiments on the KITTI dataset have demonstrated the effectiveness of the adversarial loss. The adversarial loss is further combined with the reprojection loss to achieve the state-of-the-art performance on the KITTI dataset. Kunhong Li 0001, Zhiheng Fu, Hanyun Wang, Zonghao Chen, Yulan Guo |
IEEE Signal Process. Lett. | 1 |