VLDB 2026 Research / reviewers in the wild / expert
Hong Zhang 0011
dblp:24/6914-11
· DBLP profile ↗
12ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0001-9790-0555ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Motion-guided semantic alignment for line art animation colorization
Ning Wang 0025, Hairui Yang, Hong Zhang 0011, Zhiyong Wang 0001, Zhihui Wang 0001 |
Pattern Recognit. | 4 |
| 2025 | STRobustNet: Efficient Change Detection via Spatial-Temporal Robust Representations in Remote SensingabstractVision transformers have achieved impressive performance in addressing spatial–temporal inconsistencies of change detection (CD) due to their ability to model long-range dependencies in bitemporal features. However, applying the self-attention operation directly to bitemporal features causes feature confusion and incurs high computational complexity. In this article, we propose a novel CD framework based on spatial–temporal robust representation (STRobustNet). To avoid directly applying self-attention to bitemporal features, we derive a collection of STRobustNets for land-use classes of interest. These representations are used to transform spatial–temporally inconsistent bitemporal features into consistent bitemporal classification representations, enhancing model efficiency and reducing feature confusion. Specifically, to fully perceive and integrate the varied appearances of the same class while minimizing deviation from the current input samples, we design a robust representation generation module (RRGModule), which utilizes the universal spatial–temporal context provided by the entire dataset and the specific spatial–temporal context from the current bitemporal images to generate robust representations for the interested land-use classes, improving robustness to spatial–temporal inconsistencies. Then, these representations are used to activate bitemporal features in different class channels, producing spatial–temporally consistent classification representations. The CD is ultimately performed using these consistent representations, effectively avoiding false detections caused by spatial–temporal inconsistencies. Experimental results demonstrate that STRobustNet achieves performance comparable to top-performing methods and offers the fastest inference speed among transformer-based methods. Code and pretrained models are available athttps://github.com/DLUTTengYH/STRobustNet. Hong Zhang 0011, Yuhang Teng, Zhihui Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Propagating Sparse Depth via Depth Foundation Model for Out-of-Distribution Depth CompletionabstractDepth completion is a pivotal challenge in computer vision, aiming at reconstructing the dense depth map from a sparse one, typically with a paired RGB image. Existing learning-based models rely on carefully prepared but limited data, leading to significant performance degradation in out-of-distribution (OOD) scenarios. Recent foundation models have demonstrated exceptional robustness in monocular depth estimation through large-scale training, and using such models to enhance the robustness of depth completion models is a promising solution. In this work, we propose a novel depth completion framework that leverages depth foundation models to attain remarkable robustness without large-scale training. Specifically, we leverage a depth foundation model to extract environmental cues, including structural and semantic context, from RGB images to guide the propagation of sparse depth information into missing regions. We further design a dual-space propagation approach, without any learnable parameters, to effectively propagate sparse depth in both 3D and 2D spaces to maintain geometric structure and local consistency. To refine the intricate structure, we introduce a learnable correction module to progressively adjust the depth prediction towards the real depth. We train our model on the NYUv2 and KITTI datasets as in-distribution datasets and extensively evaluate the framework on 16 other datasets. Our framework performs remarkably well in the OOD scenarios and outperforms existing state-of-the-art depth completion methods. Our models are released in https://github.com/shenglunch/PSD. Shenglun Chen, Xinzhu Ma, Hong Zhang 0011, Zhihui Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Real-Time Depth Completion With Multimodal Feature AlignmentabstractAs a key problem in computer vision, depth completion aims to recover dense depth maps from sparse ones [generally derived from light detection and ranging (LiDAR)]. Most methods introduce synchronous RGB images and leverage multimodal fusion to integrate multimodal features from these modalities to describe the complete scene. However, their different natural characteristics lead to inconsistency in features, potentially impacting the effectiveness of multimodal feature fusion. To address this issue, we propose a feature alignment network (FANet) that introduces an alignment scheme to enhance the consistency between multimodal features. This scheme aligns the modality-invariant semantic context, which is invariant to changes in modality and represents the correlation between a pixel and its surroundings. Specifically, we first design an asymmetric context extraction (ACE) module to extract modality-invariant semantic contexts from multimodal features within limited GPU memory, and then pull them closer to improve consistency. Crucially, our alignment scheme is only applied during the training phase, and no additional computation cost is incurred in the inference phase. Moreover, we introduce a simple yet effective refinement module to refine estimated results via residual learning based on intermediate depth maps and sparse depth maps. Extensive experiments on KITTI and VOID datasets demonstrate that our method achieves competitive performance against typical real-time methods. In addition, we embed the proposed alignment scheme and refinement module into other methods to demonstrate their effectiveness. Shenglun Chen, Xinzhu Ma, Hong Zhang 0011, Baoli Sun, Zhihui Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | CustomDepth: Customizing point-wise depth categories for depth completion
Shenglun Chen, Xinchen Ye, Hong Zhang 0011, Zhihui Wang 0001 |
Pattern Recognit. Lett. | 3 |
| 2024 | Learning Pixel-Wise Continuous Depth Representation via Clustering for Depth CompletionabstractDepth completion is a long-standing challenge in computer vision, where classification-based methods have made tremendous progress in recent years. However, most existing classification-based methods rely on pre-defined pixel-shared and discrete depth values as depth categories. This representation fails to capture the continuous depth values that conform to the real depth distribution, leading to depth smearing in boundary regions. To address this issue, we revisit depth completion from the clustering perspective and propose a novel clustering-based framework called CluDe which focuses on learning the pixel-wise and continuous depth representation. The key idea of CluDe is to iteratively update the pixel-shared and discrete depth representation to its corresponding pixel-wise and continuous counterpart, driven by the real depth distribution. Specifically, CluDe first utilizes depth value clustering to learn a set of depth centers as the depth representation. While these depth centers are pixel-shared and discrete, they are more in line with the real depth distribution compared to pre-defined depth categories. Then, CluDe estimates offsets for these depth centers, enabling their dynamic adjustment along the depth axis of the depth distribution to generate the pixel-wise and continuous depth representation. Extensive experiments demonstrate that CluDe successfully reduces depth smearing around object boundaries by utilizing pixel-wise and continuous depth representation. Furthermore, CluDe achieves state-of-the-art performance on the VOID datasets and outperforms classification-based methods on the KITTI dataset. Shenglun Chen, Hong Zhang 0011, Xinzhu Ma, Zhihui Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Denser is Better:cost distribution super-resolution network for more accurate sub-pixel disparityabstractThe low-quality cost distribution obtained by simple upsampling leads to disparity maps with many outliers and low sub-pixel accuracy. We propose the Cost Distribution Super-Resolution Network (CDSRNet), which directly extracts high-resolution cost distribution from the low-resolution 4D cost volume.The similarity extraction module of CDSRNet decomposes the task of estimating high-resolution cost distribution into multiple subtasks and completes each subtask by a specific translation block, ensuring high discrimination of the predicted cost distribution. The inter-block aggregation module aggregates information from other subtasks, obtaining a voting volume with global information for correcting errors in the current subtask. Experiment results demonstrate that the proposed method significantly reduces the ratio of outliers and improves the sub-pixel accuracy. Hong Zhang 0011, Shenglun Chen, Zhihui Wang 0001, Wanli Ouyang |
ICME | 1 |
| 2023 | Feature enhancement network for stereo matching
Shenglun Chen, Hong Zhang 0011, Baoli Sun, Xinchen Ye, Zhihui Wang 0001 |
Image Vis. Comput. | 2 |
| 2022 | MonoDistill: Learning Spatial Features for Monocular 3D Object Detection
Zhiyu Chong, Xinzhu Ma, Hong Zhang 0011, Yuxin Yue, Zhihui Wang 0001, Wanli Ouyang |
ICLR | 3 |
| 2022 | The Farther the Better: Balanced Stereo Matching via Depth-Based Sampling and Adaptive Feature RefinementabstractExisting stereo matching methods achieve satisfactory average accuracy on a whole predicted disparity map under common global metrics, but ignore the fine-grained performance at region level, especially for far regions in the scene, which is more crucial in actual auto-driving scenarios. There are two factors accounting for this problem: 1) Depth resolution. Existing methods use disparity-based sampling to extract matching candidates uniformly according to the disparity range, but leads to sparser sampling density at far regions than that of close regions in terms of depth range, resulting in low depth resolution in far regions. 2) Feature discriminability. Limited image resolution and inferior feature extraction at far regions result in the obtained features with low discrimination, which influences the subsequent matching process between stereo images. To improve the estimation accuracy of far regions and thus achieve a balanced performance at region level, we design a novel two-stage Balanced Stereo Matching Network (BSMNet) to address the above problems. The coarse stage of BSMNet introduces a direct depth-based sampling strategy, which generates matching candidates according to scene depth instead of disparity, thus improving the depth resolution and obtaining initial depth map with more balanced accuracy. Then, a depth refinement stage is proposed to solve the problem of low feature discriminability and further optimizes the initial depth map obtained from the coarse stage. It selects matching candidates and computes their similarity scores from a carefully designed adaptive feature volume guided by a learnable scale map, thus making the final estimation more accurate. Different from the existing methods that construct stereo matching based on disparity prediction, our proposed pipeline is to directly optimize on depth information. Experiments show that our BSMNet can obtain an obvious performance improvement at far regions without discarding that at close regions, so as to largely outperform existing state-of-the-art methods. Hong Zhang 0011, Xinchen Ye, Shenglun Chen, Zhihui Wang 0001, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Cost Affinity Learning Network for Stereo MatchingabstractExisting stereo matching methods mainly tend to directly aggregate features output from Convolutional Neural Network to obtain more discriminative cost features, but ignore the affinity of each element in the cost feature which also plays a key role in enhancing the cost feature. In this work, we propose a novel cost affinity learning network(CAL-Net) whose Affinity Enhanced Module(AEM) extracts the affinity of the elements in the cost feature and reconstructs a more discriminative feature. In addition, CAL-Net designs a Disparity Weight Loss(DWL) to guide training. Specifically, AEM takes the advatange of the self-attention mechanism to learn internal affinity between different elements and exploits it to reconstruct the cost feature for emphasizing informative elements. DWL calculates the adaptive weight according to disparity error. As the error decreases, the weight gradually increases and enables the network to gradually transit from the pixel level disparity to sub-pixel level. Experiments demonstrate that CAL-Net boosts the performance, especially in textureless and reflective regions, and achieves better results on Scene Flow and KITTI 2012 benchmarks than some typical related methods. Shenglun Chen, Baopu Li, Wei Wang 0335, Hong Zhang 0011, Zhihui Wang 0001 |
ICASSP | 4 |
| 2020 | Geometry and context guided refinement for stereo matchingabstractThe disparity refinement phase of existing end‐to‐end stereo matching networks refines the disparity by learning the mapping from the concatenated coarse disparity and corresponding features to fine disparity. It depends on the scenarios' characteristics, such as the distribution of disparity and semantic categories contained in the domain, which makes the network fail to work on unseen domain. In this paper, we propose a geometry and context guided refinement network (GCGR‐Net) containing a Fine Matching module and an Upsampling module. GCGR‐Net learns to utilize pixels' relationship to get high resolution dense disparity, which is independent of the data's content. The Fine Matching module performs a minimum range search based on the relationship between the possible matching pixel pairs, i.e. the called geometry information, to recover the internal structure of the object. The Upsampling module obtains context information, the relationship between central pixel and the pixels in its neighborhood, to upsample the lower resolution disparity. The final disparity map is obtained step by step through an iterative refinement model. Experiment results show that our method not only has good performance in the training scenarios, but also outperforms previous methods on the unseen domain without fine‐tuning. Hong Zhang 0011, Zhihui Wang 0001, Yuxin Yue, Shenglun Chen |
IET Image Process. | 1 |