Zehua Sheng

dblp:299/8234 · also Ze-Hua Sheng · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2025
0000-0002-1721-9143ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2025 Learning Arbitrary-Scale RAW Image Downscaling with Wavelet-based Recurrent Reconstruction
abstract
Image downscaling is critical for efficient storage and transmission of high-resolution (HR) images. Existing learning-based methods focus on performing downscaling within the sRGB domain, which typically suffers from blurred details and unexpected artifacts. RAW images, with their unprocessed photonic information, offer greater flexibility but lack specialized downscaling frameworks. In this paper, we propose a wavelet-based recurrent reconstruction framework that leverages the information lossless attribute of wavelet transformation to fulfill the arbitrary-scale RAW image downscaling in a coarse-to-fine manner, in which the Low-Frequency Arbitrary-Scale Downscaling Module (LASDM) and the High-Frequency Prediction Module (HFPM) are proposed to preserve structural and textural integrity of the reconstructed low-resolution (LR) RAW images, alongside an energy-maximization loss to align high-frequency energy between HR and LR domain. Furthermore, we introduce the Realistic Non-Integer RAW Downscaling (Real-NIRD) dataset, featuring a non-integer downscaling factor of 1.3×, and incorporate it with publicly available datasets with integer factors (2×, 3×, 4×) for comprehensive benchmarking arbitrary-scale image downscaling purposes. Extensive experiments demonstrate that our method outperforms existing state-of-the-art competitors both quantitatively and visually. The code and dataset will be released at https://github.com/RenYangSCU/ASRD.
Yang Ren 0001, Hai Jiang 0006, Wei Li 0075, Menglong Yang, Heng Zhang 0042, Zehua Sheng, Qingsheng Ye, Shuaicheng Liu
ACM Multimedia6
2025 STARNet: Low-light video enhancement using spatio-temporal consistency aggregation
Zehua Sheng, Si-Yuan Cao, Runmin Zhang, Beinan Yu, Chenghao Zhang 0002, Bailin Yang
Pattern Recognit.2
2025 TFDet: Target-Aware Fusion for RGB-T Pedestrian Detection
abstract
Pedestrian detection plays a critical role in computer vision as it contributes to ensuring traffic safety. Existing methods that rely solely on RGB images suffer from performance degradation under low-light conditions due to the lack of useful information. To address this issue, recent multispectral detection approaches have combined thermal images to provide complementary information and have obtained enhanced performances. Nevertheless, few approaches focus on the negative effects of false positives (FPs) caused by noisy fused feature maps. Different from them, we comprehensively analyze the impacts of FPs on detection performance and find that enhancing feature contrast can significantly reduce these FPs. In this article, we propose a novel target-aware fusion strategy for multispectral pedestrian detection, named TFDet. The target-aware fusion strategy employs a fusion-refinement paradigm. In the fusion phase, we reveal the parallel- and cross-channel similarities in RGB and thermal features and learn an adaptive receptive field to collect useful information from both features. In the refinement phase, we use a segmentation branch to discriminate the pedestrian features from the background features. We propose a correlation-maximum loss function to enhance the contrast between the pedestrian features and background features. As a result, our fusion strategy highlights pedestrian-related features and suppresses unrelated ones, generating more discriminative fused features. TFDet achieves state-of-the-art performance on two multispectral pedestrian benchmarks, KAIST and LLVIP, with absolute gains of 0.65% and 4.1% over the previous best approaches, respectively. TFDet can easily extend to multiclass object detection scenarios. It outperforms the previous best approaches on two multispectral object detection benchmarks, FLIR and M3FD, with absolute gains of 2.2% and 1.9%, respectively. Importantly, TFDet has comparable inference efficiency to the previous approaches and has remarkably good detection performance even under low-light conditions, which is a significant advancement for ensuring road safety. The code will be made publicly available at https://github.com/XueZ-phd/TFDet.git.
Jiacheng Ying, Zehua Sheng, Heng Yu 0001, Chunguang Li 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 Recurrent Homography Estimation Using Homography-Guided Image Warping and Focus Transformer
abstract
We propose the Recurrent homography estimation framework using Homography-guided image Warping and Focus transformer (FocusFormer), named RHWF. Both being appropriately absorbed into the recurrent framework, the homography-guided image warping progressively enhances the feature consistency and the attention-focusing mechanism in FocusFormer aggregates the intra-inter correspondence in a global→nonlocal→local manner. Thanks to the above strategies, RHWF ranks top in accuracy on a variety of datasets, including the challenging cross-resolution and cross-modal ones. Meanwhile, benefiting from the recurrent framework, RHWF achieves parameter efficiency despite the transformer architecture. Compared to previous state-of-the-art approaches LocalTrans and IHN, RHWF reduces the mean average corner error (MACE) by about 70% and 38.1% on the MSCOCO dataset, while saving the parameter costs by 86.5% and 24.6%. Similar to the previous works, RHWF can also be arranged in 1-scale for efficiency and 2-scale for accuracy, with the 1-scale RHWF already outperforming most of the previous methods. Source code is available at https://github.com/imdump178/RHWF.
Si-Yuan Cao, Runmin Zhang, Lun Luo, Beinan Yu, Zehua Sheng, Junwei Li 0009
CVPR5
2023 Structure Aggregation for Cross-Spectral Stereo Image Guided Denoising
abstract
To obtain clean images with salient structures from noisy observations, a growing trend in current denoising studies is to seek the help of additional guidance images with high signal-to-noise ratios, which are often acquired in different spectral bands such as near infrared. Although previous guided denoising methods basically require the input images to be well-aligned, a more common way to capture the paired noisy target and guidance images is to exploit a stereo camera system. However, current studies on cross-spectral stereo matching cannot fully guarantee the pixel-level registration accuracy, and rarely consider the case of noise contamination. In this work, for the first time, we propose a guided denoising framework for cross-spectral stereo images. Instead of aligning the input images via conventional stereo matching, we aggregate structures from the guidance image to estimate a clean structure map for the noisy target image, which is then used to regress the final denoising result with a spatially variant linear representation model. Based on this, we design a neural network, called as SANet, to complete the entire guided denoising process. Experimental results show that, our SANet can effectively transfer structures from an unaligned guidance image to the restoration result, and outperforms state-of-the-art denoisers on various stereo image datasets. Besides, our structure aggregation strategy also shows its potential to handle other unaligned guided restoration tasks such as super-resolution and deblurring. The source code is available at https://github.com/lustrouselixir/SANet.
Zehua Sheng, Zhu Yu 0001, Xiongwei Liu, Si-Yuan Cao, Yuqi Liu 0005, Huaqi Zhang
CVPR1
2023 Aggregating Feature Point Cloud for Depth Completion
abstract
Guided depth completion aims to recover dense depth maps by propagating depth information from the given pixels to the remaining ones under the guidance of RGB images. However, most of the existing methods achieve this using a large number of iterative refinements or stacking repetitive blocks. Due to the limited receptive field of conventional convolution, the generalizability with respect to different sparsity levels of input depth maps is impeded. To tackle these problems, we propose a feature point cloud aggregation framework to directly propagate 3D depth information between the given points and the missing ones. We extract 2D feature map from images and transform the sparse depth map to point cloud to extract sparse 3D features. By regarding the extracted features as two sets of feature point clouds, the depth information for a target location can be reconstructed by aggregating adjacent sparse 3D features from the known points using cross attention. Based on this, we design a neural network, called as PointDC, to complete the entire depth information reconstruction process. Experimental results show that, our PointDC achieves superior or competitive results on the KITTI benchmark and NYUv2 dataset. In addition, the proposed PointDC demonstrates its higher generalizability to different sparsity levels of the input depth maps and cross-dataset evaluation.
Zhu Yu 0001, Zehua Sheng, Lun Luo, Si-Yuan Cao, Huaqi Zhang
ICCV2
2023 Region-aware RGB and near-infrared image fusion
Jiacheng Ying, Can Tong, Zehua Sheng, Bo-Wen Yao, Si-Yuan Cao, Heng Yu 0001
Pattern Recognit.3
2023 Guided Colorization Using Mono-Color Image Pairs
abstract
Compared to color images captured by conventional RGB cameras, monochrome (mono) images usually have higher signal-to-noise ratios (SNR) and richer textures due to the lack of color filter arrays in mono cameras. Therefore, using a mono-color stereo dual-camera system, we can integrate the lightness information of target monochrome images with the color information of guidance RGB images to accomplish image enhancement in a colorization manner. In this work, based on two assumptions, we introduce a novel probabilistic-concept guided colorization framework. First, adjacent contents with similar luminance are likely to have similar colors. By lightness matching, we can utilize colors of the matched pixels to estimate the target color value. Second, by matching multiple pixels from the guidance image, if more of these matched pixels have similar luminance values to the target one, we can estimate colors with more confidence. Based on the statistical distribution of multiple matching results, we retain the reliable color estimates as initial dense scribbles and then propagate them to the rest of the mono image. However, for a target pixel, the color information provided by its matching results is quite redundant. Hence, we introduce a patch sampling strategy to accelerate the colorization process. Based on the analysis of the posteriori probability distribution of the sampling results, we can use much fewer matches for color estimation and reliability assessment. To alleviate incorrect color propagation in the sparsely scribbled regions, we generate extra color seeds according to the existed scribbles to guide the propagation process. Experimental results show that, our algorithm can efficiently and effectively restore color images with higher SNR and richer details from the mono-color image pairs, and achieves good performance in solving the color bleeding problem.
Zehua Sheng, Bo-Wen Yao, Huaqi Zhang
IEEE Trans. Image Process.1
2023 Frequency-Domain Deep Guided Image Denoising
abstract
Despite the tremendous advances in denoising techniques, it's still challenging to restore a clean image with salient structures based on one noisy observation, especially at high noise levels. In this work, we propose a frequency-domain guided denoising algorithm to conduct denoising with the help of a well-aligned guidance image. Thanks to their structural correlations, the frequency characteristics of the guidance image can indicate whether the frequency coefficients of the noisy target image are contributed by noise or textures. Therefore, the explicit frequency decomposition enables our denoising model to avoid over-smoothing detailed contents. However, as two input images are usually captured in different fields, their structures are not always consistent. Therefore, we model guided denoising with an optimization problem which considers both the representation model of the guidance image and the fidelity to the noisy target. Further, we design a convolutional neural network, called as FGDNet, to explore the optimal solution. Due to the visual masking phenomenon, human eyes are sensitive to noise in the flat areas, but may not perceive noise around edges or textures. Therefore, we expect to remove as much noise as possible to guarantee the spatial smoothness of flat contents, while also preserving high-frequency structures. Through frequency decomposition, our model can process the low-frequency and high-frequency contents separately. We also adopt a frequency-relevant loss function to train the network. Experimental results show that, compared with state-of-the-art guided and non-guided denoisers, our FGDNet achieves higher denoising accuracy and better visual quality in both flat and texture-rich regions.
Zehua Sheng, Xiongwei Liu, Si-Yuan Cao, Huaqi Zhang
IEEE Trans. Multim.1
2022 Iterative Deep Homography Estimation
abstract
We propose Iterative Homography Network, namely IHN, a new deep homography estimation architecture. Different from previous works that achieve iterative refinement by network cascading or untrainable IC-LK iterator; the iterator of IHN has tied weights and is completely trainable. IHN achieves state-of-the-art accuracy on several datasets including challenging scenes. We propose 2 versions of IHN: (1) IHN for static scenes, (2) IHN-mov for dynamic scenes with moving objects. Both versions can be arranged in 1-scale for efficiency or 2-scale for accuracy. We show that the basic 1-scale IHN already outperforms most of the existing methods. On a variety of datasets, the 2-scale IHN outperforms all competitors by a large gap. We introduce IHN-mov by producing an inlier mask to further improve the estimation accuracy of moving-objects scenes. We experimentally show that the iterative framework of IHN can achieve 95% error reduction while considerably saving network parameters. When processing sequential image pairs, IHN can achieve 32.7 fps, which is about 8× the speed of IC-LK iterator: Source code is available at https://github.com/imdump178/IHN.
Si-Yuan Cao, Jianxin Hu, Zehua Sheng
CVPR3
2022 Frequency-Relevant Residual Learning for Multi-Modal Image Denoising
abstract
Recently, multi-modal image processing has shown its great potential in boosting denoising performance in terms of both accuracy and visual quality. However, current studies generally face the challenge of achieving a good balance between noise removal and detail preservation. In this work, we introduce patch-wise frequency decomposition into convolutional neural networks and propose a novel multi-modal image denoising (MID) algorithm. Integrating the frequency-domain information of images from different modalities, our network first predicts a frequency-relevant residual and then regresses the denoised result using a learnable reconstruction kernel. Benefiting from the distinctive properties of noise and true signals as well as the correlation between multi-modal images in the frequency domain, the proposed algorithm can effectively remove noise and simultaneously reconstruct fine details. Extensive experiments demonstrate the superiority and generalizability of our algorithm over state-of-the-art competing algorithms on various MID tasks, including near-infrared guided RGB image denoising, flash guided no-flash image denoising, and RGB guided depth image denoising. Code is available at https://github.com/liuxw11/FRL.
Xiongwei Liu, Zehua Sheng
ICIP2
2022 FocusNet: Classifying better by focusing on confusing classes
Zehua Sheng
Pattern Recognit.2