EDBT 2026 Demo / reviewers in the wild / expert
Heng Zhou 0006
dblp:19/225-6
· DBLP profile ↗
15ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0003-2770-2785ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Self-Image and Cross-Image Consistency Learning for Remote Sensing Burned Area SegmentationabstractThe increasing frequency of global wildfires has led to the destruction of vast forests and wetlands. Non-contact remote sensing technologies provide an effective means for accurate burned area segmentation (BAS). However, existing BAS methods often treat each image independently, focusing primarily on local pixel contexts while neglecting the broader semantic consistency of burned regions across different scenes. The lack of global context modeling limits their robustness, as burned areas typically exhibit distinctive and consistent visual characteristics such as color and texture across diverse environments. To address this limitation, we propose a Self-image and Cross-image Consistency Learning (SCCL) framework, which captures both local pixel-level relationships within a single image and global semantic dependencies across multiple images. By enforcing consistent and compact representations of burned regions within and across images, SCCL enhances segmentation robustness under varying weather and terrain conditions. Additionally, to refine boundary delineation between burned and unburned areas, we introduce a Burned Edge Injector (BEI) and an Edge-Injected Decoder (EID). We further construct two large-scale BAS benchmark datasets, BAS-AUS and BAS-EUR, for comprehensive evaluation. Experiments on these benchmarks demonstrate that our method achieves state-of-the-art performance, significantly outperforming previous approaches, with MAE reduced to 0.017 and 0.016, respectively. The new BAS benchmarks and code are available at https://github.com/VisionVerse/SCCL. Heng Zhou 0006, Chengyang Li 0001, Chunna Tian, Yongqiang Xie, Zhongbo Li, Xiaojun Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Multi-Scale Frequency Enhancement Network for Blind Image Deblurring
Yawen Xiang, Heng Zhou 0006, Chengyang Li 0001, Zhongbo Li, Yongqiang Xie |
IET Image Process. | 2 |
| 2025 | Lightweight Spatial-Channel-Frequency Network for RGB-Thermal Salient Object DetectionabstractRGB-Thermal (RGB-T) salient object detection aims to accurately locate salient regions by integrating complementary information from visible and thermal modalities. However, existing methods often struggle to fully leverage the spatial structures and semantic differences in the frequency domain across modalities, while also incurring high computational costs. In this paper, we propose a lightweight spatial-channel-frequency information mining framework (SCF-Net) for RGB-T salient object detection. Specifically, we first introduce a local-global complementary aggregation (LGCA) module that effectively fuses local textures, global semantics, and modality-aware features to enhance cross-layer information consistency and semantic alignment. Furthermore, we design a triple cue mining (TCM) module to jointly explore spatial, channel, and frequency-domain salient cues, enabling the extraction of both intra-modal details and cross-modal complementary information. Our SCF-Net contains only 8.39 M parameters and supports real-time inference. Extensive experiments conducted on VT821, VT1000, and VT5000 three datasets demonstrate that our method consistently outperforms existing mainstream models while maintaining high operational efficiency, showing strong effectiveness and robustness across a variety of challenging scenarios. Heng Zhou 0006, Wanting Hong, Xiaoxiong Liu, Xiaojun Wu 0001 |
IEEE Signal Process. Lett. | 1 |
| 2025 | Multi-Weather Restoration: An Efficient Prompt-Guided Convolution ArchitectureabstractAddressing degraded weather conditions plays a vital role in practical applications. Many existing restoration approaches are limited to specific weather types, which limits their applicability to different weather scenarios. Advanced technologies, encompassing Transformer and diffusion model, have been harnessed to confront this challenge. However, these methods often heighten network complexity and prolong inference duration. To this end, we present MW-ConvNet, a U-shaped convolution-based network for multi-weather restoration. Specifically, the MW-Enc block and MW-Dec block are introduced to achieve simple yet strong feature extraction, which rely entirely on traditional 2D convolution. To improve adaptability to multiple weather conditions, a prompt generation module is designed to generate a representative weather prompt at the encoder’s terminus. Drawing inspiration from style transfer, the weather prompt is used to guide the decoder learning through a progressive restoration procedure. For future high-fidelity restoration, we introduce frequency separation through wavelet pooling blocks in encoder phase and corresponding up-sampling blocks in decoder phase. The segregated treatment of low-frequency and high-frequency features curbs the loss of textural information during network computation. It also future improves the quality and accuracy of generated weather prompt. Extensive experiments demonstrate that the proposed MW-ConvNet obtains superior performance compared to state-of-the-art methods across both weather-specific and real-world restoration tasks. Significantly, our method achieves an impressive inference speed of 0.12 seconds per$256\times 256$image, outpacing transformer-based and diffusion-based models. Chengyang Li 0001, Fangwei Sun, Heng Zhou 0006, Yongqiang Xie, Zhongbo Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Deformation-Resilient Multigranularity Learning for Unaligned RGB-T Semantic SegmentationabstractRGB-Thermal semantic segmentation (SS) aims to combine visual light and thermal images to determine the semantic category for each pixel and create an object mask. While existing methods typically rely on well-aligned RGB-T image pairs, real-world RGB-T pairs are often unaligned, and pixel-by-pixel alignment is both challenging and time-consuming. To address this critical issue, we introduce a new unaligned RGB-T SS benchmark and propose the deformation-resilient multigranularity learning (DML) method. DML explores the spatial consistency and modal complementarity of RGB-T and mitigates the interference of warped modalities by aligning multimodal features in a coarse-to-fine multigranularity strategy. Specifically, DML constructs a deformation-aware complementary feature enhancer (DCFE), which consists of deformation-aware feature alignment (DFA) and complementary feature aggregation (CFA) modules. DFA enhances the spatial alignment of RGB-T by estimating the deformation field of warped features. Then, CFA aggregates complementary contexts of modal differences across multiple scales to produce deformation-resilient and robust RGB-T feature representations. Finally, we design the multigranularity mask refinement engine (MMFE), which combines class-agnostic saliency prediction (CSP) and class-aware edge generation (CEG) auxiliary tasks to provide useful boundary and positional cues for SS decoders. The MMFE enhances semantic alignment and interclass separability, yielding object masks with sharp boundaries. Quantitative and qualitative experiments on aligned and unaligned datasets validate the effectiveness of our proposed DML, consistently outperforming existing methods designed for aligned RGB-T data. The new unaligned RGB-T SS benchmark and code are available at https://github.com/VisionVerse/Unaligned-RGBT-Semantic-Segmentation. Heng Zhou 0006, Chengyang Li 0001, Chunna Tian, Yongqiang Xie, Zhongbo Li, Xiaojun Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Deep learning in motion deblurring: current status, benchmarks and future prospects
Yawen Xiang, Heng Zhou 0006, Chengyang Li 0001, Fangwei Sun, Zhongbo Li, Yongqiang Xie |
Vis. Comput. | 2 |
| 2024 | Frequency-aware feature aggregation network with dual-task consistency for RGB-T salient object detection
Heng Zhou 0006, Chunna Tian, Chengyang Li 0001, Yongqiang Xie, Zhongbo Li |
Pattern Recognit. | 1 |
| 2024 | An Implicit-Explicit Prototypical Alignment Framework for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning methods have been explored to mitigate the scarcity of pixel-level annotation in medical image segmentation tasks. Consistency learning, serving as a mainstream method in semi-supervised training, suffers from low efficiency and poor stability due to inaccurate supervision and insufficient feature representation. Prototypical learning is one potential and plausible way to handle this problem due to the nature of feature aggregation in prototype calculation. However, the previous works have not fully studied how to enhance the supervision quality and feature representation using prototypical learning under the semi-supervised condition. To address this issue, we propose an implicit-explicit alignment (IEPAlign) framework to foster semi-supervised consistency training. In specific, we develop an implicit prototype alignment method based on dynamic multiple prototypes on-the-fly. And then, we design a multiple prediction voting strategy for reliable unlabeled mask generation and prototype calculation to improve the supervision quality. Afterward, to boost the intra-class consistency and inter-class separability of pixel-wise features in semi-supervised segmentation, we construct a region-aware hierarchical prototype alignment, which transmits information from labeled to unlabeled and from certain regions to uncertain regions. We evaluate IEPAlign on three medical image segmentation tasks. The extensive experimental results demonstrate that the proposed method outperforms other popular semi-supervised segmentation methods and achieves comparable performance with fully-supervised training methods. Chunna Tian, Xinbo Gao 0001, Heng Zhou 0006, Zhicheng Jiao |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | M2FNet: Mask-Guided Multi-Level Fusion for RGB-T Pedestrian DetectionabstractRGB-Thermal pedestrian detection has shown many notable advantages in various lighting and weather conditions by combining the information from RGB-T images. Due to distinct imaging principles, RGB-T modalities consist of modality-specific and modality-consistent information. However, most existing RGB-T pedestrian detection methods indiscriminately integrate these two types of information, which leads to the pollution of modality information. To address this issue, we propose a novel mask-guided multi-level fusion network (M2FNet) for RGB-T pedestrian detection. M2FNet independently explores consistent and specific features in RGB-T modalities at three different levels, utilizing pixel-level positional information in masks to exclusively focus on pedestrian-related features. Specifically, at the feature extraction level, we selectively embed cross-modality differential compensation (CDC) modules and design the bidirectional multiscale fusion (BMF) module to fully utilize the complementary modality-specific information and enhance the precision of predicted pedestrian masks. At the feature fusion level, the mask-guided global consistency mining (MGCM) module is introduced to capture intra-modal and inter-modal consistent information of pedestrians, which generates highly discriminative RGB-T features. Finally, to further reduce inter-modal differences, we propose a mask-guided pixel-level decision fusion (MPDF) strategy to dynamically weight the RGB-T predictions. Extensive experiments and comparisons demonstrate that our proposed M2FNet, with different backbones, outperforms the state-of-the-art detectors on both publicly available KAIST and CVC-14 RGB-T pedestrian detection datasets. Shiguo Chen, Chunna Tian, Heng Zhou 0006 |
IEEE Trans. Multim. | 4 |
| 2023 | Self-aware and Cross-Sample Prototypical Learning for Semi-supervised Medical Image Segmentation
Chunna Tian, Heng Zhou 0006, Xin Li 0079, Fan Yang 0054, Zhicheng Jiao |
MICCAI (2) | 4 |
| 2023 | Detection-Friendly Dehazing: Object Detection in Real-World Hazy ScenesabstractAdverse weather conditions in real-world scenarios lead to performance degradation of deep learning-based detection models. A well-known method is to use image restoration methods to enhance degraded images before object detection. However, how to build a positive correlation between these two tasks is still technically challenging. The restoration labels are also unavailable in practice. To this end, taking the hazy scene as an example, we propose a union architecture BAD-Net that connects the dehazing module and detection module in an end-to-end manner. Specifically, we design a two-branch structure with an attention fusion module for fully combining hazy and dehazing features. This reduces bad impacts on the detection module when the dehazing module performs poorly. Besides, we introduce a self-supervised haze robust loss that enables the detection module to deal with different degrees of haze. Most importantly, an interval iterative data refinement training strategy is proposed to guide the dehazing module learning with weak supervision. BAD-Net improves further detection performance through detection-friendly dehazing. Extensive experiments on RTTS and VOChaze datasets show that BAD-Net achieves higher accuracy compared to the recent state-of-the-art methods. It is a robust detection framework for bridging the gap between low-level dehazing and high-level detection. Chengyang Li 0001, Heng Zhou 0006, Caidong Yang, Yongqiang Xie, Zhongbo Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Rethinking referring relationships from a perspective of mask-level relational reasoning
Chengyang Li 0001, Gangyi Tian, Heng Zhou 0006 |
Pattern Recognit. | 5 |
| 2023 | Model-driven self-aware self-training framework for label noise-tolerant medical image segmentation
Chunna Tian, Xinbo Gao 0001, Yanyu Ye, Heng Zhou 0006, Zhuo Tong |
Signal Process. | 6 |
| 2023 | Position-Aware Relation Learning for RGB-Thermal Salient Object DetectionabstractSalient object detection (SOD) is an important task in computer vision that aims to identify visually conspicuous regions in images. RGB-Thermal SOD combines two spectra to achieve better segmentation results. However, most existing methods for RGB-T SOD use boundary maps to learn sharp boundaries, which lead to sub-optimal performance as they ignore the interactions between isolated boundary pixels and other confident pixels. To address this issue, we propose a novel position-aware relation learning network (PRLNet) for RGB-T SOD. PRLNet explores the distance and direction relationships between pixels by designing an auxiliary task and optimizing the feature structure to strengthen intra-class compactness and inter-class separation. Our method consists of two main components: A signed distance map auxiliary module (SDMAM), and a feature refinement approach with direction field (FRDF). SDMAM improves the encoder feature representation by considering the distance relationship between foreground-background pixels and boundaries, which increases the inter-class separation between foreground and background features. FRDF rectifies the features of boundary neighborhoods by exploiting the features inside salient objects. It utilizes the direction relationship of object pixels to enhance the intra-class compactness of salient features. In addition, we constitute a transformer-based decoder to decode multispectral feature representation. Experimental results on three public RGB-T SOD datasets demonstrate that our proposed method not only outperforms the state-of-the-art methods, but also can be integrated with different backbone networks in a plug-and-play manner. Ablation study and visualizations further prove the validity and interpretability of our method. Heng Zhou 0006, Chunna Tian, Chengyang Li 0001, Yongqiang Xie, Zhongbo Li |
IEEE Trans. Image Process. | 1 |
| 2022 | Multispectral Fusion Transformer Network for RGB-Thermal Urban Scene Semantic SegmentationabstractSemantic segmentation plays a vital role in autonomous vehicles. Fusing the rich details of RGB image and the illumination robustness of thermal image has great potential to improve the performance of RGB-T semantic segmentation. In multispectral feature fusion, the current main methods are less effective in the characterization of correlations and complementarities of RGB-T. In order to generate robust cross-spectral fusion features, we propose a multispectral fusion transformer network (MFTNet). Specifically, we first design an MFT module to handle the intraspectra correlation and the interspectra complementarity of RGB-T in the multispectral fusion encoder. MFT effectively enhances the RGB-T feature representation under various challenges. Then, an optimization strategy with progressive deep supervision (PDS) loss is proposed to directly supervise the upper and lower layers of the decoder. This strategy can guide the decoder to achieve precise segmentation in a coarse-to-fine manner. Finally, plenty of experimental results prove the effectiveness of our method. On the MFNet dataset, MFNet achieved 74.7 mAcc and 57.3 mIoU, outperforming the state-of-the-art methods. Heng Zhou 0006, Chunna Tian, Qizheng Huo, Yongqiang Xie, Zhongbo Li |
IEEE Geosci. Remote. Sens. Lett. | 1 |