VLDB 2026 Research / reviewers in the wild / expert
Zhongbo Li
dblp:34/7884
· DBLP profile ↗
15ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Label co-occurrence guided hash for multi-label image retrieval
Yongqiang Xie, Xiumei Wang 0002, Zhongbo Li, Peitao Cheng |
Expert Syst. Appl. | 4 |
| 2026 | Instruction-conditioned visual token sparsification for efficient vision-language model inference
Danjun Liu, Yongqiang Xie, Zhongbo Li |
Knowl. Based Syst. | 6 |
| 2026 | Exploring Self-Image and Cross-Image Consistency Learning for Remote Sensing Burned Area SegmentationabstractThe increasing frequency of global wildfires has led to the destruction of vast forests and wetlands. Non-contact remote sensing technologies provide an effective means for accurate burned area segmentation (BAS). However, existing BAS methods often treat each image independently, focusing primarily on local pixel contexts while neglecting the broader semantic consistency of burned regions across different scenes. The lack of global context modeling limits their robustness, as burned areas typically exhibit distinctive and consistent visual characteristics such as color and texture across diverse environments. To address this limitation, we propose a Self-image and Cross-image Consistency Learning (SCCL) framework, which captures both local pixel-level relationships within a single image and global semantic dependencies across multiple images. By enforcing consistent and compact representations of burned regions within and across images, SCCL enhances segmentation robustness under varying weather and terrain conditions. Additionally, to refine boundary delineation between burned and unburned areas, we introduce a Burned Edge Injector (BEI) and an Edge-Injected Decoder (EID). We further construct two large-scale BAS benchmark datasets, BAS-AUS and BAS-EUR, for comprehensive evaluation. Experiments on these benchmarks demonstrate that our method achieves state-of-the-art performance, significantly outperforming previous approaches, with MAE reduced to 0.017 and 0.016, respectively. The new BAS benchmarks and code are available at https://github.com/VisionVerse/SCCL. Heng Zhou 0006, Chengyang Li 0001, Chunna Tian, Yongqiang Xie, Zhongbo Li, Xiaojun Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Multi-Scale Frequency Enhancement Network for Blind Image Deblurring
Yawen Xiang, Heng Zhou 0006, Chengyang Li 0001, Zhongbo Li, Yongqiang Xie |
IET Image Process. | 5 |
| 2025 | Multi-Weather Restoration: An Efficient Prompt-Guided Convolution ArchitectureabstractAddressing degraded weather conditions plays a vital role in practical applications. Many existing restoration approaches are limited to specific weather types, which limits their applicability to different weather scenarios. Advanced technologies, encompassing Transformer and diffusion model, have been harnessed to confront this challenge. However, these methods often heighten network complexity and prolong inference duration. To this end, we present MW-ConvNet, a U-shaped convolution-based network for multi-weather restoration. Specifically, the MW-Enc block and MW-Dec block are introduced to achieve simple yet strong feature extraction, which rely entirely on traditional 2D convolution. To improve adaptability to multiple weather conditions, a prompt generation module is designed to generate a representative weather prompt at the encoder’s terminus. Drawing inspiration from style transfer, the weather prompt is used to guide the decoder learning through a progressive restoration procedure. For future high-fidelity restoration, we introduce frequency separation through wavelet pooling blocks in encoder phase and corresponding up-sampling blocks in decoder phase. The segregated treatment of low-frequency and high-frequency features curbs the loss of textural information during network computation. It also future improves the quality and accuracy of generated weather prompt. Extensive experiments demonstrate that the proposed MW-ConvNet obtains superior performance compared to state-of-the-art methods across both weather-specific and real-world restoration tasks. Significantly, our method achieves an impressive inference speed of 0.12 seconds per$256\times 256$image, outpacing transformer-based and diffusion-based models. Chengyang Li 0001, Fangwei Sun, Heng Zhou 0006, Yongqiang Xie, Zhongbo Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Revisiting Source-Free Domain Adaptation Object Detection in ThresholdsabstractSource-free domain adaptive object detection (SFOD) aims to transfer models pre-trained on the source domain to the unlabeled target domain without requiring access to the source data. Most existing SFOD methods leverage pseudo-labels for self-supervised training in the target domain. We investigate the limitations of threshold techniques to obtain high-quality pseudo-labels. In response, we design the Sequential SourceFree domain adaptive Object Detection (S-SFOD) algorithm, which enhances the quality of pseudo-labels at both the image and instance levels. At the image level, we reconstruct the training dataset, prioritizing the training of images that yield more reliable pseudo-labels to help the model acquire valuable target domain knowledge in the initial training stages. At the instance level, we introduce an adaptive local-global threshold method to balance the quality and quantity of pseudo-labels by dynamically adjusting the thresholds based on the model's learning progress. By improving the quality of pseudo-labels through these complementary techniques at both the image and instance levels, we effectively transfer knowledge from the source domain to the target domain. Extensive experiments on multiple cross-domain object detection datasets demonstrate that our proposed method outperforms current state-of-the-art SFOD algorithms. The code and model will be released. Yuchen Dong, Chengyang Li 0001, Yongqiang Xie, Zhongbo Li |
IEEE Trans. Multim. | 4 |
| 2025 | Deformation-Resilient Multigranularity Learning for Unaligned RGB-T Semantic SegmentationabstractRGB-Thermal semantic segmentation (SS) aims to combine visual light and thermal images to determine the semantic category for each pixel and create an object mask. While existing methods typically rely on well-aligned RGB-T image pairs, real-world RGB-T pairs are often unaligned, and pixel-by-pixel alignment is both challenging and time-consuming. To address this critical issue, we introduce a new unaligned RGB-T SS benchmark and propose the deformation-resilient multigranularity learning (DML) method. DML explores the spatial consistency and modal complementarity of RGB-T and mitigates the interference of warped modalities by aligning multimodal features in a coarse-to-fine multigranularity strategy. Specifically, DML constructs a deformation-aware complementary feature enhancer (DCFE), which consists of deformation-aware feature alignment (DFA) and complementary feature aggregation (CFA) modules. DFA enhances the spatial alignment of RGB-T by estimating the deformation field of warped features. Then, CFA aggregates complementary contexts of modal differences across multiple scales to produce deformation-resilient and robust RGB-T feature representations. Finally, we design the multigranularity mask refinement engine (MMFE), which combines class-agnostic saliency prediction (CSP) and class-aware edge generation (CEG) auxiliary tasks to provide useful boundary and positional cues for SS decoders. The MMFE enhances semantic alignment and interclass separability, yielding object masks with sharp boundaries. Quantitative and qualitative experiments on aligned and unaligned datasets validate the effectiveness of our proposed DML, consistently outperforming existing methods designed for aligned RGB-T data. The new unaligned RGB-T SS benchmark and code are available at https://github.com/VisionVerse/Unaligned-RGBT-Semantic-Segmentation. Heng Zhou 0006, Chengyang Li 0001, Chunna Tian, Yongqiang Xie, Zhongbo Li, Xiaojun Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Deep learning in motion deblurring: current status, benchmarks and future prospects
Yawen Xiang, Heng Zhou 0006, Chengyang Li 0001, Fangwei Sun, Zhongbo Li, Yongqiang Xie |
Vis. Comput. | 5 |
| 2024 | Frequency-aware feature aggregation network with dual-task consistency for RGB-T salient object detection
Heng Zhou 0006, Chunna Tian, Chengyang Li 0001, Yongqiang Xie, Zhongbo Li |
Pattern Recognit. | 6 |
| 2023 | Reliable Function Computation Offloading in Cloud-Edge Collaborative Network
Shaonan Li, Yongqiang Xie, Zhongbo Li, Yumeng Tian |
ICA3PP (2) | 3 |
| 2023 | Detection-Friendly Dehazing: Object Detection in Real-World Hazy ScenesabstractAdverse weather conditions in real-world scenarios lead to performance degradation of deep learning-based detection models. A well-known method is to use image restoration methods to enhance degraded images before object detection. However, how to build a positive correlation between these two tasks is still technically challenging. The restoration labels are also unavailable in practice. To this end, taking the hazy scene as an example, we propose a union architecture BAD-Net that connects the dehazing module and detection module in an end-to-end manner. Specifically, we design a two-branch structure with an attention fusion module for fully combining hazy and dehazing features. This reduces bad impacts on the detection module when the dehazing module performs poorly. Besides, we introduce a self-supervised haze robust loss that enables the detection module to deal with different degrees of haze. Most importantly, an interval iterative data refinement training strategy is proposed to guide the dehazing module learning with weak supervision. BAD-Net improves further detection performance through detection-friendly dehazing. Extensive experiments on RTTS and VOChaze datasets show that BAD-Net achieves higher accuracy compared to the recent state-of-the-art methods. It is a robust detection framework for bridging the gap between low-level dehazing and high-level detection. Chengyang Li 0001, Heng Zhou 0006, Caidong Yang, Yongqiang Xie, Zhongbo Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Position-Aware Relation Learning for RGB-Thermal Salient Object DetectionabstractSalient object detection (SOD) is an important task in computer vision that aims to identify visually conspicuous regions in images. RGB-Thermal SOD combines two spectra to achieve better segmentation results. However, most existing methods for RGB-T SOD use boundary maps to learn sharp boundaries, which lead to sub-optimal performance as they ignore the interactions between isolated boundary pixels and other confident pixels. To address this issue, we propose a novel position-aware relation learning network (PRLNet) for RGB-T SOD. PRLNet explores the distance and direction relationships between pixels by designing an auxiliary task and optimizing the feature structure to strengthen intra-class compactness and inter-class separation. Our method consists of two main components: A signed distance map auxiliary module (SDMAM), and a feature refinement approach with direction field (FRDF). SDMAM improves the encoder feature representation by considering the distance relationship between foreground-background pixels and boundaries, which increases the inter-class separation between foreground and background features. FRDF rectifies the features of boundary neighborhoods by exploiting the features inside salient objects. It utilizes the direction relationship of object pixels to enhance the intra-class compactness of salient features. In addition, we constitute a transformer-based decoder to decode multispectral feature representation. Experimental results on three public RGB-T SOD datasets demonstrate that our proposed method not only outperforms the state-of-the-art methods, but also can be integrated with different backbone networks in a plug-and-play manner. Ablation study and visualizations further prove the validity and interpretability of our method. Heng Zhou 0006, Chunna Tian, Chengyang Li 0001, Yongqiang Xie, Zhongbo Li |
IEEE Trans. Image Process. | 7 |
| 2022 | Multispectral Fusion Transformer Network for RGB-Thermal Urban Scene Semantic SegmentationabstractSemantic segmentation plays a vital role in autonomous vehicles. Fusing the rich details of RGB image and the illumination robustness of thermal image has great potential to improve the performance of RGB-T semantic segmentation. In multispectral feature fusion, the current main methods are less effective in the characterization of correlations and complementarities of RGB-T. In order to generate robust cross-spectral fusion features, we propose a multispectral fusion transformer network (MFTNet). Specifically, we first design an MFT module to handle the intraspectra correlation and the interspectra complementarity of RGB-T in the multispectral fusion encoder. MFT effectively enhances the RGB-T feature representation under various challenges. Then, an optimization strategy with progressive deep supervision (PDS) loss is proposed to directly supervise the upper and lower layers of the decoder. This strategy can guide the decoder to achieve precise segmentation in a coarse-to-fine manner. Finally, plenty of experimental results prove the effectiveness of our method. On the MFNet dataset, MFNet achieved 74.7 mAcc and 57.3 mIoU, outperforming the state-of-the-art methods. Heng Zhou 0006, Chunna Tian, Qizheng Huo, Yongqiang Xie, Zhongbo Li |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2012 | Comparison and optimization of packet loss recovery methods based on AMR-WB for VoIP
Zhongbo Li, Stefan Bruhn, Jing Wang 0037, Jingming Kuang 0001 |
Speech Commun. | 1 |
| 2009 | Analytical and Experimental Comparison of Packet Loss Recovery Methods Based on AMR-WB for VoIPabstractForward error control (FEC) and multiple description coding (MDC) are two classical techniques to resist packet loss for voice over IP (VoIP). AMR-WB codec has been standardized for wideband speech conversational applications and has widely potential applications in the migration of wireless or wired networks toward a single converged IP network. However, how to choose the optimal FEC or MDC for AMR-WB in different loss rate conditions is an unexplored option. In this paper, we compare the performance of different FEC and MDC techniques for AMR-WB codec both analytically and experimentally. Based on the comparison results, some practical configurations of FEC and MDC for AMR-WB codec are obtained. Zhongbo Li, Stefan Bruhn, Jingming Kuang 0001 |
ICC | 1 |