EDBT 2026 Demo / reviewers in the wild / expert
Hao Shen 0006
dblp:26/2210-6
· DBLP profile ↗
15ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-1945-4016ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A novel multi-granularity context-adaptive downsampling convolution for low-resolution images
Zejun Gu, Zhong-Qiu Zhao, Hao Shen 0006, Zhao Zhang 0001, De-Shuang Huang |
Inf. Sci. | 3 |
| 2026 | Human-Structure-Aware Token Position Embedding for Tokenized Pose EstimationabstractTokenized pose estimation (TPE) has demonstrated remarkable performance in lightweight human pose estimation (HPE) models. However, existing TPE methods typically initialize keypoint tokens randomly, without explicitly incorporating human structure priors. These priors play a vital role in HPE by effectively mitigating common challenges such as occlusion and ambiguity. To this end, we propose a Structure-Aware Keypoint Position Embedding (SAKPE). This embedding explicitly encodes inherent structural properties of the human body, such as symmetry and order, into the positional coordinates of keypoint tokens. It also employs learnable scale and offset factors to adapt to diverse human poses, thereby fully exploiting the geometric constraints among keypoints. Furthermore, to better leverage the positional relationships among patch tokens, we introduce a Layer-adaptive Hybrid Patch Position Embedding (LHPPE). It dynamically fuses absolute and relative position embeddings of patch tokens based on attention distributions across Transformer layers, enabling the model to learn both absolute and relative positional information adaptively. Taking the two together, we propose a novel position embedding method for pose estimation, named Human-structure-aware Token Position Embedding (HTPE). It significantly improves the performance of various TPE models. Extensive experiments on COCO, CrowdPose, and OCHuman show that HTPE achieves state-of-the-art (SOTA) performance among lightweight methods, with a negligible increase in parameters and FLOPs. Notably, it demonstrates consistent improvements under occlusion,, achieving up to 3.3 AP gains. The source code can be found in https://github.com/guzejungithub/HTPE. Zejun Gu, Zhong-Qiu Zhao, Henghui Ding, Hao Shen 0006, Zhenhua Tang 0001, Zhao Zhang 0001, De-Shuang Huang |
IEEE Trans. Image Process. | 4 |
| 2026 | Cross-Domain Knowledge Distillation for Low-Resolution Human Pose EstimationabstractIn practical applications of human pose estimation, low-resolution inputs frequently occur, and existing state-of-the-art models perform poorly with low-resolution images. This work focuses on boosting the performance of low-resolution models by distilling knowledge from a high-resolution model. However, we face the challenge of feature size mismatch and class number mismatch when applying knowledge distillation to networks with different input resolutions. To address this issue, we propose a novel cross-domain knowledge distillation (CDKD) framework. In this framework, we construct a scale-adaptive projector ensemble (SAPE) module to spatially align feature maps between models of varying input resolutions. It adopts a projector ensemble to map low-resolution features into multiple common spaces and adaptively merges them based on multi-scale information to match high-resolution features. Additionally, we construct a cross-class alignment (CCA) module to solve the problem of the mismatch of class numbers. By combining an easy-to-hard training (ETHT) strategy, the CCA module further enhances the distillation performance. The effectiveness and efficiency of our approach are demonstrated by extensive experiments on three common benchmark datasets: MPII, COCO, and Crowdpose. Zejun Gu, Zhong-Qiu Zhao, Henghui Ding, Hao Shen 0006, Zhao Zhang 0001, De-Shuang Huang |
IEEE Trans. Multim. | 4 |
| 2025 | Spatial Frequency Modulation Network for Efficient Image DehazingabstractCurrently, two main research lines in efficient context modeling for image dehazing are tailoring effective feature modulation mechanisms and utilizing the Fourier transform more precisely. The former is usually based on self-scale features that ignore complementary cross-scale/level features, and the latter tends to overlook regions with pronounced haze degradation and intricate structures. This paper introduces a novel spatial and frequency modulation perspective to synergistically investigate contextual feature modeling for efficient image dehazing. Specifically, we delicately develop a Spatial Frequency Modulator (SFM) equipped with a Cross-Scale Modulator (CSM) and Frequency Modulator (FM) to implement intra-block feature modulation. The CSM progressively aggregates hierarchical features across different scales, employing them for spatial self-modulation, and the FM subsequently adopts a dual-branch design to focus more on the crucial areas with severe haze and complex structures for reconstruction. Further, we propose a Cross-Level Modulator (CLM) to facilitate inter-block feature mutual modulation, enhancing seamless interaction between features at different depths and layers. Integrating the above-developed modules into the U-Net architecture, we construct a two-stage spatial frequency modulation network (SFMN). Extensive quantitative and qualitative evaluations showcase the superior performance and efficiency of the proposed SFMN over recent state-of-the-art image dehazing methods. The source code can be found in https://github.com/it-hao/SFMN. Hao Shen 0006, Henghui Ding, Yulun Zhang 0001, Zhong-Qiu Zhao, Xudong Jiang 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Adaptive branch selection for accelerate image super-resolution
Zhong-Qiu Zhao, Hao Shen 0006, Xiufeng Liu 0005 |
Vis. Comput. | 3 |
| 2024 | A Semi-Supervised Nighttime Dehazing Baseline with Spatial-Frequency Aware and Realistic Brightness ConstraintabstractExisting research based on deep learning has extensively explored the problem of daytime image dehazing. However, few studies have considered the characteristics of nighttime hazy scenes. There are two distinctions between nighttime and daytime haze. First, there may be multiple active col-ored light sources with lower illumination intensity in night-time scenes, which may cause haze, glow and noise with localized, coupled and frequency inconsistent characteris-tics. Second, due to the domain discrepancy between simulated and real-world data, unrealistic brightness may occur when applying a dehazing model trained on simulated data to real-world data. To address the above two issues, we propose a semi-supervised model for real-world nighttime dehazing. First, the spatial attention and frequency spectrum filtering are implemented as a spatial-frequency do-main information interaction module to handle the first is-sue. Second, a pseudo-label-based retraining strategy and a local window-based brightness loss for semi-supervised training process is designed to suppress haze and glow while achieving realistic brightness. Experiments on public benchmarks validate the effectiveness of the proposed method and its superiority over state-of-the-art methods. The source code and Supplementary Materials are placed in the https://github.com/Xiaofeng-life/SFSNiD. Xiaofeng Cong, Jie Gui, Jing Zhang 0037, Junming Hou, Hao Shen 0006 |
CVPR | 5 |
| 2024 | Contextual Feature Modulation Network for Efficient Super-Resolution
Wandi Zhang, Hao Shen 0006, Weidong Tian 0001, Zhong-Qiu Zhao |
ICIC (6) | 2 |
| 2024 | Global routing between capsules
Hao Shen 0006, Zhong-Qiu Zhao, Yi Yang 0036, Zhao Zhang 0001 |
Pattern Recognit. | 2 |
| 2024 | Rethinking Pan-Sharpening via Spectral-Band ModulationabstractPan-sharpening aims to super-resolve the low-resolution (LR) multispectral (MS) image under the guidance of a high-resolution (HR) panchromatic (PAN) image. Existing deep learning (DL)-based pan-sharpening methods usually adhere to a common philosophy of learning complementary information between MS and PAN images. Despite remarkable advances, few studies consider the band-private characteristics which differ greatly from band to band. An ideal MS image, however, is jointly determined by its diverse spectral bands, thus the accurate restoration of every band will benefit the pan-sharpening performance. In this work, we propose a novel yet effective solution to reconstruct the HRMS image by explicitly modulating every spectral band under the conditions of the PAN image. As a result, we design a spatially-adaptive spectral modulation network, dubbed SSMNet, which consists of three core designs: source-aware spectral modulator (SSM), cross-band information aggregation (CBIA) module, and cross-stage feature integration (CSFI) module. The first predicts a series of spatially-adaptive kernels to capture the local information of every spectral band. Followed by, the second is responsible for facilitating the information communication among various bands to guarantee continuous spectral representations. Furthermore, the third attends to integrate the cross-stage output features to produce the pan-sharpened result. In addition, we also introduce the histogram loss to constrain the band-wise distribution of the final fused products. Extensive experiments demonstrate that our SSMNet achieves favorable performance against other state-of-the-art (SOTA) methods on multiple satellite datasets. The code is available athttps://github.com/ez4lionky/SSMNet/. Junming Hou, Xiaofeng Cong, Hao Shen 0006, Zhuochen Lou, Liang-Jian Deng, Jian Wei You |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Spatial-Frequency Adaptive Remote Sensing Image Dehazing With Mixture of ExpertsabstractThe feature modulation mechanism has been demonstrated to be particularly well-suited for efficient network design and is rarely explored in remote sensing dehazing tasks. Moreover, we observe distinct patterns in haze distribution across the low-frequency (LF) and high-frequency (HF) components of haze images from various datasets. However, existing research rarely investigated the potential solution in the frequency domain. In response, we propose a novel spatial-frequency adaptive network (SFAN), which is mainly built by the proposed mixture of modulation experts (MoME) and decoupled frequency learning block (DFLB). Different from the fixed feature modulation design used in other tasks, the MoME adopts the mixture-of-expert mechanism to dynamically learn diverse contextual features of various granularities and scales in a sample-adaptive manner and then utilize them to perform elementwise local feature modulation. This pure convolution architecture enables our network to have superior performance and efficiency tradeoffs. Furthermore, the DFLB is devised to facilitate the LF global haze removal and reconstruction of HF local texture information. At the micro level, we first utilize a mask extractor (ME) to generate the frequency mask from the input hazy image, then employ a dual-branch decoupled learning unit to boost frequency learning, and finally develop a mixture of fusion experts (MoFE) to achieve HF and LF feature interaction. Extensive experiments on publicly available dehazing datasets demonstrate that our network performs superior performance while incurring lower computational costs. Compared to the state-of-the-art approach (DEA-Net), SFAN achieves, an average, 0.83-dB PSNR improvement on five remote sensing datasets but consumes only 51% of the FLOPs. The code will be available athttps://github.com/it-hao/SFAN. Hao Shen 0006, Henghui Ding, Yulun Zhang 0001, Xiaofeng Cong, Zhong-Qiu Zhao, Xudong Jiang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Adaptive Dynamic Filtering Network for Image DenoisingabstractIn image denoising networks, feature scaling is widely used to enlarge the receptive field size and reduce computational costs. This practice, however, also leads to the loss of high-frequency information and fails to consider within-scale characteristics. Recently, dynamic convolution has exhibited powerful capabilities in processing high-frequency information (e.g., edges, corners, textures), but previous works lack sufficient spatial contextual information in filter generation. To alleviate these issues, we propose to employ dynamic convolution to improve the learning of high-frequency and multi-scale features. Specifically, we design a spatially enhanced kernel generation (SEKG) module to improve dynamic convolution, enabling the learning of spatial context information with a very low computational complexity. Based on the SEKG module, we propose a dynamic convolution block (DCB) and a multi-scale dynamic convolution block (MDCB). The former enhances the high-frequency information via dynamic convolution and preserves low-frequency information via skip connections. The latter utilizes shared adaptive dynamic kernels and the idea of dilated convolution to achieve efficient multi-scale feature extraction. The proposed multi-dimension feature integration (MFI) mechanism further fuses the multi-scale features, providing precise and contextually enriched feature representations. Finally, we build an efficient denoising network with the proposed DCB and MDCB, named ADFNet. It achieves better performance with low computational complexity on real-world and synthetic Gaussian noisy datasets. The source code is available at https://github.com/it-hao/ADFNet. Hao Shen 0006, Zhong-Qiu Zhao, Wandi Zhang |
AAAI | 1 |
| 2023 | Mutual Information-driven Triple Interaction Network for Efficient Image DehazingabstractMulti-stage architectures have exhibited efficacy in image dehazing, which usually decomposes a challenging task into multiple more tractable sub-tasks and progressively estimates latent hazy-free images. Despite the remarkable progress, existing methods still suffer from the following shortcomings: (1) limited exploration of frequency domain information; (2) insufficient information interaction; (3) severe feature redundancy. To remedy these issues, we propose a novel Mutual Information-driven Triple interaction Network (MITNet) based on spatial-frequency dual domain information and two-stage architecture. To be specific, the first stage, named amplitude-guided haze removal, aims to recover the amplitude spectrum of the hazy images for haze removal. And the second stage, named phase-guided structure refined, devotes to learning the transformation and refinement of the phase spectrum. To facilitate the information exchange between two stages, an Adaptive Triple Interaction Module (ATIM) is developed to simultaneously aggregate cross-domain, cross-scale, and cross-stage features, where the fused features are further used to generate content-adaptive dynamic filters so that applying them to enhance global context representation. In addition, we impose the mutual information minimization constraint on paired scale encoder and decoder features from both stages. Such an operation can effectively reduce information redundancy and enhance cross-stage feature complementarity. Extensive experiments on multiple public datasets exhibit that our MITNet performs superior performance with lower model complexity. The code and models are available at https://github.com/it-hao/MITNet. Hao Shen 0006, Zhong-Qiu Zhao, Yulun Zhang 0001, Zhao Zhang 0001 |
ACM Multimedia | 1 |
| 2022 | Joint operation and attention block search for lightweight image restoration
Hao Shen 0006, Zhong-Qiu Zhao, Wenrui Liao, Weidong Tian 0001, De-Shuang Huang |
Pattern Recognit. | 1 |
| 2021 | Residual Attention Block Search for Lightweight Image Super-ResolutionabstractRecently, lightweight neural networks with different manual designs have presented a promising performance in single image super-resolution (SR). However, these designs rely on too much expert experience. To address this issue, we focus on searching a lightweight block for efficient and accurate image SR. Due to the frequent use of various residual blocks and attention mechanisms in SR methods, we propose the residual attention search block (RASB) which combines an operation search block (OSB) with an attention search block (ASB). The former is used to explore the suitable operation at the proper position, and the latter is applied to discover the optimal connection of various attention mechanisms. Moreover, we build the modified residual attention network (MRAN) with stacked found blocks and a refinement module. Extensive experiments demonstrate that our MRAN achieves a better trade-off against the state-of-the-art methods in terms of accuracy and model complexity. Wenrui Liao, Zhong-Qiu Zhao, Hao Shen 0006, Weidong Tian 0001 |
ICME | 3 |
| 2020 | Mid-Weight Image Super-Resolution with Bypass Connection Attention NetworkabstractDeeper networks have limited improvements for image super-resolution (SR), and are much more difficult to train. The main reason is that these networks consist of many stacked building blocks which can produce many redundant features. Besides, most of SR methods neglect the fact that different features contain various types of information with varying degrees of contributions to image reconstruction, and thus lack sufficient representational capability. Taking these issues into account, we propose a mid-weight bypass connection attention network (BCAN) with more powerful representational capability but fewer parameters. In detail, we design a novel bypass connection attention module (BCAM), which consists of several bypass connection attention blocks (BCABs), enhancing high contribution information and suppressing redundant information. Further, we embed a mixed residual attention unit (MRAU) in each BCAB, which is composed of a channel attention unit and a spatial attention unit. After obtaining all hierarchical features, we propose an adaptive feature fusion module (AFFM), which can effectively combine hierarchical features based on different contributions of each BCAM. Experiments on benchmark datasets with various degradation models show that our BCAN can achieve better performance than existing state-of-the-art methods. Hao Shen 0006, Zhong-Qiu Zhao |
ECAI | 1 |