VLDB 2026 Research / reviewers in the wild / expert
Hengsheng Zhang
dblp:234/7935
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0001-6738-3462ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OmniScaleSR: Unleashing Scale-Controlled Diffusion Prior for Faithful and Realistic Arbitrary-Scale Image Super-ResolutionabstractArbitrary-scale super-resolution (ASSR) overcomes the limitation of traditional super-resolution (SR) that works only at a fixed scale (e.g., ×4), enabling a single model to achieve arbitrary-scale SR. Most ASSR methods explicitly incorporate implicit neural representation (INR) to achieve ASSR, but INR’s inherently regression-driven feature extraction and aggregation nature restricts their capacity to synthesize meticulous details, leading to low realism. Recently, diffusion-based realistic image super-resolution (Real-ISR) methods leverage the pre-trained diffusion prior and have shown promising results at ×4 scale. We find that they could also achieve ASSR because the powerful pre-trained diffusion prior implicitly employs SR scale adaptation by encouraging the model to always generate high-realism images. However, due to the lack of explicit SR scale controls, the model fails to effectively manage the diffusion behavior according to different SR scales, causing either excessive hallucination or blurry results, especially for ultra-high magnification. To address these limitations, we proposeOmniScaleSR, a novel diffusion-based realistic arbitrary-scale super-resolution (Real-ASSR) method to achieve both high fidelity and high-realism ASSR. We introduce explicit diffusion-native SR scale controls, which could be elegantly coupled with the implicit scale adaptation, unleashing scale-controlled diffusion prior to dynamically managing the diffusion behavior in a content- and scale-aware manner. Furthermore, we incorporate multi-domain fidelity enhancement designs to achieve more faithful reconstruction. Extensive experiments on both bicubic degradation benchmarks and real-world datasets demonstrate that OmniScaleSR consistently outperforms state-of-the-art methods in terms of both fidelity and perceptual realism, with especially strong performance under high-magnification scenarios. Codes will be at https://github.com/chaixinning/OmniScaleSR. Xinning Chai, Zhengxue Cheng, Hengsheng Zhang, Yingsheng Qin, Yucai Yang, Rong Xie 0004, Li Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Diff-Restorer: Unleashing Visual Prompts for Diffusion-Based Universal Image RestorationabstractImage restoration aims to recover high-quality images from degraded observations, yet real-world degradations are complex, coupled, and difficult to model. Existing task-specific methods struggle to generalize beyond predefined degradation types, while recent all-in-one or prompt-based methods still face three key challenges: (1) they rely on task-specific training or fixed prompt pools, limiting adaptability to real-world and mixed degradations; (2) human-instruction or implicit-prompt mechanisms make them difficult to use in practice; and (3) they often fail to balance structural fidelity and perceptual realism. To address these issues, we propose Diff-Restorer, a diffusion-based universal image restoration framework that unifies diverse degradation handling within a single model. Diff-Restorer adaptively extracts decoupled visual prompts from a visual-language model (CLIP), including clear semantic and degradation embeddings. The clear semantic embeddings serve as content prompts to guide the diffusion model for generation, improving perceptual quality. The degradation embeddings as the task identifier modulate the Image-guided Control Module to generate structure control, ensuring faithfulness. Furthermore, we design a Task-aware Decoder to perform structural correction and convert the latent code to the pixel domain. Extensive experiments on various single, real-world, and mixed degradation tasks show that Diff-Restorer outperforms state-of-the-art methods in terms of generality, realism, and fidelity. Hengsheng Zhang, Xinning Chai, Zhengxue Cheng, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | SSP-IR: Semantic and Structure Priors for Diffusion-Based Realistic Image RestorationabstractRealistic image restoration is a crucial task in computer vision, and diffusion-based models for image restoration have garnered significant attention due to their ability to produce realistic results. Restoration can be seen as a controllable generation conditioning on priors. However, due to the severity of image degradation, existing diffusion-based restoration methods cannot fully exploit priors from low-quality images and still have many challenges in perceptual quality, semantic fidelity, and structure accuracy. Based on the challenges, we introduce a novel image restoration method, SSP-IR. Our approach aims to fully exploit semantic and structure priors from low-quality images to guide the diffusion model in generating semantically faithful and structurally accurate natural restoration results. Specifically, we integrate the visual comprehension capabilities of Multimodal Large Language Models (explicit) and the visual representations of the original image (implicit) to acquire accurate semantic prior. To extract degradation-independent structure prior, we introduce a Processor with RGB and FFT constraints to extract structure prior from the low-quality images, guiding the diffusion model and preventing the generation of unreasonable artifacts. Lastly, we employ a multi-level attention mechanism to integrate the acquired semantic and structure priors. The qualitative and quantitative results demonstrate that our method outperforms other state-of-the-art methods overall on both synthetic and real-world datasets. Our project page ishttps://zyhrainbow.github.io/projects/SSP-IR. Hengsheng Zhang, Zhengxue Cheng, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Hdrtvformer: Efficient Sdrtv-to-Hdrtv via Affine Transformation and Spatial-Aware TransformerabstractRecent works on reconstructing HDR videos in display format (HDRTV) suffer from high computational and memory requirements because they learn the SDRTV-to-HDRTV mapping directly in 4K resolution. This paper proposes an efficient SDRTV-to-HDRTV model (HDRTVFormer) that decomposes the HDRTV restoration into SDRTV-to-HDRTV Domain Mapping and HDRTV Refinement. SDRTV-to-HDRTV Domain Mapping is an affine transformation-based model that learns SDRTV-to-HDRTV affine coefficients in low-resolution space, achieving rapid processing times. To enhance the accuracy of the predicted affine coefficients, the model introduces global information-modulated feature extraction blocks and a detail guidance upsampling module. For HDRTV Refinement, we propose a spatial-aware Transformer to refine the luminance and color details. We modify the self-attention and feed-forward network of Transformer blocks to improve efficiency and feature representations. Experimental results have demonstrated that our method outperforms other state-of-the-art works in performance and efficiency. Hengsheng Zhang, Xinning Chai, Rong Xie 0004, Li Song 0001 |
ICASSP | 1 |
| 2023 | Dual-Head Fusion Network for Image EnhancementabstractImage enhancement algorithms have made great progress recently. However, most existing methods tend to construct a uniform enhancer for the color transformation of all pixels and ignore the local context information which is significant for photographs, causing unsatisfactory results. To solve these issues, we propose a novel dual-head fusion network for image enhancement, which synthetically considers both global scenario and local content information. Our network consists of four lightweight modules. We first develop a dual-head feature extraction module to extract the global condition vector and spatial context map. After that, we propose a context-aware retouching module and a global color rendering module to generate latent results. Finally, we employ the spatial attention based fusion module to adaptively aggregate the latent results. Experiments on public datasets show that our method consistently achieves the best results compared with SOTA methods both quantitatively and qualitatively. Hengsheng Zhang, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
ICASSP | 2 |
| 2022 | A Codec Information Assisted Framework for Efficient Compressed Video Super-Resolution
Hengsheng Zhang, Xueyi Zou, Jiaming Guo, Youliang Yan, Rong Xie 0004, Li Song 0001 |
ECCV (17) | 1 |
| 2022 | MLS-GAN: Multi-Level Semantic Guided Image ColorizationabstractImage colorization predicts plausible color versions of given grayscale images. Recently, several methods incorporate image semantics to assist image colorization and have shown impressive performance. To further exploit and take full advantage of more semantic information, in this paper, we propose a Multi-Level Semantic guided Generative Adversarial Network (MLS-GAN) for image colorization. Specifically, we utilize three different levels of semantics to guide the colorization process: image level, segmentation level and contextual level. Image-level classification semantics is used to learn category and high-level semantics, ensuring the reasonability of color results. At the segmentation level, multi-scale saliency map semantics is extracted to provide figure-background separation information, which can efficiently alleviate semantic confusion, especially for images with complex backgrounds. Furthermore, we novelly use non-local blocks to capture long-range semantic dependencies at the contextual level. Experiments show that our method enhances color consistency and can produce more vivid color in visually important regions, outperforming state-of-the-art methods qualitatively and quantitatively. Xinning Chai, Xibei Liu, Hengsheng Zhang, Li Song 0001, Liean Cao |
ICIP | 3 |
| 2022 | Joint resource management for mobility supported federated learning in Internet of Vehicles
Ge Wang 0006, Fangmin Xu, Hengsheng Zhang, Chenglin Zhao |
Future Gener. Comput. Syst. | 3 |
| 2022 | Learning to Optimize User Association and Spectrum Allocation With Partial Observation in mmWave-Enabled UAV NetworksabstractTo support large-scale unmanned aerial vehicle (UAV) networks with both payload communication (PC) packets and control and non-payload communication (CNPC) packets, millimeter wave (mmWave) communication is convinced to be a promising solution. However, efficient user association and spectrum allocation are still challenging in mmWave-enabled UAV networks considering that the network state is dynamic and the status information of each UAV is incomplete. In this paper, we investigate a joint UAV association and spectrum allocation problem under a hybrid mmWave sharing paradigm, where PC packets are transmitted over both licensed and pooled bands in a shared manner to achieve high throughput and CNPC packets are transmitted over licensed band in an exclusive manner to guarantee high reliability. To this end, we introduce a strategic form game that can characterize network stochastic states and individual partial observations to reformulate the problem, and then propose a counterfactual regret minimization scheme to achieve its correlated equilibrium (CE). Benefited from the updating mechanism on randomobservation-actionpairs, the designed scheme can converge to the corresponding CE solutions for the two types of packets with partial observation. Finally, our simulation results demonstrate the superior performance of the proposed scheme over the baseline schemes. Chaoqiong Fan, Changyang She, Hengsheng Zhang, Bin Li 0002, Chenglin Zhao, Dusit Niyato |
IEEE Trans. Wirel. Commun. | 3 |