Heng Liu 0002

dblp:59/3260-2 · DBLP profile ↗
← Back
23ranked-venue papers
11as first author
13since 2021 · last 2026
0000-0001-7563-2676ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Fine-grained contrastive and complete fusion for multi-view clustering
Shudong Hou, Yongchi Fan, Heng Liu 0002
Neurocomputing4
2026 Topology-aware contrastive learning for attributed graph clustering
Weizhi Zhao, Heng Liu 0002, Shudong Hou
Neurocomputing2
2026 ChipDiff: Staged diffusion model with loss gradient guidance for Chinese ink painting style transfer
Heng Liu 0002, Yongzhen Wang 0001, Bingwen Hu, Yang Wang 0023
Pattern Recognit.1
2026 WDMamba: When Wavelet Degradation Prior Meets Vision Mamba for Image Dehazing
abstract
In this paper, we reveal a novel haze-specific wavelet degradation prior observed through wavelet transform analysis, which shows that haze-related information predominantly resides in low-frequency components. Exploiting this insight, we propose a novel dehazing framework, WDMamba, which decomposes the image dehazing task into two sequential stages: low-frequency restoration followed by detail enhancement. This coarse-to-fine strategy enables WDMamba to effectively capture features specific to each stage of the dehazing process, resulting in high-quality restored images. Specifically, in the low-frequency restoration stage, we integrate Mamba blocks to reconstruct global structures with linear complexity, efficiently removing overall haze and producing a coarse restored image. Thereafter, the detail enhancement stage reinstates fine-grained information that may have been overlooked during the previous phase, culminating in the final dehazed output. Furthermore, to enhance detail retention and achieve more natural dehazing, we introduce a self-guided contrastive regularization during network training. By utilizing the coarse restored output as a hard negative example, our model learns more discriminative representations, substantially boosting the overall dehazing performance. Extensive evaluations on public dehazing benchmarks demonstrate that our method surpasses state-of-the-art approaches both qualitatively and quantitatively. Code is available at https://github.com/SunJ000/WDMamba.
Heng Liu 0002, Yongzhen Wang 0001, Xiao-Ping Zhang 0002, Mingqiang Wei
IEEE Trans. Circuits Syst. Video Technol.2
2026 CLIP-SR: Collaborative Linguistic and Image Processing for Super-Resolution
abstract
Convolutional Neural Networks (CNNs) have significantly advanced Image Super-Resolution (SR), yet most CNN-based methods rely solely on pixel-based transformations, often leading to artifacts and blurring, particularly under severe downsampling rates (e.g., 8× or 16×). The recently developed text-guided SR approaches leverage textual descriptions to enhance their detail restoration capabilities but frequently struggle with effectively performing alignment, resulting in semantic inconsistencies. To address these challenges, we propose a multi-modal semantic enhancement framework that integrates textual semantics with visual features, effectively mitigating semantic mismatches and detail losses in highly degraded low-resolution (LR) images. Our method enables realistic, high-quality SR to be performed at large upscaling factors, with a maximum scaling ratio of 16×. The framework integrates both text and image inputs using the prompt predictor, the Text-Image Fusion Block (TIFBlock), and the Iterative Refinement Module, leveraging Contrastive Language-Image Pretraining (CLIP) features to guide a progressive enhancement process with fine-grained alignment. This synergy produces high-resolution outputs with sharp textures and strong semantic coherence, even at substantial scaling factors. Extensive comparative experiments and ablation studies validate the effectiveness of our approach. Furthermore, by leveraging textual semantics, our method offers a degree of super-resolution editability, allowing for controlled enhancements while preserving semantic consistency.
Bingwen Hu, Heng Liu 0002, Zhedong Zheng, Ping Liu 0004
IEEE Trans. Multim.2
2026 Real-Scene Image Dehazing via Laplacian Pyramid-Based Conditional Diffusion Model
abstract
Recent diffusion models have demonstrated exceptional efficacy across various image restoration tasks, but still suffer from time-consuming and substantial computational resource consumption. To address these challenges, we present LPCDiff, a novel Laplacian Pyramid-based Conditional Diffusion model designed for real-scene image dehazing. LPCDiff leverages the Laplacian pyramid decomposition to decouple the input image into two components: the low-resolution low-pass image and the high-frequency residuals. These components are subsequently reconstructed through a diffusion model and a well-designed high-frequency residual recovery module. With such a strategy, LPCDiff can substantially accelerate inference speed and reduce computational costs without sacrificing image fidelity. In addition, the framework empowers the model to capture intrinsic high-frequency details and low-frequency structural information within the image, resulting in sharper and more realistic haze-free outputs. Moreover, to extract more valuable information from the limited training data, we introduce a low-frequency refinement module to further enhance the intricate details of the final dehazed images. Through extensive experimentation, our method significantly outperforms 12 state-of-the-art approaches on three real-world and one synthetic image dehazing benchmarks. Code is available athttps://github.com/yz-wang/LPCDiff.
Yongzhen Wang 0001, Heng Liu 0002, Xiao-Ping Zhang 0002, Mingqiang Wei
IEEE Trans. Multim.3
2025 Unsupervised Cross-Modal Person Search via Progressive Diverse Text Generation
abstract
While text-based person search (TBPS) has achieved notable progress in recent years, existing methods heavily rely on laboriously annotated and well-aligned pedestrian image-text pairs, incurring prohibitive annotation costs. To overcome this limitation, we propose to train a TBPS model using pure images without any annotations. To tackle this challenging problem, we propose an unsupervised cross-modal person search framework via Progressive Diverse Text Generation (PSPD), leveraging large pre-trained models as assistants. Particularly, PSPD features three modules: Progressive Diverse Text Generation (PDTG), Fine-grained Saliency Region Alignment (FSRA) and Cross-Modal pseudo Label Correction (CMLC), allowing training with only unannotated images. The PDTG generates and dynamically adjusts prompts to produce accurate, diverse textual descriptions in multiple styles. The FSRA then uses large language models to generate fine-grained attributes and achieves cross-modal fine-grained semantic alignment. Additionally, the CMLC is applied to eliminate pseudo label noise through dual mutual nearest-neighbor matching, combined with distance-based judgment and a voting mechanism. Experimental results demonstrate the effectiveness of our method in unsupervised settings across various text-based person search datasets. Source code is at https://github.com/flychen321/PSPD.
Jielong He, Heng Liu 0002, Yaxiong Wang
ACM Multimedia4
2025 Few-Shot Referring Video Single- and Multi-Object Segmentation Via Cross-Modal Affinity with Instance Sequence Matching
Heng Liu 0002, Mingqi Gao 0003, Xiantong Zhen, Feng Zheng 0001, Yang Wang 0023
Int. J. Comput. Vis.1
2025 Effective multi-view representation learning for single-view attributed graph clustering
Heng Liu 0002, Weizhi Zhao, Zhou Bao, Mingquan Ye, Caifeng Shan
Knowl. Based Syst.1
2025 Steerable Graph Neural Network on Point Clouds via Second-Order Random Walks
abstract
Point cloud analysis, arising from computer graphics, remains a fundamental but challenging problem, mainly due to the non-Euclidean property of point cloud data modality. With the snap increase in the amount and breadth of related research in deep learning for graphs, many important works come in the form of graphs representing the point clouds. In this paper, we present a sampling adaptive graph convolutional network that combines the powerful representation ability of random walk subgraph searching and the essential success of the Fisher vector. Extending from those existing graph representation learning or embedding methods with multi-hop neighbor random searching, we sample multi-scale walk fields by using asteerableexploration-exploitationsecond order random walk, which endows our model with the most flexibility compared with the original first order random walk. To encode each-scale walk field consisting of several walk paths, specifically, we characterize these paths of walk field by Gaussian mixture models (GMMs) so as to better analogize the standard CNNs on Euclidean modality. Each Gaussian component implicitly defines a direction and all of them properly encode thespatial layoutof walk fields after the gradient projecting to the space of Gaussian parameters, i.e. the Fisher vectors. Thereby, we introduce and name our deep graph convolutional network asPointFisher. Comprehensive evaluations on several public datasets well demonstrate the superiority of our proposed learning method over other state-of-the-arts for point cloud classification and segmentation.
Xianglin Guo, Heng Liu 0002, Haoran Xie 0001, Gary Cheng 0001, Fu Lee Wang
IEEE Trans. Multim.3
2023 Learning Cross-Modal Affinity for Referring Video Object Segmentation Targeting Limited Samples
abstract
Referring video object segmentation (RVOS), as a supervised learning task, relies on sufficient annotated data for a given scene. However, in more realistic scenarios, only minimal annotations are available for a new scene, which poses significant challenges to existing RVOS methods. With this in mind, we propose a simple yet effective model with a newly designed cross-modal affinity (CMA) module based on a Transformer architecture. The CMA module builds multimodal affinity with a few samples, thus quickly learning new semantic information, and enabling the model to adapt to different scenarios. Since the proposed method targets limited samples for new scenes, we generalize the problem as - few-shot referring video object segmentation (FS-RVOS). To foster research in this direction, we build up a new FS-RVOS benchmark based on currently available datasets. The benchmark covers a wide range and includes multiple situations, which can maximally simulate real-world scenarios. Extensive experiments show that our model adapts well to different scenarios with only a few samples, reaching state-of-the-art performance on the benchmark. On Mini-Ref-YouTube-VOS, our model achieves an average performance of 53.1 ${\mathcal{J}}$ and 54.8 ${\mathcal{F}}$, which are 10% better than the baselines. Furthermore, we show impressive results of 77.7 ${\mathcal{J}}$ and 74.8 ${\mathcal{F}}$ on Mini-Ref-SAIL-VOS, which are significantly better than the baselines. Code is publicly available at https://github.com/hengliusky/Few_shot_RVOS.
Mingqi Gao 0003, Heng Liu 0002, Xiantong Zhen, Feng Zheng 0001
ICCV3
2023 Perception consistency ultrasound image super-resolution via self-supervised CycleGAN
Heng Liu 0002, Jianyong Liu, Shudong Hou, Tao Tao 0005, Jungong Han
Neural Comput. Appl.1
2022 Progressive Residual Learning With Memory Upgrade for Ultrasound Image Blind Super-Resolution
abstract
For clinical medical diagnosis and treatment, image super-resolution (SR) technology will be helpful to improve the ultrasonic imaging quality so as to enhance the accuracy of disease diagnosis. However, due to the differences of sensing devices or transmission media, the resolution degradation process of ultrasound imaging in real scenes is uncontrollable, especially when the blur kernel is usually unknown. This issue makes current end-to-end SR networks poor performance when applied to ultrasonic images. Aiming to achieve effective SR in real ultrasound medical scenes, in this work, we propose a blind deep SR method based on progressive residual learning and memory upgrade. Specifically, we estimate the accurate blur kernel from the spatial attention map block of low resolution (LR) ultrasound image through a multi-label classification network, then we construct three modules-up- sampling (US) module, residual learning (RL) model and memory upgrading (MU) model for ultrasound image blind SR. The US module is designed to upscale the input information and the up-sampled residual result will be used for SR reconstruction. The RL module is employed to approximate the original LR and continuously generate the updated residual and feed it to the next US module. The last MU module can store all progressively learned residuals, which offers increased interactions between the US and RL modules, augmenting the details recovery. Extensive experiments and evaluations on the benchmark CCA-US and US-CASE datasets demonstrate the proposed approach achieves better performance against the state-of-the-art methods.
Heng Liu 0002, Jianyong Liu, Caifeng Shan
IEEE J. Biomed. Health Informatics1
2019 Pseudo Label Guided Subspace Learning for Multi-view Data
Shudong Hou, Heng Liu 0002, Xiujun Wang
PRCV (3)2
2019 Deep Feature-Preserving Based Face Hallucination: Feature Discrimination Versus Pixels Approximation
Heng Liu 0002, Jungong Han, Shudong Hou
PRCV (2)2
2019 Survey on GAN-based face hallucination with its model development
abstract
Face hallucination aims to produce a high‐resolution face image from an input low‐resolution face image, which is of great importance for many practical face applications, such as face recognition and face verification. Since the structure of the face image is complex and sensitive, obtaining a super‐resolved face image is more difficult than generic image super‐resolution. Recently, with great success in the high‐level face recognition task, deep learning methods, especially generative adversarial networks (GANs), have also been applied to the low‐level vision task – face hallucination. This work is to provide a model evolvement survey on GAN‐based face hallucination. The principles of image resolution degradation and GAN‐based learning are presented firstly. Then, a comprehensive review of the state‐of‐art GAN‐based face hallucination methods is provided. Finally, the comparisons of these GAN‐based face hallucination methods and the discussions of the related issues for future research direction are also provided.
Heng Liu 0002, Jungong Han, Yuezhong Chu, Tao Tao 0005
IET Image Process.1
2019 Single image super-resolution using multi-scale deep encoder-decoder with phase congruency edge map guidance
Heng Liu 0002, Zilin Fu, Jungong Han, Ling Shao 0001, Shudong Hou, Yuezhong Chu
Inf. Sci.1
2019 Sparse regularized discriminative canonical correlation analysis for multi-view semi-supervised learning
Shudong Hou, Heng Liu 0002, Quan-Sen Sun
Neural Comput. Appl.2
2019 Are mid-air dynamic gestures applicable to user identification?
Heng Liu 0002, Liangliang Dai, Shudong Hou, Jungong Han, Hongshen Liu
Pattern Recognit. Lett.1
2018 Single image super-resolution using a deep encoder-decoder symmetrical network with iterative back projection
Heng Liu 0002, Jungong Han, Shudong Hou, Ling Shao 0001, Yue Ruan
Neurocomputing1
2018 Single satellite imagery simultaneous super-resolution and colorization using multi-task deep neural networks
Heng Liu 0002, Zilin Fu, Jungong Han, Ling Shao 0001, Hongshen Liu
J. Vis. Commun. Image Represent.1
2018 End-to-end video background subtraction with 3d convolutional neural networks
Dimitrios Sakkos, Heng Liu 0002, Jungong Han, Ling Shao 0001
Multim. Tools Appl.2
2017 Large size single image fast defogging and the real time video defogging FPGA architecture
Heng Liu 0002, Dongdong Huang, Shudong Hou, Yue Ruan
Neurocomputing1