VLDB 2026 Research / reviewers in the wild / expert
Hanling Zhang
dblp:34/5279
· DBLP profile ↗
32ranked-venue papers
7as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 10 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Face forgery detection via identification of evident tampered regions and multi-view analysisabstractIn AI-synthesized faces, there usually exist prominent natural features, which poses a huge challenge for face forgery detection. In this work, we propose a Region-Aware Deep Neural Network (RDNN). RDNN calculates the tampering possibility of each face region based on the features learned from each region and selects the region with the highest tampering possibility as the detection result. Then, a new Latent Cue Capture Loss (LCCL) is designed to train RDNN to capture those fake face features ignored by traditional loss functions. Besides, by leveraging RDNN to locate forgeries, we propose a deepfake detection strategy namely RDNN-based Multi-Perspective Deepfake Detection (RMDD), to keep the advantages of RDNN while improving the detection robustness. Specifically, RMDD uses RDNN to locate suspected forgeries in the original and horizontally flipped faces, and mines local features in the vicinity of these suspected forgeries. Finally, the detection result is acquired by integrating the above detection clues. Ablation experiments verify the ability of RDNN to locate manipulated traces and the contribution of each component in RMDD. Moreover, experimental results demonstrate that RMDD has excellent detection accuracy and generalization ability. Hanling Zhang, Gaobo Yang |
Neurocomputing | 2 |
| 2026 | Deepfake detection with dual-mode swin transformer: Multi-scale feature learning and local ambiguity mitigation
Gaobo Yang, Hanling Zhang |
J. Inf. Secur. Appl. | 3 |
| 2025 | Dlfr-Gen: Diffusion-Based Video Generation With Dynamic Latent Frame Rate
Zhihang Yuan, Yuzhang Shang, Hanling Zhang, Siyuan Wang 0002, Shengen Yan, Guohao Dai 0001, Yu Wang 0002 |
ICCV | 4 |
| 2025 | DiTFastAttnV2: Head-Wise Attention Compression for Multi-Modality Diffusion TransformersabstractText-to-image generation models, especially Multimodal Diffusion Transformers (MMDiT), have shown remarkable progress in generating high-quality images. However, these models often face significant computational bottlenecks, particularly in attention mechanisms, which hinder their scalability and efficiency. In this paper, we introduce DiTFastAttnV2, a post-training compression method designed to accelerate attention in MMDiT. Through an in-depth analysis of MMDiT's attention patterns, we identify key differences from prior DiT-based methods and propose head-wise arrow attention and caching mechanisms to dynamically adjust attention heads, effectively bridging this gap. We also design an Efficient Fused Kernel for further acceleration. By leveraging local metric methods and optimization techniques, our approach significantly reduces the search time for optimal compression schemes to just minutes while maintaining generation quality. Furthermore, with the customized kernel, DiTFastAttnV2 achieves a 68% reduction in attention FLOPs and 1.5x end-to-end speedup on 2K image generation without compromising visual fidelity. Hanling Zhang, Rundong Su, Zhihang Yuan, Pengtao Chen, Mingzhu Shen, Yibo Fan, Shengen Yan, Guohao Dai 0001, Yu Wang 0002 |
ICCV | 1 |
| 2025 | DLFR-VAE: Dynamic Latent Frame Rate VAE for Video GenerationabstractIn this paper, we propose the Dynamic Latent Frame Rate VAE (DLFR-VAE), a training-free paradigm that can make use of adaptive temporal compression in latent space. While existing video generative models apply fixed compression rates via pretrained VAE, we observe that real-world video content exhibits substantial temporal non-uniformity, with high-motion segments containing more information than static scenes. Based on this insight, DLFR-VAE dynamically adjusts the latent frame rate according to the content complexity. Specifically, DLFR-VAE comprises two core innovations: (1) a Dynamic Latent Frame Rate Scheduler that partitions videos into temporal chunks and adaptively determines optimal frame rates based on information-theoretic content complexity, and (2) a training-free adaptation mechanism that transforms pretrained VAE architectures to dynamic VAE that can process features with variable frame rates. Our simple but effective DLFR-VAE can function as a plug-and-play module, seamlessly integrating with existing video generation models and accelerating the video generation process. Zhihang Yuan, Siyuan Wang 0002, Yuzhang Shang, Hanling Zhang, Tongcheng Fang, Shengen Yan, Guohao Dai 0001, Yu Wang 0002 |
ACM Multimedia | 4 |
| 2025 | Decision variables to be discovered in modelling high-dimensional omics data for cancer studiesabstractHigh-dimensional omics data are often contaminated by sources of unwanted variations caused by platforms, batches, or other external factors. These interferences and noise can obscure critical signals related to cancer. Contaminated data are modeled as a combination of variables derived from the phenotype of interest (POI) and confounding factors. To identify these variables, a novel method called Decision Variable Analysis (DVA) is proposed. The novelty of DVA is to iteratively extract independent decisive variables for modeling the data. Specifically, a priori knowledge introduced as the definite variable linked with POI is removed from data through a residual operation. The number of variables is estimated from the residual matrix based on the zero gradient of singular values, rather than relying on random matrix theory or principal components analysis, which can produce unreliable results when the number of features exceeds the number of samples. Applications of DVA to both synthetic and real data demonstrate superior performance in identifying variables compared to conventional approaches. Improvements offered by DVA are illustrated across high-dimensional omics datasets, particularly those with smaller sample sizes relative to the number of features on different platforms. The results indicate that DVA is an effective method for dissecting sources of variation in high-dimensional data with disturbances. Weike Lu, Hanling Zhang, Jie Xie 0002 |
Intell. Data Anal. | 5 |
| 2024 | DiTFastAttn: Attention Compression for Diffusion Transformer ModelsabstractDiffusion Transformers (DiT) excel at image and video generation but face computational challenges due to the quadratic complexity of self-attention operators. We propose DiTFastAttn, a post-training compression method to alleviate the computational bottleneck of DiT.
We identify three key redundancies in the attention computation during DiT inference: (1) spatial redundancy, where many attention heads focus on local information; (2) temporal redundancy, with high similarity between the attention outputs of neighboring steps; (3) conditional redundancy, where conditional and unconditional inferences exhibit significant similarity. We propose three techniques to reduce these redundancies: (1) $\textit{Window Attention with Residual Sharing}$ to reduce spatial redundancy; (2) $\textit{Attention Sharing across Timesteps}$ to exploit the similarity between steps; (3) $\textit{Attention Sharing across CFG}$ to skip redundant computations during conditional generation. Zhihang Yuan, Hanling Zhang, Lu Pu, Xuefei Ning, Linfeng Zhang 0001, Tianchen Zhao, Shengen Yan, Guohao Dai 0001, Yu Wang 0002 |
NeurIPS | 2 |
| 2024 | RCNet: Related Context-Driven Network with Hierarchical Attention for Salient Object Detection
Chenxing Xia, Kuanching Li, Bin Ge 0001, Hanling Zhang |
Expert Syst. Appl. | 5 |
| 2024 | A real-time camera-based gaze-tracking system involving dual interactive modes and its application in gaming
He Zhang 0019, Hanling Zhang |
Multim. Syst. | 3 |
| 2024 | A two-stage fake face image detection algorithm with expanded attention
Hanling Zhang, Gaobo Yang, Zhiqing Guo, Jiyou Chen |
Multim. Tools Appl. | 2 |
| 2023 | A review of micro-expression spotting: methods and challenges
He Zhang 0019, Hanling Zhang |
Multim. Syst. | 3 |
| 2022 | TransFGU: A Top-Down Approach to Fine-Grained Unsupervised Semantic Segmentation
Zhaoyuan Yin, Pichao Wang, Fan Wang 0019, Hanling Zhang, Hao Li 0030, Rong Jin 0001 |
ECCV (29) | 5 |
| 2022 | Multi-Modality Diversity Fusion Network with Swintransformer for RGB-D Salient Object DetectionabstractMulti-modality complementary information brings new impetus and innovation to saliency object detection (SOD). However, most existing RGB-D SOD methods either indiscriminately handle RGB features and depth features or only take depth features as additional information of RGB subnet-work, ignoring the different roles of two modalities for SOD tasks. To tackle this issue, we propose a novel multi-modality diversity fusion network with SwinTransformer (M2DFNet) for RGB-D SOD from the perspective of the different status of multi-modality, which adequately explores the roles of RGB and depth modalities. To this end, a triple-diversity supervision mechanism (TDSM) and a diversity fusion module (DFM) are designed to parse the function of different modalities. Besides, we designed a dense decoder (DSD) to integrate multi-scale features and transfer gain information from top to bottom, which can improve the performance of SOD. Extensive experiments on five benchmark datasets demonstrate that the proposed M2DFNet outperforms 17 other state-of-the-art (SOTA) RGB-D SOD methods. Songsong Duan, Chenxing Xia, Xiuju Gao, Bin Ge 0001, Hanling Zhang, Kuanching Li |
ICIP | 5 |
| 2022 | Emcenet: Efficient Multi-Scale Context Exploration Network for Salient Object DetectionabstractMulti-scale context is crucial for the accurate salient object detection (SOD) in the real-world scenes. Although current contextual information-based SOD methods have achieved great progress, they may fail to generate precise saliency maps due to their seldom considering the correlation of different scale context during the extraction process. To address these issues, we propose an Efficient Multi-Scale Context Exploration Network (EMCENet) for SOD. Specifically, a progressive multi-scale context extraction (PMCE) module is designed to progressively capture strongly correlated multi-scale context by using multi-receptive-field convolution operations. Afterwards, a hierarchical feature hybrid interaction (HFHI) module is introduced to generate powerful feature representations by adaptively aggregating multi-level features in a hybrid interaction strategy. Extensive experimental results on six public datasets demonstrate that the proposed EMCENet method without any post-processing performs favorably against 13 state-of-the-art SOD methods. Chenxing Xia, Xiuju Gao, Bin Ge 0001, Hanling Zhang, Kuanching Li |
ICIP | 5 |
| 2022 | A Review of Micro-expression Recognition based on Deep LearningabstractMicro-expression has the characteristics of spontaneity, low intensity, and short duration, which reflects a real personal emotion. Therefore, micro-expression recognition (MER) has been applied widely in lie detection, depression analysis, human-computer interaction systems, and commercial negotiation. Micro-expressions usually occur when people attempt to cover up their true feelings, especially in high-stake environments. In the early stage, the study of micro-expressions was mainly from a psychological point of view and required a very specialized skill. MER based on deep learning is a hot research direction recently, which generally includes several stages, such as image preprocessing, feature extraction, and emotion classification. In this paper, we first introduce the problems and challenges MER encountered. Then we present the commonly used micro-expression datasets and methods of image preprocessing. Next, we describe the MER methods based on deep learning in recent years and classify them according to the network structure. Afterward, we present the evaluation metrics and protocol and compare different algorithms on the composite dataset. Finally, we conclude and provide a prospect of the future work of MER. He Zhang 0019, Hanling Zhang |
IJCNN | 2 |
| 2022 | RAMT-GAN: Realistic and accurate makeup transfer with generative adversarial network
Qiang-Lin Yuan, Hanling Zhang |
Image Vis. Comput. | 2 |
| 2022 | HDNet: Multi-Modality Hierarchy-Aware Decision Network for RGB-D Salient Object DetectionabstractRGB-D Salient object detection (SOD) is a pixel-level dense prediction task, which can highlight the prominent object in the scene. Recently, Convolution Neural Network (CNN) is widely applied in SOD to generate multi-level features, which are complementary to each other. However, most methods ignore the unique characteristics of multi-level features (high-level and low-level features). Given the effective employment of multi-level features, we propose a novel multi-modality hierarchy-aware decision network (HDNet) by embedding a Swin Transformer as an encoder. The proposed HDNet contains three primary designs: (1) a Swin Transformer encoder is employed instead of a CNN to learn long-range dependencies; (2) a hierarchy-aware feature decision mechanism (HFDM) is proposed to exploit effective local detail cues of low-level features and global semantic information of high-level features, which consists of two sub-modules, namely low-hierarchy edge module (LEM) and high-hierarchy region module (HRM); (3) a decision-based fusion module (DFM) is designed to fuse RGB and depth features under the attribute of multi-level features generated from HFDM. Experiments on five public benchmarks verify that our framework has better performance than the other 18 state-of-the-art algorithms. Chengxing Xia, Songsong Duan, Bin Ge 0001, Hanling Zhang, Kuanching Li |
IEEE Signal Process. Lett. | 4 |
| 2022 | A Low-Rank Tensor Decomposition Model With Factors Prior and Total Variation for Impulsive Noise RemovalabstractImage restoration is a long-standing problem in signal processing and low-level computer vision. Previous studies have shown that imposing a low-rank Tucker decomposition (TKD) constraint could produce impressive performances. However, the TKD-based schemes may lead to the overfitting/underfitting problem because of incorrectly predefined ranks. To address this issue, we prove that the$n$-rank is upper bounded by the rank of each Tucker factor matrix. Using this relationship, we propose a formulation by imposing the nuclear norm regularization on the latent factors of TKD, which can avoid the burden of rank selection and reduce the computational cost when dealing with large-scale tensors. In this formulation, we adopt the Minimax Concave Penalty to remove the impulsive noise instead of the$l_{1}$-norm which may deviate from both the data-acquisition model and the prior model. Moreover, we employ an anisotropic total variation regularization to explore the piecewise smooth structure in both spatial and spectral domains. To solve this problem, we design the symmetric Gauss-Seidel (sGS) based alternating direction method of multipliers (ADMM) algorithm. Compared to the directly extended ADMM, our algorithm can achieve higher accuracy since more structural information is utilized. Finally, we conduct experiments on the three kinds of datasets, numerical results demonstrate the superiority of the proposed method, especially, the average PSNR of the proposed method can improve about 1~5dB for each noise level of color images. Xin Tian 0008, Kun Xie 0001, Hanling Zhang |
IEEE Trans. Image Process. | 3 |
| 2022 | Moment is Important: Language-Based Video Moment Retrieval via Adversarial LearningabstractThe newly emerging language-based video moment retrieval task aims at retrieving a target video moment from an untrimmed video given a natural language as the query. It is more applicable in reality since it is able to accurately localize a specific video moment, as compared to traditional whole video retrieval. In this work, we propose a novel solution to thoroughly investigate the language-based video moment retrieval issue under the adversarial learning. The key of our solution is to formulate the language-based video moment retrieval task as an adversarial learning problem with two tightly connected components. Specifically, a reinforcement learning is employed as a generator to produce a set of possible video moments. Meanwhile, a multi-task learning is utilized as a discriminator, which integrates inter-modal and intra-modal in a unified framework by employing a sequential update strategy. Finally, the generator and the discriminator are mutually reinforced in the adversarial learning, which is able to jointly optimize the performance of both video moment ranking and video moment localization. Extensive experimental results on two challenging benchmarks, i.e., Charades-STA and TACoS datasets, have well demonstrated the effectiveness and rationality of our proposed solution. Meanwhile, on the larger and unbiased datasets, i.e., ActivityNet Captions and ActivityNet-CD, our proposed framework exhibits excellent robustness. Yawen Zeng, Da Cao, Shaofei Lu, Hanling Zhang, Jiao Xu 0001, Zheng Qin 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | Learning To Recommend Frame for Interactive Video Object Segmentation in the WildabstractThis paper proposes a framework for the interactive video object segmentation (VOS) in the wild where users can choose some frames for annotations iteratively. Then, based on the user annotations, a segmentation algorithm refines the masks. The previous interactive VOS paradigm selects the frame with some worst evaluation metric, and the ground truth is required for calculating the evaluation metric, which is impractical in the testing phase. In contrast, in this paper, we advocate that the frame with the worst evaluation metric may not be exactly the most valuable frame that leads to the most performance improvement across the video. Thus, we formulate the frame selection problem in the interactive VOS as a Markov Decision Process, where an agent is learned to recommend the frame under a deep reinforcement learning framework. The learned agent can automatically determine the most valuable frame, making the interactive setting more practical in the wild. Experimental results on the public datasets show the effectiveness of our learned agent without any changes to the underlying VOS algorithms. Our data, code, and models are available at https://github.com/svip-lab/IVOS-W. Zhaoyuan Yin, Jia Zheng 0002, Weixin Luo, Shenhan Qian, Hanling Zhang, Shenghua Gao |
CVPR | 5 |
| 2020 | Occlusion-handling tracker based on discriminative correlation filtersabstractVisual object tracking (VOT) based on discriminative correlation filters (DCF) has received great attention due to its higher computational efficiency and better robustness. However, DCF‐based methods suffer from the problem of model contamination. The tracker will drift into the background due to the uncertainties brought by shifting among peaks, which will further lead to the issues of model degradation. To deal with occlusions, a novel Occlusion‐Handling Tracker Based on Discriminative Correlation Filters (OHDCF) framework is proposed for online visual object tracking, where an occlusion‐handling strategy is integrated into the spatial–temporal regularized correlation filters (STRCF). The occlusion‐handling tracker follows a hybrid approach to handle partial occlusion and complete occlusion. Specifically, we first present a function to determine whether occlusion occurs. Then, the proposed filter uses block‐based and feature‐matching methods to determine whether an object is partially occluded or completely occluded. Following this, we use different methods to track the target. Extensive experiments have performed on OTB‐100, Temple‐Color‐128, VOT‐2016 and VOT‐2018 datasets, the results show that OHDCF achieves promising performance compared to other state‐of‐the‐art trackers. On VOT‐2018, OHDCF significantly outperforms STRCF from the challenge with a relative gain of 4.8 in EAO and a gain of 4.6 in Accuracy. Hanling Zhang |
IET Image Process. | 2 |
| 2020 | Exploiting background divergence and foreground compactness for salient object detection
Chenxing Xia, Hanling Zhang, Xiuju Gao, Keqin Li 0001 |
Neurocomputing | 2 |
| 2019 | Video-Based Cross-Modal Recipe RetrievalabstractAs a natural extension of image-based cross-modal recipe retrieval, retrieving a specific video given a recipe as the query is seldom explored. There are various temporal and spatial elements hidden in cooking videos. In addition, current image-based cross-modal recipe retrieval approaches mostly emphasize the understanding of textual and visual content independently. Such methods overlook the interaction between textual and visual content. In this work, we innovatively propose a new problem of video-based cross-modal recipe retrieval and thoroughly investigate this issue under the attention paradigm. In particular, we firstly exploit a parallel-attention network to independently learn the representations of videos and recipes. Next, a co-attention network is proposed to explicitly emphasize the cross-modal interactive features between videos and recipes. Meanwhile, a cross-modal fusion sub-network is proposed to learn both the independent and collaborative dynamics, which can enhance the associated representation of videos and recipes. Last but not the least, the embedding vectors of videos and recipes stemming from joint network are optimized with a pairwise ranking loss. Extensive experiments on a self-collected dataset have verified the effectiveness and rationality of our proposed solution. Da Cao, Zhiwang Yu, Hanling Zhang, Jiansheng Fang, Liqiang Nie, Qi Tian 0001 |
ACM Multimedia | 3 |
| 2019 | Action recognition based on multi-stage jointly training convolutional network
Hanling Zhang, Chenxing Xia, Xiuju Gao |
Multim. Tools Appl. | 1 |
| 2018 | Saliency detection by aggregating complementary background template with foreground informationabstractThis paper proposes an unsupervised bottom-up saliency detection approach by exploiting novel background template and foreground information. First, a discriminative feature vector is extracted from each super-pixel to cover regional color, contrast and texture information. Then we apply it to get a background based saliency map based on a background template. In order to get more accurate saliency map, we select highly confident compact foreground seeds to compute a foreground based saliency map. After fusing the two saliency maps, the integrated map is refined to achieve the final result. Experimental results show that the proposed algorithm generates high-quality saliency maps against the state-off-the-art saliency detection methods on four publicly available datasets. Hanling Zhang, Chenxing Xia, Jianhua Cui |
CASA | 1 |
| 2017 | Robust saliency detection via corner information and an energy functionabstractIn this study, the authors propose a distinctive bottom‐up visual saliency detection algorithm based on a new background prior and a new reinforcement. Inspired by genetic algorithm, the final map is obtained with three steps. First of all, the authors construct a background‐based saliency map by manifold ranking via superior image corners selected by convex‐hull as background prior, which is different from most of the existing background prior‐based methods treated all image boundaries as background. Then, a better result is obtained by ranking the relevance of the image elements with foreground seeds extracted from the preliminary saliency map. Furthermore, a novel optimisation framework is introduced with the intention of refining the map, which integrates an energy function with a guided filter. Experimental results on three public datasets indicate that the proposed method performs favourably against the state‐of‐the‐art algorithms. Hanling Zhang, Chenxing Xia, Xiuju Gao |
IET Comput. Vis. | 1 |
| 2017 | Combining depth-skeleton feature with sparse coding for action recognition
Hanling Zhang, Ping Zhong 0001, Jiale He, Chenxing Xia |
Neurocomputing | 1 |
| 2017 | Combining multi-layer integration algorithm with background prior and label propagation for saliency detection
Chenxing Xia, Hanling Zhang, Xiuju Gao |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Aggregating complementary boundary contrast with smoothing for salient region detection
Ruihui Li, Jianrui Cai, Hanling Zhang, Taihong Wang |
Vis. Comput. | 3 |
| 2016 | A novel optimization framework for salient object detection
Hanling Zhang, Liyuan Zhuo, Vincent Havyarimana |
Vis. Comput. | 1 |
| 2015 | Robust visual tracking based on structured sparse representation model
Hanling Zhang, Gaobo Yang |
Multim. Tools Appl. | 1 |
| 2015 | Saliency detection with color contrast based on boundary information and neighbors
Hanling Zhang |
Vis. Comput. | 2 |