EDBT 2026 Demo / reviewers in the wild / expert
Zheng He 0001
dblp:61/678-1
· DBLP profile ↗
27ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0002-7700-0901ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SVRMAE: Enhancing surveillance video super-resolution through separation masking and MAE pretraining
Zheng He 0001, Gang Ye, Wenqian Zhu |
Neurocomputing | 2 |
| 2026 | OACI: Object-aware contextual integration for image captioning
Shuhan Xu, Mengya Han, Wei Yu 0004, Zheng He 0001, Xin Zhou 0003, Yong Luo 0002 |
Knowl. Based Syst. | 4 |
| 2025 | MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-ReadingabstractLip-reading is to utilize the visual information of the speaker’s lip movements to recognize words and sentences. Existing event-based lip-reading solutions integrate different frame rate branches to learn spatio-temporal features of varying granularities. However, aggregating events into event frames inevitably leads to the loss of fine-grained temporal information within frames. To remedy this drawback, we propose a novel framework termed Multi-view Temporal Granularity aligned Aggregation (MTGA). Specifically, we first present a novel event representation method, namely time-segmented voxel graph list, where the most significant local voxels are temporally connected into a graph list. Then we design a spatio-temporal fusion module based on temporal granularity alignment, where the global spatial features extracted from event frames, together with the local relative spatial and temporal features contained in voxel graph list are effectively aligned and integrated. Finally, we design a temporal aggregation module that incorporates positional encoding, which enables the capture of local absolute spatial and global temporal information. Experiments demonstrate that our method outperforms both the event-based and video-based lip-reading counterparts. Yong Luo 0002, Wei Yu 0004, Zheng He 0001, Jialie Shen 0001 |
AAAI | 6 |
| 2025 | ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language ModelsabstractXuxu Liu, Siyuan Liang, Mengya Han, Yong Luo, Aishan Liu, Xiantao Cai, Zheng He, Dacheng Tao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xuxu Liu, Siyuan Liang 0004, Mengya Han, Yong Luo 0002, Aishan Liu, Xiantao Cai, Zheng He 0001, Dacheng Tao |
ACL (1) | 7 |
| 2025 | EgoNet: An Unified Egocentric Active Speaker Detection Framework for both Camera Wearer and Visible CandidatesabstractActive Speaker Detection (ASD) aims to determine whether each candidate in a video frame is speaking. The egocentric dataset Ego4D introduces unique challenges for this task, such as dynamic shooting angles that cause candidates to frequently leave the sight, leading to temporal discontinuities. Additionally, Ego4D poses a novel task: detecting the speaking activities of the camera wearer, who never appears in the field of view. Existing methods treat these two tasks separately, and treat candidates out of sight as noise. In contrast, we propose EgoNet, a framework that uniformly models all candidates, including those not visible. By capturing interactions among all candidates and modeling broader temporal context, EgoNet reduces uncertainty and improves performance in egocentric active speaker detection. Yongqian Li, Xin Zhou 0003, Zheng He 0001, Wei Yu 0004, Yong Luo 0002 |
ICASSP | 3 |
| 2025 | A Dual Stream Visual Tokenizer for LLM Image GenerationabstractWe proposes a novel visual tokenizer by combining high-level semantic tokens and low-level pixel tokens to represent images, aiming to address the challenges of image-to-sequence conversion for Large Language Models (LLMs). Existing visual tokenizers, such as VQ-VAE and diffusion-based models, either struggle with token explosion as image resolution increases or fail to capture detailed structural information. Our method introduces a dual-token system: high-level semantic tokens capture the main content of the image, while low-level pixel tokens preserve structural details. By integrating these tokens in a hybrid architecture, we leverage a VQ-VAE branch to generate low-resolution guidance and a diffusion process to reconstruct high-resolution images with both semantic coherence and structural accuracy. This approach significantly reduces the number of required tokens and enhances image reconstruction quality, offering an efficient solution for tasks like image generation and understanding based on LLMs. Yongqian Li, Yong Luo 0002, Xiantao Cai, Zheng He 0001, Zhennan Meng, Nidong Wang, Yunlin Chen |
IJCAI | 4 |
| 2025 | Open-Vocabulary Fine-Grained Hand Action DetectionabstractIn this work, we address the new challenge of open-vocabulary fine-grained hand action detection, which aims to recognize hand actions from both known and novel categories using textual descriptions. Traditional hand action detection methods are limited to closed-set detection, making it difficult for them to generalize to new, unseen hand action categories. While current open-vocabulary detection (OVD) methods are effective at detecting novel objects, they face challenges with fine-grained action recognition, particularly when data is limited and heterogeneous. This often leads to poor generalization and performance bias between base and novel categories. To address these issues, we propose a novel approach, Open-FGHA (Open-vocabulary Fine-Grained Hand Action), which learns to distinguish fine-grained features across multiple modalities from limited heterogeneous data. It then identifies optimal matching relationships among these features, enabling accurate open-vocabulary fine-grained hand action detection. Specifically, we introduce three key components: Hierarchical Heterogeneous Low-Rank Adaptation, Bidirectional Selection and Fusion Mechanism, and Cross-Modality Query Generator. These components work in unison to enhance the alignment and fusion of multimodal fine-grained features. Extensive experiments demonstrate that Open-FGHA outperforms existing OVD methods, showing its strong potential for open-vocabulary hand action detection. The source code is available at OV-FGHAD. Ting Zhe, Mengya Han, Xiaoshuai Hao, Yong Luo 0002, Zheng He 0001, Xiantao Cai, Jing Zhang 0037 |
IJCAI | 5 |
| 2025 | DualRCS-BEV: Dynamic RCS Modeling with Dual-Branch Guidance for Efficient 3D Object Detection via Radar-Camera BEV Fusion
Zheng He 0001, Gang Ye, Wenqian Zhu |
PRCV (11) | 2 |
| 2024 | Advancing Surveillance Video Clarity and Transmission: A Real-Time Video Super-Resolution Model with Background Information Awareness
Zheng He 0001, Gang Ye, Wenqian Zhu |
PRCV (9) | 2 |
| 2024 | Bidirectional scale-aware upsampling network for arbitrary-scale video super-resolution
Laigan Luo, Benshun Yi, Zhongyuan Wang 0001, Zheng He 0001 |
Image Vis. Comput. | 4 |
| 2024 | Efficient lightweight network for video super-resolution
Laigan Luo, Benshun Yi, Zhongyuan Wang 0001, Peng Yi 0002, Zheng He 0001 |
Neural Comput. Appl. | 5 |
| 2024 | Omniscient Video Super-Resolution with Explicit-Implicit AlignmentabstractWhen considering the temporal relationships, most previous video super-resolution (VSR) methods follow the iterative or recurrent framework. The iterative framework adopts neighboring low-resolution (LR) frames from a sliding window, while the recurrent framework utilizes the output generated in the previous SR procedure. The hybrid framework combines them but still cannot fully leverage the temporal relationships. Meanwhile, the existing methods are limited in the receptive field of the optical flow or lack semantic constrains on motion information. In this work, we propose an omniscient framework to fully explore the temporal relationships in the video, which encompasses both LR frames and SR outputs from the past, present, and future. The omniscient framework is more generic because the iterative, recurrent, and hybrid frameworks can be regarded as its special cases. Besides, when addressing the motion information, most previous VSR methods adopt the explicit motion estimation and compensation, while many recent methods turn to implicit alignment. In implicit alignment methods, because basic non-local means suffers from heavy computational costs, we improve it by capturing the non-local correlations in a relatively local manner to reduce the complexity. Moreover, we integrate the explicit and implicit methods into an explicit-implicit alignment module to better utilize motion information. We have conducted extensive experiments on public datasets, which show that our method is superior over the state-of-the-art methods in objective metrics, subjective visual quality, and complexity. In particular, on datasets of Vid4 and UDM10, our method improves PSNR by 0.19 dB, 0.49 dB against the most advanced method BasicVSR++, respectively. Peng Yi 0002, Zhongyuan Wang 0002, Laigan Luo, Kui Jiang, Zheng He 0001, Junjun Jiang, Tao Lu 0001, Jiayi Ma 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Transferring fashion to surveillance with weak labels
Zheng He 0001, Chao Liang 0001, Jun Chen 0001, Chia-Wen Lin, Dapeng Tao |
Neural Comput. Appl. | 2 |
| 2022 | ECL: Exclusive Curriculum Learning for Video Super-ResolutionabstractVideo super-resolution (VSR) problem has gained a soaring development along with deep learning methods. However, the further progress requires the blessing of more complex architectures. Unlike them, this paper promotes VSR performance from a new perspective of sample difficulty. We propose an exclusive curriculum learning strategy for VSR, which can improve the representation power without noticeable computation increment. Specifically, this paper memorizes the performance track of every sample and calculate a customized weight for each sample according to it. In this way, the model can automatically concentrate on the easy samples first and gradually focus on the hard ones. Experimental analysis on training process and benchmark datasets demonstrate that our method can substantially boost the performance with a superior convergence speed and a limited number of parameters. Sicheng Hu, Zhongyuan Wang 0001, Peng Yi 0002, Zheng He 0001, Jinsheng Xiao, Jing Xiao 0004 |
ICME | 4 |
| 2022 | Face hallucination based on degradation analysis for robust manifold
Ruimin Hu, Zheng He 0001, Chao Liang 0001, Zhongyuan Wang 0001 |
Neurocomputing | 3 |
| 2022 | Two-stage unsupervised facial image quality measurement
Guangcheng Wang, Zhongyuan Wang 0001, Baojin Huang, Kui Jiang, Zheng He 0001, Hancheng Zhu, Jinsheng Xiao, Xin Tian 0006 |
Inf. Sci. | 5 |
| 2022 | Rethinking Lightweight: Multiple Angle Strategy for Efficient Video Action RecognitionabstractVideo action recognition task involves modeling spatiotemporal information, and efficiency is critical to capture spatiotemporal dependencies in the video. Most existing models rely on optical flow information to capture the dynamic visual tempos between consecutive video frames. Although impressive performance can be achieved by combining optical flow with RGB, the time-consuming nature of optical flow computation cannot be ignored. Moreover, 3D CNN has successfully modeled spatiotemporal information, yet the enormous computational volume is unsuitable for real-time action recognition. In this letter, we propose a novel lightweight video feature extraction strategy that achieves better recognition performance with lower FLOPs. In particular, we perform convolution on the video cube from three orthogonal angles to learn its appearance and motion features. Compared with the computational volume of 3D CNN, our proposed method is more economical and thus meets the lightweight requirements. Extensive experimental results on public Something Something-V1$\&$V2 and Diving48 datasets show our approach achieves the state-of-the-art performance. Jianyu Chen 0008, Zhongyuan Wang 0001, Kangli Zeng, Zheng He 0001, Zixiang Xiong |
IEEE Signal Process. Lett. | 4 |
| 2022 | Reference-Free DIBR-Synthesized Video Quality Metric in Spatial and Temporal DomainsabstractDepth image-based rendering (DIBR) techniques play an important role in free viewpoint videos (FVVs), which have a wide range of applications including immersive entertainment, remote monitoring, education, etc. FVVs are usually synthesized by DIBR techniques in a “blind” environment (without a reference video). Thus, an effective reference-free synthesized video quality assessment (VQA) metric is vital. At present, many image quality assessment (IQA) algorithms for DIBR-synthesized images have been proposed, but limited researches have been concerned about the quality assessment of DIBR-synthesized videos. To this end, this paper proposes a novel reference-free VQA method for synthesized videos, which operates in Spatial and Temporal Domains, dubbed as STD. The design fundamental of the proposed STD metric considers the effects of two major distortions introduced by DIBR techniques on the visual quality of synthesized videos. First, considering the geometric distortion introduced by DIBR technologies can increase high-frequency contents of the synthesized frame, the influence of the geometric distortion on the visual quality of a synthesized video can be effectively evaluated by estimating high-frequency energies of each synthesized frame in spatial domain. Second, temporal inconsistency caused by DIBR techniques brings the temporal flicker distortion, which is one of the most annoying artifacts in DIBR-synthesized videos. In temporal domain, we quantify temporal inconsistency by measuring motion differences between consecutive frames. Specifically, optical flow method is first used to estimate the motion field between adjacent frames. Then, we calculate the structural similarity of adjacent optical flow fields and further adopt the structural similarity value to weight the pixel differences of adjacent optical flow fields. Experiments show that the above two features are able to well perceive the visual quality of DIBR-synthesized videos. Furthermore, since the two features are extracted from spatial and temporal domains, respectively, we integrate them using a linear weighting strategy to obtain our STD metric, which proves advantageous over two components and the competing state-of-the-art I/VQA methods. The source code is available athttps://github.com/wgc-vsfm/DIBR-video-quality-assessment. Guangcheng Wang, Zhongyuan Wang 0001, Ke Gu 0001, Kui Jiang, Zheng He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Fast and Robust Loop-Closure Detection via Convolutional Auto-Encoder and Motion ConsensusabstractLoop-closure detection is an indispensable module in the visual simultaneous localization and mapping (vSLAM) system. It typically consists of three main steps: image representation, loop-closure candidate selection, and loop-closure event verification. This article proposes a novel approach for loop-closure detection. In particular, we first introduce a lightweight convolutional auto-encoder network trained by the deep perceptual similarity loss for image representation. We then propose an image-to-sequence selection approach based on place sequence division and distance-weighted voting for loop-closure candidate selection. Furthermore, we propose a motion vector consensus constraint to improve locality preserving matching, which can be used for efficient loop-closure event verification that is robust for various complex environments. Extensive experiments have been conducted on four publicly available datasets. The results demonstrate that our method is able to achieve better recall performance than the state-of-the-art and meet the real-time requirement of vSLAM systems. Jiayi Ma 0001, Shenyue Wang, Kaining Zhang, Zheng He 0001, Jun Huang 0008, Xiaoguang Mei |
IEEE Trans. Ind. Informatics | 4 |
| 2021 | A Tilt-Angle Face Dataset And Its ValidationabstractSince the surveillance cameras are usually mounted at a high position to overlook targets, tilt-angle faces on overhead view are common in the public video surveillance environment. Face recognition approaches based on deep learning models have achieved excellent performance, but there remains a large gap for the overlooking surveillance scenarios. The results of face recognition depend not only on the structure of the model, but also on the completeness and diversity of the training samples. The existing multi-pose face datasets do not cover complete top-view face samples, and the models trained by them thus cannot provide satisfactory accuracy. To this end, this paper pioneers a multi-view tilt-angle face dataset (TFD), which is collected with an elaborately devised overhead capture equipment. TFD contains 11,124 face images from 927 subjects, covering a variety of tilt angles on the overhead view. To verify the validity of the constructed dataset, we further conduct comprehensive face detection and recognition experiments using the corresponding models trained by WiderFace, Webface and our TFD, respectively. Experimental results show that our TFD substantially promotes the face detection and recognition accuracy under the top-view situation. TFD is available at https://github.com/huang1204510135/D FD. Nanxi Wang, Zhongyuan Wang 0001, Zheng He 0001, Baojin Huang, Liguo Zhou, Zhen Han 0002 |
ICIP | 3 |
| 2021 | Silicone mask face anti-spoofing detection based on visual saliency and facial motion
Guangcheng Wang, Zhongyuan Wang 0001, Kui Jiang, Baojin Huang, Zheng He 0001, Ruimin Hu |
Neurocomputing | 5 |
| 2020 | Masked Face Recognition with Identification AssociationabstractIn the crime scene, criminals often consciously conceal their facial identity through face-masked disguise, which poses a huge challenge to identity recognition. Existing disguised face recognition techniques aiming for light even slight occlusions are completely invalid for face-masked identification. To this end, this paper proposes a masked face recognition method based on person re-identification association, which converts the masked face recognition problem into an association uncovering problem between the masked face and the appearing faces of the same person. Based on the characteristics that person re-identification technique does not rely solely on facial information, it first takes advantages of re-identification to establish the association between face-masked pedestrians and face-unveiled pedestrians. It further provides an effective face image quality assessment to select the most identifiable faces for subsequent recognition from a variety of appearing candidate faces. Finally, the selected high-quality recognizable faces are used to replace masked faces for identification. The comparison experiments with the existing disguise face recognition methods show its superiority in terms of accuracy. Zhongyuan Wang 0001, Zheng He 0001, Nanxi Wang, Xin Tian 0006, Tao Lu 0001 |
ICTAI | 3 |
| 2020 | Lightweight Progressive Residual Clique Network for Image Super-ResolutionabstractDeeper and wider convolutional neural networks (CNN) hava been widely applied to the single image super-resolution (SR) task for its appealing performance. However, enormous parametric memory footprint hinders its real-time application on mobile devices, especially in the energy-sensitive environment. In this work, we take both the reconstruction performance and efficiency into consideration and propose a lightweight progressive residual clique network (PRCN) for image SR. PRCN is built on the two-stage residual channel separation block (RCSB) and long-skip connections. First, we divide the input into four channel groups to differently learn texture details, immediately followed by a primary fusion to establish cross-channel correspondence in the first stage. Then we perform a further fusion on the outputs of the first stage to constitute a clique for the refinement in the second stage. Meanwhile, we employ SENet to improve the outputs of the second stage with the separate features of the first stage. This design not only enforces the correlation across channels, but also allows fewer densely connected blocks. Experimental results on public datasets show that PRCN outperforms state-of-the-art methods in terms of performance and complexity. Baojin Huang, Zheng He 0001, Zhongyuan Wang 0001, Kui Jiang, Guangcheng Wang |
ICTAI | 2 |
| 2020 | Low-quality watermarked face inpainting with discriminative residual learningabstractMost existing image inpainting methods assume that the location of the repair area (watermark) is known, but this assumption does not always hold. In addition, the actual watermarked face is in a compressed low-quality form, which is very disadvantageous to the repair due to compression distortion effects. To address these issues, this paper proposes a low-quality watermarked face inpainting method based on joint residual learning with cooperative discriminant network. We first employ residual learning based global inpainting and facial features based local inpainting to render clean and clear faces under unknown watermark positions. Because the repair process may distort the genuine face, we further propose a discriminative constraint network to maintain the fidelity of repaired faces. Experimentally, the average PSNR of inpainted face images is increased by 4.16dB, and the average SSIM is increased by 0.08. TPR is improved by 16.96% when FPR is 10% in face verification. Zheng He 0001, Xueli Wei, Kangli Zeng, Zhen Han 0002, Qin Zou 0001, Zhongyuan Wang 0001 |
MMAsia | 1 |
| 2020 | Ultra-dense GAN for satellite imagery super-resolution
Zhongyuan Wang 0001, Kui Jiang, Peng Yi 0002, Zhen Han 0002, Zheng He 0001 |
Neurocomputing | 5 |
| 2019 | Identifying Users by Asynchronous Mobility TrajectoriesabstractWith the popularity of location-based services and applications, a large amount of mobility data has been generated. Identity recognition through mobile trajectory information, especially asynchronous trajectory data has arisen great concerns in social security prevention and control. This paper advocates an identification resolution method based on the most frequently distributed TOP-N regions regarding user trajectories. This method first finds TOP-N regions whose trajectory points are most frequently distributed so as to reduce the computational complexity. It then combines probabilistic deviation and angle cosine to calculate TOP-N region similarity between two trajectories to identify the same user. We conducted extensive experiments on two real GPS trajectory datasets GeoLife and Cabspotting and comprehensively discussed the experimental results. The experimental results show that this method is substantially effective and efficiency for user identification. Mengjun Qi, Zhongyuan Wang 0001, Zheng He 0001, Tao Lu 0001 |
IGARSS | 3 |
| 2010 | Shared Virtual Presentation Board for e-Communication on the WebELS PlatformabstractIn this paper, a shared virtual presentation board (VPB) for e-Communication application on the Web-based e-Learning System (WebELS) platform is introduced. WebELS is a general-purpose e-Learning system to support flexibility and globalization of higher education in science and technology. In WebELS, the Meeting module consists of online slide presentation and video meeting, which the combination of both creates a so-called virtual room for e-Communication applications where meeting participants convene via the Internet. Online presentation features synchronized remote control for scrolling function, zooming function, cursor movement, shifting of slides back and forth, and even controlling the playback of video embedded on the slide. It also features online annotation to enhance the versatility and usefulness of online presentation. Online annotation allows the presenter to overwrite figures, draw objects or write mathematical equations to further elaborate what is being presented in a synchronized manner. This paper discusses the features of online presentation and the development of a virtual presentation board (VPB). VPB is a shared object that resides on the WebELS server system and is periodically accessed by WebELS client system in order that attendee’s presentation viewer can replicate that of the presenter. Arjulie John Berena, Zheng He 0001, Pao Sriprasertsuk, Sila Chunwijitra, Eiji Okano, Haruki Ueno |
ICCE | 2 |