Dan Zeng 0002

dblp:06/6575-2 · DBLP profile ↗
← Back
38ranked-venue papers
8as first author
32since 2021 · last 2026
0000-0002-9036-7791ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 17 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CapeNext: Rethinking and Refining Dynamic Support Information for Category-Agnostic Pose Estimation
abstract
Recent research in Category-Agnostic Pose Estimation (CAPE) has adopted fixed textual keypoint description as semantic prior for two-stage pose matching frameworks. While this paradigm enhances robustness and flexibility by disentangling the dependency of support images, our critical analysis reveals two inherent limitations of static joint embedding: (1) polysemy-induced cross-category ambiguity during the matching process(e.g., the concept "leg" exhibiting divergent visual manifestations across humans and furniture), and (2) insufficient discriminability for fine-grained intra-category variations (e.g., posture and fur discrepancies between a sleeping white cat and a standing black cat). To overcome these challenges, we propose a new framework that innovatively integrates hierarchical cross-modal interaction with dual-stream feature refinement, enhancing the joint embedding with both class-level and instance-specific cues from textual description and specific images. Experiments on the MP-100 dataset demonstrate that, regardless of the network backbone, CapeNext consistently outperforms state-of-the-art CAPE methods by a large margin.
Dan Zeng 0002, Shuiwang Li, Qijun Zhao, Qiaomu Shen, Bo Tang 0016
AAAI2
2026 Learning motion blur robust vision transformers for real-time UAV tracking
You Wu 0009, Xucheng Wang, Dan Zeng 0002, Hengzhou Ye, Xiaolan Xie 0002, Qijun Zhao, Shuiwang Li
Expert Syst. Appl.3
2026 SMTrack: End-to-End Trained Spiking Neural Networks for Multi-Object Tracking in RGB Videos
abstract
Brain-inspired Spiking Neural Networks (SNNs) leverage a sparse, event-driven computational paradigm and have shown great potential for low-power object tracking. However, most existing SNN-based object tracking rely on event camera data, whereas traditional RGB video remains the dominant input modality in real-world applications. Research on RGB-based SNN multi-object tracking, particularly directly trained deep SNN models, is still in its infancy. To address this, we propose SMTrack, the first directly trained deep SNN framework for end-to-end multi-object tracking on standard RGB data. To handle the challenges caused by scale and density variations among objects, we introduce an Adaptive Scale-aware Normalized Wasserstein Distance Loss (Asa-NWDLoss), which dynamically adjusts the normalization factor based on the average object size within each training batch. For the identity association stage, we integrate the TrackTrack to maintain robust and consistent trajectory tracking. Extensive experiments on BEE24, MOT17, MOT20, and DanceTrack demonstrate that SMTrack achieves comparable performance to mainstream ANN-based approaches with only a few time steps. Codes is available at https://github.com/OpenCodeGithub/SMTrack.
Pengzhi Zhong, Dan Zeng 0002, Qihua Zhou, Feixiang He, Shuiwang Li
IEEE Internet Things J.3
2026 UniTrack: Unifying day and night tracking with continual learning
Jiwei Mo, Feixiang He, Pengzhi Zhong, Qijun Zhao, Dan Zeng 0002, Shuiwang Li, Xianhao Shen
Inf. Sci.5
2026 Learning an Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking
abstract
Transformer-based models have improved visual tracking, but most still cannot run in real time on resource-limited devices, especially for unmanned aerial vehicle (UAV) tracking. To achieve a better balance between performance and efficiency, we propose AVTrack, an adaptive computation tracking framework that adaptively activates transformer blocks through an Activation Module (AM), which dynamically optimizes the ViT architecture by selectively engaging relevant components. To address extreme viewpoint variations, we propose to learn view-invariant representations via mutual information (MI) maximization. In addition, we propose AVTrack-MD, an enhanced tracker incorporating a novel MI maximization-based multi-teacher knowledge distillation framework. Leveraging multiple off-the-shelf AVTrack models as teachers, we maximize the MI between their aggregated softened features and the corresponding softened feature of the student model, improving the generalization and performance of the student, especially under noisy conditions. Extensive experiments show that AVTrack-MD achieves performance comparable to AVTrack’s performance while reducing model complexity and boosting average tracking speed by over 17%. Codes is available at https://github.com/wuyou3474/AVTrack.
You Wu 0009, Xucheng Wang, Xiangyang Yang 0001, Hengzhou Ye, Dan Zeng 0002, Qijun Zhao, Shuiwang Li
IEEE Trans. Circuits Syst. Video Technol.7
2026 Toward Real-Time UAV Tracking With Adaptive and Background-Aware Vision Transformers
abstract
In Unmanned Aerial Vehicle (UAV) tracking, discriminative correlation filters (DCF) are popular for their speed, making them ideal for real-time use with limited resources. Recently, lightweight convolutional neural networks (CNNs) have offered a new approach. Through filter pruning, these CNNs maintain high accuracy and efficiency, making them a strong alternative, especially for greater precision. Despite these advancements, the potential of pure vision transformers (ViTs) in UAV tracking remains largely untapped, especially based on the paradigm of conditional computation. In this work, we introduce an adaptive and background-aware Vision Transformer (Aba-ViT) and leverage it to develop a real-time UAV tracking framework called Aba-ViTrack. The proposed Aba-ViT exploits an adaptive and background-aware token computation method to reduce inference time. This approach adaptively discards tokens based on learned halting probabilities, which a priori are higher for background tokens than target ones. To further improve efficiency, this paper proposes a novel classroom-style learning (CSL) approach, where robustness knowledge is transmitted vertically from teacher to students, and generalization capability is enhanced through horizontal mutual learning among students. This method is used to compress Aba-ViTrack, resulting in Aba-ViTrack++. The upgraded version achieves a better balance between accuracy and efficiency in real-time UAV tracking. This version achieves a better balance between accuracy and efficiency for real-time UAV tracking. Extensive experiments on six UAV tracking benchmarks demonstrate that the proposed method achieves state-of-the-art performance in UAV tracking. The code is available at https://github.com/xyyang317/Aba-ViTrack.
Xiangyang Yang 0001, Dan Zeng 0002, Xucheng Wang, Hengzhou Ye, Xiaolan Xie 0002, Qijun Zhao, Jihua Zhu, Shuiwang Li
IEEE Trans. Circuits Syst. Video Technol.2
2025 Learning Occlusion-Robust Vision Transformers for Real-Time UAV Tracking
abstract
Single-stream architectures using Vision Transformer (ViT) backbones show great potential for real-time UAV tracking recently. However, frequent occlusions from obstacles like buildings and trees expose a major drawback: these models often lack strategies to handle occlusions effectively. New methods are needed to enhance the occlusion resilience of single-stream ViT models in aerial tracking. In this work, we propose to learn Occlusion-Robust Representations (ORR) based on ViTs for UAV tracking by enforcing an invariance of the feature representation of a target with respect to random masking operations modeled by a spatial Cox process. Hopefully, this random masking approximately simulates target occlusions, thereby enabling us to learn ViTs that are robust to target occlusion for UAV tracking. This framework is termed ORTrack. Additionally, to facilitate real-time applications, we propose an Adaptive Feature-Based Knowledge Distillation (AFKD) method to create a more compact tracker, which adaptively mimics the behavior of the teacher model ORTrack according to the task’s difficulty. This student model, dubbed ORTrack-D, retains much of ORTrack’s performance while offering higher efficiency. Extensive experiments on multiple benchmarks validate the effectiveness of our method, demonstrating its state-of-the-art performance. Codes is available at https://github.com/wuyou3474/ORTrack.
You Wu 0009, Xucheng Wang, Xiangyang Yang 0001, Dan Zeng 0002, Hengzhou Ye, Shuiwang Li
CVPR5
2025 MambaNUT: Nighttime UAV Tracking via Mamba-based Adaptive Curriculum Learning
abstract
Harnessing low-light enhancement and domain adaptation, nighttime UAV tracking has made substantial strides. However, over-reliance on image enhancement, limited high-quality nighttime data, and a lack of integration between daytime and nighttime trackers hinder the development of an end-to-end trainable framework. Additionally, current ViT-based trackers demand heavy computational resources due to their reliance on the self-attention mechanism. In this paper, we propose a novel pure Mamba-based tracking framework (MambaNUT) that employs a state space model with linear complexity as its backbone, incorporating a single-stream architecture that integrates feature learning and template-search coupling within Vision Mamba. We introduce an adaptive curriculum learning (ACL) approach that dynamically adjusts sampling strategies and loss weights, thereby improving the model’s ability of generalization. Our ACL is composed of two levels of curriculum schedulers: (1) sampling scheduler that transforms the data distribution from imbalanced to balanced, as well as from easier (daytime) to harder (nighttime) samples; (2) loss scheduler that dynamically assigns weights based on the size of the training set and IoU of individual instances. Exhaustive experiments on multiple nighttime UAV tracking benchmarks demonstrate that the proposed MambaNUT achieves state-of-the-art performance while requiring lower computational costs. The code will be available at https://github.com/wuyou3474/MambaNUT.
You Wu 0009, Xiangyang Yang 0001, Xucheng Wang, Hengzhou Ye, Dan Zeng 0002, Shuiwang Li
IROS5
2025 Adaptively bypassing vision transformer blocks for efficient visual tracking
Xiangyang Yang 0001, Dan Zeng 0002, Xucheng Wang, You Wu 0009, Hengzhou Ye, Qijun Zhao, Shuiwang Li
Pattern Recognit.2
2025 Exploiting rank-based filter pruning for real-time UAV tracking
Xucheng Wang, Dan Zeng 0002, Qijun Zhao, Shuiwang Li
Signal Process. Image Commun.2
2024 Towards Labeling-free Fine-grained Animal Pose Estimation
Dan Zeng 0002, Shuiwang Li, Qijun Zhao, Qiaomu Shen, Bo Tang 0016
ACM Multimedia1
2024 GenUDC: High Quality 3D Mesh Generation With Unsigned Dual Contouring Representation
Ruowei Wang, Dan Zeng 0002, Xueqi Ma, Zixiang Xu, Jianwei Zhang 0013, Qijun Zhao
ACM Multimedia3
2024 Robust partial face recognition using multi-label attributes
abstract
Partial face recognition (PFR) is challenging as the appearance of the face changes significantly with occlusion. In particular, these occlusions can be due to any item and may appear in any position that seriously hinders the extraction of discriminative features. Existing methods deal with PFR either by training a deep model with existing face databases containing limited occlusion types or by extracting un-occluded features directly from face regions without occlusions. Limited training data (i.e., occlusion type and diversity) can not cover the real-occlusion situations, and thus training-based methods can not learn occlusion robust discriminative features. The performance of occlusion region-based method is bounded by occlusion detection. Different from limited training data and occlusion region-based methods, we propose to use multi-label attributes for Partial Face Recognition (Attr4PFR). A novel data augmentation is proposed to solve limited training data and generate occlusion attributes. Apart from occlusion attributes, we also include soft biometric attributes and semantic attributes to explore more rich attributes to combat the loss caused by occlusions. To train our Attr4PFR, we propose an implicit attributes loss combined with a softmax loss to enforce Attr4PFR to learn discriminative features. As multi-label attributes are our auxiliary signal in the training phase, we do not need them in the inference. Extensive experiments on public benchmark AR and IJB-C databases show our method is 3% and 2.3% improvement compared to the state-of-the-art.
Gaoli Sang, Dan Zeng 0002, Raymond N. J. Veldhuis, Luuk J. Spreeuwers
Intell. Data Anal.2
2024 Face super resolution with a high frequency highway
abstract
Abstract Face shape priors such as landmarks, heatmaps, and parsing maps are widely used to improve face super resolution (SR). It is observed that face priors provide locations of high‐frequency details in key facial areas such as the eyes and mouth. However, existing methods fail to effectively exploit the high‐frequency information by using the priors as either constraints or inputs. This paper proposes a novel high frequency highway () framework to better utilize prior information for face SR, which dynamically decomposes the final SR face into a coarse SR face and a high frequency (HF) face. The coarse SR face is reconstructed from a low‐resolution face via a texture branch, using only pixel‐wise reconstruction loss. Meanwhile, the HF face is directly generated from face priors via an HF branch that employs the proposed inception–hourglass model. As a result, allows the face priors to have a direct impact on the SR face by adding the outputs of both branches as the final result and provides an extra face editing function. Extensive experiments show that significantly outperforms state‐of‐the‐art face SR methods, is general for different texture branch models and face priors, and is robust to dataset mismatch and pose variations.
Dan Zeng 0002, Xiao Yan 0002, Weibao Fu, Qiaomu Shen, Raymond N. J. Veldhuis, Bo Tang 0016
IET Image Process.1
2024 Rethinking Dual-Stream Super-Resolution Semantic Learning in Medical Image Segmentation
abstract
Image segmentation is fundamental task for medical image analysis, whose accuracy is improved by the development of neural networks. However, the existing algorithms that achieve high-resolution performance require high-resolution input, resulting in substantial computational expenses and limiting their applicability in the medical field. Several studies have proposed dual-stream learning frameworks incorporating a super-resolution task as auxiliary. In this paper, we rethink these frameworks and reveal that the feature similarity between tasks is insufficient to constrain vessels or lesion segmentation in the medical field, due to their small proportion in the image. To address this issue, we propose a DS2F (Dual-Stream Shared Feature) framework, including a Shared Feature Extraction Module (SFEM). Specifically, we present Multi-Scale Cross Gate (MSCG) utilizing multi-scale features as a novel example of SFEM. Then we define a proxy task and proxy loss to enable the features focus on the targets based on the assumption that a limited set of shared features between tasks is helpful for their performance. Extensive experiments on six publicly available datasets across three different scenarios are conducted to verify the effectiveness of our framework. Furthermore, various ablation studies are conducted to demonstrate the significance of our DS2F.
Zhongxi Qiu, Xiaoshan Chen, Dan Zeng 0002, Qingyong Hu, Jiang Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Learning Target-Aware Vision Transformers for Real-Time UAV Tracking
abstract
In recent years, the field of unmanned aerial vehicle (UAV) tracking has grown rapidly, finding numerous applications across various industries. While the discriminative correlation filters (DCF)-based trackers remain the most efficient and widely used in the UAV tracking, recently lightweight convolutional neural network (CNN)-based trackers using filter pruning have also demonstrated impressive efficiency and precision. However, the performance of these lightweight CNN-based trackers is still far from satisfactory. In the generic visual tracking, emerging vision transformer (ViT)-based trackers have shown great success by using cross-attention instead of correlation operation, enabling more effective capturing of relationships between the target and the search image. But to best of the authors’ knowledge, the UAV tracking community has not yet well explored the potential of ViTs for more effective and efficient template-search coupling for UAV tracking. In this article, we propose an efficient ViT-based tracking framework for real-time UAV tracking. Our framework integrates feature learning and template-search coupling into an efficient one-stream ViT to avoid an extra heavy relation modeling module. However, we observe that it tends to weaken the target information through transformer blocks due to the significantly more background tokens. To address this problem, we propose to maximize the mutual information (MI) between the template image and its feature representation produced by the ViT. The proposed method is dubbed TATrack. In addition, to further enhance efficiency, we introduce a novel MI maximization-based knowledge distillation, which strikes a better trade-off between accuracy and efficiency. Exhaustive experiments on five benchmarks show that the proposed tracker achieves state-of-the-art performance in UAV tracking. Code is released at:https://github.com/xyyang317/TATrack.
Shuiwang Li, Xiangyang Yang 0001, Xucheng Wang, Dan Zeng 0002, Hengzhou Ye, Qijun Zhao
IEEE Trans. Geosci. Remote. Sens.4
2024 $\mathsf {CheetahTraj}$CheetahTraj: Efficient Visualization for Large Trajectory Dataset With Quality Guarantee
abstract
Visualizing large-scale trajectory dataset is a core subroutine for many applications. However, rendering all trajectories could result in severe visual clutter and incur long visualization delays due to large data volume. Naively sampling the trajectories reduces visualization time but usually harms visual quality, i.e., the generated visualizations may look substantially different from the exact ones without sampling. In this paper, we propose$\mathsf {CheetahTraj}$, a principled sampling framework that achieves both high visualization quality and low visualization latency. We first define thevisual quality functionmeasuring the similarity between two visualizations, based on which we formulate the quality optimal sampling problem (${\sf QOSP}$). To solve${\sf QOSP}$, we design theVisualQualityGuaranteedSampling algorithms, which reduce visual clutter while guaranteeing visual quality by considering both trajectory data distribution and human perception properties. We also develop a quad-tree-based index ($\mathsf {InvQuad}$) that allows using trajectory samples computed offline for interactive online visualization. Extensive experiments including case-, user-, and quantitative-studies are conducted on three real-world trajectory datasets, and the results show that$\mathsf {CheetahTraj}$consistently provides higher visual quality and better efficiency than baseline methods. Compared with visualizing all trajectories,$\mathsf {CheetahTraj}$reduces the visualization latency by up to 3 orders of magnitude while avoiding visual clutter.
Qiaomu Shen, Chaozu Zhang, Xiao Yan 0002, Dan Zeng 0002, Wei Zeng 0004, Bo Tang 0016
IEEE Trans. Knowl. Data Eng.5
2024 QEVIS: Multi-Grained Visualization of Distributed Query Execution
abstract
Distributed query processing systems such as Apache Hive and Spark are widely-used in many organizations for large-scale data analytics. Analyzing and understanding the query execution process of these systems are daily routines for engineers and crucial for identifying performance problems, optimizing system configurations, and rectifying errors. However, existing visualization tools for distributed query execution are insufficient because (i) most of them (if not all) do not provide fine-grained visualization (i.e., the atomic task level), which can be crucial for understanding query performance and reasoning about the underlying execution anomalies, and (ii) they do not support proper linkages between system status and query execution, which makes it difficult to identify the causes of execution problems. To tackle these limitations, we propose QEVIS, which visualizes distributed query execution process with multiple views that focus on different granularities and complement each other. Specifically, we first devise a query logical plan layout algorithm to visualize the overall query execution progress compactly and clearly. We then propose two novel scoring methods to summarize the anomaly degrees of the jobs and machines during query execution, and visualize the anomaly scores intuitively, which allow users to easily identify the components that are worth paying attention to. Moreover, we devise a scatter plot-based task view to show a massive number of atomic tasks, where task distribution patterns are informative for execution problems. We also equip QEVIS with a suite of auxiliary views and interaction methods to support easy and effective cross-view exploration, which makes it convenient to track the causes of execution problems. QEVIS has been used in the production environment of our industry partner, and we present three use cases from real-world applications and user interview to demonstrate its effectiveness. QEVIS is open-source at https://github.com/DBGroup-SUSTech/QEVIS.
Qiaomu Shen, Zhengxin You, Xiao Yan 0002, Chaozu Zhang, Dan Zeng 0002, Jianbin Qin, Bo Tang 0016
IEEE Trans. Vis. Comput. Graph.6
2023 Adaptive and Background-Aware Vision Transformer for Real-Time UAV Tracking
abstract
While discriminative correlation filters (DCF)-based trackers prevail in UAV tracking for their favorable efficiency, lightweight convolutional neural network (CNN)-based trackers using filter pruning have also demonstrated remarkable efficiency and precision. However, the use of pure vision transformer models (ViTs) for UAV tracking remains unexplored, which is a surprising finding given that ViTs have been shown to produce better performance and greater efficiency than CNNs in image classification. In this paper, we propose an efficient ViT-based tracking framework, Aba-ViTrack, for UAV tracking. In our framework, feature learning and template-search coupling are integrated into an efficient one-stream ViT to avoid an extra heavy relation modeling module. The proposed Aba-ViT exploits an adaptive and background-aware token computation method to reduce inference time. This approach adaptively discards tokens based on learned halting probabilities, which a priori are higher for background tokens than target ones. Extensive experiments on six UAV tracking benchmarks demonstrate that the proposed Aba-ViTrack achieves state-of-the-art performance in UAV tracking. Code is available at https://github.com/xyyang317/Aba-ViTrack.
Shuiwang Li, Xiangxyang Yang, Dan Zeng 0002, Xucheng Wang
ICCV3
2023 Towards Discriminative Representations with Contrastive Instances for Real-Time UAV Tracking
abstract
Maintaining high efficiency and high precision are two fundamental challenges in UAV tracking due to the constraints of computing resources, battery capacity, and UAV maximum load. Discriminative correlation filters (DCF)-based trackers can yield high efficiency on a single CPU but with inferior precision. Lightweight Deep learning (DL)-based trackers can achieve a good balance between efficiency and precision but performance gains are limited by the compression rate. High compression rate often leads to poor discriminative representations. To this end, this paper aims to enhance the discriminative power of feature representations from a new feature-learning perspective. Specifically, we attempt to learn more disciminative representations with contrastive instances for UAV tracking in a simple yet effective manner, which not only requires no manual annotations but also allows for developing and deploying a lightweight model. We are the first to explore contrastive learning for UAV tracking. Extensive experiments on four UAV benchmarks, including UAV123@10fps, DTB70, UAVDT and VisDrone2018, show that the proposed DRCI tracker significantly outperforms state-of-the-art UAV tracking methods.
Dan Zeng 0002, Mingliang Zou, Xucheng Wang, Shuiwang Li
ICME1
2023 Analyzing and Combating Attribute Bias for Face Restoration
abstract
Face restoration (FR) recovers high resolution (HR) faces from low resolution (LR) faces and is challenging due to its ill-posed nature. With years of development, existing methods can produce quality HR faces with realistic details. However, we observe that key facial attributes (e.g., age and gender) of the restored faces could be dramatically different from the LR faces and call this phenomenon attribute bias, which is fatal when using FR for applications such as surveillance and security. Thus, we argue that FR should consider not only image quality as in existing works but also attribute bias. To this end, we thoroughly analyze attribute bias with extensive experiments and find that two major causes are the lack of attribute information in LR faces and bias in the training data. Moreover, we propose the DebiasFR framework to produce HR faces with high image quality and accurate facial attributes. The key design is to explicitly model the facial attributes, which also allows to adjust facial attributes for the output HR faces. Experiment results show that DebiasFR has comparable image quality but significantly smaller attribute bias when compared with state-of-the-art FR methods.
Zelin Li 0002, Dan Zeng 0002, Xiao Yan 0002, Qiaomu Shen, Bo Tang 0016
IJCAI2
2023 Data-Scarce Animal Face Alignment via Bi-Directional Cross-Species Knowledge Transfer
abstract
Animal face alignment is challenging due to large intra- and inter-species variations and a scarcity of labeled data. Existing studies circumvent this problem by directly finetuning a human face alignment model or focusing on animal-specific face alignment~(e.g., horse, sheep). In this paper, we propose Cross-Species Knowledge Transfer, Meta-CSKT, for animal face alignment, which consists of a base network and an adaptation network. Two networks continuously complement each other through the bi-directional cross-species knowledge transfer. This is motivated by observing knowledge sharing among animals. Meta-CSKT uses a circuit feedback mechanism to improve the base network with the cognitive differences of the adaptation network between few-shot labeled and large-scale unlabeled data. In addition, we propose a positive example mining method to identify positives, semi-hard positives, and hard negatives in unlabeled data to mitigate the scarcity of labeled data and facilitate Meta-CSKT learning. Experiments show that Meta-CSKT outperforms state-of-the-art methods by a large margin on the horse facial keypoint dataset and Japanese Macaque Species dataset, while achieving comparable results to state-of-the-art methods on large-scale labeled AnimalWeb~(e.g., 18K), using only a few labeled images~(e.g., 40)1.
Dan Zeng 0002, Shanchuan Hong, Shuiwang Li, Qiaomu Shen, Bo Tang 0016
ACM Multimedia1
2023 Cascaded face super-resolution with shape and identity priors
abstract
Abstract Despite impressive progress in face super‐resolution (SR), it is an open challenge to reconstruct a reliable SR face that preserves authentic facial characteristics. Here, the problem of super‐resolving low‐resolution (LR) faces to high‐resolution (HR) ones is addressed. To tackle the ill‐posed nature of face SR, the cascaded super‐resolution network (CSRNet) is proposed to utilize shape and identity priors jointly and progressively, the first to explore multiple priors. Specifically, CSRNet adopts a cascaded structure to transform an LR face to HR face progressively via multiple stages. At each stage, CSRNet forces its output face image to match both the shape priors and identity priors extracted from the ground‐truth HR face. The shape priors estimated in one stage are merged into the inputs of its subsequent stage to provide rich information for the face SR. To generate realistic yet discriminative faces, the cascaded super‐resolution generative adversarial network (CSRGAN) is also proposed to incorporate the adversarial loss and identification loss into CSRNet. Extensive experiments on popular benchmarks show that the CSRNet and CSRGAN outperform existing face SR state‐of‐the‐art methods, both quantitatively and qualitatively, and detailed ablation studies show the advantage of this method.
Dan Zeng 0002, Zelin Li 0002, Xiao Yan 0002, Xinshao Wang, Jiang Liu 0001, Bo Tang 0016
IET Image Process.1
2022 Learning Disentangled Representation in Pruning for Real-Time UAV Tracking
Siyu Ma, Yuting Liu 0004, Dan Zeng 0002, Yaxin Liao, Shuiwang Li
ACML3
2022 GHive: accelerating analytical query processing in apache hive via CPU-GPU heterogeneous computing
abstract
As a popular distributed data warehouse system, Apache Hive has been widely used for big data analytics in many organizations. Meanwhile, exploiting the massive parallelism of GPU to accelerate online analytical processing (OLAP) has been extensively explored in the database community. In this paper, we present GHive, which enhances CPU-based Hive via CPU-GPU heterogeneous computing. GHive is designed for the business intelligence applications and provides the same API as Hive for compatibility. To run SQL queries jointly on both CPU and GPU, GHive comes with three key techniques: (i) a novel data model gTable, which is column-based and enables efficient data movement between CPU memory and GPU memory; (ii) a GPU-based operator library Panda, which provides a complete set of SQL operators with extensively optimized GPU implementations; (iii) a hardware-aware MapReduce job placement scheme, which puts jobs judiciously on either GPU or CPU via a cost-based approach. In the experiments, we observe that GHive outperforms Hive in both query processing speed and operating expense on the Star Schema Benchmark (SSB).
Bo Tang 0016, Jiashu Zhang, Yangshen Deng, Xiao Yan 0002, Xinying Zheng, Qiaomu Shen, Dan Zeng 0002, Zunyao Mao, Chaozu Zhang, Zhengxin You, Runzhe Jiang, Fang Wang 0012, Man Lung Yiu, Huan Li 0003, Mingji Han, Zhenghai Luo
SoCC8
2022 Face2Exp: Combating Data Biases for Facial Expression Recognition
abstract
Facial expression recognition (FER) is challenging due to the class imbalance caused by data collection. Existing studies tackle the data bias problem using only labeled facial expression dataset. Orthogonal to existing FER methods, we propose to utilize large unlabeled face recognition (FR) datasets to enhance FER. However, this raises another data bias problem—the distribution mismatch between FR and FER data. To combat the mismatch, we propose the Meta-Face2Exp framework, which consists of a base network and an adaptation network. The base network learns prior expression knowledge on class-balanced FER data while the adaptation network is trained to fit the pseudo labels of FR data generated by the base model. To combat the mismatch between FR and FER data, Meta-Face2Exp uses a circuit feedback mechanism, which improves the base network with the feedback from the adaptation network. Experiments show that our MetaFace2Exp achieves comparable accuracy to state-of-the-art FER methods with 10% of the labeled FER data utilized by the baselines. We also demonstrate that the circuit feedback mechanism successfully eliminates data bias11Code is available at link: https://github.com/danzeng1990/Face2Exp..
Dan Zeng 0002, Xiao Yan 0002, Yuting Liu 0004, Bo Tang 0016
CVPR1
2022 CheetahKG: A Demonstration for Core-based Top-$k$ Frequent Pattern Discovery on Knowledge Graphs
abstract
Knowledge graphs capture the complex relationships among various entities, which can be found in various real world applications, e.g., Amazon product graph, Freebase, and COVID-19. To facilitate the knowledge graph analytical tasks, a system that supports interactive and efficient query processing is always in demand. In this demonstration, we develop a prototype system, CheetahKG, that embeds with our state-of-the-art query processing engine for the top-$k$frequent pattern discovery. Such discovered patterns can be used for two purposes, (i) identifying related patterns and (ii) guiding knowledge exploration. In the demonstration sessions, the attendees will be invited to test the efficiency and effectiveness of the query engine and use the discovered patterns to analyze knowledge graphs on CheetahKG.
Bo Tang 0016, Qiandong Tang, Qiaomu Shen, Leong Hou U, Xiao Yan 0002, Dan Zeng 0002
ICDE8
2022 Rank-Based Filter Pruning for Real-Time UAV Tracking
abstract
Unmanned aerial vehicle (UAV) tracking has wide poten-tial applications in such as agriculture, navigation, and public security. However, the limitations of computing resources, battery capacity, and maximum load of UAV hinder the de-ployment of deep learning-based tracking algorithms on UAV. Consequently, discriminative correlation filters (DCF) track-ers stand out in the UAV tracking community because of their high efficiency. However, their precision is usually much lower than trackers based on deep learning. Model compression is a promising way to narrow the gap (i.e., effciency, precision) between DCF- and deep learning- based trackers, which has not caught much attention in UAV tracking. In this paper, we propose the P-SiamFC++ tracker, which is the first to use rank-based filter pruning to compress the SiamFC++ model, achieving a remarkable balance between efficiency and precision. Our method is general and may encourage further studies on UAV tracking with model compression. Extensive experiments on four UAV benchmarks, including UAV123@10fps, DTB70, UAVDT and Vistrone2018, show that P-SiamFC++ tracker significantly outperforms state-of-the-art UAV tracking methods.
Xucheng Wang, Dan Zeng 0002, Qijun Zhao, Shuiwang Li
ICME2
2022 SuperVessel: Segmenting High-Resolution Vessel from Low-Resolution Retinal Image
Zhongxi Qiu, Dan Zeng 0002, Jiang Liu 0001
PRCV (2)3
2022 GHive: A Demonstration of GPU-Accelerated Query Processing in Apache Hive
abstract
As a distributed, fault-tolerant data warehouse system for large-scale data analytics, Apache Hive has been used for various applications in many organizations (e.g., Facebook, Amazon, and Huawei). Exploiting the large degrees of parallelism of GPU to improve the performance of online analytical processing (OLAP) in database system is a common practice in the industry. Meanwhile, it is a common practice to exploit the large degrees of parallelism of GPU to improve the performance of online analytical processing (OLAP) in database systems. This demo presents GHive, which enables Apache Hive to accelerate OLAP queries by jointly utilizing CPU and GPU in intelligent and efficient ways. The takeaways for SIGMOD attendees include: (1) the superior performance of GHive compared with vanilla Hive that only uses CPU; (2) intuitive visualizations of execution statistics for Hive and GHive to understand where the acceleration of GHive comes from; (3) detailed profiling of the time taken by each operator on CPU and GPU to show the advantages of GPU execution.
Bo Tang 0016, Jiashu Zhang, Yangshen Deng, Xinying Zheng, Qiaomu Shen, Xiao Yan 0002, Dan Zeng 0002, Zunyao Mao, Chaozu Zhang, Zhengxin You, Runzhe Jiang, Fang Wang 0012, Man Lung Yiu, Huan Li 0003, Mingji Han, Zhenghai Luo
SIGMOD Conference8
2022 Combating spatial redundancy with spectral norm attention in convolutional learners
Jiansheng Fang, Dan Zeng 0002, Xiao Yan 0002, Yubing Zhang, Hongbo Liu 0007, Bo Tang 0016, Ming Yang 0039, Jiang Liu 0001
Neurocomputing2
2021 Combating Ambiguity for Hash-Code Learning in Medical Instance Retrieval
abstract
When encountering a dubious diagnostic case, medical instance retrieval can help radiologists make evidence-based diagnoses by finding images containing instances similar to a query case from a large image database. The similarity between the query case and retrieved similar cases is determined by visual features extracted from pathologically abnormal regions. However, the manifestation of these regions often lacks specificity, i.e., different diseases can have the same manifestation, and different manifestations may occur at different stages of the same disease. To combat the manifestation ambiguity in medical instance retrieval, we propose a novel deep framework called Y-Net, encoding images into compact hash-codes generated from convolutional features by feature aggregation. Y-Net can learn highly discriminative convolutional features by unifying the pixel-wise segmentation loss and classification loss. The segmentation loss allows exploring subtle spatial differences for good spatial-discriminability while the classification loss utilizes class-aware semantic information for good semantic-separability. As a result, Y-Net can enhance the visual features in pathologically abnormal regions and suppress the disturbing of the background during model training, which could effectively embed discriminative features into the hash-codes in the retrieval stage. Extensive experiments on two medical image datasets demonstrate that Y-Net can alleviate the ambiguity of pathologically abnormal regions and its retrieval performance outperforms the state-of-the-art method by an average of 9.27% on the returned list of 10.
Jiansheng Fang, Huazhu Fu, Dan Zeng 0002, Xiao Yan 0002, Yuguang Yan, Jiang Liu 0001
IEEE J. Biomed. Health Informatics3
2020 Joint Face Alignment and 3D Face Reconstruction with Application to Face Recognition
abstract
Face alignment and 3D face reconstruction are traditionally accomplished as separated tasks. By exploring the strong correlation between 2D landmarks and 3D shapes, in contrast, we propose a joint face alignment and 3D face reconstruction method to simultaneously solve these two problems for 2D face images of arbitrary poses and expressions. This method, based on a summation model of 3D faces and cascaded regression in 2D and 3D shape spaces, iteratively and alternately applies two cascaded regressors, one for updating 2D landmarks and the other for 3D shape. The 3D shape and the landmarks are correlated via a 3D-to-2D mapping matrix, which is updated in each iteration to refine the location and visibility of 2D landmarks. Unlike existing methods, the proposed method can fully automatically generate both pose-and-expression-normalized (PEN) and expressive 3D faces and localize both visible and invisible 2D landmarks. Based on the PEN 3D faces, we devise a method to enhance face recognition accuracy across poses and expressions. Both linear and nonlinear implementations of the proposed method are presented and evaluated in this paper. Extensive experiments show that the proposed method can achieve the state-of-the-art accuracy in both face alignment and 3D face reconstruction, and benefit face recognition owing to its reconstructed PEN 3D face.
Feng Liu 0037, Qijun Zhao, Xiaoming Liu 0002, Dan Zeng 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 Combined training strategy for low-resolution face recognition with limited application-specific data
abstract
Application‐specific data for certain biometric applications are often not sufficiently available. The authors present a solution for face recognition with limited application‐specific data. Existing methods often use a classifier with convolutional neural networks (CNNs) as feature extractors. The CNNs are trained with massive general (i.e. not application specific) data and the classifier is trained with application‐specific data. Alternatively, the authors propose a combined training strategy to train the classifier on a balanced mixture of general and application‐specific data, such that the recognition performance is maximised. The proposed method largely alleviates the needs for application‐specific data. To prove its effectiveness, they apply the proposed method to low‐resolution face recognition. Specifically, they use the heterogeneous joint Bayesian (HJB) classifier that is capable of comparing features from the same modality but with different characteristics. To further boost performance, the authors augment the training data by pre‐processing it to resemble application‐specific data. They conducted extensive experiments on challenging datasets, namely, SCface and COX. The results show that the proposed method improves the true match rate on SCface at a false match rate of 10% by ∼11% and the true match rate on COX at a false match rate of 1% by ∼12%.
Dan Zeng 0002, Luuk J. Spreeuwers, Raymond N. J. Veldhuis, Qijun Zhao
IET Image Process.1
2018 Disentangling Features in 3D Face Shapes for Joint Face Reconstruction and Recognition
abstract
This paper proposes an encoder-decoder network to disentangle shape features during 3D face reconstruction from single 2D images, such that the tasks of reconstructing accurate 3D face shapes and learning discriminative shape features for face recognition can be accomplished simultaneously. Unlike existing 3D face reconstruction methods, our proposed method directly regresses dense 3D face shapes from single 2D images, and tackles identity and residual (i.e., non-identity) components in 3D face shapes explicitly and separately based on a composite 3D face shape model with latent representations. We devise a training process for the proposed network with a joint loss measuring both face identification error and 3D face shape reconstruction error. To construct training data we develop a method for fitting 3D morphable model (3DMM) to multiple 2D images of a subject. Comprehensive experiments have been done on MICC, BU3DFE, LFW and YTF databases. The results show that our method expands the capacity of 3DMM for capturing discriminative shape features and facial detail, and thus outperforms existing methods both in 3D face reconstruction accuracy and in face recognition accuracy.
Feng Liu 0013, Ronghang Zhu, Dan Zeng 0002, Qijun Zhao, Xiaoming Liu 0002
CVPR3
2017 Examplar coherent 3D face reconstruction from forensic mugshot database
Dan Zeng 0002, Qijun Zhao, Shuqin Long, Jing Li 0060
Image Vis. Comput.1
2017 On 3D face reconstruction via cascaded regression in shape space
abstract
Cascaded regression has been recently applied to reconstruct 3D faces from single 2D images directly in shape space, and has achieved state-of-the-art performance. We investigate thoroughly such cascaded regression based 3D face reconstruction approaches from four perspectives that are not well been studied: (1) the impact of the number of 2D landmarks; (2) the impact of the number of 3D vertices; (3) the way of using standalone automated landmark detection methods; (4) the convergence property. To answer these questions, a simplified cascaded regression based 3D face reconstruction method is devised. This can be integrated with standalone automated landmark detection methods and reconstruct 3D face shapes that have the same pose and expression as the input face images, rather than normalized pose and expression. An effective training method is also proposed by disturbing the automatically detected landmarks. Comprehensive evaluation experiments have been carried out to compare to other 3D face reconstruction methods. The results not only deepen the understanding of cascaded regression based 3D face reconstruction approaches, but also prove the effectiveness of the proposed method.
Feng Liu 0013, Dan Zeng 0002, Jing Li 0060, Qijun Zhao
Frontiers Inf. Technol. Electron. Eng.2
2016 Joint Face Alignment and 3D Face Reconstruction
Feng Liu 0013, Dan Zeng 0002, Qijun Zhao, Xiaoming Liu 0002
ECCV (5)2