VLDB 2026 Research / reviewers in the wild / expert
Xiaoguang Tu
dblp:132/0758
· DBLP profile ↗
17ranked-venue papers
7as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Three-dimensional server deployment optimization in multi-UAV-assisted edge networks
Guilin Yuan, Xiaoguang Tu |
CCF Trans. Pervasive Comput. Interact. | 5 |
| 2026 | Joint optimization of UAV trajectory, RIS selection, and offloading strategy for RIS-assisted UAV-MEC systems based on DRL
Xiaoguang Tu |
Comput. Commun. | 5 |
| 2026 | Secure Task Scheduling and Trajectory Optimization for UAV-Assisted Mobile Edge Computing Based on DDPG and DPFLabstractWith the development of the Internet of Things (IoT), Unmanned Aerial Vehicle-Mobile Edge Computing (UAV-MEC) plays a critical role in enhancing network performance and data processing capabilities. However, significant challenges that impede system performance in UAV-MEC systems include task scheduling inefficiency, privacy leakage, and insufficient trust evaluation. To address these issues, we propose a novel framework that synergistically integrates a UAV trust evaluation model, Differential Privacy Federated Learning (DPFL), and the Deep Deterministic Policy Gradient (DDPG) algorithm. The core innovation lies in creating a security-aware DRL agent, where the trust model guides the DDPG’s decisions to inherently balance efficiency with security, while the entire collaborative learning process is protected by DPFL. By incorporating trust evaluation and privacy protection mechanisms, the proposed method achieves secure task scheduling with minimal latency under the condition of high trust values, while considering UAV energy constraints and successfully optimizing UAV trajectories. Simulation results demonstrate that, compared to traditional Deep Reinforcement Learning (DRL) algorithms, the proposed algorithm significantly reduces the cost by 4.12% and exhibits higher stability and convergence. Specifically, under varying conditions such as user counts, random task sizes, UAV computational capacities, and levels of UAV intrusion power, the costs decreased by up to 9.56%, 5.49%, 3.77%, and 4.43%, respectively. Overall, the proposed scheme exhibits significant improvements in user privacy protection, system efficiency, and adaptability. Xiaoguang Tu |
IEEE Internet Things J. | 5 |
| 2026 | Small object detection in UAV aerial imagery through domain consistency optimization for cross-resolution semantic alignment
Xiaoguang Tu, Zeng Gao, Rubin He |
Pattern Recognit. | 1 |
| 2026 | Phys-EdiGAN: A privacy-preserving method for editing physiological signals in facial videos
Xiaoguang Tu, Zhiyi Niu, Juhang Yin, Zhaoxin Fan, Jian Zhao 0006 |
Pattern Recognit. | 1 |
| 2025 | Trust-aware task offloading for cost-effective UAV-based edge computing based on reinforcement learning
Kemeng Lin, Xiaoguang Tu |
Neural Comput. Appl. | 4 |
| 2023 | Low-Light Image Enhancement by Learning Contrastive Representations in Spatial and Frequency DomainsabstractImages taken under low-light conditions tend to suffer from poor visibility, which can decrease image quality and even reduce the performance of the downstream tasks. It is hard for a CNN-based method to learn generalized features that can recover normal images from the ones under various unknow low-light conditions. In this paper, we propose to incorporate the contrastive learning into an illumination correction network to learn abstract representations to distinguish various low-light conditions in the representation space, with the purpose of enhancing the generalizability of the network. Considering that light conditions can change the frequency components of the images, the representations are learned and compared in both spatial and frequency domains to make full advantage of the contrastive learning. The proposed method is evaluated on LOL and LOL-V2 datasets, the results show that the proposed method achieves better qualitative and quantitative results compared with other state-of-the-arts. Yi Huang 0025, Xiaoguang Tu, Gui Fu, Bokai Liu, Ziliang Feng |
ICME | 2 |
| 2022 | Joint Face Image Restoration and Frontalization for RecognitionabstractIn real-world scenarios, many factors may harm face recognition performance,e.g., large pose, bad illumination, low resolution, blur and noise. To address these challenges, previous efforts usually first restore the low-quality faces to high-quality ones and then perform face recognition. However, most of these methods are stage-wise, which is sub-optimal and deviates from the reality. In this paper, we address all these challenges jointly for unconstrained face recognition. We propose anMulti-DegradationFaceRestoration (MDFR) model to restore frontalized high-quality faces from the given low-quality ones under arbitrary facial poses, with three distinct novelties. First, MDFR is a well-designed encoder-decoder architecture which extracts feature representation from an input face image with arbitrary low-quality factors and restores it to a high-quality counterpart. Second, MDFR introduces a pose residual learning strategy along with a 3D-basedPoseNormalizationModule (PNM), which can perceive the pose gap between the input initial pose and its real-frontal pose to guide the face frontalization. Finally, MDFR can generate frontalized high-quality face images by a single unified network, showing a strong capability of preserving face identity. Qualitative and quantitative experiments on both controlled and in-the-wild benchmarks demonstrate the superiority of MDFR over state-of-the-art methods on both face frontalization and face restoration. Xiaoguang Tu, Jian Zhao 0006, Wenjie Ai, Guodong Guo, Zhifeng Li 0001, Wei Liu 0005, Jiashi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Image-to-Video Generation via 3D Facial DynamicsabstractWe present a versatile model, FaceAnime, for various video generation tasks from still images. Video generation from a single face image is an interesting problem and usually tackled by utilizing Generative Adversarial Networks (GANs) to integrate information from the input face image and a sequence of sparse facial landmarks. However, the generated face images usually suffer from quality loss, image distortion, identity change, and expression mismatching due to the weak representation capacity of the facial landmarks. In this paper, we propose to “imagine” a face video from a single face image according to the reconstructed 3D face dynamics, aiming to generate a realistic and identity-preserving face video, with precisely predicted pose and facial expression. The 3D dynamics reveal changes of the facial expression and motion, and can serve as a strong prior knowledge for guiding highly realistic face video generation. In particular, we explore face video prediction and exploit a well-designed 3D dynamic prediction network to predict a 3D dynamic sequence for a single face image. The 3D dynamics are then further rendered by the sparse texture mapping algorithm to recover structural details and sparse textures for generating face frames. Our model is versatile for various AR/VR and entertainment applications, such as face video retargeting and face video prediction. Superior experimental results have well demonstrated its effectiveness in generating high-fidelity, identity-preserving, and visually pleasant face video clips from a single source face image. Xiaoguang Tu, Yingtian Zou, Jian Zhao 0006, Wenjie Ai, Jian Dong 0011, Yuan Yao 0011, Zhikang Wang, Guodong Guo, Zhifeng Li 0001, Wei Liu 0005, Jiashi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Robust Video-Based Person Re-Identification by Hierarchical MiningabstractVideo-based person re-identification (Re-ID) aims at retrieving the person through the video sequences across non-overlapping cameras. Some characteristics of pedestrians are not consecutive across frames due to the variations of viewpoints, postures, and occlusions over time. However, existing methods ignore such data peculiarity and the networks tend to only learn those salient consecutive characteristics among frames in video sequences. As a result, the learned representations fail to cover all the characteristics of pedestrians, thus lacking integrity and discrimination. To tackle this problem, we present a novel deep architecture termed Hierarchical Mining Network (HMN), which mines as many pedestrians’ characteristics by referring to the temporal and intra-class knowledge. It consists of a novel Attentive Temporal Module (ATM) and a Dynamic Supervising Branch (DSB), with a Balancing Triplet Loss (BTL) assisting the training. The proposed ATM, with pedestrian perceiving capacity, is capable of evaluating each activation of features through temporal analysis, so that the temporally scattered characteristics of pedestrians can be better aggregated and the contaminated ones can be eliminated. Then, the DSB along with the BTL further enhances the integrity of representations by multiple supervision. Specifically, the DSB perceives the diversities of intra-class samples in each mini-batch and generates targeted supervising signals for them, in which process the BTL guarantees the signals with smaller intra-class variations and larger inter-class variations. Comprehensive experiments on two video-based datasets, i.e., MARS, and DukeMTMC-VideoReID, demonstrate the contribution of each component and the superiority of the proposed HMN over the state-of-the-arts. Benchmarking our model on three popular image-based datasets, i.e., Market1501, DukeMTMC-Reid, and MSMT17 additionally verifies the promising generalizability of the proposed DSB and BTL. Zhikang Wang, Lihuo He, Xiaoguang Tu, Jian Zhao 0006, Xinbo Gao 0001, Shengmei Shen, Jiashi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | 3D Face Reconstruction From A Single Image Assisted by 2D Face Images in the Wildabstract3D face reconstruction from a single image is an important task in many multimedia applications. Recent works typically learn a CNN-based 3D face model that regresses coefficients of a 3D Morphable Model (3DMM) from 2D images to perform 3D face reconstruction. However, the shortage of training data with 3D annotations considerably limits performance of these methods. To alleviate this issue, we propose a novel 2D-Assisted Learning (2DAL) method that can effectively use “in the wild” 2D face images with noisy landmark information to substantially improve 3D face model learning. Specifically, taking the sparse 2D facial landmark heatmaps as additional information, 2DAL introduces four novel self-supervision schemes that view the 2D landmark and 3D landmark prediction as a self-mapping process, including the landmark self-prediction consistency for 2D and 3D faces respectively, cycle-consistency over the 2D landmark prediction and self-critic over the predicted 3DMM coefficients based on landmark prediction. Using these four self-supervision schemes, 2DAL significantly relieves the demands for the the conventional paired 2D-to-3D annotations and gives much higher-quality 3D face models without requiring any additional 3D annotations. Experiments on AFLW2000-3D, AFLW-LFPA and Florence benchmarks show that our method outperforms state-of-the-arts for both 3D face reconstruction and dense face alignment by a large margin. Xiaoguang Tu, Jian Zhao 0006, Mei Xie, Zihang Jiang, Akshaya Balamurugan, Yao Luo, Yang Zhao 0003, Lingxiao He, Zheng Ma 0005, Jiashi Feng |
IEEE Trans. Multim. | 1 |
| 2020 | Single Image Super-Resolution Via Residual Neuron Attention NetworksabstractDeep Convolutional Neural Networks (DCNNs) have achieved impressive performance in Single Image Super-Resolution (SISR). To further improve the performance, existing CNN-based methods generally focus on designing deeper architecture of the network. However, we argue blindly increasing network's depth is not the most sensible way. In this paper, we propose a novel end-to-end Residual Neuron Attention Networks (RNAN) for more efficient and effective SISR. Structurally, our RNAN is a sequential integration of the well-designed Global Context-enhanced Residual Groups (GCRGs), which extracts super-resolved features from coarse to fine. Our GCRG is designed with two novelties. Firstly, the Residual Neuron Attention (RNA) mechanism is proposed in each block of GCRG to reveal the relevance of neurons for better feature representation. Furthermore, the Global Context (GC) block is embedded into RNAN at the end of each GCRG for effectively modeling the global contextual information. Experiments results demonstrate that our RNAN achieves the comparable results with state-of-the-art methods in terms of both quantitative metrics and visual quality, however, with simplified network architecture. Wenjie Ai, Xiaoguang Tu, Shilei Cheng, Mei Xie |
ICIP | 2 |
| 2020 | Learning Generalizable and Identity-Discriminative Representations for Face Anti-SpoofingabstractFace anti-spoofing aims to detect presentation attack to face recognition--based authentication systems. It has drawn growing attention due to the high security demand. The widely adopted CNN-based methods usually well recognize the spoofing faces when training and testing spoofing samples display similar patterns, but their performance would drop drastically on testing spoofing faces of novel patterns or unseen scenes, leading to poor generalization performance. Furthermore, almost all current methods treat face anti-spoofing as a prior step to face recognition, which prolongs the response time and makes face authentication inefficient. In this article, we try to boost the generalizability and applicability of face anti-spoofing methods by designing a new generalizable face authentication CNN (GFA-CNN) model with three novelties. First, GFA-CNN introduces a simple yet effective total pairwise confusion loss for CNN training that properly balances contributions of all spoofing patterns for recognizing the spoofing faces. Second, it incorporate a fast domain adaptation component to alleviate negative effects brought by domain variation. Third, it deploys filter diversification learning to make the learned representations more adaptable to new scenes. In addition, the proposed GFA-CNN works in a multi-task manner—it performs face anti-spoofing and face recognition simultaneously. Experimental results on five popular face anti-spoofing and face recognition benchmarks show that GFA-CNN outperforms previous face anti-spoofing methods on cross-test protocols significantly and also well preserves the identity information of input face images. Xiaoguang Tu, Zheng Ma 0005, Jian Zhao 0006, Guodong Du 0004, Mei Xie, Jiashi Feng |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2019 | Multi-Prototype Networks for Unconstrained Set-based Face RecognitionabstractIn this paper, we address the challenging unconstrained set-based face recognition problem where each subject face is instantiated by a set of media (images and videos) instead of a single image. Naively aggregating information from all the media within a set would suffer from the large intra-set variance caused by heterogeneous factors (e.g., varying media modalities, poses and illumination) and fail to learn discriminative face representations. A novel Multi-Prototype Network (MP- Net) model is thus proposed to learn multiple prototype face representations adaptively from the media sets. Each learned prototype is representative for the subject face under certain condition in terms of pose, illumination and media modality. Instead of handcrafting the set partition for prototype learn- ing, MPNet introduces a Dense SubGraph (DSG) learning sub-net that implicitly untangles inconsistent media and learns a number of representative prototypes. Qualitative and quantitative experiments clearly demonstrate the superiority of the proposed model over state-of-the-arts. Jian Zhao 0006, Jianshu Li, Xiaoguang Tu, Fang Zhao 0006, Yuan Xin, Junliang Xing, Hengzhu Liu, Shuicheng Yan, Jiashi Feng |
IJCAI | 3 |
| 2019 | Detecting multi-oriented text with corner-based region proposals
Linjie Deng, Yanxiang Gong, Yi Lin 0006, Jingwen Shuai, Xiaoguang Tu, Yuefei Zhang, Zheng Ma 0005, Mei Xie |
Neurocomputing | 5 |
| 2018 | Supervoxel Segmentation and Bias Correction of MR Image with Intensity Inhomogeneity
Chongjin Zhu, Jie-Zhi Cheng, Xiaoguang Tu, Daiqiang Chen, Bin Sun 0006, Yachun Gao, Mei Xie |
Neural Process. Lett. | 5 |
| 2017 | Illumination normalization based on correction of large-scale components for face recognition
Xiaoguang Tu, Mei Xie, Zheng Ma 0005 |
Neurocomputing | 1 |