Guosheng Hu

dblp:98/7676 · DBLP profile ↗
← Back
74ranked-venue papers
10as first author
37since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 8 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 45 · 7 first-author · 20 since 2021Security and privacy · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Not All Directions Matter: Towards Structured and Task-Aware Low-Rank Model Adaptation
abstract
Xi Xiao, Chenrui Ma, Yunbei Zhang, Chen Liu, Zhuxuanzi Wang, Yanshu Li, Lin Zhao, Guosheng Hu, Tianyang Wang, Hao Xu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xi Xiao 0003, Chenrui Ma, Yunbei Zhang, Chen Liu 0020, Zhuxuanzi Wang, Yanshu Li, Guosheng Hu, Tianyang Wang 0004
ACL (1)8
2026 Uncertainty-Guided Cross-Modal Distillation for Category-Level Object Pose Estimation
abstract
Recent years have seen significant advancements in category-level object pose estimation, largely driven by multimodal (RGB-D) approaches. Despite their success, depth-only methods remain widely adopted in practical applications due to their superior computational efficiency and ease of deployment. However, these methods typically suffer from a noticeable performance gap compared to multimodal methods. To bridge this gap, we propose a novel framework, Cross-Modal Uncertainty Distillation for Pose Estimation (CMUD-Pose), which transfers discriminative knowledge from an RGB-D teacher to a depth-only student network. Furthermore, to mitigate overfitting induced by the modality gap, we propose Cross-Modal Uncertainty Distillation (CMUD), which utilizes a learned uncertainty-aware weighting mechanism for adaptively assigning importance to training samples. By incorporating uncertainty into the distillation process, CMUD allows the student model to focus selectively on reliable and transferable cross-modal features. Extensive experiments on the REAL275 and CAMERA25 benchmarks show that our method significantly improves the performance of depth-only pose estimation models.
Tianfu Wang 0003, Guosheng Hu, Hongguang Wang
IEEE Signal Process. Lett.2
2025 Beyond the Answer: Advancing Multi-Hop QA with Fine-Grained Graph Reasoning and Evaluation
abstract
Recent advancements in large language models (LLMs) have significantly improved the performance of multi-hop question answering (MHQA) systems. Despite the success of MHQA systems, the evaluation of MHQA is not deeply investigated. Existing evaluations mainly focus on comparing the final answers of the reasoning method and given ground-truths. We argue that the reasoning process should also be evaluated because wrong reasoning process can also lead to the correct final answers. Motivated by this, we propose a “Planner-Executor-Reasoner” (PER) architecture, which forms the core of the Plan-anchored Data Preprocessing (PER-DP) and the Plan-guided Multi-Hop QA (PER-QA).The former provides the ground-truth of intermediate reasoning steps and final answers, and the latter offers them of a reasoning method. Moreover, we design a fine-grained evaluation metric called Plan-aligned Stepwise Evaluation (PSE), which evaluates the intermediate reasoning steps from two aspects: planning and solving. Extensive experiments on ten types of questions demonstrate competitive reasoning performance, improved explainability of the MHQA system, and uncover issues such as “fortuitous reasoning continuance” and “latent reasoning suspension” in RAG-based MHQA systems. Besides, we also demonstrate the potential of our approach in data contamination scenarios.
Qichuan Liu, Chentao Zhang, Chenfeng Zheng, Guosheng Hu
ACL (1)4
2025 DPL++: Advancing the Network Performance via Image and Label Perturbations
abstract
Recent advances in supervised learning have predominantly focused on regularizations, optimizers, and architectures, yet the potential of simultaneously optimizing data distributions and supervisory signals for training samples remains underexplored. In this paper, we propose a novel paradigm that leverages the benefits of image perturbations for rectifying data distributions. Our method, called DPL (Deep Perturbation Learning), introduces new insights into utilizing image perturbations and focuses on improving generalizability on normal samples, rather than resisting adversarial attacks. DPL formulates a differentiable function w.r.t. image perturbations and implements an alternative optimization process that seamlessly integrates with downstream tasks. However, the limitations of DPL stem from the inefficiency in employing differentiable targets caused by the exclusive optimization of image perturbations, while neglecting the critical role of supervisory signals in training effectiveness. These lead to the excessive necessity of DPL iterations and yield inferior performance-cost trade-off. To track this, we extend DPL to DPL++ with synchronous optimization for image perturbations and label perturbations. In our DPL++ paradigm, the post-hoc application of perturbations to images and labels endows amendments toward both data distributions and supervisory signals, significantly furthering the generalizability of models over various benchmarks. Crucially, the proposed synchronous optimization process shares key differentiable objectives to reduce computational complexity, thereby achieving enhanced effectiveness within fewer optimization iterations. Theoretically, as a generic and flexible approach, DPL++ can be applied to a variety of backbone architectures (e.g., ResNet, DenseNet, and ViT) and downstream tasks (e.g., image classification and object detection). To validate the efficacy of DPL++, we conduct extensive performance experiments and in-depth analytical studies on 2 visual tasks over 5 mainstream benchmarks across 13 backbone networks. The comprehensive results verify the superiority of DPL++ over DPL and demonstrate its promising capabilities for advancing decision-making capacity, risk minimization, class distinguishability, and training convergence.
Zifan Song, Guosheng Hu, Shuguang Dou, Cairong Zhao
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Explainability-based knowledge distillation
Tianli Sun, Haonan Chen 0003, Guosheng Hu, Cairong Zhao
Pattern Recognit.3
2025 Multi-definition Deepfake detection via semantics reduction and cross-domain training
Cairong Zhao, Chutian Wang, Zifan Song, Guosheng Hu, Duoqian Miao 0001
Pattern Recognit.4
2025 Learning Label Perturbations
abstract
Supervised learning typically uses hard labels for annotations, which may not fully capture the underlying distribution of the data. In the literature, label smoothing is a method that can reduce overconfidence and enhance the model generalization by using a weighted average of one-hot vectors and the uniform distribution, but it does not extract the intrinsic information from the data. Another approach is knowledge distillation, which uses the predicted probability distribution of a teacher network trained with hard labels as soft labels for a student network. However, this method lacks a theoretical explanation. In this work, we draw inspiration from the influence function and propose a post-hoc label perturbation learning method called Deep Soft Label Learning (DSLL). This method iteratively leverages the inherent information present in both the model and data to theoretically determine optimal labels for classification and regression problems. Our experiments demonstrate that DSLL consistently enhances model performance across various tasks, including image classification and object detection.
Zifan Song, Guosheng Hu, Cairong Zhao
IEEE Signal Process. Lett.3
2025 AdvMixUp: Adversarial MixUp Regularization for Deep Learning
abstract
Deep neural networks (DNNs) have shown significant progress in many application fields. However, overfitting remains a significant challenge in their development. While existing data-augmentation techniques such as MixUp have been successful in preventing overfitting, they often fail to generate hard mixed samples near the decision boundary, impeding model optimization. In this article, we present adversarial MixUp (AdvMixUp), a novel sample-dependent method for regularizing DNNs. AdvMixUp addresses this issue by incorporating adversarial training (AT) to create sample-dependent and feature-level interpolation masks, generating more challenging mixed samples. These virtual samples enable DNNs to learn more robust features, ultimately reducing overfitting. Empirical evaluations on CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet demonstrate that AdvMixUp outperforms existing MixUp variants.
Jun Fu 0001, Xianrui Ji, Dexiong Chen, Guosheng Hu, Shuang Li 0008, Xiating Feng
IEEE Trans. Neural Networks Learn. Syst.4
2024 Diverse Person: Customize Your Own Dataset for Text-Based Person Search
abstract
Text-based person search is a challenging task aimed at locating specific target pedestrians through text descriptions. Recent advancements have been made in this field, but there remains a deficiency in datasets tailored for text-based person search. The creation of new, real-world datasets is hindered by concerns such as the risk of pedestrian privacy leakage and the substantial costs of annotation. In this paper, we introduce a framework, named Diverse Person (DP), to achieve efficient and high-quality text-based person search data generation without involving privacy concerns. Specifically, we propose to leverage available images of clothing and accessories as reference attribute images to edit the original dataset images through diffusion models. Additionally, we employ a Large Language Model (LLM) to produce annotations that are both high in quality and stylistically consistent with those found in real-world datasets. Extensive experimental results demonstrate that the baseline models trained with our DP can achieve new state-of-the-art results on three public datasets, with performance improvements up to 4.82%, 2.15%, and 2.28% on CUHK-PEDES, ICFG-PEDES, and RSTPReid in terms of Rank-1 accuracy, respectively.
Zifan Song, Guosheng Hu, Cairong Zhao
AAAI2
2024 Gradient-Guided Modality Decoupling for Missing-Modality Robustness
abstract
Multimodal learning with incomplete input data (missing modality) is very practical and challenging. In this work, we conduct an in-depth analysis of this challenge and find that modality dominance has a significant negative impact on the model training, greatly degrading the missing modality performance. Motivated by Grad-CAM, we introduce a novel indicator, gradients, to monitor and reduce modality dominance which widely exists in the missing-modality scenario. In aid of this indicator, we present a novel Gradient-guided Modality Decoupling (GMD) method to decouple the dependency on dominating modalities. Specifically, GMD removes the conflicted gradient components from different modalities to achieve this decoupling, significantly improving the performance. In addition, to flexibly handle modal-incomplete data, we design a parameter-efficient Dynamic Sharing (DS) framework which can adaptively switch on/off the network parameters based on whether one modality is available. We conduct extensive experiments on three popular multimodal benchmarks, including BraTS 2018 for medical segmentation, CMU-MOSI, and CMU-MOSEI for sentiment analysis. The results show that our method can significantly outperform the competitors, showing the effectiveness of the proposed solutions. Our code is released here: https://github.com/HaoWang420/Gradient-guided-Modality-Decoupling.
Shengda Luo, Guosheng Hu
AAAI3
2024 Neighborhood-Enhanced 3D Human Pose Estimation with Monocular LiDAR in Long-Range Outdoor Scenes
abstract
3D human pose estimation (3HPE) in large-scale outdoor scenes using commercial LiDAR has attracted significant attention due to its potential for real-life applications. However, existing LiDAR-based methods for 3HPE primarily rely on recovering 3D human poses from individual point clouds, and the coherence cues present in the neighborhood are not sufficiently harnessed. In this work, we explore spatial and contexture coherence cues contained in the neighborhood that lead to great performance improvements in 3HPE. Specifically, firstly, we deeply investigate the 3D neighbor in the background (3BN) which serves as a spatial coherence cue for inferring reliable motion since it provides physical laws to limit motion targets. Secondly, we introduce a novel 3D scanning neighbor (3SN) generated during the data collection and 3SN implies structural edge coherence cues. We use 3SN to overcome the degradation of performance and data quality caused by the sparsity-varying properties of LiDAR point clouds. In order to effectively model the complementation between these distinct cues and build consistent temporal relationships across human motions, we propose a new transformer-based module called the CoherenceFuse module. Extensive experiments were conducted on publicly available datasets, namely LidarHuman26M, CIMI4D, SLOPER4D and Waymo Open Dataset v2.0, showcase the superiority and effectiveness of our proposed method. In particular, when compared with LidarCap on the LidarHuman26M dataset, our method demonstrates a reduction of 7.08mm in the average MPJPE metric, along with a decrease of 16.55mm in the MPJPE metric for distances exceeding 25 meters. The code and models are available at https://github.com/jingyi-zhang/Neighborhood-enhanced-LidarCap.
Qihong Mao, Guosheng Hu, Cheng Wang 0003
AAAI3
2024 Object Pose Estimation via the Aggregation of Diffusion Features
abstract
Estimating the pose of objects from images is a crucial task of 3D scene understanding, and recent approaches have shown promising results on very large benchmarks. However, these methods experience a significant performance drop when dealing with unseen objects. We believe that it results from the limited generalizability of image features. To address this problem, we have an in-depth analysis on the features of diffusion models, e.g. Stable Diffusion, which hold substantial potential for modeling unseen objects. Based on this analysis, we then innovatively introduce these diffusion features for object pose estimation. To achieve this, we propose three distinct architectures that can effectively capture and aggregate diffusion features of different granularity, greatly improving the generalizability of object pose estimation. Our approach outperforms the state-of-the-art methods by a considerable margin on three popular benchmark datasets, LM, O-LM, and T-LESS. In particular, our method achieves higher accuracy than the previous best arts on unseen objects: 98.2% vs. 93.5% on Unseen LM, 85.9% vs. 76.3% on Unseen O-LM, showing the strong generalizability of our method. Our code is released at https://github.com/Tianfu18/diff-feats-pose.
Tianfu Wang 0003, Guosheng Hu, Hongguang Wang
CVPR2
2024 DiffLoc: Diffusion Model for Outdoor LiDAR Localization
abstract
Absolute pose regression (APR) estimates global pose in an end-to-end manner, achieving impressive results in learn-based LiDAR localization. However, compared to the top-performing methods reliant on 3D-3D correspondence matching, APR's accuracy still has room for improvement. We recognize APR's lack of robust features learning and iterative denoising process leads to suboptimal results. In this paper, we propose DiffLoc, a novel framework that formulates LiDAR localization as a conditional generation of poses. First, we propose to utilize the foundation model and static-object-aware pool to learn robust features. Second, we incorporate the iterative denoising process into APR via a diffusion model conditioned on the learned geometrically robust features. In addition, due to the unique nature of diffusion models, we propose to adapt our models to two additional applications: (1) using multiple inferences to evaluate pose uncertainty, and (2) seamlessly introducing geometric constraints on denoising steps to improve prediction accuracy. Extensive experiments conducted on the Oxford Radar RobotCar and NCLT datasets demonstrate that DiffLoc outperforms better than the state-of-the-art methods. Especially on the NCLT dataset, we achieve 35% and 34.7% improvement on position and orientation accuracy, respectively. Our code is released at https://github.com/liw95/DiffLoc.
Wen Li 0005, Yuyang Yang, Shangshu Yu, Guosheng Hu, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003
CVPR4
2024 Unsupervised Exposure Correction
Ruodai Cui, Li Niu 0002, Guosheng Hu
ECCV (6)3
2024 Learning Scene-Pedestrian Graph for End-to-End Person Search
abstract
Person search aims to find specific persons from visual scenes, including two subtasks, pedestrian detection, and person reidentification. The dominant fashion in this area is end-to-end networks that focus on analyzing the foreground (i.e., pedestrian) while ignoring the background (i.e., scene) information. However, the scene information often offers useful clues for person search. For example, pedestrians normally appear on the road rather than the top of a tree, and pedestrians appearing at the same location are likely to have similar occlusions. The interplay between the pedestrians and scenes can potentially improve the performance. In this article, a novel scene-pedestrian graph (SPG) is proposed, which can explicitly model the interplay between the pedestrians and scenes. To polish the quality of pedestrian bounding boxes, we pioneer a strategy of using the high-quality pedestrian bounding box to guide the low-quality one in the same scene. In addition, we design a contextual and temporal graph matching algorithm to effectively utilize the contextual and temporal information present in the constructed SPG to improve the performance of pedestrian matching. Benefiting from the robustness on complex scenes, our model achieves promising performance over the state-of-the-art methods on two popular person search benchmarks, CUHK-SYSU and PRW.
Zifan Song, Cairong Zhao, Guosheng Hu, Duoqian Miao 0001
IEEE Trans. Ind. Informatics3
2024 NIDALoc: Neurobiologically Inspired Deep LiDAR Localization
abstract
Absolute pose regression has shown great potential in LiDAR localization, which learns to regress 6-DoF LiDAR poses through deep networks. However, recent regression methods suffer from scene ambiguities in challenging scenarios, leading to inaccurate and unstable localization. Inspired by neurobiological localization mechanisms, i.e., the firing mechanism of place cells, head-direction cells, and grid cells in mammalian brains, we propose a novel LiDAR localization framework called NIDALoc to achieve more robust and accurate results. First, we propose a Hebbian memory module, motivated by place cells, to preserve historical information, which helps refine local view features to reduce scene ambiguities. Specifically, the memory module stores scene information and then recalls it when revisiting an old place. Second, we propose a novel pose constrained framework, consisting of an orientation classification task and a grid center regression task, to regularize orientation and position estimation, respectively. The framework based on head-direction cells and grid cells constrains the absolute pose regression to reduce wrong predictions. Extensive experiments on two outdoor datasets demonstrate the effectiveness of NIDALoc, which outperforms state-of-the-art localization methods, especially in large-scale challenging scenes. The source code is available on the project website at https://github.com/PSYZ1234/NIDALoc.
Shangshu Yu, Xiaotian Sun 0005, Wen Li 0005, Chenglu Wen, Yunuo Yang, Bailu Si, Guosheng Hu, Cheng Wang 0003
IEEE Trans. Intell. Transp. Syst.7
2024 Explainability of Speech Recognition Transformers via Gradient-Based Attention Visualization
abstract
In vision Transformers, attention visualization methods are used to generate heatmaps highlighting the class-corresponding areas in input images, which offers explanations on how the models make predictions. However, it is not so applicable for explaining automatic speech recognition (ASR) Transformers. An ASR Transformer makes a particular prediction for every input token to form a sentence, but a vision Transformer only makes an overall classification for the input data. Therefore, traditional attention visualization methods may fail in ASR Transformers. In this work, we propose a novel attention visualization method in ASR Transformers and try to explain which frames of the audio result in the output text. Inspired by the model explainability, we also explore ways of improving the effectiveness of the ASR model. Comparing with other Transformer attention visualization methods, our method is more efficient and intuitively understandable, which unravels the attention calculation from information flow of Transformer attention modules. In addition, we demonstrate the utilization of visualization result in three ways: (1) We visualize attention with respect to connectionist temporal classification (CTC) loss to train an ASR model with adversarial attention erasing regularization, which effectively decreases the word error rate (WER) of the model and improves its generalization capability. (2) We visualize the attention on some specific words, interpreting the model by effectively demonstrating the semantic and grammar relationships between these words. (3) Similarly, we analyze how the model manage to distinguish homophones, using contrastive explanation with respect to homophones.
Tianli Sun, Haonan Chen 0003, Guosheng Hu, Lianghua He, Cairong Zhao
IEEE Trans. Multim.3
2024 Deep Metric Learning Based on Meta-Mining Strategy With Semiglobal Information
abstract
Recently, deep metric learning (DML) has achieved great success. Some existing DML methods propose adaptive sample mining strategies, which learn to weight the samples, leading to interesting performance. However, these methods suffer from a small memory (e.g., one training batch), limiting their efficacy. In this work, we introduce a data-driven method, meta-mining strategy with semiglobal information (MMSI), to apply meta-learning to learn to weight samples during the whole training, leading to an adaptive mining strategy. To introduce richer information than one training batch only, we elaborately take advantage of the validation set of meta-learning by implicitly adding additional validation sample information to training. Furthermore, motivated by the latest self-supervised learning, we introduce a dictionary (memory) that maintains very large and diverse information. Together with the validation set, this dictionary presents much richer information to the training, leading to promising performance. In addition, we propose a new theoretical framework that can formulate pairwise and tripletwise metric learning loss functions in a unified framework. This framework brings new insights to society and facilitates us to generalize our MMSI to many existing DML methods. We conduct extensive experiments on three public datasets, CUB200-2011, Cars-196, and Stanford Online Products (SOP). Results show that our method can achieve the state of the art or very competitive performance. Our source codes have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/MMSI.
Xi Jiang 0001, Sheng Liu 0009, Xili Dai, Guosheng Hu, Xingguo Huang, Yazhou Yao, Guosen Xie, Ling Shao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Cross-Modal Distillation for Speaker Recognition
abstract
Speaker recognition achieved great progress recently, however, it is not easy or efficient to further improve its performance via traditional solutions: collecting more data and designing new neural networks. Aiming at the fundamental challenge of speech data, i.e. low information density, multimodal learning can mitigate this challenge by introducing richer and more discriminative information as input for identity recognition. Specifically, since the face image is more discriminative than the speech for identity recognition, we conduct multimodal learning by introducing a face recognition model (teacher) to transfer discriminative knowledge to a speaker recognition model (student) during training. However, this knowledge transfer via distillation is not trivial because the big domain gap between face and speech can easily lead to overfitting. In this work, we introduce a multimodal learning framework, VGSR (Vision-Guided Speaker Recognition). Specifically, we propose a MKD (Margin-based Knowledge Distillation) strategy for cross-modality distillation by introducing a loose constrain to align the teacher and student, greatly reducing overfitting. Our MKD strategy can easily adapt to various existing knowledge distillation methods. In addition, we propose a QAW (Quality-based Adaptive Weights) module to weight input samples via quantified data quality, leading to a robust model training. Experimental results on the VoxCeleb1 and CN-Celeb datasets show our proposed strategies can effectively improve the accuracy of speaker recognition by a margin of 10% ∼ 15%, and our methods are very robust to different noises.
Yufeng Jin, Guosheng Hu, Haonan Chen 0003, Duoqian Miao 0001, Liang Hu 0001, Cairong Zhao
AAAI2
2023 SGLoc: Scene Geometry Encoding for Outdoor LiDAR Localization
abstract
LiDAR-based absolute pose regression estimates the global pose through a deep network in an end-to-end manner, achieving impressive results in learning-based localization. However, the accuracy of existing methods still has room to improve due to the difficulty of effectively encoding the scene geometry and the unsatisfactory quality of the data. In this work, we propose a novel LiDAR localization frame-work, SGLoc, which decouples the pose estimation to point cloud correspondence regression and pose estimation via this correspondence. This decoupling effectively encodes the scene geometry because the decoupled correspondence regression step greatly preserves the scene geometry, leading to significant performance improvement. Apart from this decoupling, we also design a tri-scale spatial feature aggregation module and inter-geometric consistency constraint loss to effectively capture scene geometry. Moreover, we empirically find that the ground truth might be noisy due to GPS/INS measuring errors, greatly reducing the pose estimation performance. Thus, we propose a pose quality evaluation and enhancement method to measure and correct the ground truth pose. Extensive experiments on the Oxford Radar RobotCar and NCLT datasets demonstrate the effectiveness of SGLoc, which outperforms state-of-the-art regression-based localization methods by 68.5% and 67.6% on position accuracy, respectively.
Wen Li 0005, Shangshu Yu, Cheng Wang 0003, Guosheng Hu, Chenglu Wen
CVPR4
2023 Deep Perturbation Learning: Enhancing the Network Performance via Image Perturbations
abstract
Image perturbation technique is widely used to generate adversarial examples to attack networks, greatly decreasing the performance of networks. Unlike the existing works, in this paper, we introduce a novel framework Deep Perturbation Learning (DPL), the new insights into understanding image perturbations, to enhance the performance of networks rather than decrease the performance. Specifically, we learn image perturbations to amend the data distribution of training set to improve the performance of networks. This optimization w.r.t data distribution is non-trivial. To approach this, we tactfully construct a differentiable optimization target w.r.t. image perturbations via minimizing the empirical risk. Then we propose an alternating optimization of the network weights and perturbations. DPL can easily be adapted to a wide spectrum of downstream tasks and backbone networks. Extensive experiments demonstrate the effectiveness of our DPL on 6 datasets (CIFAR-10, CIFAR100, ImageNet, MS-COCO, PASCAL VOC, and SBD) over 3 popular vision tasks (image classification, object detection, and semantic segmentation) with different backbone architectures (e.g., ResNet, MobileNet, and ViT).
Zifan Song, Guosheng Hu, Cairong Zhao
ICML3
2023 ISTVT: Interpretable Spatial-Temporal Video Transformer for Deepfake Detection
abstract
With the rapid development of Deepfake synthesis technology, our information security and personal privacy have been severely threatened in recent years. To achieve a robust Deepfake detection, researchers attempt to exploit the joint spatial-temporal information in the videos, like using recurrent networks and 3D convolutional networks. However, these spatial-temporal models remain room to improve. Another general challenge for spatial-temporal models is that people do not clearly understand what these spatial-temporal models really learn. To address these two challenges, in this paper, we propose an Interpretable Spatial-Temporal Video Transformer (ISTVT), which consists of a novel decomposed spatial-temporal self-attention and a self-subtract mechanism to capture spatial artifacts and temporal inconsistency for robust Deepfake detection. Thanks to this decomposition, we propose to interpret ISTVT by visualizing the discriminative regions for both spatial and temporal dimensions via the relevance (the pixel-wise importance on the input) propagation algorithm. We conduct extensive experiments on large-scale datasets, including FaceForensics++, FaceShifter, DeeperForensics, Celeb-DF, and DFDC datasets. Our strong performance of intra-dataset and cross-dataset Deepfake detection demonstrates the effectiveness and robustness of our method, and our visualization-based interpretability offers people insights into our model.
Cairong Zhao, Chutian Wang, Guosheng Hu, Haonan Chen 0003, Chun Liu 0003, Jinhui Tang 0001
IEEE Trans. Inf. Forensics Secur.3
2023 STCLoc: Deep LiDAR Localization With Spatio-Temporal Constraints
abstract
LiDAR localization is of great importance to autonomous vehicles and robotics. Absolute pose regression, directly estimating the mapping from a scene to a 6-DoF pose, has achieved impressive results in learning-based localization. Different from traditional map-based methods, it does not need a pre-built 3D map during inference. However, current regression networks typically suffer from scene ambiguities, especially in challenging traffic environments, leading to large wrong predictions (e.g., outliers) and limited applications. To address this problem, a novel LiDAR localization framework with spatio-temporal constraints is proposed, termed STCLoc, to reduce scene ambiguities and achieve more accurate localization. First, we propose to regularize regression in the spatial dimension with a novel classification task to reduce outliers. Specifically, the classification task categorizes the point cloud in terms of position and orientation and then couples it with the regression task to conduct multi-task learning. Second, to learn discriminative features to reduce scene ambiguities, we propose using attention-based feature aggregation to capture the correlation in LiDAR sequences. We conduct extensive experiments on two benchmark datasets, where the localization takes 97ms on each dataset. Results show that our model outperforms state-of-the-art methods by 43.33%/36.76% (position/orientation) on the Oxford Radar RobotCar dataset, verifying the effectiveness of our method. The source code is available on the project website athttps://github.com/PSYZ1234/STCLoc.
Shangshu Yu, Cheng Wang 0003, Yitai Lin, Chenglu Wen, Ming Cheng 0002, Guosheng Hu
IEEE Trans. Intell. Transp. Syst.6
2023 Discrepancy-Guided Domain-Adaptive Data Augmentation
abstract
Data augmentation has been observed playing a crucial role in achieving better generalization in many machine learning tasks, especially in unsupervised domain adaptation (DA). It is particularly effective on visual object recognition tasks as images are high-dimensional with an enormous range of variations that can be simulated. Existing data augmentation techniques, however, are not explicitly designed to address the differences between different domains. Expert knowledge about the data is required, as well as manual efforts in finding the optimal parameters. In this article, we propose a novel domain-adaptive augmentation method by making use of a state-of-the-art style transfer method and domain discrepancy measurement. Specifically, we measure the discrepancy between source and target domains, and use it as a guide to augment the original source samples using style transferred source-to-target samples. The proposed domain-adaptive augmentation method is data and model agnostic that can be easily incorporated with state-of-the-art DA algorithms. We show empirically that, by using this domain-adaptive augmentation, we are able to gradually reduce the discrepancy between the source and target samples, and further boost the adaptation performance using different DA algorithms on three popular domain adaption datasets.
Jian Gao 0018, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
IEEE Trans. Neural Networks Learn. Syst.3
2022 Boosting Active Learning via Improving Test Performance
abstract
Central to active learning (AL) is what data should be selected for annotation. Existing works attempt to select highly uncertain or informative data for annotation. Nevertheless, it remains unclear how selected data impacts the test performance of the task model used in AL. In this work, we explore such an impact by theoretically proving that selecting unlabeled data of higher gradient norm leads to a lower upper-bound of test loss, resulting in a better test performance. However, due to the lack of label information, directly computing gradient norm for unlabeled data is infeasible. To address this challenge, we propose two schemes, namely expected-gradnorm and entropy-gradnorm. The former computes the gradient norm by constructing an expected empirical loss while the latter constructs an unsupervised loss with entropy. Furthermore, we integrate the two schemes in a universal AL framework. We evaluate our method on classical image classification and semantic segmentation tasks. To demonstrate its competency in domain applications and its robustness to noise, we also validate our method on a cellular imaging analysis task, namely cryo-Electron Tomography subtomogram classification. Results demonstrate that our method achieves superior performance against the state of the art. We refer readers to https://arxiv.org/pdf/2112.05683.pdf for the full version of this paper which includes the appendix and source code link.
Tianyang Wang 0004, Xingjian Li 0002, Pengkun Yang, Guosheng Hu, Siyu Huang, Cheng-Zhong Xu 0001, Min Xu 0009
AAAI4
2022 Efficient One-Stage Video Object Detection by Exploiting Temporal Consistency
Guanxiong Sun, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
ECCV (35)3
2022 TDViT: Temporal Dilated Video Transformer for Dense Video Tasks
Guanxiong Sun, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
ECCV (35)3
2022 RCANet: Row-Column Attention Network for Semantic Segmentation
abstract
Establishing high-order interactions among pixels and object parts is one of the most fundamental problems in semantic segmentation. The recent proposals are based on non-local methods which utilize the self-attention mechanism to capture the long-range correlations. However, non-local methods could be very expensive, both theoretically and experimentally. Moreover, non-local methods are typically designed to address spatial correlations rather than feature correlations across channels. In this work, we propose a Row-Column Attention Network (RCANet) to encode globally contextual information. It consists of a row-wise intra-channel attention module and a column-wise intra-channel attention module, followed by a cross-channel interaction module. We conduct experiments on two datasets: Cityscapes and ADE20K. The results show that our method is comparable to the state-of-the-art methods for semantic segmentation.
Bingxu Lu, Qinghua Hu, Yu Wang 0106, Guosheng Hu
ICASSP4
2022 Multi-Definition Video Deepfake Detection via Semantics Reduction and Cross-Domain Training
abstract
The recent development of Deepfake videos directly threatens our information security and personal privacy. Although lots of previous works have made much progress on the Deepfake detection, we empirically find that the existing approaches do not perform well on the low definition (LD) and crossdefinition (high and low) videos. To address this problem, in this paper, we follow two motivations: (1) high-level semantics reduction and (2) cross-domain training. For (1), we propose the Facial Structure Destruction and Adversarial Jigsaw Loss to reduce our model to learn high-level semantics and focus on learning low-level discriminative information; For (2), we propose a domain generalization method based on adversarial learning. We conduct extensive experiments on the FaceForensics++ dataset. Results show the great effectiveness of our method and we also achieve very competitive performance against state-of-the-art methods.
Chutian Wang, Cairong Zhao, Guosheng Hu
ICME3
2022 Self-attention neural architecture search for semantic image segmentation
Zhenkun Fan, Guosheng Hu, Xin Sun 0003, Gaige Wang, Junyu Dong, Chi Su
Knowl. Based Syst.2
2022 MetaMixUp: Learning Adaptive Interpolation Policy of MixUp With Metalearning
abstract
MixUp is an effective data augmentation method to regularize deep neural networks via random linear interpolations between pairs of samples and their labels. It plays an important role in model regularization, semisupervised learning (SSL), and domain adaption. However, despite its empirical success, its deficiency of randomly mixing samples has poorly been studied. Since deep networks are capable of memorizing the entire data set, the corrupted samples generated by vanilla MixUp with a badly chosen interpolation policy will degrade the performance of networks. To overcome overfitting to corrupted samples, inspired by metalearning (learning to learn), we propose a novel technique of learning to a mixup in this work, namely, MetaMixUp. Unlike the vanilla MixUp that samples interpolation policy from a predefined distribution, this article introduces a metalearning-based online optimization approach to dynamically learn the interpolation policy in a data-adaptive way (learning to learn better). The validation set performance via metalearning captures the noisy degree, which provides optimal directions for interpolation policy learning. Furthermore, we adapt our method for pseudolabel-based SSL along with a refined pseudolabeling strategy. In our experiments, our method achieves better performance than vanilla MixUp and its variants under SL configuration. In particular, extensive experiments show that our MetaMixUp adapted SSL greatly outperforms MixUp and many state-of-the-art methods on CIFAR-10 and SVHN benchmarks under the SSL configuration.
Zhijun Mai, Guosheng Hu, Dexiong Chen, Fumin Shen, Heng Tao Shen
IEEE Trans. Neural Networks Learn. Syst.2
2021 MAMBA: Multi-level Aggregation via Memory Bank for Video Object Detection
abstract
State-of-the-art video object detection methods maintain a memory structure, either a sliding window or a memory queue, to enhance the current frame using attention mechanisms. However, we argue that these memory structures are not efficient or sufficient because of two implied operations: (1) concatenating all features in memory for enhancement, leading to a heavy computational cost; (2) frame-wise memory updating, preventing the memory from capturing more temporal information. In this paper, we propose a multi-level aggregation architecture via memory bank called MAMBA. Specifically, our memory bank employs two novel operations to eliminate disadvantages of existing methods: (1) light-weight key-set construction which can significantly reduce the computational cost; (2) fine-grained feature-wise updating strategy which enables our method to utilize knowledge from the whole video. To better enhance features from complementary levels, i.e., feature maps and proposals, we further propose a generalized enhancement operation (GEO) to aggregate multi-level features in a unified manner. We conduct extensive evaluations on the challenging ImageNetVID dataset. Compared with existing state-of-the-art methods, our method achieves superior performance in terms of both speed and accuracy. More remarkably, MAMBA achieves mAP of 83.7%/84.6% at 12.6/9.1 FPS with ResNet-101.
Guanxiong Sun, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
AAAI3
2021 OPANAS: One-Shot Path Aggregation Network Architecture Search for Object Detection
abstract
Recently, neural architecture search (NAS) has been exploited to design feature pyramid networks (FPNs) and achieved promising results for visual object detection. Encouraged by the success, we propose a novel One-Shot Path Aggregation Network Architecture Search (OPANAS) algorithm, which significantly improves both searching efficiency and detection accuracy. Specifically, we first introduce six heterogeneous information paths to build our search space, namely top-down, bottom-up, fusing-splitting, scale-equalizing, skip-connect and none. Second, we propose a novel search space of FPNs, in which each FPN candidate is represented by a densely-connected directed acyclic graph (each node is a feature pyramid and each edge is one of the six heterogeneous information paths). Third, we propose an efficient one-shot search method to find the optimal path aggregation architecture; specifically, we first train a super-net and then find the optimal candidate with an evolutionary algorithm. Experimental results demonstrate the efficacy of the proposed OPANAS for object detection: (1) OPANAS is more efficient than state-of-the-art methods (e.g., NAS-FPN and Auto-FPN) at significantly smaller searching cost (e.g., only 4 GPU days on MS-COCO); (2) the optimal architecture found by OPANAS significantly improves main-stream detectors including RetinaNet, Faster R-CNN and Cascade R-CNN, by 2.3∼3.2 % mAP compared to their FPN counterparts; and (3) a new state-of-the-art accuracy-speed trade-off (52.2 % mAP at 7.6 FPS) is achieved at smaller training costs than comparable recent arts. Code will be released at https://github.com/VDIGPKU/OPANAS.
Tingting Liang, Yongtao Wang, Zhi Tang 0001, Guosheng Hu, Haibin Ling
CVPR4
2021 Refining Single Low-Quality Facial Depth Map by Lightweight and Efficient Deep Model
abstract
Consumer depth sensors have become increasingly common, however, the data are rather coarse and noisy, which is problematic to delicate tasks, such as 3D face modeling and 3D face recognition. In this paper, we present a novel and lightweight 3D Face Refinement Model (3D-FRM), to effectively and efficiently improve the quality of such single facial depth maps. 3D-FRM has an encoder-decoder structure, where the encoder applies depth-wise, point-wise convolutions and the fusion of features of different receptive fields to capture original discriminative information, and the decoder exploits sub-pixel convolutions and the combination of low- and high-level features to achieve strong shape recovery. We also propose a joint loss function to smooth facial surfaces and preserve their identities. In addition, we contribute a large dataset with low- and high-quality 3D face pairs to facilitate this research. Extensive experiments are conducted on the Bosphorus and Lock3DFace datasets, and results show the competency of the proposed method at ameliorating both visual quality and recognition accuracy. Code and data will be available at https://github.com/muyouhang/3D-FRM.
Guodong Mu, Di Huang 0001, Weixin Li 0001, Guosheng Hu, Yunhong Wang 0001
IJCB4
2021 DPT: Deformable Patch-based Transformer for Visual Recognition
abstract
Transformer has achieved great success in computer vision, while how to split patches in an image remains a problem. Existing methods usually use a fixed-size patch embedding which might destroy the semantics of objects. To address this problem, we propose a new Deformable Patch (DePatch) module which learns to adaptively split the images into patches with different positions and scales in a data-driven way rather than using predefined fixed patches. In this way, our method can well preserve the semantics in patches. The DePatch module can work as a plug-and-play module, which can easily be incorporated into different transformers to achieve an end-to-end training. We term this DePatch-embedded transformer as Deformable Patch-based Transformer (DPT) and conduct extensive evaluations of DPT on image classification and object detection. Results show DPT can achieve 81.8% top-1 accuracy on ImageNet classification, and 43.7% box AP with RetinaNet, 44.3% with Mask R-CNN on MSCOCO object detection. Code has been made available at: https://github.com/CASIA-IVA-Lab/DPT.
Zhiyang Chen 0002, Yousong Zhu, Chaoyang Zhao, Guosheng Hu, Wei Zeng 0006, Jinqiao Wang, Ming Tang 0001
ACM Multimedia4
2021 Semi-Supervised Face Frontalization in the Wild
abstract
Synthesizing a frontal view face from a single nonfrontal image, i.e. face frontalization, is a task of practical importance in a wide range of facial image analysis applications. However, to train the frontalization model in a supervised manner, most existing face frontalization methods rely on the availability of nonfrontal-frontal face pairs (typically from the Multi-PIE dataset) captured in a constrained environment. Such approaches, in return, limit the generalizability of their application to unconstrained scenarios. Unfortunately, although a large amount of in-the-wild face datasets are available, they cannot easily be utilized for face frontalization training since the nonfrontal and frontal facial images are not paired. To train a frontalization network which generalizes well to both constrained and unconstrained environments, we propose a semi-supervised learning framework which effectively uses both (labeled) indoor and (unlabeled) outdoor faces. Specifically, to achieve this goal, this article presents a Cycle-Consistent Face Frontalization Generative Adversarial Network (CCFF-GAN) which consists of both (1) the supervised and (2) the unsupervised components. For (1), we use the indoor paired (labeled) data to learn a roughly accurate frontalization network which may not generalize well to outdoor (in-the-wild) scenarios. For (2), to cope with the generalization issue, the unsupervised part uses the unpaired (unlabeled) images under the perceptual cycle consistency constraint in the semantic feature space to generalize the network from controlled (indoor) to uncontrolled (outdoor) environment. Extensive experiments demonstrate the effectiveness of the proposed method in comparison with the state-of-the-art face frontalization methods, especially under the in-the-wild scenarios.
Zhihong Zhang 0001, Ruiyang Liang, Xu Chen 0020, Xuexin Xu, Guosheng Hu, Wangmeng Zuo, Edwin R. Hancock
IEEE Trans. Inf. Forensics Secur.5
2021 SP-GAN: Self-Growing and Pruning Generative Adversarial Networks
abstract
This article presents a new Self-growing and Pruning Generative Adversarial Network (SP-GAN) for realistic image generation. In contrast to traditional GAN models, our SP-GAN is able to dynamically adjust the size and architecture of a network in the training stage by using the proposed self-growing and pruning mechanisms. To be more specific, we first train two seed networks as the generator and discriminator; each contains a small number of convolution kernels. Such small-scale networks are much easier and faster to train than large-capacity networks. Second, in the self-growing step, we replicate the convolution kernels of each seed network to augment the scale of the network, followed by fine-tuning the augmented/expanded network. More importantly, to prevent the excessive growth of each seed network in the self-growing stage, we propose a pruning strategy that reduces the redundancy of an augmented network, yielding the optimal scale of the network. Finally, we design a new adaptive loss function that is treated as a variable loss computational process for the training of the proposed SP-GAN model. By design, the hyperparameters of the loss function can dynamically adapt to different training stages. Experimental results obtained on a set of data sets demonstrate the merits of the proposed method, especially in terms of the stability and efficiency of network training. The source code of the proposed SP-GAN method is publicly available at https://github.com/Lambert-chen/SPGAN.git.
Xiaoning Song, Zhenhua Feng 0001, Guosheng Hu, Dongjun Yu, Xiaojun Wu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2020 Imbalance Robust Softmax for Deep Embeeding Learning
Hao Zhu 0010, Guosheng Hu, Neil Robertson 0002
ACCV (5)3
2020 Training Noise-Robust Deep Neural Networks via Meta-Learning
abstract
Label noise may significantly degrade the performance of Deep Neural Networks (DNNs). To train noise-robust DNNs, Loss correction (LC) approaches have been introduced. LC approaches assume the noisy labels are corrupted from clean (ground-truth) labels by an unknown noise transition matrix T. The backbone DNNs and T can be trained separately, where T is approximated with prior knowledge. For example, T is constructed by stacking the maximum or mean predictions of the samples from each class. In this work, we pro- pose a new loss correction approach, named as Meta Loss Correction (MLC), to directly learn T from data via the meta-learning framework. The MLC is model-agnostic and learns T from data rather than heuristically approximates it using prior knowledge. Extensive evaluations are conducted on computer vision (MNIST, CIFAR-10, CIFAR-100, Clothing1M) and natural language processing (Twitter) datasets. The experimental results show that MLC achieves very competitive performance against state-of-the-art approaches.
Zhen Wang 0033, Guosheng Hu, Qinghua Hu
CVPR2
2020 Reducing Distributional Uncertainty by Mutual Information Maximisation and Transferable Feature Learning
Jian Gao 0018, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
ECCV (23)3
2020 Differentiable Automatic Data Augmentation
Yonggang Li 0001, Guosheng Hu, Yongtao Wang, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang
ECCV (22)2
2020 Learning Flow-Based Feature Warping for Face Frontalization with Illumination Inconsistent Supervision
Yuxiang Wei 0001, Ming Liu 0018, Haolin Wang 0004, Ruifeng Zhu, Guosheng Hu, Wangmeng Zuo
ECCV (12)5
2020 Adaptive Variance Based Label Distribution Learning for Facial Age Estimation
Xin Wen 0005, Biying Li, Haiyun Guo, Zhiwei Liu 0004, Guosheng Hu, Ming Tang 0001, Jinqiao Wang
ECCV (23)5
2020 CRSSC: Salvage Reusable Samples from Noisy Data for Robust Learning
abstract
Due to the existence of label noise in web images and the high memorization capacity of deep neural networks, training deep fine-grained (FG) models directly through web images tends to have an inferior recognition ability. In the literature, to alleviate this issue, loss correction methods try to estimate the noise transition matrix, but the inevitable false correction would cause severe accumulated errors. Sample selection methods identify clean ("easy") samples based on the fact that small losses can alleviate the accumulated errors. However, "hard" and mislabeled examples that can both boost the robustness of FG models are also dropped. To this end, we propose a certainty-based reusable sample selection and correction approach, termed as CRSSC, for coping with label noise in training deep FG models with web images. Our key idea is to additionally identify and correct reusable samples, and then leverage them together with clean examples to update the networks. We demonstrate the superiority of the proposed approach from both theoretical and experimental perspectives.
Zeren Sun, Xian-Sheng Hua 0001, Yazhou Yao, Xiu-Shen Wei, Guosheng Hu, Jian Zhang 0002
ACM Multimedia5
2020 Exploring ubiquitous relations for boosting classification and localization
Xin Sun 0003, Changrui Chen, Junyu Dong, Guosheng Hu
Knowl. Based Syst.5
2020 Attention-Based Two-Stream Convolutional Networks for Face Spoofing Detection
abstract
Since the human face preserves the richest information for recognizing individuals, face recognition has been widely investigated and achieved great success in various applications in the past decades. However, face spoofing attacks (e.g., face video replay attack) remain a threat to modern face recognition systems. Though many effective methods have been proposed for anti-spoofing, we find that the performance of many existing methods is degraded by illuminations. It motivates us to develop illumination-invariant methods for anti-spoofing. In this paper, we propose a two-stream convolutional neural network (TSCNN), which works on two complementary spaces: RGB space (original imaging space) and multi-scale retinex (MSR) space (illumination-invariant space). Specifically, the RGB space contains the detailed facial textures, yet it is sensitive to illumination; MSR is invariant to illumination, yet it contains less detailed facial information. In addition, the MSR images can effectively capture the high-frequency information, which is discriminative for face spoofing detection. Images from two spaces are fed to the TSCNN to learn the discriminative features for anti-spoofing. To effectively fuse the features from two sources (RGB and MSR), we propose an attention-based fusion method, which can effectively capture the complementarity of two features. We evaluate the proposed framework on various databases, i.e., CASIA-FASD, REPLAY-ATTACK, and OULU, and achieve very competitive performance. To further verify the generalization capacity of the proposed strategies, we conduct cross-database experiments, and the results show the great effectiveness of our method.
Haonan Chen 0003, Guosheng Hu, Zhen Lei 0001, Yaowu Chen, Neil Robertson 0002, Stan Z. Li
IEEE Trans. Inf. Forensics Secur.2
2020 Learning Symmetry Consistent Deep CNNs for Face Completion
abstract
Deep convolutional networks (CNNs) have achieved great success in face completion to generate plausible facial structures. These methods, however, are limited in maintaining global consistency among face components and recovering fine facial details. On the other hand, reflectional symmetry is a prominent property of face images and benefits face analysis and consistency modeling, yet remaining uninvestigated in deep face completion. In this work, we leverage two kinds of symmetry-enforcing modules to form a symmetry-consistent CNN model (i.e., SymmFCNet) for effective face completion. For missing pixels on only one of the half-faces, an illumination-reweighted warping subnet is developed to guide the warping and illumination reweighting of the other half-face. As for missing pixels on both of half-faces, we present a generative reconstruction subnet together with a perceptual symmetry loss to enforce symmetry consistency of recovered structures. The SymmFCNet is constructed by stacking generative reconstruction subnet upon illumination-reweighted warping subnet, and can be learned in an end-to-end manner. Experiments show that SymmFCNet can generate globally consistent results on images with synthetic and real occlusions, and performs favorably against state-of-the-arts.
Xiaoming Li 0002, Guosheng Hu, Jieru Zhu, Wangmeng Zuo, Meng Wang 0001, Lei Zhang 0006
IEEE Trans. Image Process.2
2019 Deep Metric Learning by Online Soft Mining and Class-Aware Attention
abstract
Deep metric learning aims to learn a deep embedding that can capture the semantic similarity of data points. Given the availability of massive training samples, deep metric learning is known to suffer from slow convergence due to a large fraction of trivial samples. Therefore, most existing methods generally resort to sample mining strategies for selecting nontrivial samples to accelerate convergence and improve performance. In this work, we identify two critical limitations of the sample mining methods, and provide solutions for both of them. First, previous mining methods assign one binary score to each sample, i.e., dropping or keeping it, so they only selects a subset of relevant samples in a mini-batch. Therefore, we propose a novel sample mining method, called Online Soft Mining (OSM), which assigns one continuous score to each sample to make use of all samples in the mini-batch. OSM learns extended manifolds that preserve useful intraclass variances by focusing on more similar positives. Second, the existing methods are easily influenced by outliers as they are generally included in the mined subset. To address this, we introduce Class-Aware Attention (CAA) that assigns little attention to abnormal data samples. Furthermore, by combining OSM and CAA, we propose a novel weighted contrastive loss to learn discriminative embeddings. Extensive experiments on two fine-grained visual categorisation datasets and two video-based person re-identification benchmarks show that our method significantly outperforms the state-of-the-art.
Xinshao Wang, Yang Hua 0001, Elyor Kodirov, Guosheng Hu, Neil Robertson 0002
AAAI4
2019 Semantic Alignment: Finding Semantically Consistent Ground-Truth for Facial Landmark Detection
abstract
Recently, deep learning based facial landmark detection has achieved great success. Despite this, we notice that the semantic ambiguity greatly degrades the detection performance. Specifically, the semantic ambiguity means that some landmarks (e.g. those evenly distributed along the face contour) do not have clear and accurate definition, causing inconsistent annotations by annotators. Accordingly, these inconsistent annotations, which are usually provided by public databases, commonly work as the ground-truth to supervise network training, leading to the degraded accuracy. To our knowledge, little research has investigated this problem. In this paper, we propose a novel probabilistic model which introduces a latent variable, i.e. the `real' ground-truth which is semantically consistent, to optimize. This framework couples two parts (1) training landmark detection CNN and (2) searching the `real' ground-truth. These two parts are alternatively optimized: the searched `real' ground-truth supervises the CNN training; and the trained CNN assists the searching of `real' ground-truth. In addition, to recover the unconfidently predicted landmarks due to occlusion and low quality, we propose a global heatmap correction unit (GHCU) to correct outliers by considering the global face shape as a constraint. Extensive experiments on both image-based (300W and AFLW) and video-based (300-VW) databases demonstrate that our method effectively improves the landmark detection accuracy and achieves the state of the art performance.
Zhiwei Liu 0004, Xiangyu Zhu 0001, Guosheng Hu, Haiyun Guo, Ming Tang 0001, Zhen Lei 0001, Neil Robertson 0002, Jinqiao Wang
CVPR3
2019 Led3D: A Lightweight and Efficient Deep Approach to Recognizing Low-Quality 3D Faces
abstract
Due to the intrinsic invariance to pose and illumination changes, 3D Face Recognition (FR) has a promising potential in the real world. 3D FR using high-quality faces, which are of high resolutions and with smooth surfaces, have been widely studied. However, research on that with low-quality input is limited, although it involves more applications. In this paper, we focus on 3D FR using low-quality data, targeting an efficient and accurate deep learning solution. To achieve this, we work on two aspects: (1) designing a lightweight yet powerful CNN; (2) generating finer and bigger training data. For (1), we propose a Multi-Scale Feature Fusion (MSFF) module and a Spatial Attention Vectorization (SAV) module to build a compact and discriminative CNN. For (2), we propose a data processing system including point-cloud recovery, surface refinement, and data augmentation (with newly proposed shape jittering and shape scaling). We conduct extensive experiments on Lock3DFace and achieve state-of-the-art results, outperforming many heavy CNNs such as VGG-16 and ResNet-34. In addition, our model can operate at a very high speed (136 fps) on Jetson TX2, and the promising accuracy and efficiency reached show its great applicability on edge/mobile devices.
Guodong Mu, Di Huang 0001, Guosheng Hu, Yunhong Wang 0001
CVPR3
2019 Ranked List Loss for Deep Metric Learning
abstract
The objective of deep metric learning (DML) is to learn embeddings that can capture semantic similarity information among data points. Existing pairwise or tripletwise loss functions used in DML are known to suffer from slow convergence due to a large proportion of trivial pairs or triplets as the model improves. To improve this, rankingmotivated structured losses are proposed recently to incorporate multiple examples and exploit the structured information among them. They converge faster and achieve state-of-the-art performance. In this work, we present two limitations of existing ranking-motivated structured losses and propose a novel ranked list loss to solve both of them. First, given a query, only a fraction of data points is incorporated to build the similarity structure. Consequently, some useful examples are ignored and the structure is less informative. To address this, we propose to build a setbased similarity structure by exploiting all instances in the gallery. The samples are split into a positive set and a negative set. Our objective is to make the query closer to the positive set than to the negative set by a margin. Second, previous methods aim to pull positive pairs as close as possible in the embedding space. As a result, the intraclass data distribution might be dropped. In contrast, we propose to learn a hypersphere for each class in order to preserve the similarity structure inside it. Our extensive experiments show that the proposed method achieves state-of-the-art performance on three widely used benchmarks.
Xinshao Wang, Yang Hua 0001, Elyor Kodirov, Guosheng Hu, Romain Garnier, Neil Robertson 0002
CVPR4
2019 Learning Discriminative and Complementary Patches for Face Recognition
abstract
The ensemble of convolutional neural networks (CNNs) has widely been used in many computer vision tasks including face recognition. Many existing ensembles of face recognition CNNs apply a two-stage pipeline to target performance improvement [10], [20], [22], [23], [29]: (1) it trains multiple CNNs separately with many face patches covering different facial areas; (2) the features derived from different models are aggregated off-line by different fusion methods. The well-known face recognition work, DeepID2 [20] trains 200 networks based on 200 arbitrarily chosen facial areas and chooses the best 25 ones to achieve impressive performance. However, it is very time-consuming to train so many networks. In addition, a brute-force like way of choosing facial patches is used without knowing which face patches are complementary and discriminative. It might be lack of generalization capability for cross-database applications. To solve that, we propose a novel end-to-end CNN ensemble architecture which automatically learns the complementary and discriminative patches for face recognition. Specifically, we propose a novel Patch Generation Engine (PGE) with Patch Search Spatial Transformer Network (PS-STN) and ROI shrunk loss to perform the patch selection process. ROI shrunk loss enlarges the distance of learned features in spatial space and feature space and learn complementary features. In order to get final aggregated feature, we use a supervised fusion module named Two Stage Discriminative Fusion Module (TSDFM) which effective to capture the global and local information and further guide the PGE to learn better patches. Extensive experiments conducted on LFW and YTF datasets show the effectiveness of our novel end-to-end ensemble method.
Zhiwei Liu 0004, Ming Tang 0001, Guosheng Hu, Jinqiao Wang
FG3
2019 Collaborative representation based face classification exploiting block weighted LBP and analysis dictionary learning
Xiaoning Song, Youming Chen, Zhenhua Feng 0001, Guosheng Hu, Tao Zhang 0010, Xiaojun Wu 0001
Pattern Recognit.4
2019 Fast SRC using quadratic optimisation in downsized coefficient solution subspace
Xiaoning Song, Guosheng Hu, Jian-Hao Luo, Zhenhua Feng 0001, Dongjun Yu, Xiaojun Wu 0001
Signal Process.2
2019 Face Frontalization Using an Appearance-Flow-Based Convolutional Neural Network
abstract
Facial pose variation is one of the major factors making face recognition (FR) a challenging task. One popular solution is to convert non-frontal faces to frontal ones on which FR is performed. Rotating faces causes facial pixel value changes. Therefore, existing CNN-based methods learn to synthesize frontal faces in color space. However, this learning problem in a color space is highly non-linear, causing the synthetic frontal faces to lose fine facial textures. In this paper, we take the view that the nonfrontal-frontal pixel changes are essentially caused by geometric transformations (rotation, translation, and so on) in space. Therefore, we aim to learn the nonfrontal-frontal facial conversion in the spatial domain rather than the color domain to ease the learning task. To this end, we propose an appearance-flow-based face frontalization convolutional neural network (A3F-CNN). Specifically, A3F-CNN learns to establish the dense correspondence between the non-frontal and frontal faces. Once the correspondence is built, frontal faces are synthesized by explicitly "moving" pixels from the non-frontal one. In this way, the synthetic frontal faces can preserve fine facial textures. To improve the convergence of training, an appearance-flow-guided learning strategy is proposed. In addition, generative adversarial network loss is applied to achieve a more photorealistic face, and a face mirroring method is introduced to handle the self-occlusion problem. Extensive experiments are conducted on face synthesis and pose invariant FR. Results show that our method can synthesize more photorealistic faces than the existing methods in both the controlled and uncontrolled lighting environments. Moreover, we achieve a very competitive FR performance on the Multi-PIE, LFW and IJB-A databases.
Zhihong Zhang 0001, Xu Chen 0020, Beizhan Wang, Guosheng Hu, Wangmeng Zuo, Edwin R. Hancock
IEEE Trans. Image Process.4
2018 Deep Multi-task Learning to Recognise Subtle Facial Expressions of Mental States
Guosheng Hu, Li Liu 0004, Yang Hua 0001, Zhihong Zhang 0001, Fumin Shen, Ling Shao 0001, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang
ECCV (12)1
2018 Deep Stock Representation Learning: From Candlestick Charts to Investment Decisions
abstract
We propose a novel investment decision strategy (IDS) based on deep learning. The performance of many IDSs is affected by stock similarity. Most existing stock similarity measurements have the problems: (a) The linear nature of many measurements cannot capture nonlinear stock dynamics; (b) The estimation of many similarity metrics (e.g. covariance) needs very long period historic data (e.g. 3K days) which cannot represent current market effectively; (c) They cannot capture translation-invariance. To solve these problems, we apply Convolutional AutoEncoder to learn a stock representation, based on which we propose a novel portfolio construction strategy by: (i) using the deeply learned representation and modularity optimisation to cluster stocks and identify diverse sectors, (ii) picking stocks within each cluster according to their Sharpe ratio (Sharpe 1994). Overall this strategy provides low-risk high-return portfolios. We use the Financial Times Stock Exchange 100 Index (FTSE 100) data for evaluation. Results show our portfolio outperforms FTSE 100 index and many well known funds in terms of total return in 2000 trading days.
Guosheng Hu, Kai Yang 0031, Flood Sung, Zhihong Zhang 0001, Neil Robertson 0002, Timothy M. Hospedales, Qiangwei Miemie
ICASSP1
2018 A Unified Neighbor Reconstruction Method for Embeddings
abstract
In this work we propose a novel and compact Neighbor Reconstruction Method (NRM) which is a unified pre-processing method for graph-based sparse spectral algorithms. This method is conducted by vector operations on a central point and its corresponding neighbor points. NRM generates new neighbor points which can capture the local space structure of the central point more appropriately than original neighbor points. With NRM, a large number of sparse spectral based nonlinear feature extraction and selection algorithms gain significant improvement. Specifically, we embedded NRM to several classical algorithms, Local Linear Embedding (LLE) [1], Laplacian Eigenmaps (LE) [2] and Unsupervised Feature Selection for Multi-cluster Data (MCFS) [3], with accuracy improvement of up to 7%, 2.6%, 2.4% on ORL, CIFAR 10, and MINST data sets respectively. We also apply NRM to a Super Resolution algorithm, A+ [5], and obtain 0.12dB improvement than original method.
Zhiling Ye, Zhihong Zhang 0001, Lu Bai 0001, Guosheng Hu, Zheng-Jian Bai, Yiqun Hu, Edwin R. Hancock
ICPR4
2018 Recovering variations in facial albedo from low resolution images
Xu Chen 0020, Zhihong Zhang 0001, Beizhan Wang, Guosheng Hu, Edwin R. Hancock
Pattern Recognit.4
2018 Dictionary Integration Using 3D Morphable Face Models for Pose-Invariant Collaborative-Representation-Based Classification
abstract
The paper presents a dictionary integration algorithm using 3D morphable face models (3DMM) for pose-invariant collaborative-representation-based face classification. To this end, we first fit a 3DMM to the 2D face images of a dictionary to reconstruct the 3D shape and texture of each image. The 3D faces are used to render a number of virtual 2D face images with arbitrary pose variations to augment the training data, by merging the original and rendered virtual samples to create an extended dictionary. Second, to reduce the information redundancy of the extended dictionary and improve the sparsity of reconstruction coefficient vectors using collaborative-representation-based classification (CRC), we exploit an on-line class elimination scheme to optimise the extended dictionary by identifying the training samples of the most representative classes for a given query. The final goal is to perform pose-invariant face classification using the proposed dictionary integration method and the on-line pruning strategy under the CRC framework. Experimental results obtained for a set of well-known face data sets demonstrate the merits of the proposed method, especially its robustness to pose variations.
Xiaoning Song, Zhenhua Feng 0001, Guosheng Hu, Josef Kittler, Xiaojun Wu 0001
IEEE Trans. Inf. Forensics Secur.3
2018 Frankenstein: Learning Deep Face Representations Using Small Data
abstract
Deep convolutional neural networks have recently proven extremely effective for difficult face recognition problems in uncontrolled settings. To train such networks, very large training sets are needed with millions of labeled images. For some applications, such as near-infrared (NIR) face recognition, such large training data sets are not publicly available and difficult to collect. In this paper, we propose a method to generate very large training data sets of synthetic images by compositing real face images in a given data set. We show that this method enables to learn models from as few as 10 000 training images, which perform on par with models trained from 500 000 images. Using our approach, we also obtain state-of-the-art results on the CASIA NIR-VIS2.0 heterogeneous face recognition data set.
Guosheng Hu, Xiaojiang Peng, Yongxin Yang, Timothy M. Hospedales, Jakob Verbeek
IEEE Trans. Image Process.1
2017 Attribute-Enhanced Face Recognition with Neural Tensor Fusion Networks
abstract
Deep learning has achieved great success in face recognition, however deep-learned features still have limited invariance to strong intra-personal variations such as large pose changes. It is observed that some facial attributes (e.g. eyebrow thickness, gender) are robust to such variations. We present the first work to systematically explore how the fusion of face recognition features (FRF) and facial attribute features (FAF) can enhance face recognition performance in various challenging scenarios. Despite the promise of FAF, we find that in practice existing fusion methods fail to leverage FAF to boost face recognition performance in some challenging scenarios. Thus, we develop a powerful tensor-based framework which formulates feature fusion as a tensor optimisation problem. It is nontrivial to directly optimise this tensor due to the large number of parameters to optimise. To solve this problem, we establish a theoretical equivalence between low-rank tensor optimisation and a two-stream gated neural network. This equivalence allows tractable learning using standard neural network optimisation tools, leading to accurate and stable optimisation. Experimental results show the fused feature works better than individual features, thus proving for the first time that facial attributes aid face recognition. We achieve state-of-the-art performance on three popular databases: MultiPIE (cross pose, lighting and expression), CASIA NIR-VIS2.0 (cross-modality environment) and LFW (uncontrolled environment).
Guosheng Hu, Yang Hua 0001, Zhihong Zhang 0001, Sankha S. Mukherjee, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang
ICCV1
2017 Efficient 3D morphable face model fitting
Guosheng Hu, Fei Yan 0001, Josef Kittler, William J. Christmas, Chi-Ho Chan, Zhenhua Feng 0001, Patrik Huber 0001
Pattern Recognit.1
2017 Half-Face Dictionary Integration for Representation-Based Classification
abstract
This paper presents a half-face dictionary integration (HFDI) algorithm for representation-based classification. The proposed HFDI algorithm measures residuals between an input signal and the reconstructed one, using both the original and the synthesized dual-column (row) half-face training samples. More specifically, we first generate a set of virtual half-face samples for the purpose of training data augmentation. The aim is to obtain high-fidelity collaborative representation of a test sample. In this half-face integrated dictionary, each original training vector is replaced by an integrated dual-column (row) half-face matrix. Second, to reduce the redundancy between the original dictionary and the extended half-face dictionary, we propose an elimination strategy to gain the most robust training atoms. The last contribution of the proposed HFDI method is the use of a competitive fusion method weighting the reconstruction residuals from different dictionaries for robust face classification. Experimental results obtained from the Facial Recognition Technology, Aleix and Robert, Georgia Tech, ORL, and Carnegie Mellon University-pose, illumination and expression data sets demonstrate the effectiveness of the proposed method, especially in the case of the small sample size problem.
Xiaoning Song, Zhenhua Feng 0001, Guosheng Hu, Xiaojun Wu 0001
IEEE Trans. Cybern.3
2016 Face Recognition Using a Unified 3D Morphable Model
Guosheng Hu, Fei Yan 0001, Chi-Ho Chan, Weihong Deng, William J. Christmas, Josef Kittler, Neil Robertson 0002
ECCV (8)1
2016 Face image super-resolution via weighted patches regression
abstract
Recently sparse representation has gained great success in face image super-resolution. The conventional sparsity-based methods enforce sparse coding on face image patches and the representation fidelity is measured by ℓ2-norm. Such a sparse coding model regularizes all facial patches equally, which however ignores the natures of facial patches, where the facial patches in the different regions (patch positions) of human face may have distinct contributions to face image reconstruction. In this paper, we propose to weight facial patches based on their discriminative abilities in regression for robust face hallucination reconstruction. Specifically, we learn the weights for facial patches according to the information entropy in each face region, so as to highlight higher frequency details in face images and the facial discriminability can be well retrieved. Furthermore, the weighted sparse coding can reasonable represent the less sparse nature of noisy images and thus remarkably boosts noise robust performance in face image super-resolution. Various experimental results on standard face databased show that our proposed method outperforms state-of-the-art methods in terms of both objective metrics and visual quality.
Zhihong Zhang 0001, Guosheng Hu, Edwin R. Hancock
ICPR3
2016 Comparison of analytical predictions of the noise floor due to static charge pump mismatch in fractional-n frequency synthesizers
abstract
Interaction between the requantizer's periodic output and charge pump mismatch nonlinearity in a fractional-N frequency synthesizer causes it to exhibit an elevated inband noise floor and spurs. In this paper, we consider three leading analytical predictions from the literature. By comparing simulation results with analytical predictions, we show that the two most common approaches fail to deal correctly with offsets, while the method based on Price's theorem works well.
Hongjia Mo, Guosheng Hu, Michael Peter Kennedy
ISCAS2
2015 The noise and spur delusion in fractional-N frequency synthesizer design
abstract
The standard design methodology for fractional-N frequency synthesizers assumes that the filtered shaped quantization noise from the requantizer is masked below the spectral envelope of the underlying integer-N synthesizer. Fractional-N frequency synthesizers are notorious for exhibiting an elevated noise floor and an unpredictable pattern of spurs. In this paper, we argue that designers should not be deluded by the overly conservative predictions of the simplified linear model but should instead consider nonlinearities as early as possible in the design process.
Michael Peter Kennedy, Hongjia Mo, Zhida Li, Guosheng Hu, Paolo Scognamiglio, Ettore Napoli
ISCAS4
2015 Cascaded Collaborative Regression for Robust Facial Landmark Detection Trained Using a Mixture of Synthetic and Real Images With Dynamic Weighting
abstract
A large amount of training data is usually crucial for successful supervised learning. However, the task of providing training samples is often time-consuming, involving a considerable amount of tedious manual work. In addition, the amount of training data available is often limited. As an alternative, in this paper, we discuss how best to augment the available data for the application of automatic facial landmark detection. We propose the use of a 3D morphable face model to generate synthesized faces for a regression-based detector training. Benefiting from the large synthetic training data, the learned detector is shown to exhibit a better capability to detect the landmarks of a face with pose variations. Furthermore, the synthesized training data set provides accurate and consistent landmarks automatically as compared to the landmarks annotated manually, especially for occluded facial parts. The synthetic data and real data are from different domains; hence the detector trained using only synthesized faces does not generalize well to real faces. To deal with this problem, we propose a cascaded collaborative regression algorithm, which generates a cascaded shape updater that has the ability to overcome the difficulties caused by pose variations, as well as achieving better accuracy when applied to real faces. The training is based on a mix of synthetic and real image data with the mixing controlled by a dynamic mixture weighting schedule. Initially, the training uses heavily the synthetic data, as this can model the gross variations between the various poses. As the training proceeds, progressively more of the natural images are incorporated, as these can model finer detail. To improve the performance of the proposed algorithm further, we designed a dynamic multi-scale local feature extraction method, which captures more informative local features for detector training. An extensive evaluation on both controlled and uncontrolled face data sets demonstrates the merit of the proposed algorithm.
Zhenhua Feng 0001, Guosheng Hu, Josef Kittler, William J. Christmas, Xiaojun Wu 0001
IEEE Trans. Image Process.2
2014 Robust face recognition by an albedo based 3D morphable model
abstract
Large pose and illumination variations are very challenging for face recognition. The 3D Morphable Model (3DMM) approach is one of the effective methods for pose and illumination invariant face recognition. However, it is very difficult for the 3DMM to recover the illumination of the 2D input image because the ratio of the albedo and illumination contributions in a pixel intensity is ambiguous. Unlike the traditional idea of separating the albedo and illumination contributions using a 3DMM, we propose a novel Albedo Based 3D Morphable Model (AB3DMM), which removes the illumination component from the images using illumination normalisation in a preprocessing step. A comparative study of different illumination normalisation methods for this step is conducted on PIE and Multi-PIE databases. The results show that overall performance of our method outperforms state-of-the-art methods.
Guosheng Hu, Chi-Ho Chan, Fei Yan 0001, William J. Christmas, Josef Kittler
IJCB1
2014 Robust compressed sensing with bounded and structured uncertainties
abstract
The robust compressed sensing problem subject to a bounded and structured perturbation in the sensing matrix is solved in two steps. The alternating direction method of multipliers (ADMM) is first applied to obtain a robust support set. Unlike the existing robust signal recovery solutions, the proposed optimisation problem is convex. The ADMM algorithm that every subproblem has a global minimum is employed to solve the optimisation problem. Then, the standard robust regularised least‐squares problem restrained to the support is solved to reduce the recovery error. The numerical tests show that the proposed approach provides a robust estimation of support set, although it is conservative to recover signal magnitudes as a result of minimising the worst‐cast data error across all bounded perturbations.
Xiangyun Qing, Guosheng Hu
IET Signal Process.2
2012 Resolution-Aware 3D Morphable Model
abstract
The 3D Morphable Model (3DMM) is currently receiving considerable attention for \nhuman face analysis. Most existing work focuses on fitting a 3DMM to high resolution \nimages. However, in many applications, fitting a 3DMM to low-resolution images \nis also important. In this paper, we propose a Resolution-Aware 3DMM (RA- \n3DMM), which consists of 3 different resolution 3DMMs: High-Resolution 3DMM \n(HR- 3DMM), Medium-Resolution 3DMM (MR-3DMM) and Low-Resolution 3DMM \n(LR-3DMM). RA-3DMM can automatically select the best model to fit the input images \nof different resolutions. The multi-resolution model was evaluated in experiments \nconducted on PIE and XM2VTS databases. The experimental results verified that HR- \n3DMM achieves the best performance for input image of high resolution, and MR- \n3DMM and LR-3DMM worked best for medium and low resolution input images, respectively. \nA model selection strategy incorporated in the RA-3DMM is proposed based \non these results. The RA-3DMM model has been applied to pose correction of face images \nranging from high to low resolution. The face verification results obtained with \nthe pose-corrected images show considerable performance improvement over the result \nwithout pose correction in all resolutions
Guosheng Hu, Chi-Ho Chan, Josef Kittler, William J. Christmas
BMVC1
2010 Genetic Algorithms with Improved Simulated Binary Crossover and Support Vector Regression for Grid Resources Prediction
Guosheng Hu, Liang Hu 0001, Qinghai Bai, Guangyu Zhao
ISNN (2)1
2010 Support Vector Regression and Ant Colony Optimization for Grid Resources Prediction
Guosheng Hu, Liang Hu 0001, Pengchao Li, Xilong Che
ISNN (2)1