EDBT 2026 Demo / reviewers in the wild / expert
Guosheng Hu
dblp:98/7676
· DBLP profile ↗
74ranked-venue papers
10as first author
37since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 8 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 45 · 7 first-author · 20 since 2021Security and privacy · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Not All Directions Matter: Towards Structured and Task-Aware Low-Rank Model AdaptationabstractXi Xiao, Chenrui Ma, Yunbei Zhang, Chen Liu, Zhuxuanzi Wang, Yanshu Li, Lin Zhao, Guosheng Hu, Tianyang Wang, Hao Xu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xi Xiao 0003, Chenrui Ma, Yunbei Zhang, Chen Liu 0020, Zhuxuanzi Wang, Yanshu Li, Guosheng Hu, Tianyang Wang 0004 |
ACL (1) | 8 |
| 2026 | Uncertainty-Guided Cross-Modal Distillation for Category-Level Object Pose EstimationabstractRecent years have seen significant advancements in category-level object pose estimation, largely driven by multimodal (RGB-D) approaches. Despite their success, depth-only methods remain widely adopted in practical applications due to their superior computational efficiency and ease of deployment. However, these methods typically suffer from a noticeable performance gap compared to multimodal methods. To bridge this gap, we propose a novel framework, Cross-Modal Uncertainty Distillation for Pose Estimation (CMUD-Pose), which transfers discriminative knowledge from an RGB-D teacher to a depth-only student network. Furthermore, to mitigate overfitting induced by the modality gap, we propose Cross-Modal Uncertainty Distillation (CMUD), which utilizes a learned uncertainty-aware weighting mechanism for adaptively assigning importance to training samples. By incorporating uncertainty into the distillation process, CMUD allows the student model to focus selectively on reliable and transferable cross-modal features. Extensive experiments on the REAL275 and CAMERA25 benchmarks show that our method significantly improves the performance of depth-only pose estimation models. Tianfu Wang 0003, Guosheng Hu, Hongguang Wang |
IEEE Signal Process. Lett. | 2 |
| 2025 | Beyond the Answer: Advancing Multi-Hop QA with Fine-Grained Graph Reasoning and EvaluationabstractRecent advancements in large language models (LLMs) have significantly improved the performance of multi-hop question answering (MHQA) systems. Despite the success of MHQA systems, the evaluation of MHQA is not deeply investigated. Existing evaluations mainly focus on comparing the final answers of the reasoning method and given ground-truths. We argue that the reasoning process should also be evaluated because wrong reasoning process can also lead to the correct final answers. Motivated by this, we propose a “Planner-Executor-Reasoner” (PER) architecture, which forms the core of the Plan-anchored Data Preprocessing (PER-DP) and the Plan-guided Multi-Hop QA (PER-QA).The former provides the ground-truth of intermediate reasoning steps and final answers, and the latter offers them of a reasoning method. Moreover, we design a fine-grained evaluation metric called Plan-aligned Stepwise Evaluation (PSE), which evaluates the intermediate reasoning steps from two aspects: planning and solving. Extensive experiments on ten types of questions demonstrate competitive reasoning performance, improved explainability of the MHQA system, and uncover issues such as “fortuitous reasoning continuance” and “latent reasoning suspension” in RAG-based MHQA systems. Besides, we also demonstrate the potential of our approach in data contamination scenarios. Qichuan Liu, Chentao Zhang, Chenfeng Zheng, Guosheng Hu |
ACL (1) | 4 |
| 2025 | DPL++: Advancing the Network Performance via Image and Label PerturbationsabstractRecent advances in supervised learning have predominantly focused on regularizations, optimizers, and architectures, yet the potential of simultaneously optimizing data distributions and supervisory signals for training samples remains underexplored. In this paper, we propose a novel paradigm that leverages the benefits of image perturbations for rectifying data distributions. Our method, called DPL (Deep Perturbation Learning), introduces new insights into utilizing image perturbations and focuses on improving generalizability on normal samples, rather than resisting adversarial attacks. DPL formulates a differentiable function w.r.t. image perturbations and implements an alternative optimization process that seamlessly integrates with downstream tasks. However, the limitations of DPL stem from the inefficiency in employing differentiable targets caused by the exclusive optimization of image perturbations, while neglecting the critical role of supervisory signals in training effectiveness. These lead to the excessive necessity of DPL iterations and yield inferior performance-cost trade-off. To track this, we extend DPL to DPL++ with synchronous optimization for image perturbations and label perturbations. In our DPL++ paradigm, the post-hoc application of perturbations to images and labels endows amendments toward both data distributions and supervisory signals, significantly furthering the generalizability of models over various benchmarks. Crucially, the proposed synchronous optimization process shares key differentiable objectives to reduce computational complexity, thereby achieving enhanced effectiveness within fewer optimization iterations. Theoretically, as a generic and flexible approach, DPL++ can be applied to a variety of backbone architectures (e.g., ResNet, DenseNet, and ViT) and downstream tasks (e.g., image classification and object detection). To validate the efficacy of DPL++, we conduct extensive performance experiments and in-depth analytical studies on 2 visual tasks over 5 mainstream benchmarks across 13 backbone networks. The comprehensive results verify the superiority of DPL++ over DPL and demonstrate its promising capabilities for advancing decision-making capacity, risk minimization, class distinguishability, and training convergence. Zifan Song, Guosheng Hu, Shuguang Dou, Cairong Zhao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Explainability-based knowledge distillation
Tianli Sun, Haonan Chen 0003, Guosheng Hu, Cairong Zhao |
Pattern Recognit. | 3 |
| 2025 | Multi-definition Deepfake detection via semantics reduction and cross-domain training
Cairong Zhao, Chutian Wang, Zifan Song, Guosheng Hu, Duoqian Miao 0001 |
Pattern Recognit. | 4 |
| 2025 | Learning Label PerturbationsabstractSupervised learning typically uses hard labels for annotations, which may not fully capture the underlying distribution of the data. In the literature, label smoothing is a method that can reduce overconfidence and enhance the model generalization by using a weighted average of one-hot vectors and the uniform distribution, but it does not extract the intrinsic information from the data. Another approach is knowledge distillation, which uses the predicted probability distribution of a teacher network trained with hard labels as soft labels for a student network. However, this method lacks a theoretical explanation. In this work, we draw inspiration from the influence function and propose a post-hoc label perturbation learning method called Deep Soft Label Learning (DSLL). This method iteratively leverages the inherent information present in both the model and data to theoretically determine optimal labels for classification and regression problems. Our experiments demonstrate that DSLL consistently enhances model performance across various tasks, including image classification and object detection. Zifan Song, Guosheng Hu, Cairong Zhao |
IEEE Signal Process. Lett. | 3 |
| 2025 | AdvMixUp: Adversarial MixUp Regularization for Deep LearningabstractDeep neural networks (DNNs) have shown significant progress in many application fields. However, overfitting remains a significant challenge in their development. While existing data-augmentation techniques such as MixUp have been successful in preventing overfitting, they often fail to generate hard mixed samples near the decision boundary, impeding model optimization. In this article, we present adversarial MixUp (AdvMixUp), a novel sample-dependent method for regularizing DNNs. AdvMixUp addresses this issue by incorporating adversarial training (AT) to create sample-dependent and feature-level interpolation masks, generating more challenging mixed samples. These virtual samples enable DNNs to learn more robust features, ultimately reducing overfitting. Empirical evaluations on CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet demonstrate that AdvMixUp outperforms existing MixUp variants. Jun Fu 0001, Xianrui Ji, Dexiong Chen, Guosheng Hu, Shuang Li 0008, Xiating Feng |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Diverse Person: Customize Your Own Dataset for Text-Based Person SearchabstractText-based person search is a challenging task aimed at locating specific target pedestrians through text descriptions. Recent advancements have been made in this field, but there remains a deficiency in datasets tailored for text-based person search. The creation of new, real-world datasets is hindered by concerns such as the risk of pedestrian privacy leakage and the substantial costs of annotation. In this paper, we introduce a framework, named Diverse Person (DP), to achieve efficient and high-quality text-based person search data generation without involving privacy concerns. Specifically, we propose to leverage available images of clothing and accessories as reference attribute images to edit the original dataset images through diffusion models. Additionally, we employ a Large Language Model (LLM) to produce annotations that are both high in quality and stylistically consistent with those found in real-world datasets. Extensive experimental results demonstrate that the baseline models trained with our DP can achieve new state-of-the-art results on three public datasets, with performance improvements up to 4.82%, 2.15%, and 2.28% on CUHK-PEDES, ICFG-PEDES, and RSTPReid in terms of Rank-1 accuracy, respectively. Zifan Song, Guosheng Hu, Cairong Zhao |
AAAI | 2 |
| 2024 | Gradient-Guided Modality Decoupling for Missing-Modality RobustnessabstractMultimodal learning with incomplete input data (missing modality) is very practical and challenging. In this work, we conduct an in-depth analysis of this challenge and find that modality dominance has a significant negative impact on the model training, greatly degrading the missing modality performance. Motivated by Grad-CAM, we introduce a novel indicator, gradients, to monitor and reduce modality dominance which widely exists in the missing-modality scenario. In aid of this indicator, we present a novel Gradient-guided Modality Decoupling (GMD) method to decouple the dependency on dominating modalities. Specifically, GMD removes the conflicted gradient components from different modalities to achieve this decoupling, significantly improving the performance. In addition, to flexibly handle modal-incomplete data, we design a parameter-efficient Dynamic Sharing (DS) framework which can adaptively switch on/off the network parameters based on whether one modality is available. We conduct extensive experiments on three popular multimodal benchmarks, including BraTS 2018 for medical segmentation, CMU-MOSI, and CMU-MOSEI for sentiment analysis. The results show that our method can significantly outperform the competitors, showing the effectiveness of the proposed solutions. Our code is released here: https://github.com/HaoWang420/Gradient-guided-Modality-Decoupling. Shengda Luo, Guosheng Hu |
AAAI | 3 |
| 2024 | Neighborhood-Enhanced 3D Human Pose Estimation with Monocular LiDAR in Long-Range Outdoor Scenesabstract3D human pose estimation (3HPE) in large-scale outdoor scenes using commercial LiDAR has attracted significant attention due to its potential for real-life applications. However, existing LiDAR-based methods for 3HPE primarily rely on recovering 3D human poses from individual point clouds, and the coherence cues present in the neighborhood are not sufficiently harnessed. In this work, we explore spatial and contexture coherence cues contained in the neighborhood that lead to great performance improvements in 3HPE. Specifically, firstly, we deeply investigate the 3D neighbor in the background (3BN) which serves as a spatial coherence cue for inferring reliable motion since it provides physical laws to limit motion targets. Secondly, we introduce a novel 3D scanning neighbor (3SN) generated during the data collection and 3SN implies structural edge coherence cues. We use 3SN to overcome the degradation of performance and data quality caused by the sparsity-varying properties of LiDAR point clouds. In order to effectively model the complementation between these distinct cues and build consistent temporal relationships across human motions, we propose a new transformer-based module called the CoherenceFuse module. Extensive experiments were conducted on publicly available datasets, namely LidarHuman26M, CIMI4D, SLOPER4D and Waymo Open Dataset v2.0, showcase the superiority and effectiveness of our proposed method. In particular, when compared with LidarCap on the LidarHuman26M dataset, our method demonstrates a reduction of 7.08mm in the average MPJPE metric, along with a decrease of 16.55mm in the MPJPE metric for distances exceeding 25 meters. The code and models are available at https://github.com/jingyi-zhang/Neighborhood-enhanced-LidarCap. Qihong Mao, Guosheng Hu, Cheng Wang 0003 |
AAAI | 3 |
| 2024 | Object Pose Estimation via the Aggregation of Diffusion FeaturesabstractEstimating the pose of objects from images is a crucial task of 3D scene understanding, and recent approaches have shown promising results on very large benchmarks. However, these methods experience a significant performance drop when dealing with unseen objects. We believe that it results from the limited generalizability of image features. To address this problem, we have an in-depth analysis on the features of diffusion models, e.g. Stable Diffusion, which hold substantial potential for modeling unseen objects. Based on this analysis, we then innovatively introduce these diffusion features for object pose estimation. To achieve this, we propose three distinct architectures that can effectively capture and aggregate diffusion features of different granularity, greatly improving the generalizability of object pose estimation. Our approach outperforms the state-of-the-art methods by a considerable margin on three popular benchmark datasets, LM, O-LM, and T-LESS. In particular, our method achieves higher accuracy than the previous best arts on unseen objects: 98.2% vs. 93.5% on Unseen LM, 85.9% vs. 76.3% on Unseen O-LM, showing the strong generalizability of our method. Our code is released at https://github.com/Tianfu18/diff-feats-pose. Tianfu Wang 0003, Guosheng Hu, Hongguang Wang |
CVPR | 2 |
| 2024 | DiffLoc: Diffusion Model for Outdoor LiDAR LocalizationabstractAbsolute pose regression (APR) estimates global pose in an end-to-end manner, achieving impressive results in learn-based LiDAR localization. However, compared to the top-performing methods reliant on 3D-3D correspondence matching, APR's accuracy still has room for improvement. We recognize APR's lack of robust features learning and iterative denoising process leads to suboptimal results. In this paper, we propose DiffLoc, a novel framework that formulates LiDAR localization as a conditional generation of poses. First, we propose to utilize the foundation model and static-object-aware pool to learn robust features. Second, we incorporate the iterative denoising process into APR via a diffusion model conditioned on the learned geometrically robust features. In addition, due to the unique nature of diffusion models, we propose to adapt our models to two additional applications: (1) using multiple inferences to evaluate pose uncertainty, and (2) seamlessly introducing geometric constraints on denoising steps to improve prediction accuracy. Extensive experiments conducted on the Oxford Radar RobotCar and NCLT datasets demonstrate that DiffLoc outperforms better than the state-of-the-art methods. Especially on the NCLT dataset, we achieve 35% and 34.7% improvement on position and orientation accuracy, respectively. Our code is released at https://github.com/liw95/DiffLoc. Wen Li 0005, Yuyang Yang, Shangshu Yu, Guosheng Hu, Chenglu Wen, Ming Cheng 0002, Cheng Wang 0003 |
CVPR | 4 |
| 2024 | Unsupervised Exposure Correction
Ruodai Cui, Li Niu 0002, Guosheng Hu |
ECCV (6) | 3 |
| 2024 | Learning Scene-Pedestrian Graph for End-to-End Person SearchabstractPerson search aims to find specific persons from visual scenes, including two subtasks, pedestrian detection, and person reidentification. The dominant fashion in this area is end-to-end networks that focus on analyzing the foreground (i.e., pedestrian) while ignoring the background (i.e., scene) information. However, the scene information often offers useful clues for person search. For example, pedestrians normally appear on the road rather than the top of a tree, and pedestrians appearing at the same location are likely to have similar occlusions. The interplay between the pedestrians and scenes can potentially improve the performance. In this article, a novel scene-pedestrian graph (SPG) is proposed, which can explicitly model the interplay between the pedestrians and scenes. To polish the quality of pedestrian bounding boxes, we pioneer a strategy of using the high-quality pedestrian bounding box to guide the low-quality one in the same scene. In addition, we design a contextual and temporal graph matching algorithm to effectively utilize the contextual and temporal information present in the constructed SPG to improve the performance of pedestrian matching. Benefiting from the robustness on complex scenes, our model achieves promising performance over the state-of-the-art methods on two popular person search benchmarks, CUHK-SYSU and PRW. Zifan Song, Cairong Zhao, Guosheng Hu, Duoqian Miao 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | NIDALoc: Neurobiologically Inspired Deep LiDAR LocalizationabstractAbsolute pose regression has shown great potential in LiDAR localization, which learns to regress 6-DoF LiDAR poses through deep networks. However, recent regression methods suffer from scene ambiguities in challenging scenarios, leading to inaccurate and unstable localization. Inspired by neurobiological localization mechanisms, i.e., the firing mechanism of place cells, head-direction cells, and grid cells in mammalian brains, we propose a novel LiDAR localization framework called NIDALoc to achieve more robust and accurate results. First, we propose a Hebbian memory module, motivated by place cells, to preserve historical information, which helps refine local view features to reduce scene ambiguities. Specifically, the memory module stores scene information and then recalls it when revisiting an old place. Second, we propose a novel pose constrained framework, consisting of an orientation classification task and a grid center regression task, to regularize orientation and position estimation, respectively. The framework based on head-direction cells and grid cells constrains the absolute pose regression to reduce wrong predictions. Extensive experiments on two outdoor datasets demonstrate the effectiveness of NIDALoc, which outperforms state-of-the-art localization methods, especially in large-scale challenging scenes. The source code is available on the project website at https://github.com/PSYZ1234/NIDALoc. Shangshu Yu, Xiaotian Sun 0005, Wen Li 0005, Chenglu Wen, Yunuo Yang, Bailu Si, Guosheng Hu, Cheng Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | Explainability of Speech Recognition Transformers via Gradient-Based Attention VisualizationabstractIn vision Transformers, attention visualization methods are used to generate heatmaps highlighting the class-corresponding areas in input images, which offers explanations on how the models make predictions. However, it is not so applicable for explaining automatic speech recognition (ASR) Transformers. An ASR Transformer makes a particular prediction for every input token to form a sentence, but a vision Transformer only makes an overall classification for the input data. Therefore, traditional attention visualization methods may fail in ASR Transformers. In this work, we propose a novel attention visualization method in ASR Transformers and try to explain which frames of the audio result in the output text. Inspired by the model explainability, we also explore ways of improving the effectiveness of the ASR model. Comparing with other Transformer attention visualization methods, our method is more efficient and intuitively understandable, which unravels the attention calculation from information flow of Transformer attention modules. In addition, we demonstrate the utilization of visualization result in three ways: (1) We visualize attention with respect to connectionist temporal classification (CTC) loss to train an ASR model with adversarial attention erasing regularization, which effectively decreases the word error rate (WER) of the model and improves its generalization capability. (2) We visualize the attention on some specific words, interpreting the model by effectively demonstrating the semantic and grammar relationships between these words. (3) Similarly, we analyze how the model manage to distinguish homophones, using contrastive explanation with respect to homophones. Tianli Sun, Haonan Chen 0003, Guosheng Hu, Lianghua He, Cairong Zhao |
IEEE Trans. Multim. | 3 |
| 2024 | Deep Metric Learning Based on Meta-Mining Strategy With Semiglobal InformationabstractRecently, deep metric learning (DML) has achieved great success. Some existing DML methods propose adaptive sample mining strategies, which learn to weight the samples, leading to interesting performance. However, these methods suffer from a small memory (e.g., one training batch), limiting their efficacy. In this work, we introduce a data-driven method, meta-mining strategy with semiglobal information (MMSI), to apply meta-learning to learn to weight samples during the whole training, leading to an adaptive mining strategy. To introduce richer information than one training batch only, we elaborately take advantage of the validation set of meta-learning by implicitly adding additional validation sample information to training. Furthermore, motivated by the latest self-supervised learning, we introduce a dictionary (memory) that maintains very large and diverse information. Together with the validation set, this dictionary presents much richer information to the training, leading to promising performance. In addition, we propose a new theoretical framework that can formulate pairwise and tripletwise metric learning loss functions in a unified framework. This framework brings new insights to society and facilitates us to generalize our MMSI to many existing DML methods. We conduct extensive experiments on three public datasets, CUB200-2011, Cars-196, and Stanford Online Products (SOP). Results show that our method can achieve the state of the art or very competitive performance. Our source codes have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/MMSI. Xi Jiang 0001, Sheng Liu 0009, Xili Dai, Guosheng Hu, Xingguo Huang, Yazhou Yao, Guosen Xie, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Cross-Modal Distillation for Speaker RecognitionabstractSpeaker recognition achieved great progress recently, however, it is not easy or efficient to further improve its performance via traditional solutions: collecting more data and designing new neural networks. Aiming at the fundamental challenge of speech data, i.e. low information density, multimodal learning can mitigate this challenge by introducing richer and more discriminative information as input for identity recognition. Specifically, since the face image is more discriminative than the speech for identity recognition, we conduct multimodal learning by introducing a face recognition model (teacher) to transfer discriminative knowledge to a speaker recognition model (student) during training. However, this knowledge transfer via distillation is not trivial because the big domain gap between face and speech can easily lead to overfitting. In this work, we introduce a multimodal learning framework, VGSR (Vision-Guided Speaker Recognition). Specifically, we propose a MKD (Margin-based Knowledge Distillation) strategy for cross-modality distillation by introducing a loose constrain to align the teacher and student, greatly reducing overfitting. Our MKD strategy can easily adapt to various existing knowledge distillation methods. In addition, we propose a QAW (Quality-based Adaptive Weights) module to weight input samples via quantified data quality, leading to a robust model training. Experimental results on the VoxCeleb1 and CN-Celeb datasets show our proposed strategies can effectively improve the accuracy of speaker recognition by a margin of 10% ∼ 15%, and our methods are very robust to different noises. Yufeng Jin, Guosheng Hu, Haonan Chen 0003, Duoqian Miao 0001, Liang Hu 0001, Cairong Zhao |
AAAI | 2 |
| 2023 | SGLoc: Scene Geometry Encoding for Outdoor LiDAR LocalizationabstractLiDAR-based absolute pose regression estimates the global pose through a deep network in an end-to-end manner, achieving impressive results in learning-based localization. However, the accuracy of existing methods still has room to improve due to the difficulty of effectively encoding the scene geometry and the unsatisfactory quality of the data. In this work, we propose a novel LiDAR localization frame-work, SGLoc, which decouples the pose estimation to point cloud correspondence regression and pose estimation via this correspondence. This decoupling effectively encodes the scene geometry because the decoupled correspondence regression step greatly preserves the scene geometry, leading to significant performance improvement. Apart from this decoupling, we also design a tri-scale spatial feature aggregation module and inter-geometric consistency constraint loss to effectively capture scene geometry. Moreover, we empirically find that the ground truth might be noisy due to GPS/INS measuring errors, greatly reducing the pose estimation performance. Thus, we propose a pose quality evaluation and enhancement method to measure and correct the ground truth pose. Extensive experiments on the Oxford Radar RobotCar and NCLT datasets demonstrate the effectiveness of SGLoc, which outperforms state-of-the-art regression-based localization methods by 68.5% and 67.6% on position accuracy, respectively. Wen Li 0005, Shangshu Yu, Cheng Wang 0003, Guosheng Hu, Chenglu Wen |
CVPR | 4 |
| 2023 | Deep Perturbation Learning: Enhancing the Network Performance via Image PerturbationsabstractImage perturbation technique is widely used to generate adversarial examples to attack networks, greatly decreasing the performance of networks. Unlike the existing works, in this paper, we introduce a novel framework Deep Perturbation Learning (DPL), the new insights into understanding image perturbations, to enhance the performance of networks rather than decrease the performance. Specifically, we learn image perturbations to amend the data distribution of training set to improve the performance of networks. This optimization w.r.t data distribution is non-trivial. To approach this, we tactfully construct a differentiable optimization target w.r.t. image perturbations via minimizing the empirical risk. Then we propose an alternating optimization of the network weights and perturbations. DPL can easily be adapted to a wide spectrum of downstream tasks and backbone networks. Extensive experiments demonstrate the effectiveness of our DPL on 6 datasets (CIFAR-10, CIFAR100, ImageNet, MS-COCO, PASCAL VOC, and SBD) over 3 popular vision tasks (image classification, object detection, and semantic segmentation) with different backbone architectures (e.g., ResNet, MobileNet, and ViT). Zifan Song, Guosheng Hu, Cairong Zhao |
ICML | 3 |
| 2023 | ISTVT: Interpretable Spatial-Temporal Video Transformer for Deepfake DetectionabstractWith the rapid development of Deepfake synthesis technology, our information security and personal privacy have been severely threatened in recent years. To achieve a robust Deepfake detection, researchers attempt to exploit the joint spatial-temporal information in the videos, like using recurrent networks and 3D convolutional networks. However, these spatial-temporal models remain room to improve. Another general challenge for spatial-temporal models is that people do not clearly understand what these spatial-temporal models really learn. To address these two challenges, in this paper, we propose an Interpretable Spatial-Temporal Video Transformer (ISTVT), which consists of a novel decomposed spatial-temporal self-attention and a self-subtract mechanism to capture spatial artifacts and temporal inconsistency for robust Deepfake detection. Thanks to this decomposition, we propose to interpret ISTVT by visualizing the discriminative regions for both spatial and temporal dimensions via the relevance (the pixel-wise importance on the input) propagation algorithm. We conduct extensive experiments on large-scale datasets, including FaceForensics++, FaceShifter, DeeperForensics, Celeb-DF, and DFDC datasets. Our strong performance of intra-dataset and cross-dataset Deepfake detection demonstrates the effectiveness and robustness of our method, and our visualization-based interpretability offers people insights into our model. Cairong Zhao, Chutian Wang, Guosheng Hu, Haonan Chen 0003, Chun Liu 0003, Jinhui Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | STCLoc: Deep LiDAR Localization With Spatio-Temporal ConstraintsabstractLiDAR localization is of great importance to autonomous vehicles and robotics. Absolute pose regression, directly estimating the mapping from a scene to a 6-DoF pose, has achieved impressive results in learning-based localization. Different from traditional map-based methods, it does not need a pre-built 3D map during inference. However, current regression networks typically suffer from scene ambiguities, especially in challenging traffic environments, leading to large wrong predictions (e.g., outliers) and limited applications. To address this problem, a novel LiDAR localization framework with spatio-temporal constraints is proposed, termed STCLoc, to reduce scene ambiguities and achieve more accurate localization. First, we propose to regularize regression in the spatial dimension with a novel classification task to reduce outliers. Specifically, the classification task categorizes the point cloud in terms of position and orientation and then couples it with the regression task to conduct multi-task learning. Second, to learn discriminative features to reduce scene ambiguities, we propose using attention-based feature aggregation to capture the correlation in LiDAR sequences. We conduct extensive experiments on two benchmark datasets, where the localization takes 97ms on each dataset. Results show that our model outperforms state-of-the-art methods by 43.33%/36.76% (position/orientation) on the Oxford Radar RobotCar dataset, verifying the effectiveness of our method. The source code is available on the project website athttps://github.com/PSYZ1234/STCLoc. Shangshu Yu, Cheng Wang 0003, Yitai Lin, Chenglu Wen, Ming Cheng 0002, Guosheng Hu |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Discrepancy-Guided Domain-Adaptive Data AugmentationabstractData augmentation has been observed playing a crucial role in achieving better generalization in many machine learning tasks, especially in unsupervised domain adaptation (DA). It is particularly effective on visual object recognition tasks as images are high-dimensional with an enormous range of variations that can be simulated. Existing data augmentation techniques, however, are not explicitly designed to address the differences between different domains. Expert knowledge about the data is required, as well as manual efforts in finding the optimal parameters. In this article, we propose a novel domain-adaptive augmentation method by making use of a state-of-the-art style transfer method and domain discrepancy measurement. Specifically, we measure the discrepancy between source and target domains, and use it as a guide to augment the original source samples using style transferred source-to-target samples. The proposed domain-adaptive augmentation method is data and model agnostic that can be easily incorporated with state-of-the-art DA algorithms. We show empirically that, by using this domain-adaptive augmentation, we are able to gradually reduce the discrepancy between the source and target samples, and further boost the adaptation performance using different DA algorithms on three popular domain adaption datasets. Jian Gao 0018, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Boosting Active Learning via Improving Test PerformanceabstractCentral to active learning (AL) is what data should be selected for annotation. Existing works attempt to select highly uncertain or informative data for annotation. Nevertheless, it remains unclear how selected data impacts the test performance of the task model used in AL. In this work, we explore such an impact by theoretically proving that selecting unlabeled data of higher gradient norm leads to a lower upper-bound of test loss, resulting in a better test performance. However, due to the lack of label information, directly computing gradient norm for unlabeled data is infeasible. To address this challenge, we propose two schemes, namely expected-gradnorm and entropy-gradnorm. The former computes the gradient norm by constructing an expected empirical loss while the latter constructs an unsupervised loss with entropy. Furthermore, we integrate the two schemes in a universal AL framework. We evaluate our method on classical image classification and semantic segmentation tasks. To demonstrate its competency in domain applications and its robustness to noise, we also validate our method on a cellular imaging analysis task, namely cryo-Electron Tomography subtomogram classification. Results demonstrate that our method achieves superior performance against the state of the art. We refer readers to https://arxiv.org/pdf/2112.05683.pdf for the full version of this paper which includes the appendix and source code link. Tianyang Wang 0004, Xingjian Li 0002, Pengkun Yang, Guosheng Hu, Siyu Huang, Cheng-Zhong Xu 0001, Min Xu 0009 |
AAAI | 4 |
| 2022 | Efficient One-Stage Video Object Detection by Exploiting Temporal Consistency
Guanxiong Sun, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002 |
ECCV (35) | 3 |
| 2022 | TDViT: Temporal Dilated Video Transformer for Dense Video Tasks
Guanxiong Sun, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002 |
ECCV (35) | 3 |
| 2022 | RCANet: Row-Column Attention Network for Semantic SegmentationabstractEstablishing high-order interactions among pixels and object parts is one of the most fundamental problems in semantic segmentation. The recent proposals are based on non-local methods which utilize the self-attention mechanism to capture the long-range correlations. However, non-local methods could be very expensive, both theoretically and experimentally. Moreover, non-local methods are typically designed to address spatial correlations rather than feature correlations across channels. In this work, we propose a Row-Column Attention Network (RCANet) to encode globally contextual information. It consists of a row-wise intra-channel attention module and a column-wise intra-channel attention module, followed by a cross-channel interaction module. We conduct experiments on two datasets: Cityscapes and ADE20K. The results show that our method is comparable to the state-of-the-art methods for semantic segmentation. Bingxu Lu, Qinghua Hu, Yu Wang 0106, Guosheng Hu |
ICASSP | 4 |
| 2022 | Multi-Definition Video Deepfake Detection via Semantics Reduction and Cross-Domain TrainingabstractThe recent development of Deepfake videos directly threatens our information security and personal privacy. Although lots of previous works have made much progress on the Deepfake detection, we empirically find that the existing approaches do not perform well on the low definition (LD) and crossdefinition (high and low) videos. To address this problem, in this paper, we follow two motivations: (1) high-level semantics reduction and (2) cross-domain training. For (1), we propose the Facial Structure Destruction and Adversarial Jigsaw Loss to reduce our model to learn high-level semantics and focus on learning low-level discriminative information; For (2), we propose a domain generalization method based on adversarial learning. We conduct extensive experiments on the FaceForensics++ dataset. Results show the great effectiveness of our method and we also achieve very competitive performance against state-of-the-art methods. Chutian Wang, Cairong Zhao, Guosheng Hu |
ICME | 3 |
| 2022 | Self-attention neural architecture search for semantic image segmentation
Zhenkun Fan, Guosheng Hu, Xin Sun 0003, Gaige Wang, Junyu Dong, Chi Su |
Knowl. Based Syst. | 2 |
| 2022 | MetaMixUp: Learning Adaptive Interpolation Policy of MixUp With MetalearningabstractMixUp is an effective data augmentation method to regularize deep neural networks via random linear interpolations between pairs of samples and their labels. It plays an important role in model regularization, semisupervised learning (SSL), and domain adaption. However, despite its empirical success, its deficiency of randomly mixing samples has poorly been studied. Since deep networks are capable of memorizing the entire data set, the corrupted samples generated by vanilla MixUp with a badly chosen interpolation policy will degrade the performance of networks. To overcome overfitting to corrupted samples, inspired by metalearning (learning to learn), we propose a novel technique of learning to a mixup in this work, namely, MetaMixUp. Unlike the vanilla MixUp that samples interpolation policy from a predefined distribution, this article introduces a metalearning-based online optimization approach to dynamically learn the interpolation policy in a data-adaptive way (learning to learn better). The validation set performance via metalearning captures the noisy degree, which provides optimal directions for interpolation policy learning. Furthermore, we adapt our method for pseudolabel-based SSL along with a refined pseudolabeling strategy. In our experiments, our method achieves better performance than vanilla MixUp and its variants under SL configuration. In particular, extensive experiments show that our MetaMixUp adapted SSL greatly outperforms MixUp and many state-of-the-art methods on CIFAR-10 and SVHN benchmarks under the SSL configuration. Zhijun Mai, Guosheng Hu, Dexiong Chen, Fumin Shen, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | MAMBA: Multi-level Aggregation via Memory Bank for Video Object DetectionabstractState-of-the-art video object detection methods maintain a memory structure, either a sliding window or a memory queue, to enhance the current frame using attention mechanisms. However, we argue that these memory structures are not efficient or sufficient because of two implied operations: (1) concatenating all features in memory for enhancement, leading to a heavy computational cost; (2) frame-wise memory updating, preventing the memory from capturing more temporal information. In this paper, we propose a multi-level aggregation architecture via memory bank called MAMBA. Specifically, our memory bank employs two novel operations to eliminate disadvantages of existing methods: (1) light-weight key-set construction which can significantly reduce the computational cost; (2) fine-grained feature-wise updating strategy which enables our method to utilize knowledge from the whole video. To better enhance features from complementary levels, i.e., feature maps and proposals, we further propose a generalized enhancement operation (GEO) to aggregate multi-level features in a unified manner. We conduct extensive evaluations on the challenging ImageNetVID dataset. Compared with existing state-of-the-art methods, our method achieves superior performance in terms of both speed and accuracy. More remarkably, MAMBA achieves mAP of 83.7%/84.6% at 12.6/9.1 FPS with ResNet-101. Guanxiong Sun, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002 |
AAAI | 3 |
| 2021 | OPANAS: One-Shot Path Aggregation Network Architecture Search for Object DetectionabstractRecently, neural architecture search (NAS) has been exploited to design feature pyramid networks (FPNs) and achieved promising results for visual object detection. Encouraged by the success, we propose a novel One-Shot Path Aggregation Network Architecture Search (OPANAS) algorithm, which significantly improves both searching efficiency and detection accuracy. Specifically, we first introduce six heterogeneous information paths to build our search space, namely top-down, bottom-up, fusing-splitting, scale-equalizing, skip-connect and none. Second, we propose a novel search space of FPNs, in which each FPN candidate is represented by a densely-connected directed acyclic graph (each node is a feature pyramid and each edge is one of the six heterogeneous information paths). Third, we propose an efficient one-shot search method to find the optimal path aggregation architecture; specifically, we first train a super-net and then find the optimal candidate with an evolutionary algorithm. Experimental results demonstrate the efficacy of the proposed OPANAS for object detection: (1) OPANAS is more efficient than state-of-the-art methods (e.g., NAS-FPN and Auto-FPN) at significantly smaller searching cost (e.g., only 4 GPU days on MS-COCO); (2) the optimal architecture found by OPANAS significantly improves main-stream detectors including RetinaNet, Faster R-CNN and Cascade R-CNN, by 2.3∼3.2 % mAP compared to their FPN counterparts; and (3) a new state-of-the-art accuracy-speed trade-off (52.2 % mAP at 7.6 FPS) is achieved at smaller training costs than comparable recent arts. Code will be released at https://github.com/VDIGPKU/OPANAS. Tingting Liang, Yongtao Wang, Zhi Tang 0001, Guosheng Hu, Haibin Ling |
CVPR | 4 |
| 2021 | Refining Single Low-Quality Facial Depth Map by Lightweight and Efficient Deep ModelabstractConsumer depth sensors have become increasingly common, however, the data are rather coarse and noisy, which is problematic to delicate tasks, such as 3D face modeling and 3D face recognition. In this paper, we present a novel and lightweight 3D Face Refinement Model (3D-FRM), to effectively and efficiently improve the quality of such single facial depth maps. 3D-FRM has an encoder-decoder structure, where the encoder applies depth-wise, point-wise convolutions and the fusion of features of different receptive fields to capture original discriminative information, and the decoder exploits sub-pixel convolutions and the combination of low- and high-level features to achieve strong shape recovery. We also propose a joint loss function to smooth facial surfaces and preserve their identities. In addition, we contribute a large dataset with low- and high-quality 3D face pairs to facilitate this research. Extensive experiments are conducted on the Bosphorus and Lock3DFace datasets, and results show the competency of the proposed method at ameliorating both visual quality and recognition accuracy. Code and data will be available at https://github.com/muyouhang/3D-FRM. Guodong Mu, Di Huang 0001, Weixin Li 0001, Guosheng Hu, Yunhong Wang 0001 |
IJCB | 4 |
| 2021 | DPT: Deformable Patch-based Transformer for Visual RecognitionabstractTransformer has achieved great success in computer vision, while how to split patches in an image remains a problem. Existing methods usually use a fixed-size patch embedding which might destroy the semantics of objects. To address this problem, we propose a new Deformable Patch (DePatch) module which learns to adaptively split the images into patches with different positions and scales in a data-driven way rather than using predefined fixed patches. In this way, our method can well preserve the semantics in patches. The DePatch module can work as a plug-and-play module, which can easily be incorporated into different transformers to achieve an end-to-end training. We term this DePatch-embedded transformer as Deformable Patch-based Transformer (DPT) and conduct extensive evaluations of DPT on image classification and object detection. Results show DPT can achieve 81.8% top-1 accuracy on ImageNet classification, and 43.7% box AP with RetinaNet, 44.3% with Mask R-CNN on MSCOCO object detection. Code has been made available at: https://github.com/CASIA-IVA-Lab/DPT. Zhiyang Chen 0002, Yousong Zhu, Chaoyang Zhao, Guosheng Hu, Wei Zeng 0006, Jinqiao Wang, Ming Tang 0001 |
ACM Multimedia | 4 |
| 2021 | Semi-Supervised Face Frontalization in the WildabstractSynthesizing a frontal view face from a single nonfrontal image, i.e. face frontalization, is a task of practical importance in a wide range of facial image analysis applications. However, to train the frontalization model in a supervised manner, most existing face frontalization methods rely on the availability of nonfrontal-frontal face pairs (typically from the Multi-PIE dataset) captured in a constrained environment. Such approaches, in return, limit the generalizability of their application to unconstrained scenarios. Unfortunately, although a large amount of in-the-wild face datasets are available, they cannot easily be utilized for face frontalization training since the nonfrontal and frontal facial images are not paired. To train a frontalization network which generalizes well to both constrained and unconstrained environments, we propose a semi-supervised learning framework which effectively uses both (labeled) indoor and (unlabeled) outdoor faces. Specifically, to achieve this goal, this article presents a Cycle-Consistent Face Frontalization Generative Adversarial Network (CCFF-GAN) which consists of both (1) the supervised and (2) the unsupervised components. For (1), we use the indoor paired (labeled) data to learn a roughly accurate frontalization network which may not generalize well to outdoor (in-the-wild) scenarios. For (2), to cope with the generalization issue, the unsupervised part uses the unpaired (unlabeled) images under the perceptual cycle consistency constraint in the semantic feature space to generalize the network from controlled (indoor) to uncontrolled (outdoor) environment. Extensive experiments demonstrate the effectiveness of the proposed method in comparison with the state-of-the-art face frontalization methods, especially under the in-the-wild scenarios. Zhihong Zhang 0001, Ruiyang Liang, Xu Chen 0020, Xuexin Xu, Guosheng Hu, Wangmeng Zuo, Edwin R. Hancock |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | SP-GAN: Self-Growing and Pruning Generative Adversarial NetworksabstractThis article presents a new Self-growing and Pruning Generative Adversarial Network (SP-GAN) for realistic image generation. In contrast to traditional GAN models, our SP-GAN is able to dynamically adjust the size and architecture of a network in the training stage by using the proposed self-growing and pruning mechanisms. To be more specific, we first train two seed networks as the generator and discriminator; each contains a small number of convolution kernels. Such small-scale networks are much easier and faster to train than large-capacity networks. Second, in the self-growing step, we replicate the convolution kernels of each seed network to augment the scale of the network, followed by fine-tuning the augmented/expanded network. More importantly, to prevent the excessive growth of each seed network in the self-growing stage, we propose a pruning strategy that reduces the redundancy of an augmented network, yielding the optimal scale of the network. Finally, we design a new adaptive loss function that is treated as a variable loss computational process for the training of the proposed SP-GAN model. By design, the hyperparameters of the loss function can dynamically adapt to different training stages. Experimental results obtained on a set of data sets demonstrate the merits of the proposed method, especially in terms of the stability and efficiency of network training. The source code of the proposed SP-GAN method is publicly available at https://github.com/Lambert-chen/SPGAN.git. Xiaoning Song, Zhenhua Feng 0001, Guosheng Hu, Dongjun Yu, Xiaojun Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Imbalance Robust Softmax for Deep Embeeding Learning
Hao Zhu 0010, Guosheng Hu, Neil Robertson 0002 |
ACCV (5) | 3 |
| 2020 | Training Noise-Robust Deep Neural Networks via Meta-LearningabstractLabel noise may significantly degrade the performance of Deep Neural Networks (DNNs). To train noise-robust DNNs, Loss correction (LC) approaches have been introduced. LC approaches assume the noisy labels are corrupted from clean (ground-truth) labels by an unknown noise transition matrix T. The backbone DNNs and T can be trained separately, where T is approximated with prior knowledge. For example, T is constructed by stacking the maximum or mean predictions of the samples from each class. In this work, we pro- pose a new loss correction approach, named as Meta Loss Correction (MLC), to directly learn T from data via the meta-learning framework. The MLC is model-agnostic and learns T from data rather than heuristically approximates it using prior knowledge. Extensive evaluations are conducted on computer vision (MNIST, CIFAR-10, CIFAR-100, Clothing1M) and natural language processing (Twitter) datasets. The experimental results show that MLC achieves very competitive performance against state-of-the-art approaches. Zhen Wang 0033, Guosheng Hu, Qinghua Hu |
CVPR | 2 |
| 2020 | Reducing Distributional Uncertainty by Mutual Information Maximisation and Transferable Feature Learning
Jian Gao 0018, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002 |
ECCV (23) | 3 |
| 2020 | Differentiable Automatic Data Augmentation
Yonggang Li 0001, Guosheng Hu, Yongtao Wang, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang |
ECCV (22) | 2 |
| 2020 | Learning Flow-Based Feature Warping for Face Frontalization with Illumination Inconsistent Supervision
Yuxiang Wei 0001, Ming Liu 0018, Haolin Wang 0004, Ruifeng Zhu, Guosheng Hu, Wangmeng Zuo |
ECCV (12) | 5 |
| 2020 | Adaptive Variance Based Label Distribution Learning for Facial Age Estimation
Xin Wen 0005, Biying Li, Haiyun Guo, Zhiwei Liu 0004, Guosheng Hu, Ming Tang 0001, Jinqiao Wang |
ECCV (23) | 5 |
| 2020 | CRSSC: Salvage Reusable Samples from Noisy Data for Robust LearningabstractDue to the existence of label noise in web images and the high memorization capacity of deep neural networks, training deep fine-grained (FG) models directly through web images tends to have an inferior recognition ability. In the literature, to alleviate this issue, loss correction methods try to estimate the noise transition matrix, but the inevitable false correction would cause severe accumulated errors. Sample selection methods identify clean ("easy") samples based on the fact that small losses can alleviate the accumulated errors. However, "hard" and mislabeled examples that can both boost the robustness of FG models are also dropped. To this end, we propose a certainty-based reusable sample selection and correction approach, termed as CRSSC, for coping with label noise in training deep FG models with web images. Our key idea is to additionally identify and correct reusable samples, and then leverage them together with clean examples to update the networks. We demonstrate the superiority of the proposed approach from both theoretical and experimental perspectives. Zeren Sun, Xian-Sheng Hua 0001, Yazhou Yao, Xiu-Shen Wei, Guosheng Hu, Jian Zhang 0002 |
ACM Multimedia | 5 |
| 2020 | Exploring ubiquitous relations for boosting classification and localization
Xin Sun 0003, Changrui Chen, Junyu Dong, Guosheng Hu |
Knowl. Based Syst. | 5 |
| 2020 | Attention-Based Two-Stream Convolutional Networks for Face Spoofing DetectionabstractSince the human face preserves the richest information for recognizing individuals, face recognition has been widely investigated and achieved great success in various applications in the past decades. However, face spoofing attacks (e.g., face video replay attack) remain a threat to modern face recognition systems. Though many effective methods have been proposed for anti-spoofing, we find that the performance of many existing methods is degraded by illuminations. It motivates us to develop illumination-invariant methods for anti-spoofing. In this paper, we propose a two-stream convolutional neural network (TSCNN), which works on two complementary spaces: RGB space (original imaging space) and multi-scale retinex (MSR) space (illumination-invariant space). Specifically, the RGB space contains the detailed facial textures, yet it is sensitive to illumination; MSR is invariant to illumination, yet it contains less detailed facial information. In addition, the MSR images can effectively capture the high-frequency information, which is discriminative for face spoofing detection. Images from two spaces are fed to the TSCNN to learn the discriminative features for anti-spoofing. To effectively fuse the features from two sources (RGB and MSR), we propose an attention-based fusion method, which can effectively capture the complementarity of two features. We evaluate the proposed framework on various databases, i.e., CASIA-FASD, REPLAY-ATTACK, and OULU, and achieve very competitive performance. To further verify the generalization capacity of the proposed strategies, we conduct cross-database experiments, and the results show the great effectiveness of our method. Haonan Chen 0003, Guosheng Hu, Zhen Lei 0001, Yaowu Chen, Neil Robertson 0002, Stan Z. Li |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | Learning Symmetry Consistent Deep CNNs for Face CompletionabstractDeep convolutional networks (CNNs) have achieved great success in face completion to generate plausible facial structures. These methods, however, are limited in maintaining global consistency among face components and recovering fine facial details. On the other hand, reflectional symmetry is a prominent property of face images and benefits face analysis and consistency modeling, yet remaining uninvestigated in deep face completion. In this work, we leverage two kinds of symmetry-enforcing modules to form a symmetry-consistent CNN model (i.e., SymmFCNet) for effective face completion. For missing pixels on only one of the half-faces, an illumination-reweighted warping subnet is developed to guide the warping and illumination reweighting of the other half-face. As for missing pixels on both of half-faces, we present a generative reconstruction subnet together with a perceptual symmetry loss to enforce symmetry consistency of recovered structures. The SymmFCNet is constructed by stacking generative reconstruction subnet upon illumination-reweighted warping subnet, and can be learned in an end-to-end manner. Experiments show that SymmFCNet can generate globally consistent results on images with synthetic and real occlusions, and performs favorably against state-of-the-arts. Xiaoming Li 0002, Guosheng Hu, Jieru Zhu, Wangmeng Zuo, Meng Wang 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 2 |
| 2019 | Deep Metric Learning by Online Soft Mining and Class-Aware AttentionabstractDeep metric learning aims to learn a deep embedding that can capture the semantic similarity of data points. Given the availability of massive training samples, deep metric learning is known to suffer from slow convergence due to a large fraction of trivial samples. Therefore, most existing methods generally resort to sample mining strategies for selecting nontrivial samples to accelerate convergence and improve performance. In this work, we identify two critical limitations of the sample mining methods, and provide solutions for both of them. First, previous mining methods assign one binary score to each sample, i.e., dropping or keeping it, so they only selects a subset of relevant samples in a mini-batch. Therefore, we propose a novel sample mining method, called Online Soft Mining (OSM), which assigns one continuous score to each sample to make use of all samples in the mini-batch. OSM learns extended manifolds that preserve useful intraclass variances by focusing on more similar positives. Second, the existing methods are easily influenced by outliers as they are generally included in the mined subset. To address this, we introduce Class-Aware Attention (CAA) that assigns little attention to abnormal data samples. Furthermore, by combining OSM and CAA, we propose a novel weighted contrastive loss to learn discriminative embeddings. Extensive experiments on two fine-grained visual categorisation datasets and two video-based person re-identification benchmarks show that our method significantly outperforms the state-of-the-art. Xinshao Wang, Yang Hua 0001, Elyor Kodirov, Guosheng Hu, Neil Robertson 0002 |
AAAI | 4 |
| 2019 | Semantic Alignment: Finding Semantically Consistent Ground-Truth for Facial Landmark DetectionabstractRecently, deep learning based facial landmark detection has achieved great success. Despite this, we notice that the semantic ambiguity greatly degrades the detection performance. Specifically, the semantic ambiguity means that some landmarks (e.g. those evenly distributed along the face contour) do not have clear and accurate definition, causing inconsistent annotations by annotators. Accordingly, these inconsistent annotations, which are usually provided by public databases, commonly work as the ground-truth to supervise network training, leading to the degraded accuracy. To our knowledge, little research has investigated this problem. In this paper, we propose a novel probabilistic model which introduces a latent variable, i.e. the `real' ground-truth which is semantically consistent, to optimize. This framework couples two parts (1) training landmark detection CNN and (2) searching the `real' ground-truth. These two parts are alternatively optimized: the searched `real' ground-truth supervises the CNN training; and the trained CNN assists the searching of `real' ground-truth. In addition, to recover the unconfidently predicted landmarks due to occlusion and low quality, we propose a global heatmap correction unit (GHCU) to correct outliers by considering the global face shape as a constraint. Extensive experiments on both image-based (300W and AFLW) and video-based (300-VW) databases demonstrate that our method effectively improves the landmark detection accuracy and achieves the state of the art performance. Zhiwei Liu 0004, Xiangyu Zhu 0001, Guosheng Hu, Haiyun Guo, Ming Tang 0001, Zhen Lei 0001, Neil Robertson 0002, Jinqiao Wang |
CVPR | 3 |
| 2019 | Led3D: A Lightweight and Efficient Deep Approach to Recognizing Low-Quality 3D FacesabstractDue to the intrinsic invariance to pose and illumination changes, 3D Face Recognition (FR) has a promising potential in the real world. 3D FR using high-quality faces, which are of high resolutions and with smooth surfaces, have been widely studied. However, research on that with low-quality input is limited, although it involves more applications. In this paper, we focus on 3D FR using low-quality data, targeting an efficient and accurate deep learning solution. To achieve this, we work on two aspects: (1) designing a lightweight yet powerful CNN; (2) generating finer and bigger training data. For (1), we propose a Multi-Scale Feature Fusion (MSFF) module and a Spatial Attention Vectorization (SAV) module to build a compact and discriminative CNN. For (2), we propose a data processing system including point-cloud recovery, surface refinement, and data augmentation (with newly proposed shape jittering and shape scaling). We conduct extensive experiments on Lock3DFace and achieve state-of-the-art results, outperforming many heavy CNNs such as VGG-16 and ResNet-34. In addition, our model can operate at a very high speed (136 fps) on Jetson TX2, and the promising accuracy and efficiency reached show its great applicability on edge/mobile devices. Guodong Mu, Di Huang 0001, Guosheng Hu, Yunhong Wang 0001 |
CVPR | 3 |
| 2019 | Ranked List Loss for Deep Metric LearningabstractThe objective of deep metric learning (DML) is to learn embeddings that can capture semantic similarity information among data points. Existing pairwise or tripletwise loss functions used in DML are known to suffer from slow convergence due to a large proportion of trivial pairs or triplets as the model improves. To improve this, rankingmotivated structured losses are proposed recently to incorporate multiple examples and exploit the structured information among them. They converge faster and achieve state-of-the-art performance. In this work, we present two limitations of existing ranking-motivated structured losses and propose a novel ranked list loss to solve both of them. First, given a query, only a fraction of data points is incorporated to build the similarity structure. Consequently, some useful examples are ignored and the structure is less informative. To address this, we propose to build a setbased similarity structure by exploiting all instances in the gallery. The samples are split into a positive set and a negative set. Our objective is to make the query closer to the positive set than to the negative set by a margin. Second, previous methods aim to pull positive pairs as close as possible in the embedding space. As a result, the intraclass data distribution might be dropped. In contrast, we propose to learn a hypersphere for each class in order to preserve the similarity structure inside it. Our extensive experiments show that the proposed method achieves state-of-the-art performance on three widely used benchmarks. Xinshao Wang, Yang Hua 0001, Elyor Kodirov, Guosheng Hu, Romain Garnier, Neil Robertson 0002 |
CVPR | 4 |
| 2019 | Learning Discriminative and Complementary Patches for Face RecognitionabstractThe ensemble of convolutional neural networks (CNNs) has widely been used in many computer vision tasks including face recognition. Many existing ensembles of face recognition CNNs apply a two-stage pipeline to target performance improvement [10], [20], [22], [23], [29]: (1) it trains multiple CNNs separately with many face patches covering different facial areas; (2) the features derived from different models are aggregated off-line by different fusion methods. The well-known face recognition work, DeepID2 [20] trains 200 networks based on 200 arbitrarily chosen facial areas and chooses the best 25 ones to achieve impressive performance. However, it is very time-consuming to train so many networks. In addition, a brute-force like way of choosing facial patches is used without knowing which face patches are complementary and discriminative. It might be lack of generalization capability for cross-database applications. To solve that, we propose a novel end-to-end CNN ensemble architecture which automatically learns the complementary and discriminative patches for face recognition. Specifically, we propose a novel Patch Generation Engine (PGE) with Patch Search Spatial Transformer Network (PS-STN) and ROI shrunk loss to perform the patch selection process. ROI shrunk loss enlarges the distance of learned features in spatial space and feature space and learn complementary features. In order to get final aggregated feature, we use a supervised fusion module named Two Stage Discriminative Fusion Module (TSDFM) which effective to capture the global and local information and further guide the PGE to learn better patches. Extensive experiments conducted on LFW and YTF datasets show the effectiveness of our novel end-to-end ensemble method. Zhiwei Liu 0004, Ming Tang 0001, Guosheng Hu, Jinqiao Wang |
FG | 3 |
| 2019 | Collaborative representation based face classification exploiting block weighted LBP and analysis dictionary learning
Xiaoning Song, Youming Chen, Zhenhua Feng 0001, Guosheng Hu, Tao Zhang 0010, Xiaojun Wu 0001 |
Pattern Recognit. | 4 |
| 2019 | Fast SRC using quadratic optimisation in downsized coefficient solution subspace
Xiaoning Song, Guosheng Hu, Jian-Hao Luo, Zhenhua Feng 0001, Dongjun Yu, Xiaojun Wu 0001 |
Signal Process. | 2 |
| 2019 | Face Frontalization Using an Appearance-Flow-Based Convolutional Neural NetworkabstractFacial pose variation is one of the major factors making face recognition (FR) a challenging task. One popular solution is to convert non-frontal faces to frontal ones on which FR is performed. Rotating faces causes facial pixel value changes. Therefore, existing CNN-based methods learn to synthesize frontal faces in color space. However, this learning problem in a color space is highly non-linear, causing the synthetic frontal faces to lose fine facial textures. In this paper, we take the view that the nonfrontal-frontal pixel changes are essentially caused by geometric transformations (rotation, translation, and so on) in space. Therefore, we aim to learn the nonfrontal-frontal facial conversion in the spatial domain rather than the color domain to ease the learning task. To this end, we propose an appearance-flow-based face frontalization convolutional neural network (A3F-CNN). Specifically, A3F-CNN learns to establish the dense correspondence between the non-frontal and frontal faces. Once the correspondence is built, frontal faces are synthesized by explicitly "moving" pixels from the non-frontal one. In this way, the synthetic frontal faces can preserve fine facial textures. To improve the convergence of training, an appearance-flow-guided learning strategy is proposed. In addition, generative adversarial network loss is applied to achieve a more photorealistic face, and a face mirroring method is introduced to handle the self-occlusion problem. Extensive experiments are conducted on face synthesis and pose invariant FR. Results show that our method can synthesize more photorealistic faces than the existing methods in both the controlled and uncontrolled lighting environments. Moreover, we achieve a very competitive FR performance on the Multi-PIE, LFW and IJB-A databases. Zhihong Zhang 0001, Xu Chen 0020, Beizhan Wang, Guosheng Hu, Wangmeng Zuo, Edwin R. Hancock |
IEEE Trans. Image Process. | 4 |
| 2018 | Deep Multi-task Learning to Recognise Subtle Facial Expressions of Mental States
Guosheng Hu, Li Liu 0004, Yang Hua 0001, Zhihong Zhang 0001, Fumin Shen, Ling Shao 0001, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang |
ECCV (12) | 1 |
| 2018 | Deep Stock Representation Learning: From Candlestick Charts to Investment DecisionsabstractWe propose a novel investment decision strategy (IDS) based on deep learning. The performance of many IDSs is affected by stock similarity. Most existing stock similarity measurements have the problems: (a) The linear nature of many measurements cannot capture nonlinear stock dynamics; (b) The estimation of many similarity metrics (e.g. covariance) needs very long period historic data (e.g. 3K days) which cannot represent current market effectively; (c) They cannot capture translation-invariance. To solve these problems, we apply Convolutional AutoEncoder to learn a stock representation, based on which we propose a novel portfolio construction strategy by: (i) using the deeply learned representation and modularity optimisation to cluster stocks and identify diverse sectors, (ii) picking stocks within each cluster according to their Sharpe ratio (Sharpe 1994). Overall this strategy provides low-risk high-return portfolios. We use the Financial Times Stock Exchange 100 Index (FTSE 100) data for evaluation. Results show our portfolio outperforms FTSE 100 index and many well known funds in terms of total return in 2000 trading days. Guosheng Hu, Kai Yang 0031, Flood Sung, Zhihong Zhang 0001, Neil Robertson 0002, Timothy M. Hospedales, Qiangwei Miemie |
ICASSP | 1 |
| 2018 | A Unified Neighbor Reconstruction Method for EmbeddingsabstractIn this work we propose a novel and compact Neighbor Reconstruction Method (NRM) which is a unified pre-processing method for graph-based sparse spectral algorithms. This method is conducted by vector operations on a central point and its corresponding neighbor points. NRM generates new neighbor points which can capture the local space structure of the central point more appropriately than original neighbor points. With NRM, a large number of sparse spectral based nonlinear feature extraction and selection algorithms gain significant improvement. Specifically, we embedded NRM to several classical algorithms, Local Linear Embedding (LLE) [1], Laplacian Eigenmaps (LE) [2] and Unsupervised Feature Selection for Multi-cluster Data (MCFS) [3], with accuracy improvement of up to 7%, 2.6%, 2.4% on ORL, CIFAR 10, and MINST data sets respectively. We also apply NRM to a Super Resolution algorithm, A+ [5], and obtain 0.12dB improvement than original method. Zhiling Ye, Zhihong Zhang 0001, Lu Bai 0001, Guosheng Hu, Zheng-Jian Bai, Yiqun Hu, Edwin R. Hancock |
ICPR | 4 |
| 2018 | Recovering variations in facial albedo from low resolution images
Xu Chen 0020, Zhihong Zhang 0001, Beizhan Wang, Guosheng Hu, Edwin R. Hancock |
Pattern Recognit. | 4 |
| 2018 | Dictionary Integration Using 3D Morphable Face Models for Pose-Invariant Collaborative-Representation-Based ClassificationabstractThe paper presents a dictionary integration algorithm using 3D morphable face models (3DMM) for pose-invariant collaborative-representation-based face classification. To this end, we first fit a 3DMM to the 2D face images of a dictionary to reconstruct the 3D shape and texture of each image. The 3D faces are used to render a number of virtual 2D face images with arbitrary pose variations to augment the training data, by merging the original and rendered virtual samples to create an extended dictionary. Second, to reduce the information redundancy of the extended dictionary and improve the sparsity of reconstruction coefficient vectors using collaborative-representation-based classification (CRC), we exploit an on-line class elimination scheme to optimise the extended dictionary by identifying the training samples of the most representative classes for a given query. The final goal is to perform pose-invariant face classification using the proposed dictionary integration method and the on-line pruning strategy under the CRC framework. Experimental results obtained for a set of well-known face data sets demonstrate the merits of the proposed method, especially its robustness to pose variations. Xiaoning Song, Zhenhua Feng 0001, Guosheng Hu, Josef Kittler, Xiaojun Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Frankenstein: Learning Deep Face Representations Using Small DataabstractDeep convolutional neural networks have recently proven extremely effective for difficult face recognition problems in uncontrolled settings. To train such networks, very large training sets are needed with millions of labeled images. For some applications, such as near-infrared (NIR) face recognition, such large training data sets are not publicly available and difficult to collect. In this paper, we propose a method to generate very large training data sets of synthetic images by compositing real face images in a given data set. We show that this method enables to learn models from as few as 10 000 training images, which perform on par with models trained from 500 000 images. Using our approach, we also obtain state-of-the-art results on the CASIA NIR-VIS2.0 heterogeneous face recognition data set. Guosheng Hu, Xiaojiang Peng, Yongxin Yang, Timothy M. Hospedales, Jakob Verbeek |
IEEE Trans. Image Process. | 1 |
| 2017 | Attribute-Enhanced Face Recognition with Neural Tensor Fusion NetworksabstractDeep learning has achieved great success in face recognition, however deep-learned features still have limited invariance to strong intra-personal variations such as large pose changes. It is observed that some facial attributes (e.g. eyebrow thickness, gender) are robust to such variations. We present the first work to systematically explore how the fusion of face recognition features (FRF) and facial attribute features (FAF) can enhance face recognition performance in various challenging scenarios. Despite the promise of FAF, we find that in practice existing fusion methods fail to leverage FAF to boost face recognition performance in some challenging scenarios. Thus, we develop a powerful tensor-based framework which formulates feature fusion as a tensor optimisation problem. It is nontrivial to directly optimise this tensor due to the large number of parameters to optimise. To solve this problem, we establish a theoretical equivalence between low-rank tensor optimisation and a two-stream gated neural network. This equivalence allows tractable learning using standard neural network optimisation tools, leading to accurate and stable optimisation. Experimental results show the fused feature works better than individual features, thus proving for the first time that facial attributes aid face recognition. We achieve state-of-the-art performance on three popular databases: MultiPIE (cross pose, lighting and expression), CASIA NIR-VIS2.0 (cross-modality environment) and LFW (uncontrolled environment). Guosheng Hu, Yang Hua 0001, Zhihong Zhang 0001, Sankha S. Mukherjee, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang |
ICCV | 1 |
| 2017 | Efficient 3D morphable face model fitting
Guosheng Hu, Fei Yan 0001, Josef Kittler, William J. Christmas, Chi-Ho Chan, Zhenhua Feng 0001, Patrik Huber 0001 |
Pattern Recognit. | 1 |
| 2017 | Half-Face Dictionary Integration for Representation-Based ClassificationabstractThis paper presents a half-face dictionary integration (HFDI) algorithm for representation-based classification. The proposed HFDI algorithm measures residuals between an input signal and the reconstructed one, using both the original and the synthesized dual-column (row) half-face training samples. More specifically, we first generate a set of virtual half-face samples for the purpose of training data augmentation. The aim is to obtain high-fidelity collaborative representation of a test sample. In this half-face integrated dictionary, each original training vector is replaced by an integrated dual-column (row) half-face matrix. Second, to reduce the redundancy between the original dictionary and the extended half-face dictionary, we propose an elimination strategy to gain the most robust training atoms. The last contribution of the proposed HFDI method is the use of a competitive fusion method weighting the reconstruction residuals from different dictionaries for robust face classification. Experimental results obtained from the Facial Recognition Technology, Aleix and Robert, Georgia Tech, ORL, and Carnegie Mellon University-pose, illumination and expression data sets demonstrate the effectiveness of the proposed method, especially in the case of the small sample size problem. Xiaoning Song, Zhenhua Feng 0001, Guosheng Hu, Xiaojun Wu 0001 |
IEEE Trans. Cybern. | 3 |
| 2016 | Face Recognition Using a Unified 3D Morphable Model
Guosheng Hu, Fei Yan 0001, Chi-Ho Chan, Weihong Deng, William J. Christmas, Josef Kittler, Neil Robertson 0002 |
ECCV (8) | 1 |
| 2016 | Face image super-resolution via weighted patches regressionabstractRecently sparse representation has gained great success in face image super-resolution. The conventional sparsity-based methods enforce sparse coding on face image patches and the representation fidelity is measured by ℓ2-norm. Such a sparse coding model regularizes all facial patches equally, which however ignores the natures of facial patches, where the facial patches in the different regions (patch positions) of human face may have distinct contributions to face image reconstruction. In this paper, we propose to weight facial patches based on their discriminative abilities in regression for robust face hallucination reconstruction. Specifically, we learn the weights for facial patches according to the information entropy in each face region, so as to highlight higher frequency details in face images and the facial discriminability can be well retrieved. Furthermore, the weighted sparse coding can reasonable represent the less sparse nature of noisy images and thus remarkably boosts noise robust performance in face image super-resolution. Various experimental results on standard face databased show that our proposed method outperforms state-of-the-art methods in terms of both objective metrics and visual quality. Zhihong Zhang 0001, Guosheng Hu, Edwin R. Hancock |
ICPR | 3 |
| 2016 | Comparison of analytical predictions of the noise floor due to static charge pump mismatch in fractional-n frequency synthesizersabstractInteraction between the requantizer's periodic output and charge pump mismatch nonlinearity in a fractional-N frequency synthesizer causes it to exhibit an elevated inband noise floor and spurs. In this paper, we consider three leading analytical predictions from the literature. By comparing simulation results with analytical predictions, we show that the two most common approaches fail to deal correctly with offsets, while the method based on Price's theorem works well. Hongjia Mo, Guosheng Hu, Michael Peter Kennedy |
ISCAS | 2 |
| 2015 | The noise and spur delusion in fractional-N frequency synthesizer designabstractThe standard design methodology for fractional-N frequency synthesizers assumes that the filtered shaped quantization noise from the requantizer is masked below the spectral envelope of the underlying integer-N synthesizer. Fractional-N frequency synthesizers are notorious for exhibiting an elevated noise floor and an unpredictable pattern of spurs. In this paper, we argue that designers should not be deluded by the overly conservative predictions of the simplified linear model but should instead consider nonlinearities as early as possible in the design process. Michael Peter Kennedy, Hongjia Mo, Zhida Li, Guosheng Hu, Paolo Scognamiglio, Ettore Napoli |
ISCAS | 4 |
| 2015 | Cascaded Collaborative Regression for Robust Facial Landmark Detection Trained Using a Mixture of Synthetic and Real Images With Dynamic WeightingabstractA large amount of training data is usually crucial for successful supervised learning. However, the task of providing training samples is often time-consuming, involving a considerable amount of tedious manual work. In addition, the amount of training data available is often limited. As an alternative, in this paper, we discuss how best to augment the available data for the application of automatic facial landmark detection. We propose the use of a 3D morphable face model to generate synthesized faces for a regression-based detector training. Benefiting from the large synthetic training data, the learned detector is shown to exhibit a better capability to detect the landmarks of a face with pose variations. Furthermore, the synthesized training data set provides accurate and consistent landmarks automatically as compared to the landmarks annotated manually, especially for occluded facial parts. The synthetic data and real data are from different domains; hence the detector trained using only synthesized faces does not generalize well to real faces. To deal with this problem, we propose a cascaded collaborative regression algorithm, which generates a cascaded shape updater that has the ability to overcome the difficulties caused by pose variations, as well as achieving better accuracy when applied to real faces. The training is based on a mix of synthetic and real image data with the mixing controlled by a dynamic mixture weighting schedule. Initially, the training uses heavily the synthetic data, as this can model the gross variations between the various poses. As the training proceeds, progressively more of the natural images are incorporated, as these can model finer detail. To improve the performance of the proposed algorithm further, we designed a dynamic multi-scale local feature extraction method, which captures more informative local features for detector training. An extensive evaluation on both controlled and uncontrolled face data sets demonstrates the merit of the proposed algorithm. Zhenhua Feng 0001, Guosheng Hu, Josef Kittler, William J. Christmas, Xiaojun Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Robust face recognition by an albedo based 3D morphable modelabstractLarge pose and illumination variations are very challenging for face recognition. The 3D Morphable Model (3DMM) approach is one of the effective methods for pose and illumination invariant face recognition. However, it is very difficult for the 3DMM to recover the illumination of the 2D input image because the ratio of the albedo and illumination contributions in a pixel intensity is ambiguous. Unlike the traditional idea of separating the albedo and illumination contributions using a 3DMM, we propose a novel Albedo Based 3D Morphable Model (AB3DMM), which removes the illumination component from the images using illumination normalisation in a preprocessing step. A comparative study of different illumination normalisation methods for this step is conducted on PIE and Multi-PIE databases. The results show that overall performance of our method outperforms state-of-the-art methods. Guosheng Hu, Chi-Ho Chan, Fei Yan 0001, William J. Christmas, Josef Kittler |
IJCB | 1 |
| 2014 | Robust compressed sensing with bounded and structured uncertaintiesabstractThe robust compressed sensing problem subject to a bounded and structured perturbation in the sensing matrix is solved in two steps. The alternating direction method of multipliers (ADMM) is first applied to obtain a robust support set. Unlike the existing robust signal recovery solutions, the proposed optimisation problem is convex. The ADMM algorithm that every subproblem has a global minimum is employed to solve the optimisation problem. Then, the standard robust regularised least‐squares problem restrained to the support is solved to reduce the recovery error. The numerical tests show that the proposed approach provides a robust estimation of support set, although it is conservative to recover signal magnitudes as a result of minimising the worst‐cast data error across all bounded perturbations. Xiangyun Qing, Guosheng Hu |
IET Signal Process. | 2 |
| 2012 | Resolution-Aware 3D Morphable ModelabstractThe 3D Morphable Model (3DMM) is currently receiving considerable attention for \nhuman face analysis. Most existing work focuses on fitting a 3DMM to high resolution \nimages. However, in many applications, fitting a 3DMM to low-resolution images \nis also important. In this paper, we propose a Resolution-Aware 3DMM (RA- \n3DMM), which consists of 3 different resolution 3DMMs: High-Resolution 3DMM \n(HR- 3DMM), Medium-Resolution 3DMM (MR-3DMM) and Low-Resolution 3DMM \n(LR-3DMM). RA-3DMM can automatically select the best model to fit the input images \nof different resolutions. The multi-resolution model was evaluated in experiments \nconducted on PIE and XM2VTS databases. The experimental results verified that HR- \n3DMM achieves the best performance for input image of high resolution, and MR- \n3DMM and LR-3DMM worked best for medium and low resolution input images, respectively. \nA model selection strategy incorporated in the RA-3DMM is proposed based \non these results. The RA-3DMM model has been applied to pose correction of face images \nranging from high to low resolution. The face verification results obtained with \nthe pose-corrected images show considerable performance improvement over the result \nwithout pose correction in all resolutions Guosheng Hu, Chi-Ho Chan, Josef Kittler, William J. Christmas |
BMVC | 1 |
| 2010 | Genetic Algorithms with Improved Simulated Binary Crossover and Support Vector Regression for Grid Resources Prediction
Guosheng Hu, Liang Hu 0001, Qinghai Bai, Guangyu Zhao |
ISNN (2) | 1 |
| 2010 | Support Vector Regression and Ant Colony Optimization for Grid Resources Prediction
Guosheng Hu, Liang Hu 0001, Pengchao Li, Xilong Che |
ISNN (2) | 1 |