EDBT 2026 Demo / reviewers in the wild / expert
Guodong Guo
dblp:92/4520
· DBLP profile ↗
190ranked-venue papers
34as first author
95since 2021 · last 2026
0000-0001-9583-0055ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 123 · 26 first-author · 62 since 2021Graphics, computer vision, multimedia, augmented reality and games · 110 · 25 first-author · 53 since 2021Security and privacy · 14 · 7 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 2 since 2021Computer networks · 5 · 4 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DiffRefiner: Coarse to Fine Trajectory Planning via Diffusion Refinement with Semantic Interaction for End to End Autonomous DrivingabstractUnlike discriminative approaches in autonomous driving that predict a fixed set of candidate trajectories of the ego vehicle, generative methods, such as diffusion models, learn the underlying distribution of future motion, enabling more flexible trajectory prediction. However, since these methods typically rely on denoising human-craft trajectory anchors or random noise, there remains significant room for improvement. In this paper, we propose DiffRefiner, a novel two-stage trajectory prediction framework. The first stage employs a transformer-based Proposal Decoder to generate coarse trajectory predictions by regressing from sensor inputs using predefined trajectory anchors. The second stage applies a Diffusion Refiner that iteratively denoises and refines these initial predictions. In this way, we enhance the performance of diffusion-based planning by incorporating a discriminative trajectory proposal module, which provides strong guidance for the generative refinement process. Furthermore, we design a fine-grained denoising decoder to enhance scene compliance, enabling more accurate trajectory prediction through enhanced alignment with the surrounding environment. Experimental results demonstrate that DiffRefiner achieves state-of-the-art performance, attaining 87.4 EPDMS on NAVSIM v2, and 87.1 DS along with 71.4 SR on Bench2Drive, thereby setting new records on both public benchmarks. The effectiveness of each component is validated via ablation studies as well. Liuhan Yin, Runkun Ju, Guodong Guo, Erkang Cheng |
AAAI | 3 |
| 2026 | LDFE: Laplacian Decoupled Feature Enhancement block for dual-stream CNN-based RGB-IR object detection
Xiaoyan Luo, Linlin Yang 0001, Haodong Zhu, Xiaorong Shi, Guodong Guo, Baochang Zhang 0001 |
Pattern Recognit. | 6 |
| 2026 | A Vision-Driven End-User Robot Programming Method Based on the Alignment Between Body Language and Motion CommandsabstractThis work is motivated by the complexity of conventional robot programming methods, which require a deep understanding of both robotics and software development. This situation presents significant challenges for nontechnical end-users who need to program a new robotic task. To address this issue, we propose a new robot programming method that enables humans to specify robot motion tasks using simple and intuitive body-language commands. To this end, we propose a triadic task decomposition framework to decompose complex robot motion tasks into three basic task units, i.e., location, operation, and motion. These units are paired with a set of body-language cues that enable end-users to specify a variety of actions and motions, providing an intuitive alternative to traditional robot programming methods. To program these units, we develop a multimodal fusion-based instruction-parsing algorithm with a dual-phase visual architecture, enabling temporally aware body-language recognition and high-accuracy cross-modal object localization. Furthermore, a semantics-adaptive task execution system is proposed, which incorporates an intent reasoning module that dynamically handles end-users' nonstandard programming sequences and incomplete task logic, enhancing programming success rates and system usability. We evaluate the proposed method on a dual-arm robotic platform and a humanoid robot through various pick-and-place tasks, and compare its performance with that of several baselines from the literature. Shipeng Lyu, Shenzeng Huo, Wanyu Ma, Guodong Guo, David Navarro-Alarcon |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2026 | Enhancing End-user Engagement in Human-Robot Interaction by Performing LLM-driven Expressive BehaviorsabstractWe introduce a novel framework with strong generalization capabilities for enabling humanoid robots to perform semantically grounded, context-aware, and physically realizable expressive behaviors, with the goal of enhancing end-user engagement in open-world human–robot interaction (HRI). To achieve this goal, we develop a new large language model–based human cognition module that interprets user dialogs to infer latent intent and generates high-level multi-modal behavior descriptions. These semantic representations are mapped to speech and motion outputs through a dual-stage embodied behavior generation pipeline. The pipeline consists of a shape adaptation module that maps human body motions into the robot’s kinematic space, followed by a motion retargeting module that generates executable joint trajectories under physical constraints. Additionally, the modular architecture enables seamless integration of state-of-the-art generative models and serves as a practical testbed for evaluating expressive behavior generation in real-world settings. We validate this system on a 58-DoF humanoid platform through both controlled video-based studies and live HRI experiments. The results show significant improvements over rule-based and handcrafted baselines in terms of expressiveness, behavioral appeal, and user engagement. This work helps bridging the gap between high-level language understanding and low-level robot control, thereby enabling scalable and human-aligned expressive behavior generation for embodied agents. Shipeng Lyu, Fangyuan Wang 0002, Guodong Guo, David Navarro-Alarcon |
ACM Trans. Hum. Robot Interact. | 4 |
| 2025 | Asymptotic Unbiased Sample Sampling to Speed Up Sharpness-Aware MinimizationabstractSharpness-Aware Minimization (SAM) has emerged as a promising approach for effectively reducing the generalization error. However, SAM incurs twice the computational cost compared to the base optimizer (e.g., SGD). We propose Asymptotic Unbiased data sampling to accelerate SAM (AUSAM), which maintains the model's generalization capacity while significantly enhancing computational efficiency. Concretely, we probabilistically sample a subset of data points beneficial for SAM optimization based on a theoretically guaranteed criterion, i.e., the Gradient Norm of each Sample (GNS). We further approximate the GNS by evaluating the difference in loss values before and after perturbation in SAM. As a plug-and-play, architecture-agnostic method, our approach consistently accelerates SAM across various tasks and networks, i.e., classification, human pose estimation, and network quantization. On CIFAR-10/100 and Tiny-ImageNet, AUSAM achieves results comparable to SAM while providing a speedup of over 70%. By adjusting hyperparameters, AUSAM can match the speed of the base optimizer while significantly surpassing the base optimizer's performance. Compared to recent dynamic data pruning methods, AUSAM is better suited for SAM and excels in maintaining performance. Additionally, AUSAM accelerates optimization in human pose estimation and model quantization without sacrificing performance, demonstrating its broad practicality. Jiaxin Deng, Junbiao Pang, Baochang Zhang 0001, Guodong Guo |
AAAI | 4 |
| 2025 | Dynamic Clustering Convolutional Neural NetworkabstractConvolutional neural networks (CNNs) have been playing a dominant role in computer vision. However, the existing approaches of using local window modeling in popular CNNs lack flexibility and hinder their ability to capture long-range dependencies of objects in an image. To overcome these limitations, we propose a novel CNN architecture, termed Dynamic Clustering Convolutional Neural Network (DCCNeXt). The proposed DCCNeXt takes a unique approach by employing global clustering to group image patches with similar semantics into clusters that are then convolved using the shared convolution kernels. To address the high computational complexity of global clustering, the feature vectors from each patch's subspace are extracted for efficient clustering, which makes the proposed model widely compatible with the downstream vision tasks. The extensive experiments of image classification, object detection, instance segmentation, and semantic segmentation on the benchmark datasets demonstrate that the proposed DCCNeXt outperforms the mainstream Convolutional Neural Networks (CNNs), Vision Transformers (ViTs), Vision Multi-layer Perceptrons (MLPs), Vision Graph Neural Networks (GNNs), and Vision Mambas. We anticipate that this study will provide a new perspective and a promising avenue for the design of convolutional neural networks. Tanzhe Li, Baochang Zhang 0001, Jiayi Lyu, Xiawu Zheng, Guodong Guo, Taisong Jin |
AAAI | 5 |
| 2025 | Instruction-Augmented Long-Horizon Planning: Embedding Grounding Mechanisms in Embodied Mobile ManipulationabstractEnabling humanoid robots to perform long-horizon mobile manipulation planning in real-world environments based on embodied perception and comprehension abilities has been a longstanding challenge. With the recent rise of large language models (LLMs), there has been a notable increase in the development of LLM-based planners. These approaches either utilize human-provided textual representations of the real world or heavily depend on prompt engineering to extract such representations, lacking the capability to quantitatively understand the environment, such as determining the feasibility of manipulating objects. To address these limitations, we present the Instruction-Augmented Long-Horizon Planning (IALP) system, a novel framework that employs LLMs to generate feasible and optimal actions based on real-time sensor feedback, including grounded knowledge of the environment, in a closed-loop interaction. Distinct from prior works, our approach augments user instructions into PDDL problems by leveraging both the abstract reasoning capabilities of LLMs and grounding mechanisms. By conducting various real-world long-horizon tasks, each consisting of seven distinct manipulatory skills, our results demonstrate that the IALP system can efficiently solve these tasks with an average success rate exceeding 80%. Our proposed method can operate as a high-level planner, equipping robots with substantial autonomy in unstructured environments through the utilization of multi-modal sensor inputs. Fangyuan Wang 0002, Shipeng Lyu, Peng Zhou 0018, Anqing Duan, Guodong Guo, David Navarro-Alarcon |
AAAI | 5 |
| 2025 | Graph Structure Refinement with Energy-based Contrastive LearningabstractGraph Neural Networks (GNNs) have recently gained widespread attention as a successful tool for analyzing graph-structured data. However, imperfect graph structure with noisy links lacks enough robustness and may damage graph representations, therefore limiting the GNNs' performance in practical tasks. Moreover, existing generative architectures fail to fit discriminative graph-related tasks. To tackle these issues, we introduce an unsupervised method based on a joint of generative training and discriminative training to learn graph structure and representation, aiming to improve the discriminative performance of generative models. We propose an Energy-based Contrastive Learning (ECL) guided Graph Structure Refinement (GSR) framework, denoted as ECL-GSR. To our knowledge, this is the first work to combine energy-based models with contrastive learning for GSR. Specifically, we leverage ECL to approximate the joint distribution of sample pairs, which increases the similarity between representations of positive pairs while reducing the similarity between negative ones. Refined structure is produced by augmenting and removing edges according to the similarity metrics among node representations. Extensive experiments demonstrate that ECL-GSR outperforms the state-of-the-art on eight benchmark datasets in node classification. ECL-GSR achieves faster training with fewer samples and memories against the leading baseline, highlighting its simplicity and efficiency in downstream tasks. Xianlin Zeng, Yufeng Wang 0004, Guodong Guo, Wenrui Ding, Baochang Zhang 0001 |
AAAI | 4 |
| 2025 | DFM: Differentiable Feature Matching for Anomaly DetectionabstractFeature matching methods for unsupervised anomaly detection have demonstrated impressive performance. Existing methods primarily rely on self-supervised training and handcrafted matching schemes for task adaptation. However, they can only achieve an inferior feature representation for anomaly detection because the feature extraction and matching modules are separately trained. To address these issues, we propose a Differentiable Feature Matching (DFM) framework for joint optimization of the feature extractor and the matching head. DFM transforms nearest-neighbor matching into a pooling-based module and embeds it within a Feature Matching Network (FMN). This design enables end-to-end feature extraction and feature matching module training, thus providing better feature representation for anomaly detection tasks. DFM is generic and can be incorporated into existing feature-matching methods. We implement DFM with various backbones and conduct extensive experiments across various tasks and datasets, demonstrating its effectiveness. Notably, we achieve state-of-the-art results in the continual anomaly detection task with instance-AUROC improvement of up to 3.9% and pixel-AP improvement of up to 5.5%. Yimi Wang, Yuguang Yang 0007, Runqi Wang, Guodong Guo, David S. Doermann, Baochang Zhang 0001 |
CVPR | 6 |
| 2025 | Efficient Low-Bit Quantization with Adaptive Scales for Multi-Task Co-TrainingabstractCo-training can achieve parameter-efficient multi-task models but remains unexplored for quantization-aware training. Our investigation shows that directly introducing co-training into existing quantization-aware training (QAT) methods results in significant performance degradation. Our experimental study identifies that the primary issue with existing QAT methods stems from the inadequate activation quantization scales for the co-training framework. To address this issue, we propose Task-Specific Scales Quantization for Multi-Task Co-Training (TSQ-MTC) to tackle mismatched quantization scales. Specifically, a task-specific learnable multi-scale activation quantizer (TLMAQ) is incorporated to enrich the representational ability of shared features for different tasks. Additionally, we find that in the deeper layers of the Transformer model, the quantized network suffers from information distortion within the attention quantizer. A structure-based layer-by-layer distillation (SLLD) is then introduced to ensure that the quantized features effectively preserve the information from their full-precision counterparts. Our extensive experiments in two co-training scenarios demonstrate the effectiveness and versatility of TSQ-MTC. In particular, we successfully achieve a 4-bit quantized low-level visual foundation model based on IPT, which attains a PSNR comparable to the full-precision model while offering a $7.99\times$ compression ratio in the $\times4$ super-resolution task on the Set5 benchmark. Linlin Yang 0001, Yanjing Li, Guodong Guo, Xianbin Cao 0001, Baochang Zhang 0001 |
ICLR | 5 |
| 2025 | Calibrated gradient descent of convolutional neural networks for embodied visual recognition
Sheng Xu 0007, Lian Zhuo, Baochang Zhang 0001, Yanjing Li, Guodong Guo |
Image Vis. Comput. | 7 |
| 2025 | MemRank: Memory-Augmented Similarity Ranking for Video-Based Depression Severity EstimationabstractDeep learning-based methods have shown substantial promise in visual depression severity estimation. Nonetheless, their effectiveness is limited by the scarce availability of labeled depression data, potentially leading to overfitting during representation learning. One feasible approach to address this issue is to incorporate, in the training objective, regularization that considers the unique characteristics of depression data. Typical regularization includes the similarity ranking through ordered consistency between visual features and their target scores. However, previous ranking methods are limited to using only samples within a mini-batch, resulting in a decreased regularization effect in depression representation learning. To address this limitation, we propose MemRank, a global similarity ranking method that operates not only on mini-batch samples but also on a well-designed feature memory, which stores smoothed and dynamically updated feature prototypes at diverse levels of depression during training. Furthermore, we show that incorporating the feature memory in the regression loss enhances the stability of training a deep regressor, leading to improved depression predictions. Empirically and analytically, we show that our MemRank outperforms alternative ranking methods and achieves state-of-the-art results on two benchmark datasets. Zeqiang Wei, Guodong Guo, Xiuzhuang Zhou |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Explicit-Implicit Subgoal Planning for Long-Horizon Tasks With Sparse RewardsabstractThe challenges inherent in long-horizon tasks in robotics persist due to the typical inefficient exploration and sparse rewards in traditional reinforcement learning approaches. To address these challenges, we have developed a novel algorithm, termed hlexplicit-implicit subgoal planning (EISP), designed to tackle long-horizon tasks through a divide-and-conquer approach. We utilize two primary criteria, feasibility and optimality, to ensure the quality of the generated subgoals. EISP consists of three components: a hybrid subgoal generator, a hindsight sampler, and a value selector. The hybrid subgoal generator uses an explicit model to infer subgoals and an implicit model to predict the final goal, inspired by way of human thinking that infers subgoals by using the current state and final goal as well as reason about the final goal conditioned on the current state and given subgoals. Additionally, the hindsight sampler selects valid subgoals from an offline dataset to enhance the feasibility of the generated subgoals. While the value selector utilizes the value function in reinforcement learning to filter the optimal subgoals from subgoal candidates. To validate our method, we conduct four long-horizon tasks in both simulation and the real world. The obtained quantitative and qualitative data indicate that our approach achieves promising performance compared to other baseline methods. These experimental results can be seen on the website https://sites.google.com/view/vaesi. Fangyuan Wang 0002, Anqing Duan, Peng Zhou 0018, Shengzeng Huo, Guodong Guo, Chenguang Yang 0001, David Navarro-Alarcon |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Fusion-Mamba for Cross-Modality Object DetectionabstractCross-modality object detection aims to fuse complementary information from different modalities to improve model performance, which achieves a wider range of applications. However, traditional cross-modality fusion methods, based on CNN or Transformer, inadequately address the issue of pseudo-target information, which causes model attention dispersion to degrade object detection performance. In this paper, we investigate a novel cross-modality fusion approach by associating cross-modal features in a hidden state space based on an improved Mamba with a gating attention mechanism. We propose theFusion-Mamba Block(FMB), designed to map cross-modal features into a hidden state space for interaction, thereby refining the model’s attention on true target areas and enhancing overall performance. The FMB comprises two key modules: State Space Channel Swapping (SSCS) module, which facilitates the fusion of shallow features, and Dual State Space Fusion (DSSF) module, which enables deep fusion and effectively suppresses pseudo-target information within the hidden state space. Our proposed method outperforms state-of-the-art approaches, achieving improvements of 5.9%, 3.5% and 2.1% mAP on$M^{3}$FD, DroneVehicle and FLIR-Aligned, respectively. To the best of our knowledge, this work establishes a new baseline for cross-modality object detection, providing a robust foundation for future research in this area. Haodong Zhu, Shaohui Lin, Xiaoyan Luo, Yunhang Shen, Guodong Guo, Baochang Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | Where You See Is What You Know: A Visual-Semantic Conceptual Explainer
Luhao Zhu, Xiangwei Kong 0001, Runsen Li, Guodong Guo |
MMAsia | 4 |
| 2024 | CLIP in Mirror: Disentangling text from visual images through reflectionabstractThe CLIP network excels in various tasks, but struggles with text-visual images i.e., images that contain both text and visual objects; it risks confusing textual and visual representations. To address this issue, we propose MirrorCLIP, a zero-shot framework, which disentangles the image features of CLIP by exploiting the difference in the mirror effect between visual objects and text in the images. Specifically, MirrorCLIP takes both original and flipped images as inputs, comparing their features dimension-wise in the latent space to generate disentangling masks. With disentangling masks, we further design filters to separate textual and visual factors more precisely, and then get disentangled representations. Qualitative experiments using stable diffusion models and class activation mapping (CAM) validate the effectiveness of our disentanglement. Moreover, our proposed MirrorCLIP reduces confusion when encountering text-visual images and achieves a substantial improvement on typographic defense, further demonstrating its superior ability of disentanglement. Our code is available at https://github.com/tcwangbuaa/MirrorCLIP Yuguang Yang 0007, Linlin Yang 0001, Shaohui Lin, Guodong Guo, Baochang Zhang 0001 |
NeurIPS | 6 |
| 2024 | Joint learning of foreground, background and edge for salient object detection
ZhiLei Chai, Guodong Guo |
Comput. Vis. Image Underst. | 4 |
| 2024 | Spatial-Temporal Attention Network for Depression Recognition from facial videos
Zhuhong Shao, Guodong Guo |
Expert Syst. Appl. | 5 |
| 2024 | NCL++: Nested Collaborative Learning for long-tailed visual recognition
Zichang Tan, Jun Li 0033, Jinhao Du, Jun Wan 0001, Zhen Lei 0001, Guodong Guo |
Pattern Recognit. | 6 |
| 2024 | Integrating Deep Facial Priors Into Landmarks for Privacy Preserving Multimodal Depression RecognitionabstractAutomatic depression diagnosis is a challenging problem, that requires integrating spatial-temporal information and extracting features from audio-visual signals. In terms of privacy protection, the development trend of recognition algorithms based on facial landmarks has created additional challenges and difficulties. In this paper, we propose an audio-visual attention network (AVA-DepressNet) for depression recognition. It is a novel multimodal framework with facial privacy protection, and uses attention-based modules to enhance audio-visual spatial and temporal features. In addition, an adversarial multistage (AMS) training strategy is developed to optimize the encoder-decoder structure. Additionally, facial structure prior knowledge is creatively used in AMS training. Our AVA-DepressNet is evaluated on popular audio-visual depression datasets: AVEC 2013, AVEC 2014, and AVEC 2017. The results show that our approach reaches the state-of-the-art performance or competitive results for depression recognition. Zhuhong Shao, Guodong Guo |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | Defending Black-Box Skeleton-Based Human Activity ClassifiersabstractSkeletal motions have been heavily relied upon for human activity recognition (HAR). Recently, a universal vulnerability of skeleton-based HAR has been identified across a variety of classifiers and data, calling for mitigation. To this end, we propose the first black-box defense method for skeleton-based HAR to our best knowledge. Our method is featured by full Bayesian treatments of the clean data, the adversaries and the classifier, leading to (1) a new Bayesian Energy-based formulation of robust discriminative classifiers, (2) a new adversary sampling scheme based on natural motion manifolds, and (3) a new post-train Bayesian strategy for black-box defense. We name our framework Bayesian Energy-based Adversarial Training or BEAT. BEAT is straightforward but elegant, which turns vulnerable black-box classifiers into robust ones without sacrificing accuracy. It demonstrates surprising and universal effectiveness across a wide range of skeletal HAR classifiers and datasets, under various attacks. Appendix and code are available. He Wang 0002, Yunfeng Diao, Zichang Tan, Guodong Guo |
AAAI | 4 |
| 2023 | Q-DETR: An Efficient Low-Bit Quantized Detection TransformerabstractThe recent detection transformer (DETR) has advanced object detection, but its application on resource-constrained devices requires massive computation and memory resources. Quantization stands out as a solution by representing the network in low-bit parameters and operations. However, there is a significant performance drop when performing low-bit quantized DETR (Q-DETR) with existing quantization methods. We find that the bottle-necks of Q-DETR come from the query information distortion through our empirical analyses. This paper addresses this problem based on a distribution rectification distillation (DRD). We formulate our DRD as a bi-level optimization problem, which can be derived by generalizing the information bottleneck (IB) principle to the learning of Q-DETR. At the inner level, we conduct a distribution alignment for the queries to maximize the self-information entropy. At the upper level, we introduce a new foreground-aware query matching scheme to effectively transfer the teacher information to distillation-desired features to minimize the conditional information entropy. Extensive experimental results show that our method performs much better than prior arts. For example, the 4-bit Q-DETR can theoretically accelerate DETR with ResNet-50 backbone by 6.6× and achieve 39.4% AP, with only 2.6% performance gaps than its real-valued counterpart on the COCO dataset11Code: https://github.com/SteveTsui/Q-DETR. Sheng Xu 0007, Yanjing Li, Mingbao Lin, Peng Gao 0007, Guodong Guo, Jinhu Lü 0001, Baochang Zhang 0001 |
CVPR | 5 |
| 2023 | Adaptive Sparse ViT: Towards Learnable Adaptive Token Pruning by Fully Exploiting Self-AttentionabstractVision transformer has emerged as a new paradigm in computer vision, showing excellent performance while accompanied by expensive computational cost. Image token pruning is one of the main approaches for ViT compression, due to the facts that the complexity is quadratic with respect to the token number, and many tokens containing only background regions do not truly contribute to the final prediction. Existing works either rely on additional modules to score the importance of individual tokens, or implement a fixed ratio pruning strategy for different input instances. In this work, we propose an adaptive sparse token pruning framework with a minimal cost. Specifically, we firstly propose an inexpensive attention head importance weighted class attention scoring mechanism. Then, learnable parameters are inserted as thresholds to distinguish informative tokens from unimportant ones. By comparing token attention scores and thresholds, we can discard useless tokens hierarchically and thus accelerate inference. The learnable thresholds are optimized in budget-aware training to balance accuracy and complexity, performing the corresponding pruning configurations for different input instances. Extensive experiments demonstrate the effectiveness of our approach. Our method improves the throughput of DeiT-S by 50% and brings only 0.2% drop in top-1 accuracy, which achieves a better trade-off between accuracy and latency than the previous methods. Xiangcheng Liu, Guodong Guo |
IJCAI | 3 |
| 2023 | DCP-NAS: Discrepant Child-Parent Neural Architecture Search for 1-bit CNNs
Yanjing Li, Sheng Xu 0007, Xianbin Cao 0001, Lian Zhuo, Baochang Zhang 0001, Tian Wang 0002, Guodong Guo |
Int. J. Comput. Vis. | 7 |
| 2023 | Few-Shot Learning with Complex-Valued Neural Networks and Dependable Learning
Runqi Wang, Baochang Zhang 0001, Guodong Guo, David S. Doermann |
Int. J. Comput. Vis. | 4 |
| 2023 | Exploring attribute localization and correlation for pedestrian attribute recognition
Dunfang Weng, Zichang Tan, Liwei Fang, Guodong Guo |
Neurocomputing | 4 |
| 2023 | Scene Text Detection Based on Multi-Dimensional Feature Fusion with Instance-Wise LossabstractOver the past few years, scene text detection has witnessed rapid progress due to the development of deep neural networks. However, segmentation-based methods may fail to detect pixels near boundary well, and scale variations of text instances may lead to small text missing. To tackle these problems, we propose a novel segmentation-based detector for scene text detection, which can improve the quality of the detected texts. Specifically, a Multi-dimensional Feature Fusion module is used to extract structural and spatial text features from the perspective of height, width and channel, which helps to improve the representation ability of the network. In order to obtain more accurate boundaries of the detected text instances, a Boundary Refinement Branch is introduced to strengthen the supervision for pixels adjacent to boundary. Meanwhile, we propose an Instance-wise Loss to deal with text instances of different scales. Extensive ablation studies validate the effectiveness of these proposed modules. Experiments on several benchmark datasets show that our method achieves better results compared with the state-of-the-art methods. Peiwen Zhu, Guodong Guo |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2023 | LQGDNet: A Local Quaternion and Global Deep Network for Facial Depression RecognitionabstractRecent visual-based depression recognition methods mostly use hand-crafted features with information lost in color channels, or deep network features with a limited performance from the finite data. In this paper, we propose a method called Local Quaternion and Global Deep Network (LQGDNet) which can combine advantages from hand-crafted and deep features. Specifically, the Quaternion XOR Asymmetrical Regional Local Gradient Coding (XOR-AR-LGC) is first designed, which encodes the facial images with local textures in the quaternion domain to keep the dependence of color channels, and integrated into the Quaternion Feature Extractor (QFE). To the best of our knowledge, it is the first attempt to use a quaternion-based method for facial depression recognition. Second, we design the Local Quaternion Representation Module (LQRM) composed of Local Deep Feature Extractor (LDFE) and QFE to output local quaternion facial features. Third, global deep facial features are encoded from the Global Deep Representation Module (GDRM) with the deep convolutional neural network. Finally, the LQGDNet integrates LQRM and GDRM with the local quaternion and global deep features and predicts the depression score. The experimental results on AVEC 2013 and AVEC 2014 show the superiority of our method compared to the state-of-the-art approaches. Zhuhong Shao, Guodong Guo |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | Vision Transformer With Attentive Pooling for Robust Facial Expression RecognitionabstractFacial Expression Recognition (FER) in the wild is an extremely challenging task. Recently, some Vision Transformers (ViT) have been explored for FER, but most of them perform inferiorly compared to Convolutional Neural Networks (CNN). This is mainly because the new proposed modules are difficult to converge well from scratch due to lacking inductive bias and easy to focus on the occlusion and noisy areas. TransFER, a representative transformer-based method for FER, alleviates this with multi-branch attention dropping but brings excessive computations. On the contrary, we present two attentive pooling (AP) modules to pool noisy features directly. The AP modules include Attentive Patch Pooling (APP) and Attentive Token Pooling (ATP). They aim to guide the model to emphasize the most discriminative features while reducing the impacts of less relevant features. The proposed APP is employed to select the most informative patches on CNN features, and ATP discards unimportant tokens in ViT. Being simple to implement and without learnable parameters, the APP and ATP intuitively reduce the computational cost while boosting the performance by ONLY pursuing the most discriminative features. Qualitative results demonstrate the motivations and effectiveness of our attentive poolings. Besides, quantitative results on six in-the-wild datasets outperform other state-of-the-art methods. Fanglei Xue, Qiangchang Wang, Zichang Tan, Zhongsong Ma, Guodong Guo |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | Divergence-Driven Consistency Training for Semi-Supervised Facial Age EstimationabstractFacial age estimation has attracted considerable attention owing to its great potential in applications. However, it still falls short of reliable age estimation due to the lack of sufficient training data with accurate age labels. Using conventional semi-supervised methods to exploit unlabeled data appears to be a good solution, but it does not yield sufficient performance gains while significantly increasing training time. Therefore, to tackle these problems, we present a Divergence-driven Consistency Training (DCT) method for enhancing both efficiency and performance in this paper. Following the idea of pseudo-labeling and consistency regularization, we assign pseudo labels predicted by the teacher model to unlabeled samples and then train the student model on labeled and unlabeled samples based on consistency regularization. Based on this, we propose two main promotions. The first is the Efficient Sample Selection (ESS) strategy, which is based on the Divergence Score to select effective samples from massive unlabeled images to reduce the training time and improve efficiency. The second is Identity Consistency (IC) regularization as the additional loss function, which introduces a high dependency of aging traits on a person. Moreover, we propose Local Prediction (LP), which is a plug-and-play component, to capture local semantics. Extensive experiments on multiple age benchmark datasets, including CACD, Morph II, MIVIA, and Chalearn LAP 2015, indicate DCT outperforms the state-of-the-art approaches significantly. Zenghao Bao, Zichang Tan, Jun Wan 0001, Xibo Ma, Guodong Guo, Zhen Lei 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | FM-ViT: Flexible Modal Vision Transformers for Face Anti-SpoofingabstractThe availability of handy multi-modal (i.e., RGB-D) sensors has brought about a surge of face anti-spoofing research. However, the current multi-modal face presentation attack detection (PAD) has two defects: (1) The framework based on multi-modal fusion requires providing modalities consistent with the training input, which seriously limits the deployment scenario. (2) The performance of ConvNet-based model on high fidelity datasets is increasingly limited. In this work, we present a pure transformer-based framework, dubbed the Flexible Modal Vision Transformer (FM-ViT), for face anti-spoofing to flexibly target any single-modal (i.e., RGB) attack scenarios with the help of available multi-modal data. Specifically, FM-ViT retains a specific branch for each modality to capture different modal information and introduces the Cross-Modal Transformer Block (CMTB), which consists of two cascaded attentions named Multi-headed Mutual-Attention (MMA) and Fusion-Attention (MFA) to guide each modal branch to mine potential features from informative patch tokens, and to learn modality-agnostic liveness features by enriching the modal information of own CLS token, respectively. Experiments demonstrate that the single model trained based on FM-ViT can not only flexibly evaluate different modal samples, but also outperforms existing single-modal frameworks by a large margin, and approaches the multi-modal frameworks introduced with smaller FLOPs and model parameters. Ajian Liu 0001, Zichang Tan, Zitong Yu, Jun Wan 0001, Yanyan Liang 0001, Zhen Lei 0001, Stan Z. Li, Guodong Guo |
IEEE Trans. Inf. Forensics Secur. | 10 |
| 2023 | Anti-UAV: A Large-Scale Benchmark for Vision-Based UAV TrackingabstractUnmanned Aerial Vehicles (UAV) have many applications in both commerce and recreation. However, irresponsibly operated UAVs will pose a threat to public safety. Therefore, developing our understanding of UAVs and their uses is of particular interest. This paper considers tracking UAVs, which provide multifaceted information around location, paths and trajectories. To facilitate research on this topic, we introduce a new benchmark, herein referred to as Anti-UAV, which provides a novel direction for UAV tracking with more than 300 video pairs containing over 580 k manually annotated bounding boxes. Addressing anti-UAV research challenges could help to design anti-UAV systems, which in turn may improve surveillance. Accordingly, we have proposed a simple yet effective approach, called dual-flow semantic consistency (DFSC) is proposed for UAV tracking. Modulated by the semantic flow across video sequences, tracker learns more robust class-level semantic information and obtains more discriminative instance-level features. Experiments highlight significant performance gain with the proposed approach over state-of-the-art trackers and the challenging aspects of Anti-UAV. The Anti-UAV benchmark and the code for the proposed approach have been made publicly available athttps://github.com/ucas-vg/Anti-UAVandhttps://github.com/ZhaoJ9014/Anti-UAV. Kuiran Wang, Xiaoke Peng, Xuehui Yu, Qiang Wang 0051, Junliang Xing, Guorong Li, Guodong Guo, Qixiang Ye, Jianbin Jiao, Jian Zhao 0006, Zhenjun Han |
IEEE Trans. Multim. | 8 |
| 2023 | TCSD: Triple Complementary Streams Detector for Comprehensive Deepfake DetectionabstractAdvancements in computer vision and deep learning have made it difficult to distinguish deepfake visual media. While existing detection frameworks have achieved significant performance on challenging deepfake datasets, these approaches consider only a single perspective. More importantly, in urban scenes, neither complex scenarios can be covered by a single view nor can the correlation between multiple datasets of information be well utilized. In this article, to mine the new view for deepfake detection and utilize the correlation of multi-view information contained in images, we propose a novel triple complementary streams detector (TCSD). First, a novel depth estimator is designed to extract depth information (DI), which has not been used in previous methods. Then, to supplement depth information for obtaining comprehensive forgery clues, we consider the incoherence between image foreground and background information (FBI) and the inconsistency between local and global information (LGI). In addition, we designed an attention-based multi-scale feature extraction (MsFE) module to extract more complementary features from DI, FBI, and LGI. Finally, two attention-based feature fusion modules are proposed to adaptively fuse information. Extensive experiment results show that the proposed approach achieves state-of-the-art performance on detecting deepfakes. Yang Yu 0039, Xiaolong Li 0001, Yao Zhao 0001, Guodong Guo |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | A Novel Reversible Data Hiding Scheme Based on Pixel-Residual HistogramabstractPrediction-error expansion (PEE) is the most popular reversible data hiding (RDH) technique due to its efficient capacity-distortion tradeoff. With the generated prediction-error histogram (PEH) and adaptively selected expansion bins, the image redundancy is well exploited by PEE. However, for the most widely used rhombus predictor, the rounding operation which groups different prediction-errors into one value is completely unnecessary. The embedding can be extended to a general case by removing the rounding operation, and more histogram bins can be derived for expansion with a new mapping mechanism. Therefore, in this article, instead of pixel prediction-error, we propose to compute the pixel residuals without the rounding operation, and a new embedding mechanism based on pixel-residual histogram (PRH) modification is devised. In PRH, four bins correspond to one bin in PEH. Then, different from the one-to-one mapping between the prediction-error and pixel modification, a four-to-one mapping between the pixel-residual and pixel modification is established, and the performance is optimized by adaptively selecting four expansion bin pairs for embedding. Since more modification selections are considered, better performance can be obtained. Moreover, the proposed scheme is extended to the two-dimensional (2D) histogram and multiple histograms based embedding, and the performance is further enhanced. The superiority of the proposed method is experimentally verified by comparing it with some state-of-the-art works. Mengyao Xiao, Xiaolong Li 0001, Yao Zhao 0001, Bin Ma 0003, Guodong Guo |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | CQA-Face: Contrastive Quality-Aware Attentions for Face RecognitionabstractFew existing face recognition (FR) models take local representations into account. Although some works achieved this by extracting features on cropped parts around face landmarks, landmark detection may be inaccurate or even fail in some extreme cases. Recently, without relying on landmarks, attention-based networks can focus on useful parts automatically. However, there are two issues: 1) It is noticed that these approaches focus on few facial parts, while missing other potentially discriminative regions. This can cause performance drops when emphasized facial parts are invisible under heavy occlusions (e.g. face masks) or large pose variations; 2) Different facial parts may appear at various quality caused by occlusion, blur, or illumination changes. In this paper, we propose contrastive quality-aware attentions, called CQA-Face, to address these two issues. First, a Contrastive Attention Learning (CAL) module is proposed, pushing models to explore comprehensive facial parts. Consequently, more useful parts can help identification if some facial parts are invisible. Second, a Quality-Aware Network (QAN) is developed to emphasize important regions and suppress noisy parts in a global scope. Thus, our CQA-Face model is developed by integrating the CAL with QAN, which extracts diverse quality-aware local representations. It outperforms the state-of-the-art methods on several benchmarks, demonstrating its effectiveness and usefulness. Qiangchang Wang, Guodong Guo |
AAAI | 2 |
| 2022 | Pale Transformer: A General Vision Transformer Backbone with Pale-Shaped AttentionabstractRecently, Transformers have shown promising performance in various vision tasks. To reduce the quadratic computation complexity caused by the global self-attention, various methods constrain the range of attention within a local region to improve its efficiency. Consequently, their receptive fields in a single attention layer are not large enough, resulting in insufficient context modeling. To address this issue, we propose a Pale-Shaped self-Attention (PS-Attention), which performs self-attention within a pale-shaped region. Compared to the global self-attention, PS-Attention can reduce the computation and memory costs significantly. Meanwhile, it can capture richer contextual information under the similar computation complexity with previous local self-attention mechanisms. Based on the PS-Attention, we develop a general Vision Transformer backbone with a hierarchical architecture, named Pale Transformer, which achieves 83.4%, 84.3%, and 84.9% Top-1 accuracy with the model size of 22M, 48M, and 85M respectively for 224x224 ImageNet-1K classification, outperforming the previous Vision Transformer backbones. For downstream tasks, our Pale Transformer backbone performs better than the recent state-of-the-art CSWin Transformer by a large margin on ADE20K semantic segmentation and COCO object detection & instance segmentation. The code will be released on https://github.com/BR-IDL/PaddleViT. Sitong Wu, Haoru Tan, Guodong Guo |
AAAI | 4 |
| 2022 | QS-Craft: Learning to Quantize, Scrabble and Craft for Conditional Human Motion Animation
Yuxin Hong, Xuelin Qian, Simian Luo, Guodong Guo, Xiangyang Xue 0001, Yanwei Fu 0001 |
ACCV (6) | 4 |
| 2022 | Bi-level Doubly Variational Learning for Energy-based Latent Variable ModelsabstractEnergy-based latent variable models (EBLVMs) are more expressive than conventional energy-based models. However, its potential on visual tasks are limited by its training process based on maximum likelihood estimate that requires sampling from two intractable distributions. In this paper, we propose Bi-level doubly variational learning (BiDVL), which is based on a new bi-level optimization framework and two tractable variational distributions to facilitate learning EBLVMs. Particularly, we lead a decoupled EBLVM consisting of a marginal energy-based distribution and a structural posterior to handle the difficulties when learning deep EBLVMs on images. By choosing a symmetric KL divergence in the lower level of our framework, a compact BiDVL for visual tasks can be obtained. Our model achieves impressive image generation performance over related works. It also demonstrates the significant capacity of testing image reconstruction and out-of-distribution detection. Ge Kan, Jinhu Lü 0001, Tian Wang 0002, Baochang Zhang 0001, Aichun Zhu, Lei Huang 0015, Guodong Guo, Hichem Snoussi |
CVPR | 7 |
| 2022 | Nested Collaborative Learning for Long-Tailed Visual RecognitionabstractThe networks trained on the long-tailed dataset vary remarkably, despite the same training settings, which shows the great uncertainty in long-tailed learning. To alleviate the uncertainty, we propose a Nested Collaborative Learning (NCL), which tackles the problem by collaboratively learning multiple experts together. NCL consists of two core components, namely Nested Individual Learning (NIL) and Nested Balanced Online Distillation (NBOD), which focus on the individual supervised learning for each single expert and the knowledge transferring among multiple experts, respectively. To learn representations more thoroughly, both NIL and NBOD are formulated in a nested way, in which the learning is conducted on not just all categories from a full perspective but some hard categories from a partial perspective. Regarding the learning in the partial perspective, we specifically select the negative categories with high predicted scores as the hard categories by using a proposed Hard Category Mining (HCM). In the NCL, the learning from two perspectives is nested, highly related and complementary, and helps the network to capture not only global and robust features but also meticulous distinguishing ability. Moreover, self-supervision is further utilized for feature enhancement. Extensive experiments manifest the superiority of our method with outperforming the state-of-the-art whether by using a single model or an ensemble. Code is available at https://github.com/Bazinga699/NCL Jun Li 0033, Zichang Tan, Jun Wan 0001, Zhen Lei 0001, Guodong Guo |
CVPR | 5 |
| 2022 | SAR-Net: Shape Alignment and Recovery Network for Category-level 6D Object Pose and Size EstimationabstractGiven a single scene image, this paper proposes a method of Category-level 6D Object Pose and Size Estimation (COPSE) from the point cloud of the target object, without external real pose-annotated training data. Specifically, beyond the visual cues in RGB images, we rely on the shape information predominately from the depth (D) channel. The key idea is to explore the shape alignment of each instance against its corresponding category-level template shape, and the symmetric correspondence of each object category for estimating a coarse 3D object shape. Our framework deforms the point cloud of the category-level template shape to align the observed instance point cloud for implicitly representing its 3D rotation. Then we model the symmetric correspondence by predicting symmetric point cloud from the partially observed point cloud. The concatenation of the observed point cloud and symmetric one reconstructs a coarse object shape, thus facilitating object center (3D translation) and 3D size estimation. Extensive experiments on the category-level NOCS benchmark demonstrate that our lightweight model still competes with state-of-the-art approaches that require labeled real-world images. We also deploy our approach to a physical Baxter robot to perform grasping tasks on unseen but category-known instances, and the results further validate the efficacy of our proposed model. Code and pre-trained models are available on the project webpage11Project webpage. https://hetolin.github.io/SAR-Net. Zichang Liu, Chilam Cheang, Yanwei Fu 0001, Guodong Guo, Xiangyang Xue 0001 |
CVPR | 5 |
| 2022 | End-to-End Human-Gaze-Target Detection with TransformersabstractIn this paper, we propose an effective and efficient method for Human-Gaze-Target (HGT) detection, i.e., gaze following. Current approaches decouple the HGT detection task into separate branches of salient object detection and human gaze prediction, employing a two-stage framework where human head locations must first be detected and then be fed into the next gaze target prediction sub-network. In contrast, we redefine the HGT detection task as detecting human head locations and their gaze targets, simultaneously. By this way, our method, named Human-Gaze-Target detection TRansformer or HGTTR, streamlines the HGT detection pipeline by eliminating all other additional components. HGTTR reasons about the relations of salient objects and human gaze from the global image context. Moreover, unlike existing two-stage methods that require human head locations as input and can predict only one human's gaze target at a time, HGTTR can directly predict the locations of all people and their gaze targets at one time in an end-to-end manner. The effectiveness and robustness of our proposed method are verified with extensive experiments on the two standard benchmark datasets, GazeFollowing and VideoAttentionTarget. Without bells and whistles, HGTTR outperforms existing state-of-the-art methods by large margins (6.4 mAP gain on GazeFollowing and 10.3 mAP gain on VideoAttentionTarget) with a much simpler architecture. Danyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo, Guangtao Zhai, Wei Shen 0002 |
CVPR | 4 |
| 2022 | Iwin: Human-Object Interaction Detection via Transformer with Irregular Windows
Danyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo, Guangtao Zhai, Wei Shen 0002 |
ECCV (4) | 4 |
| 2022 | Anti-retroactive Interference for Lifelong Learning
Runqi Wang, Yuxiang Bao, Baochang Zhang 0001, Jianzhuang Liu, Wentao Zhu 0001, Guodong Guo |
ECCV (24) | 6 |
| 2022 | Recurrent Bilinear Optimization for Binary Neural Networks
Sheng Xu 0007, Yanjing Li, Teli Ma, Baochang Zhang 0001, Peng Gao 0007, Yu Qiao 0001, Jinhu Lü 0001, Guodong Guo |
ECCV (24) | 9 |
| 2022 | Dynamic Group Transformer: A General Vision Transformer Backbone with Dynamic Group AttentionabstractRecently, Transformers have shown promising performance in various vision tasks. To reduce the quadratic computation complexity caused by each query attending to all keys/values, various methods have constrained the range of attention within local regions, where each query only attends to keys/values within a hand-crafted window. However, these hand-crafted window partition mechanisms are data-agnostic and ignore their input content, so it is likely that one query maybe attend to irrelevant keys/values. To address this issue, we propose a Dynamic Group Attention (DG-Attention), which dynamically divides all queries into multiple groups and selects the most relevant keys/values for each group. Our DG-Attention can flexibly model more relevant dependencies without any spatial constraint that is used in hand-crafted window based attention. Built on the DG-Attention, we develop a general vision transformer backbone named Dynamic Group Transformer (DGT). Extensive experiments show that our models can outperform the state-of-the-art methods on multiple common vision tasks, including image classification, semantic segmentation, object detection, and instance segmentation. Guodong Guo |
IJCAI | 4 |
| 2022 | Region-level Contrastive and Consistency Learning for Semi-Supervised Semantic SegmentationabstractCurrent semi-supervised semantic segmentation methods mainly focus on designing pixel-level consistency and contrastive regularization. However, pixel-level regularization is sensitive to noise from pixels with incorrect predictions, and pixel-level contrastive regularization has a large memory and computational cost. To address the issues, we propose a novel region-level contrastive and consistency learning framework (RC^2L) for semi-supervised semantic segmentation. Specifically, we first propose a Region Mask Contrastive (RMC) loss and a Region Feature Contrastive (RFC) loss to accomplish region-level contrastive property. Furthermore, Region Class Consistency (RCC) loss and Semantic Mask Consistency (SMC) loss are proposed for achieving region-level consistency. Based on the proposed region-level contrastive and consistency regularization, we develop a region-level contrastive and consistency learning framework (RC^2L) for semi-supervised semantic segmentation, and evaluate our RC^2L on two challenging benchmarks (PASCAL VOC 2012 and Cityscapes), outperforming the state-of-the-art. Jianrong Zhang, Chuanghao Ding, Guodong Guo |
IJCAI | 5 |
| 2022 | CATrans: Context and Affinity Transformer for Few-Shot SegmentationabstractFew-shot segmentation (FSS) aims to segment novel categories given scarce annotated support images. The crux of FSS is how to aggregate dense correlations between support and query images for query segmentation while being robust to the large variations in appearance and context. To this end, previous Transformer-based methods explore global consensus either on context similarity or affinity map between support-query pairs. In this work, we effectively integrate the context and affinity information via the proposed novel Context and Affinity Transformer (CATrans) in a hierarchical architecture. Specifically, the Relation-guided Context Transformer (RCT) propagates context information from support to query images conditioned on more informative support features. Based on the observation that a huge feature distinction between support and query pairs brings barriers for context knowledge transfer, the Relation-guided Affinity Transformer (RAT) measures attention-aware affinity as auxiliary information for FSS, in which the self-affinity is responsible for more reliable cross-affinity. We conduct experiments to demonstrate the effectiveness of the proposed model, outperforming the state-of-the-art methods. Sitong Wu, Guodong Guo |
IJCAI | 4 |
| 2022 | One-step Low-Rank Representation for ClusteringabstractExisting low-rank representation-based methods adopt a two-step framework, which must employ an extra clustering method to gain labels after representation learning. In this paper, a novel one-step representation-based method, i.e., One-step Low-Rank Representation (OLRR), is proposed to capture multi-subspace structures for clustering. OLRR integrates the low-rank representation model and clustering into a unified framework. Thus it can jointly learn the low-rank subspace structure embedded in the database and gain the clustering results. In particular, by approximating the representation matrix with two same clustering indicator matrices, OLRR can directly show the probability of samples belonging to each cluster. Further, a probability penalty is introduced to ensure that the samples with smaller distances are more inclined to be in the same cluster, thus enhancing the discrimination of the clustering indicator matrix and resulting in a more favorable clustering performance. Moreover, to enhance the robustness against noise, OLRR uses the probability to guide denoising and then performs representation learning and clustering in a recovered clean space. Extensive experiments well demonstrate the robustness and effectiveness of OLRR. Our code is publicly available at: https://github.com/fuzhiqiang1230/OLRR. Zhiqiang Fu, Yao Zhao 0001, Dongxia Chang, Yiming Wang 0007, Jie Wen 0001, Xingxing Zhang 0001, Guodong Guo |
ACM Multimedia | 7 |
| 2022 | Q-ViT: Accurate and Fully Quantized Low-bit Vision TransformerabstractThe large pre-trained vision transformers (ViTs) have demonstrated remarkable performance on various visual tasks, but suffer from expensive computational and memory cost problems when deployed on resource-constrained devices. Among the powerful compression approaches, quantization extremely reduces the computation and memory consumption by low-bit parameters and bit-wise operations. However, low-bit ViTs remain largely unexplored and usually suffer from a significant performance drop compared with the real-valued counterparts. In this work, through extensive empirical analysis, we first identify the bottleneck for severe performance drop comes from the information distortion of the low-bit quantized self-attention map. We then develop an information rectification module (IRM) and a distribution guided distillation (DGD) scheme for fully quantized vision transformers (Q-ViT) to effectively eliminate such distortion, leading to a fully quantized ViTs. We evaluate our methods on popular DeiT and Swin backbones. Extensive experimental results show that our method achieves a much better performance than the prior arts. For example, our Q-ViT can theoretically accelerates the ViT-S by 6.14x and achieves about 80.9% Top-1 accuracy, even surpassing the full-precision counterpart by 1.0% on ImageNet dataset. Our codes and models are attached on https://github.com/YanjingLi0202/Q-ViT Yanjing Li, Sheng Xu 0007, Baochang Zhang 0001, Xianbin Cao 0001, Peng Gao 0007, Guodong Guo |
NeurIPS | 6 |
| 2022 | SKFlow: Learning Optical Flow with Super KernelsabstractOptical flow estimation is a classical yet challenging task in computer vision. One of the essential factors in accurately predicting optical flow is to alleviate occlusions between frames. However, it is still a thorny problem for current top-performing optical flow estimation methods due to insufficient local evidence to model occluded areas. In this paper, we propose the Super Kernel Flow Network (SKFlow), a CNN architecture to ameliorate the impacts of occlusions on optical flow estimation. SKFlow benefits from the super kernels which bring enlarged receptive fields to complement the absent matching information and recover the occluded motions. We present efficient super kernel designs by utilizing conical connections and hybrid depth-wise convolutions. Extensive experiments demonstrate the effectiveness of SKFlow on multiple benchmarks, especially in the occluded areas. Without pre-trained backbones on ImageNet and with a modest increase in computation, SKFlow achieves compelling performance and ranks $\textbf{1st}$ among currently published methods on the Sintel benchmark. On the challenging Sintel clean and final passes (test), SKFlow surpasses the best-published result in the unmatched areas ($7.96$ and $12.50$) by $9.09\%$ and $7.92\%$. The code is available at https://github.com/littlespray/SKFlow. Shangkun Sun, Yuanqi Chen, Yu Zhu 0006, Guodong Guo |
NeurIPS | 4 |
| 2022 | Perceptual Attacks of No-Reference Image Quality Models with Human-in-the-LoopabstractNo-reference image quality assessment (NR-IQA) aims to quantify how humans perceive visual distortions of digital images without access to their undistorted references. NR-IQA models are extensively studied in computational vision, and are widely used for performance evaluation and perceptual optimization of man-made vision systems. Here we make one of the first attempts to examine the perceptual robustness of NR-IQA models. Under a Lagrangian formulation, we identify insightful connections of the proposed perceptual attack to previous beautiful ideas in computer vision and machine learning. We test one knowledge-driven and three data-driven NR-IQA methods under four full-reference IQA models (as approximations to human perception of just-noticeable differences). Through carefully designed psychophysical experiments, we find that all four NR-IQA models are vulnerable to the proposed perceptual attack. More interestingly, we observe that the generated counterexamples are not transferable, manifesting themselves as distinct design flows of respective NR-IQA methods. Source code are available at https://github.com/zwx8981/PerceptualAttack_BIQA. Weixia Zhang, Dingquan Li, Xiongkuo Min, Guangtao Zhai, Guodong Guo, Xiaokang Yang 0001, Kede Ma |
NeurIPS | 5 |
| 2022 | Scene text detection by adaptive feature selection with text scale-aware loss
Wenli Luo, ZhiLei Chai, Guodong Guo |
Appl. Intell. | 4 |
| 2022 | EAN: Event Adaptive Network for Enhanced Action Recognition
Yuan Tian 0017, Yichao Yan, Guangtao Zhai, Guodong Guo |
Int. J. Comput. Vis. | 4 |
| 2022 | Towards Compact 1-bit CNNs via Bayesian Learning
Junhe Zhao, Sheng Xu 0007, Baochang Zhang 0001, Jiaxin Gu, David S. Doermann, Guodong Guo |
Int. J. Comput. Vis. | 6 |
| 2022 | Multi-scale feature aggregation and boundary awareness network for salient object detection
Jianzhe Wang, ZhiLei Chai, Guodong Guo |
Image Vis. Comput. | 4 |
| 2022 | Data-adaptive binary neural networks for efficient object detection and recognition
Junhe Zhao, Sheng Xu 0007, Runqi Wang, Baochang Zhang 0001, Guodong Guo, David S. Doermann, Dianmin Sun |
Pattern Recognit. Lett. | 5 |
| 2022 | Transformer-Based Feature Compensation and Aggregation for DeepFake DetectionabstractDeepfake detection has attracted increasing attention in recent years. In this paper, we propose a transformer-based framework with feature compensation and aggregation (Trans-FCA) to extract rich forgery cues for deepfake detection. To compensate local features for transformers, we propose a Locality Compensation Block (LCB) containing a Global-Local Cross-Attention (GLCA) to attentively fuse global transformer features and local convolutional features. To aggregate features of all layers for capturing comprehensive and various fake flaws, we propose Multi-head Clustering Projection (MCP) and Frequency-guided Fusion Module (FFM), where the MCP attentively reduces redundant features into a few concentrated clusters, and the FFM interacts all clustered features under the guidance of frequency cues. In Trans-FCA, besides global cues captured by transformer architecture, local details and rich forgery defects are also captured using the proposed fetaure compensation and aggregation. Extensive experiments show our method outperforms the state-of-the-art methods on both intra-dataset and cross-dataset testings (with AUCs of 99.85% on FaceForensics++ and 78.57% on Celeb-DF), which clearly demonstrates the superiority of our Trans-FCA for deepfake detection. Zichang Tan, Zhichao Yang 0008, Changtao Miao, Guodong Guo |
IEEE Signal Process. Lett. | 4 |
| 2022 | Fine-Grained Image Classification With Global Information and Adaptive Compensation LossabstractFine-grained image classification differs from traditional image classification in that the former needs to divide subclasses under a basic level of categories. Previous works always focus on how to locate discriminative parts of objects, but we find that the global and background information of objects neglected by them is also valuable in some situations. This letter proposes a method to combine the global information and discriminative parts information of objects to do classification, which includes three modules: (1) Activation map based crop-erase module localizes objects while avoiding localization bias due to excessive bias of the network to learn one discriminative part. (2) Part attention module helps learning discriminative part features of objects. (3) Two-level fusion module gives consideration to the global and local information of objects and some potentially effective background information. Meanwhile, we propose an adaptive compensation loss to distinguish easily confused categories. Experiments show that our method achieves state-of-the-art performance on three open benchmarks. Shuting Miao, ZhiLei Chai, Guodong Guo |
IEEE Signal Process. Lett. | 4 |
| 2022 | Facial Depression Recognition by Deep Joint Label Distribution and Metric LearningabstractWhile existing prediction models built on popular deep architectures have shown promising results in facial depression recognition, they still lack sufficient discriminative power due to the issues of 1) limited amount of labeled depression data for deep representation learning and, 2) large variation in facial expression across different persons of the same depression score and the subtle difference in facial expression across different depression levels. In this article, we formulate the facial depression recognition as a label distribution learning (LDL) problem, and propose a deep joint label distribution and metric learning (DJ-LDML) method to address these issues. In DJ-LDML, LDL exploits label relevance inherent in depression data to implicitly increase the amount of training data associated with each depression level without actually enlarging the dataset, while deep metric learning (DML) aims at learning a deep ordinal embedding with a specifically designed label-aware histogram loss, allowing semantics similarity between video sequences (described by ordinal labels) to be preserved for discriminative feature learning. The two learning modules in our DJ-LDML work collaboratively to enhance the representation ability and discriminative power of the deeply learned spatiotemporal feature, leading to improved depression prediction. We empirically evaluate our method on two benchmark datasets and the results demonstrate the effectiveness of our formulation. Xiuzhuang Zhou, Zeqiang Wei, Min Xu 0003, Shan Qu, Guodong Guo |
IEEE Trans. Affect. Comput. | 5 |
| 2022 | Joint Face Image Restoration and Frontalization for RecognitionabstractIn real-world scenarios, many factors may harm face recognition performance,e.g., large pose, bad illumination, low resolution, blur and noise. To address these challenges, previous efforts usually first restore the low-quality faces to high-quality ones and then perform face recognition. However, most of these methods are stage-wise, which is sub-optimal and deviates from the reality. In this paper, we address all these challenges jointly for unconstrained face recognition. We propose anMulti-DegradationFaceRestoration (MDFR) model to restore frontalized high-quality faces from the given low-quality ones under arbitrary facial poses, with three distinct novelties. First, MDFR is a well-designed encoder-decoder architecture which extracts feature representation from an input face image with arbitrary low-quality factors and restores it to a high-quality counterpart. Second, MDFR introduces a pose residual learning strategy along with a 3D-basedPoseNormalizationModule (PNM), which can perceive the pose gap between the input initial pose and its real-frontal pose to guide the face frontalization. Finally, MDFR can generate frontalized high-quality face images by a single unified network, showing a strong capability of preserving face identity. Qualitative and quantitative experiments on both controlled and in-the-wild benchmarks demonstrate the superiority of MDFR over state-of-the-art methods on both face frontalization and face restoration. Xiaoguang Tu, Jian Zhao 0006, Wenjie Ai, Guodong Guo, Zhifeng Li 0001, Wei Liu 0005, Jiashi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Image-to-Video Generation via 3D Facial DynamicsabstractWe present a versatile model, FaceAnime, for various video generation tasks from still images. Video generation from a single face image is an interesting problem and usually tackled by utilizing Generative Adversarial Networks (GANs) to integrate information from the input face image and a sequence of sparse facial landmarks. However, the generated face images usually suffer from quality loss, image distortion, identity change, and expression mismatching due to the weak representation capacity of the facial landmarks. In this paper, we propose to “imagine” a face video from a single face image according to the reconstructed 3D face dynamics, aiming to generate a realistic and identity-preserving face video, with precisely predicted pose and facial expression. The 3D dynamics reveal changes of the facial expression and motion, and can serve as a strong prior knowledge for guiding highly realistic face video generation. In particular, we explore face video prediction and exploit a well-designed 3D dynamic prediction network to predict a 3D dynamic sequence for a single face image. The 3D dynamics are then further rendered by the sparse texture mapping algorithm to recover structural details and sparse textures for generating face frames. Our model is versatile for various AR/VR and entertainment applications, such as face video retargeting and face video prediction. Superior experimental results have well demonstrated its effectiveness in generating high-fidelity, identity-preserving, and visually pleasant face video clips from a single source face image. Xiaoguang Tu, Yingtian Zou, Jian Zhao 0006, Wenjie Ai, Jian Dong 0011, Yuan Yao 0011, Zhikang Wang, Guodong Guo, Zhifeng Li 0001, Wei Liu 0005, Jiashi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2022 | ChaLearn Looking at People: IsoGD and ConGD Large-Scale RGB-D Gesture RecognitionabstractThe ChaLearn large-scale gesture recognition challenge has run twice in two workshops in conjunction with the International Conference on Pattern Recognition (ICPR) 2016 and International Conference on Computer Vision (ICCV) 2017, attracting more than 200 teams around the world. This challenge has two tracks, focusing on isolated and continuous gesture recognition, respectively. It describes the creation of both benchmark datasets and analyzes the advances in large-scale gesture recognition based on these two datasets. In this article, we discuss the challenges of collecting large-scale ground-truth annotations of gesture recognition and provide a detailed analysis of the current methods for large-scale isolated and continuous gesture recognition. In addition to the recognition rate and mean Jaccard index (MJI) as evaluation metrics used in previous challenges, we introduce the corrected segmentation rate (CSR) metric to evaluate the performance of temporal segmentation for continuous gesture recognition. Furthermore, we propose a bidirectional long short-term memory (Bi-LSTM) method, determining video division points based on skeleton points. Experiments show that the proposed Bi-LSTM outperforms state-of-the-art methods with an absolute improvement of 8.1% (from 0.8917 to 0.9639) of CSR. Jun Wan 0001, Chi Lin 0002, Longyin Wen, Yunan Li 0001, Qiguang Miao, Sergio Escalera, Gholamreza Anbarjafari, Isabelle Guyon, Guodong Guo, Stan Z. Li |
IEEE Trans. Cybern. | 9 |
| 2022 | Contrastive Context-Aware Learning for 3D High-Fidelity Mask Face Presentation Attack DetectionabstractFace presentation attack detection (PAD) is essential to secure face recognition systems primarily from high-fidelity mask attacks. Most existing 3D mask PAD benchmarks suffer from several drawbacks: 1) a limited number of mask identities, types of sensors, and a total number of videos; 2) low-fidelity quality of facial masks. Basic deep models and remote photoplethysmography (rPPG) methods achieved acceptable performance on these benchmarks but still far from the needs of practical scenarios. To bridge the gap to real-world applications, we introduce a large-scale High-Fidelity Mask dataset, namely HiFiMask. Specifically, a total amount of 54,600 videos are recorded from 75 subjects with 225 realistic masks by 7 new kinds of sensors. Along with the dataset, we propose a novel Contrastive Context-aware Learning (CCL) framework. CCL is a new training methodology for supervised PAD tasks, which is able to learn by leveraging rich contexts accurately (e.g., subjects, mask material and lighting) among pairs of live faces and high-fidelity mask attacks. Extensive experimental evaluations on HiFiMask and three additional 3D mask datasets demonstrate the effectiveness of our method. The codes and dataset will be released soon. Ajian Liu 0001, Zitong Yu, Jun Wan 0001, Anyang Su, Zichang Tan, Sergio Escalera, Junliang Xing, Yanyan Liang 0001, Guodong Guo, Zhen Lei 0001, Stan Z. Li |
IEEE Trans. Inf. Forensics Secur. | 11 |
| 2022 | Hierarchical Frequency-Assisted Interactive Networks for Face Manipulation DetectionabstractRecently, face manipulation techniques have caused increasing trust concerns in our society. Although current face manipulation detection methods achieve impressive performance regarding intra-dataset evaluation, they are struggling to improve the generalization and robustness ability. To address this issue, we propose a novel Hierarchical Frequency-assisted Interactive Networks (HFI-Net) to explore comprehensive frequency-related forgery cues for face manipulation detection. At first, we formulate HFI-Net as a dual-branch network to take full advantage of both CNN and transformer for capturing local details and global context information, respectively. Considering the forged faces are easy to show flaws in the frequency domain, a novel Frequency-based Feature Refinement (FFR) module is proposed to learn frequency-based attention from RGB features. FFR module emphasizes forgery cues and suppresses the pristine semantics information by keeping middle-high frequency features while discarding the low-frequency ones. Based on FFR, we further develop a co-sharing Global-Local Interaction (GLI) module to conduct frequency-assisted interactions while capturing complementarity among dual branches. Lastly, we further implement the GLI module in each stage of the network to effectively explore multi-level frequency artifacts. Extensive experiments are conducted on several popular benchmarks including FaceForensics++, Celeb-DF, DeepFake-TIMIT, DFDC, UADFV, and DeeperForensics-1.0, which shows that our model outperforms the state-of-the-art, especially in unseen datasets, manipulations, and perturbations evaluation. Changtao Miao, Zichang Tan, Qi Chu 0001, Nenghai Yu, Guodong Guo |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | Learning Multi-Granularity Temporal Characteristics for Face Anti-SpoofingabstractFace anti-spoofing (FAS) is essential for securing face recognition systems. Despite the decent performance, few existing works fully leverage temporal information. This would inevitably lead to inferior performance because real and fake faces tend to share highly similar spatial appearances, while important temporal features between consecutive frames are neglected. In this work, we propose a temporal transformer network (TTN) to learn multi-granularity temporal characteristics for FAS. It mainly consists of temporal difference attentions (TDA), a pyramid temporal aggregation (PTA), and a temporal depth difference loss (TDL). Firstly, the vision transformer (ViT) is used as the backbone where comprehensive local patches are utilized to provide subtle differences between live and spoof faces. Then, instead of learning temporal features on global faces which may miss some important local cues, the TDA is developed to extract motion-sensitive cues on each of the comprehensive local patches. Moreover, the TDA is inserted into different layers of the ViT, learning multi-scale motion-sensitive local cues to improve the FAS performance. Secondly, it is observed that different subjects may have different visual tempos in some actions, making it necessary to model different temporal speeds. Our PTA aggregates temporal features at various tempos, which could build short-range and long-range relations among multiple frames. Thirdly, depth maps for real parts may change continuously, while they remain zeros for spoof regions. In order to locate motion features on facial parts, the TDL is proposed to guide the network to locate spoof facial parts where motion patterns between neighboring frames are set as the ground truth. To the best of our knowledge, this work is the first attempt to learn temporal characteristics via transformers. Both qualitative and quantitative results on several challenging tasks demonstrate the usefulness and effectiveness of our proposed methods. Qiangchang Wang, Weihong Deng, Guodong Guo |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Cross-Batch Hard Example Mining With Pseudo Large Batch for ID vs. Spot Face RecognitionabstractIn our daily life, a large number of activities require identity verification, e.g., ePassport gates. Most of those verification systems recognize who you are by matching the ID document photo (ID face) to your live face image (spot face). The ID vs. Spot (IvS) face recognition is different from general face recognition where each dataset usually contains a small number of subjects and sufficient images for each subject. In IvS face recognition, the datasets usually contain massive class numbers (million or more) while each class only has two image samples (one ID face and one spot face), which makes it very challenging to train an effective model (e.g., excessive demand on GPU memory if conducting the classification on such massive classes, hardly capture the effective features for bisample data of each identity, etc.). To avoid the excessive demand on GPU memory, a two-stage training method is developed, where we first train the model on the dataset in general face recognition (e.g., MS-Celeb-1M) and then employ the metric learning losses (e.g., triplet and quadruplet losses) to learn the features on IvS data with million classes. To extract more effective features for IvS face recognition, we propose two novel algorithms to enhance the network by selecting harder samples for training. Firstly, a Cross-Batch Hard Example Mining (CB-HEM) is proposed to select the hard triplets from not only the current mini-batch but also past dozens of mini-batches (for convenience, we use batch to denote a mini-batch in the following), which can significantly expand the space of sample selection. Secondly, a Pseudo Large Batch (PLB) is proposed to virtually increase the batch size with a fixed GPU memory. The proposed PLB and CB-HEM can be employed simultaneously to train the network, which dramatically expands the selecting space by hundreds of times, where the very hard sample pairs especially the hard negative pairs can be selected for training to enhance the discriminative capability. Extensive comparative evaluations conducted on multiple IvS benchmarks demonstrate the effectiveness of the proposed method. Zichang Tan, Ajian Liu 0001, Jun Wan 0001, Hao Li 0030, Zhen Lei 0001, Guodong Guo, Stan Z. Li |
IEEE Trans. Image Process. | 6 |
| 2022 | Consensus Feature Network for Scene ParsingabstractScene parsing is challenging as it aims to assign one of the semantic categories to each pixel in scene images. Thus, pixel-level features are desired for scene parsing. However, classification networks are dominated by the discriminative portion, so directly applying classification networks to scene parsing will result in inconsistent parsing predictions within one instance and among instances of the same category. To address this problem, we propose two transform units to learn pixel-level consensus features. One is an Instance Consensus Transform (ICT) unit to learn the instance-level consensus features by aggregating features within the same instance. The other is a Category Consensus Transform (CCT) unit to pursue category-level consensus features through keeping the consensus of features among instances of the same category in scene images. The proposed ICT and CCT units are lightweight, data-driven and end-to-end trainable. The features learned by the two units are more coherent in both instance-level and category-level. Furthermore, we present the Consensus Feature Network (CFNet) based on the proposed ICT and CCT units, and demonstrate the effectiveness of each component in our method by performing extensive ablation experiments. Finally, our proposed CFNet achieves competitive performance on four datasets, including Cityscapes, Pascal Context, CamVid, and COCO Stuff. Sheng Tang, Rui Zhang 0040, Guodong Guo |
IEEE Trans. Multim. | 4 |
| 2022 | SMGEA: A New Ensemble Adversarial Attack Powered by Long-Term Gradient MemoriesabstractDeep neural networks are vulnerable to adversarial attacks. More importantly, some adversarial examples crafted against an ensemble of source models transfer to other target models and, thus, pose a security threat to black-box applications (when attackers have no access to the target models). Current transfer-based ensemble attacks, however, only consider a limited number of source models to craft an adversarial example and, thus, obtain poor transferability. Besides, recent query-based black-box attacks, which require numerous queries to the target model, not only come under suspicion by the target model but also cause expensive query cost. In this article, we propose a novel transfer-based black-box attack, dubbed serial-minigroup-ensemble-attack (SMGEA). Concretely, SMGEA first divides a large number of pretrained white-box source models into several "minigroups." For each minigroup, we design three new ensemble strategies to improve the intragroup transferability. Moreover, we propose a new algorithm that recursively accumulates the "long-term" gradient memories of the previous minigroup to the subsequent minigroup. This way, the learned adversarial information can be preserved, and the intergroup transferability can be improved. Experiments indicate that SMGEA not only achieves state-of-the-art black-box attack ability over several data sets but also deceives two online black-box saliency prediction systems in real world, i.e., DeepGaze-II (https://deepgaze.bethgelab.org/) and SALICON (http://salicon.net/demo/). Finally, we contribute a new code repository to promote research on adversarial attack and defense over ubiquitous pixel-to-pixel computer vision tasks. We share our code together with the pretrained substitute model zoo at https://github.com/CZHQuality/AAA-Pix2pix. Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Xiongkuo Min, Guodong Guo, Patrick Le Callet |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2022 | BiRe-ID: Binary Neural Network for Efficient Person Re-IDabstractPerson re-identification (Re-ID) has been promoted by the significant success of convolutional neural networks (CNNs). However, the application of such CNN-based Re-ID methods depends on the tremendous consumption of computation and memory resources, which affects its development on resource-limited devices such as next generation AI chips. As a result, CNN binarization has attracted increasing attention, which leads to binary neural networks (BNNs). In this article, we propose a new BNN-based framework for efficient person Re-ID (BiRe-ID). In this work, we discover that the significant performance drop of binarized models for Re-ID task is caused by the degraded representation capacity of kernels and features. To address the issues, we propose the kernel and feature refinement based on generative adversarial learning (KR-GAL and FR-GAL) to enhance the representation capacity of BNNs. We first introduce an adversarial attention mechanism to refine the binarized kernels based on their real-valued counterparts. Specifically, we introduce a scale factor to restore the scale of 1-bit convolution. And we employ an effective generative adversarial learning method to train the attention-aware scale factor. Furthermore, we introduce a self-supervised generative adversarial network to refine the low-level features using the corresponding high-level semantic information. Extensive experiments demonstrate that our BiRe-ID can be effectively implemented on various mainstream backbones for the Re-ID task. In terms of the performance, our BiRe-ID surpasses existing binarization methods by significant margins, at the level even comparable with the real-valued counterparts. For example, on Market-1501, BiRe-ID achieves 64.0% mAP on ResNet-18 backbone, with an impressive 12.51× speedup in theory and 11.75× storage saving. In particular, the KR-GAL and FR-GAL methods show strong generalization on multiple tasks such as Re-ID, image classification, object detection, and 3D point cloud processing. Sheng Xu 0007, Baochang Zhang 0001, Jinhu Lü 0001, Guodong Guo, David S. Doermann |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2021 | TRQ: Ternary Neural Networks With Residual QuantizationabstractTernary neural networks (TNNs) are potential for network acceleration by reducing the full-precision weights in network to ternary ones, e.g., {-1,0,1}. However, existing TNNs are mostly calculated based on rule-of-thumb quantization methods by simply thresholding operations, which causes a significant accuracy loss. In this paper, we introduce a stem-residual framework which provides new insight into Ternary quantization, termed Residual Quantization (TRQ), to achieve more powerful TNNs. Rather than directly thresholding operations, TRQ recursively performs quantization on full-precision weights for a refined reconstruction by combining the binarized stem and residual parts. With such a unique quantization process, TRQ endows the quantizer with high flexibility and precision. Our TRQ is generic, which can be easily extended to multiple bits through recursively encoded residual for a better recognition accuracy. Extensive experimental results demonstrate that the proposed method yields great recognition accuracy while being accelerated. Wenrui Ding, Chunlei Liu 0001, Baochang Zhang 0001, Guodong Guo |
AAAI | 5 |
| 2021 | POEM: 1-bit Point-wise Operations based on Expectation-Maximization for Efficient Point Cloud Processing
Sheng Xu 0007, Junhe Zhao, Yanjing Li, Baochang Zhang 0001, Guodong Guo |
BMVC | 5 |
| 2021 | LAE : Long-Tailed Age Estimation
Zenghao Bao, Zichang Tan, Yu Zhu 0006, Jun Wan 0001, Xibo Ma, Zhen Lei 0001, Guodong Guo |
CAIP (2) | 7 |
| 2021 | Depth-Conditioned Dynamic Message Propagation for Monocular 3D Object DetectionabstractThe objective of this paper is to learn context- and depth- aware feature representation to solve the problem of monocular 3D object detection. We make following contributions: (i) rather than appealing to the complicated pseudo-LiDAR based approach, we propose a depth-conditioned dynamic message propagation (DDMP) network to effectively integrate the multi-scale depth information with the image context; (ii) this is achieved by first adaptively sampling context-aware nodes in the image context and then dynamically predicting hybrid depth-dependent filter weights and affinity matrices for propagating information; (Hi) by augmenting a center-aware depth encoding (CDE) task, our method successfully alleviates the inaccurate depth prior; (iv) we thoroughly demonstrate the effectiveness of our proposed approach and show state-of-the-art results among the monocular-based approaches on the KITTI benchmark dataset. Particularly, we rank 1stin the highly competitive KITTI monocular 3D object detection track on the submission day (November 16th, 2020). Code and models are released at https: //github.com/fudan-zvg/DDMP Li Wang 0033, Liang Du 0004, Xiaoqing Ye, Yanwei Fu 0001, Guodong Guo, Xiangyang Xue 0001, Jianfeng Feng, Li Zhang 0040 |
CVPR | 5 |
| 2021 | Supervised Contrastive Learning for Facial Kinship RecognitionabstractVision-based kinship recognition aims to determine whether the face images have a kin relation. Compared to traditional solutions, the vision-based kinship recognition methods have the advantages of lower cost and being easy to implement. Therefore, such technique can be widely employed in lots of scenarios including missing children search and automatic management of family album. The Recognizing Families in the Wild (RFIW) Data Challenge provides a platform for evaluation of different kinship recognition approaches with ranked results. We propose a supervised contrastive learning approach to address three different kinship recognition tracks (i.e., kinship verification, tri-subject verification, and large-scale search-and-retrieval) announced in the RFIW 2021 with the 2021 FG. Our results on three tracks of 2021 RFIW challenge achieve the highest ranking, which demonstrate the superiority of the proposed solution. Ximiao Zhang, Min Xu 0003, Xiuzhuang Zhou, Guodong Guo |
FG | 4 |
| 2021 | Looking here or there? Gaze Following in 360-Degree ImagesabstractGaze following, i.e., detecting the gaze target of a human subject, in 2D images has become an active topic in computer vision. However, it usually suffers from the out of frame issue due to the limited field-of-view (FoV) of 2D images. In this paper, we introduce a novel task, gaze following in 360-degree images which provide an omnidirectional FoV and can alleviate the out of frame issue. We collect the first dataset, "GazeFollow360"1, for this task, containing around 10,000 360-degree images with complex gaze behaviors under various scenes. Existing 2D gaze following methods suffer from performance degradation in 360degree images since they may use the assumption that a gaze target is in the 2D gaze sight line. However, this assumption is no longer true for long-distance gaze behaviors in 360-degree images, due to the distortion brought by sphere-to-plane projection. To address this challenge, we propose a 3D sight line guided dual-pathway framework, to detect the gaze target within a local region (here) and from a distant region (there), parallelly. Specifically, the local region is obtained as a 2D cone-shaped field along the 2D projection of the sight line starting at the human subject’s head position, and the distant region is obtained by searching along the sight line in 3D sphere space. Finally, the location of the gaze target is determined by fusing the estimations from both the local region and the distant region. Experimental results show that our method achieves significant improvements over previous 2D gaze following methods on our GazeFollow360 dataset. Wei Shen 0002, Zhongpai Gao, Yucheng Zhu, Guangtao Zhai, Guodong Guo |
ICCV | 6 |
| 2021 | Self-Conditioned Probabilistic Learning of Video RescalingabstractBicubic downscaling is a prevalent technique used to reduce the video storage burden or to accelerate the downstream processing speed. However, the inverse upscaling step is non-trivial, and the downscaled video may also deteriorate the performance of downstream tasks. In this paper, we propose a self-conditioned probabilistic framework for video rescaling to learn the paired downscaling and upscaling procedures simultaneously. During the training, we decrease the entropy of the information lost in the downscaling by maximizing its probability conditioned on the strong spatial-temporal prior information within the downscaled video. After optimization, the downscaled video by our framework preserves more meaningful information, which is beneficial for both the upscaling step and the downstream tasks, e.g., video action recognition task. We further extend the framework to a lossy video compression system, in which a gradient estimator for non-differential industrial lossy codecs is proposed for the end-to-end training of the whole system. Extensive experimental results demonstrate the superiority of our approach on video rescaling, video compression, and efficient action recognition tasks. Yuan Tian 0017, Guo Lu, Xiongkuo Min, Zhaohui Che, Guangtao Zhai, Guodong Guo |
ICCV | 6 |
| 2021 | IDARTS: Interactive Differentiable Architecture SearchabstractDifferentiable Architecture Search (DARTS) improves the efficiency of architecture search by learning the architecture and network parameters end-to-end. However, the intrinsic relationship between the architecture’s parameters is neglected, leading to a sub-optimal optimization process. The reason lies in the fact that the gradient descent method used in DARTS ignores the coupling relationship of the parameters and therefore degrades the optimization. In this paper, we address this issue by formulating DARTS as a bi-linear optimization problem and introducing an Interactive Differentiable Architecture Search (IDARTS). We first develop a backtracking backpropagation process, which can decouple the relationships of different kinds of parameters and train them in the same framework. The backtracking method coordinates the training of different parameters that fully explore their interaction and optimize training. We present experiments on the CIFAR10 and ImageNet datasets that demonstrate the efficacy of the IDARTS approach by achieving a top-1 accuracy of 76.52% on ImageNet without additional search cost vs. 75.8% with the state-of-the-art PC-DARTS. Runqi Wang, Baochang Zhang 0001, Tian Wang 0002, Guodong Guo, David S. Doermann |
ICCV | 5 |
| 2021 | TransFER: Learning Relation-aware Facial Expression Representations with TransformersabstractFacial expression recognition (FER) has received increasing interest in computer vision. We propose the Trans-FER model which can learn rich relation-aware local representations. It mainly consists of three components: Multi-Attention Dropping (MAD), ViT-FER, and Multi-head Self-Attention Dropping (MSAD). First, local patches play an important role in distinguishing various expressions, however, few existing works can locate discriminative and diverse local patches. This can cause serious problems when some patches are invisible due to pose variations or viewpoint changes. To address this issue, the MAD is proposed to randomly drop an attention map. Consequently, models are pushed to explore diverse local patches adaptively. Second, to build rich relations between different local patches, the Vision Transformers (ViT) are used in FER, called ViT-FER. Since the global scope is used to reinforce each local patch, a better representation is obtained to boost the FER performance. Thirdly, the multi-head self-attention allows ViT to jointly attend to features from different information subspaces at different positions. Given no explicit guidance, however, multiple self-attentions may extract similar relations. To address this, the MSAD is proposed to randomly drop one self-attention module. As a result, models are forced to learn rich relations among diverse local patches. Our proposed TransFER model outperforms the state-of-the-art methods on several FER benchmarks, showing its effectiveness and usefulness. Fanglei Xue, Qiangchang Wang, Guodong Guo |
ICCV | 3 |
| 2021 | Uncertainty-aware Binary Neural NetworksabstractBinary Neural Networks (BNN) are promising machine learning solutions for deployment on resource-limited devices. Recent approaches to training BNNs have produced impressive results, but minimizing the drop in accuracy from full precision networks is still challenging. One reason is that conventional BNNs ignore the uncertainty caused by weights that are near zero, resulting in the instability or frequent flip while learning. In this work, we investigate the intrinsic uncertainty of vanishing near-zero weights, making the training vulnerable to instability. We introduce an uncertainty-aware BNN (UaBNN) by leveraging a new mapping function called certainty-sign (c-sign) to reduce these weights' uncertainties. Our c-sign function is the first to train BNNs with a decreasing uncertainty for binarization. The approach leads to a controlled learning process for BNNs. We also introduce a simple but effective method to measure the uncertainty-based on a Gaussian function. Extensive experiments demonstrate that our method improves multiple BNN methods by maintaining stability of training, and achieves a higher performance over prior arts. Junhe Zhao, Linlin Yang 0001, Baochang Zhang 0001, Guodong Guo, David S. Doermann |
IJCAI | 4 |
| 2021 | Energy-Based Supervised Hashing for Multimorbidity Image Retrieval
Xiuzhuang Zhou, Zeqiang Wei, Guodong Guo |
MICCAI (5) | 4 |
| 2021 | Domain-Aware SE Network for Sketch-based Image Retrieval with Multiplicative Euclidean Margin SoftmaxabstractThis paper proposes a novel approach for Sketch-Based Image Retrieval (SBIR), for which the key is to bridge the gap between sketches and photos in terms of the data representation. Inspired by channel-wise attention explored in recent years, we present a Domain-Aware Squeeze-and-Excitation (DASE) network, which seamlessly incorporates the prior knowledge of sample sketch or photo into SE module and make the SE module capable of emphasizing appropriate channels according to domain signal. Accordingly, the proposed network can switch its mode to achieve a better domain feature with lower intra-class discrepancy. Moreover, while previous works simply focus on minimizing intra-class distance and maximizing inter-class distance, we introduce a loss function, named Multiplicative Euclidean Margin Softmax (MEMS), which introduces multiplicative Euclidean margin into feature space and ensure that the maximum intra-class distance is smaller than the minimum inter-class distance. This facilitates learning a highly discriminative feature space and ensures a more accurate image retrieval result. Extensive experiments are conducted on two widely used SBIR benchmark datasets. Our approach achieves better results on both datasets, surpassing the state-of-the-art methods by a large margin. Gao Huang 0001, Wenming Yang, Guodong Guo, Yanwei Fu 0001 |
ACM Multimedia | 5 |
| 2021 | CASIA-SURF CeFA: A Benchmark for Multi-modal Cross-ethnicity Face Anti-spoofingabstractThe issue of ethnic bias has proven to affect the performance of face recognition in previous works, while it still remains to be vacant in face anti-spoofing. Therefore, in order to study the ethnic bias for face anti-spoofing, we introduce the largest CASIA-SURF Cross-ethnicity Face Anti-spoofing (CeFA) dataset, covering 3 ethnicities, 3 modalities, 1,607 subjects, and 2D plus 3D attack types. Five protocols are introduced to measure the affect under varied evaluation conditions, such as cross-ethnicity, unknown spoofs or both of them. As our knowledge, CASIA-SURF CeFA is the first dataset including explicit ethnic labels in current released datasets. Then, we propose a novel multi-modal fusion method as a strong baseline to alleviate the ethnic bias, which employs a partially shared fusion strategy to learn complementary information from multiple modalities. Extensive experiments have been conducted on the proposed dataset to verify its significance and generalization capability for other existing datasets, i.e., CASIA-SURF, OULU-NPU and SiW datasets. The dataset is available at https://sites.google.com/qq.com/face-anti-spoofing/welcome/challengecvpr2020?authuser=0. Ajian Liu 0001, Zichang Tan, Jun Wan 0001, Sergio Escalera, Guodong Guo, Stan Z. Li |
WACV | 5 |
| 2021 | A new approach to finding the extra connectivity of graphs
Qiang Zhu 0003, Fang Ma, Guodong Guo, Dajin Wang |
Discret. Appl. Math. | 3 |
| 2021 | Binarized Neural Architecture Search for Efficient Object Recognition
Lian Zhuo, Baochang Zhang 0001, Xiawu Zheng, Jianzhuang Liu, Rongrong Ji, David S. Doermann, Guodong Guo |
Int. J. Comput. Vis. | 8 |
| 2021 | Rectified Binary Convolutional Networks with Generative Adversarial Learning
Chunlei Liu 0001, Wenrui Ding, Baochang Zhang 0001, Jianzhuang Liu, Guodong Guo, David S. Doermann |
Int. J. Comput. Vis. | 6 |
| 2021 | Cascaded Split-and-Aggregate Learning with Feature Recombination for Pedestrian Attribute Recognition
Yang Yang 0062, Zichang Tan, Prayag Tiwari, Hari Mohan Pandey, Jun Wan 0001, Zhen Lei 0001, Guodong Guo, Stan Z. Li |
Int. J. Comput. Vis. | 7 |
| 2021 | A survey on dorsal hand vein biometrics
Wei Jia 0001, Bob Zhang 0001, Yang Zhao 0002, Lunke Fei, Wenxiong Kang, Di Huang 0001, Guodong Guo |
Pattern Recognit. | 8 |
| 2021 | Video-Based Depression Level Analysis by Encoding Deep Spatiotemporal FeaturesabstractAs a serious mood disorder problem, depression causes severe symptoms that affect how people feel, think, and handle daily activities, such as sleeping, eating, or working. In this paper, a novel framework is proposed to estimate the Beck Depression Inventory II (BDI-II) values from video data, which uses a 3D convolutional neural network to automatically learn the spatiotemporal features at two different scales of the face regions. Then, a Recurrent Neural Network (RNN) is used to learn further from the sequence of the spatiotemporal information. This formulation, called RNN-C3D, can model the local and global spatiotemporal information from consecutive face expressions, in order to predict the depression levels. Experiments on the AVEC2013 and AVEC2014 depression datasets show that our proposed approach is promising, when compared to the state-of-the-art visual-based depression analysis methods. Mohamad Al Jazaery, Guodong Guo |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | 3D Face Anti-Spoofing With Factorized Bilinear CodingabstractWe have witnessed rapid advances in both face presentation attack models and presentation attack detection (PAD) in recent years. When compared with widely studied 2D face presentation attacks, 3D face spoofing attacks are more challenging because face recognition systems are more easily confused by the 3D characteristics of materials similar to real faces. In this work, we tackle the problem of detecting these realistic 3D face presentation attacks and propose a novel anti-spoofing method from the perspective of fine-grained classification. Our method, based on factorized bilinear coding of multiple color channels (namely MC_FBC), targets at learning subtle fine-grained differences between real and fake images. By extracting discriminative and fusing complementary information from RGB and YCbCr spaces, we have developed a principled solution to 3D face spoofing detection. A large-scale wax figure face database (WFFD) with both images and videos has also been collected as super realistic attacks to facilitate the study of 3D face presentation attack detection. Extensive experimental results show that our proposed method achieves the state-of-the-art performance on both our own WFFD and other face spoofing databases under various intra-database and inter-database testing scenarios. Shan Jia, Xin Li 0005, Chuanbo Hu, Guodong Guo, Zhengquan Xu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Face Anti-Spoofing via Adversarial Cross-Modality TranslationabstractFace Presentation Attack Detection (PAD) approaches based on multi-modal data have been attracted increasingly by the research community. However, they require multi-modal face data consistently involved in both the training and testing phases. It would severely limit the applicability due to the most Face Anti-spoofing (FAS) systems are only equipped with Visible (VIS) imaging devices, i.e., RGB cameras. Therefore, how to use other modality (i.e., Near-Infrared (NIR)) to assist the performance improvement of VIS-based PAD is significant for FAS. In this work, we first discuss the big gap of performances among different modalities even though the same backbone network is applied. Then, we propose a novel Cross-modal Auxiliary (CMA) framework for the VIS-based FAS task. The main trait of CMA is that the performance can be greatly improved with the help of other modality while no other modality is required in the testing stage. The proposed CMA consists of a Modality Translation Network (MT-Net) and a Modality Assistance Network (MA-Net). The former aims to close the visible gap between different modalities via a generative model that maps inputs from one modality (i.e., RGB) to another (i.e., NIR). The latter focuses on how to use the translated modality (i.e., target modality) and RGB modality (i.e., source modality) together to train a discriminative PAD model. Extensive experiments are conducted to demonstrate that the proposed framework can push the state-of-the-art (SOTA) performances on both multi-modal datasets (i.e., CASIA-SURF, CeFA, and WMCA) and RGB-based datasets (i.e., OULU-NPU, and SiW). Ajian Liu 0001, Zichang Tan, Jun Wan 0001, Yanyan Liang 0001, Zhen Lei 0001, Guodong Guo, Stan Z. Li |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2021 | DSA-Face: Diverse and Sparse Attentions for Face Recognition Robust to Pose Variation and OcclusionabstractLearning local representations is important for face recognition (FR). Recent attention-based networks emphasize few facial parts, while ignoring other potentially discriminative ones. This is more serious when there are large pose variations, occlusions (e.g. face masks), or other image quality changes. To address this, we propose Diverse and Sparse Attentions, called DSA-Face. First, a divergence loss is designed to explicitly encourage the diversity among multiple attention maps by maximizing the Euclidean distance between every pair attention maps. As a result, a Pairwise Self-Contrastive Attention (PSCA) is developed to locate diverse facial parts which provide comprehensive descriptions. Second, an Attention Sparsity Loss (ASL) is proposed to encourage sparse responses in attention maps where only discriminative parts are emphasized while distracted regions (e.g. background or face masks) are discouraged. Built upon the PSCA and ASL, the DSA-Face model is developed to learn diverse and sparse attentions, which can extract diverse discriminative local representations and suppress the focus on noisy regions. Due to the pandemic of the COVID-19, the task of masked face matching is now very important, and our model can handle this much better than previous methods, demonstrating its effectiveness and usefulness. Moreover, our model outperforms the state-of-the-art methods on several other FR benchmarks, showing that it is also general to address various challenges in FR. Qiangchang Wang, Guodong Guo |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Adversarial Attack Against Deep Saliency Models Powered by Non-Redundant PriorsabstractSaliency detection is an effective front-end process to many security-related tasks, e.g. automatic drive and tracking. Adversarial attack serves as an efficient surrogate to evaluate the robustness of deep saliency models before they are deployed in real world. However, most of current adversarial attacks exploit the gradients spanning the entire image space to craft adversarial examples, ignoring the fact that natural images are high-dimensional and spatially over-redundant, thus causing expensive attack cost and poor perceptibility. To circumvent these issues, this paper builds an efficient bridge between the accessible partially-white-box source models and the unknown black-box target models. The proposed method includes two steps: 1) We design a new partially-white-box attack, which defines the cost function in the compact hidden space to punish a fraction of feature activations corresponding to the salient regions, instead of punishing every pixel spanning the entire dense output space. This partially-white-box attack reduces the redundancy of the adversarial perturbation. 2) We exploit the non-redundant perturbations from some source models as the prior cues, and use an iterative zeroth-order optimizer to compute the directional derivatives along the non-redundant prior directions, in order to estimate the actual gradient of the black-box target model. The non-redundant priors boost the update of some "critical" pixels locating at non-zero coordinates of the prior cues, while keeping other redundant pixels locating at the zero coordinates unaffected. Our method achieves the best tradeoff between attack ability and perturbation redundancy. Finally, we conduct a comprehensive experiment to test the robustness of 18 state-of-the-art deep saliency models against 16 malicious attacks, under both of white-box and black-box settings, which contributes a new robustness benchmark to the saliency community for the first time. Zhaohui Che, Ali Borji, Guangtao Zhai, Suiyi Ling, Jing Li 0026, Yuan Tian 0017, Guodong Guo, Patrick Le Callet |
IEEE Trans. Image Process. | 7 |
| 2021 | AAN-Face: Attention Augmented Networks for Face RecognitionabstractConvolutional neural networks are capable of extracting powerful representations for face recognition. However, they tend to suffer from poor generalization due to imbalanced data distributions where a small number of classes are over-represented (e.g. frontal or non-occluded faces) and some of the remaining rarely appear (e.g. profile or heavily occluded faces). This is the reason why the performance is dramatically degraded in minority classes. For example, this issue is serious for recognizing masked faces in the scenario of ongoing pandemic of the COVID-19. In this work, we propose an Attention Augmented Network, called AAN-Face, to handle this issue. First, an attention erasing (AE) scheme is proposed to randomly erase units in attention maps. This well prepares models towards occlusions or pose variations. Second, an attention center loss (ACL) is proposed to learn a center for each attention map, so that the same attention map focuses on the same facial part. Consequently, discriminative facial regions are emphasized, while useless or noisy ones are suppressed. Third, the AE and the ACL are incorporated to form the AAN-Face. Since the discriminative parts are randomly removed by the AE, the ACL is encouraged to learn different attention centers, leading to the localization of diverse and complementary facial parts. Comprehensive experiments on various test datasets, especially on masked faces, demonstrate that our AAN-Face models outperform the state-of-the-art methods, showing the importance and effectiveness. Qiangchang Wang, Guodong Guo |
IEEE Trans. Image Process. | 2 |
| 2021 | Multi-human Parsing with a Graph-based Generative Adversarial ModelabstractHuman parsing is an important task in human-centric image understanding in computer vision and multimedia systems. However, most existing works on human parsing mainly tackle the single-person scenario, which deviates from real-world applications where multiple persons are present simultaneously with interaction and occlusion. To address such a challenging multi-human parsing problem, we introduce a novel multi-human parsing model named MH-Parser, which uses a graph-based generative adversarial model to address the challenges of close-person interaction and occlusion in multi-human parsing. To validate the effectiveness of the new model, we collect a new dataset named Multi-Human Parsing (MHP), which contains multiple persons with intensive person interaction and entanglement. Experiments on the new MHP dataset and existing datasets demonstrate that the proposed method is effective in addressing the multi-human parsing problem compared with existing solutions in the literature. Jianshu Li, Jian Zhao 0006, Congyan Lang, Yidong Li, Yunchao Wei, Guodong Guo, Terence Sim, Shuicheng Yan, Jiashi Feng |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2021 | Adaptive cross-fusion learning for multi-modal gesture recognitionabstractGesture recognition has attracted significant attention because of its wide range of potential applications. Although multi-modal gesture recognition has made significant progress in recent years, a popular method still is simply fusing prediction scores at the end of each branch, which often ignores complementary features among different modalities in the early stage and does not fuse the complementary features into a more discriminative feature. This paper proposes an Adaptive Cross-modal Weighting (ACmW) scheme to exploit complementarity features from RGB-D data in this study. The scheme learns relations among different modalities by combining the features of different data streams. The proposed ACmW module contains two key functions: (1) fusing complementary features from multiple streams through an adaptive one-dimensional convolution; and (2) modeling the correlation of multi-stream complementary features in the time dimension. Through the effective combination of these two functional modules, the proposed ACmW can automatically analyze the relationship between the complementary features from different streams, and can fuse them in the spatial and temporal dimensions. Extensive experiments validate the effectiveness of the proposed method, and show that our method outperforms state-of-the-art methods on IsoGD and NVGesture. Benjia Zhou, Jun Wan 0001, Yanyan Liang 0001, Guodong Guo |
Virtual Real. Intell. Hardw. | 4 |
| 2020 | Relation-Aware Pedestrian Attribute Recognition with Graph Convolutional NetworksabstractIn this paper, we propose a new end-to-end network, named Joint Learning of Attribute and Contextual relations (JLAC), to solve the task of pedestrian attribute recognition. It includes two novel modules: Attribute Relation Module (ARM) and Contextual Relation Module (CRM). For ARM, we construct an attribute graph with attribute-specific features which are learned by the constrained losses, and further use Graph Convolutional Network (GCN) to explore the correlations among multiple attributes. For CRM, we first propose a graph projection scheme to project the 2-D feature map into a set of nodes from different image regions, and then employ GCN to explore the contextual relations among those regions. Since the relation information in the above two modules is correlated and complementary, we incorporate them into a unified framework to learn both together. Experiments on three benchmarks, including PA-100K, RAP, PETA attribute datasets, demonstrate the effectiveness of the proposed JLAC. Zichang Tan, Yang Yang 0062, Jun Wan 0001, Guodong Guo, Stan Z. Li |
AAAI | 4 |
| 2020 | Few-Shot Learning with Complex-valued Neural Networks
Baochang Zhang 0001, Guodong Guo |
BMVC | 3 |
| 2020 | Hierarchical Pyramid Diverse Attention Networks for Face RecognitionabstractDeep learning has achieved a great success in face recognition (FR), however, few existing models take hierarchical multi-scale local features into consideration. In this work, we propose a hierarchical pyramid diverse attention (HPDA) network. First, it is observed that local patches would play important roles in FR when the global face appearance changes dramatically. Some recent works apply attention modules to locate local patches automatically without relying on face landmarks. Unfortunately, without considering diversity, some learned attentions tend to have redundant responses around some similar local patches, while neglecting other potential discriminative facial parts. Meanwhile, local patches may appear at different scales due to pose variations or large expression changes. To alleviate these challenges, we propose a pyramid diverse attention (PDA) to learn multi-scale diverse local representations automatically and adaptively. More specifically, a pyramid attention is developed to capture multi-scale features. Meanwhile, a diverse learning is developed to encourage models to focus on different local patches and generate diverse local features. Second, almost all existing models focus on extracting features from the last convolutional layer, lacking of local details or small-scale face parts in lower layers. Instead of simple concatenation or addition, we propose to use a hierarchical bilinear pooling (HBP) to fuse information from multiple layers effectively. Thus, the HPDA is developed by integrating the PDA into the HBP. Experimental results on several datasets show the effectiveness of the HPDA, compared to the state-of-the-art methods. Qiangchang Wang, He Zheng, Guodong Guo |
CVPR | 4 |
| 2020 | Cogradient Descent for Bilinear OptimizationabstractConventional learning methods simplify the bilinear model by regarding two intrinsically coupled factors independently, which degrades the optimization procedure. One reason lies in the insufficient training due to the asynchronous gradient descent, which results in vanishing gradients for the coupled variables. In this paper, we introduce a Cogradient Descent algorithm (CoGD) to address the bilinear problem, based on a theoretical framework to coordinate the gradient of hidden variables via a projection function. We solve one variable by considering its coupling relationship with the other, leading to a synchronous gradient descent to facilitate the optimization procedure. Our algorithm is applied to solve problems with one variable under the sparsity constraint, which is widely used in the learning paradigm. We validate our CoGD considering an extensive set of applications including image reconstruction, inpainting, and network pruning. Experiments show that it improves the state-of-the-art by a significant margin. Lian Zhuo, Baochang Zhang 0001, Linlin Yang 0001, Qixiang Ye, David S. Doermann, Rongrong Ji, Guodong Guo |
CVPR | 8 |
| 2020 | GINet: Graph Interaction Network for Scene Parsing
Yu Zhu 0006, Ming Wu 0001, Zhanyu Ma, Guodong Guo |
ECCV (17) | 7 |
| 2020 | 3DPC-Net: 3D Point Cloud Network for Face Anti-spoofingabstractFace anti-spoofing plays a vital role in face recognition systems. Most deep learning-based methods directly use 2D images assisted with temporal information (i.e., motion, rPPG) or pseudo-3D information (i.e., Depth). The main drawback of the mentioned methods is that another extra network is needed to generate the depth/rPPG information to assist the backbone network for face anti-spoofing. Different from these methods, we propose a novel method named 3D Point Cloud Network (3DPC-Net). It is an encoder-decoder network that can predict the 3DPC maps to discriminate live faces from spoofing ones. The main traits of the proposed method are that: 1) It is the first time that 3DPC is used for face anti-spoofing; 2) 3DPC-Net is simple and effective and it only relies on 3DPC supervision. Extensive experiments on four databases (i.e., Oulu-NPU, SiW, CASIA-FASD, Replay Attack) have demonstrated that the 3DPC-Net is comparative to the state-of-the-art methods. Jun Wan 0001, Yi Jin 0001, Ajian Liu 0001, Guodong Guo, Stan Z. Li |
IJCB | 5 |
| 2020 | Visual BMI estimation from face images using a label distribution based method
Min Jiang 0003, Guodong Guo, Guowang Mu |
Comput. Vis. Image Underst. | 2 |
| 2020 | Crowd counting by the dual-branch scale-aware network with ranking loss constraintsabstractImage crowd counting is a challenging problem. This study proposes a new deep learning method that estimates crowd counting for the congested scene. The proposed network is composed of two major components: the first ten layers of VGG16 are used as the backbone network, and a dual‐branch (named as Branch_S and Branch_D) network is proposed to be the second part of the network. Branch_S extracts low‐level information (head blob) through a shallow fully convolutional network and Branch_D uses a deep fully convolutional network to extract high‐level context features (faces and body). Features learnt from the two different branches can handle the problem of scale variation due to perspective effects and image size differences. Features of different scales extracted from the two branches are fused to generate predicted density map. On the basis of the fact that an original graph must contain more or equal number of persons than any of its sub‐images, a ranking loss function utilising the constraint relationship inside an image is proposed. Moreover, the ranking loss is combined with Euclidean loss as the final loss function. Our approach is evaluated on three benchmark datasets, and better results are achieved compared with the state‐of‐the‐art works. Fangfang Yan, ZhiLei Chai, Guodong Guo |
IET Comput. Vis. | 4 |
| 2020 | Computational approach to body mass index estimation from dressed people in 3D spaceabstractBody mass index (BMI) defines as a person's weight divided by the square of height (BMI ), which is an important indicator of the health condition. The authors study BMI estimation from the three‐dimensional (3D) visual data by measuring the correlation between the estimated body volume and BMIs, and then develop an efficient BMI computation method. Their approach consists of body weight and height estimation from normally dressed people in 3D space. To address the influence of loose clothes on body volume estimation, two clothes models are developed to make the volume estimation more accurate. A new RGB‐D video dataset is collected for this study, and the reconstructed 3D data are provided by the KinectFusion on depth data. Experimental results show the effectiveness of the approach to work on normal conditions of dressed people. The mean absolute error of the estimated BMI can achieve 2.54 in their experiments. Min Jiang 0003, Guodong Guo |
IET Image Process. | 3 |
| 2020 | Flow driven attention network for video salient object detectionabstractSalient object detection has been revolutionised by convolutional neural network (CNN) recently. However, it is hard to transfer the state‐of‐the‐art still‐image based saliency detectors to videos directly, owing to the neglect of temporal contexts between frames. In this study, the authors propose a flow‐driven attention network (FDAN) to exploit motion information for video salient object detection. FDAN consists of an appearance feature extractor, a motion‐guided attention module and a saliency map regression module. It extracts the appearance feature per frame, refines appearance feature with optical flow and infers the ultimate saliency map, respectively. Motion‐guided attention module is the core of FDAN, which extracts motion information in the form of attention. This attention mechanism is a two‐branch CNN, fusing optical flow and appearance features. In addition, a shortcut connection is applied to the attention multiplied feature map for noise suppression intensively. Experimental results show that the proposed method can achieve performance on par with the state‐of‐the‐art method flow‐guided recurrent neural encoder on challenging benchmarks of Densely Annotated Video Segmentation and Freiburg–Berkeley Motion Segmentation while being two times faster in detection. Feng Zhou 0006, Hui Shuai, Qingshan Liu 0001, Guodong Guo |
IET Image Process. | 4 |
| 2020 | CR-Net: A Deep Classification-Regression Network for Multimodal Apparent Personality Analysis
Yunan Li 0001, Jun Wan 0001, Qiguang Miao, Sergio Escalera, Huijuan Fang, Huizhou Chen, Xiangda Qi, Guodong Guo |
Int. J. Comput. Vis. | 8 |
| 2020 | Face presentation attack detection in mobile scenarios: A comprehensive evaluation
Shan Jia, Guodong Guo, Zhengquan Xu, Qiangchang Wang |
Image Vis. Comput. | 2 |
| 2020 | A survey on 3D mask presentation attack detection and countermeasures
Shan Jia, Guodong Guo, Zhengquan Xu |
Pattern Recognit. | 2 |
| 2020 | Visually Interpretable Representation Learning for Depression Recognition from Facial ImagesabstractRecent evidence in mental health assessment have demonstrated that facial appearance could be highly indicative of depressive disorder. While previous methods based on the facial analysis promise to advance clinical diagnosis of depressive disorder in a more efficient and objective manner, challenges in visual representation of complex depression pattern prevent widespread practice of automated depression diagnosis. In this paper, we present a deep regression network termed DepressNet to learn a depression representation with visual explanation. Specifically, a deep convolutional neural network equipped with a global average pooling layer is first trained with facial depression data, which allows for identifying salient regions of input image in terms of its severity score based on the generated depression activation map (DAM). We then propose a multi-region DepressNet, with which multiple local deep regression models for different face regions are jointly leaned and their responses are fused to improve the overall recognition performance. We evaluate our method on two benchmark datasets, and the results show that our method significantly boosts state-of-the-art performance of the visual-based depression recognition. Most importantly, the DAM induced by our learned deep model may help reveal the visual depression pattern on faces and understand the insights of automated depression diagnosis. Xiuzhuang Zhou, Guodong Guo |
IEEE Trans. Affect. Comput. | 4 |
| 2020 | Task-Oriented Feature-Fused Network With Multivariate Dataset for Joint Face AnalysisabstractDeep multitask learning for face analysis has received increasing attentions. From literature, most existing methods focus on optimizing a main task by jointly learning several auxiliary tasks. It is challenging to consider the performance of each task in a multitask framework due to the following reasons: 1) different face tasks usually rely on different levels of semantic features; 2) each task has different learning convergence rate, which could affect the whole performance when joint training; and 3) multitask model needs rich label information for efficient training, but existing facial datasets provide limited annotations. To address these issues, we propose a task-oriented feature-fused network (TFN) for simultaneously solving face detection, landmark localization, and attribute analysis. In this network, a task-oriented feature-fused block is designed to learn task-specific feature combinations; then, an alternative multitask training scheme is presented to optimize each task with considering of their different learning capacities. We also present a large-scale face dataset called JFA in support of proposed method, which provides multivariate labels, including face bounding box, 68 facial landmarks, and 3 attribute labels (i.e., apparent age, gender, and ethnicity). The experimental results suggest that the TFN outperforms several multitask models on the JFA dataset. Furthermore, our approach achieves competitive performances on WIDER FACE and 300W dataset, and obtains state-of-the-art results for gender recognition on the MORPH II dataset. Xuxin Lin, Jun Wan 0001, Yiliang Xie, Chi Lin 0002, Yanyan Liang 0001, Guodong Guo, Stan Z. Li |
IEEE Trans. Cybern. | 7 |
| 2020 | LS-CNN: Characterizing Local Patches at Multiple Scales for Face RecognitionabstractFaces in the wild may contain pose variations, age changes, and with different qualities which significantly enlarge the intra-class variations. Although great progresses have been made in face recognition, few existing works could learn local and multi-scale representations together. In this work, we propose a new model, called Local and multi-Scale Convolutional Neural Networks (LS-CNN). First, since similar discriminative face regions may occur at different scales, it is necessary to learn multi-scale features. To this aim, we introduce a new backbone network, namely Harmonious multi-Scale Network (HSNet), which extracts rich multi-scale features from two harmonious perspectives: utilization of different kernel sizes in a single layer, and concatenation of multi-scale feature maps from different layers. Second, identifying similar local patches is important when global face appearances have dramatic changes. Meanwhile, different face regions have different discriminative abilities. To capture critical local similarities and weigh adaptively on different local patches, a spatial attention is proposed. Third, channels have different convolutional kernels which can detect different features with various importance. Besides, hierarchical channels concatenated from different layers contain diverse information: channels from low layers describe local details or small-scale parts, and channels in high layers represent high-level abstraction or large-scale parts. To emphasize important channels and suppress less informative ones automatically, channel attention is used. Due to the complementary characteristics of channel attention and spatial attention, they are fused to form the Dual Face Attentions (DFA). To the best of our knowledge, this is the first effort to employ attentions for the general face recognition task. The LS-CNN is developed by incorporating DFA into HSNet model. Experimental results on various face matching tasks show its capability of learning complex data distributions. Qiangchang Wang, Guodong Guo |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | How is Gaze Influenced by Image Transformations? Dataset and ModelabstractData size is the bottleneck for developing deep saliency models, because collecting eye-movement data is very time-consuming and expensive. Most of current studies on human attention and saliency modeling have used high-quality stereotype stimuli. In real world, however, captured images undergo various types of transformations. Can we use these transformations to augment existing saliency datasets? Here, we first create a novel saliency dataset including fixations of 10 observers over 1900 images degraded by 19 types of transformations. Second, by analyzing eye movements, we find that observers look at different locations over transformed versus original images. Third, we utilize the new data over transformed images, called data augmentation transformation (DAT), to train deep saliency models. We find that label-preserving DATs with negligible impact on human gaze boost saliency prediction, whereas some other DATs that severely impact human gaze degrade the performance. These label-preserving valid augmentation transformations provide a solution to enlarge existing saliency datasets. Finally, we introduce a novel saliency model based on generative adversarial networks (dubbed GazeGAN). A modified U-Net is utilized as the generator of the GazeGAN, which combines classic "skip connection" with a novel "center-surround connection" (CSC) module. Our proposed CSC module mitigates trivial artifacts while emphasizing semantic salient regions, and increases model nonlinearity, thus demonstrating better robustness against transformations. Extensive experiments and comparisons indicate that GazeGAN achieves state-of-the-art performance over multiple datasets. We also provide a comprehensive comparison of 22 saliency models on various transformed scenes, which contributes a new robustness benchmark to saliency community. Our code and dataset are available at. Zhaohui Che, Ali Borji, Guangtao Zhai, Xiongkuo Min, Guodong Guo, Patrick Le Callet |
IEEE Trans. Image Process. | 5 |
| 2020 | Aggregation Signature for Small Object TrackingabstractSmall object tracking becomes an increasingly important task, which however has been largely unexplored in computer vision. The great challenges stem from the facts that: 1) small objects show extreme vague and variable appearances, and 2) they tend to be lost easier as compared to normal-sized ones due to the shaking of lens. In this paper, we propose a novel aggregation signature suitable for small object tracking, especially aiming for the challenge of sudden and large drift. We make three-fold contributions in this work. First, technically, we propose a new descriptor, named aggregation signature, based on saliency, able to represent highly distinctive features for small objects. Second, theoretically, we prove that the proposed signature matches the foreground object more accurately with a high probability. Third, experimentally, the aggregation signature achieves a high performance on multiple datasets, outperforming the state-of-the-art methods by large margins. Moreover, we contribute with two newly collected benchmark datasets, i.e., small90 and small112, for visually small object tracking. The datasets will be available in https://github.com/bczhangbczhang/. Chunlei Liu 0001, Wenrui Ding, Vittorio Murino, Baochang Zhang 0001, Jungong Han, Guodong Guo |
IEEE Trans. Image Process. | 7 |
| 2020 | Dual-Path Attention Network for Compressed Sensing Image ReconstructionabstractAlthough deep neural network methods achieved much success in compressed sensing image reconstruction in recent years, they still have some issues, especially in preserving texture details. In this paper, we propose a new dual-path attention network for compressed sensing image reconstruction, which is composed of a structure path, a texture path and a texture attention module. Motivated by the classical paradigm of image structure-texture decomposition, the structure path aims to reconstruct the dominant structure component of the original image, and the texture path targets at recovering the remaining texture details. To better bridge the information between two paths, the texture attention module is designed to deliver the useful structure information to the texture path and predict the texture region, thereby facilitating the recovery of texture details. Two paths are optimized with a unified loss function. In the testing phase, given the measurement vector of a new image, it can be well reconstructed by carrying out the well trained dual-path attention network and integrating the outputs of the structure path and the texture path. Experimental results on the SET5, SET11 and BSD68 testing datasets demonstrate that the proposed method achieves comparable or better results compared with some state-of-the-art deep learning based methods and conventional iterative optimization based methods in terms of reconstruction quality and robustness to noise. Yubao Sun, Qingshan Liu 0001, Bo Liu 0005, Guodong Guo |
IEEE Trans. Image Process. | 5 |
| 2020 | Learning Non-Locally Regularized Compressed Sensing Network With Half-Quadratic SplittingabstractDeep learning-based Compressed Sensing (CS) reconstruction attracts much attention in recent years, due to its significant superiority of reconstruction quality. Its success is mainly attributed to the employment of a large dataset for pre-training the network to learn a reconstruction mapping. In this paper, we propose a non-locally regularized compressed sensing network for reconstructing image sequences, which can achieve high reconstruction quality without pre-training. Specifically, the proposed method attempts to learn a deep network prior for the reconstruction of an individual instance under the constraint that the network output can well match the given CS measurement. The non-local prior is designed to guide the network to capture the long-range dependencies by exploiting the self-similarities among images, and it can also make the network noise-aware. In order to deal with the compound of non-local prior and deep network prior, we construct a half-quadratic splitting based optimization method for network learning, in which the two priors are decoupled into two simple sub-problems by introducing an auxiliary variable and a quadratic fidelity constraint. Extensive experimental results demonstrate that our method is competitive to the popular methods, including sparsity prior based methods and deep learning based methods, even better than them in the cases of low measurement rates. Yubao Sun, Qingshan Liu 0001, Xiao-Tong Yuan, Guodong Guo |
IEEE Trans. Multim. | 6 |
| 2020 | WiderPerson: A Diverse Dataset for Dense Pedestrian Detection in the WildabstractPedestrian detection has achieved significant progress with the availability of existing benchmark datasets. However, there is a gap in the diversity and density between real world requirements and current pedestrian detection benchmarks: first, most existing datasets are taken from a vehicle driving through the regular traffic scenario, usually leading to insufficient diversity; second, crowd scenarios with highly occluded pedestrians are still underrepresented, resulting in low density. To narrow this gap and facilitate future pedestrian detection research, we introduce a large and diverse dataset named WiderPerson for dense pedestrian detection in the wild. This dataset involves five types of annotations in a wide range of scenarios, no longer limited to the traffic scenario. There are a total of 13 382 images with 399 786 annotations, that is, 29.87 annotations per image, which means this dataset contains dense pedestrians with various kinds of occlusions. Hence, pedestrians in the proposed dataset are extremely challenging due to large variations in the scenario and occlusion, which is suitable to evaluate pedestrian detectors in the wild. We introduce an improved Faster R-CNN and the vanilla RetinaNet to serve as baselines for the new pedestrian detection benchmark. Several experiments are conducted on previous datasets including Caltech-USA and CityPersons to analyze the generalization capabilities of the proposed dataset, and we achieve state-of-the-art performances on these previous datasets without bells and whistles. Finally, we analyze common failure cases and find the classification ability of pedestrian detector needs to be improved to reduce false alarm and misdetection rates. The proposed dataset is available at http://www.cbsr.ia.ac.cn/users/sfzhang/WiderPerson. Yiliang Xie, Jun Wan 0001, Hansheng Xia, Stan Z. Li, Guodong Guo |
IEEE Trans. Multim. | 6 |
| 2019 | Bayesian Optimized 1-Bit CNNsabstractDeep convolutional neural networks (DCNNs) have dominated the recent developments in computer vision through making various record-breaking models. However, it is still a great challenge to achieve powerful DCNNs in resource-limited environments, such as on embedded devices and smart phones. Researchers have realized that 1-bit CNNs can be one feasible solution to resolve the issue; however, they are baffled by the inferior performance compared to the full-precision DCNNs. In this paper, we propose a novel approach, called Bayesian optimized 1-bit CNNs (denoted as BONNs), taking the advantage of Bayesian learning, a well-established strategy for hard problems, to significantly improve the performance of extreme 1-bit CNNs. We incorporate the prior distributions of full-precision kernels and features into the Bayesian framework to construct 1-bit CNNs in an end-to-end manner, which have not been considered in any previous related methods. The Bayesian losses are achieved with a theoretical support to optimize the network simultaneously in both continuous and discrete spaces, aggregating different losses jointly to improve the model capacity. Extensive experiments on the ImageNet and CIFAR datasets show that BONNs achieve the best classification performance compared to state-of-the-art 1-bit CNNs. Jiaxin Gu, Junhe Zhao, Baochang Zhang 0001, Jianzhuang Liu, Guodong Guo, Rongrong Ji |
ICCV | 6 |
| 2019 | Generalized Zero-Shot Vehicle Detection in Remote Sensing Imagery via Coarse-to-Fine FrameworkabstractVehicle detection and recognition in remote sensing images are challenging, especially when only limited training data are available to accommodate various target categories. In this paper, we introduce a novel coarse-to-fine framework, which decomposes vehicle detection into segmentation-based vehicle localization and generalized zero-shot vehicle classification. Particularly, the proposed framework can well handle the problem of generalized zero-shot vehicle detection, which is challenging due to the requirement of recognizing vehicles that are even unseen during training. Specifically, a hierarchical DeepLab v3 model is proposed in the framework, which fully exploits fine-grained features to locate the target on a pixel-wise level, then recognizes vehicles in a coarse-grained manner. Additionally, the hierarchical DeepLab v3 model is beneficially compatible to combine the generalized zero-shot recognition. To the best of our knowledge, there is no publically available dataset to test comparative methods, we therefore construct a new dataset to fill this gap of evaluation. The experimental results show that the proposed framework yields promising results on the imperative yet difficult task of zero-shot vehicle detection and recognition. Yongtan Luo, Liujuan Cao, Baochang Zhang 0001, Guodong Guo, Cheng Wang 0003, Jonathan Li 0001, Rongrong Ji |
IJCAI | 5 |
| 2019 | Rectified Binary Convolutional Networks for Enhancing the Performance of 1-bit DCNNsabstractBinarized convolutional neural networks (BCNNs) are widely used to improve memory and computation efficiency of deep convolutional neural networks (DCNNs) for mobile and AI chips based applications. However, current BCNNs are not able to fully explore their corresponding full-precision models, causing a significant performance gap between them. In this paper, we propose rectified binary convolutional networks (RBCNs), towards optimized BCNNs, by combining full-precision kernels and feature maps to rectify the binarization process in a unified framework. In particular, we use a GAN to train the 1-bit binary network with the guidance of its corresponding full-precision model, which significantly improves the performance of BCNNs. The rectified convolutional layers are generic and flexible, and can be easily incorporated into existing DCNNs such as WideResNets and ResNets. Extensive experiments demonstrate the superior performance of the proposed RBCNs over state-of-the-art BCNNs. In particular, our method shows strong generalization on the object tracking task. Chunlei Liu 0001, Wenrui Ding, Xin Xia 0005, Baochang Zhang 0001, Jianzhuang Liu, Bohan Zhuang, Guodong Guo |
IJCAI | 8 |
| 2019 | Deeply-learned Hybrid Representations for Facial Age EstimationabstractIn this paper, we propose a novel unified network named Deep Hybrid-Aligned Architecture for facial age estimation. It contains global, local and global-local branches. They are jointly optimized and thus can capture multiple types of features with complementary information. In each branch, we employ a separate loss for each sub-network to extract the independent features and use a recurrent fusion to explore correlations among those region features. Considering that the pose variations may lead to misalignment in different regions, we design an Aligned Region Pooling operation to generate aligned region features. Moreover, a new large age dataset named Web-FaceAge owning more than 120K samples is collected under diverse scenes and spanning a large age range. Experiments on five age benchmark datasets, including Web-FaceAge, Morph, FG-NET, CACD and Chalearn LAP 2015, show that the proposed method outperforms the state-of-the-art approaches significantly. Zichang Tan, Yang Yang 0062, Jun Wan 0001, Guodong Guo, Stan Z. Li |
IJCAI | 4 |
| 2019 | A survey on deep learning based face recognition
Guodong Guo |
Comput. Vis. Image Underst. | 1 |
| 2019 | Multiview discriminative marginal metric learning for makeup face verificationabstractMakeup face verification in the wild is an important research problem for its popularization in real-world. However, little effort has been made to tackle it in computer vision. In this research, we first build a new database, i.e., Facial Beauty Database (FBD), which contains paired facial images of 8933 subjects without and with makeup in different real-world scenarios. To the best of our knowledge, FBD is the largest makeup face database to date compared with existing databases for facial makeup research. Moreover, we propose a new discriminative marginal metric learning (DMML) algorithm to deal with this problem in the wild. Inspired by the fact that interclass marginal faces are usually more discriminative than interclass nonmarginal faces in learning the discriminative metric space, we use the interclass marginal faces to depict the discriminative information. Simultaneously, we wish that those interclass marginal faces without makeup relations are separated from each other as far as possible, so that more discriminative information between facial images without and with makeup can be exploited for verification. Furthermore, since multiple features could provide comprehensive information in describing the facial representations from diverse points of view and extract more informative cues from facial images, we also introduce a multiview discriminative marginal metric learning (MDMML) algorithm by effectively learning a robust metric space such that multiple features from different points of view can be integrated to effectively enhance the performance of makeup face verification. Experimental results on two real-world makeup face databases are utilized to show the effectiveness of our method and the possibility of verifying the makeup relations from facial images in real-world. Lining Zhang, Hubert P. H. Shum, Li Liu 0004, Guodong Guo, Ling Shao 0001 |
Neurocomputing | 4 |
| 2019 | EMBDN: An Efficient Multiclass Barcode Detection Network for Complicated EnvironmentsabstractThis article presents a novel method for efficient barcodes detection in real and complicated environments using a convolutional neural network (CNN)-based model. The method is developed as a preprocess-module of existing decoders to enhance decoding rates. Our method is trained as an end-to-end model to determine accurate locations of four barcode vertexes. Our method consists of four modules: 1) base net module; 2) region proposals generator; 3) classification and regression module; and 4) distortion removal module. The feature of barcodes extracted from the base net is fed to the next module. Region proposals are generated and selected as region of interest (ROI). Then the ROI are forward propagated to the classification and regression module to determine the positions and shapes of the barcodes. Finally, the distortion removal module is used to remove the geometric distortion according to regression parameters acquired from the previous step. The accurate position and distorted barcodes shape can be determined and corrected by our method. We validate our method on a challenging large-scale dataset in experiments. Compared with the previous methods, our method provides an end-to-end solution to determine accurate locations of barcode vertexes, which shows an excellent performance on detection accuracy. In addition, our method can enhance decoding rate through distortion removal. Jun Jia, Guangtao Zhai, Zhongpai Gao, Zehao Zhu, Xiongkuo Min, Xiaokang Yang 0001, Guodong Guo |
IEEE Internet Things J. | 8 |
| 2019 | On visual BMI analysis from facial images
Min Jiang 0003, Guodong Guo |
Image Vis. Comput. | 3 |
| 2019 | Benchmarking deep learning techniques for face recognition
Qiangchang Wang, Guodong Guo |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Sparse approximation to discriminant projection learning and application to image classification
Yu-Feng Yu 0001, Chuan-Xian Ren, Min Jiang 0003, Man-Yu Sun, Dao-Qing Dai, Guodong Guo |
Pattern Recognit. | 6 |
| 2019 | Body Weight Analysis From Human Body ImagesabstractHuman body images encode plenty of useful biometric information, such as pupil color, gender, and weight. Among these, body weight is a good indicator of health conditions. Motivated by recent health science studies, this paper investigates the feasibility of analyzing body weight from two-dimensional (2D) frontal view human body images. The widely used body mass index (BMI) is employed as a measure of body weight. To investigate the problems at different levels of difficulties, three feasibility problems, from easy to hard, are studied. More specifically, a framework is developed for analyzing body weight from human body images. Computation of five anthropometric features is proposed for body weight characterization. Correlation is analyzed between the extracted anthropometric features and the BMI values, which validates the usability of the selected features. A visual-body-to-BMI dataset is collected and cleaned to facilitate the study, which contains 5900 images of 2950 subjects along with the labels corresponding gender, height, and weight. Some interesting results are obtained, demonstrating the feasibility of analyzing body weight from 2D body images. In addition, the proposed method outperforms two state-of-art facial image-based weight analysis approaches in most cases. Min Jiang 0003, Guodong Guo |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2019 | Deep Manifold Structure Transfer for Action RecognitionabstractWhile intrinsic data structure in subspace provides useful information for visual recognition, it has not yet been well studied in deep feature learning for action recognition. In this paper, we introduce a new spatio-temporal manifold network (STMN) that leverages data manifold structures to regularize deep action feature learning, aiming at simultaneously minimizing the intra-class variations of learned deep features and alleviating the over-fitting problem. To this end, the manifold prior is imposed from the top layer of a convolutional neural network (CNN), and is propagated across convolutional layers during forward-backward propagation. The observed correspondence of manifold structures in the data space and feature space validates that the manifold priori can be transferred across CNN layers. STMN theoretically recasts the problem of transferring the data structure prior into the deep learning architectures as a projection over the manifold via an embedding method, which can be easily solved by an Alternating Direction Method of Multipliers and Backward Propagation (ADMM-BP) algorithm. STMN is generic in the sense that it can be plugged into various backbone architectures to learn more discriminative representation for action recognition. Extensive experimental results show that our method achieves comparable or even better performance as compared with the state-of-the-art approaches on four benchmark datasets. Ce Li 0002, Baochang Zhang 0001, Chen Chen 0001, Qixiang Ye, Jungong Han, Guodong Guo, Rongrong Ji |
IEEE Trans. Image Process. | 6 |
| 2019 | Attention-Based Pedestrian Attribute AnalysisabstractRecognizing the pedestrian attributes in surveillance scenes is an inherently challenging task, especially for the pedestrian images with large pose variations, complex backgrounds, and various camera viewing angles. To select important and discriminative regions or pixels against the variations, three attention mechanisms are proposed, including parsing attention, label attention, and spatial attention. Those attentions aim at accessing effective information by considering problems from different perspectives. To be specific, the parsing attention extracts discriminative features by learning not only where to turn attention to but also how to aggregate features from different semantic regions of human bodies, e.g., head and upper body. The label attention aims at targetedly collecting the discriminative features for each attribute. Different from the parsing and label attention mechanisms, the spatial attention considers the problem from a global perspective, aiming at selecting several important and discriminative image regions or pixels for all attributes. Then, we propose a joint learning framework formulated in a multi-task-like way with these three attention mechanisms learned concurrently to extract complementary and correlated features. This joint learning framework is named Joint Learning of Parsing attention, Label attention, and Spatial attention for Pedestrian Attributes Analysis (JLPLS-PAA, for short). Extensive comparative evaluations conducted on multiple large-scale benchmarks, including PA-100K, RAP, PETA, Market-1501, and Duke attribute datasets, further demonstrate the effectiveness of the proposed JLPLS-PAA framework for pedestrian attribute analysis. Zichang Tan, Yang Yang 0062, Jun Wan 0001, Hanyuan Hang, Guodong Guo, Stan Z. Li |
IEEE Trans. Image Process. | 5 |
| 2019 | Quality Evaluation of Image Dehazing Methods Using Synthetic Hazy ImagesabstractTo enhance the visibility and usability of images captured in hazy conditions, many image dehazing algorithms (DHAs) have been proposed. With so many image DHAs, there is a need to evaluate and compare these DHAs. Due to the lack of the reference haze-free images, DHAs are generally evaluated qualitatively using real hazy images. But it is possible to perform quantitative evaluation using synthetic hazy images since the reference haze-free images are available and full-reference (FR) image quality assessment (IQA) measures can be utilized. In this paper, we follow this strategy and study DHA evaluation using synthetic hazy images systematically. We first build a synthetic haze removing quality (SHRQ) database. It consists of two subsets: regular and aerial image subsets, which include 360 and 240 dehazed images created from 45 and 30 synthetic hazy images using 8 DHAs, respectively. Since aerial imaging is an important application area of dehazing, we create an aerial image subset specifically. We then carry out subjective quality evaluation study on these two subsets. We observe that taking DHA evaluation as an exact FR IQA process is questionable, and the state-of-the-art FR IQA measures are not effective for DHA evaluation. Thus, we propose a DHA quality evaluation method by integrating some dehazing-relevant features, including image structure recovering, color rendition, and over-enhancement of low-contrast areas. The proposed method works for both types of images, but we further improve it for aerial images by incorporating its specific characteristics. Experimental results on two subsets of the SHRQ database validate the effectiveness of the proposed measures. Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Yucheng Zhu, Jiantao Zhou 0001, Guodong Guo, Xiaokang Yang 0001, Xin-Ping Guan, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2019 | Multi-Channel Decomposition in Tandem With Free-Energy Principle for Reduced-Reference Image Quality AssessmentabstractThe visual quality of perceptions is highly correlated with the mechanisms of the human brain and visual system. Recently, the free-energy principle, which has been widely researched in brain theory and neuroscience, is introduced to quantize the perception, action, and learning in human brain. In the field of image quality assessment (IQA), on one hand, the free-energy principle can resort to the internal generative model to simulate the visual stimulus of the human beings. On the other hand, abundant psychological and neurobiological studies reveal that different frequency and orientation components of one visual stimulus arouse different neurons in the striate cortex, and the striate cortex processes visual information in the cerebral cortex. Motivated by these two aspects, a novel reduce-reference IQA metric called the multi-channel free-energy based reduced-reference quality metric is proposed in this paper. First, a two-level discrete Haar wavelet transform is used to decompose the input reference and distorted images. Next, to simulate the generative model in the human brain, the sparse representation is leveraged to extract the free-energy-based features in subband images. Finally, the overall quality metric is obtained through the support vector regressor. Extensive experimental comparisons on four benchmark image quality databases (LIVE, CSIQ, TID2008, and TID2013) demonstrate that the proposed method is highly competitive with the representative reduced-reference and classical full-reference models. Wenhan Zhu, Guangtao Zhai, Xiongkuo Min, Menghan Hu, Jing Liu 0002, Guodong Guo, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 6 |
| 2018 | What Is the Challenge for Deep Learning in Unconstrained Face Recognition?abstractRecently deep learning has become dominant in face recognition and many other artificial intelligence areas. We raise a question: Can deep learning truly solve the face recognition problem? If not, what is the challenge for deep learning methods in face recognition? We think that the face image quality issue might be one of the challenges for deep learning, especially in unconstrained face recognition. To investigate the problem, we partition face images into different qualities, and evaluate the recognition performance, using the state-of-the-art deep networks. Some interesting results are obtained, and our studies can show directions to promote the deep learning methods towards high-accuracy and practical use in solving the hard problem of unconstrained face recognition. Guodong Guo |
FG | 1 |
| 2018 | Attributes in Multiple Facial ImagesabstractFacial attribute recognition is conventionally computed from a single image. In practice, each subject may have multiple face images. Taking the eye size as an example, it should not change, but it may have different estimation in multiple images, which would make a negative impact on face recognition. Thus, how to compute these attributes corresponding to each subject rather than each single image is a profound work. To address this question, we deploy deep training for facial attributes prediction, and we explore the inconsistency issue among the attributes computed from each single image. Then, we develop two approaches to address the inconsistency issue. Experimental results show that the proposed methods can handle facial attribute estimation on either multiple still images or video frames, and can correct the incorrectly annotated labels. The experiments are conducted on two large public databases with annotations of facial attributes. Guodong Guo |
FG | 2 |
| 2018 | Parallelizing Hartley transform with Hadoop for fast detection of glass defectsabstractSummary Glass defect detection methods based on grating projection can effectively detect various glass defects. The Fourier transform in general can be used as an online processing method for detecting glass defects based on fringe images. Processing fringe images with Fourier transform needs a large amount of computation as Fourier transform is a complex computation method. In order to reduce the amount of computation, an improved fringe image processing method based on the Hartley transform is proposed in this paper. To further speed up the computation process, the Hartley transform is parallelized with Hadoop, which is a major computing technology in support of data intensive applications. Experimental results show that the parallel Hartley transform significantly reduces computation complexity in detection of glass defects. Maozhen Li 0001, Yong Jin 0004, Zhaoba Wang, Guodong Guo |
Concurr. Comput. Pract. Exp. | 5 |
| 2018 | Efficient Group-n Encoding and Decoding for Facial Age EstimationabstractDifferent ages are closely related especially among the adjacent ages because aging is a slow and extremely non-stationary process with much randomness. To explore the relationship between the real age and its adjacent ages, an age group-n encoding (AGEn) method is proposed in this paper. In our model, adjacent ages are grouped into the same group and each age corresponds to n groups. The ages grouped into the same group would be regarded as an independent class in the training stage. On this basis, the original age estimation problem can be transformed into a series of binary classification sub-problems. And a deep Convolutional Neural Networks (CNN) with multiple classifiers is designed to cope with such sub-problems. Later, a Local Age Decoding (LAD) strategy is further presented to accelerate the prediction process, which locally decodes the estimated age value from ordinal classifiers. Besides, to alleviate the imbalance data learning problem of each classifier, a penalty factor is inserted into the unified objective function to favor the minority class. To compare with state-of-the-art methods, we evaluate the proposed method on FG-NET, MORPH II, CACD and Chalearn LAP 2015 databases and it achieves the best performance. Zichang Tan, Jun Wan 0001, Zhen Lei 0001, Ruicong Zhi, Guodong Guo, Stan Z. Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | Automated Depression Diagnosis Based on Deep Networks to Encode Facial Appearance and DynamicsabstractAs a severe psychiatric disorder disease, depression is a state of low mood and aversion to activity, which prevents a person from functioning normally in both work and daily lives. The study on automated mental health assessment has been given increasing attentions in recent years. In this paper, we study the problem of automatic diagnosis of depression. A new approach to predict the Beck Depression Inventory II (BDI-II) values from video data is proposed based on the deep networks. The proposed framework is designed in a two stream manner, aiming at capturing both the facial appearance and dynamics. Further, we employ joint tuning layers that can implicitly integrate the appearance and dynamic information. Experiments are conducted on two depression databases, AVEC2013 and AVEC2014. The experimental results show that our proposed approach significantly improve the depression prediction performance, compared to other visual-based approaches. Yu Zhu 0006, Zhuhong Shao, Guodong Guo |
IEEE Trans. Affect. Comput. | 4 |
| 2018 | Auxiliary Demographic Information Assisted Age Estimation With Cascaded StructureabstractOwing to the variations including both intrinsic and extrinsic factors, age estimation remains a challenging problem. In this paper, five cascaded structure frameworks are proposed for age estimation based on convolutional neural networks. All frameworks are learned and guided by auxiliary demographic information, since other demographic information (i.e., gender and race) is beneficial for age prediction. Each cascaded structure framework is embodied in a parent network and several subnetworks. For example, one of the applied framework is a gender classifier trained by gender information, and then two subnetworks are trained by the male and female samples, respectively. Furthermore, we use the features extracted from the cascaded structure frameworks with Gaussian process regression that can boost the performance further for age estimation. Experimental results on the MORPH II and CACD datasets have gained superior performances compared to the state-of-the-art methods. The mean absolute error is significantly reduced from 3.63 to 2.93 years under the same test protocol on the MORPH II dataset. Jun Wan 0001, Zichang Tan, Zhen Lei 0001, Guodong Guo, Stan Z. Li |
IEEE Trans. Cybern. | 4 |
| 2017 | Unconstrained Face Detection and Open-Set Face Recognition ChallengeabstractFace detection and recognition benchmarks have shifted toward more difficult environments. The challenge presented in this paper addresses the next step in the direction of automatic detection and identification of people from outdoor surveillance cameras. While face detection has shown remarkable success in images collected from the web, surveillance cameras include more diverse occlusions, poses, weather conditions and image blur. Although face verification or closed-set face identification have surpassed human capabilities on some datasets, open-set identification is much more complex as it needs to reject both unknown identities and false accepts from the face detector. We show that unconstrained face detection can approach high detection rates albeit with moderate false accept rates. By contrast, open-set face recognition is currently weak and requires much more attention. Manuel Günther, Peiyun Hu, Christian Herrmann 0001, Chi-Ho Chan, Min Jiang 0003, Shufan Yang, Akshay Raj Dhamija, Deva Ramanan, Jürgen Beyerer, Josef Kittler, Mohamad Al Jazaery, Mohammad Iqbal Nouyed, Guodong Guo, Cezary Stankiewicz, Terrance E. Boult |
IJCB | 13 |
| 2017 | Fast single image dehazing based on a regression model
Zhong Luan, Xiuzhuang Zhou, Zhuhong Shao, Guodong Guo, Xiaoming Liu 0002 |
Neurocomputing | 5 |
| 2017 | Leveraging multiple cues for recognizing family photos
Xiaolong Wang 0006, Guodong Guo, Michele Merler, Noel Codella, M. V. Rohith, John R. Smith, Chandra Kambhamettu |
Image Vis. Comput. | 2 |
| 2016 | Sky detection by effective context inference
Zhong Luan, Xiuzhuang Zhou, Guodong Guo |
Neurocomputing | 5 |
| 2016 | Explore Efficient Local Features from RGB-D Data for One-Shot Learning Gesture RecognitionabstractAvailability of handy RGB-D sensors has brought about a surge of gesture recognition research and applications. Among various approaches, one shot learning approach is advantageous because it requires minimum amount of data. Here, we provide a thorough review about one-shot learning gesture recognition from RGB-D data and propose a novel spatiotemporal feature extracted from RGB-D data, namely mixed features around sparse keypoints (MFSK). In the review, we analyze the challenges that we are facing, and point out some future research directions which may enlighten researchers in this field. The proposed MFSK feature is robust and invariant to scale, rotation and partial occlusions. To alleviate the insufficiency of one shot training samples, we augment the training samples by artificially synthesizing versions of various temporal scales, which is beneficial for coping with gestures performed at varying speed. We evaluate the proposed method on the Chalearn gesture dataset (CGD). The results show that our approach outperforms all currently published approaches on the challenging data of CGD, such as translated, scaled and occluded subsets. When applied to the RGB-D datasets that are not one-shot (e.g., the Cornell Activity Dataset-60 and MSR Daily Activity 3D dataset), the proposed feature also produces very promising results under leave-one-out cross validation or one-shot learning. Jun Wan 0001, Guodong Guo, Stan Z. Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Robust watermarking using orthogonal Fourier-Mellin moments and chaotic map for double images
Zhuhong Shao, Xilin Liu 0003, Guodong Guo |
Signal Process. | 5 |
| 2016 | Still-to-Video Face Matching Using Multiple Geodesic FlowsabstractStill-to-video (S2V) face recognition has recently attracted attention from researchers because of its great applications in real-world scenarios. In S2V FR, still images are usually of high quality, captured from cooperative users under controlled environment, such as mugshots, while video clips may be acquired with low resolutions and low quality, from non-cooperative users under uncontrolled environment. Because of those significant differences, we interpret the S2V FR as a heterogeneous matching problem, and propose an approach aiming at building multiple “bridges” between those two heterogeneous face modalities. Considering the unbalanced distributions and large diversities between two modalities, we propose to exploit a Grassmann manifold learning method to construct subspaces in between to find connections (or transitions) between the still images and video clips. Multiple geodesic flows are generated connecting the subspace of still images and the clustered subspace centers of videos, which are representative and robust to characterize the relationship between still images and video frames. Extensive experiments are conducted on two large scale benchmark databases, COX-S2V and PaSC, with different recognition tasks: face identification and verification. The experimental results show that the proposed approach outperforms the state-of-the-art methods under the same experimental settings. Yu Zhu 0006, Yan Li 0014, Guowang Mu, Shiguang Shan, Guodong Guo |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2015 | TriViews: A general framework to use 3D depth data effectively for action recognition
Wenbin Chen 0001, Guodong Guo |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Automated Depression Diagnosis Based on Facial Dynamic Analysis and Sparse CodingabstractDepression is a severe psychiatric disorder preventing a person from functioning normally in both work and daily lives. Currently, diagnosis of depression requires extensive participation from clinical experts. It has drawn much attention to develop an automatic system for efficient and reliable diagnosis of depression. Under the influence of depression, visual-based behavior disorder is readily observable. This paper presents a novel method of exploring facial region visual-based nonverbal behavior analysis for automatic depression diagnosis. Dynamic feature descriptors are extracted from facial region subvolumes, and sparse coding is employed to implicitly organize the extracted feature descriptors for depression diagnosis. Discriminative mapping and decision fusion are applied to further improve the accuracy of visual-based diagnosis. The integrated approach has been tested on the AVEC2013 depression database and the best visual-based mean absolute error/root mean square error results have been achieved. Lingyun Wen, Xin Li 0005, Guodong Guo, Yu Zhu 0006 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2015 | Fusing Multiple Features for Depth-Based Action RecognitionabstractHuman action recognition is a very active research topic in computer vision and pattern recognition. Recently, it has shown a great potential for human action recognition using the three-dimensional (3D) depth data captured by the emerging RGB-D sensors. Several features and/or algorithms have been proposed for depth-based action recognition. A question is raised: Can we find some complementary features and combine them to improve the accuracy significantly for depth-based action recognition? To address the question and have a better understanding of the problem, we study the fusion of different features for depth-based action recognition. Although data fusion has shown great success in other areas, it has not been well studied yet on 3D action recognition. Some issues need to be addressed, for example, whether the fusion is helpful or not for depth-based action recognition, and how to do the fusion properly. In this article, we study different fusion schemes comprehensively, using diverse features for action characterization in depth videos. Two different levels of fusion schemes are investigated, that is, feature level and decision level. Various methods are explored at each fusion level. Four different features are considered to characterize the depth action patterns from different aspects. The experiments are conducted on four challenging depth action databases, in order to evaluate and find the best fusion methods generally. Our experimental results show that the four different features investigated in the article can complement each other, and appropriate fusion methods can improve the recognition accuracies significantly over each individual feature. More importantly, our fusion-based action recognition outperforms the state-of-the-art approaches on these challenging databases. Yu Zhu 0006, Wenbin Chen 0001, Guodong Guo |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2014 | A Study on Cross-Population Age EstimationabstractWe study the problem of cross-population age estimation. Human aging is determined by the genes and influenced by many factors. Different populations, e.g., males and females, Caucasian and Asian, may age differently. Previous research has discovered the aging difference among different populations, and reported large errors in age estimation when crossing gender and/or ethnicity. In this paper we propose novel methods for cross-population age estimation with a good performance. The proposed methods are based on projecting the different aging patterns into a common space where the aging patterns can be correlated even though they come from different populations. The projections are also discriminative between age classes due to the integration of the classical discriminant analysis technique. Further, we study the amount of data needed in the target population to learn a cross-population age estimator. Finally, we study the feasibility of multi-source cross-population age estimation. Experiments are conducted on a large database of more than 21, 000 face images selected from the MORPH. Our studies are valuable to significantly reduce the burden of training data collection for age estimation on a new population, utilizing existing aging patterns even from different populations. Guodong Guo |
CVPR | 1 |
| 2014 | A study on the influence of body weight changes on face recognitionabstractOverweight and obesity is quite common in the modern society, which can result in many severe health problems. Thus weight loss has become a major event for many people to have a healthy living. A question is then raised for Biometrics or identity management: Is there any influence on face recognition when the facial shapes are varied, caused by body weight changes? No previous research has addressed this issue, to the best of our knowledge. In this paper, we study the influence of body weight changes on face recognition. Both synthesized and real face images are assembled as the databases to facilitate our study. Empirically, we found that large body weight alterations can significantly reduce the matching accuracy of the face recognition system. This is a new exploration to the biometrics society. Then we study if the influence of weight changes can be reduced to improve the face recognition performance. The partial least squares (PLS) method is applied for this purpose. Our preliminary results show that it is feasible to develop algorithms to address the influence of facial adiposity variation on face recognition, caused by weight changes. Lingyun Wen, Guodong Guo, Xin Li 0005 |
IJCB | 2 |
| 2014 | A framework for joint estimation of age, gender and ethnicity on a large database
Guodong Guo, Guowang Mu |
Image Vis. Comput. | 1 |
| 2014 | Evaluating spatiotemporal interest point features for depth-based action recognition
Yu Zhu 0006, Wenbin Chen 0001, Guodong Guo |
Image Vis. Comput. | 3 |
| 2014 | A survey on still image based human action recognition
Guodong Guo, Alice Lai |
Pattern Recognit. | 1 |
| 2014 | Optimal local community detection in social networks based on density drop of subgraphs
Xingqin Qi, Wenliang Tang, Yezhou Wu, Guodong Guo, Eddie Fuller, Cun-Quan Zhang |
Pattern Recognit. Lett. | 4 |
| 2014 | Relating Diagnosability, Strong Diagnosability and Conditional Diagnosability of Strong NetworksabstractAn interconnection network’s diagnosability is an important measure of its self-diagnostic capability. Based on the classical notion of diagnosability, strong diagnosability and conditional diagnosability were proposed later to better reflect the networks’ self-diagnostic capability under more realistic assumptions. In this paper, we study a class of interconnection networks called strong networks, which are$n$-regular,$(n - 1)$-connected, and with$cn$-number no more than$n - 3$. We build a relationship among the three diagnosability measures for strong networks. Under both PMC and${\rm MM}^{\ast}$models, given a strong network$G$with diagnosability$t$, we prove that$G$is strongly$t$-diagnosable if and only if$G$’s conditional diagnosability is greater than$t$. A simple check can show that almost all well-known regular interconnection networks are strong networks. The significance of this paper’s result is that it reveals an important relationship between strong and conditional diagnosabilities, and the proof of strong diagnosability for many interconnection networks under${\rm MM}^{\ast}$or PMC model is not necessary if their conditional diagnosability can be shown to be strictly larger than their diagnosability. Qiang Zhu 0003, Guodong Guo, Dajin Wang |
IEEE Trans. Computers | 2 |
| 2014 | Face Authentication With Makeup ChangesabstractRecent studies have shown that facial cosmetics have an impact on face recognition. To develop a face recognition system that is robust to facial makeup, we propose performing correlation mapping between makeup and nonmakeup faces on features extracted from local patches. Three methods are explored to learn the correlations. We also study the problem of makeup detection. Four categories of features are proposed to characterize cosmetics, including skin color tone, skin smoothness, texture, and highlight. A patch selection scheme and discriminative mapping are presented to enhance the performance of makeup detection. A complete system is then developed for face verification utilizing the makeup detection result. Experimental results show that our system is robust to cosmetics in face authentication. An accuracy of about 80.0% can be achieved on a database of about 500 pairs of makeup and nonmakeup face images. Guodong Guo, Lingyun Wen, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | A New Approach for 2D-3D Heterogeneous Face RecognitionabstractThis paper proposes a novel scheme for face recognition from visible images to depth images. In our proposed technique, we adopt Partial Least Square (PLS) to handle correlation mapping between 2D to 3D. A considerable performance improvement is observed compared to using Canonical Correlation Analysis (CCA). To further improve the performance, a fusion scheme based on PLS and CCA is advocated. We evaluate the advocated approach on a popular face dataset-FRGCV2.0. Experimental results demonstrate that the proposed scheme is an effective approach to perform 2D-3D face recognition. Xiaolong Wang 0006, Vincent Ly, Guodong Guo, Chandra Kambhamettu |
ISM | 3 |
| 2013 | A computational approach to body mass index prediction from face images
Lingyun Wen, Guodong Guo |
Image Vis. Comput. | 2 |
| 2013 | A Study on Visible to Infrared Action RecognitionabstractHuman action recognition is important in image and video processing with many applications. With the development of sensor technology, different cameras can be used for action acquisition, e.g., infrared cameras. Is it possible to adapt the visible light action recognizers to a new modality or domain? In this paper, we study the feasibility to adapt the action recognizer learned from visible light spectrum to infrared. A preliminary result is obtained on a large database based on an adaptive learning method, demonstrating the potential to perform cross-spectral action recognition. Yu Zhu 0006, Guodong Guo |
IEEE Signal Process. Lett. | 2 |
| 2013 | Facial Expression Recognition Influenced by Human AgingabstractFacial expression recognition (FER) is an active research topic in computer vision. However, there is no study yet to discover whether FER is affected by human aging, from a computational perspective. We perform a computational study of FER within and across age groups and compare the FER accuracies. Two databases from the psychology society are introduced to the computer vision community and used for our study. We found that the FER is influenced significantly by human aging, and we analyze the influence and interpret it from a computational viewpoint. Next, we propose some schemes to reduce the influence of aging on FER and evaluate the effectiveness in dealing with lifespan FER. Guodong Guo, Xin Li 0005 |
IEEE Trans. Affect. Comput. | 1 |
| 2012 | A New Projection Space for Separation of Specular-Diffuse Reflection Components in Color Images
Zhaowei Cai, Longyin Wen, Zhen Lei 0001, Guodong Guo, Stan Z. Li |
ACCV (4) | 5 |
| 2012 | A study on human age estimation under facial expression changesabstractIn this paper, we study human age estimation in face images under significant expression changes. We will address two issues: (1) Is age estimation affected by facial expression changes and how significant is the influence? (2) How to develop a robust method to perform age estimation undergoing various facial expression changes? This systematic study will not only discover the relation between age estimation and expression changes, but also contribute a robust solution to solve the problem of cross-expression age estimation. This study is an important step towards developing a practical and robust age estimation system that allows users to present their faces naturally (with various expressions) rather than constrained to the neutral expression only. Two databases originally captured in the Psychology community are introduced to Computer Vision, to quantitatively demonstrate the influence of expression changes on age estimation, and evaluate the proposed framework and corresponding methods for cross-expression age estimation. Guodong Guo, Xiaolong Wang 0006 |
CVPR | 1 |
| 2012 | Automatic Classification of Teeth in Bitewing Dental Images Using OLPPabstractTeeth classification is an important component in building an Automated Dental Identification System (ADIS) as part of creating a data structure that guides tooth-to-tooth matching. This aids in avoiding illogical comparisons that both inefficiently consume the limited computational resources and mislead decision-making. We tackle this problem by using low computational-cost, appearance-based Orthogonal Locality Preserving Projection (OLPP) algorithm to assign an initial class, i.e. molar or premolar to the teeth in bitewing dental images. After this initial classification, we use a string matching technique, based on teeth neighborhood rules, to validate initial teeth-classes and thus assign each tooth a number corresponding to its location in the dental chart. On a large dataset of bitewing films that contain 622 teeth, the proposed approach achieves classification accuracy of 89% and teeth class validation enhances the overall teeth classification accuracy to 92%. Nourdin Al-sherif, Guodong Guo, Hany H. Ammar |
ISM | 2 |
| 2012 | A New Approach to Teeth SegmentationabstractTeeth segmentation is one of the important components in building an Automated Dental Identification System (ADIS). The extraction of the teeth from their corresponding dental radiographs is called teeth segmentation. Dental radiographs may suffer from poor teeth image quality, low contrast and uneven exposure that complicate the task of teeth segmentation. To achieve a good performance in segmentation, the teeth images are preprocessed by a two-step thresholding technique, which starts with an iterative thresholding followed by an adaptive thresholding to binarize the teeth images. Then, we propose to adapt the seam carving technique on the binary images, using both horizontal and vertical seams, to separate each individual tooth. The proposed method is evaluated experimentally and compared to other algorithms. The results show that our new approach achieves the lowest failure rate among all existing methods, and the highest optimality among all of the fully automated approaches reported in the literature. Nourdin Al-sherif, Guodong Guo, Hany H. Ammar |
ISM | 2 |
| 2011 | Simultaneous dimensionality reduction and human age estimation via kernel partial least squares regressionabstractHuman age estimation has recently become an active research topic in computer vision and pattern recognition, because of many potential applications in reality. In this paper we propose to use the kernel partial least squares (KPLS) regression for age estimation. The KPLS (or linear PLS) method has several advantages over previous approaches: (1) the KPLS can reduce feature dimensionality and learn the aging function simultaneously in a single learning framework, instead of performing each task separately using different techniques; (2) the KPLS can find a small number of latent variables, e.g., 20, to project thousands of features into a very low-dimensional subspace, which may have great impact on real-time applications; and (3) the KPLS regression has an output vector that can contain multiple labels, so that several related problems, e.g., age estimation, gender classification, and ethnicity estimation can be solved altogether. This is the first time that the kernel PLS method is introduced and applied to solve a regression problem in computer vision with high accuracy. Experimental results on a very large database show that the KPLS is significantly better than the popular SVM method, and outperform the state-of-the-art approaches in human age estimation. Guodong Guo, Guowang Mu |
CVPR | 1 |
| 2011 | Digital anti-aging in face imagesabstractWe study a problem called digital facial anti-aging, which aims at making human faces look younger in digital photos. This novel problem is different from the traditional age synthesis where new faces are synthesized at older ages using other example face images. In contrast, our facial anti-aging works on a single digital photo without using any other faces. The proposed system contains several modules. First, an input color face image is decomposed into three layers: face structure, aging detail, and color. Second, the specular highlight in the original face image is recovered. Third, the face structure, color, and the recovered specular highlight are combined to form an anti-aging face image. Further, the system can deal with hair coloring and eyebrow change if needed, since these factors influence human judges of ages. Our approach keeps facial identities and delivers high-quality outputs. We show that the anti-aging system can be built based on adapting and integrating the state-of-the-art methods that were originally proposed to solve other problems. The anti-aging faces are evaluated by humans to demonstrate the performance. Guodong Guo |
ICCV | 1 |
| 2010 | Cross-Age Face Recognition on a Very Large Database: The Performance versus Age Intervals and Improvement Using Soft Biometric TraitsabstractFacial aging can degrade the face recognition performance dramatically. Traditional face recognition studies focus on dealing with pose, illumination, and expression (PIE) changes. Considering a large span of age difference, the influence of facial aging could be very significant compared to the PIE variations. How big the aging influence could be? What is the relation between recognition accuracy and age intervals? Can soft biometrics be used to improve the face recognition performance under age variations? In this paper we address all these issues. First, we investigate the face recognition performance degradation with respect to age intervals between the probe and gallery images on a very large database which contains about 55,000 face images of more than 13,000 individuals. Second, we study if soft biometric traits, e.g., race, gender, height, and weight, could be used to improve the cross-age face recognition accuracies, and how useful each of them could be. Guodong Guo, Guowang Mu, Karl Ricanek |
ICPR | 1 |
| 2010 | Age Synthesis and Estimation via Faces: A SurveyabstractHuman age, as an important personal trait, can be directly inferred by distinct patterns emerging from the facial appearance. Derived from rapid advances in computer graphics and machine vision, computer-based age synthesis and estimation via faces have become particularly prevalent topics recently because of their explosively emerging real-world applications, such as forensic art, electronic customer relationship management, security control and surveillance monitoring, biometrics, entertainment, and cosmetology. Age synthesis is defined to rerender a face image aesthetically with natural aging and rejuvenating effects on the individual face. Age estimation is defined to label a face image automatically with the exact age (year) or the age group (year range) of the individual face. Because of their particularity and complexity, both problems are attractive yet challenging to computer-based application system designers. Large efforts from both academia and industry have been devoted in the last a few decades. In this paper, we survey the complete state-of-the-art techniques in the face image-based age synthesis and estimation topics. Existing models, popular algorithms, system performances, technical difficulties, popular face aging databases, evaluation protocols, and promising future directions are also provided with systematic discussions. Yun Fu 0001, Guodong Guo, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Gender from Body: A Biologically-Inspired Approach with Manifold Learning
Guodong Guo, Guowang Mu, Yun Fu 0001 |
ACCV (3) | 1 |
| 2009 | Human age estimation using bio-inspired featuresabstractWe investigate the biologically inspired features (BIF) for human age estimation from faces. As in previous bio-inspired models, a pyramid of Gabor filters are used at all positions of the input image for the S1units. But unlike previous models, we find that the pre-learned prototypes for the S2layer and then progressing to C2cannot work well for age estimation. We also propose to use Gabor filters with smaller sizes and suggest to determine the number of bands and orientations in a problem-specific manner, rather than using a predefined number. More importantly, we propose a new operator “STD” to encode the aging subtlety on faces. Evaluated on the large database YGA with 8,000 face images and the public available FG-NET database, our approach achieves significant improvements in age estimation accuracy over the state-of-the-artmethods. By applying our system to some Internet face images, we show the robustness of our method and the potential of cross-race age estimation, which has not been explored by any studies before. Guodong Guo, Guowang Mu, Yun Fu 0001, Thomas S. Huang |
CVPR | 1 |
| 2009 | A study on automatic age estimation using a large databaseabstractIn this paper we study some problems related to human age estimation using a large database. First, we study the influence of gender on age estimation based on face representations that combine biologically-inspired features with manifold learning techniques. Second, we study age estimation using smaller gender and age groups rather than on all ages. Significant error reductions are observed in both cases. Based on these results, we designed three frameworks for automatic age estimation that exhibit high performance. Unlike previous methods that require manual separation of males and females prior to age estimation, our work is the first to estimate age automatically on a large database. Furthermore, a data fusion approach is proposed using one of the frameworks, which gives an age estimation error more than 40% smaller than previous methods. Guodong Guo, Guowang Mu, Yun Fu 0001, Charles R. Dyer, Thomas S. Huang |
ICCV | 1 |
| 2008 | Head pose estimation: Classification or regression?abstractHead pose estimation has many useful applications in practice. How to estimate the head pose automatically and robustly is still a challenging problem. In pose estimation, different pose angles can be used as regression values or viewed as different class labels. Thus a question is raised in our study: which is proper for pose estimation - classification or regression? We investigate representative classification and regression methods on the same problem to see any difference. A method that combines regression and classification approaches is also examined. Preliminary experiments show some interesting results which might prompt further exploration of related issues in pose estimation. Guodong Guo, Yun Fu 0001, Charles R. Dyer, Thomas S. Huang |
ICPR | 1 |
| 2008 | Locally Adjusted Robust Regression for Human Age EstimationabstractAutomatic human age estimation has considerable potential applications in human computer interaction and multimedia communication. However, the age estimation problem is challenging. We design a locally adjusted robust regressor (LARR) for learning and prediction of human ages. The novel approach reduces the age estimation errors significantly over all previous methods. Experiments on two aging databases show the success of the proposed method for human aging estimation. Guodong Guo, Yun Fu 0001, Thomas S. Huang, Charles R. Dyer |
WACV | 1 |
| 2008 | Iris Extraction Based on Intensity Gradient and Texture DifferenceabstractBiometrics has become more and more important in security applications. In comparison with many other biometric features, iris recognition has very high recognition accuracy. Successful iris recognition depends largely on correct iris localization, however, the performance of current techniques for iris localization still leaves room for improvement. To improve the iris localization performance, we propose a novel method that optimally utilizes both the intensity gradient and texture difference. Experimental results demonstrate that our new approach gives much better results than previous approaches. In order to make the iris boundary more accurate, we present a new issue called model selection and propose a method to choose between ellipse/circle and circle/circle models. Furthermore, we propose a dome model to compute mask images and remove eyelid occlusions in the unwrapped images rather than in the original eye images with a least commitment strategy. Guodong Guo, Michael J. Jones 0001 |
WACV | 1 |
| 2008 | Image-Based Human Age Estimation by Manifold Learning and Locally Adjusted Robust RegressionabstractEstimating human age automatically via facial image analysis has lots of potential real-world applications, such as human computer interaction and multimedia communication. However, it is still a challenging problem for the existing computer vision systems to automatically and effectively estimate human ages. The aging process is determined by not only the person's gene, but also many external factors, such as health, living style, living location, and weather conditions. Males and females may also age differently. The current age estimation performance is still not good enough for practical use and more effort has to be put into this research direction. In this paper, we introduce the age manifold learning scheme for extracting face aging features and design a locally adjusted robust regressor for learning and prediction of human ages. The novel approach improves the age estimation accuracy significantly over all previous methods. The merit of the proposed approaches for image-based age estimation is shown by extensive experiments on a large internal age database and the public available FG-NET database. Guodong Guo, Yun Fu 0001, Charles R. Dyer, Thomas S. Huang |
IEEE Trans. Image Process. | 1 |
| 2007 | Patch-based Image Correlation with Rapid FilteringabstractThis paper describes a patch-based approach for rapid image correlation or template matching. By representing a template image with an ensemble of patches, the method is robust with respect to variations such as local appearance variation, partial occlusion, and scale changes. Rectangle filters are applied to each image patch for fast filtering based on the integral image representation. A new method is developed for feature dimension reduction by detecting the "salient" image structures given a single image. Experiments on a variety images show the success of the method in dealing with different variations in the test images. In terms of computation time, the approach is faster than traditional methods by up to two orders of magnitude and is at least three times faster than a fast implementation of normalized cross correlation. Guodong Guo, Charles R. Dyer |
CVPR | 1 |
| 2005 | Linear Combination Representation for Outlier Detection in Motion TrackingabstractIn this paper we show that Ullman and Basri's linear combination (LC) representation, which was originally proposed for alignment-based object recognition, can be used for outlier detection in motion tracking with an affine camera. For this task LC can be realized either on image frames or feature trajectories, and therefore two methods are developed which we call linear combination of frames and linear combination of trajectories. For robust estimation of the linear combination coefficients, the support vector regression (SVR) algorithm is used and compared with the RANSAC method. SVR based on quadratic programming optimization can efficiently deal with more than 50 percent outliers and delivers more consistent results than RANSAC in our experiments. The linear combination representation can use SVR in a straightforward manner while previous factorization-based or subspace separation methods cannot. Experimental results are presented using real video sequences to demonstrate the effectiveness of our LC+SVR approaches, including a quantitative comparison of SVR and RANSAC. Guodong Guo, Charles R. Dyer, Zhengyou Zhang |
CVPR (2) | 1 |
| 2005 | Learning from examples in the small sample case: face expression recognitionabstractExample-based learning for computer vision can be difficult when a large number of examples to represent each pattern or object class is not available. In such situations, learning from a small number of samples is of practical value. To study this issue, the task of face expression recognition with a small number of training images of each expression is considered. A new technique based on linear programming for both feature selection and classifier training is introduced. A pairwise framework for feature selection, instead of using all classes simultaneously, is presented. Experimental results compare the method with three others: a simplified Bayes classifier, support vector machine, and AdaBoost. Finally, each algorithm is analyzed and a new categorization of these algorithms is given, especially for learning from examples in the small sample case. Guodong Guo, Charles R. Dyer |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2003 | Simultaneous Feature Selection and Classifier Training via Linear Programming: A Case Study for Face Expression RecognitionabstractA linear programming technique is introduced that jointly performs feature selection and classifier training so that a subset of features is optimally selected together with the classifier. Because traditional classification methods in computer vision have used a two-step approach: feature selection followed by classifier training, feature selection has often been ad hoc using heuristics or requiring a time-consuming forward and backward search process. Moreover, it is difficult to determine which features to use and how many features to use when these two steps are separated. The linear programming technique used in this paper, which we call feature selection via linear programming (FSLP), can determine the number of features and which features to use in the resulting classification function based on recent results in optimization. We analyze why FSLP can avoid the curse of dimensionality problem based on margin analysis. As one demonstration of the performance of this FSLP technique for computer vision tasks, we apply it to the problem of face expression recognition. Recognition accuracy is compared with results using support vector machines, the AdaBoost algorithm, and a Bayes classifier. Guodong Guo, Charles R. Dyer |
CVPR (1) | 1 |
| 2003 | Content-based audio classification and retrieval by support vector machinesabstractSupport vector machines (SVMs) have been recently proposed as a new learning algorithm for pattern recognition. In this paper, the SVMs with a binary tree recognition strategy are used to tackle the audio classification problem. We illustrate the potential of SVMs on a common audio database, which consists of 409 sounds of 16 classes. We compare the SVMs based classification with other popular approaches. For audio retrieval, we propose a new metric, called distance-from-boundary (DFB). When a query audio is given, the system first finds a boundary inside which the query pattern is located. Then, all the audio patterns in the database are sorted by their distances to this boundary. All boundaries are learned by the SVMs and stored together with the audio database. Experimental comparisons for audio retrieval are presented to show the superiority of this novel metric to other similarity measures. Guodong Guo, Stan Z. Li |
IEEE Trans. Neural Networks | 1 |
| 2002 | Learning similarity measure for natural image retrieval with relevance feedbackabstractA new scheme of learning similarity measure is proposed for content-based image retrieval (CBIR). It learns a boundary that separates the images in the database into two clusters. Images inside the boundary are ranked by their Euclidean distances to the query. The scheme is called constrained similarity measure (CSM), which not only takes into consideration the perceptual similarity between images, but also significantly improves the retrieval performance of the Euclidean distance measure. Two techniques, support vector machine (SVM) and AdaBoost from machine learning, are utilized to learn the boundary. They are compared to see their differences in boundary learning. The positive and negative examples used to learn the boundary are provided by the user with relevance feedback. The CSM metric is evaluated in a large database of 10009 natural images with an accurate ground truth. Experimental results demonstrate the usefulness and effectiveness of the proposed similarity measure for image retrieval. Guodong Guo, Anil K. Jain 0001, Wei-Ying Ma, HongJiang Zhang |
IEEE Trans. Neural Networks | 1 |
| 2001 | Learning Similarity Measure for Natural Image Retrieval with Relevance FeedbackabstractA new scheme of learning similarity measure is proposed for content-based image retrieval (CBIR). It learns a boundary that separates the images in the database into two parts. Images on the positive side of the boundary are ranked by their Euclidean distances to the query. The scheme is called restricted similarity measure (RSM), which not only takes into consideration the perceptual similarity between images, but also significantly improves the retrieval performance based on the Euclidean distance measure. Two techniques, support vector machine and AdaBoost, are utilized to learn the boundary, and compared with respect to their performance in boundary learning. The positive and negative examples used to learn the boundary are provided by the user with relevance feedback. The RSM metric is evaluated on a large database of 10,009 natural images with an accurate ground truth. Experimental results demonstrate the usefulness and effectiveness of the proposed similarity measure for image retrieval. Guodong Guo, Anil K. Jain 0001, Wei-Ying Ma, HongJiang Zhang |
CVPR (1) | 1 |
| 2001 | Distance-from-boundary as a metric for texture image retrievalabstractA new metric is proposed for texture image retrieval, which is based on the signed distance of the images in the database to a boundary chosen by the query. This novel metric has three advantages: (1) the boundary distance measures are relatively insensitive to the sample distributions; (2) the same retrieval results can be obtained with respect to different (but visually similar) queries; (3) retrieval performance can be improved. The boundaries are obtained by using a statistical learning algorithm called support vector machine (SVM), and hence the boundaries can be simply represented by some vectors and their combination coefficients. Experimental results on the Brodatz texture database indicate that a significantly better retrieval performance can be achieved as compared to the traditional Euclidean distance-based approach. This technique can be further developed to learn pattern similarities among different texture classes and used in relevance feedback. Guodong Guo, HongJiang Zhang, Stan Z. Li |
ICASSP | 1 |
| 2001 | Pairwise Face RecognitionabstractWe develop a pairwise classification framework for face recognition, in which a C class face recognition problem is divided into a set of C(C-1)/2 two class problems. Such a problem decomposition not only leads to a set of simpler classification problems to be solved, thereby increasing overall classification accuracy, but also provides a framework for independent feature selection for each pair of classes. A simple feature ranking strategy is used to select a small subset of the features for each pair of classes. Furthermore, we evaluate two classification methods under the pairwise comparison framework: the Bayes classifier and the AdaBoost. Experiments on a large face database with 1079 face images of 137 individuals indicate that 20 features are enough to achieve a relatively high recognition accuracy, which demonstrates the effectiveness of the pairwise recognition framework. Guodong Guo, HongJiang Zhang, Stan Z. Li |
ICCV | 1 |
| 2001 | Boosting For Content-Based Audio Classification And Retrieval: An EvaluationabstractIn this paper, we evaluate a recently proposed algorithm in machine learning called AdaBoost for content-based audio classification and retrieval. AdaBoost is a kind of large margin classifiers and is efficient for on-line learning. Our focus is to evaluate its classification and retrieval accuracy as compared with other methods. The Muscle Fish audio database of 409 sounds is used for the evaluation with perceptual and cepstral features. 1. Guodong Guo, HongJiang Zhang, Stan Z. Li |
ICME | 1 |
| 2001 | Support vector machines for face recognition
Guodong Guo, Stan Z. Li, Kap Luk Chan |
Image Vis. Comput. | 1 |
| 2000 | Learning Similarity for Texture Image Retrieval
Guodong Guo, Stan Z. Li, Kap Luk Chan |
ECCV (1) | 1 |
| 2000 | Face Recognition by Support Vector MachinesabstractSupport vector machines (SVM) have been recently proposed as a new technique for pattern recognition. SVM with a binary tree recognition strategy are used to tackle the face recognition problem. We illustrate the potential of SVM on the Cambridge ORL face database, which consists of 400 images of 40 individuals, containing quite a high degree of variability in expression, pose, and facial details. We also present the recognition experiment on a larger face database of 1079 images of 137 individuals. We compare the SVM-based recognition with the standard eigenface approach using the nearest center classification (NCC) criterion. Guodong Guo, Stan Z. Li, Kap Luk Chan |
FG | 1 |
| 2000 | Bayesian learning, global competition and unsupervised image segmentation
Guodong Guo, Songde Ma |
Pattern Recognit. Lett. | 1 |
| 1998 | Unsupervised Segmentation of Color Images
Guodong Guo, Songde Ma |
ICIP (3) | 1 |
| 1998 | Unsupervised segmentation based on multi-resolution analysis, robust statistics and majority game theoryabstractAn unsupervised model-based image segmentation technique requires the model parameters for the various image classes in an observed image to be estimated directly from the image. The accuracy of the segmentation depends on the correct estimation of the parameters, as well as on the correct labeling of the pixels. In this work, the parameters are estimated by a multiresolution analysis on the histogram and a robust estimator using least median of squares. The labeling process is based on majority game theory. The method is tested in various synthetic and real images, showing its effectiveness. Guodong Guo, Songde Ma |
ICPR | 1 |