Di Huang 0001

dblp:45/780-1 · DBLP profile ↗
← Back
189ranked-venue papers
13as first author
108since 2021 · last 2026
0000-0002-2412-9330ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 132 · 7 first-author · 70 since 2021Artificial intelligence and machine learning · 118 · 7 first-author · 74 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 first-author · 2 since 2021Security and privacy · 10 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
YearPublicationVenuePosition
2026 RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
abstract
Pre-trained Vision-Language Models (VLMs), e.g. CLIP, have become essential tools in multimodal transfer learning. However, fine-tuning VLMs in few-shot scenarios poses significant challenges in balancing task-specific adaptation and generalization in the obtained model. Meanwhile, current researches have predominantly focused on prompt-based adaptation methods, leaving adapter-based approaches underexplored and revealing notable performance gaps. To address these challenges, we introduce a novel Reconstruction-based Multimodal Adapter (RMAdapter), which leverages a dual-branch architecture. Unlike conventional single-branch adapters, RMAdapter consists of: (1) an adaptation branch that injects task-specific knowledge through parameter-efficient fine-tuning, and (2) a reconstruction branch that preserves general knowledge by reconstructing latent space features back into the original feature space. This design facilitates a dynamic balance between general and task-specific knowledge. Importantly, although RMAdapter introduces an additional reconstruction branch, it is carefully optimized to remain lightweight. By computing reconstruction loss locally at each layer and sharing projection modules, the overall computational overhead is kept minimal. A consistency constraint is also incorporated to better regulate the trade-off between discriminability and generalization. We comprehensively evaluate the effectiveness of RMAdapter on three representative tasks: generalization to new categories, generalization to new target datasets, and domain generalization. Without relying on data augmentation or duplicate prompt designs, our RMAdapter consistently outperforms state-of-the-art approaches across all evaluation metrics.
Weixin Li 0001, Di Huang 0001
AAAI5
2026 SceneGenesis: 3D Scene Synthesis via Semantic Structural Priors and Mesh-Guided Video-Geometry Fusion
abstract
Generating high-quality, controllable, and structurally consistent 3D scenes in complex multi-object environments remains a fundamental challenge. We present SceneGenesis, a unified framework that synthesizes 3D scenes by combining semantic structural priors with mesh-guided video–geometry fusion. SceneGenesis first employs large language models to convert textual descriptions into category-aware object specifications, which are transformed into structured meshes using procedural approximations and pretrained asset generators, enabling precise layout control and scalable scene construction. To obtain rich and style-controllable appearances, SceneGenesis generates multi-view video representations conditioned on the initialized structure. A mesh-guided video–geometry fusion module then consolidates video evidence with mesh priors through mesh-conditioned fragment initialization, progressive geometric refinement, and structure-aware optimization, substantially improving global geometric fidelity and visual realism. Experiments demonstrate that SceneGenesis supports flexible style variation and object-level editing while achieving strong controllability, scalability, and structural quality.
Yueming Zhao, Hongyu Yang 0001, Di Huang 0001
AAAI3
2026 Reference patch momentum distillation for open-vocabulary semantic segmentation
Qingjie Liu 0001, Di Huang 0001
Sci. China Inf. Sci.4
2026 Label-informed knowledge integration: Advancing visual prompt for VLMs adaptation
Yunhong Wang 0001, Guodong Wang 0006, Yingjie Gao 0001, Xiuguo Bao, Di Huang 0001
Comput. Vis. Image Underst.7
2026 Learning storage-efficient 3D Gaussian head avatars from monocular videos via parametric adaptation and material decomposition
Guohao Li 0010, Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001
Frontiers Comput. Sci.3
2026 Vision-based 3D occupancy prediction in autonomous driving: a review and outlook
abstract
Abstract In recent years, autonomous driving has garnered escalating attention for its potential to relieve drivers’ burdens and improve driving safety. Vision-based 3D occupancy prediction, which predicts the spatial occupancy status and semantics of 3D voxel grids around the autonomous vehicle from image inputs, is an emerging perception task suitable for cost-effective perception system of autonomous driving. Although numerous studies have demonstrated the greater advantages of 3D occupancy prediction over object-centric perception tasks, there is still a lack of a dedicated review focusing on this rapidly developing field. In this paper, we first introduce the background of vision-based 3D occupancy prediction and discuss the challenges in this task. Second, we conduct a comprehensive survey of the progress in vision-based 3D occupancy prediction from three aspects: feature enhancement, deployment friendliness and label efficiency, and provide an in-depth analysis of the potentials and challenges of each category of methods. Finally, we present a summary of prevailing research trends and propose some inspiring future outlooks. To provide a valuable reference for researchers, a regularly updated collection of related papers, datasets, and codes is organized at github.com/zya3d/Awesome-3D-Occupancy-Prediction website.
Yanan Zhang 0005, Jinqing Zhang, Zengran Wang, Di Huang 0001
Frontiers Comput. Sci.5
2026 Towards unbiased source-free object detection via vision foundation models
Zhi Cai, Yingjie Gao 0001, Yanan Zhang 0005, Xinzhu Ma, Di Huang 0001
Pattern Recognit.5
2026 AST-Adapter: Parameter-Efficient Video-to-Video Transfer Learning With Adaptive Spatiotemporal Information Bias
abstract
Leveraging video pre-trained models for video downstream tasks has recently emerged with promising performance. Except for the full fine-tuning paradigm, parameter-efficient transfer learning (PETL) exists as a promising way and has not yet been fully explored in video-to-video transfer learning. While current PETL approaches succeed to reduce parameter quantity and computation cost, they overlook the critical spatiotemporal property in video modality. In this paper, we first propose a novel metric to quantify the spatiotemporal information bias across video datasets and uncover its impact on transfer preferences through systematic analysis. Based on the analysis, we introduce an innovative parameter-efficient transfer learning method, named Adaptive SpatioTemporal Adapter (AST-Adapter). Our approach automatically adjusts layer-wise architectures with different spatiotemporal adapter modules to exploit the intrinsics of downstream tasks to achieve adaptive spatiotemporal learning, thus delivering robustness and generalization. Extensive experiments on five datasets across action recognition and action detection task show that AST-Adapter surpasses both video-to-video and image-to-video approaches, whilst keeping the advantage of parameter efficiency. Notably, AST-Adapter achieves 89.9% on Kinectics-400, 80.3% on HMDB51 and 97.7% on UCF101 while introduces only 1% to 10% tunable parameters. Our code is available at https://github.com/hhhhhpy/AST-Adapter.
Puyue Hou, Guohao Li 0010, Zhi Cai, Di Huang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Parameter-Efficient Tuning for Fine-Grained Recognition via Channel-Wise Importance Equalization and Diversity Navigation
abstract
Parameter-efficient tuning (PET) has achieved promising performance on various downstream vision tasks. Despite their effectiveness for general classification, existing PET approaches neglect the over-concentration of channel-wise saliency and the feature redundancy of pre-trained models during fine-tuning, thus leaving much room for improvement when applied to the downstream fine-grained recognition tasks. To address these issues, we propose a novel parameter-efficient tuning approach tailored for fine-grained recognition (FG-PET). Specifically, FG-PET first employs a Channel-wise Importance Equalization (CIE) module. It suppresses the concentrated salient channels while strengthening the remaining majority ones during fine-tuning, notably mitigating the over-concentrated saliency, thus evoking more channels within pre-trained models to deliver abundant local visual clues. Furthermore, FG-PET develops an Efficient Navigator for Diversity (EFIND) by introducing a center-based loss and orthogonal constraints on features generated from distinct attention heads. It alleviates the redundancy between different attention maps, thus enforcing the models to explore diverse subtle visual differences in various discriminative local regions, which are critical for fine-grained recognition. Extensive experimental results on five public fine-grained benchmarks based on distinct ViT models demonstrate that the proposed method remarkably boosts the performance of existing PET approaches, and generalizes well to general classification tasks. The source code is available at FG-PET.
Hanwen Zhong, Jiaxin Chen 0002, Di Huang 0001, Yunhong Wang 0001
IEEE Trans. Image Process.4
2026 Is Diversity All You Need for Scalable Robotic Manipulation?
abstract
Data scaling has driven remarkable success in foundation models for Natural Language Processing (NLP) and Computer Vision (CV), yet the principles of effective data scaling in robotic manipulation remain insufficiently understood. In this work, we investigate the nuanced role of data diversity in robot learning by examining three critical dimensions-task (what to do), embodiment (which robot to use), and expert (who demonstrates)-challenging the conventional intuition of “more diverse is better”. Throughout extensive experiments on various robot platforms, we reveal that (1) task diversity proves more critical than per-task demonstration quantity, with scene diversity playing a more important role than skill diversity for robustness and generalization under distribution shifts; (2) multi-embodiment pre-training data is non-essential for cross-embodiment transfer-models trained on high-quality single-embodiment data can efficiently transfer to different platforms, showing desirable scaling property during fine-tuning and its potential of replacing large-scale multi-embodiment pre-training; and (3) expert diversity, arising from individual operational preferences and stochastic variations in human demonstrations, can be confounding to policy learning, with action rate multimodality emerging as a key contributing factor. Based on this insight, we propose a distribution debiasing method to mitigate action rate ambiguity, the yielding GO-1-Pro achieves substantial performance gains of 15%, equivalent to using 2.5× pre-training data. Collectively, these findings provide new perspectives and offer practical guidance on how to scale robotic manipulation datasets effectively. The code will be released.
Modi Shi, Li Chen 0008, Chiming Liu, Guanghui Ren, Ping Luo 0002, Di Huang 0001, Maoqing Yao, Hongyang Li 0001
IEEE Trans. Robotics8
2025 Micro-macro Wavelet-based Gaussian Splatting for 3D Reconstruction from Unconstrained Images
abstract
3D reconstruction from unconstrained image collections presents substantial challenges due to varying appearances and transient occlusions. In this paper, we introduce Micro-macro Wavelet-based Gaussian Splatting (MW-GS), a novel approach designed to enhance 3D reconstruction by disentangling scene representations into global, refined, and intrinsic components. The proposed method features two key innovations: Micro-macro Projection, which allows Gaussian points to capture details from feature maps across multiple scales with enhanced diversity; and Wavelet-based Sampling, which leverages frequency domain information to refine feature representations and significantly improve the modeling of scene appearances. Additionally, we incorporate a Hierarchical Residual Fusion Network to seamlessly integrate these features. Extensive experiments demonstrate that MW-GS delivers state-of-the-art rendering performance, surpassing existing methods.
Chengxin Lv, Hongyu Yang 0001, Di Huang 0001
AAAI4
2025 Unveiling the Knowledge of CLIP for Training-Free Open-Vocabulary Semantic Segmentation
abstract
Training-free open-vocabulary semantic segmentation aims to explore the potential of frozen vision-language models (VLM) for segmentation tasks. Recent works reform the inference process of CLIP and utilize the features from the final layer to reconstruct dense representations for segmentation, demonstrating promising performance. However, the final layer tends to prioritize global components over local representations, leading to suboptimal robustness and effectiveness of existing methods. In this paper, we propose CLIPSeg, a novel training-free framework that fully exploits the diverse knowledge across layers in CLIP for dense predictions. Our study unveils two key discoveries: Firstly, the features in the middle layers exhibit high locality awareness and feature coherence compared to the final layer, based on which we propose the coherence enhanced residual attention module that generates semantic-aware attention. Secondly, despite not being directly aligned with the text, the deep layers capture valid local semantics that complement those in the final layer. Leveraging this insight, we introduce the deep semantic integration module to boost the patch semantics in the final block. Experiments conducted on 9 segmentation benchmarks with various CLIP models demonstrate that CLIPSeg consistently outperforms all training-free methods by substantial margins, e.g., a 7.8 % improvement in average mIoU for CLIP with a ViT-L backbone, and competes with learning-based counterparts in generalizing to novel concepts in an efficient way.
Guodong Wang 0006, Qingjie Liu 0001, Di Huang 0001
AAAI5
2025 3D²-Actor: Learning Pose-Conditioned 3D-Aware Denoiser for Realistic Gaussian Avatar Modeling
abstract
Advancements in neural implicit representations and differentiable rendering have markedly improved the ability to learn animatable 3D avatars from sparse multi-view RGB videos. However, current methods that map observation space to canonical space often face challenges in capturing pose-dependent details and generalizing to novel poses. While diffusion models have demonstrated remarkable zero-shot capabilities in 2D image generation, their potential for creating animatable 3D avatars from 2D inputs remains underexplored. In this work, we introduce 3D²-Actor, a novel approach featuring a pose-conditioned 3D-aware human modeling pipeline that integrates iterative 2D denoising and 3D rectifying steps. The 2D denoiser, guided by pose cues, generates detailed multi-view images that provide the rich feature set necessary for high-fidelity 3D reconstruction and pose rendering. Complementing this, our Gaussian-based 3D rectifier renders images with enhanced 3D consistency through a two-stage projection strategy and a novel local coordinate representation. Additionally, we propose an innovative sampling strategy to ensure smooth temporal continuity across frames in video synthesis. Our method effectively addresses the limitations of traditional numerical solutions in handling ill-posed mappings, producing realistic and animatable 3D human avatars. Experimental results demonstrate that 3D²-Actor excels in high-fidelity avatar modeling and robustly generalizes to novel poses.
Zichen Tang, Hongyu Yang 0001, Hanchen Zhang, Jiaxin Chen 0002, Di Huang 0001
AAAI5
2025 APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision Transformers
abstract
Vision Transformers (ViTs) have become one of the most commonly used backbones for vision tasks. Despite their remarkable performance, they often suffer significant accuracy drops when quantized for practical deployment, particularly by post-training quantization (PTQ) under ultra-low bits. Recently, reconstruction-based PTQ methods have shown promising performance in quantizing Convolutional Neural Networks (CNNs). However, they fail when applied to ViTs, primarily due to the inaccurate estimation of output importance and the substantial accuracy degradation in quantizing post-GELU activations. To address these issues, we propose APHQ-ViT, a novel PTQ approach based on importance estimation with Average Perturbation Hessian (APH). Specifically, we first thoroughly analyze the current approximation approaches with Hessian loss, and propose an improved average perturbation Hessian loss. To deal with the quantization of the post-GELU activations, we design an MLP Reconstruction (MR) method by replacing the GELU function in MLP with ReLU and reconstructing it by the APH loss on a small unlabeled calibration set. Extensive experiments demonstrate that APHQ-ViT using linear quantizers outperforms existing PTQ methods by substantial margins in 3-bit and 4-bit across different vision tasks. The source code is available at https://github.com/GoatWu/APHQ-ViT.
Zhuguanyu Wu, Jiaxin Chen 0002, Jinyang Guo 0002, Di Huang 0001, Yunhong Wang 0001
CVPR5
2025 CoSDH: Communication-Efficient Collaborative Perception via Supply-Demand Awareness and Intermediate-Late Hybridization
abstract
Multi-Agent collaborative perception enhances perceptual capabilities by utilizing information from multiple agents and is considered a fundamental solution to the problem of weak single-vehicle perception in autonomous driving. However, existing collaborative perception methods face a dilemma between communication efficiency and perception accuracy. To address this issue, we propose a novel communication-efficient collaborative perception framework based on supply-demand awareness and intermediate-late hybridization, dubbed as CoSDH. By modeling the supply-demand relationship between agents, the framework refines the selection of collaboration regions, reducing unnecessary communication cost while maintaining accuracy. In addition, we innovatively introduce the intermediate-late hybrid collaboration mode, where late-stage collaboration compensates for the performance degradation in collaborative perception under low communication bandwidth. Extensive experiments on multiple datasets, including both simulated and real-world scenarios, demonstrate that CoSDH achieves state-of- the-art detection accuracy and optimal bandwidth tradeoffs, delivering superior detection precision under real communication bandwidths, thus proving its effectiveness and practical applicability. The code will be released at https://github.com/Xu2729/CoSDH.
Yanan Zhang 0005, Zhi Cai, Di Huang 0001
CVPR4
2025 Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models
abstract
In this paper, we present Diffusion-4K, a novel framework for direct ultra-high-resolution image synthesis using text-to-image diffusion models. The core advancements include: (1) Aesthetic-4K Benchmark: addressing the absence of a publicly available 4K image synthesis dataset, we construct Aesthetic-4K, a comprehensive benchmark for ultra-high-resolution image generation. We curated a high-quality 4K dataset with carefully selected images and captions generated by GPT-4o. Additionally, we introduce GLCM Score and Compression Ratio metrics to evaluate fine details, combined with holistic measures such as FID, Aesthetics and CLIPScore for a comprehensive assessment of ultra-high-resolution images. (2) Wavelet-based Fine-tuning: we propose a wavelet-based fine-tuning approach for direct training with photorealistic 4K images, applicable to various latent diffusion models, demonstrating its effectiveness in synthesizing highly detailed 4K images. Consequently, Diffusion-4K achieves impressive performance in high-quality image synthesis and text prompt adherence, especially when powered by modern large-scale diffusion models (e.g., SD3-2B and Flux-12B). Extensive experimental results from our benchmark demonstrate the superiority of Diffusion-4K in ultra-high-resolution image synthesis. Code is available at https://github.com/zhang0jhon/diffusion-4k.
Qiuyu Huang, Junjie Liu 0003, Xiefan Guo, Di Huang 0001
CVPR5
2025 Towards Training-free Anomaly Detection with Vision and Language Foundation Models
abstract
Anomaly detection is valuable for real-world applications, such as industrial quality inspection. However, most approaches focus on detecting local structural anomalies while neglecting compositional anomalies incorporating logical constraints. In this paper, we introduce LogSAD, a novel multi-modal framework that requires no training for both Logical and Structural Anomaly Detection. First, we propose a match-of-thought architecture that employs advanced large multi-modal models (i.e. GPT-4V) to generate matching proposals, formulating interests and compositional rules of thought for anomaly detection. Second, we elaborate on multi-granularity anomaly detection, consisting of patch tokens, sets of interests, and composition matching with vision and language foundation models. Subsequently, we present a calibration module to align anomaly scores from different detectors, followed by integration strategies for the final decision. Consequently, our approach addresses both logical and structural anomaly detection within a unified framework and achieves state-of-the-art results without the need for training, even when compared to supervised approaches, highlighting its robustness and effectiveness. Code is available at https://github.com/zhang0jhon/LogSAD.
Guodong Wang 0006, Yizhou Jin, Di Huang 0001
CVPR4
2025 Generating Editable Head Avatars with 3D Gaussian GANs
abstract
Generating animatable and editable 3D head avatars is essential for various applications in computer vision and graphics. Traditional 3D-aware generative adversarial networks (GANs), often using implicit fields like Neural Radiance Fields (NeRF), achieve photo-realistic and view-consistent 3D head synthesis. However, these methods face limitations in deformation flexibility and editability, hindering the creation of lifelike and easily modifiable 3D heads. We propose a novel approach that enhances the editability and animation control of 3D head avatars by incorporating 3D Gaussian Splatting (3DGS) as an explicit 3D representation. This method enables easier illumination control and improved editability. Central to our approach is the Editable Gaussian Head (EG-Head) model, which combines a 3D Morphable Model (3DMM) with texture maps, allowing precise expression control and flexible texture editing for accurate animation while preserving identity. To capture complex non-facial geometries like hair, we use an auxiliary set of 3DGS and tri-plane features. Extensive experiments demonstrate that our approach delivers high-quality 3D-aware synthesis with state-of-the-art controllability. Our code and models are available at https://github.com/liguohao96/EGG3D.
Guohao Li 0010, Hongyu Yang 0001, Yifang Men, Di Huang 0001, Weixin Li 0001, Ruijie Yang, Yunhong Wang 0001
ICASSP4
2025 ShortFT: Diffusion Model Alignment via Shortcut-Based Fine-Tuning
Xiefan Guo, Miaomiao Cui, Liefeng Bo, Di Huang 0001
ICCV4
2025 GIP: Gated Interaction Prompt for Parameter Efficient Vision-Language Fine-Tuning
abstract
Existing Parameter Efficient Fine-Tuning (PEFT) methods in vision-language (VL) domains, primarily adapted from single-modality approaches, face limitations in modeling cross-modal interactions. These methods often unify visual and textual features without explicit modality-specific processing, or rely on unidirectional interaction, leading to suboptimal task adaptation for pre-trained Vision-Language Models (VLMs). To address this issue, we propose a Gated Interaction Prompt (GIP) module as a plug-and-play adaptation to existing PEFT methods, which effectively enhances the two-way interaction between visual and textual features. Our GIP module integrates learnable prompts alongside visual and textual features into the attention layers of VLMs, serving as a bridge for cross-modal interaction. Furthermore, GIP introduces task-specific gating mechanisms to regulate and adapt the influence of prompts across different tasks, thereby further enhancing model performance. Extensive experiments on four VL tasks demonstrate that our approach can seamlessly integrate with existing methods and achieves significant performance improvements with minimal impact on parameter counts and computational costs. With only a 0.02% increase in trainable parameters, our method achieves performance gains of 0.6%, 0.8%, and 1.2% across four tasks—when applied to VL-PET, VL-Adapter, and LoRA, respectively.
Weixin Li 0001, Di Huang 0001
ICIP5
2025 Progressive Parameter Efficient Transfer Learning for Semantic Segmentation
abstract
Parameter Efficient Transfer Learning (PETL) excels in downstream classification fine-tuning with minimal computational overhead, demonstrating its potential within the pre-train and fine-tune paradigm. However, recent PETL methods consistently struggle when fine-tuning for semantic segmentation tasks, limiting their broader applicability. In this paper, we identify that fine-tuning for semantic segmentation requires larger parameter adjustments due to shifts in semantic perception granularity. Current PETL approaches are unable to effectively accommodate these shifts, leading to significant performance degradation. To address this, we introduce ProPETL, a novel approach that incorporates an additional midstream adaptation to progressively align pre-trained models for segmentation tasks. Through this process, ProPETL achieves state-of-the-art performance on most segmentation benchmarks and, for the first time, surpasses full fine-tuning on the challenging COCO-Stuff10k dataset. Furthermore, ProPETL demonstrates strong generalization across various pre-trained models and scenarios, highlighting its effectiveness and versatility for broader adoption in segmentation tasks. Code is available at: https://github.com/weeknan/ProPETL.
Huiqun Wang, Yaoyan Zheng, Di Huang 0001
ICLR4
2025 DreamScape: 3D Scene Creation via Gaussian Splatting joint Correlation Modeling
abstract
Recent advances in text-to-3D creation integrate the potent prior of Diffusion Models from text-to-image generation into 3D domain. Nevertheless, generating 3D scenes with multiple objects remains challenging. Therefore, we present DreamScape, a method for generating 3D scenes from text. Utilizing Gaussian Splatting for 3D representation, DreamScape introduces 3D Gaussian Guide that encodes semantic primitives, spatial transformations and relationships from text using LLMs, enabling local-to-global optimization. Progressive scale control is tailored during local object generation, addressing training instability issue arising from simple blending in the global optimization stage. Collision relationships between objects are modeled at the global level to mitigate biases in LLMs priors, ensuring physical correctness. Additionally, to generate pervasive objects like rain and snow distributed extensively across the scene, we design specialized sparse initialization and densification strategy. Experiments demonstrate that DreamScape achieves state-of-the-art performance, enabling high-fidelity, controllable 3D scene generation.
Yueming Zhao, Xuening Yuan, Hongyu Yang 0001, Di Huang 0001
ICME4
2025 Feature Perturbation Agent based Adversarial Attack Method for Weakly Supervised Video Anomaly Detection
abstract
Weakly supervised video anomaly detection (WS-VAD) techniques, based on video backbone models, are widely used in surveillance but are vulnerable to adversarial attacks. However, directly applying existing methods causes high memory consumption and low efficiency, and adversarial attacks on WS-VAD models have yet to be specifically studied. In this paper, we pioneer to propose a two-staged Feature Perturbation Agent based Adversarial Attack (FPAgent) method for WS-VAD. To better deceive detection models, we explore the deceivable feature spaces. To describe the locations of the deceivable feature spaces, we propose a feature perturbation agent, which also transforms the complex video-level attack into a simple segment-level attack. Besides, we propose a perturbation guider strategy to guide the feature vectors into the deceivable feature spaces, by computing the perturbation from the first segment of each video. The experiments have verified the effectiveness, as well as the attack efficiency and low memory consumption of our method.
Zhen Yang 0037, Yuanfang Guo, Ruijie Yang, Di Huang 0001, Jiantao Zhou 0001
ISCAS4
2025 Test-Time Adaptive Object Detection with Foundation Model
abstract
In recent years, test-time adaptive object detection has attracted increasing attention due to its unique advantages in online domain adaptation, which aligns more closely with real-world application scenarios. However, existing approaches heavily rely on source-derived statistical characteristics while making the strong assumption that the source and target domains share an identical category space. In this paper, we propose the first foundation model-powered test-time adaptive object detection method that eliminates the need for source data entirely and overcomes traditional closed-set limitations. Specifically, we design a Multi-modal Prompt-based Mean-Teacher framework for vision-language detector-driven test-time adaptation, which incorporates text and visual prompt tuning to adapt both language and vision representation spaces on the test data in a parameter-efficient manner. Correspondingly, we propose a Test-time Warm-start strategy tailored for the visual prompts to effectively preserve the representation capability of the vision branch. Furthermore, to guarantee high-quality pseudo-labels in every test batch, we maintain an Instance Dynamic Memory (IDM) module that stores high-quality pseudo-labels from previous test samples, and propose two novel strategies-Memory Enhancement and Memory Hallucination-to leverage IDM's high-quality instances for enhancing original predictions and hallucinating images without available pseudo-labels, respectively. Extensive experiments on cross-corruption and cross-dataset benchmarks demonstrate that our method consistently outperforms previous state-of-the-art methods, and can adapt to arbitrary cross-domain and cross-category target data. Code is available at https://github.com/gaoyingjay/ttaod_foundation.
Yingjie Gao 0001, Yanan Zhang 0005, Zhi Cai, Di Huang 0001
NeurIPS4
2025 Implicit Modeling for Transferability Estimation of Vision Foundation Models
abstract
Transferability estimation identifies the best pre-trained models for downstream tasks without incurring the high computational cost of full fine-tuning. This capability facilitates deployment and advances the pre-training and fine-tuning paradigm. However, existing methods often struggle to accurately assess transferability for emerging pre-trained models with diverse architectures, training strategies, and task alignments. In this work, we propose Implicit Transferability Modeling (ITM), a novel framework that implicitly models each model’s intrinsic transferability, coupled with a Divide-and-Conquer Variational Approximation (DVA) strategy to efficiently approximate embedding space evolution. This design enables generalization across a broader range of models and downstream tasks. Extensive experiments on a comprehensive benchmark—spanning extensive training regimes and a wider variety of model types—demonstrate that ITM consistently outperforms existing methods in terms of stability, effectiveness, and efficiency.
Yaoyan Zheng, Huiqun Wang, Di Huang 0001
NeurIPS4
2025 Anomaly-aware self-supervised feature learning for weakly supervised video anomaly detection
Zhen Yang 0037, Guodong Wang 0006, Yuanfang Guo, Xiuguo Bao, Di Huang 0001
Comput. Vis. Image Underst.5
2025 CMAE-3D: Contrastive Masked AutoEncoders for Self-Supervised 3D Object Detection
Yanan Zhang 0005, Jiaxin Chen 0002, Di Huang 0001
Int. J. Comput. Vis.3
2025 Correction: CMAE-3D: Contrastive Masked AutoEncoders for Self-Supervised 3D Object Detection
Yanan Zhang 0005, Jiaxin Chen 0002, Di Huang 0001
Int. J. Comput. Vis.3
2025 ImFace++: A Sophisticated Nonlinear 3D Morphable Face Model With Implicit Neural Representations
abstract
Accurate representations of 3D faces are of paramount importance in various computer vision and graphics applications. However, the challenges persist due to the limitations imposed by data discretization and model linearity, which hinder the precise capture of identity and expression clues in current studies. This paper presents a novel 3D morphable face model, named ImFace++, to learn a sophisticated and continuous space with implicit neural representations. ImFace++ first constructs two explicitly disentangled deformation fields to model complex shapes associated with identities and expressions, respectively, which simultaneously facilitate automatic learning of point-to-point correspondences across diverse facial shapes. To capture more sophisticated facial details, a refinement displacement field within the template space is further incorporated, enabling fine-grained learning of individual-specific facial details. Furthermore, a Neural Blend-Field is designed to reinforce the representation capabilities through adaptive blending of an array of local fields. In addition to ImFace++, we devise an improved learning strategy to extend expression embeddings, allowing for a broader range of expression variations. Comprehensive qualitative and quantitative evaluation demonstrates that ImFace++ significantly advances the state-of-the-art in terms of both face reconstruction fidelity and correspondence accuracy.
Mingwu Zheng, Hongyu Yang 0001, Liming Chen 0002, Di Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 CoupleFER: Dynamic Cross-Modal Fusion via Prompt Learning for Improved 2D+3D FER
abstract
The integration of 2D texture information and 3D geometric data has shown great promise in advancing the accuracy and robustness of 2D+3D facial expression recognition (FER) systems. Traditional methods in this domain often rely on projecting 3D data onto 2D maps, which limits the effective utilization of critical 3D features. To address this, we introduce CoupleFER, a novel approach that utilizes a cross-modal fusion strategy by combining image-based and point cloud-based networks. Unlike conventional multi-modal fusion methods, CoupleFER introduces the Cross-Modal Prompt Fusion (CouPle) module, enabling dynamic and interactive fusion between the two branches at every layer. This allows 2D texture information to serve as a guiding prompt, thereby enhancing the performance of the 3D FER branch. To further boost robustness and generalization, we propose a dual-level supervision mechanism, which imposes constraints at both the cluster and sample levels during training. Extensive experiments on the widely used BU-3DFE and Bosphorus datasets demonstrate that CoupleFER outperforms state-of-the-art methods, achieving superior recognition accuracy. Ablation studies validate the importance of each key component of the framework, underscoring its potential to significantly improve the performance of 2D + 3D FER systems, and robustness tests demonstrate its stability.
Hebeizi Li, Hongyu Yang 0001, Di Huang 0001
IEEE Trans. Affect. Comput.3
2025 A Micro-Expression Recognition Network Based on Attention Mechanism and Motion Magnification
abstract
Micro-expressions (MEs) are spontaneous facial movements that reveal an individual’s genuine emotions and play a crucial role in various domains, including lie detection, criminal analysis, mental health treatment, national security, and others. Micro-expression recognition is a highly complex aspect within the domain of affective computing, aimed at identifying subtle facial motions that are difficult for humans to discern accurately. To model the subtle facial muscle motions and the brief duration of MEs, we propose a robust micro-expression recognition (MER) network, named the attention mechanism-based motion magnification guided micro-expression recognition network (AM-MM-MER). This network consists of two primary components: the ST-MEMM network, which enhances subtle motions in micro-expression videos to reveal imperceptible facial muscle motions, and the AM-MER, which focuses on facial landmarks related to micro-expressions and incorporates novel landmark positions to extract the underlying relationships among these landmarks, thereby reducing interference from video magnification and irrelevant identity features. Extensive analysis on the CASME II and SAMM datasets demonstrates the high accuracy and effectiveness of the proposed network, achieving superior results compared to state-of-the-art methods. Ablation studies further illustrate the robustness of the proposed network.
Falin Wu, Yu Xia 0020, Boyi Ma, Tianyang Hu 0003, Jingyao Yang, Haoxin Li, Di Huang 0001
IEEE Trans. Affect. Comput.7
2025 ALD-GCN: Graph Convolutional Networks With Attribute-Level Defense
abstract
Graph Neural Networks(GNNs), such as Graph Convolutional Network, have exhibited impressive performance on various real-world datasets. However, many researches have confirmed that deliberately designed adversarial attacks can easily confuse GNNs on the classification of target nodes (targeted attacks) or all the nodes (global attacks). According to our observations, different attributes tend to be differently treated when the graph is attacked. Unfortunately, most of the existing defense methods can only defend at the graph or node level, which ignores the diversity of different attributes within each node. To address this limitation, we propose to leverage a new property, named Attribute-level Smoothness (ALS), which is defined based on the local differences of graph. We then propose a novel defense method, named GCN with Attribute-level Defense (ALD-GCN), which utilizes the ALS property to provide attribute-level protection to each attributes. Extensive experiments on real-world graphs have demonstrated the superiority of the proposed work and the potentials of our ALS property in the attacks.
Yuanfang Guo, Junfu Wang, Shihao Nie, Liang Yang 0002, Di Huang 0001, Yunhong Wang 0001
IEEE Trans. Big Data6
2025 Industrial Foundation Model
abstract
Recently, foundation models (such as ChatGPT) have emerged with powerful learning, understanding, and generalization abilities, showcasing tremendous potential to revolutionarily promote modern industry. Despite significant advancements in various fields, existing general foundation models face challenges in industry when dealing with the data of specialized modalities, the tasks of varying-scenario with multiple processes, and the requirements of trustworthy output, which makes industrial foundation model (IFM) a necessity. This article proposes a system architecture of termed IFMsys, including model training, model adaptation, and model application. Specifically, in model training, a base model is constructed by pretraining on multimodal industrial data and fine-tuning with fundamental industrial mechanisms. In model adaptation, the base model is developed into a series of task-oriented and domain-specific IFMs through fine-tuning with representative tasks and domain knowledge. In model application, an industrial agent-centric collaboration system and a comprehensive application framework of IFM are proposed to enhance the industrial product lifecycle applications. In addition, a prototype system of the IFM, namely, MetaIndux, is delivered, with application examples presented in typical industrial tasks. Finally, future research directions and open issues of IFM are prospected. We hope this article will inspire the advancements in the theories, technologies, and applications in this emerging research field of IFM.
Lei Ren 0001, Haiteng Wang, Jiabao Dong, Zidi Jia, Shixiang Li, Yuqing Wang 0007, Yuanjun Laili, Di Huang 0001, Lin Zhang 0009, Bo Hu Li 0001
IEEE Trans. Cybern.8
2025 Cross-Modal Contrastive Masked AutoEncoder for Compressed Video Pre-Training
abstract
In this paper, we propose a novel Transformer based approach, namely Cross-modal Contrastive Masked AutoEncoder (C2MAE), to Self-Supervised Learning (SSL) on compressed videos. A unified Transformer encoder is employed to discover relationships of visual tokens from RGBs, motion vectors and residuals. A hybrid SSL framework is proposed, which combines the complementary advantages of Masked Image Modeling (MIM) and Contrastive Learning (CL) pretext tasks, for powerful representation learning. The MIM branch extends VideoMAE by a new Fine-Grained Motion-aware Masking (FGMM) strategy and a modified Multi-modal Reconstruction (MR) task, where FGMM computes motion saliency maps as motion priors to guide the masks so that it well fits for the data properties in the compressed domain and the MR task highlights the reconstruction of raw videos by joint representations from corresponding compressed videos in addition to that in each single modality. The CL branch introduces the Contrastive Cross-modal Learning (CCL) module, and the features from a compressed video clip and the ones from its raw video counterpart are compared instead of widely used augmented data. Due to these designs, C2MAE significantly enhances interactions across modalities to compensate the sparsity of I-frames and the coarse and noisy nature of P-frames, thus delivering much stronger pre-trained models. Extensive experiments are conducted on the UCF-101, HMDB-51 and Kinetics-400 benchmarks with state-of-the-art results reported, demonstrating its effectiveness.
Jiaxin Chen 0002, Guohao Li 0010, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001
IEEE Trans. Image Process.6
2025 Sharing Task-Relevant Information in Visual Prompt Tuning by Cross-Layer Dynamic Connection
abstract
Recent progress has shown great potential of visual prompt tuning (VPT) when adapting pre-trained vision transformers to various downstream tasks. However, most existing solutions independently optimize prompts at each layer, thereby neglecting the usage of task-relevant information encoded in prompt tokens across layers. Additionally, existing prompt structures are prone to interference from task-irrelevant noise in input images, which can adversely affect the sharing of task-relevant information. In this paper, we propose a novel VPT approach, SVPT. It innovatively incorporates a cross-layer dynamic connection (CDC) for input prompt tokens from adjacent layers, enabling effective sharing of task-relevant information. Furthermore, we design a dynamic aggregation (DA) module that facilitates selective sharing of information between layers. The combination of CDC and DA enhances the flexibility of the attention process within the VPT framework. Building upon these foundations, SVPT introduces an attentive enhancement (AE) mechanism that automatically identifies salient image tokens and refines them with prompt tokens in an additive manner. Extensive experiments on 24 image classification and semantic segmentation benchmarks clearly demonstrate the advantages of the proposed SVPT, compared to the state-of-the-art counterparts.
Jiaxin Chen 0002, Di Huang 0001
IEEE Trans. Image Process.3
2025 Multi-Grained Contrastive Learning for Text-Supervised Open-Vocabulary Semantic Segmentation
abstract
Learning open-vocabulary semantic segmentation (OVSS) from text supervision has recently received increasing attention for its promising potential in real-world applications. However, only with image-level supervision, it struggles to achieve dense and robust cross-modal alignment and thus limits pixel-level predictions. In this article, we present a novel approach to this task with M ulti- G rained C ross-modal C ontrastive L earning, named MGCCL. Specifically, unlike current solutions restricted by coarse image/object-text alignment, MGCCL constructs pseudo multi-granular semantic correspondences at the object-, part-, and pixel-level and collaborates with hard sampling strategies to conduct cross-modal contrastive learning, significantly facilitating fine-grained alignment. Further, we develop an adaptive semantic unit which flexibly harnesses the learned multi-grained cross-modal alignment capabilities to effectively mitigate the under- and over-segmentation issues arising from the per-group and per-pixel units. Extensive experiments over a broad suite of eight segmentation benchmarks show that our approach delivers significant advancements over state-of-the-art counterparts, demonstrating its effectiveness.
Pu Ge, Guodong Wang 0006, Qingjie Liu 0001, Di Huang 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Align-DETR: Enhancing End-to-end Object Detection with Aligned Loss
Zhi Cai, Guodong Wang 0006, Zheng Ge, Xiangyu Zhang 0005, Di Huang 0001
BMVC7
2024 Initno: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization
abstract
Recent strides in the development of diffusion models, ex-emplified by advancements such as Stable Diffusion, have underscored their remarkable prowess in generating visu-ally compelling images. However, the imperative of achieving a seamless alignment between the generated image and the provided prompt persists as a formidable challenge. This paper traces the root of these difficulties to invalid initial noise, and proposes a solution in the form of Initial Noise Optimization (INITNO), a paradigm that refines this noise. Considering text prompts, not all random noises are effective in synthesizing semantically-faithful images. We design the cross-attention response score and the selfattention conflict score to evaluate the initial noise, bifurcating the initial latent space into valid and invalid sectors. A strategically crafted noise optimization pipeline is developed to guide the initial noise towards valid regions. Our method, validated through rigorous experimentation, shows a commendable proficiency in generating images in strict accordance with text prompts. Our code is available at https://github.com/xiefan-guo/initno.
Xiefan Guo, Jinlin Liu, Miaomiao Cui, Jiankai Li, Hongyu Yang 0001, Di Huang 0001
CVPR6
2024 Generalizing 6-DoF Grasp Detection via Domain Prior Knowledge
abstract
We focus on the generalization ability of the 6-DoF grasp detection method in this paper. While learning-based grasp detection methods can predict grasp poses for unseen ob-jects using the grasp distribution learned from the training set, they often exhibit a significant performance drop when encountering objects with diverse shapes and struc-tures. To enhance the grasp detection methods' general-ization ability, we incorporate domain prior knowledge of robotic grasping, enabling better adaptation to objects with significant shape and structure differences. More specifi-cally, we employ the physical constraint regularization during the training phase to guide the model towards predicting grasps that comply with the physical rule on grasping. For the unstable grasp poses predicted on novel objects, we design a contact-score joint optimization using the pro-jection contact map to refine these poses in cluttered sce-narios. Extensive experiments conducted on the GraspNet-1 billion benchmark demonstrate a substantial performance gain on the novel object set and the real-world grasping experiments also demonstrate the effectiveness of our gen-eralizing 6-DoF grasp detection method. Code is available at https://github.com/mahaoxiang822/Generalizing-Grasp.
Modi Shi, Boyang Gao, Di Huang 0001
CVPR4
2024 Crowd-SAM: SAM as a Smart Annotator for Object Detection in Crowded Scenes
Zhi Cai, Yingjie Gao 0001, Yaoyan Zheng, Di Huang 0001
ECCV (69)5
2024 Multi-modal Relation Distillation for Unified 3D Representation Learning
Huiqun Wang, Yiping Bao, Panwang Pan, Ruijie Yang, Di Huang 0001
ECCV (33)7
2024 AdaLog: Post-training Quantization for Vision Transformers with Adaptive Logarithm Quantizer
Zhuguanyu Wu, Jiaxin Chen 0002, Hanwen Zhong, Di Huang 0001, Yunhong Wang 0001
ECCV (27)4
2024 DrFER: Learning Disentangled Representations for 3D Facial Expression Recognition
abstract
Facial Expression Recognition (FER) has consistently been a focal point in the field of facial analysis. In the context of existing methodologies for 3D FER or 2D+3D FER, the extraction of expression features often gets entangled with identity information, compromising the distinctiveness of these features. To tackle this challenge, we introduce the innovative DrFER method, which brings the concept of disentangled representation learning to the field of 3D FER. DrFER employs a dual-branch framework to effectively disentangle expression information from identity information. Diverging from prior disentanglement endeavors in the 3D facial domain, we have carefully reconfigured both the loss functions and network structure to make the overall framework adaptable to point cloud data. This adaptation enhances the capability of the framework in recognizing facial expressions, even in cases involving varying head poses. Extensive evaluations conducted on the BU-3DFE and Bosphorus datasets substantiate that DrFER surpasses the performance of other 3D FER methods.
Hebeizi Li, Hongyu Yang 0001, Di Huang 0001
FG3
2024 3D Face Modeling via Weakly-Supervised Disentanglement Network Joint Identity-Consistency Prior
abstract
Generative 3D face models featuring disentangled controlling factors hold immense potential for diverse applications in computer vision and computer graphics. However, previous 3D face modeling methods face a challenge as they demand specific labels to effectively disentangle these factors. This becomes particularly problematic when integrating multiple 3D face datasets to improve the generalization of the model. Addressing this issue, this paper introduces a Weakly-Supervised Disentanglement Framework, denoted as WSDF, to facilitate the training of controllable 3D face models without an overly stringent labeling requirement. Adhering to the paradigm of Variational Autoencoders (VAEs), the proposed model achieves disentanglement of identity and expression controlling factors through a two-branch encoder equipped with dedicated identity-consistency prior. It then faithfully re-entangles these factors via a tensor-based combination mechanism. Notably, the introduction of the Neutral Bank allows precise acquisition of subject-specific information using only identity labels, thereby averting degeneration due to insufficient supervision. Additionally, the framework incorporates a label-free second-order loss function for the expression factor to regulate deformation space and eliminate extraneous information, resulting in enhanced disentanglement. Extensive experiments have been conducted to substantiate the superior performance of WSDF. Our code is available at https://github.com/liguoha096/WSDF.
Guohao Li 0010, Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001
FG3
2024 Progressive Self-supervised Representation Learning for 3D Facial Expression Recognition
abstract
Facial expression recognition (FER) is a critical area of research in face analysis. While 2D data has been extensively used, 3D data offers inherent advantages, such as increased resilience to illumination and pose variations. However, the limited size of current 3D FER datasets significantly constrains the performance of 3D FER methods. To overcome this challenge, we propose a novel self-supervised pre-training scheme by leveraging large-scale external 3D data, followed by fine-tuning on 3D FER datasets. Our approach starts with self-supervised learning on a large-scale 3D point cloud object dataset, specifically ShapeNet. We then move on to the FaceScape dataset, which is primarily used for morphable face prediction. To enhance robustness, we integrate synthetic data before fine-tuning on specific FER datasets. This multi-stage process allows the model to progressively learn 3D facial expression representations from coarse to fine. For this purpose, we utilize Point-MAE, a leading self-supervised model for representation learning. To enhance its ability for FER task, we further incorporate facial priors in the masking and point sampling steps, leveraging the distinctive characteristics of facial data. Our method achieves state-of-the-art performance on both BU-3DFE and Bosphorus datasets, matching or surpassing results achieved by other 2D+3D FER techniques.
Hebeizi Li, Hongyu Yang 0001, Di Huang 0001
IJCB3
2024 Towards Generalizable Referring Image Segmentation Via Target Prompt And Visual Coherence
abstract
Referring image segmentation (RIS) aims to segment objects in an image conditioning on free-form text descriptions. Despite the overwhelming progress, it still remains challenging for current approaches to perform well on cases with various text expressions or with unseen visual entities, limiting its further application. In this paper, we present a novel RIS approach, which substantially improves the generalization ability by addressing the two dilemmas mentioned above. Specially, to deal with unconstrained texts, we propose to boost a given expression with an explicit and crucial prompt, which complements the expression in a unified context, facilitating target capturing in the presence of linguistic style changes. Furthermore, we introduce a multi-modal fusion aggregation module with visual guidance from a powerful pretrained model to leverage spatial relations and pixel coherences to handle the incomplete target masks and false positive irregular clumps which often appear on unseen visual entities. Extensive experiments are conducted in the zero-shot cross-dataset settings and the proposed approach achieves consistent gains compared to the state-of-the-art, e.g., $4.15 \%$, $5.45 \%$, and $4.64 \%$ mIoU increase on RefCOCO, RefCOCO+ and ReferIt respectively, demonstrating its effectiveness.
Pu Ge, Shichao Fan, Qingjie Liu 0001, Di Huang 0001, Yunhong Wang 0001
ICIP6
2024 Rotation Has Two Sides: Evaluating Data Augmentation for Deep One-class Classification
abstract
One-class classification (OCC) involves predicting whether a new data is normal or anomalous based solely on the data from a single class during training. Various attempts have been made to learn suitable representations for OCC within a self-supervised framework. Notably, discriminative methods that use geometric visual transformations, such as rotation, to generate pseudo-anomaly samples have exhibited impressive detection performance. Although rotation is commonly viewed as a distribution-shifting transformation and is widely used in the literature, the cause of its effectiveness remains a mystery. In this study, we are the first to make a surprising observation: there exists a strong linear relationship (Pearson's Correlation, $r > 0.9$) between the accuracy of rotation prediction and the performance of OCC. This suggests that a classifier that effectively distinguishes different rotations is more likely to excel in OCC, and vice versa. The root cause of this phenomenon can be attributed to the transformation bias in the dataset, where representations learned from transformations already present in the dataset tend to be less effective, making it essential to accurately estimate the transformation distribution before utilizing pretext tasks involving these transformations for reliable self-supervised representation learning. To the end, we propose a novel two-stage method to estimate the transformation distribution within the dataset. In the first stage, we learn general representations through standard contrastive pre-training. In the second stage, we select potentially semantics-preserving samples from the entire augmented dataset, which includes all rotations, by employing density matching with the provided reference distribution. By sorting samples based on semantics-preserving versus shifting transformations, we achieve improved performance on OCC benchmarks.
Guodong Wang 0006, Yunhong Wang 0001, Xiuguo Bao, Di Huang 0001
ICLR4
2024 Fast Textile Pilling Classification Based on a Lightweight Network and 3D Point Clouds
abstract
Point clouds have demonstrated extensive application prospects in various fields, including research related to the evaluation of textile pilling. We collect 3D point cloud data in the actual test environment of textiles, which has been organized and named the TextileNet dataset. To the best of our knowledge, it is the first publicly available 3D point cloud dataset in the field of textile pilling assessment. Based on the Non-parametric Network for 3D point cloud analysis (Point-NN), we construct a Few-parameter Network called Point-FN for experiments on the TextileNet dataset. Experimental results indicate that under conditions with a parameter count of only 0.5M and FLOPs of 1.7G, Point-FN achieves an Overall Accuracy (OA) of 91.1% and a Mean per-class Accuracy (MA) of 93.0%. Moreover, under the testing conditions of a single RTX 2080Ti GPU, Point-FN demonstrates an inference speed of 164 FPS. Testing results on other publicly available datasets also validate the competitive performance of Point-FN. The proposed TextileNet dataset will be publicly available.
Yizhou Jin, Qingjie Liu 0001, Di Huang 0001, Yunhong Wang 0001
ICME7
2024 Sim-to-Real Grasp Detection with Global-to-Local RGB-D Adaptation
abstract
This paper focuses on the sim-to-real issue of RGB-D grasp detection and formulates it as a domain adaptation problem. In this case, we present a global-to-local method to address hybrid domain gaps in RGB and depth data and insufficient multi-modal feature alignment. First, a self-supervised rotation pre-training strategy is adopted to deliver robust initialization for RGB and depth networks. We then propose a global-to-local alignment pipeline with individual global domain classifiers for scene features of RGB and depth images as well as a local one specifically working for grasp features in the two modalities. In particular, we propose a grasp prototype adaptation module, which aims to facilitate fine-grained local feature alignment by dynamically updating and matching the grasp prototypes from the simulation and real-world scenarios throughout the training process. Due to such designs, the proposed method substantially reduces the domain shift and thus leads to consistent performance improvements. Extensive experiments are conducted on the GraspNet-Planar benchmark and physical environment, and superior results are achieved which demonstrate the effectiveness of our method. Code is available at https://github.com/mahaoxiang822/GL-MSDA.
Ran Qin, Modi Shi, Boyang Gao, Di Huang 0001
ICRA5
2024 PS-TTL: Prototype-based Soft-labels and Test-Time Learning for Few-shot Object Detection
abstract
In recent years, Few-Shot Object Detection (FSOD) has gained widespread attention and made significant progress due to its ability to build models with a good generalization power using extremely limited annotated data. The fine-tuning based paradigm is currently dominating this field, where detectors are initially pre-trained on base classes with sufficient samples and then fine-tuned on novel ones with few samples, but the scarcity of labeled samples of novel classes greatly interferes precisely fitting their data distribution, thus hampering the performance. To address this issue, we propose a new framework for FSOD, namely Prototype-based Soft-labels and Test-Time Learning (PS-TTL). Specifically, we design a Test-Time Learning (TTL) module that employs a mean-teacher network for self-training to discover novel instances from test data, allowing detectors to learn better representations and classifiers for novel classes. Furthermore, we notice that even though relatively low-confidence pseudo-labels exhibit classification confusion, they still tend to recall foreground. We thus develop a Prototype-based Soft-labels (PS) strategy through assessing similarities between low-confidence pseudo-labels and category prototypes as soft-labels to unleash their potential, which substantially mitigates the constraints posed by few-shot samples. Extensive experiments on both the VOC and COCO benchmarks show that PS-TTL achieves the state-of-the-art, highlighting its effectiveness. The code and model are available at https://github.com/gaoyingjay/PS-TTL.
Yingjie Gao 0001, Yanan Zhang 0005, Ziyue Huang 0001, Nanqing Liu, Di Huang 0001
ACM Multimedia5
2024 Active Perception for Grasp Detection via Neural Graspness Field
abstract
This paper tackles the challenge of active perception for robotic grasp detection in cluttered environments. Incomplete 3D geometry information can negatively affect the performance of learning-based grasp detection methods, and scanning the scene from multiple views introduces significant time costs. To achieve reliable grasping performance with efficient camera movement, we propose an active grasp detection framework based on the Neural Graspness Field (NGF), which models the scene incrementally and facilitates next-best-view planning. Constructed in real-time as the camera moves, the NGF effectively models the grasp distribution in 3D space by rendering graspness predictions from each view. For next-best-view planning, we aim to reduce the uncertainty of the NGF through a graspness inconsistency-guided policy, selecting views based on discrepancies between NGF outputs and a pre-trained graspness network. Additionally, we present a neural graspness sampling method that decodes graspness values from the NGF to improve grasp pose detection results. Extensive experiments on the GraspNet-1Billion benchmark demonstrate significant performance improvements compared to previous works. Real-world experiments show that our method achieves a superior trade-off between grasping performance and time costs.
Modi Shi, Boyang Gao, Di Huang 0001
NeurIPS4
2024 Transforming Vision Transformer: Towards Efficient Multi-Task Asynchronous Learner
abstract
Multi-Task Learning (MTL) for Vision Transformer aims at enhancing the model capability by tackling multiple tasks simultaneously. Most recent works have predominantly focused on designing Mixture-of-Experts (MoE) structures and integrating Low-Rank Adaptation (LoRA) to efficiently perform multi-task learning. However, their rigid combination hampers both the optimization of MoE and the effectiveness of reparameterization of LoRA, leading to sub-optimal performance and low inference speed. In this work, we propose a novel approach dubbed Efficient Multi-Task Learning (EMTAL) by transforming a pre-trained Vision Transformer into an efficient multi-task learner during training, and reparameterizing the learned structure for efficient inference. Specifically, we firstly develop the MoEfied LoRA structure, which decomposes the pre-trained Transformer into a low-rank MoE structure and employ LoRA to fine-tune the parameters. Subsequently, we take into account the intrinsic asynchronous nature of multi-task learning and devise a learning Quality Retaining (QR) optimization mechanism, by leveraging the historical high-quality class logits to prevent a well-trained task from performance degradation. Finally, we design a router fading strategy to integrate the learned parameters into the original Transformer, archiving efficient inference. Extensive experiments on public benchmarks demonstrate the superiority of our method, compared to the state-of-the-art multi-task learning approaches.
Hanwen Zhong, Jiaxin Chen 0002, Di Huang 0001, Yunhong Wang 0001
NeurIPS4
2024 FIFAWC: a dataset with detailed annotation and rich semantics for group activity recognition
Duoxuan Pei, Di Huang 0001, Yunhong Wang 0001
Frontiers Comput. Sci.2
2024 SA3WT: Adaptive Wavelet-Based Transformer with Self-Paced Auto Augmentation for Face Forgery Detection
Hongyu Yang 0001, Binghui Chen, Di Huang 0001
Int. J. Comput. Vis.5
2024 FG-AGR: Fine-Grained Associative Graph Representation for Facial Expression Recognition in the Wild
abstract
Facial expression recognition (FER) in the wild is challenging due to various unconstrained conditions, i.e., occlusions and head pose variations. Previous methods tend to improve the performance of facial expression recognition through resorting to holistic methods or coarse local-based methods, while ignoring the local fine-grained feature structure knowledge and the correlation between features. In this paper, we propose a Fine-Grained Association Graph Representation (FG-AGR) framework which can capture the local fine-grained facial expression representation. Firstly, an Adaptive Salient Region Induction (ASRI) is designed for adaptively highlighting the local saliency regions of facial expressions combined with spatial location information. Based on this, a Local Fine-grained Feature Extraction (LFFE) based on Visual Transformers is introduced to further extract fine but discriminative fine-grained features of saliency regions. Thirdly, an Adaptive Graph Association Reasoning (AGAR) based on Graph Convolutional Network is constructed to learn associated fine-grained feature combinations. Extensive experiments demonstrate that our FG-AGR achieves superior performance compared to the state-of-the-art methods with 90.81% on RAF-DB, 64.91% on AffectNet-7, 60.69% on AffectNet-8 and 91.09% on FERPlus.
Chunlei Li 0002, Xiao Li 0047, Di Huang 0001, Zhoufeng Liu
IEEE Trans. Circuits Syst. Video Technol.4
2024 Towards Video Anomaly Detection in the Real World: A Binarization Embedded Weakly-Supervised Network
abstract
In this letter, we pioneer to propose a binarization embedded weakly-supervised video anomaly detection (BE-WSVAD) method by constructing a binarized GCN-based anomaly detection module. Compared to the existing weakly-supervised video anomaly detection (WS-VAD) methods, BE-WSVAD focuses on the detection efficiency, which is ignored by the existing literature yet vital in real applications. Specifically, to improve the detection performance of the binary anomaly detection module, we propose a binary network augmentation strategy in the training process. Due to the weakly supervision mechanism, the videos employed in the training process are usually lengthy, in which the lengthy-input dependencies tend to be exploited to improve the detection performance with extra memory consumption. Then, we propose the short-input inference modes, which can largely reduce the desired length of the input video. Experimental results demonstrate the superiority of our BE-WSVAD in terms of the memory and computational consumptions while giving comparable accuracies.
Zhen Yang 0037, Yuanfang Guo, Junfu Wang, Di Huang 0001, Xiuguo Bao, Yunhong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Deep Common Feature Mining for Efficient Video Semantic Segmentation
abstract
Recent advancements in video semantic segmentation have made substantial progress by exploiting temporal correlations. Nevertheless, persistent challenges, including redundant computation and the reliability of the feature propagation process, underscore the need for further innovation. In response, we present Deep Common Feature Mining (DCFM), a novel approach strategically designed to address these challenges by leveraging the concept of feature sharing. DCFM explicitly decomposes features into two complementary components. The common representation extracted from a key-frame furnishes essential high-level information to neighboring non-key frames, allowing for direct re-utilization without feature propagation. Simultaneously, the independent feature, derived from each video frame, captures rapidly changing information, providing frame-specific clues crucial for segmentation. To achieve such decomposition, we employ a symmetric training strategy tailored for sparsely annotated data, empowering the backbone to learn a robust high-level representation enriched with common information. Additionally, we incorporate a self-supervised loss function to reinforce intra-class feature similarity and enhance temporal consistency. Experimental evaluations on the VSPW and Cityscapes datasets demonstrate the effectiveness of our method, showing a superior balance between accuracy and efficiency. The implementation is available athttps://github.com/BUAAHugeGun/DCFM.
Yaoyan Zheng, Hongyu Yang 0001, Di Huang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 OTAMatch: Optimal Transport Assignment With PseudoNCE for Semi-Supervised Learning
abstract
In semi-supervised learning (SSL), many approaches follow the effective self-training paradigm with consistency regularization, utilizing threshold heuristics to alleviate label noise. However, such threshold heuristics lead to the underutilization of crucial discriminative information from the excluded data. In this paper, we present OTAMatch, a novel SSL framework that reformulates pseudo-labeling as an optimal transport (OT) assignment problem and simultaneously exploits data with high confidence to mitigate the confirmation bias. Firstly, OTAMatch models the pseudo-label allocation task as a convex minimization problem, facilitating end-to-end optimization with all pseudo-labels and employing the Sinkhorn-Knopp algorithm for efficient approximation. Meanwhile, we incorporate epsilon-greedy posterior regularization and curriculum bias correction strategies to constrain the distribution of OT assignments, improving the robustness with noisy pseudo-labels. Secondly, we propose PseudoNCE, which explicitly exploits pseudo-label consistency with threshold heuristics to maximize mutual information within self-training, significantly boosting the balance of convergence speed and performance. Consequently, our proposed approach achieves competitive performance on various SSL benchmarks. Specifically, OTAMatch substantially outperforms the previous state-of-the-art SSL algorithms in realistic and challenging scenarios, exemplified by a notable 9.45% error rate reduction over SoftMatch on ImageNet with 100K-label split, underlining its robustness and effectiveness.
Junjie Liu 0003, Debang Li, Qiuyu Huang, Jiaxin Chen 0002, Di Huang 0001
IEEE Trans. Image Process.6
2024 Deep Learning for Time-Series Prediction in IIoT: Progress, Challenges, and Prospects
abstract
Time-series prediction plays a crucial role in the Industrial Internet of Things (IIoT) to enable intelligent process control, analysis, and management, such as complex equipment maintenance, product quality management, and dynamic process monitoring. Traditional methods face challenges in obtaining latent insights due to the growing complexity of IIoT. Recently, the latest development of deep learning provides innovative solutions for IIoT time-series prediction. In this survey, we analyze the existing deep learning-based time-series prediction methods and present the main challenges of time-series prediction in IIoT. Furthermore, we propose a framework of state-of-the-art solutions to overcome the challenges of time-series prediction in IIoT and summarize its application in practical scenarios, such as predictive maintenance, product quality prediction, and supply chain management. Finally, we conclude with comments on possible future directions for the development of time-series prediction to enable extensible knowledge mining for complex tasks in IIoT.
Lei Ren 0001, Zidi Jia, Yuanjun Laili, Di Huang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Learning Polysemantic Spoof Trace: A Multi-Modal Disentanglement Network for Face Anti-spoofing
abstract
Along with the widespread use of face recognition systems, their vulnerability has become highlighted. While existing face anti-spoofing methods can be generalized between attack types, generic solutions are still challenging due to the diversity of spoof characteristics. Recently, the spoof trace disentanglement framework has shown great potential for coping with both seen and unseen spoof scenarios, but the performance is largely restricted by the single-modal input. This paper focuses on this issue and presents a multi-modal disentanglement model which targetedly learns polysemantic spoof traces for more accurate and robust generic attack detection. In particular, based on the adversarial learning mechanism, a two-stream disentangling network is designed to estimate spoof patterns from the RGB and depth inputs, respectively. In this case, it captures complementary spoofing clues inhering in different attacks. Furthermore, a fusion module is exploited, which recalibrates both representations at multiple stages to promote the disentanglement in each individual modality. It then performs cross-modality aggregation to deliver a more comprehensive spoof trace representation for prediction. Extensive evaluations are conducted on multiple benchmarks, demonstrating that learning polysemantic spoof traces favorably contributes to anti-spoofing with more perceptible and interpretable results.
Kaicheng Li, Hongyu Yang 0001, Binghui Chen, Di Huang 0001
AAAI6
2023 Adaptive Sparse Convolutional Networks with Global Context Enhancement for Faster Object Detection on Drone Images
abstract
Object detection on drone images with low-latency is an important but challenging task on the resource-constrained unmanned aerial vehicle (UAV) platform. This paper investigates optimizing the detection head based on the sparse convolution, which proves effective in balancing the accuracy and efficiency. Nevertheless, it suffers from inadequate integration of contextual information of tiny objects as well as clumsy control of the mask ratio in the presence of foreground with varying scales. To address the issues above, we propose a novel global context-enhanced adaptive sparse convolutional network (CEASC). It first develops a context-enhanced group normalization (CE-GN) layer, by replacing the statistics based on sparsely sampled features with the global contextual ones, and then designs an adaptive multi-layer masking strategy to generate optimal mask ratios at distinct scales for compact foreground coverage, promoting both the accuracy and efficiency. Extensive experimental results on two major benchmarks, i.e. VisDrone and UAVDT, demonstrate that CEASC remarkably reduces the GFLOPs and accelerates the inference procedure when plugging into the typical state-of-the-art detection frameworks (e.g. RetinaNet and GFL V1) with competitive performance. Code is available at https://github.com/Cuogeihong/CEASC.
Bowei Du, Yecheng Huang, Jiaxin Chen 0002, Di Huang 0001
CVPR4
2023 NeuFace: Realistic 3D Neural Face Rendering from Multi-View Images
abstract
Realistic face rendering from multi-view images is beneficial to various computer vision and graphics applications. Due to complex spatially-varying reflectance properties and geometry characteristics of faces, however, it remains challenging to recover 3D facial representations both faithfully and efficiently in the current studies. This paper presents a novel 3D face rendering model, namely NeuFace, to learn accurate and physically-meaningful underlying 3D representations by neural rendering techniques. It naturally in-corporates the neural BRDFs into physically based rendering, capturing sophisticated facial geometry and appearance clues in a collaborative manner. Specifically, we introduce an approximated BRDF integration and a simple yet new low-rank prior, which effectively lower the ambiguities and boost the performance of the facial BRDFs. Extensive experiments are performed to demonstrate the superiority of NeuFace in human face rendering, along with a decent generalization ability to common objects. Code is released at NeuFace.
Mingwu Zheng, Hongyu Yang 0001, Di Huang 0001
CVPR4
2023 OcTr: Octree-Based Transformer for 3D Object Detection
abstract
A key challenge for LiDAR-based 3D object detection is to capture sufficient features from large scale 3D scenes especially for distant or/and occluded objects. Albeit recent efforts made by Transformers with the long sequence modeling capability, they fail to properly balance the accuracy and efficiency, suffering from inadequate receptive fields or coarse-grained holistic correlations. In this paper, we propose an Octree-based Transformer, named OcTr, to address this issue. It first constructs a dynamic octree on the hierarchical feature pyramid through conducting self-attention on the top level and then recursively propagates to the level below restricted by the octants, which captures rich global context in a coarse-to-fine manner while maintaining the computational complexity under control. Furthermore, for enhanced foreground perception, we propose a hybrid positional embedding, composed of the semantic-aware positional embedding and attention mask, to fully exploit semantic and geometry clues. Extensive experiments are conducted on the Waymo Open Dataset and KITTI Dataset, and OcTr reaches newly state-of-the-art results.
Yanan Zhang 0005, Jiaxin Chen 0002, Di Huang 0001
CVPR4
2023 Weakly-Supervised Photo-realistic Texture Generation for 3D Face Reconstruction
abstract
Although much progress has been made recently in 3D face reconstruction, most previous work has been devoted to predicting accurate and fine-grained 3D shapes. In contrast, relatively little work has focused on generating high-fidelity face textures. Compared with the prosperity of photo-realistic 2D face image generation, high-fidelity 3D face texture generation has yet to be studied. In this paper, we propose a novel UV map generation model that predicts the UV map from a single face image. The model consists of a UV sampler and a UV generator. By selectively sampling the input face image's pixels and adjusting their relative locations, the UV sampler generates an incomplete UV map that could faithfully reconstruct the original face. Missing textures in the incomplete UV map are further full-filled by the UV generator. The training is based on pseudo ground truth blended by the 3DMM texture and the input face texture, thus weakly supervised. To deal with the artifacts in the imperfect pseudo UV map, multiple UV map and face image discriminators are leveraged.
Xiangnan Yin, Di Huang 0001, Zehua Fu, Yunhong Wang 0001, Liming Chen 0002
FG2
2023 Segmentation-Reconstruction-Guided Facial Image De-occlusion
abstract
Occlusions are very common in face images in the wild, leading to the degraded performance of face-related tasks. Although much effort has been devoted to removing occlusions from face images, the varying shapes and textures of occlusions still challenge the robustness of current methods. As a result, current methods either rely on manual occlusion masks or only apply to specific occlusions. This paper proposes a novel face de-occlusion model based on face segmentation and 3D face reconstruction, which is robust to arbitrary kinds of face occlusions. The proposed model consists of a 3D face reconstruction module, a face segmentation module, and an image generation module. With the face prior and the occlusion mask predicted by the first two, respectively, the image generation module can faithfully recover the missing facial textures. To supervise the training, we further build a large occlusion dataset, with both manually labeled and synthetic occlusions. Qualitative and quantitative results demonstrate the effectiveness and robustness of the proposed method.
Xiangnan Yin, Di Huang 0001, Zehua Fu, Yunhong Wang 0001, Liming Chen 0002
FG2
2023 Unilaterally Aggregated Contrastive Learning with Hierarchical Augmentation for Anomaly Detection
abstract
Anomaly detection (AD), aiming to find samples that deviate from the training distribution, is essential in safety-critical applications. Though recent self-supervised learning based attempts achieve promising results by creating virtual outliers, their training objectives are less faithful to AD which requires a concentrated inlier distribution as well as a dispersive outlier distribution. In this paper, we propose Unilaterally Aggregated Contrastive Learning with Hierarchical Augmentation (UniCon-HA), taking into account both the requirements above. Specifically, we explicitly encourage the concentration of inliers and the dispersion of virtual outliers via supervised and unsupervised contrastive losses, respectively. Considering that standard contrastive data augmentation for generating positive views may induce outliers, we additionally introduce a soft mechanism to re-weight each augmented inlier according to its deviation from the inlier distribution, to ensure a purified concentration. Moreover, to prompt a higher concentration, inspired by curriculum learning, we adopt an easy-to-hard hierarchical augmentation strategy and perform contrastive aggregation at different depths of the network based on the strengths of data augmentation. Our method is evaluated under three AD settings including unlabeled one-class, unlabeled multi-class, and labeled multi-class, demonstrating its consistent superiority over other competitors.
Guodong Wang 0006, Yunhong Wang 0001, Jie Qin 0004, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001
ICCV6
2023 Denoising Diffusion Autoencoders are Unified Self-supervised Learners
abstract
Inspired by recent advances in diffusion models, which are reminiscent of denoising autoencoders, we investigate whether they can acquire discriminative representations for classification via generative pre-training. This paper shows that the networks in diffusion models, namely denoising diffusion autoencoders (DDAE), are unified self-supervised learners: by pre-training on unconditional image generation, DDAE has already learned strongly linear-separable representations within its intermediate layers without auxiliary encoders, thus making diffusion pre-training emerge as a general approach for generative-and-discriminative dual learning. To validate this, we conduct linear probe and finetuning evaluations. Our diffusion-based approach achieves 95.9% and 50.0% linear evaluation accuracies on CIFAR-10 and Tiny-ImageNet, respectively, and is comparable to contrastive learning and masked autoencoders for the first time. Transfer learning from ImageNet also confirms the suitability of DDAE for Vision Transformers, suggesting the potential to scale DDAEs as unified foundation models. Code is available at github.com/FutureXiang/ddae.
Weilai Xiang, Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001
ICCV3
2023 DR-Tune: Improving Fine-tuning of Pretrained Visual Models by Distribution Regularization with Semantic Calibration
abstract
The visual models pretrained on large-scale benchmarks encode general knowledge and prove effective in building more powerful representations for downstream tasks. Most existing approaches follow the fine-tuning paradigm, either by initializing or regularizing the downstream model based on the pretrained one. The former fails to retain the knowledge in the successive fine-tuning phase, thereby prone to be over-fitting, and the latter imposes strong constraints to the weights or feature maps of the downstream model without considering semantic drift, often incurring insufficient optimization. To deal with these issues, we propose a novel fine-tuning framework, namely distribution regularization with semantic calibration (DR-Tune). It employs distribution regularization by enforcing the downstream task head to decrease its classification error on the pretrained feature distribution, which prevents it from over-fitting while enabling sufficient training of downstream encoders. Furthermore, to alleviate the interference by semantic drift, we develop the semantic calibration (SC) module to align the global shape and class centers of the pretrained and downstream feature distributions. Extensive experiments on widely used image classification datasets show that DR-Tune consistently improves the performance when combing with various backbones under different pretraining strategies. Code is available at: https://github.com/weeknan/DR-Tune.
Jiaxin Chen 0002, Di Huang 0001
ICCV3
2023 RGB-D Grasp Detection via Depth Guided Learning with Cross-modal Attention
abstract
Planar grasp detection is one of the most fundamental tasks to robotic manipulation, and the recent progress of consumer-grade RGB-D sensors enables delivering more comprehensive features from both the texture and shape modalities. However, depth maps are generally of a relatively lower quality with much stronger noise compared to RGB images, making it challenging to acquire grasp depth and fuse multi-modal clues. To address the two issues, this paper proposes a novel learning based approach to RGB-D grasp detection, namely Depth Guided Cross-modal Attention Network (DGCAN). To better leverage the geometry information recorded in the depth channel, a complete 6-dimensional rectangle representation is adopted with the grasp depth dedicatedly considered in addition to those defined in the common 5-dimensional one. The prediction of the extra grasp depth substantially strengthens feature learning, thereby leading to more accurate results. Moreover, to reduce the negative impact caused by the discrepancy of data quality in two modalities, a Local Cross-modal Attention (LCA) module is designed, where the depth features are refined according to cross-modal relations and concatenated to the RGB ones for more sufficient fusion. Extensive simulation and physical evaluations are conducted and the experimental results highlight the superiority of the proposed approach.
Ran Qin, Boyang Gao, Di Huang 0001
ICRA4
2023 MIEP: Channel Pruning with Multi-granular Importance Estimation for Object Detection
abstract
This paper investigates compressing a pre-trained deep object detector to a lightweight one by channel pruning, which has proved effective and flexible in promoting efficiency. However, the majority of existing works trim channels based on a monotonous criterion for general purposes, i.e., the importance to the task-specific loss. They are prone to overly prune intermediate layers and simultaneously leave large intra-layer redundancy, severely deteriorating the detection accuracy. To address the issues above, we propose a novel channel pruning approach with multi-granular importance estimation (MIEP), consisting of the Feature-level Object-sensitive Importance (FOI) and the Intra-layer Redundancy-aware Importance (IRI). The former puts large weights on channels that are critical for object representation through the guidance of object features from the pre-trained model, and mitigates over-pruning when combined with the task-specific loss. The latter groups highly correlated channels based on clustering, which are subsequently pruned with priority to decrease redundancy. Extensive experiments on the COCO and VOC benchmarks demonstrate that MIEP remarkably outperforms the state-of-the-art channel pruning approaches, achieves a better balance between accuracy and efficiency compared to lightweight object detectors, and generalizes well to various detection frameworks (e.g., Faster-RCNN and FSAF) and tasks (e.g., classification).
Liangwei Jiang, Jiaxin Chen 0002, Di Huang 0001, Yunhong Wang 0001
ACM Multimedia3
2023 Compressed Video Prompt Tuning
abstract
Compressed videos offer a compelling alternative to raw videos, showing the possibility to significantly reduce the on-line computational and storage cost. However, current approaches to compressed video processing generally follow the resource-consuming pre-training and fine-tuning paradigm, which does not fully take advantage of such properties, making them not favorable enough for widespread applications. Inspired by recent successes of prompt tuning techniques in computer vision, this paper presents the first attempt to build a prompt based representation learning framework, which enables effective and efficient adaptation of pre-trained raw video models to compressed video understanding tasks. To this end, we propose a novel prompt tuning approach, namely Compressed Video Prompt Tuning (CVPT), emphatically dealing with the challenging issue caused by the inconsistency between pre-training and downstream data modalities. Specifically, CVPT replaces the learnable prompts with compressed modalities (\emph{e.g.} Motion Vectors and Residuals) by re-parameterizing them into conditional prompts followed by layer-wise refinement. The conditional prompts exhibit improved adaptability and generalizability to instances compared to conventional individual learnable ones, and the Residual prompts enhance the noisy motion cues in the Motion Vector prompts for further fusion with the visual cues from I-frames. Additionally, we design Selective Cross-modal Complementary Prompt (SCCP) blocks. After inserting them into the backbone, SCCP blocks leverage semantic relations across diverse levels and modalities to improve cross-modal interactions between prompts and input flows. Extensive evaluations on HMDB-51, UCF-101 and Something-Something v2 demonstrate that CVPT remarkably outperforms the state-of-the-art counterparts, delivering a much better balance between accuracy and efficiency.
Jiaxin Chen 0002, Xiuguo Bao, Di Huang 0001
NeurIPS4
2023 Beyond 3DMM: Learning to Capture High-Fidelity 3D Face Shape
abstract
3D Morphable Model (3DMM) fitting has widely benefited face analysis due to its strong 3D priori. However, previous reconstructed 3D faces suffer from degraded visual verisimilitude due to the loss of fine-grained geometry, which is attributed to insufficient ground-truth 3D shapes, unreliable training strategies and limited representation power of 3DMM. To alleviate this issue, this paper proposes a complete solution to capture the personalized shape so that the reconstructed shape looks identical to the corresponding person. Specifically, given a 2D image as the input, we virtually render the image in several calibrated views to normalize pose variations while preserving the original image geometry. A many-to-one hourglass network serves as the encode-decoder to fuse multiview features and generate vertex displacements as the fine-grained geometry. Besides, the neural network is trained by directly optimizing the visual effect, where two 3D shapes are compared by measuring the similarity between the multiview images rendered from the shapes. Finally, we propose to generate the ground-truth 3D shapes by registering RGB-D images followed by pose and shape augmentation, providing sufficient data for network training. Experiments on several challenging protocols demonstrate the superior reconstruction accuracy of our proposal on the face shape.
Xiangyu Zhu 0001, Chang Yu 0001, Di Huang 0001, Zhen Lei 0001, Hao Wang 0074, Stan Z. Li
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Facial Expression Animation by Landmark Guided Residual Module
abstract
We study the problem of facial expression animation from a still image according to a driving video. This is a challenging task as expression motions are non-rigid and very subtle to be captured. Existing methods mostly fail to model these subtle expression motions, leading to the lack of details in their animation results. In this paper, we propose a novel facial expression animation method based on generative adversarial learning. To capture the subtle expression motions, Landmark guided Residual Module (LRM) is proposed to model detailed facial expression features. Specifically, residual learning is conducted at both coarse and fine levels conditioned on facial landmark heatmaps and landmark points respectively. Furthermore, we employ a consistency discriminator to ensure the temporal consistency of the generated video sequence. In addition, a novel metric named Emotion Consistency Metric is proposed to evaluate the consistency of facial expressions in the generated sequences with those in the driving videos. Experiments on MUG-Face, Oulu-CASIA and CAER datasets show that the proposed method can generate arbitrary expression motions on the source still image effectively, which are more photo-realistic and consistent with the driving video compared with results of state-of-the-art methods.
Yunhong Wang 0001, Weixin Li 0001, Zhengyin Du, Di Huang 0001
IEEE Trans. Affect. Comput.5
2023 Group Activity Representation Learning With Long-Short States Predictive Transformer
abstract
The research goal of this paper is to learn the group activity representations in a self-supervised fashion instead of through the use of conventional methods that rely on manually annotated labels. It is essential for this task to better describe the complex group states and their future transitions. To this end, we propose a long-short state predictive Transformer (LSSPT), which mines the meaningful spatiotemporal features of group activities by predicting the future group states with long- and short-term historical state dynamics. LSSPT consists of an encoder that models diverse spatiotemporal state representations in the observation, together with a decoder that exploits rich dynamic patterns by attending to both the short-term spatial context and long-term history state evolutions to predict future group states. Furthermore, we consider the distinguishability and consistency of the predicted states and introduce a joint learning mechanism to optimize the models, enabling LSSPT to describe more reliable state transitions. Finally, extensive experiments are carried out to evaluate the learned representation on downstream tasks on the Volleyball, Collective Activity and VolleyTactic datasets, which showcases the method’s state-of-the-art performance over the existing self-supervised learning approaches.
Longteng Kong, Duoxuan Pei, Zhaofeng He 0001, Di Huang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Key Role Guided Transformer for Group Activity Recognition
abstract
Group Activity Recognition (GAR) is a challenging task, where modeling spatio-temporal relationships among participants plays a fundamental role. To address this issue, we propose a novel end-to-end trainable network, termed Key Role Guided Transformer (KRGFormer). Different from current methods that concurrently take all individuals into account for global reasoning, it captures crucial contextual information by emphasizing a set of key individuals in a coarse-to-fine manner considering that group activities are usually dominated by them. A Key Individual-aware Block (KIaBlock) is designed to select relevant individuals and enhance their relationships with the reservation of global dependencies of the entire group. The representations are then iteratively refined by deploying multiple stacked KIaBlocks, leading to a stronger discriminative power to distinguish group activities. Moreover, along with general data augmentation schemes, several “actor-centric” ones are presented to relieve the over-fitting risk, which further boost the performance. We extensively evaluate the proposed approach on the Volleyball, VolleyTactic and NBA datasets, and the experimental results demonstrate its superiority.
Duoxuan Pei, Di Huang 0001, Longteng Kong, Yunhong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 MCTAN: A Novel Multichannel Temporal Attention-Based Network for Industrial Health Indicator Prediction
abstract
Health indicator prediction, such as remaining useful life prediction and product quality prediction, is an important aspect of industrial intelligence. It is essential to process the massive multichannel industrial time series collected from the Industrial Internet of Things for the industrial health indicator prediction. At present, there are still three issues that need to be considered for industrial health indicator prediction. First, it is difficult to directly connect the distant positions in the industrial time series to extract the temporal relations, which decreases the efficiency of extracting the potential long-distance temporal relations and training networks. Second, it should be fully considered that data from different channels have different contributions. Equally dealing with the contributions of each channel will weaken the representational ability of prediction networks. Third, the loss function deals with early predictions and delay predictions equally, which will lead to high risks caused by delay predictions. In this article, for these issues, a novel multichannel temporal attention-based network (MCTAN) is proposed for industrial health indicator prediction, which can weigh contributions of different channels through the channel attention while avoiding the loss of the temporal information and directly connect each time series position to the local fields of the sequence through the multi-head local attention mechanism to efficiently extract potential long-distance temporal relations. Then, a weighted mean square error loss function differently dealing with early predictions and delay predictions by setting dynamic weights is presented to reduce delay predictions. Next, to deal with the above-mentioned issues systematically, a framework combining data preprocessing and MCTAN collaboratively is introduced to predict industrial health indicators through multichannel time series. Finally, the experiments are carried out on the commercial modular aero-propulsion system simulation dataset to measure the performances, including the accuracy of industrial health indicator predictions and the inference speed.
Lei Ren 0001, Yuxin Liu 0004, Di Huang 0001, Keke Huang, Chunhua Yang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 iDARTS: Improving DARTS by Node Normalization and Decorrelation Discretization
abstract
Differentiable ARchiTecture Search (DARTS) uses a continuous relaxation of network representation and dramatically accelerates Neural Architecture Search (NAS) by almost thousands of times in GPU-day. However, the searching process of DARTS is unstable, which suffers severe degradation when training epochs become large, thus limiting its application. In this article, we claim that this degradation issue is caused by the imbalanced norms between different nodes and the highly correlated outputs from various operations. We then propose an improved version of DARTS, namely iDARTS, to deal with the two problems. In the training phase, it introduces node normalization to maintain the norm balance. In the discretization phase, the continuous architecture is approximated based on the similarity between the outputs of the node and the decorrelated operations rather than the values of the architecture parameters. Extensive evaluation is conducted on CIFAR-10 and ImageNet, and the error rates of 2.25% and 24.7% are reported within 0.2 and 1.9 GPU-day for architecture search, respectively, which shows its effectiveness. Additional analysis also reveals that iDARTS has the advantage in robustness and generalization over other DARTS-based counterparts.
Huiqun Wang, Ruijie Yang, Di Huang 0001, Yunhong Wang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 UFPMP-Det: Toward Accurate and Efficient Object Detection on Drone Imagery
abstract
This paper proposes a novel approach to object detection on drone imagery, namely Multi-Proxy Detection Network with Unified Foreground Packing (UFPMP-Det). To deal with the numerous instances of very small scales, different from the common solution that divides the high-resolution input image into quite a number of chips with low foreground ratios to perform detection on them each, the Unified Foreground Packing (UFP) module is designed, where the sub-regions given by a coarse detector are initially merged through clustering to suppress background and the resulting ones are subsequently packed into a mosaic for a single inference, thus significantly reducing overall time cost. Furthermore, to address the more serious confusion between inter-class similarities and intra-class variations of instances, which deteriorates detection performance but is rarely discussed, the Multi-Proxy Detection Network (MP-Det) is presented to model object distributions in a fine-grained manner by employing multiple proxy learning, and the proxies are enforced to be diverse by minimizing a Bag-of-Instance-Words (BoIW) guided optimal transport loss. By such means, UFPMP-Det largely promotes both the detection accuracy and efficiency. Extensive experiments are carried out on the widely used VisDrone and UAVDT datasets, and UFPMP-Det reports new state-of-the-art scores at a much higher speed, highlighting its advantages. The code is available at https://github.com/PuAnysh/UFPMP-Det.
Yecheng Huang, Jiaxin Chen 0002, Di Huang 0001
AAAI3
2022 ACGNet: Action Complement Graph Network for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization (WTAL) in untrimmed videos has emerged as a practical but challenging task since only video-level labels are available. Existing approaches typically leverage off-the-shelf segment-level features, which suffer from spatial incompleteness and temporal incoherence, thus limiting their performance. In this paper, we tackle this problem from a new perspective by enhancing segment-level representations with a simple yet effective graph convolutional network, namely action complement graph network (ACGNet). It facilitates the current video segment to perceive spatial-temporal dependencies from others that potentially convey complementary clues, implicitly mitigating the negative effects caused by the two issues above. By this means, the segment-level features are more discriminative and robust to spatial-temporal variations, contributing to higher localization accuracies. More importantly, the proposed ACGNet works as a universal module that can be flexibly plugged into different WTAL frameworks, while maintaining the end-to-end training fashion. Extensive experiments are conducted on the THUMOS'14 and ActivityNet1.2 benchmarks, where the state-of-the-art results clearly demonstrate the superiority of the proposed approach.
Zichen Yang, Jie Qin 0004, Di Huang 0001
AAAI3
2022 ABPN: Adaptive Blend Pyramid Network for Real-Time Local Retouching of Ultra High-Resolution Photo
abstract
Photo retouching finds many applications in various fields. However, most existing methods are designed for global retouching and seldom pay attention to the local region, while the latter is actually much more tedious and time-consuming in photography pipelines. In this paper, we propose a novel adaptive blend pyramid network, which aims to achieve fast local retouching on ultra high-resolution photos. The network is mainly composed of two components: a context-aware local retouching layer (LRL) and an adaptive blend pyramid layer (BPL). The LRL is designed to implement local retouching on low-resolution images, giving full consideration of the global context and local texture information, and the BPL is then developed to progressively expand the low-resolution results to the higher ones, with the help of the proposed adaptive blend module and refining module. Our method outperforms the existing methods by a large margin on two local photo retouching tasks and exhibits excellent performance in terms of running speed, achieving real-time inference on 4K images with a single NVIDIA Tesla P100 GPU. Moreover, we introduce the first high-definition cloth retouching dataset CRHD-3K to promote the research on local photo retouching. The dataset is available at https://github.com/youngLbw/crhd-3K.
Biwen Lei, Xiefan Guo, Hongyu Yang 0001, Miaomiao Cui, Xuansong Xie, Di Huang 0001
CVPR6
2022 Entropy-based Active Learning for Object Detection with Progressive Diversity Constraint
abstract
Active learning is a promising alternative to alleviate the issue of high annotation cost in the computer vision tasks by consciously selecting more informative samples to label. Active learning for object detection is more challenging and existing efforts on it are relatively rare. In this paper, we propose a novel hybrid approach to address this problem, where the instance-level uncertainty and diversity are jointly considered in a bottom-up manner. To balance the computational complexity, the proposed approach is designed as a two-stage procedure. At the first stage, an Entropy-based Non-Maximum Suppression (ENMS) is presented to estimate the uncertainty of every image, which performs NMS according to the entropy in the feature space to remove predictions with redundant information gains. At the second stage, a diverse prototype (DivProto) strategy is explored to ensure the diversity across images by progressively converting it into the intra-class and inter-class diversities of the entropy-based class-specific prototypes. Extensive experiments are conducted on MS COCO and Pascal VOC, and the proposed approach achieves state of the art results and significantly outperforms the other counter-parts, highlighting its superiority.
Jiaxin Chen 0002, Di Huang 0001
CVPR3
2022 Target-Relevant Knowledge Preservation for Multi-Source Domain Adaptive Object Detection
abstract
Domain adaptive object detection (DAOD) is a promising way to alleviate performance drop of detectors in new scenes. Albeit great effort made in single source domain adaptation, a more generalized task with multiple source domains remains not being well explored, due to knowledge degradation during their combination. To address this issue, we propose a novel approach, namely target-relevant knowledge preservation (TRKP), to unsupervised multi-source DAOD. Specifically, TRKP adopts the teacher-student framework, where the multi-head teacher network is built to extract knowledge from labeled source domains and guide the student network to learn detectors in unlabeled target domain. The teacher network is further equipped with an adversarial multi-source disentanglement (AMSD) module to preserve source domain-specific knowledge and simultaneously perform cross-domain alignment. Besides, a holistic target-relevant mining (HTRM) scheme is developed to re-weight the source images according to the source-target relevance. By this means, the teacher network is enforced to capture target-relevant knowledge, thus benefiting decreasing domain shift when mentoring object detection in the target domain. Extensive experiments are conducted on various widely used benchmarks with new state-of-the-art scores reported, highlighting the effectiveness.
Jiaxin Chen 0002, Mengzhe He, Yiru Wang 0003, Bo Li 0114, Bingqi Ma, Weihao Gan, Wei Wu 0021, Yali Wang 0001, Di Huang 0001
CVPR10
2022 CAT-Det: Contrastively Augmented Transformer for Multimodal 3D Object Detection
abstract
In autonomous driving, LiDAR point-clouds and RGB images are two major data modalities with complementary cues for 3D object detection. However, it is quite difficult to sufficiently use them, due to large inter-modal discrepancies. To address this issue, we propose a novel framework, namely Contrastively Augmented Transformer for multi-modal 3D object Detection (CAT-Det). Specifically, CAT-Det adopts a two-stream structure consisting of a Pointformer (PT) branch, an Imageformer (IT) branch along with a Cross-Modal Transformer (CMT) module. PT, IT and CMT jointly encode intra-modal and inter-modal long-range contexts for representing an object, thus fully exploring multi-modal information for detection. Furthermore, we propose an effective One-way Multimodal Data Augmentation (OMDA) approach via hierarchical contrastive learning at both the point and object levels, significantly improving the accuracy only by augmenting point-clouds, which is free from complex generation of paired samples of the two modalities. Extensive experiments on the KITTI benchmark show that CAT-Det achieves a new state-of-the-art, highlighting its effectiveness.
Yanan Zhang 0005, Jiaxin Chen 0002, Di Huang 0001
CVPR3
2022 ImFace: A Nonlinear 3D Morphable Face Model with Implicit Neural Representations
abstract
Precise representations of 3D faces are beneficial to various computer vision and graphics applications. Due to the data discretization and model linearity, however, it remains challenging to capture accurate identity and expression clues in current studies. This paper presents a novel 3D morphable face model, namely ImFace, to learn a nonlinear and continuous space with implicit neural representations. It builds two explicitly disentangled deformation fields to model complex shapes associated with identities and expressions, respectively, and designs an improved learning strategy to extend embeddings of expressions to allow more diverse changes. We further introduce a Neural Blend-Field to learn sophisticated details by adaptively blending a series of local fields. In addition to ImFace, an effective pre-processing pipeline is proposed to address the issue of watertight input requirement in implicit representations, enabling them to work with common facial surfaces for the first time. Extensive experiments are performed to demonstrate the superiority of ImFace.
Mingwu Zheng, Hongyu Yang 0001, Di Huang 0001, Liming Chen 0002
CVPR3
2022 Motion Sensitive Contrastive Learning for Self-supervised Video Representation
Jingcheng Ni, Jie Qin 0004, Boxun Li, Di Huang 0001
ECCV (35)7
2022 Video Anomaly Detection by Solving Decoupled Spatio-Temporal Jigsaw Puzzles
Guodong Wang 0006, Yunhong Wang 0001, Jie Qin 0004, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001
ECCV (10)6
2022 Non-Deterministic Face Mask Removal Based on 3d Priors
abstract
This paper presents a novel image inpainting framework for face mask removal. Although current methods have demonstrated their impressive ability in recovering damaged face images, they suffer from two main problems: the dependence on manually labeled missing regions and the deterministic result corresponding to each input. The proposed approach tackles these problems by integrating a multi-task 3D face reconstruction module with a face inpainting module. Given a masked face image, the former predicts a 3DMM-based reconstructed face together with a binary occlusion map, providing dense geometrical and textural priors that greatly facilitate the inpainting task of the latter. By gradually controlling the 3D shape parameters, our method generates high-quality dynamic in-painting results with different expressions and mouth movements. Qualitative and quantitative experiments verify the effectiveness of the proposed method. Our code: https://github.com/face3d0725/face_de_mask
Xiangnan Yin, Di Huang 0001, Liming Chen 0002
ICIP2
2022 Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
abstract
Compressed video action recognition has recently drawn growing attention, since it remarkably reduces the storage and computational cost via replacing raw videos by sparsely sampled RGB frames and compressed motion cues (e.g., motion vectors and residuals). However, this task severely suffers from the coarse and noisy dynamics and the insufficient fusion of the heterogeneous RGB and motion modalities. To address the two issues above, this paper proposes a novel framework, namely Attentive Cross-modal Interaction Network with Motion Enhancement (MEACI-Net). It follows the two-stream architecture, i.e. one for the RGB modality and the other for the motion modality. Particularly, the motion stream employs a multi-scale block embedded with a denoising module to enhance representation learning. The interaction between the two streams is then strengthened by introducing the Selective Motion Complement (SMC) and Cross-Modality Augment (CMA) modules, where SMC complements the RGB modality with spatio-temporally attentive local motion features and CMA further combines the two modalities with selective feature augmentation. Extensive experiments on the UCF-101, HMDB-51 and Kinetics-400 benchmarks demonstrate the effectiveness and efficiency of MEACI-Net.
Jiaxin Chen 0002, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001
IJCAI5
2022 Multi-view Gait Video Synthesis
abstract
This paper investigates a new fine-grained video generation task, namely Multi-view Gait Video Synthesis, where the generation model works on a video of a walking human of arbitrary viewpoint and creates multi-view renderings of the subject. This task is particularly challenging, as it requires synthesizing visually plausible results, while simultaneously preserving discriminative gait cues subject to identification. To tackle the challenge caused by the entanglement of viewpoint, texture, and body structure, we present a network with two collaborative branches to decouple the novel view rendering process into two streams for human appearances (texture) and silhouettes (structure), respectively. Additionally, the prior knowledge of person re-identification and gait recognition is incorporated into the training loss for more adequate and accurate dynamic details. Experimental results show that the presented method is able to achieve promising success rates when attacking state-of-the-art gait recognition models. Furthermore, the method can improve gait recognition systems by effective data augmentation. To the best of our knowledge, this is the first task to manipulate views for human videos with person-specific behavioral constraints.
Weilai Xiang, Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001
ACM Multimedia3
2022 GridNet: efficiently learning deep hierarchical representation for 3D point cloud understanding
Huiqun Wang, Di Huang 0001, Yunhong Wang 0001
Frontiers Comput. Sci.2
2022 Spatio-Temporal Player Relation Modeling for Tactic Recognition in Sports Videos
abstract
Tactic recognition in sports videos is a challenging task. To address this, we present a novel spatio-temporal relation modeling approach, which captures both detailed player interactions and long-range group dynamics in tactics. In spatial modeling, we propose an Adaptive Graph Convolutional Network (A-GCN), and it represents individual and common patterns of data through local and global graphs to learn diverse player interactions. In temporal modeling, we propose an Attentive Temporal Convolutional Network (A-TCN) and with spatial configurations as input, it builds group dynamics and is robust to redundant content by considering sequence dependencies. Due to adaptive interaction and attentive dynamics modeling, our approach is able to comprehensively describe team cooperation over time in a tactic. We extensively evaluate the proposed approach on the Volleyball dataset and a newly collected VolleyTactic dataset, and the experimental results show its advantage.
Longteng Kong, Duoxuan Pei, Di Huang 0001, Yunhong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 MRDet: A Multihead Network for Accurate Rotated Object Detection in Aerial Images
abstract
Objects in aerial images usually have arbitrary orientations and are densely located over the ground, making them extremely challenge to be detected. Many of the recent developed methods attempt to solve these issues by estimating an extra orientation parameter and placing dense anchors, which will result in high model complexity and computational costs. In this article, we propose an arbitrary-oriented region proposal network (AO-RPN) to generate oriented proposals transformed from horizontal anchors. The AO-RPN is very efficient with only a few amounts of parameters increase than the original RPN. Furthermore, to obtain accurate bounding boxes, we decouple the detection task into multiple subtasks and propose a multihead network to accomplish them. Each head is specially designed to learn the features optimal for the corresponding task, which allows our network to detect objects accurately. We name it multihead rotated object detector (MRDet). We evaluate the performance of the proposed MRDet on two challenging benchmarks, i.e., DOTA and HRSC2016, and compare it with several state-of-the-art methods. Our method achieves very promising results, which clearly demonstrates its effectiveness. Code has been available athttps://github.com/qinr/MRDet.
Ran Qin, Qingjie Liu 0001, Guangshuai Gao, Di Huang 0001, Yunhong Wang 0001
IEEE Trans. Geosci. Remote. Sens.4
2021 PC-RGNN: Point Cloud Completion and Graph Neural Network for 3D Object Detection
abstract
LiDAR-based 3D object detection is an important task for autonomous driving and current approaches suffer from sparse and partial point clouds caused by distant and occluded objects. In this paper, we propose a novel two-stage framework, namely PC-RGNN, which deals with these challenges by two specific solutions. On the one hand, we introduce a point cloud completion module to recover high-quality proposals of dense points and entire view with original structures preserved. On the other hand, a graph neural network module, is designed, which comprehensively captures relations among points by the local-global attention mechanism as well as the multi-scale graph based context aggregation and substantially strengthens encoded features. Extensive experiments on the KITTI benchmark show that the proposed approach outperforms the previous state-of-the-art baselines by remarkable margins, highlighting its effectiveness.
Yanan Zhang 0005, Di Huang 0001, Yunhong Wang 0001
AAAI2
2021 Boundary Guided Context Aggregation for Semantic Segmentation
Hongyu Yang 0001, Di Huang 0001
BMVC3
2021 Student-Teacher Feature Pyramid Matching for Anomaly Detection
Guodong Wang 0006, Shumin Han, Errui Ding, Di Huang 0001
BMVC4
2021 Expression-Latent-Space-guided GAN for Facial Expression Animation based on Discrete Labels
abstract
Facial expression animation aims to synthesize face images that correspond to the target expression in a continuum. This is a challenging task because the animation not only cares about the smooth transition in the generated sequence but also needs to take the facial expression and identity details into consideration. Most existing expression animation methods resort to continuous expression labels, e.g. Action Units (AUs) or landmark sequences. Compared with discrete expression labels, the annotations of a considerable part of them are ambiguous and prone to errors. However, how to animate facial expression conditioned on discrete expression labels is less investigated and existing methods cannot generate satisfactory facial details. To tackle these problems, we propose an end-to-end Expression-Latent-Space-guided Generative Adversarial Network (ELS-GAN) model, which utilizes discrete expression labels as input to generate images with expected expressions, and employs expression latent space learning to control the expression changing process. An expression ranking loss is also proposed to strengthen expression intensity learning during generation. Moreover, we put forward a Self-Attention Generator to synthesize face images with fine details by considering both local areas and the long-range dependency of different areas. Extensive experiments show that our method can generate continuous intermediate expression between source and target expressions only conditioned on discrete labels and superior results are achieved compared with state-of-the-art methods.
Weixin Li 0001, Di Huang 0001
FG3
2021 Human-Aware Coarse-to-Fine Online Action Detection
abstract
In this work, we propose a two-stage framework to efficiently and effectively detect actions on-the-fly. An action location network (ALN) is developed in the first stage to judge whether the current frame is action-related, while the second stage involves an action classification network (ACN) to further identify the action category. In this way, irrelevant negative frames are quickly discarded and actions are detected as early as they occur. Moreover, we highlight human areas at both the stages by respectively incorporating a human detector and a human mask layer. As a result, more accurate spatial-temporal windows of actions are detected, based on which more robust features are extracted for classification. Experimental results on two popular benchmarks demonstrate the superior performance of the proposed approach.
Zichen Yang, Di Huang 0001, Jie Qin 0004, Yunhong Wang 0001
ICASSP2
2021 Refining Single Low-Quality Facial Depth Map by Lightweight and Efficient Deep Model
abstract
Consumer depth sensors have become increasingly common, however, the data are rather coarse and noisy, which is problematic to delicate tasks, such as 3D face modeling and 3D face recognition. In this paper, we present a novel and lightweight 3D Face Refinement Model (3D-FRM), to effectively and efficiently improve the quality of such single facial depth maps. 3D-FRM has an encoder-decoder structure, where the encoder applies depth-wise, point-wise convolutions and the fusion of features of different receptive fields to capture original discriminative information, and the decoder exploits sub-pixel convolutions and the combination of low- and high-level features to achieve strong shape recovery. We also propose a joint loss function to smooth facial surfaces and preserve their identities. In addition, we contribute a large dataset with low- and high-quality 3D face pairs to facilitate this research. Extensive experiments are conducted on the Bosphorus and Lock3DFace datasets, and results show the competency of the proposed method at ameliorating both visual quality and recognition accuracy. Code and data will be available at https://github.com/muyouhang/3D-FRM.
Guodong Mu, Di Huang 0001, Weixin Li 0001, Guosheng Hu, Yunhong Wang 0001
IJCB2
2021 Image Inpainting via Conditional Texture and Structure Dual Generation
abstract
Deep generative approaches have recently made considerable progress in image inpainting by introducing structure priors. Due to the lack of proper interaction with image texture during structure reconstruction, however, current solutions are incompetent in handling the cases with large corruptions, and they generally suffer from distorted results. In this paper, we propose a novel two-stream network for image inpainting, which models the structure-constrained texture synthesis and texture-guided structure reconstruction in a coupled manner so that they better leverage each other for more plausible generation. Furthermore, to enhance the global consistency, a Bi-directional Gated Feature Fusion (Bi-GFF) module is designed to exchange and combine the structure and texture information and a Contextual Feature Aggregation (CFA) module is developed to refine the generated contents by region affinity learning and multi-scale feature aggregation. Qualitative and quantitative experiments on the CelebA, Paris StreetView and Places2 datasets demonstrate the superiority of the proposed method. Our code is available at https://github.com/Xiefan-Guo/CTSDG.
Xiefan Guo, Hongyu Yang 0001, Di Huang 0001
ICCV3
2021 PR-GCN: A Deep Graph Convolutional Network with Point Refinement for 6D Pose Estimation
abstract
RGB-D based 6D pose estimation has recently achieved remarkable progress, but still suffers from two major limitations: (1) ineffective representation of depth data and (2) insufficient integration of different modalities. This paper proposes a novel deep learning approach, namely Graph Convolutional Network with Point Refinement (PR-GCN), to simultaneously address the issues above in a unified way. It first introduces the Point Refinement Network (PRN) to polish 3D point clouds, recovering missing parts with noise removed. Subsequently, the Multi-Modal Fusion Graph Convolutional Network (MMF-GCN) is presented to strengthen RGB-D combination, which captures geometry-aware inter-modality correlation through local information propagation in the graph convolutional network. Extensive experiments are conducted on three widely used benchmarks, and state-of-the-art performance is reached. Besides, it is also shown that the proposed PRN and MMF-GCN modules are well generalized to other frameworks.
Guangyuan Zhou, Huiqun Wang, Jiaxin Chen 0002, Di Huang 0001
ICCV4
2021 Multi-Scale Background Suppression Anomaly Detection In Surveillance Videos
abstract
Video anomaly detection has been widely applied in various surveillance systems for public security. However, the existing weakly supervised video anomaly detection methods tend to ignore the interference of the background frames and possess limited ability to extract effective temporal information among the video snippets. In this paper, a multi-scale background suppression based anomaly detection (MSBSAD) method is proposed to suppress the interference of the background frames. We propose a multi-scale temporal convolution module to effectively extract more temporal information among the video snippets for the anomaly events with different durations. A modified hinge loss is constructed in the suppression branch to help our model to better differentiate the abnormal samples from the confusing samples. Experiments on UCF Crime demonstrate the superiority of our MS-BSAD method in the video anomaly detection task.
Yuanfang Guo, Jinjie Wei, Xiuguo Bao, Di Huang 0001
ICIP5
2021 Double-Dot Network for Antipodal Grasp Detection
abstract
This paper proposes a new deep learning approach to antipodal grasp detection, named Double-Dot Network (DD-Net). It follows the recent anchor-free object detection framework, which does not depend on empirically pre-set anchors and thus allows more generalized and flexible prediction on unseen objects. Specifically, unlike the widely used 5-dimensional rectangle, the gripper configuration is defined as a pair of fingertips. An effective CNN architecture is introduced to localize such fingertips, and with the help of auxiliary centers for refinement, it accurately and robustly infers grasp candidates. Additionally, we design a specialized loss function to measure the quality of grasps, and in contrast to the IoU scores of bounding boxes adopted in object detection, it is more consistent to the grasp detection task. Both the simulation and robotic experiments are executed and state of the art accuracies are achieved, showing that DD-Net is superior to the counterparts in handling unseen objects.
Yangtao Zheng, Boyang Gao, Di Huang 0001
IROS4
2021 Identity-aware Graph Memory Network for Action Detection
abstract
Action detection plays an important role in high-level video understanding and media interpretation. Many existing studies fulfill this spatio-temporal localization by modeling the context, capturing the relationship of actors, objects, and scenes conveyed in the video. However, they often universally treat all the actors without considering the consistency and distinctness between individuals, leaving much room for improvement. In this paper, we explicitly highlight the identity information of the actors in terms of both long-term and short-term context through a graph memory network, namely identity-aware graph memory network (IGMN). Specifically, we propose the hierarchical graph neural network (HGNN) to comprehensively conduct long-term relation modeling within the same identity as well as between different ones. Regarding short-term context, we develop a dual attention module (DAM) to generate identity-aware constraint to reduce the influence of interference by the actors of different identities. Extensive experiments on the challenging AVA dataset demonstrate the effectiveness of our method, which achieves state-of-the-art results on AVA v2.1 and v2.2.
Jingcheng Ni, Jie Qin 0004, Di Huang 0001
ACM Multimedia3
2021 Latent Memory-augmented Graph Transformer for Visual Storytelling
abstract
Visual storytelling aims to automatically generate a human-like short story given an image stream. Most existing works utilize either scene-level or object-level representations, neglecting the interaction among objects in each image and the sequential dependency between consecutive images. In this paper, we present a novel Latent Memory-augmented Graph Transformer~(LMGT ), a Transformer based framework for visual story generation. LMGT directly inherits the merits from the Transformer, which is further enhanced with two carefully designed components, i.e., a graph encoding module and a latent memory unit. Specifically, the graph encoding module exploits the semantic relationships among image regions and attentively aggregates critical visual features based on the parsed scene graphs. Furthermore, to better preserve inter-sentence coherence and topic consistency, we introduce an augmented latent memory unit that learns and records highly summarized latent information as the story line from the image stream and the sentence history. Experimental results on three widely-used datasets demonstrate the superior performance of LMGT over the state-of-the-art methods.
Mengshi Qi, Jie Qin 0004, Di Huang 0001, Yi Yang 0001, Jiebo Luo 0001
ACM Multimedia3
2021 Intensity enhancement via GAN for multimodal face expression recognition
Hongyu Yang 0001, Kangkang Zhu, Di Huang 0001, Hebeizi Li, Yunhong Wang 0001, Liming Chen 0002
Neurocomputing3
2021 Learning Continuous Face Age Progression: A Pyramid of GANs
abstract
The two underlying requirements of face age progression, i.e., aging accuracy and identity permanence, are not well studied in the literature. This paper presents a novel generative adversarial network based approach to address the issues in a coupled manner. It separately models the constraints for the intrinsic subject-specific characteristics and the age-specific facial changes with respect to the elapsed time, ensuring that the generated faces present desired aging effects while keeping personalized properties stable. To render photo-realistic facial details, high-level age-specific features conveyed by the synthesized face are estimated by a pyramidal adversarial discriminator at multiple scales, which simulates the aging effects in a finer way. Further, an adversarial learning scheme is introduced to simultaneously train a single generator and multiple parallel discriminators, resulting in smooth continuous face aging sequences. The proposed method is applicable even in the presence of variations in pose, expression, makeup, etc., achieving remarkably vivid aging effects. Quantitative evaluations by a COTS face recognition system demonstrate that the target age distributions are accurately recovered, and 99.88 and 99.98 percent age progressed faces can be correctly verified at 0.001 percent FAR after age transformations of approximately 28 and 23 years elapsed time on the MORPH and CACD databases, respectively. Both visual and quantitative assessments show that the approach advances the state-of-the-art.
Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 A survey on dorsal hand vein biometrics
Wei Jia 0001, Bob Zhang 0001, Yang Zhao 0002, Lunke Fei, Wenxiong Kang, Di Huang 0001, Guodong Guo
Pattern Recognit.7
2021 Spatio-Temporal Encoder-Decoder Fully Convolutional Network for Video-Based Dimensional Emotion Recognition
abstract
Video-based dimensional emotion recognition aims to map human affect into the dimensional emotion space based on visual signals, which is a fundamental challenge in affective computing and human-computer interaction. In this paper, we present a novel encoder-decoder framework to tackle this problem. It adopts a fully convolutional design with the cascaded 2D convolution based spatial encoder and 1D convolution based temporal encoder-decoder for joint spatio-temporal modeling. In particular, to address the key issue of capturing discriminative long-term dynamic dependency, our temporal model, referred to as Temporal Hourglass Convolutional Neural Network (TH-CNN), extracts contextual relationship through integrating both low-level encoded and high-level decoded clues. Temporal Intermediate Supervision (TIS) is then introduced to enhance affective representations generated by TH-CNN under a multi-resolution strategy, which guides TH-CNN to learn macroscopic long-term trend and refined short-term fluctuations progressively. Furthermore, thanks to TH-CNN and TIS, knowledge learnt from the intermediate layers also makes it possible to offer customized solutions to different applications by adjusting the decoder depth. Extensive experiments are conducted on three benchmark databases (RECOLA, SEWA and OMG) and superior results are shown compared to state-of-the-art methods, which indicates the effectiveness of the proposed approach.
Zhengyin Du, Suowei Wu, Di Huang 0001, Weixin Li 0001, Yunhong Wang 0001
IEEE Trans. Affect. Comput.3
2020 Distraction-Aware Feature Learning for Human Attribute Recognition via Coarse-to-Fine Attention Mechanism
abstract
Recently, Human Attribute Recognition (HAR) has become a hot topic due to its scientific challenges and application potentials, where localizing attributes is a crucial stage but not well handled. In this paper, we propose a novel deep learning approach to HAR, namely Distraction-aware HAR (Da-HAR). It enhances deep CNN feature learning by improving attribute localization through a coarse-to-fine attention mechanism. At the coarse step, a self-mask block is built to roughly discriminate and reduce distractions, while at the fine step, a masked attention branch is applied to further eliminate irrelevant regions. Thanks to this mechanism, feature learning is more accurate, especially when heavy occlusions and complex backgrounds exist. Extensive experiments are conducted on the WIDER-Attribute and RAP databases, and state-of-the-art results are achieved, demonstrating the effectiveness of the proposed approach.
Mingda Wu, Di Huang 0001, Yuanfang Guo, Yunhong Wang 0001
AAAI2
2020 Cross-domain Object Detection through Coarse-to-Fine Feature Adaptation
abstract
Recent years have witnessed great progress in deep learning based object detection. However, due to the domain shift problem, applying off-the-shelf detectors to an unseen domain leads to significant performance drop. To address such an issue, this paper proposes a novel coarse-to-fine feature adaptation approach to cross-domain object detection. At the coarse-grained stage, different from the rough image-level or instance-level feature alignment used in the literature, foreground regions are extracted by adopting the attention mechanism, and aligned according to their marginal distributions via multi-layer adversarial learning in the common feature space. At the fine-grained stage, we conduct conditional distribution alignment of foregrounds by minimizing the distance of global prototypes with the same category but from different domains. Thanks to this coarse-to-fine feature adaptation, domain knowledge in foreground regions can be effectively transferred. Extensive experiments are carried out in various cross-domain detection scenarios. The results are state-of-the-art, which demonstrate the broad applicability and effectiveness of the proposed approach.
Yangtao Zheng, Di Huang 0001, Yunhong Wang 0001
CVPR2
2020 Multi-scale Positive Sample Refinement for Few-Shot Object Detection
Di Huang 0001, Yunhong Wang 0001
ECCV (16)3
2020 Beyond 3DMM Space: Towards Fine-Grained 3D Face Reconstruction
Xiangyu Zhu 0001, Fan Yang 0062, Di Huang 0001, Chang Yu 0001, Hao Wang 0074, Jianzhu Guo, Zhen Lei 0001, Stan Z. Li
ECCV (8)3
2020 3D Face Mask Anti-spoofing via Deep Fusion of Dynamic Texture and Shape Clues
abstract
Face anti-spoofing has recently become more important to the wide application of Face Recognition (FR) techniques. Compared to Spoofing Attacks (SAs) of printed photos and replayed videos, 3D masks bring more challenges to FR systems. This paper proposes a novel approach to 3D face mask anti-spoofing, namely Multi-Modal Dynamics Fusion Network (MM-DFN), and different from the overwhelming majority of the methods in the literature that only employ RGB data, it highlights the credit of the geometry information delivered by depth sensors or reconstructed from RGB images. Dynamic texture and shape clues are respectively encoded by a two-branch deep CNN model at different rates so that discriminative details are sufficiently captured, and they are combined at intervals for more comprehensive description. Moreover, a 3D model guided data augmentation method is applied to generate a diversity of samples with various poses, which further enhances the anti-spoofing model. The proposed method is extensively evaluated on three public benchmarks, i.e. 3DMAD, HKBU-MARs V1 and SMAD, and the results achieved are state-of-the-art, demonstrating its effectiveness for this issue.
Weixin Li 0001, Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001
FG4
2020 Pixel Sampling for Style Preserving Face Pose Editing
abstract
The existing auto-encoder based face pose editing methods primarily focus on modeling the identity preserving ability during pose synthesis, but are less able to preserve the image style properly, which refers to the color, brightness, saturation, etc. In this paper, we take advantage of the well-known frontal/profile optical illusion and present a novel two-stage approach to solve the aforementioned dilemma, where the task of face pose manipulation is cast into face inpainting. By selectively sampling pixels from the input face and slightly adjust their relative locations with the proposed “Pixel Attention Sampling” module, the face editing result faithfully keeps the identity information as well as the image style unchanged. By leveraging high-dimensional embedding at the inpainting stage, finer details are generated. Further, with the 3D facial landmarks as guidance, our method is able to manipulate face pose in three degrees of freedom, i.e., yaw, pitch, and roll, resulting in more flexible face pose editing than merely controlling the yaw angle as usually achieved by the current state-of-the-art. Both the qualitative and quantitative evaluations validate the superiority of the proposed approach.
Xiangnan Yin, Di Huang 0001, Hongyu Yang 0001, Zehua Fu, Yunhong Wang 0001, Liming Chen 0002
IJCB2
2020 Intensity Enhancement Via Gan for Multimodal Facial Expression Recognition
abstract
Face expression recognition (FER) on low intensity is not well studied in the literature. This paper investigates this new problem and presents a novel Generative Adversarial Network (GAN) based multimodal approach to it. The method models the tasks of intensity enhancement and expression recognition jointly, ensuring that the synthesize faces not only present expression of high intensity, but also truly contribute to promoting the performance of FER. Extensive experiments are conducted on the BU-3DFE and BU-4DFE datasets. State-of-the-art FER performance clearly validates the effectiveness of the proposed method.
Kangkang Zhu, Yunhong Wang 0001, Hongyu Yang 0001, Di Huang 0001, Liming Chen 0002
ICIP4
2020 Face and Gesture Analysis for Health Informatics
abstract
The goal of Face and Gesture Analysis for Health Informatics's workshop is to share and discuss the achievements as well as the challenges in using computer vision and machine learning for automatic human behavior analysis and modeling for clinical research and healthcare applications. The workshop aims to promote current research and support growth of multidisciplinary collaborations to advance this groundbreaking research. The meeting gathers scientists working in related areas of computer vision and machine learning, multi-modal signal processing and fusion, human centered computing, behavioral sensing, assistive technologies, and medical tutoring systems for healthcare applications and medicine.
Zakia Hammal, Di Huang 0001, Kevin Bailly, Liming Chen 0002, Mohamed Daoudi
ICMI2
2020 Towards Practical Compressed Video Action Recognition: A Temporal Enhanced Multi-Stream Network
abstract
Current compressed video action recognition methods are mainly based on complete data. However, in a real transmission scenario, the compressed video packets are usually disorderly received and even lost due to network jitters or congestion. To recognize actions in early phases with limited packets, e.g. for quickly forecasting possible potential risks, in this paper, we propose a Temporal Enhanced Multi-Stream Network (TEMSN) towards practical compressed video action recognition. First, we make use of three modalities in the compressed domain as complementary cues and build a multi-stream network to capture rich information from compressed video packets. Second, we design a temporal enhanced module based on an Encoder-Decoder structure, which is applied to each stream to infer missing packets, generating more accurate action dynamics. Thanks to the multiple modalities and their temporal enhancement, our approach better models actions with partial available compressed video packets. Experiments on the HMDB-51 and UCF-101 datasets validate its effectiveness and efficiency.
Longteng Kong, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001, Yunhong Wang 0001
ICPR5
2020 Few-Shot Ensemble Learning for Video Classification with SlowFast Memory Networks
abstract
In the era of big data, few-shot learning has recently received much attention in multimedia analysis and computer vision due to its appealing ability of learning from scarce labeled data. However, it has been largely underdeveloped in the video domain, which is even more challenging due to the huge spatial-temporal variability of video data. In this paper, we address few-shot video classification by learning an ensemble of SlowFast networks augmented with memory units. Specifically, we introduce a family of few-shot learners based on SlowFast networks which are used to extract informative features at multiple rates, and we incorporate a memory unit into each network to enable encoding and retrieving crucial information instantly. Furthermore, we propose a choice controller network to leverage the diversity of few-shot learners by learning to adaptively assign a confidence score to each SlowFast memory network, leading to a strong classifier for enhanced prediction. Experimental results on two widely-adopted video datasets demonstrate the effectiveness of the proposed method, as well as its superior performance over the state-of-the-art approaches.
Mengshi Qi, Jie Qin 0004, Xiantong Zhen, Di Huang 0001, Yi Yang 0001, Jiebo Luo 0001
ACM Multimedia4
2020 Local Discriminant Direction Binary Pattern for Palmprint Representation and Recognition
abstract
Direction-based methods are the most powerful and popular palmprint recognition methods. However, there is no existing work that completely analyzes the essential differences among different direction-based methods and explores the most discriminant direction representation of a palmprint. In this paper, we attempt to establish the connection between the direction feature extraction model and the discriminability of direction features, and we propose a novel exponential and Gaussian fusion model (EGM) to characterize the discriminative power of different directions. The EGM can provide us with a new insight into the optimal direction feature selection of palmprints. Moreover, we propose a local discriminant direction binary pattern (LDDBP) to completely represent the direction features of a palmprint. Guided by the EGM, the most discriminant directions can be exploited to form the LDDBP-based descriptor for palmprint representation and recognition. Extensive experiment results conducted on four widely used palmprint databases demonstrate the superiority of the proposed LDDBP method over the state-of-the-art direction-based methods.
Lunke Fei, Bob Zhang 0001, Yong Xu 0001, Di Huang 0001, Wei Jia 0001, Jie Wen 0001
IEEE Trans. Circuits Syst. Video Technol.4
2020 A Joint Framework for Athlete Tracking and Action Recognition in Sports Videos
abstract
Sports video analysis has received increasing attention in recent years. Athlete tracking and action recognition are its two major issues that are highly related to each other; however, they are individually considered and processed in the existing studies. In this paper, we propose a joint framework for athlete tracking and action recognition in sports videos. In athlete tracking, we propose a scaling and occlusion robust tracker, named scaling and occlusion robust compressive tracking (CT), to localize the position of specific athlete in each frame. It follows the approach of CT but extends it in two aspects, i.e., scale refinement as well as occlusion recovery. For the former, an objectness method, edge box, is adopted to generate proposals, which replace the fixed sampling boxes in CT and better fit the scales of the candidate objects. For the latter, a candidate obstruction-based solution is presented, which brings in additional trackers to detect possible obstructions and to relocate the target as occlusion ends. Regarding action recognition, we propose a long-term recurrent region-guided convolutional network, which recognizes pre-defined actions by modeling discriminative temporal cues of the tracking results. We employ SPP-net to extract the robust feature of the tracked region of each frame. The features of all the frames are then fed into a stack of recurrent sequence models to capture the long-term region-level information. We extensively evaluate the proposed approach on a newly collected sports video benchmark and on the off-the-shelf UIUC2 dataset, and the experimental results clearly show its effectiveness.
Longteng Kong, Di Huang 0001, Jie Qin 0004, Yunhong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Long-Term Action Dependence-Based Hierarchical Deep Association for Multi-Athlete Tracking in Sports Videos
abstract
Tracking multiple athletes in sports videos is a very challenging Multi-Object Tracking (MOT) task, as athletes generally share high similarity in appearance with large deformations. In this paper, unlike the existing hand-crafted solutions, we propose a novel and effective approach to this issue, which hierarchically associates detections of the same identity through discriminative and robust deep features. First, in detection association, we make use of athlete appearances and poses instead of traditional position cues to generate short tracklets for better initialization. Second, in tracklet association, a new deep architecture, namely Siamese Tracklet Affinity Networks (STAN), is presented, which is able to bi-directionally simulate the unseen dynamics of actions, comprehensively models the long-term action dependences, and sequentially estimates their affinity. Such hierarchical association is finally solved as a minimum-cost network flow problem. We extensively evaluate the proposed approach on the APIDIS, NCAA Basketball and VolleyTrack (newly collected) databases, and the experimental results show its advantages.
Longteng Kong, Di Huang 0001, Yunhong Wang 0001
IEEE Trans. Image Process.2
2020 Pay Attention to Them: Deep Reinforcement Learning-Based Cascade Object Detection
abstract
This paper proposes a novel and effective approach, namely pay attention to them (PAT), to general object detection, which integrates the bottom-up single-shot convolutional neural networks (CNNs) and a top-down operating strategy. PAT starts by routinely applying a CNN regression detector to the entire input image. It then conducts refinement, which locates a sub-region that probably contains relevant objects through an intelligent agent built with an attentional mechanism and zooms it in to launch the detector again. This refining step is repeated in a cascaded way, where all the bounding boxes produced are scaled according to the original resolution and the sub-marginal and overlapping parts are wiped out to generate the final output. Due to such progressive processing, PAT improves the detection accuracy, especially for the objects of small sizes. Extensive experiments are conducted on the Pascal VOC and MS COCO benchmarks, and the results show that PAT is able to improve the representative baseline detectors, i.e., single shot multibox detector, YOLOv2, and Faster regions with CNN features, with remarkable accuracy gains [about 2%-5% mean Average Precision (mAP)], which demonstrates its competency.
Di Huang 0001, Yunhong Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2019 Adaptive NMS: Refining Pedestrian Detection in a Crowd
abstract
Pedestrian detection in a crowd is a very challenging issue. This paper addresses this problem by a novel Non-Maximum Suppression (NMS) algorithm to better refine the bounding boxes given by detectors. The contributions are threefold: (1) we propose adaptive-NMS, which applies a dynamic suppression threshold to an instance, according to the target density; (2) we design an efficient subnetwork to learn density scores, which can be conveniently embedded into both the single-stage and two-stage detectors; and (3) we achieve state of the art results on the CityPersons and CrowdHuman benchmarks.
Di Huang 0001, Yunhong Wang 0001
CVPR2
2019 Led3D: A Lightweight and Efficient Deep Approach to Recognizing Low-Quality 3D Faces
abstract
Due to the intrinsic invariance to pose and illumination changes, 3D Face Recognition (FR) has a promising potential in the real world. 3D FR using high-quality faces, which are of high resolutions and with smooth surfaces, have been widely studied. However, research on that with low-quality input is limited, although it involves more applications. In this paper, we focus on 3D FR using low-quality data, targeting an efficient and accurate deep learning solution. To achieve this, we work on two aspects: (1) designing a lightweight yet powerful CNN; (2) generating finer and bigger training data. For (1), we propose a Multi-Scale Feature Fusion (MSFF) module and a Spatial Attention Vectorization (SAV) module to build a compact and discriminative CNN. For (2), we propose a data processing system including point-cloud recovery, surface refinement, and data augmentation (with newly proposed shape jittering and shape scaling). We conduct extensive experiments on Lock3DFace and achieve state-of-the-art results, outperforming many heavy CNNs such as VGG-16 and ResNet-34. In addition, our model can operate at a very high speed (136 fps) on Jetson TX2, and the promising accuracy and efficiency reached show its great applicability on edge/mobile devices.
Guodong Mu, Di Huang 0001, Guosheng Hu, Yunhong Wang 0001
CVPR2
2019 Encoding Visual Behaviors with Attentive Temporal Convolution for Depression Prediction
abstract
Depression is a common and serious medical illness which has a wide negative impact on individuals, families, and society. Automatic Depression Detection (ADD) is increasingly demanded for human healthcare thanks to its objectiveness, convenience, and low cost. Considering that the duration of depressive symptoms varies among different identities and treatment phases, it is essential for ADD methods to have the capability to capture information at various temporal scales. However, most existing ADD methods cannot generate rich contextual cues or utilize long-range temporal dependency effectively. In this paper, we propose a novel approach for depression recognition based on visual behaviors, which employs Atrous Residual Temporal Convolutional Network (DepArt-Net) as well as temporal fusion to capture the long-range dynamic depressive cues. First, the proposed atrous temporal convolution generates multi-scale contextual features from low-level visual behaviors, which are further strengthened by residual blocks across different convolution groups. Second, we introduce the attention mechanism in temporal feature fusion stage, and with the learned attentive distribution, more discriminative video-level depression representation can be acquired. Experimental results on the DAIC-WOZ benchmark demonstrate the effectiveness of the proposed approach and its superiority over other state-of-the-art methods.
Zhengyin Du, Weixin Li 0001, Di Huang 0001, Yunhong Wang 0001
FG3
2019 Discriminative Attention-based Convolutional Neural Network for 3D Facial Expression Recognition
abstract
3D Facial Expression Recognition (FER) is an active research area in computer vision. Although previous methods report promising results, two key issues still remain to be solved. On the one hand, different facial areas contribute unequally to performing various expressions, but most existing methods extract features from the entire 3D surface. On the other hand, the difference between expressions varies, while previous methods generally treat different emotions equally, making some of them extremely hard to be distinguished. To solve these problems, we propose a novel approach for 3D FER, namely Discriminative Attention-based Convolution Neural Network (DA-CNN), to generate more comprehensive expression related representations. DA-CNN introduces an attention module to the CNN models, which helps the deep model selectively focus on emotional salient regions in a learnable way. Furthermore, a novel loss named Dimensional Distribution (DD) loss is proposed to model the inter-expression relationship. Supervised by DD loss, DA-CNN can generate more discriminative expression representation. Extensive experiments are conducted on BU-3DFE dataset, and the results show that DA-CNN achieves significant improvement over the state-of-the-art.
Kangkang Zhu, Zhengyin Du, Weixin Li 0001, Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
FG4
2019 Continuous Emotion Recognition in Videos by Fusing Facial Expression, Head Pose and Eye Gaze
abstract
Continuous emotion recognition is of great significance in affective computing and human-computer interaction. Most of existing methods for video based continuous emotion recognition utilize facial expression. However, besides facial expression, other clues including head pose and eye gaze are also closely related to human emotion, but have not been well explored in continuous emotion recognition task. On the one hand, head pose and eye gaze could result in different degrees of credibility of facial expression features. On the other hand, head pose and eye gaze carry emotional clues themselves, which are complementary to facial expression. Accordingly, in this paper we propose two ways to incorporate these two clues into continuous emotion recognition. They are respectively an attention mechanism based on head pose and eye gaze clues to guide the utilization of facial features in continuous emotion recognition, and an auxiliary line which helps extract more useful emotion information from head pose and eye gaze. Experiments are conducted on the Recola dataset, a database for continuous emotion recognition, and the results show that our framework outperforms other state-of-the-art methods due to the full use of head pose and eye gaze clues in addition to facial expression for continuous emotion recognition.
Suowei Wu, Zhengyin Du, Weixin Li 0001, Di Huang 0001, Yunhong Wang 0001
ICMI4
2019 Pedestrian Attribute Recognition via Hierarchical Multi-task Learning and Relationship Attention
abstract
Pedestrian Attribute Recognition (PAR) is an important task in surveillance video analysis. In this paper, we propose a novel end-to-end hierarchical deep learning approach to PAR. The proposed network introduces semantic segmentation into PAR and formulates it as a multi-task learning problem, which brings in pixel-level supervision in feature learning for attribute localization. According to the spatial properties of local and global attributes, we present a two stage learning mechanism to decouple coarse attribute localization and fine attribute recognition into successive phases within a single model, which strengthens feature learning. Besides, we design an attribute relationship attention module to efficiently capture and emphasize the latent relations among different attributes, further enhancing the discriminative power of the feature. Extensive experiments are conducted and very competitive results are reached on the RAP and PETA databases, indicating the effectiveness and superiority of the proposed approach.
Lian Gao, Di Huang 0001, Yuanfang Guo, Yunhong Wang 0001
ACM Multimedia2
2019 Attacking Gait Recognition Systems via Silhouette Guided GANs
abstract
This paper investigates a new attack method to gait recognition systems. Different from typical spoofing attacks that require impostors to mimic certain clothing or walking styles, it proposes to intercept the video stream captured by the on-site camera and replace it with synthesized samples. To this end, we present a novel Generative Adversarial Network (GAN) based approach, which is able to render a faked video from the source walking sequence of a specified subject and the target scene image with both good visual effects and sufficient discriminative details. A new generator architecture is built, where the features of the source foreground sequence and the target background image are combined at multiple scales, making the synthesized video vivid. To fool recognition systems, the silhouette-conditioned losses are specially designed to constrain the static and dynamic consistency between the subjects in the source and generated videos. The person re-identification similarity based triplet loss is exploited to guide the generator, which keeps the personalized appearance properties stable. The edge and flow-related losses further regulate the generation of the attacking video. Two state-of-the-art gait recognition systems are used for evaluation, namely GaitSet and CNN-Gait, and we analyze their performance under attacking. Both the visual fidelity and attacking ability of the generated videos validate the effectiveness of the proposed method.
Meijuan Jia, Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001
ACM Multimedia3
2019 Magnifying Subtle Facial Motions for Effective 4D Expression Recognition
abstract
In this paper, an effective approach is proposed for automatic 4D Facial Expression Recognition (FER). It combines two growing but disparate ideas in the domain of computer vision, i.e., computing spatial facial deformations using a Riemannian method and magnifying them by a temporal filtering technique. Key frames highly related to facial expressions are first extracted from a long 4D video through a spectral clustering process, forming the Onset-Apex-Offset flow. It is then analyzed to capture the spatial deformations based on Dense Scalar Fields (DSF), where registration and comparison of neighboring 3D faces are jointly led. The generated temporal evolution of these deformations is further fed into a magnification method to amplify facial activities over time. The proposed approach allows revealing subtle deformations and thus improves the emotion classification performance. Experiments are conducted on the BU-4DFE and BP-4D databases, and competitive results are achieved compared to the state-of-the-art.
Qingkai Zhen, Di Huang 0001, Hassen Drira, Boulbaba Ben Amor, Yunhong Wang 0001, Mohamed Daoudi
IEEE Trans. Affect. Comput.2
2019 Hierarchical Image Segmentation Ensemble for Objectness in RGB-D Images
abstract
Objectness has recently become a standard step in many computer vision tasks. Among various techniques, those based on hierarchical image segmentation play a fundamental role for developments in new data modalities. In this paper, we address the problem of objectness in RGB-D images and propose a novel and effective approach, namely, hierarchical image segmentation ensemble (HISE). Different from existing image segmentation based methods that generate object segments or proposals largely by heuristics or empirical rules, HISE learns superpixel mergings with a hierarchical tree-structured ensemble, where individual merging models of the ensemble are formed by traversing different paths of the tree, and where both the merging accuracy and proposal diversity are emphasized. Furthermore, we use efficient feature measurements that support easy integration of additional clues. Extensive experiments conducted on the benchmark NYU-v2 RGB-D and SUN RGB-D data sets show the competency of our proposed method.
Huiqun Wang, Di Huang 0001, Kui Jia, Yunhong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2019 Expression Robust 3D Facial Landmarking via Progressive Coarse-to-Fine Tuning
abstract
Facial landmarking is a fundamental task in automatic machine-based face analysis. The majority of existing techniques for such a problem are based on 2D images; however, they suffer from illumination and pose variations that may largely degrade landmarking performance. The emergence of 3D data theoretically provides an alternative to overcome these weaknesses in the 2D domain. This article proposes a novel approach to 3D facial landmarking, which combines both the advantages of feature-based methods as well as model-based ones in a progressive three-stage coarse-to-fine manner (initial, intermediate, and fine stages). For the initial stage, a few fiducial landmarks (i.e., the nose tip and two inner eye corners) are robustly detected through curvature analysis, and these points are further exploited to initialize the subsequent stage. For the intermediate stage, a statistical model is learned in the feature space of three normal components of the facial point-cloud rather than the smooth original coordinates, namely Active Normal Model (ANM). For the fine stage, cascaded regression is employed to locally refine the landmarks according to their geometry attributes. The proposed approach can accurately localize dozens of fiducial points on each 3D face scan, greatly surpassing the feature-based ones, and it also improves the state of the art of the model-based ones in two aspects: sensitivity to initialization and deficiency in discrimination. The proposed method is evaluated on the BU-3DFE, Bosphorus, and BU-4DFE databases, and competitive results are achieved in comparison with counterparts in the literature, clearly demonstrating its effectiveness.
Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
ACM Trans. Multim. Comput. Commun. Appl.2
2018 Learning Face Age Progression: A Pyramid Architecture of GANs
abstract
The two underlying requirements of face age progression, i.e. aging accuracy and identity permanence, are not well studied in the literature. In this paper, we present a novel generative adversarial network based approach. It separately models the constraints for the intrinsic subject-specific characteristics and the age-specific facial changes with respect to the elapsed time, ensuring that the generated faces present desired aging effects while simultaneously keeping personalized properties stable. Further, to generate more lifelike facial details, high-level age-specific features conveyed by the synthesized face are estimated by a pyramidal adversarial discriminator at multiple scales, which simulates the aging effects in a finer manner. The proposed method is applicable to diverse face samples in the presence of variations in pose, expression, makeup, etc., and remarkably vivid aging effects are achieved. Both visual fidelity and quantitative evaluations show that the approach advances the state-of-the-art.
Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001, Anil K. Jain 0001
CVPR2
2018 Receptive Field Block Net for Accurate and Fast Object Detection
Di Huang 0001, Yunhong Wang 0001
ECCV (11)2
2018 Automatic 4D Facial Expression Recognition Using Dynamic Geometrical Image Network
abstract
In this paper, we propose a novel Dynamic Geometrical Image Network (DGIN) for automatic 4D Facial Expression Recognition (FER). Given a 3D video represented as a sequence of face scans, we first estimate their differential geometry quantities and generate geometrical images, including Depth Images (DPI), three Normal Component Images (NCI) and Shape Index Images (SII). These geometrical images are then fed into DGIN for end-to-end training and prediction. DGIN consists of a short-term temporal pooling layer for dynamic geometric image generation, several repetitions of convolution+ReLU+pooling layers for facial spatial feature extraction, and a long-term temporal pooling layer for dynamic feature map fusion, followed by fully connected layers and a joint loss layer. During the training phase, the two-stage longterm and short-term sliding window scheme is introduced for data augmentation and temporal pooling. Meanwhile, a joint loss integrating both the cross-entropy loss and the triplet loss is used to achieve more discriminative expression features. In the testing phase, only the short-term sliding window scheme is applied to the whole video sequence of certain geometric images, whose outputs further go through the deep net for expression similarity measurement. The final result is achieved by fusing the predicted expression scores of different types of geometrical images. Experimental results reported on the BU- 4DFE database demonstrate the effectiveness of the proposed approach.
Weijian Li 0001, Di Huang 0001, Huibin Li 0001, Yunhong Wang 0001
FG2
2018 Hierarchical Attention and Context Modeling for Group Activity Recognition
abstract
Group activity recognition in videos is a challenging task, with two major issues, i.e. attending to those persons and their body parts that contribute significantly to the activity, and modeling contextual person structures in the group. Most previous approaches fail to provide a practical solution to jointly address both issues, however. In this paper, we propose to simultaneously deal with both issues via a hierarchical attention and context modeling framework based on Long Short-Term Memory (LSTM) networks. For the former, we propose `Hierarchical Attention Networks' applied at the part/person level, capable of attending distinctively to different persons and their body parts. For the latter, we build `Hierarchical Context Networks' that take the attentively pooled person-level features as input and recurrently model intra/inter-group contextual structures. The attentive and contextual representations are concatenated and fed into another LSTM to generate high-level discriminative temporal representations for group activity recognition. Extensive experiments on two widely-used group activity datasets demonstrate the effectiveness and superiority of the proposed framework.
Longteng Kong, Jie Qin 0004, Di Huang 0001, Yunhong Wang 0001, Luc Van Gool
ICASSP3
2018 Automatic Facial Attractiveness Prediction by Deep Multi-Task Learning
abstract
Facial Attractiveness Prediction (FAP) is a useful yet challenging problem in the domain of computer vision. In this paper, we propose a deep learning based approach. Different from the existing deep methods, the proposed one models both the texture and shape clues within a multi-task learning framework consisting of attractiveness score prediction and fiducial landmark localization, thus highlighting both of their roles in assessing attractiveness of faces. Considering that the training data are not extensive, a lightweight CNN is designed to jointly learn the facial representation, landmark location, and facial attractiveness score. The proposed method is evaluated on the SCUT-FBP database, and a prediction correlation 0.92, is delivered, which shows the effectiveness of our method. Furthermore, two additional experiments in terms of comparison between facial images before and after make-up or beautification are conducted. The results also prove the advantage of the proposed method.
Lian Gao, Weixin Li 0001, Zehua Huang, Di Huang 0001, Yunhong Wang 0001
ICPR4
2018 Facial Expression Synthesis by U-Net Conditional Generative Adversarial Networks
abstract
High-level manipulation of facial expressions in images such as expression synthesis is challenging because facial expression changes are highly non-linear, and vary depending on the facial appearance. Identity of the person should also be well preserved in the synthesized face. In this paper, we propose a novel U-Net Conditioned Generative Adversarial Network (UC-GAN) for facial expression generation. U-Net helps retain the property of the input face, including the identity information and facial details. We also propose an identity preserving loss, which further improves the performance of our model. Both qualitative and quantitative experiments are conducted on the Oulu-CASIA and KDEF datasets, and the results show that our method can generate faces with natural and realistic expressions while preserve the identity information. Comparison with the state-of-the-art approaches also demonstrates the competency of our method.
Weixin Li 0001, Guodong Mu, Di Huang 0001, Yunhong Wang 0001
ICMR4
2018 Fast and Light Manifold CNN based 3D Facial Expression Recognition across Pose Variations
abstract
This paper proposes a novel approach to 3D Facial Expression Recognition (FER), and it is based on a Fast and Light Manifold CNN model, namely FLM-CNN. Different from current manifold CNNs, FLM-CNN adopts a human vision inspired pooling structure and a multi-scale encoding strategy to enhance geometry representation, which highlights shape characteristics of expressions and runs efficiently. Furthermore, a sampling tree based preprocessing method is presented, and it sharply saves memory when applied to 3D facial surfaces, without much information loss of original data. More importantly, due to the property of manifold CNN features of being rotation-invariant, the proposed method shows a high robustness to pose variations. Extensive experiments are conducted on BU-3DFE, and state-of-the-art results are achieved, indicating its effectiveness.
Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
ACM Multimedia2
2018 REVT: Robust and Efficient Visual Tracking by Region-Convolutional Regression Network
Di Huang 0001, Yunhong Wang 0001
MMM (1)2
2018 Texture and Geometry Scattering Representation-Based Facial Expression Recognition in 2D+3D Videos
abstract
Facial Expression Recognition (FER) is one of the most important topics in the domain of computer vision and pattern recognition, and it has attracted increasing attention for its scientific challenges and application potentials. In this article, we propose a novel and effective approach to FER using multi-model two-dimensional (2D) and 3D videos, which encodes both static and dynamic clues by scattering convolution network. First, a shape-based detection method is introduced to locate the start and the end of an expression in videos; segment its onset, apex, and offset states; and sample the important frames for emotion analysis. Second, the frames in Apex of 2D videos are represented by scattering, conveying static texture details. Those of 3D videos are processed in a similar way, but to highlight static shape details, several geometric maps in terms of multiple order differential quantities, i.e., Normal Maps and Shape Index Maps, are generated as the input of scattering, instead of original smooth facial surfaces. Third, the average of neighboring samples centred at each key texture frame or shape map in Onset is computed, and the scattering features extracted from all the average samples of 2D and 3D videos are then concatenated to capture dynamic texture and shape cues, respectively. Finally, Multiple Kernel Learning is adopted to combine the features in the 2D and 3D modalities and compute similarities to predict the expression label. Thanks to the scattering descriptor, the proposed approach not only encodes distinct local texture and shape variations of different expressions as by several milestone operators, such as SIFT, HOG, and so on, but also captures subtle information hidden in high frequencies in both channels, which is quite crucial to better distinguish expressions that are easily confused. The validation is conducted on the BU-4DFE and BP-4D databa ses, and the accuracies reached are very competitive, indicating its competency for this issue.
Yongqiang Yao, Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
ACM Trans. Multim. Comput. Commun. Appl.2
2017 Feature map pooling for cross-view gait recognition based on silhouette sequence images
abstract
In this paper, we develop a novel convolutional neural network based approach to extract and aggregate useful information from gait silhouette sequence images instead of simply representing the gait process by averaging silhouette images. The network takes a pair of arbitrary length sequence images as inputs and extracts features for each silhouette independently. Then a feature map pooling strategy is adopted to aggregate sequence features. Subsequently, a network which is similar to Siamese network is designed to perform recognition. The proposed network is simple and easy to implement and can be trained in an end-to-end manner Cross-view gait recognition experiments are conducted on OU-ISIR large population dataset. The results demonstrate that our network can extract and aggregate features from silhouette sequence effectively. It also achieves significant equal error rates and comparable identification rates when compared with the state of the art.
Yunhong Wang 0001, Zheng Liu 0014, Qingjie Liu 0001, Di Huang 0001
IJCB5
2017 Facial aging simulation via tensor completion and metric learning
abstract
Facial aging simulation is one of the most challenging issues in automatic machine based face analysis, where the most essential requirements are (i) human identity should remain stable in texture synthesis and (ii) the texture synthesised is expected to accord with human cognitive perception in aging. In this study, the authors propose a tensor completion based method to transform the simulation task to a standard matrix completion one. To protect human dependent characteristics during texture synthesis, the proposed method processes the two major components, i.e. identity and age, in different channels. Furthermore, they incorporate prior information in such a process, assuming that the textures of different subjects in the same age group are similar and similar looking people tend to age in similar ways, and the metric learning technique is adopted to measure the similarity between identities so that the faces that have the highest similarities with the one in the test image are assigned bigger weights in texture generation. In addition, shape deformation is also considered to make the synthesised images more natural. Experimental results achieved on the FG‐NET database demonstrate the effectiveness of the proposed method.
Di Huang 0001, Yunhong Wang 0001, Hongyu Yang 0001
IET Comput. Vis.2
2017 Local feature approach to dorsal hand vein recognition by Centroid-based Circular Key-point Grid and fine-grained matching
Di Huang 0001, Renke Zhang, Yunhong Wang 0001
Image Vis. Comput.1
2016 Cost-Sensitive Two-Stage Depression Prediction Using Dynamic Visual Clues
Xingchen Ma, Di Huang 0001, Yunhong Wang 0001
ACCV (2)2
2016 Scaling and occlusion robust athlete tracking in sports videos
abstract
This paper proposes a novel approach to athlete tracking in sports videos. It follows the framework of Compressive Tracking (CT), but extends it by two manners, i.e. scale refinement as well as occlusion recovery. For the former, an objectness method, namely Edge Box (EB) is adopted to generate proposals, replacing the fixed sampling box in CT, which better fits the scales of the candidate objects. For the latter, a candidate obstruction based solution is presented, which makes use of additional trackers to detect possible obstructions especially the ones possessing highly similar appearances as the target one, and relocate the target as occlusion ends. Therefore, the proposed method inherits the advantage of CT in robust object modelling and fast processing speed, and embodies the tolerance to occlusion and scaling. We evaluate the proposed method on a collection of videos of beach volleyball games, and the experimental results and the comparison with recent advanced trackers highlight its effectiveness.
Jianghu Lu, Di Huang 0001, Yunhong Wang 0001, Longteng Kong
ICASSP2
2016 Hand dorsal vein recognition by matching Width Skeleton Models
abstract
This paper proposes a novel and efficient shape-based approach for hand dorsal vein recognition. A coarse-to-fine segmentation method is first introduced to precisely detect the boundaries of the vein areas. A generalized graph model, namely Width Skeleton Model (WSM), is built then, which takes both the topology of the vein network and the width of the vessel into account, thereby achieving more comprehensive geometric representation and conveying more discriminative cues for identification. The models of different samples are further efficiently compared through a new matching scheme for similarity measurement, based on which the identity of the individual is finally decided. We evaluate the proposed approach on the NCUT database, and the rank-one recognition rate reaches 99.31%, which is superior to the state of the arts, clearly illustrating its competency.
Di Huang 0001, Renke Zhang, Yunhong Wang 0001, Xianbo Xie
ICIP2
2016 Magnifying subtle facial motions for 4D Expression Recognition
abstract
In this paper, we propose an effective approach for automatic 4D Facial Expression Recognition (FER). The flow of 3D facial scans is first modeled to capture spatial deformations based on the recently-developed Riemannian approach, namely Dense Scalar Fields (DSF), where registration and comparison of neighboring 3D face frames are jointly led. The deformations are then fed into a temporal filtering based magnification step to amplify the slight facial actions over time. The proposed method allows revealing subtle (hidden) deformations which enhances the performance in classification. We evaluate our approach on the BU-4DFE dataset, and the state-of-art accuracy up to 94.18% is achieved, which is superior to the top one so far reported, clearly demonstrating its effectiveness.
Qingkai Zhen, Di Huang 0001, Yunhong Wang 0001, Hassen Drira, Boulbaba Ben Amor, Mohamed Daoudi
ICPR2
2016 Dorsal hand vein recognition via hierarchical combination of texture and shape clues
Di Huang 0001, Xiangrong Zhu 0002, Yunhong Wang 0001, David Zhang 0001
Neurocomputing1
2016 Face Aging Effect Simulation Using Hidden Factor Analysis Joint Sparse Representation
abstract
Face aging simulation has received rising investigations nowadays, whereas it still remains a challenge to generate convincing and natural age-progressed face images. In this paper, we present a novel approach to such an issue using hidden factor analysis joint sparse representation. In contrast to the majority of tasks in the literature that integrally handle the facial texture, the proposed aging approach separately models the person-specific facial properties that tend to be stable in a relatively long period and the age-specific clues that gradually change over time. It then transforms the age component to a target age group via sparse reconstruction, yielding aging effects, which is finally combined with the identity component to achieve the aged face. Experiments are carried out on three face aging databases, and the results achieved clearly demonstrate the effectiveness and robustness of the proposed method in rendering a face with aging effects. In addition, a series of evaluations prove its validity with respect to identity preservation and aging effect generation.
Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001, Yuan Yan Tang
IEEE Trans. Image Process.2
2016 Muscular Movement Model-Based Automatic 3D/4D Facial Expression Recognition
abstract
Facial expression is an important channel for human nonverbal communication. This paper presents a novel and effective approach to automatic 3D/4D facial expression recognition based on the muscular movement model (MMM). In contrast to most of existing methods, the MMM deals with such an issue in the viewpoint of anatomy. It first automatically segments the input 3D face (frame) by localizing the corresponding points within each muscular region of the reference using iterative closest normal point. A set of features with multiple differential quantities, including coordinate, normal, values, are then extracted to describe the geometry deformation of each segmented region. Meanwhile, we analyze the importance of these muscular areas, and a score level fusion strategy is exploited to optimize their weights by the genetic algorithm in the learning step. The support vector machine and the hidden Markov model are finally used to predict the expression label in 3D and 4D, respectively. The experiments are conducted on the BU-3DFE and BU-4DFE databases, and the results achieved clearly demonstrate the effectiveness of the proposed method.
Qingkai Zhen, Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
IEEE Trans. Multim.2
2015 Muscular Movement Model Based Automatic 3D Facial Expression Recognition
Qingkai Zhen, Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
MMM (1)2
2015 An efficient multimodal 2D + 3D feature-based approach to automatic facial expression recognition
Huibin Li 0001, Huaxiong Ding, Di Huang 0001, Yunhong Wang 0001, Xi Zhao 0001, Jean-Marie Morvan, Liming Chen 0002
Comput. Vis. Image Underst.3
2015 Towards 3D Face Recognition in the Real: A Registration-Free Approach Using Fine-Grained Matching of 3D Keypoint Descriptors
Huibin Li 0001, Di Huang 0001, Jean-Marie Morvan, Yunhong Wang 0001, Liming Chen 0002
Int. J. Comput. Vis.2
2015 Dither modulation of significant amplitude difference for wavelet based robust watermarking
Chunlei Li 0004, Zhaoxiang Zhang 0001, Yunhong Wang 0001, Bin Ma 0004, Di Huang 0001
Neurocomputing5
2015 Semi-fragile self-recoverable watermarking algorithm based on wavelet group quantization and double authentication
Chunlei Li 0002, Zhoufeng Liu, Di Huang 0001
Multim. Tools Appl.5
2015 Hand-Dorsa Vein Recognition by Matching Local Features of Multisource Keypoints
abstract
As an emerging biometric for people identification, the dorsal hand vein has received increasing attention in recent years due to the properties of being universal, unique, permanent, and contactless, and especially its simplicity of liveness detection and difficulty of forging. However, the dorsal hand vein is usually captured by near-infrared (NIR) sensors and the resulting image is of low contrast and shows a very sparse subcutaneous vascular network. Therefore, it does not offer sufficient distinctiveness in recognition particularly in the presence of large population. This paper proposes a novel approach to hand-dorsa vein recognition through matching local features of multiple sources. In contrast to current studies only concentrating on the hand vein network, we also make use of person dependent optical characteristics of the skin and subcutaneous tissue revealed by NIR hand-dorsa images and encode geometrical attributes of their landscapes, e.g., ridges, valleys, etc., through different quantities, such as cornerness and blobness, closely related to differential geometry. Specifically, the proposed method adopts an effective keypoint detection strategy to localize features on dorsal hand images, where the speciality of absorption and scattering of the entire dorsal hand is modeled as a combination of multiple (first-, second-, and third-) order gradients. These features comprehensively describe the discriminative clues of each dorsal hand. This method further robustly associates the corresponding keypoints between gallery and probe samples, and finally predicts the identity. Evaluated by extensive experiments, the proposed method achieves the best performance so far known on the North China University of Technology (NCUT) Part A dataset, showing its effectiveness. Additional results on NCUT Part B illustrate its generalization ability and robustness to low quality data.
Di Huang 0001, Yinhang Tang, Liming Chen 0002, Yunhong Wang 0001
IEEE Trans. Cybern.1
2015 DERF: Distinctive Efficient Robust Features From the Biological Modeling of the P Ganglion Cells
abstract
Studies in neuroscience and biological vision have shown that the human retina has strong computational power, and its information representation supports vision tasks on both ventral and dorsal pathways. In this paper, a new local image descriptor, termed distinctive efficient robust features (DERF), is derived by modeling the response and distribution properties of the parvocellular-projecting ganglion cells in the primate retina. DERF features exponential scale distribution, exponential grid structure, and circularly symmetric function difference of Gaussian (DoG) used as a convolution kernel, all of which are consistent with the characteristics of the ganglion cell array found in neurophysiology, anatomy, and biophysics. In addition, a new explanation for local descriptor design is presented from the perspective of wavelet tight frames. DoG is naturally a wavelet, and the structure of the grid points array in our descriptor is closely related to the spatial sampling of wavelets. The DoG wavelet itself forms a frame, and when we modulate the parameters of our descriptor to make the frame tighter, the performance of the DERF descriptor improves accordingly. This is verified by designing a tight frame DoG, which leads to much better performance. Extensive experiments conducted in the image matching task on the multiview stereo correspondence data set demonstrate that DERF outperforms state of the art methods for both hand-crafted and learned descriptors, while remaining robust and being much faster to compute.
Dawei Weng, Yunhong Wang 0001, Mingming Gong, Dacheng Tao, Hui Wei 0001, Di Huang 0001
IEEE Trans. Image Process.6
2014 A coarse-to-fine approach to robust 3D facial landmarking via curvature analysis and Active Normal Model
abstract
Facial landmarking is a fundamental step in machine-based face analysis. The majority of existing techniques handle such an issue based on 2D images; however, they suffer from illumination and pose variations that largely degrade landmarking performance. The emergence of 3D data provides us with an alternative to overcome these unsolved problems in the 2D domain. This paper proposes a novel approach to 3D facial landmarking, combining both the advantages of feature based methods as well as model based ones in a coarse-to-fine manner. For the coarse stage, three fiducial landmarks (the nose tip and two inner eye corners) are robustly detected through curvature analysis, and these points are further employed to initialize the subsequent model fitting. For the fine stage, a statistical model is constructed based on the normal information including the x, y, and z components of the facial point-cloud rather than the smooth coordinate information, thereby namely Active Normal Model (ANM), to highlight its shape characteristics for final landmark prediction. The proposed approach accurately localizes 83 fiducial points on each 3D face model, greatly surpassing those of feature based ones, while improving the state of the art model based ones in two aspects, i.e. sensitivity to initialization and deficiency in discrimination. Evaluated on the BU-3DFE database, very competitive results are achieved in comparison with those in the literature, clearly demonstrating its effectiveness.
Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
IJCB2
2014 Age invariant face recognition based on texture embedded discriminative graph model
abstract
In an automatic face recognition system, it still remains a challenge to improve the robustness to aging. In this paper, we present a novel approach to address age invariant face recognition, by formulating it as a graph matching problem. In contrast to the majority of tasks in the literature that only make use of robust texture features, this method generates a graph from a set of fiducial landmarks of each face, which captures the texture clues that tend to be stable in a period as well as the common facial geometry configuration. The nodes of the graph denote the texture of a face area around a landmark, and the edges correspond to the geometry topology of the face. For each area, the age invariant texture information is extracted by a discriminative and compact feature encoded in the Local Gabor Binary Pattern Histogram Sequence (LGBPHS) projected in an LDA subspace. An objective function is then designed to match graphs for the purpose of registration and identification. Experiments are carried out on the FG-NET Aging database, and the results achieved outperform the state of the art ones, which clearly demonstrate the effectiveness and robustness of the proposed method in face recognition across age variations.
Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001
IJCB2
2014 Action recognition based on kinematic representation of video data
abstract
The local space-time feature is an effective way to represent video data and achieves state-of-the-art performance in action recognition. However, in majority of cases, it only captures the static or dynamic cues of the image sequence. In this paper, we propose a novel kinematic descriptor, namely Static and Dynamic fEature Velocity (SDEV), which models the changes of both static and dynamic information with time for action recognition. It is not only discriminative itself, but also complementary to the existing descriptors, thus leading to more comprehensive representation of actions by their combination. Evaluated on two public databases, i.e. UCF sports and Olympic Sports, the results clearly illustrate the competency of SDEV.
Di Huang 0001, Yunhong Wang 0001, Jie Qin 0004
ICIP2
2014 3D assisted face recognition via progressive pose estimation
abstract
Most existing pose-independent Face Recognition (FR) techniques take advantage of 3D model to guarantee the naturalness while normalizing or simulating pose variations. Two nontrivial problems to be tackled are accurate measurement of pose parameters and computational efficiency. In this paper, we introduce an effective and efficient approach to estimate human head pose, which fundamentally ameliorates the performance of 3D aided FR systems. The proposed method works in a progressive way: firstly, a random forest (RF) is constructed utilizing synthesized images derived from 3D models; secondly, the classification result obtained by applying well-trained RF on a probe image is considered as the preliminary pose estimation; finally, this initial pose is transferred to shape-based 3D morphable model (3DMM) aiming at definitive pose normalization. Using such a method, similarity scores between frontal view gallery set and pose-normalized probe set can be computed to predict the identity. Experimental results achieved on the UHDB dataset outperform the ones so far reported. Additionally, it is much less time-consuming than prevailing 3DMM based approaches.
Wuming Zhang, Di Huang 0001, Dimitris Samaras, Jean-Marie Morvan, Yunhong Wang 0001, Liming Chen 0002
ICIP2
2014 Video face recognition via combination of real-time local features and temporal-spatial cues
abstract
Video‐based face recognition has attracted much attention and made great progress in the past decade. However, it still encounters two main problems, which are efficiently representing faces in frames and sufficiently exploiting temporal–spatial constraints between frames. The authors investigate the existing real‐time features for face description, and compare their performance. Moreover, a novel approach is proposed to model temporal–spatial information which is then combined with real‐time features to further enforce the consistent constraints between frames to improve the recognition performance. The experiments are validated on three video face databases and the results demonstrate that temporal–spatial cues combined with the most powerful real‐time features largely improve the recognition rate.
Gaopeng Gou, Di Huang 0001, Yunhong Wang 0001
IET Comput. Vis.2
2014 Expression-robust 3D face recognition via weighted sparse representation of multi-scale and multi-component local normal patterns
Huibin Li 0001, Di Huang 0001, Jean-Marie Morvan, Liming Chen 0002, Yunhong Wang 0001
Neurocomputing2
2014 Local circular patterns for multi-modal facial gender and ethnicity classification
Di Huang 0001, Huaxiong Ding, Yunhong Wang 0001, Guangpeng Zhang, Liming Chen 0002
Image Vis. Comput.1
2014 Secure multimodal biometric authentication with wavelet quantization based fingerprint watermarking
Bin Ma 0004, Yunhong Wang 0001, Chunlei Li 0004, Zhaoxiang Zhang 0001, Di Huang 0001
Multim. Tools Appl.5
2014 HSOG: A Novel Local Image Descriptor Based on Histograms of the Second-Order Gradients
abstract
Recent investigations on human vision discover that the retinal image is a landscape or a geometric surface, consisting of features such as ridges and summits. However, most of existing popular local image descriptors in the literature, e.g., scale invariant feature transform (SIFT), histogram of oriented gradient (HOG), DAISY, local binary Patterns (LBP), and gradient location and orientation histogram, only employ the first-order gradient information related to the slope and the elasticity, i.e., length, area, and so on of a surface, and thereby partially characterize the geometric properties of a landscape. In this paper, we introduce a novel and powerful local image descriptor that extracts the histograms of second-order gradients (HSOGs) to capture the curvature related geometric properties of the neural landscape, i.e., cliffs, ridges, summits, valleys, basins, and so on. We conduct comprehensive experiments on three different applications, including the problem of local image matching, visual object categorization, and scene classification. The experimental results clearly evidence the discriminative power of HSOG as compared with its first-order gradient-based counterparts, e.g., SIFT, HOG, DAISY, and center-symmetric LBP, and the complementarity in terms of image representation, demonstrating the effectiveness of the proposed local descriptor.
Di Huang 0001, Chao Zhu 0003, Yunhong Wang 0001, Liming Chen 0002
IEEE Trans. Image Process.1
2013 HSOG: a novel local descriptor based on histograms of second order gradients for object categorization
abstract
This paper presents a novel local image descriptor for object categorization that extracts the Histograms of the Second Order Gradients and is thereby named as HSOG. The HSOG descriptor is in contrast to the widely used ones in the literature, e.g. SIFT, DAISY, HOG, LBP, etc., which are based on the first order gradient information. The contributions of this work can be summarized as: (1) the design of HSOG; (2) the prove of its discriminative power and its complementation to the first order gradient based descriptors; (3) the analysis of performance variation caused by different parameter settings; and (4) the multi-scale extension which further improves the categorization accuracy. The experimental results achieved on the Caltech 101 and Caltech 256 databases clearly highlight the effectiveness of the proposed approach.
Di Huang 0001, Chao Zhu 0003, Charles-Edmond Bichot, Yunhong Wang 0001, Liming Chen 0002
ICMR1
2013 View-Invariant Discriminative Projection for Multi-View Gait-Based Human Identification
abstract
Existing methods for multi-view gait-based identification mainly focus on transforming the features of one view to the features of another view, which is technically sound but has limited practical utility. In this paper, we propose a view-invariant discriminative projection (ViDP) method, to improve the discriminative ability of multi-view gait features by a unitary linear projection. It is implemented by iteratively learning the low dimensional geometry and finding the optimal projection according to the geometry. By virtue of ViDP, the multi-view gait features can be directly matched without knowing or estimating the viewing angles. The ViDP feature projected from gait energy image achieves promising performance in the experiments of multi-view gait-based identification. We suggest that it is possible to construct a gait-based identification system for arbitrary probe views, by incorporating the information of gallery data with sufficient viewing angles. In addition, ViDP performs even better than the state-of-the-art view transformation methods, which are trained for the combination of gallery and probe viewing angles in every evaluation.
Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001, James J. Little, Di Huang 0001
IEEE Trans. Inf. Forensics Secur.5
2013 Learning the Spherical Harmonic Features for 3-D Face Recognition
abstract
In this paper, a competitive method for 3-D face recognition (FR) using spherical harmonic features (SHF) is proposed. With this solution, 3-D face models are characterized by the energies contained in spherical harmonics with different frequencies, thereby enabling the capture of both gross shape and fine surface details of a 3-D facial surface. This is in clear contrast to most 3-D FR techniques which are either holistic or feature based, using local features extracted from distinctive points. First, 3-D face models are represented in a canonical representation, namely, spherical depth map, by which SHF can be calculated. Then, considering the predictive contribution of each SHF feature, especially in the presence of facial expression and occlusion, feature selection methods are used to improve the predictive performance and provide faster and more cost-effective predictors. Experiments have been carried out on three public 3-D face datasets, SHREC2007, FRGC v2.0, and Bosphorus, with increasing difficulties in terms of facial expression, pose, and occlusion, and which demonstrate the effectiveness of the proposed method.
Peijiang Liu, Yunhong Wang 0001, Di Huang 0001, Zhaoxiang Zhang 0001, Liming Chen 0002
IEEE Trans. Image Process.3
2012 Hand Vein Recognition Based on Oriented Gradient Maps and Local Feature Matching
Di Huang 0001, Yinhang Tang, Liming Chen 0002, Yunhong Wang 0001
ACCV (4)1
2012 Recognizing Occluded 3D Faces Using an Efficient ICP Variant
abstract
This paper proposes an efficient variant of the Iterative Closest Point (ICP) algorithm for 3D face recognition in the presence of occlusion. The new ICP variant improves the original one in two aspects: the computational efficiency and the robustness to occlusion changes. For the former one, a facial surface is firstly described as a Spherical Depth Map (SDM), based on which uniform down-sampling can be conveniently applied to remove redundant vertices, aiming to decrease the consumed time of ICP. For the latter one, since occlusions can be considered as face outliers, a rejection strategy is embedded into ICP to eliminate their impacts. The proposed method is validated in face verification and identification scenarios on the Bosphorus database, and the experimental results clearly demonstrate its effectiveness and efficiency.
Peijiang Liu, Yunhong Wang 0001, Di Huang 0001, Zhaoxiang Zhang 0001
ICME3
2012 3D facial expression recognition via multiple kernel learning of Multi-Scale Local Normal Patterns
Huibin Li 0001, Liming Chen 0002, Di Huang 0001, Yunhong Wang 0001, Jean-Marie Morvan
ICPR3
2012 Enhancing biometric security with wavelet quantization watermarking based two-stage multimodal authentication
Bin Ma 0004, Chunlei Li 0004, Yunhong Wang 0001, Zhaoxiang Zhang 0001, Di Huang 0001
ICPR5
2012 Hand-dorsa vein recognition based on multi-level keypoint detection and local feature matching
Yinhang Tang, Di Huang 0001, Yunhong Wang 0001
ICPR2
2012 Facial image-based gender classification using Local Circular Patterns
Di Huang 0001, Yunhong Wang 0001, Guangpeng Zhang
ICPR2
2012 A Novel Video Face Clustering Algorithm Based on Divide and Conquer Strategy
Gaopeng Gou, Di Huang 0001, Yunhong Wang 0001
PRICAI2
2012 A Hybrid Local Feature for Face Recognition
Gaopeng Gou, Di Huang 0001, Yunhong Wang 0001
PRICAI2
2012 3-D Face Recognition Using eLBP-Based Facial Description and Local Feature Hybrid Matching
abstract
This paper presents an effective method for 3-D face recognition using a novel geometric facial representation along with a local feature hybrid matching scheme. The proposed facial surface description is based on a set of facial depth maps extracted by multiscale extended Local Binary Patterns (eLBP) and enables an efficient and accurate description of local shape changes; it thus enhances the distinctiveness of smooth and similar facial range images generated by preprocessing steps. The following matching strategy is SIFT-based and performs in a hybrid way that combines local and holistic analysis, robustly associating the keypoints between two facial representations of the same subject. As a result, the proposed approach proves robust to facial expression variations, partial occlusions, and moderate pose changes, and the last property makes our system registration-free for nearly frontal face models. The proposed method was experimented on three public datasets, i.e. FRGC v2.0, Gavab, and Bosphorus. It displays a rank-one recognition rate of 97.6% and a verification rate of 98.4% at a 0.001 FAR on the FRGC v2.0 database without any face alignment. Additional experiments on the Bosphorus dataset further highlight the advantages of the proposed method with regard to expression changes and external partial occlusions. The last experiment carried out on the Gavab database demonstrates that the entire system can also deal with faces under large pose variations and even partially occluded ones, when only aided by a coarse alignment process.
Di Huang 0001, Mohsen Ardabilian, Yunhong Wang 0001, Liming Chen 0002
IEEE Trans. Inf. Forensics Secur.1
2011 A novel geometric facial representation based on multi-scale extended local binary patterns
abstract
In this study, we present a novel geometric representation for 3D faces in order to enhance distinctiveness of generally smooth range images. This novel face representation is based on Multi-Scale Extended Local Binary Patterns (ELBP) and enables accurate and fast description of local shape variations on range faces. When associated with the proposed SIFT-based local feature matching scheme, this novel geometric facial representation shows its discriminative power in 3D face recognition, displaying a rank-one recognition rate up to 97.2% and a verification rate of 98.4% at a 0.001 FAR respectively on the FRGC v2.0 database. Moreover, costly registration is not needed thanks to the relative tolerance of the proposed representation and the SIFT methodology to moderate pose changes as the ones existing in FRGC v2.0. Finally, additional experiments demonstrate that the entire system is also robust to facial expression variations.
Di Huang 0001, Mohsen Ardabilian, Yunhong Wang 0001, Liming Chen 0002
FG1
2011 Learning weighted sparse representation of encoded facial normal information for expression-robust 3D face recognition
abstract
This paper proposes a novel approach for 3D face recognition by learning weighted sparse representation of encoded facial normal information. To comprehensively describe 3D facial surface, three components, in X, Y, and Z-plane respectively, of normal vector are encoded locally to their corresponding normal pattern histograms. They are finally fed to a sparse representation classifier enhanced by learning based spatial weights. Experimental results achieved on the FRGC v2.0 database prove that the proposed encoded normal information is much more discriminative than original normal information. Moreover, the patch based weights learned using the FRGC v1.0 and Bosphorus datasets also demonstrate the importance of each facial physical component for 3D face recognition.
Huibin Li 0001, Di Huang 0001, Jean-Marie Morvan, Liming Chen 0002
IJCB2
2011 Twins 3D face recognition challenge
abstract
Existing 3D face recognition algorithms have achieved high enough performances against public datasets like FRGC v2, that it is difficult to achieve further significant increases in recognition performance. However, the 3D TEC dataset is a more challenging dataset which consists of 3D scans of 107 pairs of twins that were acquired in a single session, with each subject having a scan of a neutral expression and a smiling expression. The combination of factors related to the facial similarity of identical twins and the variation in facial expression makes this a challenging dataset. We conduct experiments using state of the art face recognition algorithms and present the results. Our results indicate that 3D face recognition of identical twins in the presence of varying facial expressions is far from a solved problem, but that good performance is possible.
Vipin Vijayan, Kevin W. Bowyer, Patrick J. Flynn, Di Huang 0001, Liming Chen 0002, Mark Hansen, Omar Ocegueda, Shishir Shah 0001, Ioannis A. Kakadiaris
IJCB4
2011 Expression robust 3D face recognition via mesh-based histograms of multiple order surface differential quantities
abstract
This paper presents a mesh-based approach for 3D face recognition using a novel local shape descriptor and a SIFT-like matching process. Both maximum and minimum curvatures estimated in the 3D Gaussian scale space are employed to detect salient points. To comprehensively characterize 3D facial surfaces and their variations, we calculate weighted statistical distributions of multiple order surface differential quantities, including histogram of mesh gradient (HoG), histogram of shape index (HoS) and histogram of gradient of shape index (HoGS) within a local neighborhood of each salient point. The subsequent matching step then robustly associates corresponding points of two facial surfaces, leading to much more matched points between different scans of a same person than the ones of different persons. Experimental results on the Bosphorus dataset highlight the effectiveness of the proposed method and its robustness to facial expression variations.
Huibin Li 0001, Di Huang 0001, Pierre Lemaire 0002, Jean-Marie Morvan, Liming Chen 0002
ICIP2
2011 A mixture of gated experts optimized using simulated annealing for 3D face recognition
abstract
A commonly accepted fact in the biometrics related domain is that fusing multiple classifiers for decision making generally leads to improved recognition performance. Meanwhile, the search for an optimal fusion strategy remains extraordinarily complex since the cardinality of the space of possible fusion schemes is exponentially proportional to the number of competing classifiers. In this paper, we propose a mixture of gated experts for 3D face recognition using an ensemble of 24 different classifiers. The mixture of gated experts is optimized using a Simulated Annealing-based algorithm. It automatically selects and fuses the most relevant similarity measurements. The experimental results of 3D face recognition achieved on the FRGC v2.0 dataset illustrate the effectiveness and stability of the proposed method. Additionally, as a learning-based method, it also has a good robustness to the variations of training database.
Wael Ben Soltana, Di Huang 0001, Mohsen Ardabilian, Liming Chen 0002, Chokri Ben Amar
ICIP2
2011 3D Face Recognition Based on Local Shape Patterns and Sparse Representation Classifier
Di Huang 0001, Karima Ouji, Mohsen Ardabilian, Yunhong Wang 0001, Liming Chen 0002
MMM (1)1
2011 Local Binary Patterns and Its Application to Facial Image Analysis: A Survey
abstract
Local binary pattern (LBP) is a nonparametric descriptor, which efficiently summarizes the local structures of images. In recent years, it has aroused increasing interest in many areas of image processing and computer vision and has shown its effectiveness in a number of applications, in particular for facial image analysis, including tasks as diverse as face detection, face recognition, facial expression analysis, and demographic classification. This paper presents a comprehensive survey of LBP methodology, including several more recent variations. As a typical application of the LBP approach, LBP-based facial image analysis is extensively reviewed, while its successful extensions, which deal with various tasks of facial image analysis, are also highlighted.
Di Huang 0001, Caifeng Shan, Mohsen Ardabilian, Yunhong Wang 0001, Liming Chen 0002
IEEE Trans. Syst. Man Cybern. Part C1
2010 Automatic Asymmetric 3D-2D Face Recognition
abstract
3D Face recognition has been considered as a major solution to deal with unsolved issues of reliable 2D face recognition in recent years, i.e. lighting and pose variations. However, 3D techniques are currently limited by their high registration and computation cost. In this paper, an asymmetric 3D-2D face recognition method is presented, enrolling in textured 3D whilst performing automatic identification using only 2D facial images. The goal is to limit the use of 3D data to where it really helps to improve face recognition accuracy. The proposed approach contains two separate matching steps: Sparse Representation Classifier (SRC) is applied to 2D-2D matching, while Canonical Correlation Analysis (CCA) is exploited to learn the mapping between range LBP faces (3D) and texture LBP faces (2D). Both matching scores are combined for the final decision. Moreover, we propose a new preprocessing pipeline to enhance robustness to lighting and pose effects. The proposed method achieves better experimental results in the FRGC v2.0 dataset than 2D methods do, but avoiding the cost and inconvenience of data acquisition and computation of 3D approaches.
Di Huang 0001, Mohsen Ardabilian, Yunhong Wang 0001, Liming Chen 0002
ICPR1
2010 Automatic 3D Facial Expression Recognition Based on a Bayesian Belief Net and a Statistical Facial Feature Model
abstract
Automatic facial expression recognition on 3D face data is still a challenging problem. In this paper we propose a novel approach to perform expression recognition automatically and flexibly by combining a Bayesian Belief Net (BBN) and Statistical facial feature models (SFAM). A novel BBN is designed for the specific problem with our proposed parameter computing method. By learning global variations in face landmark configuration (morphology) and local ones in terms of texture and shape around landmarks, morphable Statistic Facial feature Model (SFAM) allows not only to perform an automatic landmarking but also to compute the belief to feed the BBN. Tested on the public 3D face expression database BU-3DFE, our automatic approach allows to recognize expressions successfully, reaching an average recognition rate over 82%.
Xi Zhao 0001, Di Huang 0001, Emmanuel Dellandréa, Liming Chen 0002
ICPR2
2009 Asymmetric 3D/2D face recognition based on LBP facial representation and canonical correlation analysis
abstract
In the recent years, 3D Face recognition has emerged as a major solution to deal with the unsolved issues for reliable 2D face recognition, i.e. lighting condition and viewpoint variations. However, 3D method is currently limited by its registration and computation cost. In this paper, we propose to investigate a solution named asymmetric face recognition scheme, enrolling people in 3D environment but performing identification in 2D. The goal is to limit the use of 3D data to where it really helps to improve recognition performances. In our approach, Local Binary Patterns (LBP) is used as an efficient facial representation for both 2D texture images and 3D range images. A weighted Chi square distance is used as matching score between the 2D LBP facial representations; Canonical Correlation Analysis (CCA) is applied to learn the mapping between LBP-based range face images (3D) and LBP facial texture images (2D). Both matching scores are further fused to obtain the final result. Compared with the traditional 2D/2D algorithms, the proposed asymmetric face recognition scheme achieves better accuracy; while avoiding the high cost of data acquisition and computation in 3D/3D approaches.
Di Huang 0001, Mohsen Ardabilian, Yunhong Wang 0001, Liming Chen 0002
ICIP1