EDBT 2026 Demo / reviewers in the wild / expert
Haoyu Chen 0001
dblp:146/8170-1
· DBLP profile ↗
36ranked-venue papers
9as first author
30since 2021 · last 2026
0000-0003-3267-2664ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 7 first-author · 21 since 2021Artificial intelligence and machine learning · 24 · 8 first-author · 21 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LatentMag: Self-Supervised 3D Magnification for Micro Expressions via Latent ExtrapolationabstractMicro-expressions (MEs) are subtle and brief facial movements that reveal genuine emotional states but are often imperceptible due to their low intensity. While motion magnification has proven effective for enhancing ME visibility in 2D settings, its extension to 3D remains largely unexplored. In this work, we presentLatentMag, the first controllable 3D micro-expression magnification framework. Unlike traditional editing methods that rely on fixed labels or expression targets, our approach models expression intensity as a relative, input-dependent signal. We adopt registered 3D meshes as our representation, enabling vertex-level correspondence and interpretable displacement analysis. To guide magnification, we introduce a geometric prior that models amplification as a spatially adaptive transformation, where the change in pairwise distance between points on the output mesh scales with that observed between the input shapes, ensuring natural, localized deformation. We operationalize this prior in a generative framework by disentangling a latent intensity code, whose extrapolation drives controllable shape amplification. Trained in a self-supervised manner using unlabeled mesh sequences, LatentMag generalizes well to unseen identities and expressions, offering a novel solution that bridges geometric interpretability with realistic 3D expression modeling. Mengting Wei, Xingxun Jiang, Haoyu Chen 0001, Yante Li, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2026 | Hybrid-Supervised Hypergraph-Enhanced Transformer for Micro-Gesture Based Emotion RecognitionabstractMicro-gestures are unconsciously performed body gestures that can convey the emotion states of humans and start to attract more research attention in the fields of human behavior understanding and affective computing as an emerging topic. However, the modeling of human emotion based on micro-gestures has not been explored sufficiently. In this work, we propose to recognize the emotion states based on the micro-gestures by reconstructing behavioral patterns with a hypergraph-enhanced Transformer in a hybrid-supervised framework. In the framework, hypergraph Transformer based encoder and decoder are separately designed by stacking the hypergraph-enhanced self-attention and multiscale temporal convolution modules. Especially, to better capture the subtle motion of micro-gestures, we construct a decoder with additional upsampling operations for a reconstruction task in a self-supervised learning manner. We further propose a hypergraph-enhanced self-attention module where the hyperedges between skeleton joints are gradually updated to present the relationships of body joints for modeling the subtle local motion. Lastly, for exploiting the relationship between the emotion states and local motion of micro-gestures, an emotion recognition head from the output of encoder is designed with a shallow architecture and learned in a supervised way. The end-to-end framework is jointly trained in a one-stage way by comprehensively utilizing self-reconstruction and supervision information. The proposed method is evaluated on two publicly available datasets, namely iMiGUE and SMG, and achieves the best performance under multiple metrics, which is superior to the existing methods. The code is available on Github (https://github.com/xiazhaoqiang/H2OFormerMicroGestureRec). Zhaoqiang Xia, Haoyu Chen 0001, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2026 | CROMBO: Cross-Modality Bootstrapping for Unified Sketch-Photo Representation LearningabstractSketch–photo recognition refers to matching hand-drawn sketches with their corresponding photos, where the performance essentially depends on how well the representations of the two modalities are aligned in the feature spaces. Existing works bluntly force models to reduce the representation discrepancy between the modalities, making the learning less effective. Besides, the current symmetric feature extraction framework prefers the photo modality for richer information while neglecting the sketch modality. Driven by these observations, we argue that, instead of forcefully wiping out the modality discrepancy, we may utilize the discrepancy to enhance model learning. Thus, we propose a Cross-Modality Bootstrapping learning framework (CROMBO) that utilizes the modality discrepancy to bootstrap cross-modality representation learning via a differentiated interaction manner. Specifically, we first present a Sketch Implicit Bootstrapping (SIB) module to magnify the recognizable elements in the photo modality by utilizing the characteristic of sketches having only contours and key details. Second, a Photo-driven Sketch Refinement (PSR) module is developed to guide the sketch representation in the shared feature extraction process by supplementing rich information from the photo modality. Moreover, we design a second-order alignment strategy to dynamically align the latent distribution of two modalities in a Hilbert space. Also, our CROMBO can learn fewer parameters by freezing the weights of shallow layers in the backbone while making no sacrifice in performance. Extensive experiments on six public datasets verify the superior performance of our CROMBO for sketch–photo-based tasks, such as sketch re-identification (Re-ID), sketch–photo face recognition, and sketch-based image retrieval. Xu Cheng 0003, Hao Yu 0015, Haoyu Chen 0001, Guoying Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | FreeNet: Liberating Depth-Wise Separable Operations for Building Faster Mobile Vision ArchitecturesabstractIn the pursuit of efficient vision architectures, substantial efforts have been devoted to optimizing operator efficiency. Depth-wise separable operators, such as DWConv, are found cheap in both FLOPs and parameters. As a result, they are increasingly incorporated into efficient backbones, trading for deeper and wider architectures to enhance performance. However, separable operators are not really fast on devices due to the discontinuous memory access requirements. In this paper, we propose FreeNets, a family of simple and efficient backbones that free the separable operation to further accelerate the running speed. We introduce sparse sampling mixers (S2-Mixer) to supersede existing separable token mixers. The S2-Mixer samples multiple segments of partially continuous signals across spatial and channel dimensions for convolutional processing, achieving extremely fast on-device speed. The sparse sampling also enables S2-Mixer to capture long-range pixel relationships from dynamic receptive fields. Furthermore, we introduce a Shift Feed-Forward Network (ShiftFFN) as a faster alternative to existing channel mixers. It utilizes a shift neck architecture that aggregates global information to shift features, enabling faster channel mixing while incorporating global pixel information. Extensive experiments demonstrate that FreeNet offers a superior accuracy-efficiency tradeoff compared to the latest efficient models. On ImageNet-1k, FreeNet-S2 outperforms the StarNet-S4 by 0.4% in top-1 accuracy, while running around 40% faster on desktop GPU and 15% faster on Mobile GPU. Hao Yu 0015, Haoyu Chen 0001, Wei Peng 0009, Xu Cheng 0003, Guoying Zhao 0001 |
AAAI | 2 |
| 2025 | From Laboratory to Real World: A New Benchmark Towards Privacy-Preserved Visible-Infrared Person Re-IdentificationabstractAiming to match pedestrian images captured under varying lighting conditions, visible-infrared person re-identification (VI-ReID) has drawn intensive research attention and achieved promising results. However, in real-world surveillance contexts, data is distributed across multiple devices/entities, raising privacy and ownership concerns that make existing centralized training impractical for VI-ReID. To tackle these challenges, we propose L2RW, a benchmark that brings VI-ReID closer to real-world applications. The rationale of L2RW is that integrating decentralized training into VI-ReID can address privacy concerns in scenarios with limited data-sharing regulation. Specifically, we design protocols and corresponding algorithms for different privacy sensitivity levels. In our new benchmark, we ensure the model training is done in the conditions that: 1) data from each camera remains completely isolated, or 2) different data entities (e.g., data controllers of a certain region) can selectively share the data. In this way, we simulate scenarios with strict privacy constraints which is closer to real-world conditions. Intensive experiments with various server-side federated algorithms are conducted, showing the feasibility of decentralized VI-ReID training. Notably, when evaluated in unseen domains (i.e., new data entities), our L2RW, trained with isolated data (privacy-preserved), achieves performance comparable to SOTAs trained with shared data (privacy-unrestricted). We hope this work offers a novel research entry for deploying VI-ReID that fits real-world scenarios and can benefit the community. Hao Yu 0015, Xu Cheng 0003, Haoyu Chen 0001, Zhaodong Sun, Guoying Zhao 0001 |
CVPR | 4 |
| 2025 | Deep Change Monitoring: A Hyperbolic Representative Learning Framework and a Dataset for Long-term Fine-grained Tree Change DetectionabstractIn environmental protection, tree monitoring plays an essential role in maintaining and improving ecosystem health. However, precise monitoring is challenging because existing datasets fail to capture continuous fine-grained changes in trees due to low-resolution images and high acquisition costs. In this paper, we introduce UAVTC, a large-scale, long-term, high-resolution dataset collected using UAVs equipped with cameras, specifically designed to detect individual Tree Changes (TCs). UAVTC includes rich annotations and statistics based on biological knowledge, offering a fine-grained view for tree monitoring. To address environmental influences and effectively model the hierarchical diversity of physiological TCs, we propose a novel Hyperbolic Siamese Network (HSN) for TC detection, enabling compact and hierarchical representations of dynamic tree changes. Extensive experiments show that HSN can effectively capture complex hierarchical changes and provide a robust solution for fine-grained TC detection. In addition, HSN generalizes well to cross-domain face anti-spoofing task, highlighting its broader significance in AI. We believe our work, combining ecological insights and interdisciplinary expertise, will benefit the community by offering a new benchmark and innovative AI technologies. Source code is available on https://github.com/liyantett/Tree-Changes-Detection-with-Siamese-Hyperbolic-network. Yante Li, Hanwen Qi, Haoyu Chen 0001, Xinlian Liang, Guoying Zhao 0001 |
CVPR | 3 |
| 2025 | Learning Binary-Antithetical Information Bottleneck for Generalizable Face Anti-SpoofingabstractWe investigate generalizable face anti-spoofing (FAS) using information bottleneck theory. As generalizable FAS aims to detect spoofing in unseen scenarios, it has recently gained significant attention. Existing methods often use adversarial strategies or auxiliary modules to learn domain-invariant features by mining data relationships from distinct source domains. However, their learned feature space may still shift for unseen data due to the spurious correlations overfitted from training domains. Our rationale is that the problem of generalized pattern learning in FAS can be framed as a unified binary-antithetical information transition process, grounded in information bottleneck theory. Specifically, we leverage mutual-information optimization to preserve the instance-level spoof-aware information while compressing domain-related information modeled from the antithetical identity distribution. This enables the model to dynamically identify domain-agnostic, minimal sufficient representations that consistently describe the live/spoof distributions while mitigating spurious correlations through cross-identity compression. In light of this, we propose a novel learning framework for FAS, named Binary-Antithetical Information Bottleneck (BIB)-FAS, which is proven to be effectively generalized to unseen scenarios without using auxiliary information (e.g., domain labels) for training. Extensive cross-domain evaluations show that BIB-FAS significantly outperforms state-of-the-art methods. The code is available at: github.com/CV-AC/BIB-FAS. Hao Yu 0015, Haoyu Chen 0001, Guoying Zhao 0001 |
ICASSP | 2 |
| 2025 | AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language ModelsabstractThe emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level—from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suffers from a lack of large-scale datasets with intensive, descriptive emotion annotations, as well as a multimodal-centric framework to maximize the potential of MLLMs for emotion understanding. To address this, we establish a new benchmark for MLLM-based emotion understanding with a novel dataset (MER-Caption) and a new model (AffectGPT). Utilizing our model-based crowd-sourcing data collection strategy, we construct the largest descriptive emotion dataset to date (by far), featuring over 2K fine-grained emotion categories across 115K samples. We also introduce the AffectGPT model, designed with pre-fusion operations to enhance multimodal integration. Finally, we present MER-UniBench, a unified benchmark with evaluation metrics tailored for typical MER tasks and the free-form, natural language output style of MLLMs. Extensive experimental results show AffectGPT's robust performance across various MER tasks. We have released both the code and the dataset to advance research and development in emotion understanding: https://github.com/zeroQiaoba/AffectGPT. Zheng Lian 0004, Haoyu Chen 0001, Lan Chen 0005, Haiyang Sun 0004, Licai Sun, Yong Ren 0006, Zebang Cheng, Bin Liu 0041, Rui Liu 0008, Xiaojiang Peng, Jiangyan Yi, Jianhua Tao 0001 |
ICML | 2 |
| 2025 | OV-MER: Towards Open-Vocabulary Multimodal Emotion RecognitionabstractMultimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to capture the inherent complexity, subtlety, and multi-appraisal nature of human emotional experiences, as demonstrated by studies in psychology and cognitive science. To overcome this limitation, we advocate for introducing the concept of open vocabulary into MER. This paradigm shift aims to enable models to predict emotions beyond a fixed label space, accommodating a flexible set of categories to better reflect the nuanced spectrum of human emotions. To achieve this, we propose a novel paradigm: Open-Vocabulary MER (OV-MER), which enables emotion prediction without being confined to predefined spaces. However, constructing a dataset that encompasses the full range of emotions for OV-MER is practically infeasible; hence, we present a comprehensive solution including a newly curated database, novel evaluation metrics, and a preliminary benchmark. By advancing MER from basic emotions to more nuanced and diverse emotional states, we hope this work can inspire the next generation of MER, enhancing its generalizability and applicability in real-world scenarios. Code and dataset are available at: https://github.com/zeroQiaoba/AffectGPT. Zheng Lian 0004, Haiyang Sun 0004, Licai Sun, Haoyu Chen 0001, Lan Chen 0005, Zhuofan Wen 0001, Hailiang Yao, Bin Liu 0041, Rui Liu 0008, Shan Liang 0007, Ya Li 0001, Jiangyan Yi, Jianhua Tao 0001 |
ICML | 4 |
| 2025 | MAC 2025: The 2nd Micro-Action Analysis Grand ChallengeabstractMicro-Actions (MAs) are a crucial form of non-verbal communication in social interactions, with promising applications in human emotion analysis. Although the topic has attracted considerable research interest, progress has been hindered by the lack of publicly available benchmark datasets. To address this gap, the Micro-Action Analysis Grand Challenge (MAC) is organized annually. This paper presents an overview of the 2nd Micro-Action Analysis Grand Challenge, held in conjunction with ACM Multimedia 2025. We provide a comprehensive summary of the challenge, including its dataset, evaluation protocol, results, and discussion. The top-ranked solutions are highlighted to offer valuable insights for researchers, and potential future directions are outlined to guide ongoing developments in this area. The goal of this grand challenge is to foster innovative research in micro-action analysis and advance research in the human-centric action understanding community. Kun Li 0008, Dan Guo 0001, Haoyu Chen 0001, Pengyu Liu 0005, Fei Wang 0073, Guoying Zhao 0001, Meng Wang 0001 |
ACM Multimedia | 4 |
| 2025 | To Remember, To Adapt, To Preempt: A Stable Continual Test-Time Adaptation Framework for Remote Physiological Measurement in Dynamic Domain ShiftsabstractRemote photoplethysmography (rPPG) aims to extract non-contact physiological signals from facial videos and has shown great potential. However, existing rPPG approaches struggle to bridge the gap between source and target domains. Recent test-time adaptation (TTA) solutions typically optimize rPPG model for the incoming test videos using self-training loss under an unrealistic assumption that the target domain remains stationary. However, time-varying factors like weather and lighting in dynamic environments often cause continual domain shifts. The erroneous gradients accumulation from these shifts may corrupt the model's key parameters for physiological information, leading to catastrophic forgetting. Therefore, We propose a physiology-related parameters freezing strategy to retain such knowledge. It isolates physiology-related and domain-related parameters by assessing the model's uncertainty to current domain and freezes the physiology-related parameters during adaptation to prevent catastrophic forgetting. Moreover, the dynamic domain shifts with various non-physiological characteristics may lead to conflicting optimization objectives during TTA, which is manifested as the over-adapted model losing its adaptability to future domains. To fix over-adaptation, we propose a preemptive gradient modification strategy. It preemptively adapts to future domains and uses the acquired gradients to modify current adaptation, thereby preserving the model's adaptability. In summary, we propose a stable continual test-time adaptation (CTTA) framework for rPPG measurement, called PhysRAP, which Remembers the past, Adapts to the present, and Preempts the future. Extensive experiments show its state-of-the-art performance, especially in domain shifts. The code is available at https://github.com/xjtucsy/PhysRAP. Shuyang Chu, Jingang Shi, Xu Cheng 0003, Haoyu Chen 0001, Xin Liu 0012, Guoying Zhao 0001 |
ACM Multimedia | 4 |
| 2025 | DSAF: Dual Space Alignment Framework for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a cross-modality retrieval task that aims to match visible and infrared pedestrian images across non-overlapped cameras. However, we observe that three crucial challenges remain inadequately addressed by existing methods: (i) limited discriminative capacity for modality-shared representation, (ii) modality misalignment, and (iii) neglect of identity consistency knowledge. To solve the above issues, we propose a novel dual space alignment framework (DSAF) to constrain the modality in two specific spaces. Specifically, for (i), we design a lightweight and plug-and-play modality invariant enhancement (MIE) module to capture fine-grained semantic information and render identity discriminative. This facilitates the establishment of correlations between visible and infrared modalities, enabling the model to learn robust modality-shared features. To tackle (ii), a dual space alignment (DSA) is introduced to conduct the pixel-level alignment in both Euclidean space and Hilbert space. DSA establishes an elastic relationship between these two spaces, remaining invariant knowledge across two spaces. To solve (iii), we propose an adaptive identity-consistent learning (AIL) to discover identity-consistent knowledge between visible and infrared modalities in a dynamic manner. Extensive experiments on mainstream VI-ReID benchmarks show the superiority and flexibility of our proposed method, achieving competitive performance on mainstream datasets. Xu Cheng 0003, Hao Yu 0015, Haoyu Chen 0001, Guoying Zhao 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Differentiable Auxiliary Learning for Sketch Re-IdentificationabstractSketch re-identification (Re-ID) seeks to match pedestrians' photos from surveillance videos with corresponding sketches. However, we observe that existing works still have two critical limitations: (i) cross- and intra-modality discrepancies hinder the extraction of modality-shared features, (ii) standard triplet loss fails to constrain latent feature distribution in each modality with inadequate samples. To overcome the above issues, we propose a differentiable auxiliary learning network (DALNet) to explore a robust auxiliary modality for Sketch Re-ID. Specifically, for (i) we construct an auxiliary modality by using a dynamic auxiliary generator (DAG) to bridge the gap between sketch and photo modalities. The auxiliary modality highlights the described person in photos to mitigate background clutter and learns sketch style through style refinement. Moreover, a modality interactive attention module (MIA) is presented to align the features and learn the invariant patterns of two modalities by auxiliary modality. To address (ii), we propose a multi-modality collaborative learning scheme (MMCL) to align the latent distribution of three modalities. An intra-modality circle loss in MMCL brings learned global and modality-shared features of the same identity closer in the case of insufficient samples within each modality. Extensive experiments verify the superior performance of our DALNet over the state-of-the-art methods for Sketch Re-ID, and the generalization in sketch-based image retrieval and sketch-photo face recognition tasks. Xu Cheng 0003, Haoyu Chen 0001, Hao Yu 0015, Guoying Zhao 0001 |
AAAI | 3 |
| 2024 | Towards Robust 3D Pose Transfer with Adversarial Learningabstract3D pose transfer that aims to transfer the desired pose to a target mesh is one of the most challenging 3D generation tasks. Previous attempts rely on well-defined parametric human models or skeletal joints as driving pose sources. However, to obtain those clean pose sources, cumbersome but necessary pre-processing pipelines are inevitable, hindering implementations of the real-time applications. This work is driven by the intuition that the robustness of the model can be enhanced by introducing adversarial samples into the training, leading to a more invulnerable model to the noisy inputs, which even can be further extended to directly handling the real-world data like raw point clouds/scans without intermediate processing. Furthermore, we propose a novel 3D pose Masked Autoencoder (3D-PoseMAE), a customized MAE that effectively learns 3D extrinsic presentations (i.e., pose). 3D-PoseMAE facilitates learning from the aspect of extrinsic attributes by simultaneously generating adversarial samples that perturb the model and learning the arbitrary raw noisy poses via a multi-scale masking strategy. Both qualitative and quantitative studies show that the transferred meshes given by our network result in much better quality. Besides, we demonstrate the strong generalizability of our method on various poses, different domains, and even raw scans. Experimental results also show meaningful insights that the intermediate adversarial samples generated in the training can success-fully attack the existing pose transfer models. Haoyu Chen 0001, Hao Tang 0005, Ehsan Adeli-Mosabbeb, Guoying Zhao 0001 |
CVPR | 1 |
| 2024 | Domain Shifting: A Generalized Solution for Heterogeneous Cross-Modality Person Re-Identification
Xu Cheng 0003, Hao Yu 0015, Haoyu Chen 0001, Guoying Zhao 0001 |
ECCV (72) | 5 |
| 2024 | Naive Data Augmentation Might Be Toxic: Data-Prior Guided Self-Supervised Representation Learning for Micro-Gesture RecognitionabstractBody gestures play an important role in nonverbal communication because they transmit emotional information. Recently, a specific group of gestures, so-called Micro-gestures (MGs), has drawn increasing research interests in the community, as they can be useful cues to interpret human inner feelings. In this study, we focused on recognizing MG via self-supervised learning from skeleton sequences with several contributions. Initially, we observed that existing data augmentation methods for skeleton data always fail in MG representation learning. Our investigation shows that the failure is caused by the inherent properties of real-world datasets, such as imbalanced/long-tail data distribution, intra-class ambiguity, and inter-class heterogeneity. Thus, we propose a novel prior-guided augmentation strategy that can preserve the original data distribution while maximizing the agreement between samples in self-supervised learning. Furthermore, we proposed a three-stream architecture of self-supervised presentation learning for micro-gestures via spatial/temporal masking to jointly enhance the learning of invariant features. Lastly, the experimental results show that our proposed method has achieved state-of-the-art performances on two public MG datasets. Atif Shah, Haoyu Chen 0001, Guoying Zhao 0001 |
FG | 2 |
| 2024 | MAC 2024: Micro-Action Analysis Grand ChallengeabstractThis is the overview paper for the Micro-Action Analysis Grand Challenge hosted at ACM Multimedia 2024. In recent years, a growing trend towards deeper understanding of human emotional states has led to a gradual shift in the attention of multimedia and computer vision researchers from macro facial expressions to whole-body micro-actions. Micro-actions are spontaneous body movements that indicate a person's true feelings and potential intentions. Yet, recognizing, distinguishing, and understanding micro-actions is challenging because they are subtle compared to normal actions. This grand challenge aims to foster innovative research in micro-action analysis and provide benchmark evaluations to advance the technology in the human-centric action understanding community. Dan Guo 0001, Kun Li 0008, Haoyu Chen 0001, Guoying Zhao 0001, Yi Yang 0001, Meng Wang 0001 |
ACM Multimedia | 4 |
| 2024 | AGIL-SwinT: Attention-guided inconsistency learning for face forgery detection
Wuti Xiong, Haoyu Chen 0001, Guoying Zhao 0001 |
Image Vis. Comput. | 2 |
| 2024 | Uncertainty-Aware Distillation for Semi-Supervised Few-Shot Class-Incremental LearningabstractGiven a model well-trained with a large-scale base dataset, few-shot class-incremental learning (FSCIL) aims at incrementally learning novel classes from a few labeled samples by avoiding overfitting, without catastrophically forgetting all encountered classes previously. Currently, semi-supervised learning technique that harnesses freely available unlabeled data to compensate for limited labeled data can boost the performance in numerous vision tasks, which heuristically can be applied to tackle issues in FSCIL, i.e., the semi-supervised FSCIL (Semi-FSCIL). So far, very limited work focuses on the Semi-FSCIL task, leaving the adaptability issue of semi-supervised learning to the FSCIL task unresolved. In this article, we focus on this adaptability issue and present a simple yet efficient Semi-FSCIL framework named uncertainty-aware distillation with class-equilibrium (UaD-ClE), encompassing two modules: uncertainty-aware distillation (UaD) and class equilibrium (ClE). Specifically, when incorporating unlabeled data into each incremental session, we introduce the ClE module that employs a class-balanced self-training (CB_ST) to avoid the gradual dominance of easy-to-classified classes on pseudo-label generation. To distill reliable knowledge from the reference model, we further implement the UaD module that combines uncertainty-guided knowledge refinement with adaptive distillation. Comprehensive experiments on three benchmark datasets demonstrate that our method can boost the adaptability of unlabeled data with the semi-supervised learning technique in FSCIL tasks. The code is available at https://github.com/yawencui/UaD-ClE. Yawen Cui, Wanxia Deng, Haoyu Chen 0001, Li Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Adaptive Adversarial Norm Space for Efficient Adversarial Training
Hui Kuurila-Zhang, Haoyu Chen 0001, Guoying Zhao 0001 |
BMVC | 2 |
| 2023 | LART: Neural Correspondence Learning with Latent Regularization Transformer for 3D Motion Transferabstract3D motion transfer aims at transferring the motion from a dynamic input sequence to a static 3D object and outputs an identical motion of the target with high-fidelity and realistic visual effects. In this work, we propose a novel 3D Transformer framework called LART for 3D motion transfer. With carefully-designed architectures, LART is able to implicitly learn the correspondence via a flexible geometry perception. Thus, unlike other existing methods, LART does not require any key point annotations or pre-defined correspondence between the motion source and target meshes and can also handle large-size full-detailed unseen 3D targets. Besides, we introduce a novel latent metric regularization on the Transformer for better motion generation. Our rationale lies in the observation that the decoded motions can be approximately expressed as linearly geometric distortion at the frame level. The metric preservation of motions could be translated to the formation of linear paths in the underlying latent space as a rigorous constraint to control the synthetic motions occurring in the construction of the latent space. The proposed LART shows a high learning efficiency with the need for a few samples from the AMASS dataset to generate motions with plausible visual effects. The experimental results verify the potential of our generative model in applications of motion transfer, content generation, temporal interpolation, and motion denoising. The code is made available: https://github.com/mikecheninoulu/LART. Haoyu Chen 0001, Hao Tang 0005, Radu Timofte, Luc Van Gool, Guoying Zhao 0001 |
NeurIPS | 1 |
| 2023 | SMG: A Micro-gesture Dataset Towards Spontaneous Body Gestures for Emotional Stress State AnalysisabstractAbstract We explore using body gestures for hidden emotional state analysis. As an important non-verbal communicative fashion, human body gestures are capable of conveying emotional information during social communication. In previous works, efforts have been made mainly on facial expressions, speech, or expressive body gestures to interpret classical expressive emotions. Differently, we focus on a specific group of body gestures, called micro-gestures (MGs), used in the psychology research field to interpret inner human feelings. MGs are subtle and spontaneous body movements that are proven, together with micro-expressions, to be more reliable than normal facial expressions for conveying hidden emotional information. In this work, a comprehensive study of MGs is presented from the computer vision aspect, including a novel spontaneous micro-gesture (SMG) dataset with two emotional stress states and a comprehensive statistical analysis indicating the correlations between MGs and emotional states. Novel frameworks are further presented together with various state-of-the-art methods as benchmarks for automatic classification, online recognition of MGs, and emotional stress state recognition. The dataset and methods presented could inspire a new way of utilizing body gestures for human emotion understanding and bring a new direction to the emotion AI community. The source code and dataset are made available: https://github.com/mikecheninoulu/SMG . Haoyu Chen 0001, Henglin Shi, Xin Liu 0012, Guoying Zhao 0001 |
Int. J. Comput. Vis. | 1 |
| 2022 | Geometry-Contrastive Transformer for Generalized 3D Pose TransferabstractWe present a customized 3D mesh Transformer model for the pose transfer task. As the 3D pose transfer essentially is a deformation procedure dependent on the given meshes, the intuition of this work is to perceive the geometric inconsistency between the given meshes with the powerful self-attention mechanism. Specifically, we propose a novel geometry-contrastive Transformer that has an efficient 3D structured perceiving ability to the global geometric inconsistencies across the given meshes. Moreover, locally, a simple yet efficient central geodesic contrastive loss is further proposed to improve the regional geometric-inconsistency learning. At last, we present a latent isometric regularization module together with a novel semi-synthesized dataset for the cross-dataset 3D pose transfer task towards unknown spaces. The massive experimental results prove the efficacy of our approach by showing state-of-the-art quantitative performances on SMPL-NPT, FAUST and our new proposed dataset SMG-3D datasets, as well as promising qualitative results on MG-cloth and SMAL datasets. It's demonstrated that our method can achieve robust 3D pose transfer and be generalized to challenging meshes from unknown spaces on cross-dataset tasks. The code and dataset are made available. Code is available: https://github.com/mikecheninoulu/CGT. Haoyu Chen 0001, Hao Tang 0005, Zitong Yu, Nicu Sebe, Guoying Zhao 0001 |
AAAI | 1 |
| 2022 | Benchmarking 3D Face De-Identification with Preserving Facial AttributesabstractPrivacy with the use of face images is becoming a major concern in civilians’ applications. Recent studies have exploited privacy protection methods by means of facial attributes editing or de-identifying face images. Altering attributes causes loss of information for facial analysis while most de-identification studies did not quantitatively evaluate how well facial attributes are preserved. Moreover, state-of-the-art face analysis utilized 3D information for better performance. Existing face privacy studies only focusing in 2D domain is a key limitation towards the compatibility of more advanced 3D face analysis. This paper presents the first study on the possibility of 3D face de-identification with preserving facial attributes. We systematically evaluate the performance of 2D/3D face/facial attribute recognition and develop 2D/3D de-identification methods with preserving facial attributes using Auto Encoder and Generative Adversarial Networks approaches. We present comprehensive and reproducible experimental results using a publicly available 3D face database with facial attribute annotations for benchmarking and further research. https://github.com/kevinhmcheng/3d-face-de-id Kevin H. M. Cheng, Zitong Yu, Haoyu Chen 0001, Guoying Zhao 0001 |
ICIP | 3 |
| 2022 | WEDAR: Webcam-based Attention Analysis via Attention Regulator Behavior Recognition with a Novel E-reading DatasetabstractHuman attention is critical yet challenging cognitive process to measure due to its diverse definitions and non-standardized evaluation. In this work, we focus on the attention self-regulation of learners, which commonly occurs as an effort to regain focus, contrary to attention loss. We focus on easy-to-observe behavioral signs in the real-world setting to grasp learners’ attention in e-reading. We collected a novel dataset of 30 learners, which provides clues of learners’ attentional states through various metrics, such as learner behaviors, distraction self-reports, and questionnaires for knowledge gain. To achieve automatic attention regulator behavior recognition, we annotated 931,440 frames into six behavior categories every second in the short clip form, using attention self-regulation from the literature study as our labels. The preliminary Pearson correlation coefficient analysis indicates certain correlations between distraction self-reports and unimodal attention regulator behaviors. Baseline model training has been conducted to recognize the attention regulator behaviors by implementing classical neural networks to our WEDAR dataset, with the highest prediction result of 75.18% and 68.15% in subject-dependent and subject-independent settings, respectively. Furthermore, we present the baseline of using attention regulator behaviors to recognize the attentional states, showing a promising performance of 89.41% (leave-five-subject-out). Our work inspires the detection & feedback loop design for attentive e-reading, connecting multimodal interaction, learning analytics, and affective computing. Yoon Lee, Haoyu Chen 0001, Guoying Zhao 0001, Marcus Specht |
ICMI | 2 |
| 2022 | Efficient Dense-Graph Convolutional Network with Inductive Prior Augmentations for Unsupervised Micro-Gesture RecognitionabstractSkeleton-based action/gesture recognition has already witnessed excellent progress on processing large-scale, laboratory-based datasets with pre-defined skeleton joint topology. However, it’s still an unsolved task when it comes to real-world scenarios with practical limitations such as small-scaled dataset sizes, few-labeled samples, and various skeleton topologies. In this paper, we work on the recognition of micro-gestures, which are subtle body gestures collected in real-world scenarios. Specifically, we utilize contrastive learning to heritage the knowledge from known large-scale datasets for enhancing the learning on fewer samples of micro-gestures. To overcome the gap caused by various domain distributions and structure topologies between the datasets, we compute skeleton representations from augmented sequences via momentum-based efficient and scalable encoders as additional inductive priors. Importantly, we propose an effective dense-graph based unsupervised architecture that resorts to a queue-based dictionary to store positive and negative keys for better contrast with queries to learn substantially efficient and discriminant patterns in the feature space. Together with cross-dataset experimental results show that our model significantly improves the accuracies on two micro-gesture datasets, SMG by 7.4% and iMiGUE by 18.41% advocating its superiority. Atif Shah, Haoyu Chen 0001, Henglin Shi, Guoying Zhao 0001 |
ICPR | 2 |
| 2021 | AniFormer: Data-driven 3D Animation with Transformer
Haoyu Chen 0001, Hao Tang 0005, Nicu Sebe, Guoying Zhao 0001 |
BMVC | 1 |
| 2021 | iMiGUE: An Identity-Free Video Dataset for Micro-Gesture Understanding and Emotion AnalysisabstractWe introduce a new dataset for the emotional artificial intelligence research: identity-free video dataset for Micro-Gesture Understanding and Emotion analysis (iMiGUE). Different from existing public datasets, iMiGUE focuses on nonverbal body gestures without using any identity information, while the predominant researches of emotion analysis concern sensitive biometric data, like face and speech. Most importantly, iMiGUE focuses on micro-gestures, i.e., unintentional behaviors driven by inner feelings, which are different from ordinary scope of gestures from other gesture datasets which are mostly intentionally performed for illustrative purposes. Furthermore, iMiGUE is designed to evaluate the ability of models to analyze the emotional states by integrating information of recognized micro-gesture, rather than just recognizing prototypes in the sequences separately (or isolatedly). This is because the real need for emotion AI is to understand the emotional states behind gestures in a holistic way. Moreover, to counter for the challenge of imbalanced sample distribution of this dataset, an unsupervised learning method is proposed to capture latent representations from the micro-gesture sequences themselves. We systematically investigate representative methods on this dataset, and comprehensive experimental results reveal several interesting insights from the iMiGUE, e.g., micro-gesture-based analysis can promote emotion understanding. We confirm that the new iMiGUE dataset could advance studies of micro-gesture and emotion AI. Xin Liu 0012, Henglin Shi, Haoyu Chen 0001, Zitong Yu, Guoying Zhao 0001 |
CVPR | 3 |
| 2021 | Intrinsic-Extrinsic Preserved GANs for Unsupervised 3D Pose TransferabstractWith the strength of deep generative models, 3D pose transfer regains intensive research interests in recent years. Existing methods mainly rely on a variety of constraints to achieve the pose transfer over 3D meshes, e.g., the need for manually encoding for shape and pose disentanglement. In this paper, we present an unsupervised approach to conduct the pose transfer between any arbitrate given 3D meshes. Specifically, a novel Intrinsic-Extrinsic Preserved Generative Adversarial Network (IEP-GAN) is presented for both intrinsic (i.e., shape) and extrinsic (i.e., pose) information preservation. Extrinsically, we propose a co-occurrence discriminator to capture the structural/pose invariance from distinct Laplacians of the mesh. Meanwhile, intrinsically, a local intrinsic-preserved loss is introduced to preserve the geodesic priors while avoiding heavy computations. At last, we show the possibility of using IEP-GAN to manipulate 3D human meshes in various ways, including pose transfer, identity swapping and pose interpolation with latent code vector arithmetic. The extensive experiments on various 3D datasets of humans, animals and hands qualitatively and quantitatively demonstrate the generality of our approach. Our proposed model produces better results and is substantially more efficient compared to recent state-of-the-art methods. Code is available: https://github.com/mikecheninoulu/Unsupervised_IEPGAN Haoyu Chen 0001, Hao Tang 0005, Henglin Shi, Wei Peng 0009, Nicu Sebe, Guoying Zhao 0001 |
ICCV | 1 |
| 2021 | Searching Multi-Rate and Multi-Modal Temporal Enhanced Networks for Gesture RecognitionabstractGesture recognition has attracted considerable attention owing to its great potential in applications. Although the great progress has been made recently in multi-modal learning methods, existing methods still lack effective integration to fully explore synergies among spatio-temporal modalities effectively for gesture recognition. The problems are partially due to the fact that the existing manually designed network architectures have low efficiency in the joint learning of multi-modalities. In this paper, we propose the first neural architecture search (NAS)-based method for RGB-D gesture recognition. The proposed method includes two key components: 1) enhanced temporal representation via the proposed 3D Central Difference Convolution (3D-CDC) family, which is able to capture rich temporal context via aggregating temporal difference information; and 2) optimized backbones for multi-sampling-rate branches and lateral connections among varied modalities. The resultant multi-modal multi-rate network provides a new perspective to understand the relationship between RGB and depth modalities and their temporal dynamics. Comprehensive experiments are performed on three benchmark datasets (IsoGD, NvGesture, and EgoGesture), demonstrating the state-of-the-art performance in both single- and multi-modality settings. The code is available at https://github.com/ZitongYu/3DCDC-NAS. Zitong Yu, Benjia Zhou, Jun Wan 0001, Pichao Wang, Haoyu Chen 0001, Xin Liu 0012, Stan Z. Li, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Learning Graph Convolutional Network for Skeleton-Based Human Action Recognition by Neural SearchingabstractHuman action recognition from skeleton data, fuelled by the Graph Convolutional Network (GCN) with its powerful capability of modeling non-Euclidean data, has attracted lots of attention. However, many existing GCNs provide a pre-defined graph structure and share it through the entire network, which can loss implicit joint correlations especially for the higher-level features. Besides, the mainstream spectral GCN is approximated by one-order hop such that higher-order connections are not well involved. All of these require huge efforts to design a better GCN architecture. To address these problems, we turn to Neural Architecture Search (NAS) and propose the first automatically designed GCN for this task. Specifically, we explore the spatial-temporal correlations between nodes and build a search space with multiple dynamic graph modules. Besides, we introduce multiple-hop modules and expect to break the limitation of representational capacity caused by one-order approximation. Moreover, a corresponding sampling- and memory-efficient evolution strategy is proposed to search in this space. The resulted architecture proves the effectiveness of the higher-order approximation and the layer-wise dynamic graph modules. To evaluate the performance of the searched model, we conduct extensive experiments on two very large scale skeleton-based action recognition datasets. The results show that our model gets the state-of-the-art results in term of given metrics. Wei Peng 0009, Xiaopeng Hong, Haoyu Chen 0001, Guoying Zhao 0001 |
AAAI | 3 |
| 2020 | Temporal Hierarchical Dictionary Guided Decoding for Online Gesture Segmentation and RecognitionabstractOnline segmentation and recognition of skeleton- based gestures are challenging. Compared with offline cases, the inference of online settings can only rely on the current few frames and always completes before whole temporal movements are performed. However, incompletely performed gestures are ambiguous and their early recognition is easy to fall into local optimum. In this work, we address the problem with a temporal hierarchical dictionary to guide the hidden Markov model (HMM) decoding procedure. The intuition is that, gestures are ambiguous with high uncertainty at early performing phases, and only become discriminate after certain phases. This uncertainty naturally can be measured by entropy. Thus, we propose a measurement called "relative entropy map" (REM) to encode this temporal context to guide HMM decoding. Furthermore, we introduce a progressive learning strategy with which neural networks could learn a robust recognition of HMM states in an iterative manner. The performance of our method is intensively evaluated on three challenging databases and achieves state-of-the-art results. Our method shows the abilities of both extracting the discriminate connotations and reducing large redundancy in the HMM transition process. It is verified that our framework can achieve online recognition of continuous gesture streams even when they are halfway performed. Haoyu Chen 0001, Xin Liu 0012, Jingang Shi, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | 3D Skeletal Gesture Recognition via Hidden States ExplorationabstractTemporal dynamics is an open issue for modeling human body gestures. A solution is resorting to the generative models, such as the hidden Markov model (HMM). Nevertheless, most of the work assumes fixed anchors for each hidden state, which make it hard to describe the explicit temporal structure of gestures. Based on the observation that a gesture is a time series with distinctly defined phases, we propose a new formulation to build temporal compositions of gestures by the low-rank matrix decomposition. The only assumption is that the gesture's "hold" phases with static poses are linearly correlated among each other. As such, a gesture sequence could be segmented into temporal states with semantically meaningful and discriminative concepts. Furthermore, different to traditional HMMs which tend to use specific distance metric for clustering and ignore the temporal contextual information when estimating the emission probability, we utilize the long short-term memory to learn probability distributions over states of HMM. The proposed method is validated on multiple challenging datasets. Experiments demonstrate that our approach can effectively work on a wide range of gestures, and achieve state-of-the-art performance. Xin Liu 0012, Henglin Shi, Xiaopeng Hong, Haoyu Chen 0001, Dacheng Tao, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Analyze Spontaneous Gestures for Emotional Stress State Recognition: A Micro-gesture Dataset and Analysis with Deep LearningabstractEmotions are central for human intelligence and should have a similar role in AI. When it comes to emotion recognition, however, analysis cues for robots were mostly limited to human facial expressions and speech. As an alternative important non-verbal communicative fashion, the body gesture is proved to be capable of conveying emotional information which should gain more attention. Inspired by recent researches on micro-expressions, in this paper, we try to explore a specific group of gestures which are spontaneously and unconsciously elicited by inner feelings. These gestures are different from common gestures for facilitating communications or to express feelings on ones own initiative and always ignored in our daily life. This kind of subtle body movements is known as `micro-gestures' (MGs). Work of interpreting the human hidden emotions via these specific gestural behaviors in unconstrained situations, however, is limited. It is because of an unclear correspondence between body movements and emotional states which need multidisciplinary efforts from computer science, psychology, and statistic researchers. To fill the gap, we built a novel Spontaneous Micro-Gesture (SMG) dataset containing 3,692 manually labeled gesture clips. The data collection from 40 participants was conducted through a story-telling game with two emotional state settings. In this paper, we explored the emotional gestures with a sign-based measurement. To verify the latent relationship between emotional states and MGs, we proposed a framework that encodes the objective gestures to a Bayesian network to infer the subjective emotional states. Our experimental results revealed that, most of the participants would do `micro-gestures' spontaneously to relieve their mental strains. We also carried out a human test on ordinary and trained people for comparison. The performance of both our framework and human beings was evaluated on 142 testing instances (71 for each emotional state) by subject-independent testing. To authors' best knowledge, this is the first presented MG dataset. Results showed that the proposed MG recognition method achieved promising performance. We also showed that MGs could be helpful cues for the recognition of hidden emotional states. Haoyu Chen 0001, Xin Liu 0012, Henglin Shi, Guoying Zhao 0001 |
FG | 1 |
| 2019 | Hidden States Exploration for 3D Skeleton-Based Gesture Recognitionabstract3D skeletal data has recently attracted wide attention in human behavior analysis for its robustness to variant scenes, while accurate gesture recognition is still challenging. The main reason lies in the high intra-class variance caused by temporal dynamics. A solution is resorting to the generative models, such as the hidden Markov model (HMM). However, existing methods commonly assume fixed anchors for each hidden state, which is hard to depict the explicit temporal structure of gestures. Based on the observation that a gesture is a time series with distinctly defined phases, we propose a new formulation to build temporal compositions of gestures by the low-rank matrix decomposition. The only assumption is that the gesture's "hold" phases with static poses are linearly correlated among each other. As such, a gesture sequence could be segmented into temporal states with semantically meaningful and discriminative concepts. Furthermore, different to traditional HMMs which tend to use specific distance metric for clustering and ignore the temporal contextual information when estimating the emission probability, the Long Short-Term Memory (LSTM) is utilized to learn probability distributions over states of HMM. The proposed method is validated on two challenging datasets. Experiments demonstrate that our approach can effectively work on a wide range of gestures and actions, and achieve state-of-the-art performance. Xin Liu 0012, Henglin Shi, Xiaopeng Hong, Haoyu Chen 0001, Dacheng Tao, Guoying Zhao 0001 |
WACV | 4 |
| 2018 | Temporal Hierarchical Dictionary with HMM for Fast Gesture RecognitionabstractIn this paper, we propose a novel temporal hierarchical dictionary with hidden Markov model (HMM) for gesture recognition task. Dictionaries with spatio-temporal elements have been commonly used for gesture recognition. However, the existing spatio-temporal dictionary based methods need the whole pre-segmented gestures for inference, thus are hard to deal with nonstationary sequences. The proposed method combines HMM with Deep Belief Networks (DBN) to tackle both gesture segmentation and recognition by the inference at the frame level. Besides, we investigate the redundancy in dictionaries and introduce the relative entropy to measure the information richness of a dictionary. Furthermore, when inferring an element, a temporal hierarchy-flat dictionary will be searched entirely every time in which the temporal structure of gestures isn't utilized sufficiently. The proposed temporal hierarchical dictionary is organized in HMM states and can limit the search range to distinct states. Our framework includes three key novel properties: (1) a temporal hierarchical structure with HMM, which makes both the HMM transition and Viterbi decoding more efficient; (2) a relative entropy model to compress the dictionary with less redundancy; (3) an unsupervised hierarchical clustering algorithm to build a hierarchical dictionary automatically. Our method is evaluated on two gesture datasets and consistently achieves state-of-the-art performance. The results indicate that the dictionary redundancy has a significant impact on the performance which can be tackled by a temporal hierarchy and an entropy model. Haoyu Chen 0001, Xin Liu 0012, Guoying Zhao 0001 |
ICPR | 1 |