Hanyu Jiang 0004

dblp:169/1040-4 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0003-9258-521XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Continuous Action Unit Intensity Modeling for Micro-Expression Recognition
abstract
Micro-Expression Recognition (MER) remains challenging due to the subtle and transient nature of facial muscle movements. While recent methods leverage Action Unit (AU) labels for MER, they often tend to ignore continuous AU intensity variations, which are critical for capturing nuanced facial expressions. To address these limitations, we propose a novel framework integrating continuous AU intensity with hierarchical motion modeling. Our approach begins with a lightweight model that regresses in-frame AU intensity values. These AU intensities are fed into our proposed Continuous AU Transformer (CAUT), which employs a temporal Transformer and a spatial Transformer to model AU evolution across frames and inter-AU dependencies. Simultaneously, a two-stage Transformer architecture extracts hierarchical optical flow features, fused with AU semantics via a multi-scale region-based fusion strategy for enhancing facial motion features. Extensive experiments demonstrate the proposed method’s state-of-the-art performance, validating the effectiveness of continuous AU intensity modeling and hierarchical feature integration for MER.
Hanyu Jiang 0004, Jiayi Lyu, Xing Lan, Jian Xue 0002
ICIP1
2025 One General Plug-In for Facial Heatmap-based Keypoint Detection
abstract
In this paper, we systematically investigate the error distribution in predicted heatmaps for face alignment, and point out that previous works are unreliable in following the rule that decodes coordinates by locating the maximum-score pixel. Our research reveals that the majority of ground-truth positions do not match that pixel but rather lie within a range of a few pixels. Building on this phenomenon, we transform the model’s objective from predicting inaccurate landmarks to identifying precise proposals with that range. We propose a simple but effective module, termed the Response Aware Module (RAM), leveraging response scores in the proposal to regress the proposal offset, which can be used as a plug-and-play layer integrated into public models. Furthermore, we present a novel Heatmap RCNN framework to exploit the distribution of multi-scale heatmaps. Extensive experiments have demonstrated that the trained RAM can be integrated seamlessly as a ready-to-use plugin with the model, yielding impressive improvements. Meanwhile, Heatmap RCNN performs far superior to SOTA results, with 3.82 NME on WFLW, 3.09 on COFW, and 2.90 on 300W.
Hanyu Jiang 0004, Jian Xue 0002, Xing Lan, Ke Lu 0002
ICME1
2025 Multimodal Emotional Talking Face Generation Based on Action Units
abstract
Talking face generation focuses on creating natural facial animations that align with the provided text or audio input. Current methods in this field primarily rely on facial landmarks to convey emotional changes. However, spatial key-points are valuable, yet limited in capturing the intricate dynamics and subtle nuances of emotional expressions due to their restricted spatial coverage. Consequently, this reliance on sparse landmarks can result in decreased accuracy and visual quality, especially when representing complex emotional states. To address this issue, we propose a novel method called Emotional Talking with Action Unit (ETAU), which seamlessly integrates facial Action Units (AUs) into the generation process. Unlike previous works that solely rely on facial landmarks, ETAU employs both Action Units and landmarks to comprehensively represent facial expressions through interpretable representations. Our method provides a detailed and dynamic representation of emotions by capturing the complex interactions among facial muscle movements. Moreover, ETAU adopts a multi-modal strategy by seamlessly integrating emotion prompts, driving videos, and target images, and by leveraging various input data effectively, it generates highly realistic and emotional talking-face videos. Through extensive evaluations across multiple datasets, including MEAD, LRW, GRID and HDTF, ETAU outperforms previous methods, showcasing its superior ability to generate high-quality, expressive talking faces with improved visual fidelity and synchronization. Moreover, ETAU exhibits a significant improvement on the emotion accuracy of the generated results, reaching an impressive average accuracy of 84% on the MEAD dataset.
Jiayi Lyu, Xing Lan, Guohong Hu, Hanyu Jiang 0004, Jinbao Wang 0001, Jian Xue 0002
IEEE Trans. Circuits Syst. Video Technol.4
2025 FoodSAM: Any Food Segmentation
abstract
In this paper, we explore the zero-shot capability of the Segment Anything Model (SAM) for food image segmentation. To address the lack of class-specific information in SAM-generated masks, we propose a novel framework, calledFoodSAM. This innovative approach integrates the coarse semantic mask with SAM-generated masks to enhance semantic segmentation quality. Besides, we recognize that the ingredients in food can be supposed as independent individuals, which motivated us to perform instance segmentation on food images. Furthermore, FoodSAM extends its zero-shot capability to encompass panoptic segmentation by incorporating an object detector, which renders FoodSAM to effectively capture non-food object information. Drawing inspiration from the recent success of promptable segmentation, we also extend FoodSAM to promptable segmentation, supporting various prompt variants. Consequently, FoodSAM emerges as an all-encompassing solution capable of segmenting food items at multiple levels of granularity. Remarkably, this pioneering framework stands as the first-ever work to achieve instance, panoptic, and promptable segmentation on food images. Extensive experiments demonstrate the feasibility and impressing performance of FoodSAM, validating SAM's potential as a prominent and influential tool within the domain of food image segmentation.
Xing Lan, Jiayi Lyu, Hanyu Jiang 0004, Kun Dong 0001, Zehai Niu, Yi Zhang 0162, Jian Xue 0002
IEEE Trans. Multim.3
2024 ETAU: Towards Emotional Talking Head Generation Via Facial Action Unit
abstract
Creating expressive talking heads is crucial for multimedia applications involving virtual human. Existing approaches predominantly rely on facial landmarks to convey emotional changes. However, these spatial keypoints struggle to capture subtle emotional intricacies due to their limited spatial coverage, consequently decreasing accuracy and visual quality, particularly in emotion representation. To address this issue, we introduce a novel method called Emotional Talking with Action Unit (ETAU), which introduces the additional facial Action Units (AUs) to generate talking head video that accurately portray the target emotions. Unlike previous works, ETAU comprehensively quantify facial expressions through Action Units, which provides a detailed and dynamic representation of emotion. To the best of our knowledge, this work pioneers the integration of Action Units for emotional talking head generation. Extensive evaluations on the MEAD dataset showcase ETAU’s state-of-the-art performance with 21.89 PSNR and 0.68 SSIM. Critically, ETAU achieves significant improvement in emotion accuracy of the generated results, reaching 84%, confirming its feasibility in representing emotional expressions.
Jiayi Lyu, Xing Lan, Guohong Hu, Hanyu Jiang 0004, Jian Xue 0002
ICME4
2024 Does Pixel Value Represent Facial Landmark Well in Heatmap?
abstract
Heatmap-based methods have dominated the face alignment task, yet the maximum response decoding scheme necessitates further reform. While some studies have attempted to compensate for prediction offsets using a post-processing module, the prediction errors induced by the maximum response decoding scheme remain challenging to rectify. In this paper, we assume that using heatmap value to denote the ground-truth probability is not accurate enough. To cure this problem, we propose DISPAL, a novel DIStribution-based Probability for fAcial Landmarks, which signifies the ground-truth probability by the similarity between the pixel’s neighbouring value distribution and Gaussian distribution. This innovative probability enables us to pinpoint the keypoint location more robustly than previous methods that rely solely on the peak score. It also exhibits remarkable generalization to complex decoding methodologies. Furthermore, we propose supervising this probability as an additional task loss to help the model learn better heatmap representation. Extensive empirical results on WFLW, 300W, and COFW datasets demonstrate that our distribution-based probability mechanism significantly surpasses original value-based probability approaches.
Xing Lan, Jiayi Lyu, Kun Dong 0001, Hanyu Jiang 0004, Qinghao Hu 0001, Jian Xue 0002
IEEE Trans. Circuits Syst. Video Technol.4