Feng Zhou 0007

dblp:21/6430-7 · DBLP profile ↗
← Back
25ranked-venue papers
11as first author
22since 2021 · last 2026
0000-0001-9184-2040ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 9 first-author · 17 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ECLIPSE: Continuous Alpha Field Modulation for Zero-Shot Educational Facial Expression Recognition
Yixiao Xu, Yulian Sheng, Junxuan Bai, Feng Zhou 0007, Ju Dai, JunJun Pan
ICIC (19)4
2026 EmoPoseFace: Head Pose Aware Speech-Driven 3D Emotional Facial Animation Using Latent Diffusion
abstract
Speech-driven 3D facial animation has notable applications in the VR domain, including virtual anchors and digital avatars, etc. However, producing facial animations that convey complex emotional expressions remains a substantial challenge. Existing methods struggle to simultaneously achieve accurate lip synchronization, natural facial expressions, and realistic emotional representation. Significantly, the impact of head pose on boosting facial emotional expressiveness has not been thoroughly investigated. To address these issues, we propose EmoPoseFace, a novel Diffusion-based network to generate speech-driven 3D emotional facial animations with synchronized head poses. Our method employs a dual-branch conditional generation architecture to separately model facial expressions and head poses, integrating emotion and head-pose conditions for coherent facial expression-pose control. In addition, we design the Global-local Facial Fine-grained Editing Module (GL-FFE), which achieves emotional enhancement of facial expressions and fine-grained facial modification, while maintains the naturalness and authenticity of facial movements. Extensive experiments demonstrate that our approach outperforms existing methods in lip-sync accuracy and emotional detail preservation. The introduction of head pose control and GL-FFE significantly expands the expressiveness of emotional virtual facial animation, and the fine-grained editing is widely approved in perceptual user studies.
Xin Zhao 0025, Ju Dai, Feng Zhou 0007, Haofei Wang 0001, Aimin Hao, Hong Qin 0001, Yang Gao 0032
IEEE Trans. Vis. Comput. Graph.3
2026 EmoDiffuser: emotional diffuser for speech-driven 3D facial animation
Xin Zhao 0025, Ju Dai, Feng Zhou 0007, Haofei Wang 0001, Aimin Hao, JunJun Pan, Yang Gao 0032
Vis. Comput.3
2026 Multi-level fusion tokens for enhanced self-supervised skeleton-based action recognition
Kaida Ning, Feng Zhou 0007, JunJun Pan, Hongwen Xu, Ju Dai
Vis. Comput.3
2026 CLIP-Hand: CLIP-based regressor for hand pose estimation and mesh recovery
Feng Zhou 0007, Shuang Ji, Pei Shen, Ju Dai, JunJun Pan, Yukun Lai, Paul L. Rosin
Vis. Comput.1
2026 Mask-aware tri-modal learning for indoor 3D object detection
Feng Zhou 0007, Kaida Ning, JunJun Pan, Jin Li 0068, Ju Dai
Vis. Comput.1
2026 AgeStyle: Dynamic age-guided motion transfer in virtual reality
abstract
Age significantly influences human motor patterns, yet existing virtual reality (VR) systems lack dynamic modelling of these variations. This paper introduces AgeStyle, a versatile framework that integrates age-guided style selection with motion style transfer to convert user-uploaded videos into interactive 3D motion models. Utilizing 2D joint detection and 3D pose estimation, AgeStyle constructs motion representations enhanced by a CLIP-driven Cross-Attention module, capturing the distinct traits of different age groups—child flexibility, adult efficiency, and elderly stability. Our system enables real-time switching between motion styles and perspectives through voice commands, offering an immersive exploration of age-related movements. Quantitative experiments on the XIA dataset demonstrate AgeStyle’s competitive performance in both content preservation and style consistency, achieving average CC and SC++ scores of 7.4 and 14.8, respectively. AgeStyle represents a meaningful advancement in VR character design, with broad potential applications in education, healthcare, rehabilitation, and interactive entertainment. For the demo, please refer to https://youtu.be/eo7Shy0Ukps . The source code of AgeStyle is available at https://github.com/codeozzz/ageStyle .
Feng Zhou 0007, Ju Dai, Sen-Zhe Xu 0001
Virtual Real. Intell. Hardw.1
2025 Position-Aware Guided Point Cloud Completion with CLIP Model
abstract
Point cloud completion aims to recover partial geometric and topological shapes caused by equipment defects or limited viewpoints. Current methods either solely rely on the 3D coordinates of the point cloud to complete it or incorporate additional images with well-calibrated intrinsic parameters to guide the geometric estimation of the missing parts. Although these methods have achieved excellent performance by directly predicting the location of complete points, the extracted features lack fine-grained information regarding the location of the missing area. To address this issue, we propose a rapid and efficient method to expand an unimodal framework into a multimodal framework. This approach incorporates a position-aware module designed to enhance the spatial information of the missing parts through a weighted map learning mechanism. In addition, we establish a Point-Text-Image triplet corpus PCI-TI and MVP-TI based on the existing unimodal point cloud completion dataset and use the pre-trained vision-language model CLIP to provide richer detail information for 3D shapes, thereby enhancing performance. Extensive quantitative and qualitative experiments demonstrate that our method outperforms state-of-the-art point cloud completion methods.
Feng Zhou 0007, Ju Dai, Lei Li 0050, Junliang Xing
AAAI1
2025 Human Motion Instruction Tuning
abstract
This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model’s ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domain-specific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction.
Lei Li 0050, Sen Jia 0003, Zhongyu Jiang, Feng Zhou 0007, Ju Dai, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang
CVPR5
2025 Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation
abstract
In 3D speech-driven facial animation generation, existing methods commonly employ pre-trained self-supervised audio models as encoders. However, due to the prevalence of phonetically similar syllables with distinct lip shapes in language, these near-homophone syllables tend to exhibit significant coupling in self-supervised audio feature spaces, leading to the averaging effect in subsequent lip motion generation. To address this issue, this paper proposes a plug-and-play semantic decorrelation module—Wav2Sem. This module extracts semantic features corresponding to the entire audio sequence, leveraging the added semantic information to decorrelate audio encodings within the feature space, thereby achieving more expressive audio features. Extensive experiments across multiple Speech-driven models indicate that the Wav2Sem module effectively decouples audio features, significantly alleviating the averaging effect of phonetically similar syllables in lip shape generation, thereby enhancing the precision and naturalness of facial animations. Our source code is available at https://github.com/wslh852/Wav2Sem.git.
Ju Dai, Xin Zhao 0025, Feng Zhou 0007, JunJun Pan, Lei Li 0050
CVPR4
2025 Chat-Driven 3D Human Pose and Shape Editing with Large Language Models
abstract
Generating and creating humanoid 3D models has received increasing attention recently due to its fundamental support for many high-level 3D applications. Although automatic 3D pose and shape reconstruction methods have achieved promising results, there are still some failure cases due to self-occlusions, viewpoint changes, and the complexity of human pose articulations. In this paper, we propose a novel way to leverage Large Language Models (LLMs) to interactively reconstruct human pose and shape based on a Skinned Multi-Person Linear (SMPL) model. We construct a mapping table to fine-tune an LLM, enabling it to understand user inputs better and output the positional information of joint points. Additionally, a simple neural network is adopted to regress the shape cues of the SMPL. We demonstrate a gallery of results of numerous poses and shapes. We validate our method via numerical evaluations, user studies, and comparisons to manually posed characters and previous work.
Feng Zhou 0007, Ju Dai, Mengxiao Zhu 0004, Yongmei Zhang, Yukun Lai, Paul L. Rosin
ICASSP1
2025 AU-Blendshape for Fine-Grained Stylized 3D Facial Expression Manipulation
abstract
While 3D facial animation has made impressive progress, challenges still exist in realizing fine-grained stylized 3D facial expression manipulation due to the lack of appropriate datasets. In this paper, we introduce the AUBlendSet, a 3D facial dataset based on AU-Blendshape representation for fine-grained facial expression manipulation across identities. AUBlendSet is a blendshape data collection based on 32 standard facial action units (AUs) across 500 identities, along with an additional set of facial postures annotated with detailed AUs. Based on AUBlendSet, we propose AUBlendNet to learn AU-Blendshape basis vectors for different character styles. AUBlendNet predicts, in parallel, the AU-Blendshape basis vectors of the corresponding style for a given identity mesh, thereby achieving stylized 3D emotional facial manipulation. We comprehensively validate the effectiveness of AUBlendSet and AUBlendNet through tasks such as stylized facial expression manipulation, speech-driven emotional facial animation, and emotion recognition data augmentation. Through a series of qualitative and quantitative experiments, we demonstrate the potential and importance of AUBlendSet and AUBlendNet in 3D facial animation tasks. To the best of our knowledge, AUBlendSet is the first dataset, and AUBlendNet is the first network for continuous 3D facial expression manipulation for any identity through facial AUs. Our source code is available at https://github.com/wslh852/AUBlendNet.git.
Ju Dai, Feng Zhou 0007, Kaida Ning, Lei Li 0050, JunJun Pan
ICCV3
2025 Learning Hanzi Character Through VR-Based Mortise-Tenon
Conglin Ma, Sen-Zhe Xu 0001, Ju Dai, Jie Liu 0022, Feng Zhou 0007
ICXR6
2025 Diffusion model with temporal constraint for 3D human pose estimation
Zhangmeng Chen, Ju Dai, JunJun Pan, Feng Zhou 0007
Vis. Comput.4
2025 Dual-path spatio-temporal Mamba for skeleton-based action recognition
Ju Dai, Feng Zhou 0007, JunJun Pan, Hongwen Xu
Vis. Comput.3
2024 AHRNET: Attention and Heatmap-Based Regressor for Hand Pose Estimation and Mesh Recovery
abstract
Estimating 3D hand pose and recovering the full hand surface mesh from a single RGB image is a challenging task due to self-occlusions, viewpoint changes, and the complexity of hand articulations. In this paper, we propose a novel framework that combines an attention mechanism with heatmap regression to accurately and efficiently predict 3D joint locations and reconstruct the hand mesh. We adopt a pooling attention module that learns to focus on relevant regions in the input image to extract better features for handling occlusions, while greatly reducing the computational cost. The multi-scale 2D heatmaps provide spatial constraints to guide the 3D vertex predictions. By exploiting the complementary strengths of sparse 2D supervision and dense mesh regression, our method accurately reconstructs hand meshes with realistic details. Extensive experiments on standard benchmarks demonstrate that the proposed method efficiently improves the performance of 3D hand pose estimation and mesh recovery. The reproducible recipes are available at https://github.com/SDiannn/AHRNET-Heatmap.
Feng Zhou 0007, Pei Shen, Ju Dai, Yukun Lai, Paul L. Rosin
ICASSP1
2024 Enhancing Real-Time Fluid Simulations with Lagrangian Methods and GPU-Based Techniques in Unity
Yanrui Sun, Feng Zhou 0007, Ju Dai
ICXR2
2024 Free editing of Shape and Texture with Deformable Net for 3D Caricature Generation
Yuanyuan lin, Ju Dai, JunJun Pan, Feng Zhou 0007, Junxuan Bai
Vis. Comput.4
2023 GFENet: Group-Free Enhancement Network for Indoor Scene 3D Object Detection
Feng Zhou 0007, Ju Dai, JunJun Pan, Mengxiao Zhu 0004, Xingquan Cai, Chen Wang 0043
CGI (3)1
2023 KD-Former: Kinematic and dynamic coupled transformer network for 3D human motion prediction
Ju Dai, Junxuan Bai, Feng Zhou 0007, JunJun Pan
Pattern Recognit.5
2022 MCGNet: Multi-Level Context-aware and Geometric-aware Network for 3D Object Detection
abstract
Hough voting based on PointNet++ [1] is effective against 3D object detection, which has been verified by VoteNet [2], H3DNet [3], etc. However, we find there is still room for improvements in two aspects. The first is that most existing methods ignores the particular significance of different format inputs and geometric primitives for predicting object proposals. The second is that the feature extracted by PointNet++ overlooks contextual information about each object. In this paper, to tackle the above issues, we introduce MCGNet to learn multi-level geometric-aware and scale-aware contextual information for 3D object detection. Specifically, our network mainly consists of the baseline module based on H3DNet, geometric-aware module, and context-aware module. The baseline module feeding with four-types inputs (Point, Edge, Surface, and Line) concentrates on extracting diversified geometric primitives, i.e., BB centers, BB face centers, and BB edge centers. The geometric-aware module is proposed to learn the different contributions among the four-types feature maps and the three geometric primitives. The context-aware module aims to establish long-range dependencies features for either four-types feature maps or three geometric primitives. Extensive experiments on two large datasets with real 3D scans, SUN RGB-D and ScanNet datasets, demonstrate that our method is effective against 3D object detection.
Keng Chen, Feng Zhou 0007, Ju Dai, Pei Shen, Xingquan Cai, Fengquan Zhang
ICIP2
2022 Scale-aware network with modality-awareness for RGB-D indoor semantic segmentation
Feng Zhou 0007, Yukun Lai, Paul L. Rosin, Fengquan Zhang
Neurocomputing1
2020 Scale-aware spatial pyramid pooling with both encoder-mask and scale-attention for semantic segmentation
Feng Zhou 0007, Xukun Shen
Neurocomputing1
2019 MSANet: multimodal self-augmentation and adversarial network for RGB-D object recognition
Feng Zhou 0007, Xukun Shen
Vis. Comput.1
2018 GRANet: Global Refinement Atrous Convolutional Neural Network for Semantic Scene Segmentation
abstract
The main problems of complex-scene understanding and semantic scene segmentation are caused by mismatched relationships, confusion categories, and inconspicuous classes. Towards above issues, we propose a global refinement atrous convolutional neural network (GRANet) for semantic scene segmentation. To enlarge the receptive field of filters, we use atrous convolution instead of the downsampling operators. To handle the challenge caused by the existence of objects at multiple scales in a scene, we adopt multiple rates atrous convolution structure. And to overcome the problem that the current semantic segmentation architecture can not make good use of global information, we propose a multiple pooling module schemes to utilize the global context information to boost the performance of our GRANet. The proposed GRANet achieves state-of-the-art performance on the SiftFlow Dataset and attains comparable performance with other state-of-the-art works on Cityscapes dataset.
Feng Zhou 0007, Xukun Shen
ICIP1