Fangzhe Nan

dblp:263/2431 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0008-1115-532XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 R²D-LPCC: Relevance-Ranking Guided Region-Adaptive Dynamic LiDAR Point Cloud Compression
abstract
Dynamic LiDAR point cloud compression (LPCC) is crucial for the efficient transmission and storage of large-scale three-dimensional data in applications such as autonomous driving. However, many existing methods, which primarily focus on compressing geometric or motion information, face a fundamental limitation: they treat all points as equally important. This approach neglects the semantic priorities of a scene, resulting in inefficient bit allocation and particularly compromising the reconstruction quality of safety-critical regions, such as pedestrians and vehicles, which are vital to downstream perception tasks. To address these limitations, we propose R²D-LPCC, a relevance-ranking framework for region adaptive LPCC that prioritizes fidelity in semantically important regions. Central to our approach is the Adaptive Relevance Learning (ARL) module, which integrates semantic context with uncertainty to evaluate regional significance and guide compression. We also introduce a Multi-scale Region-Adaptive Transform (MRAT) module to enhance semantic feature modeling and preserve fine-grained details in key areas. Additionally, we develop an adaptive multi-modal motion estimation module to improve motion prediction in complex three-dimensional environments. Extensive experiments conducted on the SemanticKITTI benchmark demonstrate that R²D-LPCC significantly surpasses ten recent state-of-the-art methods, achieving a 45.48% BD-rate gain over the previous leading method, Unicorn, and a 98.58% gain over the GPCC standard, while ensuring superior reconstruction quality in semantically important regions.
Fangzhe Nan, Frederick W. B. Li, Gary K. L. Tam, Zhaoyi Jiang, Bailin Yang, Jingke Cui, Changshuo Wang 0001
AAAI1
2025 Speech-Driven 3D Facial Animation with Regional Attention for Style Capture
Bailin Yang, Fangzhe Nan
CASA3
2025 MP-DPCC: A Motion Proxy-Based Dynamic Point Cloud Compression Framework
abstract
The increasing data volume and the demand for real-time transmission highlight the necessity for efficient compression of dynamic point cloud data. Existing methods primarily focus on reducing inter-frame redundancy by calculating per-point motion information, overlooking the computational and storage costs involved. In this paper, we propose a novel motion proxy-based dynamic point cloud compression framework to enhance the efficiency and accuracy of motion information utilization. Specifically, we introduce a feature proxy module to adaptively locate proxy points, which represent the overall motion through the motion of proxy points. Additionally, a motion enhancement module is employed to refine motion details and prevent local information loss caused by dense motion trajectories. Extensive experiments demonstrate the superiority of our approach. Compared with baseline methods, it achieves an average BD-rate improvement of 12.07% (D1) and 11.90% (D2).
Zhaoyi Jiang, Cao Song, Fangzhe Nan, Bailin Yang
ICASSP4
2025 Multi-modal Dynamic Point Cloud Geometric Compression Based on Bidirectional Recurrent Scene Flow
abstract
Deep learning methods have recently shown significant promise in compressing the geometric features of point clouds. However, challenges arise when consecutive point clouds contain holes, resulting in incomplete information that complicates motion estimation. To our knowledge, most existing dynamic point cloud compression methods have largely overlooked this critical issue. Moreover, these methods typically employ a multi-scale single-pass approach for motion estimation, performing only one estimation at each scale. This limits accuracy and adversely impacts compression performance. To address these challenges, we propose a dynamic point cloud compression model called M2BR-DPCC (Multi-Modal Multi-Scale Bidirectional Recursion for Dynamic Point Cloud Compression). Our method introduces two key innovations. First, we integrate both point cloud and image data as inputs, leveraging a multi-modal feature representation completion (MFRepC) approach to align information across modalities. This addresses the issue of missing data in point clouds by using complementary information from images. Second, we implement a multi-scale bidirectional recursive (MSBR) motion estimation method. This module iteratively refines motion flows in both forward and backward directions, progressively enhancing point cloud features and improving motion estimation accuracy. Experimental results on widely used datasets, including MVUB and 8iVFB, demonstrate the effectiveness of our approach. Compared to existing methods, M2BR-DPCC achieves superior performance, with an average BD-rate improvement of 95.23% over V-PCC, 12.92% over D-DPCC, and 16.16% over patchDPCC. These results underscore the potential of leveraging multi-modal data and bidirectional refinement for dynamic point cloud compression.
Fangzhe Nan, Frederick W. B. Li, Zhuoyue Wang, Gary K. L. Tam, Zhaoyi Jiang, DongZheng DongZheng, Bailin Yang
ICASSP1
2025 Seeing the Overlooked: Bio-Visual Inspired Weak Saliency Feedback Transformer for Person Re-identification
abstract
The domain gap between pretraining data (e.g., ImageNet, LUPerson) and downstream ReID datasets often leads to suboptimal performance when directly fine-tuning pretrained models. While existing methods attempt to bridge this gap by incorporating additional modalities (e.g., text, 3D data) or visual cues (e.g., pose, body masks), these approaches introduce two key limitations: (1) they may distract the model with irrelevant factors like background clutter or clothing variations, and (2) they inevitably increase computational overhead during inference. To address these issues, we propose the Weak Saliency Feedback Transformer (WSFFormer), inspired by the feedback mechanisms in biological visual systems. Unlike traditional one-way feature propagation, WSFFormer employs an adaptive feedback loop during training to enhance low-response regions, enabling the model to capture richer and more discriminative features. The WSFFormer introduces three key components: (1) The Lateral Feedback Module (LFM) mimics retinal lateral inhibition by adaptively suppressing high-response regions and amplifying weak discriminative features, forcing attention on subtle details; (2) The Progressive Feedback Module (PFM) refines feedback through deep-to-shallow closed-loop propagation, blending high-level semantics with spatial details; (3) The Feedback Sensitive Entropy Loss (FSE Loss) optimizes target-domain adaptation by quantifying divergence between forward and feedback-corrected features. Experiments on holistic/occluded ReID benchmarks show WSFFormer outperforms ViT/Swin-based SOTA methods without extra inference cost.
Changshuo Wang 0001, Shuting He, Fangzhe Nan, Prayag Tiwari
ACM Multimedia4
2025 CLPFusion: A Latent Diffusion Model Framework for Realistic Chinese Landscape Painting Style Transfer
abstract
ABSTRACT This study focuses on transforming real‐world scenery into Chinese landscape painting masterpieces through style transfer. Traditional methods using convolutional neural networks (CNNs) and generative adversarial networks (GANs) often yield inconsistent patterns and artifacts. The rise of diffusion models (DMs) presents new opportunities for realistic image generation, but their inherent noise characteristics make it challenging to synthesize pure white or black images. Consequently, existing DM‐based methods struggle to capture the unique style and color information of Chinese landscape paintings. To overcome these limitations, we propose CLPFusion, a novel framework that leverages pre‐trained diffusion models for artistic style transfer. A key innovation is the Bidirectional State Space Models‐CrossAttention (BiSSM‐CA) module, which efficiently learns and retains the distinct styles of Chinese landscape paintings. Additionally, we introduce two latent space feature adjustment methods, Latent‐AdaIN and Latent‐WCT, to enhance style modulation during inference. Experiments demonstrate that CLPFusion produces more realistic and artistic Chinese landscape paintings than existing approaches, showcasing its effectiveness and uniqueness in the field.
Jiahui Pan 0001, Frederick W. B. Li, Bailin Yang, Fangzhe Nan
Comput. Animat. Virtual Worlds4
2025 Talking Face Generation With Lip and Identity Priors
abstract
ABSTRACT Speech‐driven talking face video generation has attracted growing interest in recent research. While person‐specific approaches yield high‐fidelity results, they require extensive training data from each individual speaker. In contrast, general‐purpose methods often struggle with accurate lip synchronization, identity preservation, and natural facial movements. To address these limitations, we propose a novel architecture that combines an alignment model with a rendering model. The rendering model synthesizes identity‐consistent lip movements by leveraging facial landmarks derived from speech, a partially occluded target face, multi‐reference lip features, and the input audio. Concurrently, the alignment model estimates optical flow using the occluded face and a static reference image, enabling precise alignment of facial poses and lip shapes. This collaborative design enhances the rendering process, resulting in more realistic and identity‐preserving outputs. Extensive experiments demonstrate that our method significantly improves lip synchronization and identity retention, establishing a new benchmark in talking face video generation.
Frederick W. B. Li, Gary K. L. Tam, Bailin Yang, Fangzhe Nan, Jia Pan 0001
Comput. Animat. Virtual Worlds5
2024 Multi-style cartoonization: Leveraging multiple datasets with generative adversarial networks
abstract
Abstract Scene cartoonization aims to convert photos into stylized cartoons. While generative adversarial networks (GANs) can generate high‐quality images, previous methods focus on individual images or single styles, ignoring relationships between datasets. We propose a novel multi‐style scene cartoonization GAN that leverages multiple cartoon datasets jointly. Our main technical contribution is a multi‐branch style encoder that disentangles representations to model styles as distributions over entire datasets rather than images. Combined with a multi‐task discriminator and perceptual losses optimizing across collections, our model achieves state‐of‐the‐art diverse stylization while preserving semantics. Experiments demonstrate that by learning from inter‐dataset relationships, our method translates photos into cartoon images with improved realism and abstraction fidelity compared to prior arts, without iterative re‐training for new styles.
Jianlu Cai, Frederick W. B. Li, Fangzhe Nan, Bailin Yang
Comput. Animat. Virtual Worlds3
2023 Multi-view frontal face image generation: A survey
abstract
Abstract Face images from different perspectives reduce the accuracy of face recognition, and the generation of frontal face images is an important research topic in the field of face recognition. To understand the development of frontal face generation models and grasp the current research hotspots and trends, existing methods based on 3D models, deep learning, and hybrid models are summarized, and the current commonly used face generation methods are introduced. Dataset, and compare the performance of existing models through experiments. The purpose of this paper is to fundamentally understand the advantages of existing frontal face generation, sort out the key issues of such generation, and look toward future development trends.
Xin Ning 0001, Fangzhe Nan, Shaohui Xu, Liping Zhang 0014
Concurr. Comput. Pract. Exp.2
2022 Image super-resolution reconstruction based on generative adversarial network model with feedback and attention mechanisms
Yongqiang Wang 0003, Fangzhe Nan, Yurong Qian
Multim. Tools Appl.3
2020 Continuous Learning of Face Attribute Synthesis
abstract
The generative adversarial network (GAN) exhibits great superiority in the face attribute synthesis task. However, existing methods have very limited effects on the expansion of new attributes. To overcome the limitations of a single network in new attribute synthesis, a continuous learning method for face attribute synthesis is proposed in this work. First, the feature vector of the input image is extracted and attribute direction regression is performed in the feature space to obtain the axes of different attributes. The feature vector is then linearly guided along the axis so that images with target attributes can be synthesized by the decoder. Finally, to make the network capable of continuous learning, the orthogonal direction modification module is used to extend the newly-added attributes. Experimental results show that the proposed method can endow a single network with the ability to learn attributes continuously, and, as compared to those produced by the current state-of-the-art methods, the synthetic attributes have higher accuracy.
Xin Ning 0001, Weijun Li 0002, Xiaoli Dong, Shaohui Xu, Fangzhe Nan, Yuanzhou Yao
ICPR5
2020 Single Image Super-Resolution Reconstruction based on the ResNeXt Network
Fangzhe Nan, Qingliang Zeng, Yanni Xing, Yurong Qian
Multim. Tools Appl.1