Huanyu Zhou

dblp:44/8489 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Robot navigation and mapping · 26% 3D vision · 26% Video understanding and tracking · 22%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
point cloud processing
1.922026
DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations · AAAI 2026
TDFANet: Encoding Sequential 4D Radar Point Clouds Using Trajectory-Guided Deformable Feature Aggregation for Place Recognition · ICRA 2025
Computer vision › 3D vision › point cloud processing › radar point cloud processing
4d radar point cloud
1.012026
DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations · AAAI 2026
Robotics › Robot navigation and mapping › localization
odometry
1.012026
DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations · AAAI 2026
Robotics › Robot navigation and mapping › localization › odometry
radar odometry
1.012026
DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations · AAAI 2026
Computer vision › Video understanding and tracking
spatiotemporal aggregation
0.912025
TDFANet: Encoding Sequential 4D Radar Point Clouds Using Trajectory-Guided Deformable Feature Aggregation for Place Recognition · ICRA 2025
Robotics › Robot navigation and mapping › place recognition
visual place recognition
0.912025
TDFANet: Encoding Sequential 4D Radar Point Clouds Using Trajectory-Guided Deformable Feature Aggregation for Place Recognition · ICRA 2025
Machine learning › Generative modeling
diffusion model
0.812024
GPD-VVTO: Preserving Garment Details in Video Virtual Try-On · ACM Multimedia 2024
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.812024
GPD-VVTO: Preserving Garment Details in Video Virtual Try-On · ACM Multimedia 2024
Visual content generation and editing › virtual try-on
video virtual try-on
0.812024
GPD-VVTO: Preserving Garment Details in Video Virtual Try-On · ACM Multimedia 2024
Visual content generation and editing
virtual try-on
0.812024
GPD-VVTO: Preserving Garment Details in Video Virtual Try-On · ACM Multimedia 2024
Computer vision › Video understanding and tracking
action recognition
0.712023
Learning Discriminative Representations for Skeleton Based Action Recognition · CVPR 2023
Machine learning › Representation and self-supervised learning › representation learning › feature extraction
discriminative feature learning
0.712023
Learning Discriminative Representations for Skeleton Based Action Recognition · CVPR 2023
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition
0.712023
Learning Discriminative Representations for Skeleton Based Action Recognition · CVPR 2023
Computer vision › Video understanding and tracking › temporal modeling
temporal consistency
0.212024
GPD-VVTO: Preserving Garment Details in Video Virtual Try-On · ACM Multimedia 2024

Methods — techniques the papers use, named apart from their topics

temporal attention · 1.5latent diffusion model · 1.5gauss-newton optimization · 1.0dual-stream backbone · 1.0differentiable neural-optimization iteration · 1.0trajectory-guided alignment · 0.9deformable feature aggregation · 0.9graph convolutional network · 0.7contrastive learning · 0.7
YearPublicationVenuePosition
2026 DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations
abstract
A novel learning-optimization-combined 4D radar odometry model, named DNOI-4DRO, is proposed in this paper. The proposed model seamlessly integrates traditional geometric optimization with end-to-end neural network training, leveraging an innovative differentiable neural-optimization iteration operator. In this framework, point-wise motion flow is first estimated using a neural network, followed by the construction of a cost function based on the relationship between point motion and pose in 3D space. The radar pose is then refined using Gauss-Newton updates. Additionally, we design a dual-stream 4D radar backbone that integrates multi-scale geometric features and clustering-based class-aware features to enhance the representation of sparse 4D radar point clouds. Extensive experiments on the VoD and Snail-Radar datasets demonstrate the superior performance of our model, which outperforms recent classical and learning-based approaches. Notably, our method even achieves results comparable to A-LOAM with mapping optimization using LiDAR point clouds as input.
Shouyi Lu, Huanyu Zhou, Guirong Zhuo
AAAI2
2026 Spatiotemporal physics-guided graph neural network for wind farm power prediction
abstract
Recent efforts to enhance the interpretability of Graph Neural Networks (GNN) have focused on boosting their trustworthiness and traceability while maintaining predictive accuracy. However, mainstream approaches largely rely on data-driven strategies, with limited integration of physical prior knowledge or critical examination of GNN interpretability within physics-constrained frameworks. This paper introduces a novel Physics-guided GNN that incorporates spatiotemporal wind farm dynamics using a cutting-edge three-dimensional wake analytical model to guide the GNN architecture. This integration ensures compliance with physical laws during training, reducing the uncertainty and complexity associated with purely data-driven learning and addressing scalability challenges. By employing spatiotemporal directed local subgraphs and physics-induced attention weight learning, the model effectively considers the spatiotemporal wake coupling processes in wind farms, enabling high-accuracy power prediction. The proposed model outperforms other neural network structures in predicting both overall wind farm power production and individual turbine output. This research offers an efficient solution for power prediction in complex wind farms and demonstrates the potential of embedding domain-specific physical knowledge into GNN across broader multi-physics scenarios. It provides valuable insights into integrating physical theories with GNNs, enhancing model precision and transparency. • A novel 3D STV-PGNN enhances spatiotemporal prediction in dynamic systems via optimal features, boosting interpretability. • A spatial time-varying directed subgraph strategy addresses wind farm challenges from complex topologies and varying inflow. • Physics-driven attention edge weights and parameter learning enhance prediction accuracy via physical consistency. • Results validate the physics-GNN integration, demonstrating gains in accuracy, robustness, interpretability and scalability.
Yingning Qiu, Yanhui Feng, Huanyu Zhou, Xue-Lu Xiong
Eng. Appl. Artif. Intell.4
2025 SkeletonMix: A Mixup-Based Data Augmentation Framework for Skeleton-Based Action Recognition
abstract
Skeleton-based human action recognition has received widespread attention for its robustness to changes in the background and appearance of actors compared to the RGB modality. Data augmentation is widely used to explicitly regularize the model to prevent overfitting, especially when the number of labeled samples is scarce. However, compared to various augmentation methods available for the RGB modality, there are fewer works on the skeleton modality, especially a model-agnostic augmentation method that can be easily integrated into multiple models. We address the problem by proposing a comprehensive data augmentation framework named SkeletonMix, which contains a pair sample selection module for mixup and random augmentations tailored for skeleton modality. SkeletonMix is a non-learning framework and can be applied to different models in a plug-and-play manner. Extensive experiments on NTU RGB+D, NTU RGB+D 120 and PKU-MMD datasets demonstrate the effectiveness of our proposed framework under limited labeled training data. Our proposed method boosts the performance by a maximum of 7.5% on scarce training data setup (5% of training data).
Zongye Zhang 0002, Huanyu Zhou, Qingjie Liu 0001, Yunhong Wang 0001
ICASSP2
2025 TDFANet: Encoding Sequential 4D Radar Point Clouds Using Trajectory-Guided Deformable Feature Aggregation for Place Recognition
abstract
Place recognition is essential for achieving closedloop or global positioning in autonomous vehicles and mobile robots. Despite recent advancements in place recognition using 2D cameras or 3D LiDAR, it remains to be seen how to use 4D radar for place recognition - an increasingly popular sensor for its robustness against adverse weather and lighting conditions. Compared to LiDAR point clouds, radar data are drastically sparser, noisier and in much lower resolution, which hampers their ability to effectively represent scenes, posing significant challenges for 4D radar-based place recognition. This work addresses these challenges by leveraging multimodal information from sequential 4D radar scans and effectively extracting and aggregating spatio-temporal features. Our approach follows a principled pipeline that comprises (1) dynamic points removal and ego-velocity estimation from velocity property, (2) bird's eye view (BEV) feature encoding on the refined point cloud, (3) feature alignment using BEV feature map motion trajectory calculated by ego-velocity, (4) multiscale spatio-temporal features of the aligned BEV feature maps are extracted and aggregated. Real-world experimental results validate the feasibility of the proposed method and demonstrate its robustness in handling dynamic environments. Source codes are available.
Shouyi Lu, Guirong Zhuo, Huanyu Zhou, Renbo Huang, Minqing Huang, Lianqing Zheng, Qiang Shu
ICRA5
2024 GPD-VVTO: Preserving Garment Details in Video Virtual Try-On
abstract
Video Virtual Try-On aims to transfer a garment onto a person in the video. Previous methods typically focus on image-based virtual try-on, but directly applying these methods to videos often leads to temporal discontinuity due to inconsistencies between frames. Limited attempts in video virtual try-on also suffer from unrealistic results and poor generalization ability. In light of previous research, we posit that the task of video virtual try-on can be decomposed into two key aspects: (1) single-frame results are realistic and natural, while retaining consistency with the garment; (2) the person's actions and the garment are coherent throughout the entire video. To address these two aspects, we propose a novel two-stage framework based on Latent Diffusion Model, namely Garment-Preserving Diffusion for Video Virtual Try-On (GPD-VVTO). In the first stage, the model is trained on single-frame data to improve the ability of generating high-quality try-on images. We integrate both low-level texture features and high-level semantic features of the garment into the denoising network to preserve garment details while ensuring a natural fit between the garment and the person. In the second stage, the model is trained on video data to enhance temporal consistency. We devise a novel Garment-aware Temporal Attention (GTA) module that incorporates garment features into temporal attention, enabling the model to maintain the fidelity to the garment during temporal modeling. Furthermore, we collect a video virtual try-on dataset containing high-resolution videos from diverse scenes, addressing the limited variety of current datasets in terms of video background and human actions. Extensive experiments demonstrate that our method outperforms existing state-of-the-art methods in both image-based and video-based virtual try-on tasks, indicating the effectiveness of our proposed framework.
Weilun Dai, Long Chan, Huanyu Zhou, Aixi Zhang, Si Liu 0001
ACM Multimedia4
2023 Learning Discriminative Representations for Skeleton Based Action Recognition
abstract
Human action recognition aims at classifying the category of human action from a segment of a video. Recently, people have dived into designing GCN-based models to extract features from skeletons for performing this task, because skeleton representations are much more efficient and robust than other modalities such as RGB frames. However, when employing the skeleton data, some important clues like related items are also discarded. It results in some ambiguous actions that are hard to be distinguished and tend to be misclassified. To alleviate this problem, we propose an auxiliary feature refinement head (FR Head), which consists of spatial-temporal decoupling and contrastive feature refinement, to obtain discriminative representations of skeletons. Ambiguous samples are dynamically discovered and calibrated in the feature space. Furthermore, FR Head could be imposed on different stages of GCNs to build a multi-level refinement for stronger supervision. Extensive experiments are conducted on NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets. Our proposed models obtain competitive results from state-of-the-art methods and can help to discriminate those ambiguous samples. Codes are available at https://github.com/zhysora/FR-Head.
Huanyu Zhou, Qingjie Liu 0001, Yunhong Wang 0001
CVPR1
2022 PanFormer: A Transformer Based Model for Pan-Sharpening
abstract
Pan-sharpening aims at producing a high-resolution (HR) multi-spectral (MS) image from a low-resolution (LR) multi-spectral (MS) image and its corresponding panchromatic (PAN) image acquired by a same satellite. Inspired by a new fashion in recent deep learning community, we propose a novel Transformer based model for pan-sharpening. We explore the potential of Transformer in image feature extraction and fusion. Following the successful development of vision transformers, we design a two-stream network with the self-attention to extract the modality-specific features from the PAN and MS modalities and apply a cross-attention module to merge the spectral and spatial features. The pan-sharpened image is produced from the enhanced fused features. Extensive experiments on GaoFen-2 and WorldView-3 images demonstrate that our Transformer based model achieves impressive results and outperforms many existing CNN based methods, which shows the great potential of introducing Transformer to the pan-sharpening task. Codes are available at https://github.com/zhysora/PanFormer.
Huanyu Zhou, Qingjie Liu 0001, Yunhong Wang 0001
ICME1
2022 Unsupervised Cycle-Consistent Generative Adversarial Networks for Pan Sharpening
abstract
Deep learning-based pan sharpening has received significant research interest in recent years. Most of the existing methods fall into the supervised learning framework in which they downsample the multispectral (MS) and panchromatic (PAN) images and regard the original MS images as ground truths to form training samples based on Wald’s protocol. Although impressive performance could be achieved, they have difficulties when generalizing to the original full-scale images due to the scale gap, which makes them lack of practicability. In this article, we propose an unsupervised generative adversarial framework that learns from the full-scale images without the ground truths to alleviate this problem. We first extract the modality-specific features from the PAN and MS images with a two-stream generator, perform fusion in the feature domain, and then reconstruct the pan-sharpened images. Furthermore, we introduce a novel hybrid loss based on the cycle-consistency and adversarial scheme to improve the performance. Comparison experiments with the state-of-the-art methods are conducted on GaoFen-2 (GF-2) and WorldView-3 satellites. Results demonstrate that the proposed method can greatly improve the pan-sharpening performance on the full-scale images, which clearly shows its practical value. Codes are available athttps://github.com/zhysora/UCGAN.
Huanyu Zhou, Qingjie Liu 0001, Dawei Weng, Yunhong Wang 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 PSGAN: A Generative Adversarial Network for Remote Sensing Image Pan-Sharpening
abstract
This article addresses the problem of remote sensing image pan-sharpening from the perspective of generative adversarial learning. We propose a novel deep neural network-based method named pansharpening GAN (PSGAN). To the best of our knowledge, this is one of the first attempts at producing high-quality pan-sharpened images with generative adversarial networks (GANs). The PSGAN consists of two components: a generative network (i.e., generator) and a discriminative network (i.e., discriminator). The generator is designed to accept panchromatic (PAN) and multispectral (MS) images as inputs and maps them to the desired high-resolution (HR) MS images, and the discriminator implements the adversarial training strategy for generating higher fidelity pan-sharpened images. In this article, we evaluate several architectures and designs, namely, two-stream input, stacking input, batch normalization layer, and attention mechanism to find the optimal solution for pan-sharpening. Extensive experiments on QuickBird, GaoFen-2, and WorldView-2 satellite images demonstrate that the proposed PSGANs not only are effective in generating high-quality HR MS images and superior to state-of-the-art methods but also generalize well to full-scale images.
Qingjie Liu 0001, Huanyu Zhou, Qizhi Xu, Yunhong Wang 0001
IEEE Trans. Geosci. Remote. Sens.2
2020 Pan-Sharpening with a CNN-Based Two Stage Ratio Enhancement Method
abstract
We propose a hybrid method combining the deep learning technique and the ratio enhancement (RE) method for pansharpening. The intuition behind is to utilize the deep learning technique to synthesize a panchromatic (PAN) image for the RE method to reduce the spectral distortion while keeping the spatial details. The method consists of two stages. First, the CNN synthesizer is optimized to generate the downsampled PAN image to guarantee the network have a good initialization. Second, CNN is integrated into the RE method and supervised by the ground truth multi-spectral (MS) to produce an ideal synthesized PAN for the RE method. We conduct experiments on various datasets and compare with widely used methods to demonstrate the superiority of the proposed method.
Huanyu Zhou, Qingjie Liu 0001, Qizhi Xu, Yunhong Wang 0001
IGARSS1
2020 Prototyping federated learning on edge computing systems
Jianlei Yang 0001, Yixiao Duan, Huanyu Zhou, Jingyuan Wang 0001, Weisheng Zhao 0001
Frontiers Comput. Sci.4
2015 Activation force-based air pollution observation station clustering
Di Huang 0006, Hong Yu 0006, Huanyu Zhou, Zhanyu Ma, Weisong Hu, Jun Guo 0002
QSHINE4