EDBT 2026 Demo / reviewers in the wild / expert
Huanyu Zhou
dblp:44/8489
· DBLP profile ↗
12ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Robot navigation and mapping · 26% 3D vision · 26% Video understanding and tracking · 22% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
point cloud processing |
1.9 | 2 | 2026 | DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations · AAAI 2026 TDFANet: Encoding Sequential 4D Radar Point Clouds Using Trajectory-Guided Deformable Feature Aggregation for Place Recognition · ICRA 2025 |
Computer vision › 3D vision › point cloud processing › radar point cloud processing
4d radar point cloud |
1.0 | 1 | 2026 | DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations · AAAI 2026 |
Robotics › Robot navigation and mapping › localization
odometry |
1.0 | 1 | 2026 | DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations · AAAI 2026 |
Robotics › Robot navigation and mapping › localization › odometry
radar odometry |
1.0 | 1 | 2026 | DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization Iterations · AAAI 2026 |
Computer vision › Video understanding and tracking
spatiotemporal aggregation |
0.9 | 1 | 2025 | TDFANet: Encoding Sequential 4D Radar Point Clouds Using Trajectory-Guided Deformable Feature Aggregation for Place Recognition · ICRA 2025 |
Robotics › Robot navigation and mapping › place recognition
visual place recognition |
0.9 | 1 | 2025 | TDFANet: Encoding Sequential 4D Radar Point Clouds Using Trajectory-Guided Deformable Feature Aggregation for Place Recognition · ICRA 2025 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | GPD-VVTO: Preserving Garment Details in Video Virtual Try-On · ACM Multimedia 2024 |
Machine learning › Generative modeling › diffusion model
latent diffusion model |
0.8 | 1 | 2024 | GPD-VVTO: Preserving Garment Details in Video Virtual Try-On · ACM Multimedia 2024 |
Visual content generation and editing › virtual try-on
video virtual try-on |
0.8 | 1 | 2024 | GPD-VVTO: Preserving Garment Details in Video Virtual Try-On · ACM Multimedia 2024 |
Visual content generation and editing
virtual try-on |
0.8 | 1 | 2024 | GPD-VVTO: Preserving Garment Details in Video Virtual Try-On · ACM Multimedia 2024 |
Computer vision › Video understanding and tracking
action recognition |
0.7 | 1 | 2023 | Learning Discriminative Representations for Skeleton Based Action Recognition · CVPR 2023 |
Machine learning › Representation and self-supervised learning › representation learning › feature extraction
discriminative feature learning |
0.7 | 1 | 2023 | Learning Discriminative Representations for Skeleton Based Action Recognition · CVPR 2023 |
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition |
0.7 | 1 | 2023 | Learning Discriminative Representations for Skeleton Based Action Recognition · CVPR 2023 |
Computer vision › Video understanding and tracking › temporal modeling
temporal consistency |
0.2 | 1 | 2024 | GPD-VVTO: Preserving Garment Details in Video Virtual Try-On · ACM Multimedia 2024 |
Methods — techniques the papers use, named apart from their topics
temporal attention · 1.5latent diffusion model · 1.5gauss-newton optimization · 1.0dual-stream backbone · 1.0differentiable neural-optimization iteration · 1.0trajectory-guided alignment · 0.9deformable feature aggregation · 0.9graph convolutional network · 0.7contrastive learning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization IterationsabstractA novel learning-optimization-combined 4D radar odometry model, named DNOI-4DRO, is proposed in this paper. The proposed model seamlessly integrates traditional geometric optimization with end-to-end neural network training, leveraging an innovative differentiable neural-optimization iteration operator. In this framework, point-wise motion flow is first estimated using a neural network, followed by the construction of a cost function based on the relationship between point motion and pose in 3D space. The radar pose is then refined using Gauss-Newton updates. Additionally, we design a dual-stream 4D radar backbone that integrates multi-scale geometric features and clustering-based class-aware features to enhance the representation of sparse 4D radar point clouds. Extensive experiments on the VoD and Snail-Radar datasets demonstrate the superior performance of our model, which outperforms recent classical and learning-based approaches. Notably, our method even achieves results comparable to A-LOAM with mapping optimization using LiDAR point clouds as input. Shouyi Lu, Huanyu Zhou, Guirong Zhuo |
AAAI | 2 |
| 2026 | Spatiotemporal physics-guided graph neural network for wind farm power predictionabstractRecent efforts to enhance the interpretability of Graph Neural Networks (GNN) have focused on boosting their trustworthiness and traceability while maintaining predictive accuracy. However, mainstream approaches largely rely on data-driven strategies, with limited integration of physical prior knowledge or critical examination of GNN interpretability within physics-constrained frameworks. This paper introduces a novel Physics-guided GNN that incorporates spatiotemporal wind farm dynamics using a cutting-edge three-dimensional wake analytical model to guide the GNN architecture. This integration ensures compliance with physical laws during training, reducing the uncertainty and complexity associated with purely data-driven learning and addressing scalability challenges. By employing spatiotemporal directed local subgraphs and physics-induced attention weight learning, the model effectively considers the spatiotemporal wake coupling processes in wind farms, enabling high-accuracy power prediction. The proposed model outperforms other neural network structures in predicting both overall wind farm power production and individual turbine output. This research offers an efficient solution for power prediction in complex wind farms and demonstrates the potential of embedding domain-specific physical knowledge into GNN across broader multi-physics scenarios. It provides valuable insights into integrating physical theories with GNNs, enhancing model precision and transparency. • A novel 3D STV-PGNN enhances spatiotemporal prediction in dynamic systems via optimal features, boosting interpretability. • A spatial time-varying directed subgraph strategy addresses wind farm challenges from complex topologies and varying inflow. • Physics-driven attention edge weights and parameter learning enhance prediction accuracy via physical consistency. • Results validate the physics-GNN integration, demonstrating gains in accuracy, robustness, interpretability and scalability. Yingning Qiu, Yanhui Feng, Huanyu Zhou, Xue-Lu Xiong |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | SkeletonMix: A Mixup-Based Data Augmentation Framework for Skeleton-Based Action RecognitionabstractSkeleton-based human action recognition has received widespread attention for its robustness to changes in the background and appearance of actors compared to the RGB modality. Data augmentation is widely used to explicitly regularize the model to prevent overfitting, especially when the number of labeled samples is scarce. However, compared to various augmentation methods available for the RGB modality, there are fewer works on the skeleton modality, especially a model-agnostic augmentation method that can be easily integrated into multiple models. We address the problem by proposing a comprehensive data augmentation framework named SkeletonMix, which contains a pair sample selection module for mixup and random augmentations tailored for skeleton modality. SkeletonMix is a non-learning framework and can be applied to different models in a plug-and-play manner. Extensive experiments on NTU RGB+D, NTU RGB+D 120 and PKU-MMD datasets demonstrate the effectiveness of our proposed framework under limited labeled training data. Our proposed method boosts the performance by a maximum of 7.5% on scarce training data setup (5% of training data). Zongye Zhang 0002, Huanyu Zhou, Qingjie Liu 0001, Yunhong Wang 0001 |
ICASSP | 2 |
| 2025 | TDFANet: Encoding Sequential 4D Radar Point Clouds Using Trajectory-Guided Deformable Feature Aggregation for Place RecognitionabstractPlace recognition is essential for achieving closedloop or global positioning in autonomous vehicles and mobile robots. Despite recent advancements in place recognition using 2D cameras or 3D LiDAR, it remains to be seen how to use 4D radar for place recognition - an increasingly popular sensor for its robustness against adverse weather and lighting conditions. Compared to LiDAR point clouds, radar data are drastically sparser, noisier and in much lower resolution, which hampers their ability to effectively represent scenes, posing significant challenges for 4D radar-based place recognition. This work addresses these challenges by leveraging multimodal information from sequential 4D radar scans and effectively extracting and aggregating spatio-temporal features. Our approach follows a principled pipeline that comprises (1) dynamic points removal and ego-velocity estimation from velocity property, (2) bird's eye view (BEV) feature encoding on the refined point cloud, (3) feature alignment using BEV feature map motion trajectory calculated by ego-velocity, (4) multiscale spatio-temporal features of the aligned BEV feature maps are extracted and aggregated. Real-world experimental results validate the feasibility of the proposed method and demonstrate its robustness in handling dynamic environments. Source codes are available. Shouyi Lu, Guirong Zhuo, Huanyu Zhou, Renbo Huang, Minqing Huang, Lianqing Zheng, Qiang Shu |
ICRA | 5 |
| 2024 | GPD-VVTO: Preserving Garment Details in Video Virtual Try-OnabstractVideo Virtual Try-On aims to transfer a garment onto a person in the video. Previous methods typically focus on image-based virtual try-on, but directly applying these methods to videos often leads to temporal discontinuity due to inconsistencies between frames. Limited attempts in video virtual try-on also suffer from unrealistic results and poor generalization ability. In light of previous research, we posit that the task of video virtual try-on can be decomposed into two key aspects: (1) single-frame results are realistic and natural, while retaining consistency with the garment; (2) the person's actions and the garment are coherent throughout the entire video. To address these two aspects, we propose a novel two-stage framework based on Latent Diffusion Model, namely Garment-Preserving Diffusion for Video Virtual Try-On (GPD-VVTO). In the first stage, the model is trained on single-frame data to improve the ability of generating high-quality try-on images. We integrate both low-level texture features and high-level semantic features of the garment into the denoising network to preserve garment details while ensuring a natural fit between the garment and the person. In the second stage, the model is trained on video data to enhance temporal consistency. We devise a novel Garment-aware Temporal Attention (GTA) module that incorporates garment features into temporal attention, enabling the model to maintain the fidelity to the garment during temporal modeling. Furthermore, we collect a video virtual try-on dataset containing high-resolution videos from diverse scenes, addressing the limited variety of current datasets in terms of video background and human actions. Extensive experiments demonstrate that our method outperforms existing state-of-the-art methods in both image-based and video-based virtual try-on tasks, indicating the effectiveness of our proposed framework. Weilun Dai, Long Chan, Huanyu Zhou, Aixi Zhang, Si Liu 0001 |
ACM Multimedia | 4 |
| 2023 | Learning Discriminative Representations for Skeleton Based Action RecognitionabstractHuman action recognition aims at classifying the category of human action from a segment of a video. Recently, people have dived into designing GCN-based models to extract features from skeletons for performing this task, because skeleton representations are much more efficient and robust than other modalities such as RGB frames. However, when employing the skeleton data, some important clues like related items are also discarded. It results in some ambiguous actions that are hard to be distinguished and tend to be misclassified. To alleviate this problem, we propose an auxiliary feature refinement head (FR Head), which consists of spatial-temporal decoupling and contrastive feature refinement, to obtain discriminative representations of skeletons. Ambiguous samples are dynamically discovered and calibrated in the feature space. Furthermore, FR Head could be imposed on different stages of GCNs to build a multi-level refinement for stronger supervision. Extensive experiments are conducted on NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets. Our proposed models obtain competitive results from state-of-the-art methods and can help to discriminate those ambiguous samples. Codes are available at https://github.com/zhysora/FR-Head. Huanyu Zhou, Qingjie Liu 0001, Yunhong Wang 0001 |
CVPR | 1 |
| 2022 | PanFormer: A Transformer Based Model for Pan-SharpeningabstractPan-sharpening aims at producing a high-resolution (HR) multi-spectral (MS) image from a low-resolution (LR) multi-spectral (MS) image and its corresponding panchromatic (PAN) image acquired by a same satellite. Inspired by a new fashion in recent deep learning community, we propose a novel Transformer based model for pan-sharpening. We explore the potential of Transformer in image feature extraction and fusion. Following the successful development of vision transformers, we design a two-stream network with the self-attention to extract the modality-specific features from the PAN and MS modalities and apply a cross-attention module to merge the spectral and spatial features. The pan-sharpened image is produced from the enhanced fused features. Extensive experiments on GaoFen-2 and WorldView-3 images demonstrate that our Transformer based model achieves impressive results and outperforms many existing CNN based methods, which shows the great potential of introducing Transformer to the pan-sharpening task. Codes are available at https://github.com/zhysora/PanFormer. Huanyu Zhou, Qingjie Liu 0001, Yunhong Wang 0001 |
ICME | 1 |
| 2022 | Unsupervised Cycle-Consistent Generative Adversarial Networks for Pan SharpeningabstractDeep learning-based pan sharpening has received significant research interest in recent years. Most of the existing methods fall into the supervised learning framework in which they downsample the multispectral (MS) and panchromatic (PAN) images and regard the original MS images as ground truths to form training samples based on Wald’s protocol. Although impressive performance could be achieved, they have difficulties when generalizing to the original full-scale images due to the scale gap, which makes them lack of practicability. In this article, we propose an unsupervised generative adversarial framework that learns from the full-scale images without the ground truths to alleviate this problem. We first extract the modality-specific features from the PAN and MS images with a two-stream generator, perform fusion in the feature domain, and then reconstruct the pan-sharpened images. Furthermore, we introduce a novel hybrid loss based on the cycle-consistency and adversarial scheme to improve the performance. Comparison experiments with the state-of-the-art methods are conducted on GaoFen-2 (GF-2) and WorldView-3 satellites. Results demonstrate that the proposed method can greatly improve the pan-sharpening performance on the full-scale images, which clearly shows its practical value. Codes are available athttps://github.com/zhysora/UCGAN. Huanyu Zhou, Qingjie Liu 0001, Dawei Weng, Yunhong Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | PSGAN: A Generative Adversarial Network for Remote Sensing Image Pan-SharpeningabstractThis article addresses the problem of remote sensing image pan-sharpening from the perspective of generative adversarial learning. We propose a novel deep neural network-based method named pansharpening GAN (PSGAN). To the best of our knowledge, this is one of the first attempts at producing high-quality pan-sharpened images with generative adversarial networks (GANs). The PSGAN consists of two components: a generative network (i.e., generator) and a discriminative network (i.e., discriminator). The generator is designed to accept panchromatic (PAN) and multispectral (MS) images as inputs and maps them to the desired high-resolution (HR) MS images, and the discriminator implements the adversarial training strategy for generating higher fidelity pan-sharpened images. In this article, we evaluate several architectures and designs, namely, two-stream input, stacking input, batch normalization layer, and attention mechanism to find the optimal solution for pan-sharpening. Extensive experiments on QuickBird, GaoFen-2, and WorldView-2 satellite images demonstrate that the proposed PSGANs not only are effective in generating high-quality HR MS images and superior to state-of-the-art methods but also generalize well to full-scale images. Qingjie Liu 0001, Huanyu Zhou, Qizhi Xu, Yunhong Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Pan-Sharpening with a CNN-Based Two Stage Ratio Enhancement MethodabstractWe propose a hybrid method combining the deep learning technique and the ratio enhancement (RE) method for pansharpening. The intuition behind is to utilize the deep learning technique to synthesize a panchromatic (PAN) image for the RE method to reduce the spectral distortion while keeping the spatial details. The method consists of two stages. First, the CNN synthesizer is optimized to generate the downsampled PAN image to guarantee the network have a good initialization. Second, CNN is integrated into the RE method and supervised by the ground truth multi-spectral (MS) to produce an ideal synthesized PAN for the RE method. We conduct experiments on various datasets and compare with widely used methods to demonstrate the superiority of the proposed method. Huanyu Zhou, Qingjie Liu 0001, Qizhi Xu, Yunhong Wang 0001 |
IGARSS | 1 |
| 2020 | Prototyping federated learning on edge computing systems
Jianlei Yang 0001, Yixiao Duan, Huanyu Zhou, Jingyuan Wang 0001, Weisheng Zhao 0001 |
Frontiers Comput. Sci. | 4 |
| 2015 | Activation force-based air pollution observation station clustering
Di Huang 0006, Hong Yu 0006, Huanyu Zhou, Zhanyu Ma, Weisong Hu, Jun Guo 0002 |
QSHINE | 4 |