Zhao-Yang Wang

dblp:190/9251 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
7since 2021 · last 2026
0009-0008-0945-702XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSystems, architecture and hardware · 1Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Mesh-Gait: A Unified Framework for Gait Recognition Through Multi-Modal Representation Learning from 2D Silhouettes
Zhao-Yang Wang, Jieneng Chen, Yuxiang Guo 0001, Jiang Liu 0014, Rama Chellappa
FG1
2026 DiffProtect: Generative adversarial examples using diffusion models for facial privacy protection
abstract
The increasingly pervasive facial recognition (FR) systems raise serious concerns about personal privacy, especially for billions of users who have publicly shared their photos on social media. To address this challenge, several adversarial attack methods have been proposed to protect individuals from being identified by unauthorized FR systems with perturbed facial images. However, these approaches suffer from poor visual quality or low attack success rates, which limit their practical utility. Recently, diffusion models have achieved tremendous success in image generation. In this work, we ask: can diffusion models be used to generate adversarial examples against FR systems to improve both visual quality and attack performance? We propose DiffProtect, a novel method leveraging a diffusion autoencoder to generate semantically meaningful perturbations on FR systems. Extensive experiments demonstrate that DiffProtect produces more natural-looking encrypted images than state-of-the-art methods while achieving significantly higher attack success rates, e.g. , 24.5 % and 25.1 % absolute improvements on the CelebA-HQ and FFHQ datasets. We further evaluate the effectiveness of DiffProtect in the real world using a commercial FR API and validate its usefulness in practice through a user study. Our code is available at https://github.com/joellliu/DiffProtect .
Jiang Liu 0014, Chun Pong Lau 0001, Zhongliang Guo 0001, Yuxiang Guo 0001, Zhao-Yang Wang, Rama Chellappa
Pattern Recognit.5
2025 UniGait: A Unified Transformer-based Multitask Framework for Gait Analysis in the Wild
abstract
Gait recognition is a rapidly emerging and significant area of biometrics, leveraging the unique walking patterns of individuals to perform personal identification and facilitate healthcare monitoring, such as elderly care, fall detection, etc. While existing gait recognition methods perform well in indoor, or short-range environments, their effectiveness diminishes significantly when applied to unconstrained outdoor scenarios. Challenges such as environmental turbulence, occlusion, varying viewing angles contribute to this performance drop. To address these challenges and enhance gait recognition accuracy in real-world settings, while also expanding the functionality of gait features for healthcare applications, we propose a unified multitask framework called UniGait. UniGait is designed to perform a comprehensive range of gait analysis tasks, including gait recognition and estimation of gait-related human attributes. UniGait is built upon a transformer-based architecture, which leverages the power of a cross-attention mechanism to simultaneously process multiple sub-tasks. This multitask learning approach allows the model to extract more robust gait features by jointly learning gait recognition and human attribute estimation, leading to improved overall performance. We report the results of extensive experiments and analysis on large-scale, real-world datasets collected under challenging conditions, including long-range (up to 1000 meters) and high-pitch angles (including UAV-based data). The results demonstrated state-of-the-art performance, highlighting the potential of UniGait for deployment in real-world applications, making it a valuable tool for a range of biometric and healthcare monitoring scenarios.
Zhao-Yang Wang, Jiang Liu 0014, Yuxiang Guo 0001, Jieneng Chen, Rama Chellappa
FG1
2025 Medical World Model
Zhao-Yang Wang, Qiuping Liu, Shuwen Sun, Kang Wang 0016, Rama Chellappa, Zongwei Zhou, Alan L. Yuille, Lei Zhu 0003, Jieneng Chen
ICCV2
2025 VM-Gait: Multi-Modal 3D Representation Based on Virtual Marker for Gait Recognition
abstract
Gait recognition plays a vital role in biometric applications by analyzing the unique characteristics of an individ-ual's walking pattern. Methods based on 2D representations, such as silhouettes and skeletons, are increasingly being developed to learn the shape features and joint dy-namic movements. Nevertheless, the effectiveness of 2D representation-based methods is impeded by factors such as changes in viewpoint, partial occlusion, and noisy en-vironments. 3D representation-based methods can complement 2D representation-based approaches by providing more precise dynamic body shapes and motion information, along with increased robustness against changes in viewpoint and partial occlusion. However, the complex-ity of acquiring accurate 3D representations and the chal-lenges associated with extracting dynamic topological features from sequences of 3D representations hinder the de-velopment of 3D representations-based methods. In this pa-per, we present VM-Gait, a novel multi-modal gait recognition framework that harnesses the advantages of integrating both 2D and 3D representations. Furthermore, we in-troduce a new 3D representation, Virtual Marker, into gait recognition to efficiently learn topological features from 3D representations, avoiding the computational complexi-ties inherent in directly learning from 3D representations like 3D meshes or 3D point clouds. Extensive experiments demonstrate that the proposed framework effectively learns and fuses discriminative information from different gait modalities, enhancing gait recognition performance.
Zhao-Yang Wang, Jiang Liu 0014, Jieneng Chen, Rama Chellappa
WACV1
2024 HyperGait: A Video-based Multitask Network for Gait Recognition and Human Attribute Estimation at Range and Altitude
abstract
Gait recognition is one of the mainstream approaches for identifying individuals when face information is not available. Most previous methods achieve good performance on structured indoor walking sequences with silhouettes provided. However, when these methods are applied to unconstrained outdoor sequences, a significant reduction in performance is inevitably observed due to factors such as turbulence, occlusion, view angle, and oversized clothing. To make gait recognition methods stable and effective for real-world settings, we extend gait-only-based approaches by introducing more useful biometric information such as gender, age, height, weight, and body mass index to cooperatively work with the gait recognition module. In this paper, we propose a video-based multitasking network for gait recognition and human attribute prediction at ranges of up to 1000 meters and high-pitch angles to mutually improve the robustness and accuracy of each task. Through a series of experiments on OU-MVLP and BRIAR datasets, we show that our multitasking network outperforms previous methods and provides more useful biometric information for human identification tasks.
Zhao-Yang Wang, Jiang Liu 0014, Ram Prabhakar Kathirvel, Chun Pong Lau 0001, Rama Chellappa
IJCB1
2024 Integrated optimization of truck appointment quotas and container relocations for multiple blocks considering transshipment containers handling
Shuang Duan, Hong-Xing Zheng, Zhao-Yang Wang
Adv. Eng. Informatics3
2020 Z-Net: an Anisotropic 3D DCNN for Medical CT Volume Segmentation
abstract
Accurate volume segmentation from the Computed Tomography (CT) scan is a common prerequisite for pre-operative planning, intra-operative guidance and quantitative assessment of therapeutic outcomes in robot-assisted Minimally Invasive Surgery (MIS). 3D Deep Convolutional Neural Network (DCNN) is a viable solution for this task, but is memory intensive. Small isotropic patches are cropped from the original and large CT volume to mitigate this issue in practice, but it may cause discontinuities between the adjacent patches and severe class-imbalances within individual sub-volumes. This paper presents a new 3D DCNN framework, namely Z-Net, to tackle the discontinuity and class-imbalance issue by preserving a full field-of-view of the objects in the XY planes using anisotropic spatial separable convolutions. The proposed Z-Net can be seamlessly integrated into existing 3D DCNNs with isotropic convolutions such as 3D U-Net and V-Net, with improved volume segmentation Intersection over Union (IoU) - up to 12.6%. Detailed validation of Z-Net is provided for CT aortic, liver and lung segmentation, demonstrating the effectiveness and practical value of Z-Net for intra-operative 3D navigation in robot-assisted MIS.
Peichao Li, Xiaoyun Zhou 0001, Zhao-Yang Wang, Guang-Zhong Yang
IROS3
2020 Instantiation-Net: 3D Mesh Reconstruction from Single 2D Image for Right Ventricle
Zhao-Yang Wang, Xiaoyun Zhou 0001, Peichao Li, Celia V. Riga, Guang-Zhong Yang
MICCAI (4)1
2020 Detecting Anomalous Bus-Driving Behaviors from Trajectories
Zhao-Yang Wang, Beihong Jin, Tingjian Ge, Taofeng Xue
J. Comput. Sci. Technol.1
2019 One-Stage Shape Instantiation from a Single 2D Image to 3D Point Cloud
Xiaoyun Zhou 0001, Zhao-Yang Wang, Peichao Li, Jian-Qing Zheng, Guang-Zhong Yang
MICCAI (4)2