VLDB 2026 Research / reviewers in the wild / expert
Yuxiang Guo 0001
dblp:215/9733-1
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0003-9325-5220ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mesh-Gait: A Unified Framework for Gait Recognition Through Multi-Modal Representation Learning from 2D Silhouettes
Zhao-Yang Wang, Jieneng Chen, Yuxiang Guo 0001, Jiang Liu 0014, Rama Chellappa |
FG | 3 |
| 2026 | DiffProtect: Generative adversarial examples using diffusion models for facial privacy protectionabstractThe increasingly pervasive facial recognition (FR) systems raise serious concerns about personal privacy, especially for billions of users who have publicly shared their photos on social media. To address this challenge, several adversarial attack methods have been proposed to protect individuals from being identified by unauthorized FR systems with perturbed facial images. However, these approaches suffer from poor visual quality or low attack success rates, which limit their practical utility. Recently, diffusion models have achieved tremendous success in image generation. In this work, we ask: can diffusion models be used to generate adversarial examples against FR systems to improve both visual quality and attack performance? We propose DiffProtect, a novel method leveraging a diffusion autoencoder to generate semantically meaningful perturbations on FR systems. Extensive experiments demonstrate that DiffProtect produces more natural-looking encrypted images than state-of-the-art methods while achieving significantly higher attack success rates, e.g. , 24.5 % and 25.1 % absolute improvements on the CelebA-HQ and FFHQ datasets. We further evaluate the effectiveness of DiffProtect in the real world using a commercial FR API and validate its usefulness in practice through a user study. Our code is available at https://github.com/joellliu/DiffProtect . Jiang Liu 0014, Chun Pong Lau 0001, Zhongliang Guo 0001, Yuxiang Guo 0001, Zhao-Yang Wang, Rama Chellappa |
Pattern Recognit. | 4 |
| 2025 | SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D ReconstructionabstractRecent efforts in Gaussian-Splat-based Novel View Synthesis can achieve photorealistic rendering; however, such capability is limited in sparse-view scenarios due to sparse initialization and over-fitting floaters. Recent progress in depth estimation and alignment can provide dense point cloud using few views; however, the resulting pose accuracy is suboptimal. In this work, we present SPARS3R, which combines the advantages of accurate pose estimation from Structure-from-Motion and dense point cloud from depth estimation. To this end, SPARS3R first performs a Global Fusion Alignment process that maps a prior dense point cloud to a sparse point cloud from Structure-from-Motion based on triangulated correspondences. RANSAC is applied during this process to distinguish inliers and outliers. SPARS3R then performs a second, Semantic Outlier Alignment step, which extracts semantically coherent regions around the outliers and performs local alignment in these regions. Along with several improvements in the evaluation process, we demonstrate that SPARS3R can achieve photorealistic rendering with sparse images and significantly outperforms existing approaches. Yutao Tang, Yuxiang Guo 0001, Deming Li, Cheng Peng 0008 |
CVPR | 2 |
| 2025 | UniGait: A Unified Transformer-based Multitask Framework for Gait Analysis in the WildabstractGait recognition is a rapidly emerging and significant area of biometrics, leveraging the unique walking patterns of individuals to perform personal identification and facilitate healthcare monitoring, such as elderly care, fall detection, etc. While existing gait recognition methods perform well in indoor, or short-range environments, their effectiveness diminishes significantly when applied to unconstrained outdoor scenarios. Challenges such as environmental turbulence, occlusion, varying viewing angles contribute to this performance drop. To address these challenges and enhance gait recognition accuracy in real-world settings, while also expanding the functionality of gait features for healthcare applications, we propose a unified multitask framework called UniGait. UniGait is designed to perform a comprehensive range of gait analysis tasks, including gait recognition and estimation of gait-related human attributes. UniGait is built upon a transformer-based architecture, which leverages the power of a cross-attention mechanism to simultaneously process multiple sub-tasks. This multitask learning approach allows the model to extract more robust gait features by jointly learning gait recognition and human attribute estimation, leading to improved overall performance. We report the results of extensive experiments and analysis on large-scale, real-world datasets collected under challenging conditions, including long-range (up to 1000 meters) and high-pitch angles (including UAV-based data). The results demonstrated state-of-the-art performance, highlighting the potential of UniGait for deployment in real-world applications, making it a valuable tool for a range of biometric and healthcare monitoring scenarios. Zhao-Yang Wang, Jiang Liu 0014, Yuxiang Guo 0001, Jieneng Chen, Rama Chellappa |
FG | 3 |
| 2025 | GaitContour: Efficient Gait Recognition Based on a Contour-Pose RepresentationabstractGait recognition holds the promise to robustly identify subjects based on walking patterns instead of appearance information. In recent years, this field has been dominated by learning methods based on two input formats: silhouette images and sparse keypoints. Compared to image-based approaches, keypoint-based methods can achieve significantly higher efficiency due to their sparsity. However, sparsity also results in information loss, thereby reducing performance. In this work, we propose a novel, keypoint-based Contour-Pose representation, which compactly encodes both body shape and parts information. We further propose a local-to-global architecture, called GaitContour, to leverage this novel representation and efficiently compute subject embedding in two stages. The first stage consists of a local transformer that extracts features from five different body regions. The second stage then aggregates the regional features to estimate a global human gait representation. Such a design significantly reduces the complexity of the attention operation and improves both efficiency and performance. Through large scale experiments, Gait-Contour is shown to perform significantly better than previous keypoint-based methods. Furthermore, the ContourPose representation also achieves new SoTA performances on fusion-based gait recognition methods. Yuxiang Guo 0001, Anshul Shah 0001, Jiang Liu 0014, Ayush Gupta 0001, Rama Chellappa, Cheng Peng 0008 |
WACV | 1 |
| 2025 | VILLS: Video-Image Learning to Learn Semantics for Person Re-IdentificationabstractPerson Re-identification is a research area with significant real world applications. Despite recent progress, existing methods face challenges in robust re-identification in the wild, e.g., by focusing only on a particular modality and on unreliable patterns such as clothing. A generalized method is highly desired, but remains elusive to achieve due to issues such as the trade-off between spatial and temporal resolution and inaccurate feature extraction. We propose VILLS (Video-Image Learning to Learn Semantics), a self-supervised method that jointly learns spatial and temporal features from images and videos. VILLS first designs a local semantic extraction module that adaptively extracts semantically consistent and robust spatial features. Then, VILLS designs a unified feature learning and adaptation module to represent image and video modalities in a consistent feature space. By Leveraging self-supervised, large-scale pre-training, VILLS establishes a new State-of-The-Art that significantly outperforms existing image and video-based methods. Siyuan Huang 0005, Ram Prabhakar, Yuxiang Guo 0001, Rama Chellappa, Cheng Peng 0008 |
WACV | 3 |
| 2025 | StimuVAR: Spatiotemporal Stimuli-Aware Video Affective Reasoning with Multimodal Large Language Models
Yuxiang Guo 0001, Yang Zhao 0024, Rama Chellappa, Shao-Yuan Lo |
Int. J. Comput. Vis. | 1 |
| 2024 | Distillation-guided Representation Learning for Unconstrained Gait RecognitionabstractGait recognition holds the promise of robustly identifying subjects based on walking patterns instead of appearance information. While previous approaches have performed well for curated indoor data, they tend to underperform in unconstrained situations, e.g. in outdoor, long distance scenes, etc. We propose a framework, termed GAit DEtection and Recognition (GADER), for human authentication in challenging outdoor scenarios. Specifically, GADER leverages a Double Helical Signature to detect segments that contain human movement and builds discriminative features through a novel gait recognition method, where only frames containing gait information are used. To further enhance robustness, GADER encodes viewpoint information in its architecture, and distills representation from an auxiliary RGB recognition model, which enables GADER to learn from silhouette and RGB data at training time. At test time, GADER only infers from the silhouette modality. We evaluate our method on multiple State-of-The-Arts(SoTA) gait baselines and demonstrate consistent improvements on indoor and outdoor datasets, especially with a significant 25.2% improvement on unconstrained, remote gait data. Yuxiang Guo 0001, Siyuan Huang 0005, Ram Prabhakar, Chun Pong Lau 0001, Rama Chellappa, Cheng Peng 0008 |
IJCB | 1 |
| 2023 | Multi-Modal Human Authentication Using Silhouettes, Gait and RGBabstractWhole-body-based human authentication is a promising approach for remote biometrics scenarios. Current literature focuses on either body recognition based on RGB images or gait recognition based on body shapes and walking patterns; both have their advantages and drawbacks. In this work, we propose Dual-Modal Ensemble (DME), which combines both RGB and silhouette data to achieve more robust performances for indoor and outdoor whole-body based recognition. Within DME, we propose GaitPattern, which is inspired by the double helical gait pattern used in traditional gait analysis. The GaitPattern contributes to robust identification performance over a large range of viewing angles. Extensive experimental results on the CASIA-B dataset demonstrate that the proposed method outperforms state-of-the-art recognition systems. We also provide experimental results using the newly collected BRIAR dataset. Yuxiang Guo 0001, Cheng Peng 0008, Chun Pong Lau 0001, Rama Chellappa |
FG | 1 |
| 2021 | Towards Predicting Vehicular Data ConsumptionabstractCombining in-car multiple sensors measuring parameters that can be used to improve both safety and efficiency with a plethora of external data sources (e.g., traffic conditions, weather) which, if properly used, can significantly improve the overall trip experience. One source that can help the navigation and provide "context awareness", especially for autonomous driving, are the High Definition (HD) maps, which have recently witnessed a tremendous growth of popularity in vehicular technology and use. As they are limited to a particular geographic area with respect to a given point along a trip, different portions need to be downloaded (and processed) on multiple occasions throughout a given trip, along with the other data from internal and external sources. We take a first step towards formalizing the problem of Predicting Map Data Consumption (PMDC) in the future time instants for a given trip, based on a (time) window from its history, and investigate the use of Long Short-Term Memory (LSTM) networks - a special type of Recurrent Neural Networks (RNN). Significant efforts were focused on generating an appropriate dataset for this study, towards which we fused the information available in multiple heterogeneous data sources. We conducted experimental observations demonstrating the benefits of the proposed approach. Andi Zang, Xiaofeng Zhu 0004, Yuxiang Guo 0001, Fan Zhou 0002, Goce Trajcevski |
MDM | 3 |