VLDB 2026 Research / reviewers in the wild / expert
Xiang Li 0028
dblp:40/1491-28
· DBLP profile ↗
19ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0002-8044-7050ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 5 since 2021Security and privacy · 6 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Robust and Efficient Continuous Face Authentication via Disentangled Representation Learning and Adaptive Identity CompletionabstractIn this paper, we focus on the problem of continuous face authentication, which aims to verify a user’s identity persistently over time from streaming or sequential facial inputs. Unlike conventional face authentication methods that operate in a one-time or set-to-set manner, our approach produces segment-wise identity decisions, enabling real-time identity monitoring in dynamic scenarios. To tackle challenges such as temporal variation and identity-irrelevant noise (e.g., pose, blur, occlusion), we propose a novel three-stage framework comprising: (1) disentangled representation learning to suppress identity-irrelevant components in pre-trained embeddings, (2) attention-based intra-segment aggregation to extract robust identity cues within each segment, and (3) inter-segment adaptive identity feature completion to incrementally integrate new segments with prior identity representations for stable and efficient prediction. Extensive experiments on IJB-B, IJB-C benchmark datasets, and our own collected dataset under real-world video conferencing conditions demonstrate that our method achieves competitive accuracy with notable improvements in inference efficiency (up to 2× speed-up) over state-of-the-art baselines. These results validate the effectiveness and practicality of our framework for real-world continuous authentication systems. Xiang Li 0028, Chi Xu 0003, Yasushi Yagi |
IJCB | 1 |
| 2025 | EEG-based User Authentication in Realistic Scenarios: From Solo Reading to Dialog GamesabstractWhile electroencephalography (EEG)-based user authentication has demonstrated strong potential in controlled laboratory conditions, its reliability in more realistic settings remains underexplored. In this work, we investigate EEG-based user authentication in realistic interactive scenarios involving cognitively and behaviorally rich tasks. We collected EEG data from subjects under two conditions: solo reading aloud and two-person dialog-based games. These scenarios represent practical, everyday activities where users produce speech and engage in turn-taking interactions, introducing non-stationarity, muscle artifacts, and attention shifts that challenge conventional EEG-based models. To evaluate the user authentication performance under these challenges, we systematically apply various deep learning backbone models and loss functions under a fair and consistent experimental protocol. Specifically, we first divide the whole EEG signal into short windows with pre-defined length and stride. For each window, a time-frequency representation is computed using a continuous wavelet transform, resulting in a sequence of 2D time-frequency maps that capture the non-stationary characteristics of EEG signals. These maps are then fed into various backbone models, such as long short term memory (LSTM) and convolutional neural network (CNN), which are trained to extract robust identity-discriminative features. To further enhance inter-subject separability, we employ the Softmax loss or ArcFace loss during training. Experimental results demonstrate that even under dynamic and less controlled conditions, EEG signals retain individual-specific patterns, yielding high authentication accuracy. These findings highlight the feasibility of extending EEG-based biometrics to more natural environments. Chi Xu 0003, Xiang Li 0028, Shuqiong Wu, Yasushi Yagi |
IJCB | 2 |
| 2024 | On Cropping for Gait Recognition: Does Constant-velocity Locomotion Assumption Improve Gait Recognition Accuracy?abstractMost of gait recognition studies focus on feature extraction and matching steps by using cropped image sequences as inputs. We usually use frame-by-frame tight bounding boxes (TBBs) obtained by pedestrian detection or instance segmentation for cropping. Cropping by the TBB, however, suffers from apparent scale changes within gait period (e.g., apparent height changes between single/double support phases), which may cause a drop in gait recognition accuracy. We therefore propose a method of cropping for gait recognition to better preserve the apparent scale by introducing constant-velocity locomotion (CVL) assumption for a short period (e.g., one second). We derive that a bounding box sequence (BBS) in the 2D image coordinate under CVL assumption in the 3D camera coordinate, is represented by non-linear interpolation between the starting and ending frames without camera calibration parameters. We then estimate BBS parameters (i.e., bounding boxes at the starting and ending frames) by generalized Hough transform with voting from pedestrian region proposals. Experiments with OU-MVLP show that the proposed cropping improves the gait recognition accuracies for both model-based and appearance-based approaches. Tappei Okimura, Xiang Li 0028, Chi Xu 0003, Yasushi Yagi |
IJCB | 2 |
| 2023 | Online Model-based Gait Age and Gender EstimationabstractThis paper presents an online human model-based framework for gait-based age and gender estimation from a sequence of monocular frames. More specifically, we fine-tune a human mesh recovery model (i.e., HMR) to estimate the shape and pose parameters of a predefined 3D human model (i.e., SMPL). We then utilize the estimated parameters to predict the age and gender of the walking subject. To make the age and gender estimation task more favorable for real-time applications, we consider estimating the corresponding probability distributions of age and gender, which preserve the prediction uncertainty. Experiments on the world’s largest multi-view gait age and gender estimation dataset showed the superiority of the proposed method compared to the existing appearance-based baseline. We implement online standalone and client-server systems based on the proposed framework to demonstrate the performance of real-time estimation. We further propose a geometric correction step to the input gait sequence for a more generalization capability of the online system. Allam Shehata, Mohamad Ammar Alsherfawi Aljazaerly, Levin Gäher, Xiang Li 0028, Yasushi Makihara, Yasushi Yagi |
IJCB | 4 |
| 2023 | Occlusion-Aware Human Mesh Model-Based Gait RecognitionabstractPartial occlusion of the human body caused by obstacles or a limited camera field of view often occurs in surveillance videos, which affects the performance of gait recognition in practice. Existing methods for gait recognition against occlusion require a bounding box or the height of a full human body as a prerequisite, which is unobserved in occlusion scenarios. In this paper, we propose an occlusion-aware model-based gait recognition method that works directly on gait videos under occlusion without the above-mentioned prerequisite. Specifically, given a gait sequence that only contains non-occluded body parts in the images, we directly fit a skinned multi-person linear (SMPL)-based human mesh model to the input images without any pre-normalization or registration of the human body. We further use the pose and shape features extracted from the estimated SMPL model for recognition purposes, and use the extracted camera parameters in the occlusion attenuation module to reduce intra-subject variation in human model fitting caused by occlusion pattern differences. Experiments on occlusion samples simulated from the OU-MVLP dataset demonstrated the effectiveness of the proposed method, which outperformed state-of-the-art gait recognition methods by about 15% rank-1 identification rate and 2% equal error rate in the identification and verification scenarios, respectively. Chi Xu 0003, Yasushi Makihara, Xiang Li 0028, Yasushi Yagi |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | PVT v2: Improved baselines with Pyramid Vision TransformerabstractTransformers have recently lead to encouraging progress in computer vision. In this work, we present new baselines by improving the original Pyramid Vision Transformer (PVT v1) by adding three designs: (i) a linear complexity attention layer, (ii) an overlapping patch embedding, and (iii) a convolutional feed-forward network. With these modifications, PVT v2 reduces the computational complexity of PVT v1 to linearity and provides significant improvements on fundamental vision tasks such as classification, detection, and segmentation. In particular, PVT v2 achieves comparable or better performance than recent work such as the Swin transformer. We hope this work will facilitate state-of-the-art transformer research in computer vision. Code is available at https://github.com/whai362/PVT . Wenhai Wang, Enze Xie, Xiang Li 0028, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu 0002, Ping Luo 0002, Ling Shao 0001 |
Comput. Vis. Media | 3 |
| 2021 | Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsabstractAlthough convolutional neural networks (CNNs) have achieved great success in computer vision, this work investigates a simpler, convolution-free backbone network use-fid for many dense prediction tasks. Unlike the recently-proposed Vision Transformer (ViT) that was designed for image classification specifically, we introduce the Pyramid Vision Transformer (PVT), which overcomes the difficulties of porting Transformer to various dense prediction tasks. PVT has several merits compared to current state of the arts. (1) Different from ViT that typically yields low-resolution outputs and incurs high computational and memory costs, PVT not only can be trained on dense partitions of an image to achieve high output resolution, which is important for dense prediction, but also uses a progressive shrinking pyramid to reduce the computations of large feature maps. (2) PVT inherits the advantages of both CNN and Transformer, making it a unified backbone for various vision tasks without convolutions, where it can be used as a direct replacement for CNN backbones. (3) We validate PVT through extensive experiments, showing that it boosts the performance of many downstream tasks, including object detection, instance and semantic segmentation. For example, with a comparable number of parameters, PVT+RetinaNet achieves 40.4 AP on the COCO dataset, surpassing ResNet50+RetinNet (36.3 AP) by 4.1 absolute AP (see Figure 2). We hope that PVT could, serre as an alternative and useful backbone for pixel-level predictions and facilitate future research. Wenhai Wang, Enze Xie, Xiang Li 0028, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu 0002, Ping Luo 0002, Ling Shao 0001 |
ICCV | 3 |
| 2021 | Real-Time Gait-Based Age Estimation and Gender Classification from a Single ImageabstractIn this paper, we propose a unified real-time framework for gait-based age estimation and gender classification that uses just a single image, which reduces the latency in video capturing compared with the existing methods based on a gait cycle. To cope with the problem of lacking motion information in the input single image, we first reconstruct a gait cycle of a silhouette sequence from the input image via a gait cycle reconstruction network. The reconstructed gait cycle is then fed into a state-of-the-art gait recognition network for feature representation learning, which is further used to obtain the class of the gender and the estimated probability distribution of integer age labels. Unlike the existing methods focusing on the gait sequences captured from the side view, the proposed method is applicable to the gait images from an arbitrary view with a single trained model, which is more suitable for real-world application scenarios (e.g., automatic access control). Stand-alone and client-server online systems were implemented based on the proposed method, which validates the real-time/online property in actual scenes. The experiments on the world's largest multi-view gait dataset demonstrate the effectiveness of the proposed method, which achieves performance improvement compared with the benchmark algorithms. Chi Xu 0003, Yasushi Makihara, Ruochen Liao, Hirotaka Niitsuma, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003 |
WACV | 5 |
| 2021 | Cross-View Gait Recognition Using Pairwise Spatial Transformer NetworksabstractIn this paper, we propose a pairwise spatial transformer network (PSTN) for cross-view gait recognition, which reduces unwanted feature mis-alignment due to view differences before a recognition step for better performance. The proposed PSTN is a unified CNN architecture that consists of a pairwise spatial transformer (PST) and subsequent recognition network (RN). More specifically, given a matching pair of gait features from different source and target views, the PST estimates a non-rigid deformation field to register the features in the matching pair into their intermediate view, which mitigates distortion by registration compared with the case of direct deformation from the source view to target view. The registered matching pair is then fed into the RN to output a dissimilarity score. Although registration may reduce not only intra-subject variations but also inter-subject variations, we can still achieve a good trade-off between them using a loss function designed to optimize recognition accuracy. Experiments on three publicly available gait datasets demonstrate that the proposed method yields superior performance for both verification and identification scenarios by combining any gait recognition network benchmarks with the PST. Chi Xu 0003, Yasushi Makihara, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | End-to-End Model-Based Gait Recognition
Xiang Li 0028, Yasushi Makihara, Chi Xu 0003, Yasushi Yagi, Shiqi Yu 0001, Mingwu Ren |
ACCV (3) | 1 |
| 2020 | Gait Recognition via Semi-supervised Disentangled Representation Learning to Identity and Covariate FeaturesabstractExisting gait recognition approaches typically focus on learning identity features that are invariant to covariates (e.g., the carrying status, clothing, walking speed, and viewing angle) and seldom involve learning features from the covariate aspect, which may lead to failure modes when variations due to the covariate overwhelm those due to the identity. We therefore propose a method of gait recognition via disentangled representation learning that considers both identity and covariate features. Specifically, we first encode an input gait template to get the disentangled identity and covariate features, and then decode the features to simultaneously reconstruct the input gait template and the canonical version of the same subject with no covariates in a semi-supervised manner to ensure successful disentanglement. We finally feed the disentangled identity features into a contrastive/triplet loss function for a verification/identification task. Moreover, we find that new gait templates can be synthesized by transferring the covariate feature from one subject to another. Experimental results on three publicly available gait data sets demonstrate the effectiveness of the proposed method compared with other state-of-the-art methods. Xiang Li 0028, Yasushi Makihara, Chi Xu 0003, Yasushi Yagi, Mingwu Ren |
CVPR | 1 |
| 2020 | Gait Recognition from a Single Image Using a Phase-Aware Gait Cycle Reconstruction Network
Chi Xu 0003, Yasushi Makihara, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003 |
ECCV (19) | 3 |
| 2020 | Gait recognition invariant to carried objects using alpha blending generative adversarial networksabstractGait recognition invariant to carried objects (COs) is very difficult in a real-life scene because the COs can have various shapes and sizes, in addition to unpredictable carrying locations (e.g., front, back, and side, or multiple locations). Therefore, in this paper, we propose a robust method for gait recognition against various COs by reconstructing a gait template without COs. A straightforward approach is to directly generate a gait template without COs given a gait template with COs as the input using a conventional generative adversarial network. There is, however, a potential risk of unnecessarily altering parts that were originally unaffected by COs (e.g., leg parts for a person carrying a backpack). Because we do not want to touch such unaffected parts in the original template, we first estimate a gait template without COs, and then blend it with the original template by an estimated alpha matte that indicates the blending parameters. We then create an alpha-blended template from the original template and the generated template without COs based on the estimated alpha matte. We use two independent generators to estimate the alpha matte and the generated template without COs. Finally, we feed the alpha-blended gait template into a state-of-the-art discrimination network for gait recognition. The experimental results on three publicly available gait databases with real-life COs demonstrate the state-of-the-art performance of the proposed method. Xiang Li 0028, Yasushi Makihara, Chi Xu 0003, Yasushi Yagi, Mingwu Ren |
Pattern Recognit. | 1 |
| 2019 | Speed-Invariant Gait Recognition Using Single-Support Gait Energy ImageabstractGait is one of the most popular behavioral biometrics because it can be authenticated at a distance from a camera without subject cooperation. Speed differences between matching pairs, however, cause significant performance drops in gait recognition, and gait mode difference (i.e., walking versus running) makes gait recognition further challenging. We therefore propose a speed-invariant gait representation called single-support GEI (SSGEI), which realizes a good trade-off between speed invariance and stability by aggregating multiple frames around single-support phases. In addition, to mitigate the pose differences between walking and running modes at single-support phases, we morph walking and running SSGEIs into intermediate SSGEIs between walking and running mode, where we exploit a free-form deformation field from the walking or running modes to the intermediate mode obtained by training data. We finally apply Gabor filtering and spatial metric learning as postprocessing for further accuracy improvement. Experiments on two publicly available datasets, the OU-ISIR Treadmill Dataset A and the CASIA-C Dataset demonstrate that the proposed method yields the state-of-the-art accuracies in both identification and verification scenarios with a low computational cost. Chi Xu 0003, Yasushi Makihara, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003 |
Multim. Tools Appl. | 3 |
| 2019 | Joint Intensity Transformer Network for Gait Recognition Robust Against Clothing and Carrying StatusabstractClothing and carrying status variations are the two key factors that affect the performance of gait recognition because people usually wear various clothes and carry all kinds of objects, while walking in their daily life. These covariates substantially affect the intensities within conventional gait representations such as gait energy images. Hence, to properly compare a pair of input gait features, an appropriate metric for joint intensity is needed in addition to the conventional spatial metric. We therefore propose a unified joint intensity transformer network for gait recognition that is robust against various clothing and carrying statuses. Specifically, the joint intensity transformer network is a unified deep learning-based architecture containing three parts: a joint intensity metric estimation net, a joint intensity transformer, and a discrimination network. First, the joint intensity metric estimation net uses a well-designed encoder-decoder network to estimate a sample-dependent joint intensity metric for a pair of input gait energy images. Subsequently, a joint intensity transformer module outputs the spatial dissimilarity of two gait energy images using the metric learned by the joint intensity metric estimation net. Third, the discrimination network is a generic convolution neural network for gait recognition. In addition, the joint intensity transformer network is designed with different loss functions depending on the gait recognition task (i.e., a contrastive loss function for the verification task and a triplet loss function for the identification task). The experiments on the world's largest datasets containing various clothing and carrying statuses demonstrate the state-of-the-art performance of the proposed method. Xiang Li 0028, Yasushi Makihara, Chi Xu 0003, Yasushi Yagi, Mingwu Ren |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2018 | Gait-based human age estimation using age group-dependent manifold learning and regressionabstractHuman age estimation from gait is expected to be an important technology for a variety of applications such as automatic customer counting for marketing research or automatic age-based access control restriction for a specific area because the gait can be observable at a distance from a camera (e.g., CCTV). Although the aging process of gait significantly differs among age groups (e.g., children, adults, and the elderly), previous studies on gait-based human age estimation employ a single age group-independent estimation model that suffers from large estimation errors when the age variation increases. We therefore propose an age group-dependent gait-based human age estimation method for better accuracy. Specifically, in the training phase, we first compose age groups that are well-separated from each other by clustering gait features along with their age labels. We then learn a classifier that classifies the gait features for multiple age groups using a directed acyclic graph support vector machine. Next, we learn an age regression model for each age group using support vector regression with a Gaussian kernel in conjunction with a manifold learning technique, i.e., orthogonal locality preserving projection, to better characterize the gait feature. In the test phase, given a gait feature, it is first classified into an age group and then its age is estimated with the age regression model of the classified age group. Experimental results on a gait database that has the world’s largest population of participants ranging from 2 to 90 years old demonstrate the state-of-the-art performance of the proposed method. Xiang Li 0028, Yasushi Makihara, Chi Xu 0003, Yasushi Yagi, Mingwu Ren |
Multim. Tools Appl. | 1 |
| 2017 | Joint Intensity and Spatial Metric Learning for Robust Gait RecognitionabstractThis paper describes a joint intensity metric learning method to improve the robustness of gait recognition with silhouette-based descriptors such as gait energy images. Because existing methods often use the difference of image intensities between a matching pair (e.g., the absolute difference of gait energies for the l1-norm) to measure a dissimilarity, large intrasubject differences derived from covariate conditions (e.g., large gait energies caused by carried objects vs. small gait energies caused by the background), may wash out subtle intersubject differences (e.g., the difference of middle-level gait energies derived from motion differences). We therefore introduce a metric on joint intensity to mitigate the large intrasubject differences as well as leverage the subtle intersubject differences. More specifically, we formulate the joint intensity and spatial metric learning in a unified framework and alternately optimize it by linear or ranking support vector machines. Experiments using the OU-ISIR treadmill data set B with the largest clothing variation and large population data set with bag, β version containing carrying status in the wild demonstrate the effectiveness of the proposed method. Yasushi Makihara, Atsuyuki Suzuki, Daigo Muramatsu, Xiang Li 0028, Yasushi Yagi |
CVPR | 4 |
| 2016 | Gait Energy Response Function for Clothing-Invariant Gait Recognition
Xiang Li 0028, Yasushi Makihara, Chi Xu 0003, Daigo Muramatsu, Yasushi Yagi, Mingwu Ren |
ACCV (2) | 1 |
| 2016 | Speed Invariance vs. Stability: Cross-Speed Gait Recognition Using Single-Support Gait Energy Image
Chi Xu 0003, Yasushi Makihara, Xiang Li 0028, Yasushi Yagi, Jianfeng Lu 0003 |
ACCV (2) | 3 |