Xin Zhang 0013

dblp:76/1584-13 · DBLP profile ↗
← Back
38ranked-venue papers
6as first author
22since 2021 · last 2026
0000-0003-1583-6401ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 Triplet longitudinal masked autoencoder for predicting individualized functional connectome development during infancy
Weiran Xia, Xin Zhang 0013, Dan Hu 0004, Xiaowei Yu 0001, Weiyan Yin, Zhengwang Wu, Li Wang 0026, Weili Lin, Gang Li 0001
Medical Image Anal.2
2025 Hand-Aware Masked Graph Convolutional Network for Skeleton-based Sign Language Recognition
abstract
Sign language recognition (SLR) presents significant challenges due to the complexities in modeling the intricate interactions between hand gestures and other body movements, as well as the presence of redundant and noisy temporal features. Existing Graph Convolutional Network (GCN)-based methods, while promising in capturing skeletal dynamics, are limited by their reliance on fixed skeletal structures and uniform frame importance, which hinder their effectiveness in addressing these issues. To overcome these challenges, we propose a novel Hand-Aware Masked Graph Convolutional Network (HAM-GCN), which introduces two key innovations: an adaptive dynamic hand graph and an adaptive masking mechanism. The hand graph effectively captures fine-grained interactions between hand gestures and body movements, enhancing the accuracy of sign language representations. The adaptive masking mechanism, driven by hand-specific confidence scores, dynamically adjusts the importance of individual frames, emphasizing informative ones while suppressing noisy or redundant features. We validate HAM-GCN through extensive experiments on three publicly available sign language datasets, achieving state-of-the-art performance and surpassing existing methods. The source code is publicly available at https://github.com/Gavin113/HAM-GCN.
Jiyu Qian, Dongzi Shi, Xin Zhang 0013
FG3
2025 Drawing Developmental Trajectory From Cortical Surface Reconstruction
Ruowen Qu, Zhongliang Liu, Zhuoyan Dai, Dongzi Shi, Sijin Yu, Tong Xiong, Shiping Liu, Xiangmin Xu 0001, Xiaofen Xing, Xin Zhang 0013
ICCV11
2025 DP-Net: A 3D Dilated Projection Framework For Precise Fetal Brain Tissue Segmentation
abstract
Precise segmentation of fetal brain tissues in MRI is essential for studying brain development and for the early diagnosis and treatment of neurological disorders. However, the complex and variable anatomy of the fetal brain, significant morphological changes at different gestational ages, and the low-quality MRI and inherent noise of fetal acquisition pose significant challenges. To address these, we propose a novel 3D Dilated Projection U-net segmentation framework, DP-Net, which incorporates large kernel convolutions, atrous convolution for receptive field expansion, and dual skip connections mechanism to enhance global semantic consistency. Specifically, we introduce a Dilated Projection Block (DPB) that leverages atrous convolution to capture global context across multiple anatomical regions without additional parameters. Furthermore, we propose a Dual Skip Connection (DSC) mechanism to maintain encoder-decoder global consistency by fusing low-level and projected high-level features, mitigating blind spots introduced by atrous convolution. Extensive experiments show that our method significantly outperforms state-of-the-art methods, demonstrating its robustness and effectiveness in addressing the challenges of fetal brain tissue segmentation.
Junpeng Tan, Mingjin Chen, Chunmei Qing, Xin Zhang 0013, Xiangmin Xu 0001
ICIP4
2025 Heterogeneous Masked Attention-Guided Path Convolution for Functional Brain Network Analysis
Jiakun Xu, Xin Zhang 0013, Tong Xiong, Shengxian Chen, Xiaofen Xing, Jindou Hao, Xiangmin Xu 0001
MICCAI (12)2
2025 Whole slide cervical cancer classification via graph attention networks and contrastive learning
Manman Fei, Xin Zhang 0013, Dongdong Chen 0003, Zhiyun Song, Qian Wang 0001, Lichi Zhang
Neurocomputing2
2024 Adaptive Global Gesture Paths and Signature Features for Skeleton-based Gesture Recognition
Dongzi Shi, Xin Zhang 0013, Tong Xiong, Hao Ni 0001
ICPR (15)2
2024 Fetal MRI Reconstruction by Global Diffusion and Consistent Implicit Representation
Junpeng Tan, Xin Zhang 0013, Chunmei Qing, Chaoxiang Yang, He Zhang 0023, Gang Li 0001, Xiangmin Xu 0001
MICCAI (7)2
2024 Cortical Surface Reconstruction from 2D MRI with Segmentation-Constrained Super-Resolution and Representation Learning
Ruowen Qu, Dongzi Shi, Tong Xiong, Xiangmin Xu 0001, Xiaofen Xing, Xin Zhang 0013
MICCAI (2)7
2024 Skeleton-Based Gesture Recognition With Learnable Paths and Signature Features
abstract
For the skeleton-based gesture recognition, graph convolutional networks (GCNs) have achieved remarkable performance since the human skeleton is a natural graph. However, the biological structure might not be the crucial one for motion analysis. Also, spatial differential information like joint distance and angle between bones may be overlooked during the graph convolution. In this article, we focus on obtaining meaningful joint groups and extracting their discriminative features by the path signature (PS) theory. Firstly, to characterize the constraints and dependencies of various joints, we propose three types of paths, i.e., spatial, temporal, and learnable path. Especially, a learnable path generation mechanism can group joints together that are not directly connected or far away, according to their kinematic characteristic. Secondly, to obtain informative and compact features, a deep integration of PS with few parameters are introduced. All the computational process is packed into two modules, i.e., spatial-temporal path signature module (ST-PSM) and learnable path signature module (L-PSM) for the convenience of utilization. They are plug-and-play modules available for any neural network like CNNs and GCNs to enhance the feature extraction ability. Extensive experiments have conducted on three mainstream datasets (ChaLearn 2013, ChaLearn 2016, and AUTSL). We achieved the state-of-the-art results with simpler framework and much smaller model size. By inserting our two modules into the several GCN-based networks, we can observe clear improvements demonstrating the great effectiveness of our proposed method.
Dongzi Shi, Chenyang Li 0007, Yu Li 0043, Hao Ni 0001, Xin Zhang 0013
IEEE Trans. Multim.7
2024 Fourier Domain Robust Denoising Decomposition and Adaptive Patch MRI Reconstruction
abstract
The sparsity of the Fourier transform domain has been applied to magnetic resonance imaging (MRI) reconstruction in k -space. Although unsupervised adaptive patch optimization methods have shown promise compared to data-driven-based supervised methods, the following challenges exist in MRI reconstruction: 1) in previous k -space MRI reconstruction tasks, MRI with noise interference in the acquisition process is rarely considered. 2) Differences in transform domains should be resolved to achieve the high-quality reconstruction of low undersampled MRI data. 3) Robust patch dictionary learning problems are usually nonconvex and NP-hard, and alternate minimization methods are often computationally expensive. In this article, we propose a method for Fourier domain robust denoising decomposition and adaptive patch MRI reconstruction (DDAPR). DDAPR is a two-step optimization method for MRI reconstruction in the presence of noise and low undersampled data. It includes the low-rank and sparse denoising reconstruction model (LSDRM) and the robust dictionary learning reconstruction model (RDLRM). In the first step, we propose LSDRM for different domains. For the optimization solution, the proximal gradient method is used to optimize LSDRM by singular value decomposition and soft threshold algorithms. In the second step, we propose RDLRM, which is an effective adaptive patch method by introducing a low-rank and sparse penalty adaptive patch dictionary and using a sparse rank-one matrix to approximate the undersampled data. Then, the block coordinate descent (BCD) method is used to optimize the variables. The BCD optimization process involves valid closed-form solutions. Extensive numerical experiments show that the proposed method has a better performance than previous methods in image reconstruction based on compressed sensing or deep learning.
Junpeng Tan, Xin Zhang 0013, Chunmei Qing, Xiangmin Xu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Look and Think: Intrinsic Unification of Self-Attention and Convolution for Spatial-Channel Specificity
abstract
Convolution and self-attention are popular paradigms and many works take them as two separate components to explore their potential combination. In this work, we consider their intrinsic properties in spatial and channel domains for vision representation. Convolution has the great property of channel-specificity to "think" to refine diverse features, and it has the weaknesses of spatial-independent to perceive different regions. Self-attention has the great property of spatial-specificity to "look" to perceive various regions, and it has the weaknesses of channel-insensible to summarize different features. With the intrinsic insight of this, we combine the spatial-specificity of self-attention and channel-specificity of convolution to effectively compensate their respective weakness. We propose a unified module, termed as SCS module, to achieve the combinative advantage of Spatial-Channel Specificity. Specifically, SCS module calculates dynamic attention weight with self-attention mechanism, followed by a weighted sum of input features similar to convolution. Extensive experiments show that SCS module improves both of the CNN and Transformer models on image classification and downstream tasks. The visualizations show the outstanding ability of SCS module for vision representation.
Honghui Lin, Yu Li 0043, Ruiyan Fang, Xin Zhang 0013
ICASSP5
2023 Prediction of Infant Cognitive Development with Cortical Surface-Based Multimodal Learning
Xin Zhang 0013, Fenqiang Zhao, Zhengwang Wu, Xinrui Yuan, Li Wang 0026, Weili Lin, Gang Li 0001
MICCAI (2)2
2023 Path-Based Heterogeneous Brain Transformer Network for Resting-State Functional Connectivity Analysis
Ruiyan Fang, Yu Li 0043, Xin Zhang 0013, Shengxian Chen, Xiangmin Xu 0001, Jieling Wu, Weili Lin, Li Wang 0026, Zhengwang Wu, Gang Li 0001
MICCAI (8)3
2023 Robust Cervical Abnormal Cell Detection via Distillation from Local-Scale Consistency Refinement
Manman Fei, Xin Zhang 0013, Maosong Cao, Zhenrong Shen 0001, Xiangyu Zhao 0003, Zhiyun Song, Qian Wang 0001, Lichi Zhang
MICCAI (6)2
2023 Predicting Diverse Functional Connectivity from Structural Connectivity Based on Multi-contexts Discriminator GAN
Xin Zhang 0013, Lu Zhang 0050, Xiangmin Xu 0001, Dajiang Zhu
MICCAI (8)2
2022 Whole Slide Cervical Cancer Screening Using Graph Attention Network and Supervised Contrastive Learning
Xin Zhang 0013, Maosong Cao, Sheng Wang 0014, Jiayin Sun, Xiangshan Fan, Qian Wang 0001, Lichi Zhang
MICCAI (2)1
2022 Path Signature Neural Network of Cortical Features for Prediction of Infant Cognitive Scores
abstract
Studies have shown that there is a tight connection between cognition skills and brain morphology during infancy. Nonetheless, it is still a great challenge to predict individual cognitive scores using their brain morphological features, considering issues like the excessive feature dimension, small sample size and missing data. Due to the limited data, a compact but expressive feature set is desirable as it can reduce the dimension and avoid the potential overfitting issue. Therefore, we pioneer the path signature method to further explore the essential hidden dynamic patterns of longitudinal cortical features. To form a hierarchical and more informative temporal representation, in this work, a novel cortical feature based path signature neural network (CF-PSNet) is proposed with stacked differentiable temporal path signature layers for prediction of individual cognitive scores. By introducing the existence embedding in path generation, we can improve the robustness against the missing data. Benefiting from the global temporal receptive field of CF-PSNet, characteristics consisted in the existing data can be fully leveraged. Further, as there is no need for the whole brain to work for a certain cognitive ability, a top K selection module is used to select the most influential brain regions, decreasing the model size and the risk of overfitting. Extensive experiments are conducted on an in-house longitudinal infant dataset within 9 time points. By comparing with several recent algorithms, we illustrate the state-of-the-art performance of our CF-PSNet (i.e., root mean square error of 0.027 with the time latency of 518 milliseconds for each sample).
Xin Zhang 0013, Hao Ni 0001, Chenyang Li 0007, Xiangmin Xu 0001, Zhengwang Wu, Li Wang 0026, Weili Lin, Gang Li 0001
IEEE Trans. Medical Imaging2
2022 Brain Connectivity Based Graph Convolutional Networks and Its Application to Infant Age Prediction
abstract
Infancy is a critical period for the human brain development, and brain age is one of the indices for the brain development status associated with neuroimaging data. The difference between the predicted age based on neuroimaging and the chronological age can provide an important early indicator of deviation from the normal developmental trajectory. In this study, we utilize the Graph Convolutional Network (GCN) to predict the infant brain age based on resting-state fMRI data. The brain connectivity obtained from rs-fMRI can be represented as a graph with brain regions as nodes and functional connections as edges. However, since the brain connectivity is a fully connected graph with features on edges, current GCN cannot be directly used for it is a node-based method for sparse graphs. Hence, we propose an edge-based Graph Path Convolution (GPC) method, which aggregates the information from different paths and can be naturally applied on dense graphs. We refer the whole model as Brain Connectivity Graph Convolutional Networks (BC-GCN). Further, two upgraded network structures are proposed by including the residual and attention modules, referred as BC-GCN-Res and BC-GCN-SE to emphasize the information of the original data and enhance influential channels. Moreover, we design a two-stage coarse-to-fine framework, which determines the age group first and then predicts the age using group-specific BC-GCN-SE models. To avoid accumulated errors from the first stage, a cross-group training strategy is adopted for the second stage regression models. We conduct experiments on infant fMRI scans from 6 to 811 days of age. The coarse-to-fine framework shows significant improvements when being applied to several models (reducing error over 10 days). Comparing with state-of-the-art methods, our proposed model BC-GCN-SE with coarse-to-fine framework reduces the mean absolute error of the prediction from >70 days to 49.9 days. The code is now available at https://github.com/SCUT-Xinlab/BC-GCN.
Yu Li 0043, Xin Zhang 0013, Jingxin Nie, Ruiyan Fang, Xiangmin Xu 0001, Zhengwang Wu, Dan Hu 0004, Li Wang 0026, Han Zhang 0002, Weili Lin, Gang Li 0001
IEEE Trans. Medical Imaging2
2021 Temporal-Spatial Deformable Pose Network For Skeleton-Based Gesture Recognition
abstract
Gesture recognition is a challenging research topic, and also has a wide range of potential applications in our daily life. With the development of hardware and advanced algorithms, we can easily extract skeleton data from video sequences and apply them for the recognition task. In this paper, we propose a novel temporal-spatial deformable pose network to leverage space and time information together. Our proposed network can automatically locate the most correlated joints across multiple frames and extract features accordingly. Additionally, we introduce a parallel multi-scale convolutional layer with different dilation rates, which can capture multi-term temporal information efficiently. We have conducted experiments on MSRC-12, ChaLearn 2013, and ChaLearn 2016 datasets and our proposed method outperforms state-of-the-art methods. Moreover, Additional experiments showed that our proposed module is more robust to handle noise data and dynamic gestures with various temporal scales.
Honghui Lin, Yu Li 0043, Xin Zhang 0013
ICIP4
2021 A Novel Unsupervised domain adaptation method for inertia-Trajectory translation of in-air handwriting
Songbin Xu, Yang Xue 0001, Xin Zhang 0013
Pattern Recognit.3
2021 HF-UNet: Learning Hierarchically Inter-Task Relevance in Multi-Task U-Net for Accurate Prostate Segmentation in CT Images
abstract
Accurate segmentation of the prostate is a key step in external beam radiation therapy treatments. In this paper, we tackle the challenging task of prostate segmentation in CT images by a two-stage network with 1) the first stage to fast localize, and 2) the second stage to accurately segment the prostate. To precisely segment the prostate in the second stage, we formulate prostate segmentation into a multi-task learning framework, which includes a main task to segment the prostate, and an auxiliary task to delineate the prostate boundary. Here, the second task is applied to provide additional guidance of unclear prostate boundary in CT images. Besides, the conventional multi-task deep networks typically share most of the parameters (i.e., feature representations) across all tasks, which may limit their data fitting ability, as the specificity of different tasks are inevitably ignored. By contrast, we solve them by a hierarchically-fused U-Net structure, namely HF-UNet. The HF-UNet has two complementary branches for two tasks, with the novel proposed attention-based task consistency learning block to communicate at each level between the two decoding branches. Therefore, HF-UNet endows the ability to learn hierarchically the shared representations for different tasks, and preserve the specificity of learned representations for different tasks simultaneously. We did extensive evaluations of the proposed method on a large planning CT image dataset and a benchmark prostate zonal dataset. The experimental results show HF-UNet outperforms the conventional multi-task network architectures and the state-of-the-art methods.
Kelei He, Chunfeng Lian, Bing Zhang 0012, Xin Zhang 0013, Xiaohuan Cao, Dong Nie, Yang Gao 0001, Dinggang Shen
IEEE Trans. Medical Imaging4
2020 Joint Image Quality Assessment and Brain Extraction of Fetal MRI Using Deep Learning
Lufan Liao, Xin Zhang 0013, Fenqiang Zhao, Tao Zhong 0002, Yuchen Pei, Xiangmin Xu 0001, Li Wang 0026, He Zhang 0023, Dinggang Shen, Gang Li 0001
MICCAI (6)2
2020 Infant Cognitive Scores Prediction with Multi-stream Attention-Based Temporal Path Signature Features
Xin Zhang 0013, Hao Ni 0001, Chenyang Li 0007, Xiangmin Xu 0001, Zhengwang Wu, Li Wang 0026, Weili Lin, Dinggang Shen, Gang Li 0001
MICCAI (7)1
2020 Attention Based Dual Branches Fingertip Detection Network and Virtual Key System
abstract
Gesture and fingertip are becoming more and more important mediums for human-computer interaction (HCI). Therefore, algorithms of gesture recognition and fingertip detection have been extensively investigated. However, problems mainly remain in how to achieve a win-win situation between speed and accuracy, and how to deal with complex interaction environment. To rectify these problems, this paper proposes an attention-based dual branches network that can efficiently fulfill both fingertip detection and gesture recognition tasks. In order to deal with complex interaction environment, we combine both channel-wise attention and spatial-wise attention into the fingertip detection model. The extensive experiments demonstrate that our novel model is both effective and efficient. In the experiment, our proposed model achieves the average fingertip detection error at around 2.8 pixels in 640×480 video frame, and the average recognition accuracy among eight gestures reaches $99%$. Moreover, the average forward time is about 8 ms. Due to the light-weight design, this model can also achieve high-efficiency performance on CPU. In addition, we design a virtual key system based on our proposed model, which can allow users to complete the "clicking" operation naturally in virtual environment. Our proposed system can perform well with a single normal RGB camera without any pre-processing (e.g., image segmentation or contour extraction), which can significantly reduce the complexity of the interaction system.
Chong Mou, Xin Zhang 0013
ACM Multimedia2
2020 Automatic fetal brain extraction from 2D in utero fetal MRI slices using deep neural network
Yishan Luo, Lin Shi 0001, Xin Zhang 0013, Ming Li 0005, Bing Zhang 0012, Defeng Wang
Neurocomputing4
2019 Skeleton-Based Gesture Recognition Using Several Fully Connected Layers with Path Signature Features and Temporal Transformer Module
abstract
The skeleton based gesture recognition is gaining more popularity due to its wide possible applications. The key issues are how to extract discriminative features and how to design the classification model. In this paper, we first leverage a robust feature descriptor, path signature (PS), and propose three PS features to explicitly represent the spatial and temporal motion characteristics, i.e., spatial PS (S PS), temporal PS (T PS) and temporal spatial PS (T S PS). Considering the significance of fine hand movements in the gesture, we propose an ”attention on hand” (AOH) principle to define joint pairs for the S PS and select single joint for the T PS. In addition, the dyadic method is employed to extract the T PS and T S PS features that encode global and local temporal dynamics in the motion. Secondly, without the recurrent strategy, the classification model still faces challenges on temporal variation among different sequences. We propose a new temporal transformer module (TTM) that can match the sequence key frames by learning the temporal shifting parameter for each input. This is a learning-based module that can be included into standard neural network architecture. Finally, we design a multi-stream fully connected layer based network to treat spatial and temporal features separately and fused them together for the final result. We have tested our method on three benchmark gesture datasets, i.e., ChaLearn 2016, ChaLearn 2013 and MSRC-12. Experimental results demonstrate that we achieve the state-of-the-art performance on skeleton-based gesture recognition with high computational efficiency.
Chenyang Li 0007, Xin Zhang 0013, Lufan Liao
AAAI2
2019 Multi-path Convolutional Neural Network based on Rectangular Kernel with Path Signature Features for Gesture Recognition
abstract
Skeleton based gesture recognition has gained more attention due to its wide application and large-scale databases availability. Recent methods designed for skeleton sequence data mainly pay attention to network architecture but ignore an essential characteristic of skeleton sequences that the temporal dimensionality of skeleton sequences is usually higher than its spatial dimensionality. Directly applying CNNs designed for image classification to skeleton-based data can not capture this unique property. Considering this fact, we propose the rectangular convolution and pooling to skeleton sequence data. Temporal features are crucial for gesture action recognition. Further, we introduce path signature features (PSF) to represent temporal variation characteristics of each joint. Moreover, there only exist a few minor distinctions between some gestures. To classify them more accurately, we add two sub-networks to extract discriminative features from two hands respectively. We evaluate our method on three major benchmark gesture datasets, i.e., ChaLearn 2013, ChaLearn 2016 and MSRC-12, and reach the state-of-the-art performance.
Lufan Liao, Xin Zhang 0013, Chenyang Li 0007
VCIP2
2019 Multi-heads Attention Graph Convolutional Networks for Skeleton-Based Action Recognition
abstract
Compared with video-based action recognition, skeleton-based methods have more compact and accurate representation. In the recent development, human skeleton is modeled as graph and graph-based deep learning method is applied for recognition. Generally, in actions, few joints are pivotal other than the whole body, like waving hands mostly related with hand and arm joints. Hence, we propose the data-driven multi-head attention model for graph convolutional networks. The attention model identifies key joints of every action by introducing two regularization terms, i.e., spatial diversity and local continuity. Further, we introduce the joint-wise second order motion information as the additional feature on the graph node, which represents the motion variation explicitly. We have tested our methods on the largest popular dataset, NTU-RGB+D, and we reach state-of-the-art performance.
Xin Zhang 0013
VCIP2
2018 A Selective Tracking and Detection Framework with Target Enhanced Feature
abstract
In the long time tracking, object representation and occlusion handling are two important challenges. We propose a novel selective tracking and detection framework in which a new probabilistic object-enhanced feature is integrated. Firstly, besides precise object appearance feature, we believe the neighboring foreground-background contrast is another key factor in the tracking. Hence we propose a foreground probability map to enhance the target and weaken the surrounding background. It is computed based on the object color distribution and its comparison with the surrounding background. Secondly, we introduce the selective tracking and detection framework that has two sets of conditions to control the detector activation and final result selection. The detector will only be activated when the tracker is not trustable, which is determined by the tracking confidence and foreground parochiality value. Then, given the tracking and detection results, the final output is selected in terms of their individual correspondence values. We have evaluated our methods on two popular benchmark datasets. Extensive experiments demonstrate that our algorithm performs favorably comparing with state-of-the-art methods.
Xinyao Ding, Xin Zhang 0013
ICPR3
2018 ICPR2018 Contest on Robust Reading for Multi-Type Web Images
abstract
Electronic commerce has infiltrated every aspect of our daily lives, which offers great convenience for shopping, advertising, etc. Text in the web images is responsible to convey essential information for consumers. Algorithms that read text in these web images can facilitate applications of various types, such as goods surveillance, products classification, and intelligent retrieval or recommendation. Despite of various existing text reading tasks, this contest introduces a novel large-scale dataset named MTWI that contains 20,000 images, which is the first dataset that is mainly constructed by Chinese and English web text. Three tasks (web text recognition, web text detection, and end-to-end web text detection and recognition) were set up for encouraging more research on the web text reading problem. The contest was held from February 2, 2018 to May 26, 2018 with 289 valid submissions from 4,282 registered teams. Throughout this report, we describe the details of this new dataset, the purposes and definitions of the tasks, the evaluation protocols, and the summaries of the results.
Mengchao He, Zhibo Yang 0003, Sheng Zhang 0024, Canjie Luo, Feiyu Gao, Qi Zheng 0002, Yongpan Wang, Xin Zhang 0013
ICPR9
2017 Video-Based Human Walking Estimation Using Joint Gait and Pose Manifolds
abstract
We study two fundamental issues about video-based human walking estimation, where the goal is to estimate 3D gait kinematics (i.e., joint positions) from 2D gait appearances (i.e., silhouettes). One is how to model the gait kinematics from different walking styles, and the other is how to represent the gait appearances captured under different views and from individuals of distinct walking styles and body shapes. Our research is conducted in three steps. First, we propose the idea of joint gait-pose manifold (JGPM), which represents gait kinematics by coupling two nonlinear variables, pose (a specific walking stage) and gait (a particular walking style) in a unified latent space. We extend the Gaussian process latent variable model (GPLVM) for JGPM learning, where two heuristic topological priors, a torus and a cylinder, are considered and several JGPMs of different degrees of freedom (DoFs) are introduced for comparative analysis. Second, we develop a validation technique and a series of benchmark tests to evaluate multiple JGPMs and recent GPLVMs in terms of their performance for gait motion modeling. It is shown that the toroidal prior is slightly better than the cylindrical one, and the JGPM of 4 DoFs that balances the toroidal prior with the intrinsic data structure achieves the best performance. Third, a JGPM-based visual gait generative model (JGPM-VGGM) is developed, where JGPM plays a central role to bridge the gap between the gait appearances and the gait kinematics. Our proposed JGPM-VGGM is learned from Carnegie Mellon University MoCap data and tested on the HumanEva-I and HumanEva-II data sets. Our experimental results demonstrate the effectiveness and competitiveness of our algorithms compared with existing algorithms.
Xin Zhang 0013, Meng Ding 0001, Guoliang Fan 0001
IEEE Trans. Circuits Syst. Video Technol.1
2015 DeepFinger: A Cascade Convolutional Neuron Network Approach to Finger Key Point Detection in Egocentric Vision with Mobile Camera
abstract
In this paper, we introduce a new approach to finger key point detection. For RGB images captured from an egocentric vision with a mobile camera, fingertip point detection remains a challenging problem due to various factors, like background complexity, illumination variety, hand shape diversity, and image blur cause by camera movements. To address these issues, we propose a bi-level cascade structure of a convolutional neuron network (CNN). The first-level CNN generates a bounding box of hand region by filtering a large proportion of complicated background information. Using the bounding box area as input, the second-level CNN including an extra branch returns accurate fingertip location with a multi-channel dataset. Our approach is the first attempt of finger key point detection from an egocentric vision with a mobile camera. The proposed method achieves satisfying and significant better results compared to previous fingertip detection methods based on handcraft features.
Yichao Huang, Xin Zhang 0013
SMC4
2015 Multiple Facial Image Editing Using Edge-Aware PDE Learning
abstract
This paper introduces a novel facial editing tool, called edge-aware mask, to achieve multiple photo-realistic rendering effects in a unified framework. The edge-aware masks facilitate three basic operations for adaptive facial editing, including region selection, edit setting and region blending. Inspired by the state-of-the-art edit propagation and partial differential equation (PDE) learning method, we propose an adaptive PDE model with facial priors for masks generation through edge-aware diffusion. The edge-aware masks can automatically fit the complex region boundary with great accuracy and produce smooth transition between different regions, which significantly improves the visual consistence of face editing and reduce the human intervention. Then, a unified and flexible facial editing framework is constructed, which consists of layer decomposition, edge-aware masks generation, and layer/mask composition. The combinations of multiple facial layers and edge-aware masks can achieve various facial effects simultaneously, including face enhancement, relighting, makeup and face blending etc. Qualitative and quantitative evaluations were performed using different datasets for different facial editing tasks. Experiments demonstrate the effectiveness and flexibility of our methods, and the comparisons with the previous methods indicate that improved results are obtained using the combination of multiple edge-aware masks.
Lingyu Liang, Xin Zhang 0013, Yong Xu 0007
Comput. Graph. Forum3
2013 Two-layer dual gait generative models for human motion estimation from a single camera
Xin Zhang 0013, Guoliang Fan 0001, Li-Shan Chou
Image Vis. Comput.1
2012 Structure-guided manifold learning for video-based motion estimation
abstract
We present a new structure-guided joint gait pose manifold (JGPM) that represents gait kinematics by two variables. One is the pose to denote a series of stages in a walking cycle and the other is the gait to reflect the individual walking styles. Coupling pose and gait variables in the same latent space, such as a torus-like JGPM, was shown promising and effective for video-based motion estimation. However, the two-step learning used in torus-like JGPM is computationally expensive and it separates the optimization of pose and gait variables. This work overcomes the limitations of the previous method by developing a new structure-guided JGPM that is able to jointly optimize four variables in the same latent space, leading to a much compact parameter set while sustaining a comparable performance on video-based motion estimation, as well as a great potential for large-scale learning.
Meng Ding 0001, Guoliang Fan 0001, Xin Zhang 0013, Li-Shan Chou
ICIP3
2010 Dual Gait Generative Models for Human Motion Estimation From a Single Camera
abstract
This paper presents a general gait representation framework for video-based human motion estimation. Specifically, we want to estimate the kinematics of an unknown gait from image sequences taken by a single camera. This approach involves two generative models, called the kinematic gait generative model (KGGM) and the visual gait generative model (VGGM), which represent the kinematics and appearances of a gait by a few latent variables, respectively. The concept of gait manifold is proposed to capture the gait variability among different individuals by which KGGM and VGGM can be integrated together, so that a new gait with unknown kinematics can be inferred from gait appearances via KGGM and VGGM. Moreover, a new particle-filtering algorithm is proposed for dynamic gait estimation, which is embedded with a segmental jump-diffusion Markov Chain Monte Carlo scheme to accommodate the gait variability in a long observed sequence. The proposed algorithm is trained from the Carnegie Mellon University (CMU) Mocap data and tested on the Brown University HumanEva data with promising results.
Xin Zhang 0013, Guoliang Fan 0001
IEEE Trans. Syst. Man Cybern. Part B1
2008 Dual generative models for human motion estimation from an uncalibrated monocular camera
abstract
We propose a new approach to estimate gait kinematics from image sequences taken by a monocular uncalibrated camera. This approach involves two generative models for gait representations in the kinematic and visual spaces, which induce two gait manifolds that characterize the gait variability in terms of the kinematics and visual appearance. A manifold topology enforcement scheme is introduced to incorporate the two gait manifolds. Moreover, a new particle filtering algorithm is proposed for dynamic gait tracking and estimation where a segmental jump-diffusion Markov Chain Monte Carlo (MCMC) technique is developed to accommodate the dynamic nature of the gait variability. The proposed algorithm is trained from CMU Mocap data and tested on the HumanEva dataset with promising results.
Xin Zhang 0013, Guoliang Fan 0001
ICPR1