Huazhong Ning

dblp:75/804 · DBLP profile ↗
← Back
26ranked-venue papers
14as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-authorArtificial intelligence and machine learning · 14 · 7 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
3D vision · 33% Face, body and person analysis · 23% Probabilistic and Bayesian machine learning · 16%
Theoretical computer science
2 papers
Algorithms and data structures · 95% Mathematical optimization · 5%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%
Network and information security
1 paper
Biometric security · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management › learned database components
learned data structures
0.112011
Learning to Search Efficiently in High Dimensions · NIPS 2011
Algorithms and data structures › similarity search
high-dimensional similarity search
0.112011
Learning to Search Efficiently in High Dimensions · NIPS 2011
Algorithms and data structures › data structure design › search structures › indexing
learned index
0.112011
Learning to Search Efficiently in High Dimensions · NIPS 2011
Algorithms and data structures
similarity search
0.112011
Learning to Search Efficiently in High Dimensions · NIPS 2011
Computer vision › 3D vision
3d human pose estimation
0.112008
Discriminative learning of visual words for 3D human pose estimation · CVPR 2008
Computer vision › Video understanding and tracking
action recognition
0.112008
Latent Pose Estimator for Continuous Action Recognition · ECCV (2) 2008
Machine learning › Representation and self-supervised learning › visual representation › image representation
bag of visual words
0.112008
Discriminative learning of visual words for 3D human pose estimation · CVPR 2008
Computer vision › Face, body and person analysis › gait analysis
gait recognition
0.122003
Silhouette Analysis-Based Gait Recognition for Human Identification · IEEE Trans. Pattern Anal. Mach. Intell. 2003
Fusion of Static and Dynamic Body Biometrics for Gait Recognition · ICCV 2003
Computer vision › 3D vision › pose estimation
monocular pose estimation
0.112008
Discriminative learning of visual words for 3D human pose estimation · CVPR 2008
Computer vision › 3D vision › motion capture
articulated body tracking
0.112006
Efficient Nonparametric Belief Propagation with Application to Articulated Body Tracking · CVPR (1) 2006
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
belief propagation
0.112006
Efficient Nonparametric Belief Propagation with Application to Articulated Body Tracking · CVPR (1) 2006
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › belief propagation
nonparametric belief propagation
0.112006
Efficient Nonparametric Belief Propagation with Application to Articulated Body Tracking · CVPR (1) 2006
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.012003
Silhouette Analysis-Based Gait Recognition for Human Identification · IEEE Trans. Pattern Anal. Mach. Intell. 2003
Computer vision › Face, body and person analysis
person re-identification
0.012003
Fusion of Static and Dynamic Body Biometrics for Gait Recognition · ICCV 2003
Computer vision › Face, body and person analysis › gait analysis › gait recognition
silhouette-based gait recognition
0.012003
Silhouette Analysis-Based Gait Recognition for Human Identification · IEEE Trans. Pattern Anal. Mach. Intell. 2003
Biometric security
gait recognition
0.012003
Automatic gait recognition based on statistical shape analysis · IEEE Trans. Image Process. 2003
Computer vision › 3D vision
pose estimation
0.012008
Latent Pose Estimator for Continuous Action Recognition · ECCV (2) 2008
Computer vision › Video understanding and tracking › motion analysis
human motion analysis
0.012003
Fusion of Static and Dynamic Body Biometrics for Gait Recognition · ICCV 2003
Computer vision › Face, body and person analysis
person identification
0.012003
Automatic gait recognition based on statistical shape analysis · IEEE Trans. Image Process. 2003

Methods — techniques the papers use, named apart from their topics

supervised learning · 0.2inverted indexing · 0.2boosted search forest · 0.2mode propagation · 0.1mixture gaussian density approximation · 0.1kernel fitting · 0.1procrustes shape analysis · 0.1background subtraction · 0.1pose regressor · 0.1latent variable model · 0.1distance metric learning · 0.1supervised classification · 0.0decision-level fusion · 0.0condensation · 0.0
YearPublicationVenuePosition
2019 Variational Training for Large-Scale Noisy-OR Bayesian Networks
Geng Ji 0001, Dehua Cheng, Huazhong Ning, Changhe Yuan, Hanning Zhou, Liang Xiong, Erik B. Sudderth
UAI3
2011 Learning to Search Efficiently in High Dimensions
abstract
High dimensional similarity search in large scale databases becomes an important challenge due to the advent of Internet. For such applications, specialized data structures are required to achieve computational efficiency. Traditional approaches relied on algorithmic constructions that are often data independent (such as Locality Sensitive Hashing) or weakly dependent (such as kd-trees, k-means trees). While supervised learning algorithms have been applied to related problems, those proposed in the literature mainly focused on learning hash codes optimized for compact embedding of the data rather than search efficiency. Consequently such an embedding has to be used with linear scan or another search algorithm. Hence learning to hash does not directly address the search efficiency issue. This paper considers a new framework that applies supervised learning to directly optimize a data structure that supports efficient large scale search. Our approach takes both search quality and computational cost into consideration. Specifically, we learn a boosted search forest that is optimized using pair-wise similarity labeled examples. The output of this search forest can be efficiently converted into an inverted indexing data structure, which can leverage modern text search infrastructure to achieve both scalability and efficiency. Experimental results show that our approach significantly outperforms the start-of-the-art learning to hash methods (such as spectral hashing), as well as state-of-the-art high dimensional search algorithms (such as LSH and k-means trees).
Zhen Li 0028, Huazhong Ning, Liangliang Cao, Tong Zhang 0001, Yihong Gong, Thomas S. Huang
NIPS2
2010 Incremental spectral clustering by efficiently updating the eigen-system
Huazhong Ning, Wei Xu 0007, Yun Chi, Yihong Gong, Thomas S. Huang
Pattern Recognit.1
2010 Human Pose Regression Through Multiview Visual Fusion
abstract
We consider the problem of estimating 3-D human body pose from visual signals within a discriminative framework. It is challenging because there is a wide gap between complex 3-D human motion and planar visual observation, which makes this a severely ill-conditioned problem. In this paper, we focus on three critical factors to tackle human body pose estimation, namely, feature extraction, learning algorithm, and camera utilization. On the feature level, we describe images using the salient interest points represented by scale-invariant feature transform (SIFT)-like descriptors, in which the position, appearance, and local structural information are encoded simultaneously. On the learning algorithm level, we propose to use Gaussian processes and multiple linear (ML) regression to model the mapping between poses and features. Fusing image information from multiple cameras in different views is of great interest to us on the camera level. We make a comprehensive evaluation on the HumanEva database and get two meaningful insights into the three crucial aspects for human pose estimation: 1) although the choice of feature is very important to the problem, once the learning algorithm becomes efficient, the choice of feature is no longer critical, and 2) the impact of information combination from multiple cameras on pose estimation is closely related to not only the quantity of image information, but also its quality. In most cases, it is true that the more information is involved, the better results can be achieved. But when the information quantity is the same, the differences in quality will lead to totally different performance. Furthermore, dense evaluations demonstrate that our approach is an accurate and robust solution to the human body pose estimation problem.
Xu Zhao 0001, Yun Fu 0001, Huazhong Ning, Yuncai Liu, Thomas S. Huang
IEEE Trans. Circuits Syst. Video Technol.3
2010 A General Framework to Detect Unsafe System States From Multisensor Data Stream
abstract
This paper proposes a general framework for detecting unsafe states of a system whose basic real-time parameters are captured by multiple sensors. Our approach is to learn a danger-level function that can be used to alert the users of dangerous situations in advance so that certain measures can be taken to avoid the collapse. The main challenge to this learning problem is the labeling issue, i.e., it is difficult to assign an objective danger level at each time step to the training data, except at the collapse points, where a definitive penalty can be assigned, and at the successful ends, where a certain reward can be assigned. In this paper, we treat the danger level as an expected future reward (a penalty is regarded as a negative reward) and use temporal difference (TD) learning to learn a function for approximating the expected future reward, given the current and historical sensor readings. The TD learning obtains the approximation by propagating the penalties/rewards observable at collapse points or successful ends to the entire feature space following some constraints. This avoids the labeling issue and naturally allows a general framework to detect unsafe states. Our approach is applied to, but not limited to, the application of monitoring driving safety, and the experimental results demonstrate the effectiveness of the approach.
Huazhong Ning, Wei Xu 0007, Yihong Gong, Thomas S. Huang
IEEE Trans. Intell. Transp. Syst.1
2009 Hierarchical Space-Time Model Enabling Efficient Search for Human Actions
abstract
We propose a five-layer hierarchical space-time model (HSTM) for representing and searching human actions in videos. From a features point of view, both invariance and selectivity are desirable characteristics, which seem to contradict each other. To make these characteristics coexist, we introduce a coarse-to-fine search and verification scheme for action searching, based on the HSTM model. Because going through layers of the hierarchy corresponds to progressively turning the knob between invariance and selectivity, this strategy enables search for human actions ranging from rapid movements of sports to subtle motions of facial expressions. The introduction of the Histogram of Gabor Orientations feature makes the searching for actions go smoothly across the hierarchical layers of the HSTM model. The efficient matching is achieved by applying integral histograms to compute the features in the top two layers. The HSTM model was tested on three selected challenging video sequences and on the KTH human action database. And it achieved improvement over other state-of-the-art algorithms. These promising results validate that the HSTM model is both selective and robust for searching human actions.
Huazhong Ning, Tony X. Han, Dirk Bernhardt-Walther, Ming Liu 0009, Thomas S. Huang
IEEE Trans. Circuits Syst. Video Technol.1
2008 Discriminative learning of visual words for 3D human pose estimation
abstract
This paper addresses the problem of recovering 3D human pose from a single monocular image, using a discriminative bag-of-words approach. In previous work, the visual words are learned by unsupervised clustering algorithms. They capture the most common patterns and are good features for coarse-grain recognition tasks like object classification. But for those tasks which deal with subtle differences such as pose estimation, such representation may lack the needed discriminative power. In this paper, we propose to jointly learn the visual words and the pose regressors in a supervised manner. More specifically, we learn an individual distance metric for each visual word to optimize the pose estimation performance. The learned metrics rescale the visual words to suppress unimportant dimensions such as those corresponding to background. Another contribution is that we design an Appearance and Position Context (APC) local descriptor that achieves both selectivity and invariance while requiring no background subtraction. We test our approach on both a quasi-synthetic dataset and a real dataset (HumanEva) to verify its effectiveness. Our approach also achieves fast computational speed thanks to the integral histograms used in APC descriptor extraction and fast inference of pose regressors.
Huazhong Ning, Wei Xu 0007, Yihong Gong, Thomas S. Huang
CVPR1
2008 Latent Pose Estimator for Continuous Action Recognition
Huazhong Ning, Wei Xu 0007, Yihong Gong, Thomas S. Huang
ECCV (2)1
2008 Efficient initialization of Mixtures of Experts for human pose estimation
abstract
This paper addresses the problem of recovering 3D human pose from a single monocular image. In the literature, Bayesian Mixtures of Experts (BME) was successfully used to represent the multimodal image-to-pose distributions. However, the expectation-maximization (EM) algorithm that learns the BME model may converge to a suboptimal local maximum. And the quality of the final solution depends largely on the initial values. In this paper, we propose an efficient initialization method for BME learning. We first partition the training set so that each subset can be well modeled by a single expert and the total regression error is minimized. Then each expert and gate of BME model is initialized on a partition subset. Our initialization method is tested on both a quasi-synthetic dataset and a real dataset (HumanEva). Results show that it greatly reduces the computational cost in training while improves testing accuracy.
Huazhong Ning, Yuxiao Hu 0001, Thomas S. Huang
ICIP1
2008 Temporal difference learning to detect unsafe system states
abstract
This paper proposes a general framework to detect unsafe states of a system whose basic realtime parameters are captured by multi-sensors. Our approach is to learn a danger level function which can be used to alert the users in advance of dangerous situations. The main challenge to this learning problem is the labelling issue, i.e., it is difficult to assign an objective danger level at each time step to the training data, except at the collapse points where a penalty can be assigned and at the successful ends where a certain reward can be assigned. In this paper, we treat the danger level as expected future reward (penalty is regarded as negative reward) and use temporal difference (TD) learning [2] to learn a function to approximate the expected future reward. The TD learning obtains the approximation by propagating the penalty/reward observable at collapse points or successful ends to the entire feature space following some constraints. Our approach is applied to, but not limited to, the application of monitoring of driving safety and the experimental results demonstrate the effectiveness of the approach.
Huazhong Ning, Wei Xu 0007, Yihong Gong, Thomas S. Huang
ICPR1
2008 Discriminative estimation of 3D human pose using Gaussian processes
abstract
In this paper, we present an efficient discriminative method for human pose estimation. This method learns a direct mapping from visual observations to human body configurations. The framework requires that the visual features should be powerful enough to discriminate the subtle differences between similar human poses. We propose to describe the image features using salient interest points that are represented by SIFT-like descriptors. The descriptor encode the position, appearance, and local structural information simultaneously. Bag-of-words representation is used to model the distribution of feature space. The descriptor can tolerate a range of illumination and position variations because it is computed on overlapped patches. We use Gaussian process regression to model the mapping from visual observations to human poses. This probabilistic regression algorithm is effective and robust to the pose estimation problem. We test our approach on the HumanEva data set. Experimental results demonstrate that our approach achieves the state of the art performance.
Xu Zhao 0001, Huazhong Ning, Yuncai Liu, Thomas S. Huang
ICPR2
2007 Searching Human Behaviors using Spatial-Temporalwords
abstract
This paper proposes an approach to searching human behaviors in videos using spatial-temporal words which are learnt from unlabelled data with various human behaviors through unsupervised learning. Both the query and the searched videos are represented by codewords frequencies, which capture the intrinsic information of motion and appearance of human behaviors. This representation further enables us to make use of integral histograms to accelerate the searching procedure. The performance also benefits from our feature representation that, through a MAX-like operation, may simulate the cortical equivalent of the machine-vision "window of analysis"(M. Riesenhuber and T. Poggio, 1999). Examples of challenging sequences with complex behaviors, including tennis and ballet, are shown.
Huazhong Ning, Yuxiao Hu 0001, Thomas S. Huang
ICIP (6)1
2007 Detecting Unsafe Driving Patterns using Discriminative Learning
abstract
We propose a discriminative learning approach for fusing multichannel sequential data with application to detect unsafe driving patterns from multi-channel driving recording data. The fusion is performed using a discriminatively trained graphical model -conditional random field (CRF). The proposed approach offers several advantage over existing information fusing approaches. First, it derives its classification power by directly modelling and maximizing the conditional probability. Second, it represents the variable dependency in an undirected graph, which is very efficient in inference. Third, it does not require to label all the training data and utilizes both labelled and unlabelled data efficiently by semi-supervised learning algorithms. The proposed approach is evaluated on driving recording data collected from driving simulator -STISIM. Experiments show it outperforms the simple discriminative classifier (SVM) and generative model (HMM).
Wei Xu 0007, Huazhong Ning, Yihong Gong, Thomas S. Huang
ICME3
2007 Incremental Spectral Clustering With Application to Monitoring of Evolving Blog Communities
abstract
In recent years, spectral clustering method has gained attentions because of its superior performance compared to other traditional clustering algorithms such as K-means algorithm. The existing spectral clustering algorithms are all off-line algorithms, i.e., they can not incrementally update the clustering result given a small change of the data set. However, the capability of incrementally updating is essential to some applications such as real time monitoring of the evolving communities of websphere or blogsphere. Unlike traditional stream data, these applications require incremental algorithms to handle not only insertion/deletion of data points but also similarity changes between existing items. This paper extends the standard spectral clustering to such evolving data by introducing the incidence vector/matrix to represent two kinds of dynamics in the same framework and by incrementally updating the eigenvalue system. Our incremental algorithm, initialized by a standard spectral clustering, continuously and efficiently updates the eigenvalue system and generates instant cluster labels, as the data set is evolving. The algorithm is applied to a blog data set. Compared with recomputation of the solution by standard spectral clustering, it achieves similar accuracy but with much lower computational cost. Close inspection into the blog content shows that the incremental approach can discover not only the stable blog communities but also the evolution of the individual multi-topic blogs.
Huazhong Ning, Wei Xu 0007, Yun Chi, Yihong Gong, Thomas S. Huang
SDM1
2006 Efficient Nonparametric Belief Propagation with Application to Articulated Body Tracking
abstract
An efficient Nonparametric Belief Propagation (NBP) algorithm is developed in this paper. While the recently proposed nonparametric belief propagation algorithm has wide applications such as articulated tracking [22, 19], superresolution [6], stereo vision and sensor calibration [10], the hardcore of the algorithm requires repeatedly sampling from products of mixture of Gaussians, which makes the algorithm computationally very expensive. To avoid the slow sampling process, we applied mixture Gaussian density approximation by mode propagation and kernel fitting [2, 7]. The products of mixture of Gaussians are approximated accurately by just a few mode propagation and kernel fitting steps, while the sampling method (e.g. Gibbs sampler) needs many samples to achieve similar approximation results. The proposed algorithm is then applied to articulated body tracking for several scenarios. The experimental results show the robustness and the efficiency of the proposed algorithm. The proposed efficient NBP algorithm also has potentials in other applications mentioned above.
Tony X. Han, Huazhong Ning, Thomas S. Huang
CVPR (1)2
2006 Improving Speaker Diarization by Cross EM Refinement
abstract
In this paper, we present a new speaker diarization system that improves the accuracy of traditional hierarchical clustering-based methods with little increase in computational cost. Our contributions are mainly two fold. First, we include a preprocessing called "local clustering" before the hierarchical clustering algorithm to merge very similar adjacent speech segments. This local clustering aims to reduce the number of segments to be clustered by the hierarchical clustering, so as to dramatically increase the processing speed. Second, we perform a postprocessing called "cross EM refinement" to purify the clusters generated by the hierarchical clustering. This algorithm is based on the idea of cross validation and EM algorithm. Our experimental evaluations show that the proposed cross EM refinement approach reduces the speaker diarization error by up to 56%, with an average reduction of 22% compared to the traditional hierarchical clustering method.
Huazhong Ning, Wei Xu 0007, Yihong Gong, Thomas S. Huang
ICME1
2006 A novel framework of text-independent speaker verification based on utterance transform and iterative cohort modeling
abstract
A novel framework for text-independent speaker verification is proposed. The framework is based on a new interpretation of Universal Background Model. The UBM in our framework actually defines a transform which maps the variable length observation into a fixed dimensional supervector(supervector space). Each speech utterance is then mapped into a point in this supervector space. The similarity measure in this vector space is progressively refined via an iterative cohort modeling scheme. The experiments on NIST 2002 corpus show the effectiveness of this new framework. Overall the EER drops from the baseline system(with T-Norm) 9.21 % to final improved system(without T-Norm) 8.07%. The new framework can effectively reduce the data dependence in the final output score which is clearly indicated in the second sets of experiments. The EER after T-Norm of final system marginally increases by relatively 1.73 % compared to the EER of baseline system drops 16.12 % relatively after T-Norm. Also, the relative improvement of DCF after T-Norm is marginal for the final improved system (2.47%) compared to 33.68 % in baseline system. It clear shows that the iterative cohort modeling effectively reduce the data dependence of the final scores, so that T-Norm will not further improve the system performance. Also, the performance of novel frame clearly increases as the iteration grows which suggest that the framework progressively refine the similarity measure on the supervector space with the iterative cohort modeling. Index Terms: speaker verification, utterance transform, iterative cohort modeling.
Ming Liu 0009, Huazhong Ning, Thomas S. Huang, Zhengyou Zhang
INTERSPEECH2
2006 A spectral clustering approach to speaker diarization
abstract
In this paper, we present a spectral clustering approach to explore the possibility of discovering structure from audio data. To apply the Ng-Jordan-Weiss (NJW) spectral clustering algorithm to speaker diarization, we propose some domain specific solutions to the open issues of this algorithm: choice of metric; selection of scaling parameter; estimation of the number of clusters. Then, a postprocessing step – “Cross EM refinement ” – is conducted to further improve the performance of spectral learning. In experiments, this approach has performance very similar to the traditional hierarchical clustering on the audio data of Japanese Parliament Panel Discussions, but it runs much faster than the latter.
Huazhong Ning, Ming Liu 0009, Hao Tang 0001, Thomas S. Huang
INTERSPEECH1
2004 Kinematics-based tracking of human walking in monocular video sequences
Huazhong Ning, Tieniu Tan, Liang Wang 0001, Weiming Hu 0004
Image Vis. Comput.1
2004 People tracking based on motion model and motion constraints with automatic initialization
Huazhong Ning, Tieniu Tan, Liang Wang 0001, Weiming Hu 0004
Pattern Recognit.1
2004 Fusion of static and dynamic body biometrics for gait recognition
abstract
Vision-based human identification at a distance has recently gained growing interest from computer vision researchers. This paper describes a human recognition algorithm by combining static and dynamic body biometrics. For each sequence involving a walker, temporal pose changes of the segmented moving silhouettes are represented as an associated sequence of complex vector configurations and are then analyzed using the Procrustes shape analysis method to obtain a compact appearance representation, called static information of body. In addition, a model-based approach is presented under a Condensation framework to track the walker and to further recover joint-angle trajectories of lower limbs, called dynamic information of gait. Both static and dynamic cues obtained from walking video may be independently used for recognition using the nearest exemplar classifier. They are fused on the decision level using different combinations of rules to improve the performance of both identification and verification. Experimental results of a dataset including 20 subjects demonstrate the feasibility of the proposed algorithm.
Liang Wang 0001, Huazhong Ning, Tieniu Tan, Weiming Hu 0004
IEEE Trans. Circuits Syst. Video Technol.2
2003 Fusion of Static and Dynamic Body Biometrics for Gait Recognition
abstract
Human identification at a distance has recently gained growing interest from computer vision researchers. This paper aims to propose a visual recognition algorithm based upon fusion of static and dynamic body biometrics. For each sequence involving a walking figure, pose changes of the segmented moving silhouettes are represented as an associated sequence of complex vector configurations, and are then analyzed using the Procrustes shape analysis method to obtain a compact appearance representation, called static information of body. Also, a model-based approach is presented under a condensation framework to track the walker and to recover joint-angle trajectories of lower limbs, called dynamic information of gait. Both static and dynamic cues are respectively used for recognition using the nearest exemplar classifier. They are also effectively fused on decision level using different combination rules to improve the performance of both identification and verification. Experimental results on a dataset including 20 subjects demonstrate the validity of the proposed algorithm.
Liang Wang 0001, Huazhong Ning, Tieniu Tan, Weiming Hu 0004
ICCV2
2003 Silhouette Analysis-Based Gait Recognition for Human Identification
abstract
Human identification at a distance has recently gained growing interest from computer vision researchers. Gait recognition aims essentially to address this problem by identifying people based on the way they walk. In this paper, a simple but efficient gait recognition algorithm using spatial-temporal silhouette analysis is proposed. For each image sequence, a background subtraction algorithm and a simple correspondence procedure are first used to segment and track the moving silhouettes of a walking figure. Then, eigenspace transformation based on principal component analysis (PCA) is applied to time-varying distance signals derived from a sequence of silhouette images to reduce the dimensionality of the input feature space. Supervised pattern classification techniques are finally performed in the lower-dimensional eigenspace for recognition. This method implicitly captures the structural and transitional characteristics of gait. Extensive experimental results on outdoor image sequences demonstrate that the proposed algorithm has an encouraging recognition performance with relatively low computational cost.
Liang Wang 0001, Tieniu Tan, Huazhong Ning, Weiming Hu 0004
IEEE Trans. Pattern Anal. Mach. Intell.3
2003 Automatic gait recognition based on statistical shape analysis
abstract
Gait recognition has recently gained significant attention from computer vision researchers. This interest is strongly motivated by the need for automated person identification systems at a distance in visual surveillance and monitoring applications. The paper proposes a simple and efficient automatic gait recognition algorithm using statistical shape analysis. For each image sequence, an improved background subtraction procedure is used to extract moving silhouettes of a walking figure from the background. Temporal changes of the detected silhouettes are then represented as an associated sequence of complex vector configurations in a common coordinate frame, and are further analyzed using the Procrustes shape analysis method to obtain mean shape as gait signature. Supervised pattern classification techniques, based on the full Procrustes distance measure, are adopted for recognition. This method does not directly analyze the dynamics of gait, but implicitly uses the action of walking to capture the structural characteristics of gait, especially the shape cues of body biometrics. The algorithm is tested on a database consisting of 240 sequences from 20 different subjects walking at 3 viewing angles in an outdoor environment. Experimental results are included to demonstrate the encouraging performance of the proposed algorithm.
Liang Wang 0001, Tieniu Tan, Weiming Hu 0004, Huazhong Ning
IEEE Trans. Image Process.4
2002 Gait recognition based on Procrustes shape analysis
abstract
Gait recognition has recently attracted increasing attention, especially in vision-based human identification-at-a-distance in visual surveillance. The paper proposes a simple but efficient gait recognition algorithm, based on statistical shape analysis. For each gait sequence, a background subtraction procedure is used to segment spatial silhouettes of the walking figures from the background. Static pose changes of these silhouettes over time are represented as a sequence of associated complex configurations in a common coordinate, and are then analyzed using the Procrustes shape analysis method to obtain a gait signature. The k-nearest neighbor classifier and the nearest exemplar classifier based on the full Procrustes distance measure are adopted for recognition. Experimental results demonstrate that the proposed algorithm has an encouraging recognition performance.
Liang Wang 0001, Huazhong Ning, Weiming Hu 0004, Tieniu Tan
ICIP (3)2
2002 Articulated Model Based People Tracking Using Motion Models
abstract
This paper focuses on acquisition of human motion data such as joint angles and velocity for applications of virtual reality, using both an articulated body model and a motion model in the CONDENSATION framework. Firstly, we learn a motion model represented by Gaussian distributions, and explore motion constraints by considering the dependency of motion parameters and represent them as conditional distributions. Both are integrated into the dynamic model to concentrate factored sampling in the areas of state-space with most posterior information. To measure the observing density with accuracy and robustness, a PEF (pose evaluation function) modeled with a radial term is proposed. We also address the issue of automatic acquisition of initial model posture and recovery from severe failures. A large number of experiments on several persons demonstrate that our approach works well.
Huazhong Ning, Liang Wang 0001, Weiming Hu 0004, Tieniu Tan
ICMI1