Guoliang Fan 0001

dblp:369/5654-1 · DBLP profile ↗
← Back
32ranked-venue papers
0as first author
9since 2021 · last 2024
0000-0002-8584-9040ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 7 since 2021Artificial intelligence and machine learning · 7Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 A Review of Depth-Based Human Motion Enhancement: Past and Present
abstract
In this article, we survey the current research trends of enhancement and denoising of depth-based motion capture data (D-Mocap) and also discuss possible future research issues. We first present the commonly used problem formulation for human motion enhancement. We then review related work and cover a broad set of methodologies including filtering based, learning based, and evolutionary based approaches. In addition, we present some important experiments-related issues, such as data creation or collection, reference data generation, and the metrics used for performance evaluation. It is our intent to provide a comprehensive tutorial and survey on the recent efforts on D-Mocap improvement, both methodologically and experimentally. By comparing the state-of-the-art methods, we also propose future research needs that could make D-Mocap more useful and relevant for real-world clinical applications.
Nate Lannan, Guoliang Fan 0001
IEEE J. Biomed. Health Informatics3
2023 A Transfer Learning-Based Smart Homecare Assistive Technology to Support Activities of Daily Living for People with Mild Dementia
abstract
People with dementia (PwD) experience widespread cognitive impairment, leading to challenges in carrying out routine tasks known as activities of daily living (ADLs) and especially more complex instrumental ADLs (IADLs). Without adequate support, PwD becomes highly susceptible to a loss of independence and vulnerability. Various assistive technologies (ATs) have been proposed to assist PwD with mild conditions to complete some IADL tasks independently. However, most existing AT devices provide limited IADL support to PwD with little or no user-specific customization. This paper presents our early development of a new homecare assistive tool for PwD with mild conditions, called CATcare (Cognitive Assistive Technology Care). Our CATcare system can be operated and configured on a smartphone or smart glasses and further customized by a caregiver according to the specific IADL needs of a care recipient. It can provide IADL-specific step-by-step cueing, prompting, and timely feedback to assist PwD in accomplishing the task of interest. Our research leverages the capabilities of transfer learning from AI models in indoor localization, object detection, and Natural Language Processing (NLP) and shows the potential and promise to develop a customizable and personalizable CATcare tool. This CATcare system is intended to improve the life quality of PwD and reduce the burden of their caregivers.
Xiaowei Chen 0005, Guoliang Fan 0001, Emily Roberts, Steven Howell Jr.
BIBE2
2023 Indoor Camera Pose Estimation From Room Layouts and Image Outer Corners
abstract
To support indoor scene understanding, room layouts have been recently introduced that define a few typical space configurations according to junctions and boundary lines. In this paper, we study camera pose estimation from eight common room layouts with at least two boundary lines that is cast as a PnL (Perspective-n-Line) problem. Specifically, the intersecting points between image borders and room layout boundaries, named image outer corners (IOCs), are introduced to create additional auxiliary lines for PnL optimization. Therefore, a new PnL-IOC algorithm is proposed which has two implementations according to the room layout types. The first one considers six layouts with more than two boundary lines where 3D correspondence estimation of IOCs creates sufficient line correspondences for camera pose estimation. The second one is an extended version to handle two challenging layouts with only two coplanar boundaries where correspondence estimation of IOCs is ill-posed due to insufficient conditions. Thus the powerful NSGA-II algorithm is embedded in PnL-IOC to estimate the correspondences of IOCs. At the last step, the camera pose is jointly optimized with 3D correspondence refinement of IOCs in the iterative Gauss-Newton algorithm. Experiment results on both simulated and real images show the advantages of the proposed PnL-IOC method on the accuracy and robustness of camera pose estimation from eight different room layouts over the existing PnL methods. The code is available at https://github.com/XiaoweiChenOSU/PnL-IOC.
Xiaowei Chen 0005, Guoliang Fan 0001
IEEE Trans. Multim.2
2022 Holistic indoor scene understanding by context-supported instance segmentation
Guoliang Fan 0001
Multim. Tools Appl.2
2022 ManhattanFusion: Online Dense Reconstruction of Indoor Scenes From Depth Sequences
abstract
We present a new framework for online dense 3D reconstruction of indoor scenes by using only depth sequences. This research is particularly useful in cases with a poor light condition or in a nearly featureless indoor environment. The lack of RGB information makes long-range camera pose estimation difficult in a large indoor environment. The key idea of our research is to take advantage of the geometric prior of Manhattan scenes in each stage of the reconstruction pipeline with the specific aim to reduce the cumulative registration error and overall odometry drift in a long sequence. This idea is further boosted by local Manhattan frame growing and the local-to-global strategy that leads to implicit loop closure handling for a large indoor scene. Our proposed pipeline, namely ManhattanFusion, starts with planar alignment and local pose optimization where the Manhattan constraints are imposed to create detailed local segments. These segments preserve intrinsic scene geometry by minimizing the odometry drift even under complex and long trajectories. The final model is generated by integrating all local segments into a global volumetric representation under the constraint of Manhattan frame-based registration across segments. Our algorithm outperforms others that use depth data only in terms of both the mean distance error and the absolute trajectory error, and it is also very competitive compared with RGB-D based reconstruction algorithms. Moreover, our algorithm outperforms the state-of-the-art in terms of the surface area coverage by 10-40 percent, largely due to the usefulness and effectiveness of the Manhattan assumption through the reconstruction pipeline.
Mahdi Yazdanpour, Guoliang Fan 0001, Weihua Sheng
IEEE Trans. Vis. Comput. Graph.2
2021 An Effective Sharpness Assessment Method For Shallow Depth-Of-Field Images
abstract
No-reference (NR) image sharpness assessment is an important issue for image quality assessment and algorithm performance evaluation. Many objective NR sharpness assessment metrics have been proposed which are often intended to be strongly associated with the human visual system (HVS). However, recent studies show that common sharpness assessment indicators may misjudge the degree of blurring for images with shallow depth of field that are often used to highlight the main subject in the view. This paper proposes an efficient no-reference objective image sharpness assessment metric based on the product of bidirectional pixel intensity differences that is computed block-by-block (PBDB). This paper contributes the following: (1) the sharpness of shallow depth-of-field images can be accurately evaluated with the proposed algorithm when traditional methods do not work well. (2) Experimental results on three public datasets demonstrate competitiveness and effectiveness of the proposed algorithm when compared with several state-of-the-art methods.
Zhixiang Duan, Guangxin Li, Guoliang Fan 0001
ICIP3
2021 Locop: Local Collaborative Object Presence For Semantic Labeling Via Score Map Re-Inference
abstract
Recent research has focused on end-to-end networks for indoor scene semantic labeling. However, in addition to learning bottom-up features, high-level knowledge could be implemented to guide the local classification. In this paper, we take advantage of trained semantic labeling networks by using the intermediate layer output as a per-category local detector and implement the context information in a network structure to boost the semantic segmentation performance. A deep learning-based re-inferencing frame work is proposed to boost any pixel-level labeling outputs using our local collaborative object presence (LoCOP) feature as the global-to-local guidance. Experimental results show that the detection accuracy is improved with our re-inference approach.
Guoliang Fan 0001
ICIP2
2021 Human Motion Enhancement via Tobit Particle Filtering and Differential Evolution
abstract
This paper proposes a novel approach to improve the quality of human motion data captured by a depth sensor. Depth-based motion capture (D-Mocap) data often suffer significant errors due to noise, self-occlusion, interference, and other algorithmic limitations. We aim to improve 3D joint trajectories to be more kinematically admissible and anthropometrically consistent. The Tobit model is incorporated with a particle filter (TPF) to handle censored measurements. We also embed the DE algorithm in the TPF, which allows particles to be re-distributed and re-weighted according to bone length consistency before re-sampling. This integration leads to a new TPF-DE algorithm that harmoniously takes advantage of kinematic and anthropometric constraints. We compare our methods with several nonlinear Kalman filters and deep learning-based methods to demonstrate the efficacy of TPF-DE on both simulated and real-world D-Mocap data.
Nate Lannan, Guoliang Fan 0001
ICIP3
2021 DrsNet: Dual-resolution semantic segmentation with rare class-oriented superpixel prior
Liangjiang Yu, Guoliang Fan 0001
Multim. Tools Appl.2
2020 Human Motion Enhancement Using Nonlinear Kalman Filter Assisted Convolutional Autoencoders
abstract
Human motion analysis is integral to many applications ranging from biomedicine to surveillance and 3D animation. The availability of consumer RGB-D sensors makes motion capture (Mocap) more prevalent and accessible in our daily lives. However, depth-based Mocap (D-Mocap) suffers significant errors and noise due to the limitation of depth sensing, self-occlusion, and many other problems. We present a novel filter-assisted deep learning approach to improve low-quality human motion data (e.g., D-Mocap) by taking advantage of the recent progress in both deep learning and nonlinear Kalman filtering. At the heart of this method is a learned motion manifold through the use of a convolutional autoencoder trained on high-quality, rich-variety CMU Mocap data which is used to recover valid human motion from corrupted input. Furthermore, the Tobit Kalman filter (TKF), proposed to handle censored measurements, is used to assist the autoencoder with more kinematic and dynamic constraints. Two structural paradigms are investigated to handle different kinds of data errors by maximizing the synergy between the two integral parts in this work. The experimental results on both simulated and real-world human motion data demonstrate the effectiveness and robustness of the proposed methods to improve the quality of noisy and erroneous Mocap data.
Nate Lannan, Guoliang Fan 0001, Jerome G. Hausselle
BIBE3
2020 A Hybrid Approach to Human Motion Enhancement under Kinematic and Anthropometric Constraints
abstract
This paper proposes a novel approach to improve the quality of human motion data captured by a depth sensor. Depth-based motion capture (D-Mocap) data suffers significant errors due to noise, self-occlusion, interference, and other algorithmic limitations. We aim to smooth and correct 3D joint trajectories to be more kinematically admissible and anthropometrically stable. Our research synergistically integrates advanced nonlinear Kalman filters (KF) with Differential Evolution (DE) algorithms into one computational flow to take advantage of the kinematic and arthrometric constraints respectively. Specifically, we compare three nonlinear KFs in terms of their effectiveness and accuracy, including the Extended Kalman filter (EKF), the Unscented Kalman filter (UKF), and the Tobit Kalman filter (TKF). Two sets of motion capture datasets are used in this research to examine the proposed algorithms in both simulated and real-world data. Furthermore, we demonstrate significant improvements in six joint angles that are often used for clinical gait assessment.
Nate Lannan, Guoliang Fan 0001, Jerome G. Hausselle
BIBE3
2019 Learning deep transmission network for efficient image dehazing
Zhigang Ling, Guoliang Fan 0001, Jianwei Gong
Multim. Tools Appl.2
2019 Topology-aware non-rigid point set registration via global-local topology preservation
Guoliang Fan 0001
Mach. Vis. Appl.2
2019 Semantics-enhanced supervised deep autoencoder for depth image-based 3D model retrieval
Ayesha Siddiqua, Guoliang Fan 0001
Pattern Recognit. Lett.2
2018 Joint optimization and perceptual boosting of global and local contrast for efficient contrast enhancement
Zhigang Ling, Guoliang Fan 0001, Yan Liang 0001, Junyi Zuo
Multim. Tools Appl.2
2018 Optimal Transmission Estimation via Fog Density Perception for Efficient Single Image Defogging
abstract
Single image defogging algorithms based on prior assumptions or constraints have captured much attention because of their simplicity and practicality. However, they still have some challenges to deal with foggy images captured under weather conditions where these assumptions or constraints may not be effective or efficient enough. In this paper, we aim to develop a novel image defogging algorithm by directly predicting the fog density of recovered images rather than adopting prior assumptions or constraints. In order to achieve this goal, two specific steps are introduced. First, we adopt three fog-relevant statistical features derived from foggy images, and further develop a simple fog density evaluator (SFDE) by creating a linear combination of these fog-relevant features. This proposed evaluator can efficiently perceive the fog density of a single image without reference to a corresponding fog-free image and has a low computational load compared with an existing method. Second, a physics-based mathematical relationship between the transmission and the fog density score of the recovered image is developed via SFDE, thus image defogging can be posed as a minimization problem on the fog density score of the recovered image. As a result, two optimal transmission models, called an optimal transmission model via SFDE (OTSFDE) and a simpler optimal transmission models via SFDE (SOTSFDE), are present to determine the key transmission map for efficient fog removal. Compared to OTSFDE, SOTSFDE has low computational complexity with slight performance degradation. Experimental results demonstrate that the proposed algorithms can effectively remove fog and are not confined by any assumptions or constraints, both quantitatively and qualitatively, compared with some existing algorithms.
Zhigang Ling, Jianwei Gong, Guoliang Fan 0001, Xiao Lu 0002
IEEE Trans. Multim.3
2017 Perception oriented transmission estimation for high quality image dehazing
Zhigang Ling, Guoliang Fan 0001, Jianwei Gong, Yaonan Wang 0001, Xiao Lu 0002
Neurocomputing2
2017 Video-Based Human Walking Estimation Using Joint Gait and Pose Manifolds
abstract
We study two fundamental issues about video-based human walking estimation, where the goal is to estimate 3D gait kinematics (i.e., joint positions) from 2D gait appearances (i.e., silhouettes). One is how to model the gait kinematics from different walking styles, and the other is how to represent the gait appearances captured under different views and from individuals of distinct walking styles and body shapes. Our research is conducted in three steps. First, we propose the idea of joint gait-pose manifold (JGPM), which represents gait kinematics by coupling two nonlinear variables, pose (a specific walking stage) and gait (a particular walking style) in a unified latent space. We extend the Gaussian process latent variable model (GPLVM) for JGPM learning, where two heuristic topological priors, a torus and a cylinder, are considered and several JGPMs of different degrees of freedom (DoFs) are introduced for comparative analysis. Second, we develop a validation technique and a series of benchmark tests to evaluate multiple JGPMs and recent GPLVMs in terms of their performance for gait motion modeling. It is shown that the toroidal prior is slightly better than the cylindrical one, and the JGPM of 4 DoFs that balances the toroidal prior with the intrinsic data structure achieves the best performance. Third, a JGPM-based visual gait generative model (JGPM-VGGM) is developed, where JGPM plays a central role to bridge the gap between the gait appearances and the gait kinematics. Our proposed JGPM-VGGM is learned from Carnegie Mellon University MoCap data and tested on the HumanEva-I and HumanEva-II data sets. Our experimental results demonstrate the effectiveness and competitiveness of our algorithms compared with existing algorithms.
Xin Zhang 0013, Meng Ding 0001, Guoliang Fan 0001
IEEE Trans. Circuits Syst. Video Technol.3
2016 Learning deep transmission network for single image dehazing
abstract
State-of-the-art single image dehazing algorithms have some challenges to deal with images captured under complex weather conditions because their assumptions usually do not hold in those situations. In this paper, we develop a deep transmission network for robust single image dehazing. This deep transmission network simultaneously copes with three color channels and local patch information to automatically explore and exploit haze-relevant features in a learning framework. We further explore different network structures and parameter settings to achieve tradeoffs between performance and speed, which shows that color channels information is the most useful haze-relevant feature rather than local information. Experiment results demonstrate that the proposed algorithm outperforms state-of-the-art methods on both synthetic and real-world datasets.
Zhigang Ling, Guoliang Fan 0001, Yaonan Wang 0001, Xiao Lu 0002
ICIP2
2016 Articulated and Generalized Gaussian Kernel Correlation for Human Pose Estimation
abstract
In this paper, we propose an articulated and generalized Gaussian kernel correlation (GKC)-based framework for human pose estimation. We first derive a unified GKC representation that generalizes the previous sum of Gaussians (SoG)-based methods for the similarity measure between a template and an observation both of which are represented by various SoG variants. Then, we develop an articulated GKC (AGKC) by integrating a kinematic skeleton in a multivariate SoG template that supports subject-specific shape modeling and articulated pose estimation for both the full body and the hands. We further propose a sequential (body/hand) pose tracking algorithm by incorporating three regularization terms in the AGKC function, including visibility, intersection penalty, and pose continuity. Our tracking algorithm is simple yet effective and computationally efficient. We evaluate our algorithm on two benchmark depth data sets. The experimental results are promising and competitive when compared with the state-of-the-art algorithms.
Meng Ding 0001, Guoliang Fan 0001
IEEE Trans. Image Process.2
2015 Multimodal topic modeling based geo-annotation for social event detection in large photo collections
abstract
Multimodal image clustering becomes an effective approach for social event detection in large photo collections. In addition to visual and textual information, geographic information can also be used to improve the detection accuracy of social events. However, not every image in a photo collection is tagged with geographic information. A topic model based approach is proposed to estimate missing geographic information in a photo which involves a supervised multimodal model to estimated the joint distribution of time, geographic, content, and textual information for a large set of photos. The photos without geographic information are annotated with a predicted geographic coordinate. We show the efficacy of the proposed approach for event detection and annotation from a large photo collection.
Bin Xu 0010, Guoliang Fan 0001
ICIP2
2015 Generalized Sum of Gaussians for Real-Time Human Pose Tracking from a Single Depth Sensor
abstract
We propose a generalized Sum-of-Gaussians (G-SoG) model for statistical 3D shape modeling that is applied to human pose tracking from a single depth sensor. G-SoG generalizes the original SoG model by involving much fewer anisotropic Gaussians yet with better flexibility and adaptability. Both SoG and G-SoG are involved for pose tracking with different roles, where the former one is used to represent observed point cloud data through an efficient Octree partitioning, and the latter one is embedded with a quaternion-based articulated skeleton to create a standard human template model. We derive a differentiable similarity function between SoG and G-SoG that can be optimized analytically not only to learn a subject-specific articulated model but also to support sequential pose tracking where two additional terms (visibility and continuity) are also involved. Our algorithm is simple yet effective and can achieve real-time performance. The experimental results on a public depth dataset are promising and competitive when compared with state-of-the-art algorithms.
Meng Ding 0001, Guoliang Fan 0001
WACV2
2015 Multilayer Joint Gait-Pose Manifolds for Human Gait Motion Modeling
abstract
We present new multilayer joint gait-pose manifolds (multilayer JGPMs) for complex human gait motion modeling, where three latent variables are defined jointly in a low-dimensional manifold to represent a variety of body configurations. Specifically, the pose variable (along the pose manifold) denotes a specific stage in a walking cycle; the gait variable (along the gait manifold) represents different walking styles; and the linear scale variable characterizes the maximum stride in a walking cycle. We discuss two kinds of topological priors for coupling the pose and gait manifolds, i.e., cylindrical and toroidal, to examine their effectiveness and suitability for motion modeling. We resort to a topologically-constrained Gaussian process (GP) latent variable model to learn the multilayer JGPMs where two new techniques are introduced to facilitate model learning under limited training data. First is training data diversification that creates a set of simulated motion data with different strides. Second is the topology-aware local learning to speed up model learning by taking advantage of the local topological structure. The experimental results on the Carnegie Mellon University motion capture data demonstrate the advantages of our proposed multilayer models over several existing GP-based motion models in terms of the overall performance of human gait motion modeling.
Meng Ding 0001, Guoliang Fan 0001
IEEE Trans. Cybern.2
2014 Joint view-identity manifold for infrared target tracking and recognition
Jiulu Gong, Guoliang Fan 0001, Liangjiang Yu, Joseph P. Havlicek, Derong Chen, Ningjun Fan
Comput. Vis. Image Underst.2
2013 Infrared target tracking, recognition and segmentation using shape-aware level set
abstract
A new probabilistic model called ATR-Seg for automated target tracking, recognition and segmentation is proposed that incorporates a shape constrained level set with a shape generative model along with motion model. The shape model involves a view-independent identity manifold and infinite identity-dependent view manifolds for multi-view and multi-target shape modeling. ATR-Seg applies the motion model to predict the state of the target (i.e., 3D position, pose and identity), and then uses a shape-aware level set energy functional to evaluate the tracking and segmentation results. A particle filtering-based method is used for sequential inference, where the level set energy functional is treated as the likelihood function. Experimental results obtained against the SENSIAC ATR database demonstrate the advantages of the proposed method compared with the two recent techniques that require target pre-segmentation via background subtraction.
Jiulu Gong, Guoliang Fan 0001, Joseph P. Havlicek, Ningjun Fan, Derong Chen
ICIP2
2013 Simultaneous target recognition, segmentation and pose estimation
abstract
We propose a simultaneous target recognition, segmentation and pose estimation algorithm for the infrared ATR task. A probabilistic framework of level set segmentation is extended by incorporating a shape generative model that provides a multi-class and multiview shape prior. This generative model involves a couplet of a view manifold and an identity manifold for general shape modeling. Then an energy function from the probabilistic level set formulation can be iteratively optimized by a shape-constrained variational method. Due to the fact that both the view and identity variables are explicitly involved in the level set optimization, the proposed method is able to accomplish recognition, segmentation, and pose estimation. Experimental results show that the proposed method outperforms two traditional methods where target recognition and pose estimation are implemented after segmentation.
Liangjiang Yu, Guoliang Fan 0001, Jiulu Gong, Joseph P. Havlicek
ICIP2
2013 Two-layer dual gait generative models for human motion estimation from a single camera
Xin Zhang 0013, Guoliang Fan 0001, Li-Shan Chou
Image Vis. Comput.2
2012 Structure-guided manifold learning for video-based motion estimation
abstract
We present a new structure-guided joint gait pose manifold (JGPM) that represents gait kinematics by two variables. One is the pose to denote a series of stages in a walking cycle and the other is the gait to reflect the individual walking styles. Coupling pose and gait variables in the same latent space, such as a torus-like JGPM, was shown promising and effective for video-based motion estimation. However, the two-step learning used in torus-like JGPM is computationally expensive and it separates the optimization of pose and gait variables. This work overcomes the limitations of the previous method by developing a new structure-guided JGPM that is able to jointly optimize four variables in the same latent space, leading to a much compact parameter set while sustaining a comparable performance on video-based motion estimation, as well as a great potential for large-scale learning.
Meng Ding 0001, Guoliang Fan 0001, Xin Zhang 0013, Li-Shan Chou
ICIP2
2012 Joint view-identity manifold for target tracking and recognition
abstract
A new joint view-identity manifold (JVIM) is proposed for multiview shape modeling that is applied to automated target tracking and recognition (ATR). This work improves our recent work where the view and identity manifolds are assumed to be independent for multi-view multi-target modeling. A local linear Gaussian process latent variable model (LL-GPLVM) is used to learn a probabilistic JVIM which can capture both inter-class and intra-class variability of 2D target shapes under arbitrary view point jointly in one coexisted latent space. A particle filter-based ATR algorithm is developed to simultaneously infer the view and identity parameters along JVIM so that target tracking and recognition can be achieved jointly in a seamlessly fashion. The experimental results using SENSIAC ATR database demonstrate the advantages of our method both qualitatively and quantitatively compared with existing methods using template matching or separate view and identity manifolds.
Jiulu Gong, Guoliang Fan 0001, Liangjiang Yu, Joseph P. Havlicek, Derong Chen
ICIP2
2012 Adaptive Kalman Filtering for Histogram-Based Appearance Learning in Infrared Imagery
abstract
Targets of interest in video acquired from imaging infrared sensors often exhibit profound appearance variations due to a variety of factors, including complex target maneuvers, ego-motion of the sensor platform, background clutter, etc., making it difficult to maintain a reliable detection process and track lock over extended time periods. Two key issues in overcoming this problem are how to represent the target and how to learn its appearance online. In this paper, we adopt a recent appearance model that estimates the pixel intensity histograms as well as the distribution of local standard deviations in both the foreground and background regions for robust target representation. Appearance learning is then cast as an adaptive Kalman filtering problem where the process and measurement noise variances are both unknown. We formulate this problem using both covariance matching and, for the first time in a visual tracking application, the recent autocovariance least-squares (ALS) method. Although convergence of the ALS algorithm is guaranteed only for the case of globally wide sense stationary process and measurement noises, we demonstrate for the first time that the technique can often be applied with great effectiveness under the much weaker assumption of piecewise stationarity. The performance advantages of the ALS method relative to the classical covariance matching are illustrated by means of simulated stationary and nonstationary systems. Against real data, our results show that the ALS-based algorithm outperforms the covariance matching as well as the traditional histogram similarity-based methods, achieving sub-pixel tracking accuracy against the well-known AMCOM closure sequences and the recent SENSIAC automatic target recognition dataset.
Vijay Venkataraman, Guoliang Fan 0001, Joseph P. Havlicek, Xin Fan 0001, Yan Zhai, Mark B. Yeary
IEEE Trans. Image Process.2
2010 Dual Gait Generative Models for Human Motion Estimation From a Single Camera
abstract
This paper presents a general gait representation framework for video-based human motion estimation. Specifically, we want to estimate the kinematics of an unknown gait from image sequences taken by a single camera. This approach involves two generative models, called the kinematic gait generative model (KGGM) and the visual gait generative model (VGGM), which represent the kinematics and appearances of a gait by a few latent variables, respectively. The concept of gait manifold is proposed to capture the gait variability among different individuals by which KGGM and VGGM can be integrated together, so that a new gait with unknown kinematics can be inferred from gait appearances via KGGM and VGGM. Moreover, a new particle-filtering algorithm is proposed for dynamic gait estimation, which is embedded with a segmental jump-diffusion Markov Chain Monte Carlo scheme to accommodate the gait variability in a long observed sequence. The proposed algorithm is trained from the Carnegie Mellon University (CMU) Mocap data and tested on the Brown University HumanEva data with promising results.
Xin Zhang 0013, Guoliang Fan 0001
IEEE Trans. Syst. Man Cybern. Part B2
2008 Dual generative models for human motion estimation from an uncalibrated monocular camera
abstract
We propose a new approach to estimate gait kinematics from image sequences taken by a monocular uncalibrated camera. This approach involves two generative models for gait representations in the kinematic and visual spaces, which induce two gait manifolds that characterize the gait variability in terms of the kinematics and visual appearance. A manifold topology enforcement scheme is introduced to incorporate the two gait manifolds. Moreover, a new particle filtering algorithm is proposed for dynamic gait tracking and estimation where a segmental jump-diffusion Markov Chain Monte Carlo (MCMC) technique is developed to accommodate the dynamic nature of the gait variability. The proposed algorithm is trained from CMU Mocap data and tested on the HumanEva dataset with promising results.
Xin Zhang 0013, Guoliang Fan 0001
ICPR2