EDBT 2026 Demo / reviewers in the wild / expert
Worapan Kusakunniran
dblp:05/7576
· DBLP profile ↗
32ranked-venue papers
17as first author
12since 2021 · last 2026
0000-0002-2896-611XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 10 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Security and privacy · 2 · 2 first-authorComputer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Minutiae-based palm photo recognition using deep neural networks
Javad Khodadoust, Raúl Monroy, Miguel Angel Medina-Pérez, Emanuela Marasco, Worapan Kusakunniran |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Cross-modality video person Re-ID with modality-aware cosine-triplet lossabstractAbstract Person Re-Identification (Re-ID) is one of the important applications for surveillance. The scenarios where we need to identify the subjects captured by night-vision (infrared) cameras are a significant challenge to the existing Re-ID techniques, where only color footage is available for comparison. This is due to large differences in the composition between color and infrared images, which results in appearance-based information becoming less reliable for Re-ID. For this reason, we hypothesized that motion information from sequences of inputs is vital for cross-modality (visible-to-infrared) Re-ID. From our initial findings, motion information from the sequence of frames significantly improved the cross-modality Re-ID performance. In addition, choices of distance metrics (Euclidean vs. cosine) have a significant effect on the overall performance. As a result, the experimental performance on SYSU-MM01 reached 72.70% in mAP and 73.27% in rank-1 accuracy and yielded significant performance gains of 28.32% in mAP and 29.14% in rank-1 accuracy over our baseline. The performance competes with the existing state-of-the-art techniques tested on the same dataset. Rangwan Kasantikul, Worapan Kusakunniran |
Multim. Tools Appl. | 2 |
| 2026 | A multibiometric system based on finger photo and palm photo
Javad Khodadoust, Raúl Monroy, Miguel Angel Medina-Pérez, Worapan Kusakunniran, Ali Mohammad Khodadoust |
Multim. Tools Appl. | 4 |
| 2025 | Detection of translucent flesh disorder and automatic grading of mangosteens in multi-view imagesabstractAbstract In this paper, convolutional neural network (CNN)-based solutions are developed for grading assessment and flesh disorder detection of mangosteens in images. The grading is set to three classes of three quality levels based on the local market, where the data were collected. In addition, three flesh disorders/status are focused in this work, including translucent flesh disorder, gamboge, and rotten. Three types of solutions are attempted in this paper. The first solution relies on the well-known CNN architectures with the transfer learning and data augmentation. The second solution is developed based on the detection model, i.e., YOLOv8. The third solution is to design a new architecture by taking into account of human expert knowledge that is used for the manual grading and detection. Multiple views of each mangosteen must be considered simultaneously for the disorder detection. Four side views should be considered together, before looking at the top and bottom views. This is a very difficult task even for the human experts. The proposed solutions are trained and evaluated on the self-collected dataset of 206 mangosteens captured under six views (i.e., top view, bottom view, and four side views). The proposed solutions could achieve the perfect accuracy of 100% for the grading and up to 78% AUC for the disorder detection. Worapan Kusakunniran, Thanandon Imaromkul, Kittinun Aukkapinyo, Kittikhun Thongkanchorn, Pimpinan Somsong, Pimsiri Tiyayon |
Neural Comput. Appl. | 1 |
| 2024 | Deep Learning for Automatic Classification of Carotenoid Associated Color PigmentationabstractThis study explores the application of deep learning models, specifically ResNet-34, ResN et-50, and EfficientNet-B0, for the automatic classification of carotenoid-associated color pigmentation in tomatoes. The dataset comprises 250 images categorized into five pigmentation levels, reflecting the varying carotenoid content. Carotenoids, such as lycopene and beta-carotene, are key pigments influencing the color of tomatoes, with deeper reds and oranges indicating higher concentrations. The models were evaluated for direct classification and regression followed by classification. Results show that EfficientNet-B0 achieved the highest accuracy in direct classification (94.00%), while ResNet-34 excelled in regression tasks (91.33%). Future research will continue exploring regression tasks to predict actual carotenoid content in tomatoes, enhancing prediction accuracy and robustness. Kasidit Ruaydee, Worapan Kusakunniran, Warangkana Srichamnong |
TENCON | 2 |
| 2024 | Automatic classification of mangosteens and ripe status in images using deep learning based approaches
Worapan Kusakunniran, Thanandon Imaromkul, Kittinun Aukkapinyo, Kittikhun Thongkanchorn, Pimpinan Somsong |
Multim. Tools Appl. | 1 |
| 2024 | Deep-learning-based head pose estimation from a single RGB image and its application to medical CROM measurementabstractAbstract For human beings, neck movement will be degraded due to aging, trauma, musculoskeletal disorders, or degenerative diseases. Cervical range of motion (CROM) measurement is one of the popular quantitative neck examinations. Despite radiography is considered as the gold standard, it suffers from invasiveness, radiation exposure, and expensiveness. Recently, vision-based methods have been applied for CROM measurement but achieve large errors and require depth camera. On the other hand, deep neural networks provide good performances on head pose estimation (HPE) from a single image, thus promising for medical CROM measurement. We propose to use CNN networks to extract pyramidal or multi-level image features, which are passed to cross-level attention modules for feature fusion and then to a modified ASPP module and a multi-bin classification/regression module for spatial-channel attention and Euler angle conversion/prediction, respectively. The proposed technique was evaluated on public datasets, such as 300W_LP, AFLW2000, and BIWI, to verify its superior performances (with mean MAE = 3.50°, 3.40°, and 2.31° for different experimental protocols) than state-of-the-art methods. Our pre-trained model was also evaluated with our own collected dataset from hospital for CROM measurement. It also achieved the lowest MAE of 4.58° among other methods and conformed with a medical standard of 5 degrees except the pitch angle (which has a MAE of 5.70°, larger than the standard and the yaw (MAE = 3.60°) and roll angles (MAE = 4.44°)). In general, HPE technique is feasible for CROM measurement and shows its advantages of speed, non-invasiveness, free of anatomical landmark and low cost of operation. Panrasee Ritthipravat, Kittisak Chotikkakamthorn, Wen-Nung Lie, Worapan Kusakunniran, Pimchanok Tuakta, Paitoon Benjapornlert |
Multim. Tools Appl. | 4 |
| 2023 | Automated tongue segmentation using deep encoder-decoder model
Worapan Kusakunniran, Punyanuch Borwarnginn, Thanandon Imaromkul, Kittinun Aukkapinyo, Kittikhun Thongkanchorn, Disathon Wattanadhirach, Sophon Mongkolluksamee, Ratchainant Thammasudjarit, Panrasee Ritthipravat, Pimchanok Tuakta, Paitoon Benjapornlert |
Multim. Tools Appl. | 1 |
| 2023 | Improving Disentangled Representation Learning for Gait Recognition Using Group SupervisionabstractIn decades, gait has been gathering extensive interest for the advantage that it can be measured from a distance without physical contact. However, for image/video-based gait recognition, its performance can be remarkably influenced by exterior factors, such as viewing angles and clothing changes. Thus, in this paper, a group-supervised disentangled representation learning network is proposed for gait recognition to extract features invariant to these factors. First, sequences are explicitly disentangled into pose, gait, appearance, and view features through a generic encoder-decoder framework. To ensure the feature adaptability and independency, a disentanglement swap module is specifically adopted during our encode-decoder process through a series of swap operations based on the feature attributes. Following the feature disentanglement, a disentanglement aggregation module is also specially proposed for pose, gait, and appearance features to enhance their effectiveness. Finally, the enhanced three features are concatenated together for gait recognition. Relevant experiments certify that compared with other disentangled representation learning-based gait recognition methods, our proposed method enables to obtain a more excellent recognition result, despite fewer gait frames being utilized. Lingxiang Yao, Worapan Kusakunniran, Peng Zhang 0057, Qiang Wu 0001, Jian Zhang 0002 |
IEEE Trans. Multim. | 2 |
| 2022 | Collaborative Feature Learning for Gait Recognition Under Cloth ChangesabstractSince gait can be utilized to identify individuals from a far distance without their interaction and coordination, recently many gait recognition methods have been proposed. However, due to a real-world scenario of clothing changes, a degradation occurs for most of these methods. Thus in this paper, a more efficient gait recognition method is proposed to address the problem of clothing variances. First, part-based gait features are formulated from two different perspectives,i.e., the separated body parts that are more robust to clothing changes and the estimated human skeleton key-point regions. It is reasonable to formulate such features for cloth-changing gait recognition, because these two perspectives are both less vulnerable to clothing changes. Given that each feature has its own advantages and disadvantages, a more efficient gait feature is generated in this paper by assembling these two features together. Moreover, since local features are more discriminative than global features, in this paper more attention is focused on the local short-range features. Also, unlike most methods, in our method we treat the estimated key-point features as a set of word embeddings, and a transformer encoder is specifically used to learn the dependence of each correlative key-points. The robustness and effectiveness of our proposed method are certified by experiments on CASIA Gait Dataset B, and it has achieved the state-of-the-art performance on this dataset. Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jingsong Xu, Jian Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Recognizing Gaits Across Walking and Running SpeedsabstractFor decades, very few methods were proposed for cross-mode (i.e., walking vs. running) gait recognition. Thus, it remains largely unexplored regarding how to recognize persons by the way they walk and run. Existing cross-mode methods handle the walking-versus-running problem in two ways, either by exploring the generic mapping relation between walking and running modes or by extracting gait features which are non-/less vulnerable to the changes across these two modes. However, for the first approach, a mapping relation fit for one person may not be applicable to another person. There is no generic mapping relation given that walking and running are two highly self-related motions. The second approach does not give more attention to the disparity between walking and running modes, since mode labels are not involved in their feature learning processes. Distinct from these existing cross-mode methods, in our method, mode labels are used in the feature learning process, and a mode-invariant gait descriptor is hybridized for cross-mode gait recognition to handle this walking-versus-running problem. Further research is organized in this article to investigate the disparity between walking and running. Running is different from walking not only in the speed variances but also, more significantly, in prominent gesture/motion changes. According to these rationales, in our proposed method, we give more attention to the differences between walking and running modes, and a robust gait descriptor is developed to hybridize the mode-invariant spatial and temporal features. Two multi-task learning-based networks are proposed in this method to explore these mode-invariant features. Spatial features describe the body parts non-/less affected by mode changes, and temporal features depict the instinct motion relation of each person. Mode labels are also adopted in the training phase to guide the network to give more attention to the disparity across walking and running modes. In addition, relevant experiments on OU-ISIR Treadmill Dataset A have affirmed the effectiveness and feasibility of the proposed method. A state-of-the-art result can be achieved by our proposed method on this dataset. Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jingsong Xu, Jian Zhang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Robust gait recognition using hybrid descriptors based on Skeleton Gait Energy Image
Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Zhenmin Tang, Wankou Yang |
Pattern Recognit. Lett. | 2 |
| 2020 | Part-based Collaborative Spatio-temporal Feature Learning for Cloth-changing Gait RecognitionabstractIn decades many gait recognition methods have been proposed using different techniques. However, due to a real-world scenario of clothing variations, a reduction of the recognition rate occurs for most of these methods. Thus in this paper, a part-based spatio-temporal feature learning method is proposed to tackle the problem of clothing variations for gait recognition. First, based on the anatomical properties, human bodies are segmented into two regions, which are affected and unaffected by clothing variations. A learning network is particularly proposed in this paper to grasp principal spatio-temporal features from those unaffected regions. Different from most part-based methods with spatial or temporal features solely being utilized, in our method these two features are associated in a more collaborative manner. Snapshots are created for each gait sequence from the H-W and T-W views. Stable spatial information is embedded in the H-W view and adequate temporal information is embedded in the T-W view. An inherent relationship exists between these two views. Thus, a collaborative spatio-temporal feature will be hybridized by concatenating these correlative spatial and temporal information. The robustness and efficiency of our proposed method are validated by experiments on CASIA Gait Dataset B and OU-ISIR Treadmill Gait Dataset B. Our proposed method can both achieve the state-of-the-art results on these two databases. Lingxiang Yao, Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Jingsong Xu |
ICPR | 2 |
| 2020 | Water Level Detection from CCTV Cameras using a Deep Learning ApproachabstractNatural disasters are a global problem that causes widespread losses and damage. A system to provide timely information is required in order to help reduce losses. Flooding is one of the major natural disasters that requires a monitoring and detection system. The traditional flood detection systems use remote sensors such as river water levels and rainfall to provide information to both disaster management professionals and the general public. There is an attempt to use visual information such as CCTV cameras to detect extreme flooding events; however, it requires human experts and consistent attention to monitor any changes. In this paper, we introduce an approach to the automatic river water level detection using deep learning to determine the water level from surveillance cameras. The model achieves 93% accuracy using a single camera location and 83% accuracy using multiple camera locations. Punyanuch Borwarnginn, Jason H. Haga, Worapan Kusakunniran |
TENCON | 3 |
| 2020 | Game-based Learning Tool for PhotographyabstractNowadays, it is convenient to gain new knowledge and skills by learning from the internet. However, some contents may be hard to understand just by reading texts. Photography knowledge is one of those, in which learner may need a large amount of time and cost to practice it. Therefore, this paper provides an alternative way of a game-based learning tool. It simulates a camera into a game combining with a story and some challenges of engagement. It is designed such that a photography knowledge can be delivered to learners through this developed game. To validate this assumption, 25 participants are asked to join our experiment. They are asked to do a pre-test, play our game, and do a post-test. Then, they are asked to answer questionnaires regarding effectiveness and benefits of the proposed game. It is shown that the participants could improve their scores from the pre-test of 44% to the post-test of 89%, regarding the understanding of photography knowledge. Also, the questionnaires' results show that our game could help the participants to gain the knowledge. Napat Romlamduan, Worapan Kusakunniran |
TENCON | 2 |
| 2020 | Biometric for Cattle Identification using Muzzle PatternsabstractSimilar to human biometrics such as faces and fingerprints, animals also have biometrics for individual identifiers. This research paper works on biometrics of cattle using images of muzzle patterns. The proposed approach begins with a training process to construct a cattle face localization model using a Haar feature-based cascade classifier. Then, the watershed technique is applied to segment a region of interest (RoI) of a muzzle area in the detected region of the cattle face. This muzzle ROI is further enhanced to make ridge lines more outstanding. The next step, using two approaches, is to extract a main feature descriptor based on a bag of histograms of oriented gradients (BoHoG) and a histogram of local binary patterns (LBP). Then, the support vector machine (SVM) is applied with the histogram intersection kernel for a final cattle identifier. The proposed method is evaluated using five different datasets including one existing cattle dataset used in previous research works, one newly collected dataset of swamp buffalo captured in a controlled environment, and three newly collected datasets of swamp buffalo captured in an outdoor field environment. This outdoor field environment includes challenges of freely moving cattle and differences in daylight. It could achieve a promising accuracy of 95% for a large dataset of 431 subjects. Worapan Kusakunniran, Anuwat Wiratsudakul, Udom Chuachan, Sarattha Kanchanapreechakorn, Thanandon Imaromkul, Noppanut Suksriupatham, Kittikhun Thongkanchorn |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2018 | Gesture Recognition for Traffic Hand-Signals Training Simulator Using KinectabstractHuman gesture recognition is a way to interpret human movement and/or posture automatically. In this paper, it is used as the main interaction for the developed traffic hand-signals training simulator, based on the Kinect skeleton tracking system. Therefore, the gestures defined in this work are traffic hand-signals used in Thailand. They consist of both static postures and dynamic movements. The recognition is trained and constructed using the rule-based system. The rules must be trained to distinguish these traffic hand-signals, based on both movement information and depth-map information of hands. Then, in a part of the simulator, the artificial intelligent techniques are applied to make it realistic and challenge. The techniques include finite state machine, pathfinding, and path following. They are implemented and used for individual vehicles in the scene. Then, the performances of these two key components of the developed system, the hand-signals recognition and the traffic simulator, are evaluated. It is shown that the system can achieve a very promising performance in both aspects of the recognition accuracy and the user satisfaction. Atid Puwatnuttasit, Worapan Kusakunniran |
TENCON | 2 |
| 2018 | Deep Trajectory Based Gait Recognition for Human Re-identificationabstractThe popular techniques of gait recognition rely on the appearance information, such as Gait Energy Image (GEI). However, they need the pre-processing stage of silhouette segmentation in a walking video. This may not be efficient when the complete silhouette could not be obtained under the cluttered walking environment. It is also sensitive to the changes of walking conditions. Thus, this paper comes up with a new solution using the dense trajectory. This technique is commonly used in the action recognition domain. In this paper, it is used to extract the gait information. The key points and their corresponding trajectories are detected. Then, HOG, HOF, MBHx, MBHy and dense trajectory are extracted from each key point as the point descriptor. In the training phase, the bag of word (BoW) are trained using the extracted point descriptors from the training gait videos. Finally, in the testing phase, the BoW is extracted for each gait video, as the gait feature. The experimental result based on the well-known CASIA gait database B shows the promising performance of the proposed method, under various views. Thunwa Sattrupai, Worapan Kusakunniran |
TENCON | 2 |
| 2018 | Game-based Enhancement for Rehabilitation Based on Action Recognition Using KinectabstractThe physical rehabilitation is a way to improve physical abilities of patients. One of the main problems is patients can get bored from repeated rehab-activities which must be worked out for a long period or even for a lifetime. This can decrease patients' motivations to do the rehabilitation. In addition, game is considered as an entertainment media that is able to use to motivate players to engage in particular activities by giving rewards in exchange. Therefore, this paper proposes the digital game-based enhancement for the Cerebral Palsy (CP) rehabilitation. The CP is the congenital disease caused by the damage on a part of the brain that controls the movement functionality. In this study, the patient must perform specific rehab-actions to control the games. Four games are developed for four rehab-actions including shoulder flexion, shoulder abduction, shoulder horizontal abduction and elbow flexion/extension. They are parts of fundamental movements needed for the Activities of Daily Living (ADL). In each game, the patient must perform the requested rehab-action naturally without carrying any controller. The Kinect is used to track the movements from the patient. Then, the rule-based system is trained and used to detect the correct rehab-action for controlling the game. Then, the experiment is performed to see possibility of using the game-based system for the rehabilitation enhancement. It shows the promising performance in terms of action recognition accuracy that is flexible enough for the patient, the play records, and the expert review. Chanat Sinpithakkul, Worapan Kusakunniran, Sunee Bovonsunthonchai, Peemongkon Wattananon |
TENCON | 2 |
| 2014 | Attribute-based learning for gait recognition using spatio-temporal interest points
Worapan Kusakunniran |
Image Vis. Comput. | 1 |
| 2014 | Recognizing Gaits on Spatio-Temporal Feature DomainabstractGait has been known as an effective biometric feature to identify a person at a distance, e.g., in video surveillance applications. Many methods have been proposed for gait recognitions from various different perspectives. It is found that these methods rely on appearance (e.g., shape contour, silhouette)-based analyses, which require preprocessing of foreground-background segmentation (FG/BG). This process not only causes additional time complexity, but also adversely influences performances of gait analyses due to imperfections of existing FG/BG methods. Besides, appearance-based gait recognitions are sensitive to several variations and partial occlusions, e.g., caused by carrying a bag and varying a cloth type. To avoid these limitations, this paper proposes a new framework to construct a new gait feature directly from a raw video. The proposed gait feature extraction process is performed in the spatio-temporal domain. The space-time interest points (STIPs) are detected by considering large variations along both spatial and temporal directions in local spatio-temporal volumes of a raw gait video sequence. Thus, STIPs are allocated, where there are significant movements of human body in both space and time. A histogram of oriented gradients and a histogram of optical flow are computed on a 3D video patch in a neighborhood of each detected STIP, as a STIP descriptor. Then, the bag-of-words model is applied on each set of STIP descriptors to construct a gait feature for representing and recognizing an individual gait. When compared with other existing methods in the literature, it has been shown that the performance of the proposed method is promising for the case of normal walking, and is outstanding for the case of partial occlusion caused by walking with carrying a bag and walking with varying a cloth type. Worapan Kusakunniran |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2014 | Recognizing Gaits Across Views Through Correlated Motion Co-ClusteringabstractHuman gait is an important biometric feature, which can be used to identify a person remotely. However, view change can cause significant difficulties for gait recognition because it will alter available visual features for matching substantially. Moreover, it is observed that different parts of gait will be affected differently by view change. By exploring relations between two gaits from two different views, it is also observed that a part of gait in one view is more related to a typical part than any other parts of gait in another view. A new method proposed in this paper considers such variance of correlations between gaits across views that is not explicitly analyzed in the other existing methods. In our method, a novel motion co-clustering is carried out to partition the most related parts of gaits from different views into the same group. In this way, relationships between gaits from different views will be more precisely described based on multiple groups of the motion co-clustering instead of a single correlation descriptor. Inside each group, a linear correlation between gait information across views is further maximized through canonical correlation analysis (CCA). Consequently, gait information in one view can be projected onto another view through a linear approximation under the trained CCA subspaces. In the end, a similarity between gaits originally recorded from different views can be measured under the approximately same view. Comprehensive experiments based on widely adopted gait databases have shown that our method outperforms the state-of-the-art. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li, Liang Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | Attribute-based learning for large scale object classificationabstractScalability to large numbers of classes is an important challenge for multi-class classification. It can often be computationally infeasible at test phase when class prediction is performed by using every possible classifier trained for each individual class. This paper proposes an attribute-based learning method to overcome this limitation. First is to define attributes and their associations with object classes automatically and simultaneously. Such associations are learned based on greedy strategy under certain conditions. Second is to learn a classifier for each attribute instead of each class. Then, these trained classifiers are used to predict classes based on their attribute representations. The proposed method also allows trade-off between test-time complexity (which grows linearly with the number of attributes) and accuracy. Experiments based on Animals-with-Attributes and ILSVRC2010 datasets have shown that the performance of our method is promising when compared with the state-of-the-art. Worapan Kusakunniran, Shin'ichi Satoh 0001, Jian Zhang 0002, Qiang Wu 0001 |
ICME | 1 |
| 2013 | A New View-Invariant Feature for Cross-View Gait RecognitionabstractHuman gait is an important biometric feature which is able to identify a person remotely. However, change of view causes significant difficulties for recognizing gaits. This paper proposes a new framework to construct a new view-invariant feature for cross-view gait recognition. Our view-normalization process is performed in the input layer (i.e., on gait silhouettes) to normalize gaits from arbitrary views. That is, each sequence of gait silhouettes recorded from a certain view is transformed onto the common canonical view by using corresponding domain transformation obtained through invariant low-rank textures (TILTs). Then, an improved scheme of procrustes shape analysis (PSA) is proposed and applied on a sequence of the normalized gait silhouettes to extract a novel view-invariant gait feature based on procrustes mean shape (PMS) and consecutively measure a gait similarity based on procrustes distance (PD). Comprehensive experiments were carried out on widely adopted gait databases. It has been shown that the performance of the proposed method is promising when compared with other existing methods in the literature. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2012 | Cross-view and multi-view gait recognitions based on view transformation model using multi-layer perceptron
Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
Pattern Recognit. Lett. | 1 |
| 2012 | Gait Recognition Under Various Viewing Angles Based on Correlated Motion RegressionabstractIt is well recognized that gait is an important biometric feature to identify a person at a distance, e.g., in video surveillance application. However, in reality, change of viewing angle causes significant challenge for gait recognition. A novel approach using regression-based view transformation model (VTM) is proposed to address this challenge. Gait features from across views can be normalized into a common view using learned VTM(s). In principle, a VTM is used to transform gait feature from one viewing angle (source) into another viewing angle (target). It consists of multiple regression processes to explore correlated walking motions, which are encoded in gait features, between source and target views. In the learning processes, sparse regression based on the elastic net is adopted as the regression function, which is free from the problem of overfitting and results in more stable regression models for VTM construction. Based on widely adopted gait database, experimental results show that the proposed method significantly improves upon existing VTM-based methods and outperforms most other baseline methods reported in the literature. Several practical scenarios of applying the proposed method for gait recognition under various views are also discussed in this paper. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | Gait Recognition Across Various Walking Speeds Using Higher Order Shape Configuration Based on a Differential Composition ModelabstractGait has been known as an effective biometric feature to identify a person at a distance. However, variation of walking speeds may lead to significant changes to human walking patterns. It causes many difficulties for gait recognition. A comprehensive analysis has been carried out in this paper to identify such effects. Based on the analysis, Procrustes shape analysis is adopted for gait signature description and relevant similarity measurement. To tackle the challenges raised by speed change, this paper proposes a higher order shape configuration for gait shape description, which deliberately conserves discriminative information in the gait signatures and is still able to tolerate the varying walking speed. Instead of simply measuring the similarity between two gaits by treating them as two unified objects, a differential composition model (DCM) is constructed. The DCM differentiates the different effects caused by walking speed changes on various human body parts. In the meantime, it also balances well the different discriminabilities of each body part on the overall gait similarity measurements. In this model, the Fisher discriminant ratio is adopted to calculate weights for each body part. Comprehensive experiments based on widely adopted gait databases demonstrate that our proposed method is efficient for cross-speed gait recognition and outperforms other state-of-the-art methods. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2011 | Pairwise Shape configuration-based PSA for gait recognition under small viewing angle changeabstractTwo main components of Procrustes Shape Analysis (PSA) are adopted and adapted specifically to address gait recognition under small viewing angle change: 1) Procrustes Mean Shape (PMS) for gait signature description; 2) Procrustes Distance (PD) for similarity measurement. Pairwise Shape Configuration (PSC) is proposed as a shape descriptor in place of existing Centroid Shape Configuration (CSC) in conventional PSA. PSC can better tolerate shape change caused by viewing angle change than CSC. Small variation of viewing angle makes large impact only on global gait appearance. Without major impact on local spatio-temporal motion, PSC which effectively embeds local shape information can generate robust view-invariant gait feature. To enhance gait recognition performance, a novel boundary re-sampling process is proposed. It provides only necessary re-sampled points to PSC description. In the meantime, it efficiently solves problems of boundary point correspondence, boundary normalization and boundary smoothness. This re-sampling process adopts prior knowledge of body pose structure. Comprehensive experiment is carried out on the CASIA gait database. The proposed method is shown to significantly improve performance of gait recognition under small viewing angle change without additional requirements of supervised learning, known viewing angle and multi-camera system, when compared with other methods in literatures. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
AVSS | 1 |
| 2011 | Speed-invariant gait recognition based on Procrustes Shape Analysis using higher-order shape configurationabstractWalking speed change is considered a typical challenge hindering reliable human gait recognition. This paper proposes a novel method to extract speed-invariant gait feature based on Procrustes Shape Analysis (PSA). Two major components of PSA, i.e., Procrustes Mean Shape (PMS) and Procrustes Distance (PD), are adopted and adapted specifically for the purpose of speed-invariant gait recognition. One of our major contributions in this work is that, instead of using conventional Centroid Shape Configuration (CSC) which is not suitable to describe individual gait when body shape changes particularly due to change of walking speed, we propose a new descriptor named Higher-order derivative Shape Configuration (HSC) which can generate robust speed-invariant gait feature. From the first order to the higher order, derivative shape configuration contains gait shape information of different levels. Intuitively, the higher order of derivative is able to describe gait with shape change caused by the larger change of walking speed. Encouraging experimental results show that our proposed method is efficient for speed-invariant gait recognition and evidently outperforms other existing methods in the literatures. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
ICIP | 1 |
| 2010 | Support vector regression for multi-view gait recognition based on local motion feature selectionabstractGait is a well recognized biometric feature that is used to identify a human at a distance. However, in real environment, appearance changes of individuals due to viewing angle changes cause many difficulties for gait recognition. This paper re-formulates this problem as a regression problem. A novel solution is proposed to create a View Transformation Model (VTM) from the different point of view using Support Vector Regression (SVR). To facilitate the process of regression, a new method is proposed to seek local Region of Interest (ROI) under one viewing angle for predicting the corresponding motion information under another viewing angle. Thus, the well constructed VTM is able to transfer gait information under one viewing angle into another viewing angle. This proposal can achieve view-independent gait recognition. It normalizes gait features under various viewing angles into a common viewing angle before similarity measurement is carried out. The extensive experimental results based on widely adopted benchmark dataset demonstrate that the proposed algorithm can achieve significantly better performance than the existing methods in literature. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
CVPR | 1 |
| 2010 | Multi-view Gait Recognition Based on Motion Regression Using Multilayer PerceptronabstractIt has been shown that gait is an efficient biometric feature for identifying a person at a distance. However, it is a challenging problem to obtain reliable gait feature when viewing angle changes because the body appearance can be different under the various viewing angles. In this paper, the problem above is formulated as a regression problem where a novel View Transformation Model (VTM) is constructed by adopting Multilayer Perceptron (MLP) as regression tool. It smoothly estimates gait feature under an unknown viewing angle based on motion information in a well selected Region of Interest (ROI) under other existing viewing angles. Thus, this proposal can normalize gait features under various viewing angles into a common viewing angle before gait similarity measurement is carried out. Encouraging experimental results have been obtained based on widely adopted benchmark database. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
ICPR | 1 |
| 2009 | Automatic Gait Recognition Using Weighted Binary Pattern on VideoabstractHuman identification by recognizing the spontaneous gait recorded in real-world setting is a tough and not yet fully resolved problem in biometrics research. Several issues have contributed to the difficulties of this task. They include various poses, different clothes, moderate to large changes of normal walking manner due to carrying diverse goods when walking, and the uncertainty of the environments where the people are walking. In order to achieve a better gait recognition, this paper proposes a new method based on Weighted Binary Pattern (WBP). WBP first constructs binary pattern from a sequence of aligned silhouettes. Then, adaptive weighting technique is applied to discriminate significances of the bits in gait signatures. Being compared with most of existing methods in the literatures, this method can better deal with gait frequency, local spatial-temporal human pose features, and global body shape statistics. The proposed method is validated on several well known benchmark databases. The extensive and encouraging experimental results show that the proposed algorithm achieves high accuracy, but with low complexity and computational time. Worapan Kusakunniran, Qiang Wu 0001, Hongdong Li, Jian Zhang 0002 |
AVSS | 1 |