Xin Ma 0001

dblp:18/6265-1 · DBLP profile ↗
← Back
32ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0003-4402-1957ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Batch self-organizing memory neural network for continual supervised learning
Jiahui Niu, Xin Ma 0001
Neural Networks2
2026 Multi-Task Learning for Gait Phase and Gait Cycle Percentage Prediction With Wearable Sensors in Frail Older Adults
abstract
Deep learning has been widely used in wearable sensors to improve accuracy in gait analysis. However, these deep learning models typically focus on single tasks, either in gait parameter estimation or gait phase detection. This study presents a novel multi-task learning framework for regression (i.e., gait cycle percentage prediction) and classification (i.e., gait phase prediction) tasks in pathological gait analysis using wearable sensors. Our framework employs a Multi-gate Mixture-of-Experts architecture to achieve soft parameter-sharing, integrating expert networks, cross-expert attention mechanisms, and dynamic routing to balance shared and task-specific representations. To reduce computational burden in wearable applications, we compare lightweight model configurations that optimize expert count and feature dimensionality. Model performance has been validated on a public dataset consisting of 158 frail older adults, demonstrating that our framework significantly outperforms single-task learning and hard parameter-sharing baselines, achieving an accuracy of 97.56% and a Mean Absolute Error (MAE) of 0.0397. Notably, the most compact lightweight configuration reduces the parameter count by nearly 98% (from 2.118 million to 0.0469 million), achieving an accuracy of 96.47% and a MAE of 0.0549. Attention mechanisms significantly enhance performance across all configurations, with improvements ranging from 17.9% to 30.4%. These findings validate the potential of lightweight multi-task approaches for real-time gait assessment, offering promising applications for clinical evaluation and rehabilitation monitoring in geriatric populations.
Zeyang Guan, Ziyun Ding, Xin Ma 0001, Yibin Li 0001, Rui Song 0002, Huanghe Zhang
IEEE J. Biomed. Health Informatics5
2025 Enhancing autism spectrum disorder early detection with parent-child dyads block-play protocol and attention-enhanced hybrid deep learning framework
Xiang Li 0131, Lizhou Fan, Hanbo Wu, Kunping Chen, Xiaoxiao Yu, Chao Che, Zhifeng Cai, Xiuhong Niu, Aihua Cao, Xin Ma 0001
Eng. Appl. Artif. Intell.10
2025 Unsupervised domain adaptation framework with global-local adversarial learning and masked image consistency for fish counting in deep-sea aquaculture
Hanchi Liu, Xin Ma 0001
Eng. Appl. Artif. Intell.2
2025 Semantic-guided modeling of spatial relation and object co-occurrence for indoor scene recognition
Chuanxin Song, Hanbo Wu, Xin Ma 0001
Expert Syst. Appl.3
2025 Transformer-based multiview spatiotemporal feature interactive fusion for human action recognition in depth videos
Hanbo Wu, Xin Ma 0001
Signal Process. Image Commun.2
2025 Neural Network-Based Nonlinear Stabilizing Control for 3-D Offshore Crane With Double-Pendulum Effect
abstract
Wave-induced ship motions pose great challenges to the design of control systems for offshore cranes, especially with double-pendulum effect and unknown system dynamics. In this paper, a neural network-based nonlinear stabilizing controller is proposed for 3D offshore cranes to accomplish boom positioning and payload swing-elimination under wave-induced ship roll and pitch motions. Critical and practical-oriented issues including double-pendulum effect, boom position limitations, actuator input dead zones, and unknown system dynamics are considered simultaneously. First, the dynamic model of 3D offshore cranes with double-pendulum effect is established by using the Lagrange’s modeling method. By combining the state variables with perturbation terms, new auxiliary variables are introduced for model transformation to simplify model analysis and controller design. Then, by analyzing the transformed model, the neural network is rationally designed to estimate the unknown system dynamics and input dead zones. In addition, barrier Lyapunov functions (BLFs) are employed to the controller ensuring that the boom operates within a safe range. The Lyapunov-based theory is utilized to rigorously prove the stability and convergence of the designed control system. Hardware experiments are elaborately designed to verify the performance of the proposed method in terms of effectiveness, robustness, and anti-disturbance. Note to Practitioners—This paper studies the stabilizing control problem of offshore crane systems. In practice, the double-pendulum effect between the payload and hook is evident, and it is necessary to simultaneously lift and rotate the boom to position and stabilize the payload smoothly and accurately under wave-induced ship motions. However, existing studies typically oversimplify the model of offshore cranes, neglecting crucial factors such as their 3D characteristics, the double-pendulum effect, and the intricate wave-induced ship motions. This limits the application of the corresponding control methods in practice. Moreover, most of the existing control methods ignore unknown model dynamics, input dead zones, etc., which is not favorable for practical applications. To address the above problems, this paper proposes a neural network-based nonlinear stabilizing controller, which can accomplish the stabilization control for 3D offshore cranes with double pendulum effect in the presence of unknown dynamics and input dead zones. The effectiveness of the proposed method is validated on the self-built hardware platform.
Xin Ma 0001
IEEE Trans Autom. Sci. Eng.3
2025 Adaptive Sliding Mode Control Based on Time-Delay Estimation for Underactuated 7-DOF Tower Crane
abstract
Tower cranes are complex multi-input multioutput underactuated mechatronics systems. The anti-swing control issue of tower crane with varying suspension cable length and double spherical pendulum effect is still open. Furthermore, the system parameters uncertainty makes it more challenging to implement anti-swing control. In this study, we present an adaptive sliding mode anti-swing control approach based on time-delay estimation for underactuated tower crane with varying suspension cable length and double spherical pendulum effect. First, we employ the Lagrange’s method to develop a seven-degree-of-freedom (7-DOF) tower crane dynamic model that comprehensively accounts for jib slewing, trolley motion, payload hoisting/lowering, and payload/hook spherical swing within a three-dimensional (3-D) space. Then, a sliding mode surface is constructed by analyzing the nonlinear coupling relationship between the unactuated states and actuated states. The time-delay estimation technique with adaptive scheme can adapt and predicate unknown system parameters online. An adaptive sliding mode anti-swing control method with time-delay estimation is designed for 7-DOF tower crane system subject to the parameter uncertainties. The convergence of the closed-loop control system is carefully demonstrated through the Lyapunov stability theory. Finally, the hardware experiments verify the anti-swing control performance and robustness of the designed adaptive sliding mode controller. The superiority of the proposed adaptive sliding mode anti-swing controller is confirmed by a decrease of at least 42.09${\%}$and 58.33${\%}$in the maximum and residual payload swing, respectively, over state-of-the-art control methods.
Xin Ma 0001, Yibin Li 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2024 SC-ViT: Semantic Contrast Vision Transformer for Scene Recognition
abstract
Scene recognition remains a challenging task in image recognition. Despite the remarkable advances made by deep learning, especially with the emergence of Convolutional Neural Networks (CNNs), scene recognition continues to face unresolved issues. Due to high complexity of scene images, merely identifying several objects in the image is insufficient for obtaining accurate results. Furthermore, the wide range of scene categories makes single-modality learning susceptible to confusion. To address these challenges, we propose an end-to-end multimodal network, SC-ViT, based on the Vision Transformer (ViT). Leveraging the powerful self-attention mechanism, our model captures visual and contextual cues from both RGB images and semantic information. The semantic information is derived through semantic segmentation, encompassing object category information and spatial layout within the scene. Combining semantic information with RGB images results in a comprehensive representation of the scene. Specifically, we utilize two branches, each equipped with self-attention mechanisms but with different structures, to extract features from RGB images and semantic information. Through a contrastive learning framework, SC-ViT aligns the feature representations of RGB and semantic modalities, enhancing their consistency and discriminative power to express the scene. Experimental evaluations on the MIT Indoor67 and SUN397 datasets demonstrate that SC-ViT outperforms state-of-the-art methods, achieving significant improvements in scene recognition.
Jiahui Niu, Xin Ma 0001
IJCNN2
2024 Inter-object discriminative graph modeling for indoor scene recognition
Chuanxin Song, Hanbo Wu, Xin Ma 0001
Knowl. Based Syst.3
2024 Semantic-embedded similarity prototype for scene recognition
Chuanxin Song, Hanbo Wu, Xin Ma 0001, Yibin Li 0001
Pattern Recognit.3
2024 Adaptive Anti-Swing Control for 7-DOF Overhead Crane With Double Spherical Pendulum and Varying Cable Length
abstract
Three dimensional (3D) overhead crane is a typical multi-input multi-output (MIMO) and underactuated mechatronic system. Complex double spherical pendulum dynamics and varying cable length increase the difficulty of the swing suppression control for overhead crane. Moreover, it would be a great challenge for considering friction uncertain and overshoot issues of the actuators. In this article, an adaptive anti-swing controller is proposed for seven degrees of freedom (7-DOF) overhead crane with double spherical pendulum and varying cable length. First, we use Lagrange’s method to establish an accurate dynamic model of 7-DOF overhead crane. The complex nonlinear dynamic model includes trolley moving, bridge moving, cable length varying, hook swing and payload swing in 3D space. Then, by analysis the system energy function, an adaptive control method is designed to control the trolley, bridge and suspension cable simultaneously. Moreover, we elaborately design some nonlinear terms (such as overshoot limiting term, adaptive term and anti-swing term), which can be used to handle the actuator overshoots, friction uncertain and unactuated swing suppression. As far as we know, it is the first nonlinear closed-loop controller for 7-DOF overhead crane without any linearization. Finally, a group of the hardware experiment results prove that the proposed adaptive anti-swing controller has better effectiveness and robustness than the existing state-of-the-art controllers.Note to Practitioners—This paper studies the anti-swing control problem of the 7-DOF overhead crane system. To improve the efficiency of overhead crane, the trolley moving, the bridge moving and the payload hoisting/lowing are carried out simultaneously in practice. Due to the complex dynamics of the double spherical pendulum and varying cable length, the anti-swing control of the 7-DOF crane remains an open problem. Moreover, most existing control methods ignore the friction uncertain and actuator overshoot, which is not feasible in practice. To tackle these problems, this paper designs a novel adaptive anti-swing controller, which can solve the issues of friction estimation, overshoot limitation, actuators positioning while suppressing the unactuated swing in the 7-DOF overhead crane system. The control performance of the proposed anti-swing controller is verified on the self-build 7-DOF overhead crane platform in the laboratory. In the future, the proposed anti-swing controller will be applied in the industrial overhead crane system to improve the control performance.
Xin Ma 0001, Yibin Li 0001
IEEE Trans Autom. Sci. Eng.2
2023 SRRM: Semantic Region Relation Model for Indoor Scene Recognition
abstract
Despite the remarkable success of convolutional neural networks in various computer vision tasks, recognizing indoor scenes still presents a significant challenge due to their complex composition. Consequently, effectively leveraging semantic information in the scene has been a key issue in advancing indoor scene recognition. Unfortunately, the accuracy of semantic segmentation has limited the effectiveness of existing approaches for leveraging semantic information. As a result, many of these approaches remain at the stage of auxiliary labeling or co-occurrence statistics, with few exploring the contextual relationships between the semantic elements directly within the scene. In this paper, we propose the Semantic Region Relationship Model (SRRM), which starts directly from the semantic information inside the scene. Specifically, SRRM adopts an adaptive and efficient approach to mitigate the negative impact of semantic ambiguity and then models the semantic region relationship to perform scene recognition. Additionally, to more comprehensively exploit the information contained in the scene, we combine the proposed SRRM with the PlacesCNN module to create the Combined Semantic Region Relation Model (CSRRM), and propose a novel information combining approach to effectively explore the complementary contents between them. CSRRM significantly outperforms the SOTA methods on the MIT Indoor 67, reduced Places365 dataset, and SUN RGB-D without retraining. The code is available at: https://github.com/ChuanxinSong/SRRM
Chuanxin Song, Xin Ma 0001
IJCNN2
2023 Multi-level channel attention excitation network for human action recognition in videos
Hanbo Wu, Xin Ma 0001, Yibin Li 0001
Signal Process. Image Commun.2
2023 Attention-Driven Appearance-Motion Fusion Network for Action Recognition
abstract
Recent years have witnessed the popularity of using a two-stream architecture and attention mechanism for action recognition with videos. However, it is time-consuming to train two separate convolutional neural networks (ConvNets), especially with the complex attention mechanism. In this paper, we present a novel architecture, termed as Appearance-Motion Fusion Network (AMFNet), to learn efficient and robust action representation from RGB and optical flow data in an end-to-end manner. AMFNet is constructed by connecting a convolutional neural network with an appearance-motion fusion block (AMFB), whose goal is to incorporate appearance and motion streams into a unified framework driven by a cross-modality attention (CMA) mechanism. More specifically, the CMA only relies on optical flow data, which consists of a Key-Frame Adaptive Selection Module (KFASM) and an Optical-Flow-Driven Spatial Attention Module (OFDSAM). The former aims to adaptively identify the discriminative key frames from a sequence, while the latter is able to guide our networks to focus on the action-relevant regions of each frame. We explore two schemes for appearance and motion streams fusion in AMFB from hierarchical and comprehensive levels. The proposed AMFNet is extensively evaluated on five action recognition data sets, including HMDB-51, UCF-101, JHMDB, Penn and Kinetics-400. Compared to the state-of-the-art methods operated at RGB and optical flow, the experimental results validate that our AMFNet achieves a comparable performance with a pure 2D-Single-ConvNet design.
Shaocan Liu, Xin Ma 0001
IEEE Trans. Multim.2
2023 Skeleton-Based Action Recognition With Select-Assemble-Normalize Graph Convolutional Networks
abstract
Skeleton-based action recognition has been substantially driven by the development of artificial intelligence technology and deep sensors. Recently, graph convolutional networks (GCNs) have achieved excellent performances in skeleton-based action recognition. However, the performances of GCN-based methods are impaired by inappropriate node partitioning strategy and obstructed long-range information flow. To solve these issues, a novel Select-Assemble-Normalize Graph Convolution Network (SAN-GCN) is proposed to model the spatio-temporal features of skeleton. First, all skeleton joints are selected as root nodes, and the neighborhoods of the root joints are assembled and normalized according to the body structure, which explicitly and interpretably expresses the spatial geometric relation of the skeleton joints. Second, we propose an attention-based assembly and normalization strategy to adaptively capture non-local joints. The adaptive assembly and normalization can avoid the dilution of key long-range features. Moreover, a bi-level aggregation strategy is introduced to learn spatio-temporal dependencies of joints, where the low-level aggregation aligns the normalized neighborhood graphs, and the high-level aggregation aggregates the features of neighbor nodes by a standard convolution kernel. In high-level aggregation, it is convenient to realize factorized spatio-temporal aggregation or unified spatio-temporal aggregation. Extensive experiments on four datasets with different numbers of action patterns demonstrate that our model achieves comparable performance with the state-of-the-art works.
Xin Ma 0001, Xiang Li 0131, Yibin Li 0001
IEEE Trans. Multim.2
2022 Skeleton-based abnormal gait recognition with spatio-temporal attention enhanced gait-structural graph convolutional networks
Xin Ma 0001, Hanbo Wu, Yibin Li 0001
Neurocomputing2
2022 Adaptive neural control for mobile manipulator systems based on adaptive state observer
Yukun Zheng, Yixiang Liu, Rui Song 0002, Xin Ma 0001, Yibin Li 0001
Neurocomputing4
2022 Spatiotemporal Multimodal Learning With 3D CNNs for Video Action Recognition
abstract
Extracting effective spatial-temporal information is significantly important for video-based action recognition. Recently 3D convolutional neural networks (3D CNNs) that could simultaneously encode spatial and temporal dynamics in videos have made considerable progress in action recognition. However, almost all existing 3D CNN-based methods recognize human actions only using RGB videos. The single modality may limit the performance capacity of 3D networks. In this paper, we extend 3D CNN to depth and pose data besides RGB data to evaluate its capacity for spatiotemporal multimodal learning for video action recognition. We propose a novel multimodal two-stream 3D network framework, which can exploit complementary multimodal information to improve the recognition performance. Specifically, we first construct two discriminative video representations under depth and pose data modalities respectively, referred as depth residual dynamic image sequence (DRDIS) and pose estimation map sequence (PEMS). DRDIS captures spatial-temporal evolution of actions in depth videos by progressively aggregating the local motion information. PEMS eliminates the interference of cluttered backgrounds and describes the spatial configuration of body parts intuitively. The multimodal two-stream 3D CNN deals with two separate data streams to learn spatiotemporal features from DRDIS and PEMS representations. Finally, the classification scores from two streams are fused for action recognition. We conduct extensive experiments on four challenging action recognition datasets. The experimental results verify the effectiveness and superiority of our proposed method.
Hanbo Wu, Xin Ma 0001, Yibin Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Autonomous cognition development with lifelong learning: A self-organizing and reflecting cognitive network
Xin Ma 0001, Rui Song 0002, Xuewen Rong, Yibin Li 0001
Neurocomputing2
2020 End-to-end multitask Siamese network with residual hierarchical attention for real-time object tracking
Wenhui Huang 0002, Jason Gu, Xin Ma 0001, Yibin Li 0001
Appl. Intell.3
2020 A self-organizing developmental cognitive architecture with interactive reinforcement learning
Xin Ma 0001, Rui Song 0002, Xuewen Rong, Xincheng Tian, Yibin Li 0001
Neurocomputing2
2020 Convolutional Networks With Channel and STIPs Attention Model for Action Recognition in Videos
abstract
With the help of convolutional neural networks (CNNs), video-based human action recognition has made significant progress. CNN features that are spatial and channel-wise can provide rich information for powerful image description. However, CNNs lack the ability to process the long-term temporal dependency of an entire video and further cannot well focus on the informative motion regions of actions. Aiming at the two problems, we propose a novel video-based action recognition framework in this paper. We first represent videos with dynamic image sequences (DISs), which effectively describe videos by modeling the local spatial-temporal dynamics and dependencies. Then a channel and spatial-temporal interest points (STIPs) attention model (CSAM) based on CNNs is proposed to focus on the discriminative channels in networks and the informative spatial motion regions of human actions. Specifically, channel attention (CA) is implemented by automatically learning channel-wise convolutional features and assigning different weights for different channels. STIPs attention (SA) is encoded by projecting the detected STIPs on frames of dynamic image sequences into the corresponding convolutional feature map space. The proposed CSAM is embedded after CNN convolutional layers to refine the feature maps, followed by global average pooling to produce effective feature representations for videos. Finally frame-level video representations are fed into an LSTM to capture the temporal dependencies and make classification. Experiments on three challenging RGB-D datasets show that our method has better performance and outperforms the state-of-the-art approaches with only depth data.
Hanbo Wu, Xin Ma 0001, Yibin Li 0001
IEEE Trans. Multim.2
2019 Quarter-Point Codeword Expansion for Product Quantization
abstract
Due to its low storage cost and high query accuracy, Product Quantization (PQ) has been widely used for approximate nearest neighbor (ANN) search. However, almost all existing PQ-based methods use the nearest clustering center as the codeword, which might not fully utilize the information of distances from data points to clustering centers. In this paper, we propose a novel codeword expansion method for PQ-based methods, called Quarter-point Codeword Expansion (QCE), by estimating the distances from the query points to the database points using the quarter points instead of the clustering centers. The distances can be computed more precisely and it will result in a lower distortion using QCE, which is also a general method could be used to improve all PQ-based methods. Extensive experiments on approximate nearest neighbor search show that PQ-based methods with QCE can outperform the state-of-the-art.
Shan An, Zhibiao Huang, Guangfu Che, Xianglong Liu 0001, Xin Ma 0001
ICME5
2019 Fast and Incremental Loop Closure Detection Using Proximity Graphs
abstract
Visual loop closure detection, which can be considered as an image retrieval task, is an important problem in SLAM (Simultaneous Localization and Mapping) systems. The frequently used bag-of-words (BoW) models can achieve high precision and moderate recall. However, the requirement for lower time costs and fewer memory costs for mobile robot applications is not well satisfied. In this paper, we propose a novel loop closure detection framework titled FILD' (Fast and Incremental Loop closure Detection), which focuses on an on-line and incremental graph vocabulary construction for fast loop closure detection. The global and local features of frames are extracted using the Convolutional Neural Networks (CNN) and SURF on the GPU, which guarantee extremely fast extraction speeds. The graph vocabulary construction is based on one type of proximity graph, named Hierarchical Navigable Small World (HNSW) graphs, which is modified to adapt to this specific application. In addition, this process is coupled with a novel strategy for real-time geometrical verification, which only keeps binary hash codes and significantly saves on memory usage. Extensive experiments on several publicly available datasets show that the proposed approach can achieve fairly good recall at 100% precision compared to other state-of-the-art methods. The source code can be downloaded at https://github.comlAnshanTJU/FILD for further studies.
Shan An, Guangfu Che, Fangru Zhou, Xianglong Liu 0001, Xin Ma 0001
IROS5
2019 Robust tracking control strategy for a quadrotor using RPD-SMC and RISE
Xin Ma 0001, Yibin Li 0001
Neurocomputing2
2018 Hybrid 3D Surface Description with Global Frames and Local Signatures of Histograms
abstract
This paper presents a novel 3D descriptor named Frame-SHOT to combine global structural frame with local surface information. Global feature descriptors are generally more descriptive for surface representation, while susceptible to occlusion and clutter. In contrast, local feature descriptors are more robust, while less discriminative due to the limit of the support region. We employ the advantages of these two methods by combining both local and global information for surface description. The Signature of Histograms of Orientation (SHOT) descriptor is used to characterize the local surface and structural frame points of the object are used to encode the global surface. We have compared the proposed Frame-SHOT descriptor with the state-of-the-art global and local methods on two public datasets. The results show that our proposed descriptor is more descriptive and robust for surface matching and recognition.
Xin Ma 0001, Xianglei Zeng
ICPR2
2017 Correlation filter-based self-paced object tracking
abstract
Object tracking is an important capability for robots tasked with interacting with humans and the environment, and it enables robots to manipulate objects. In object tracking, selecting samples to learn a robust and efficient appearance model is a challenging task. Model learning determines both the strategy and frequency of model updating, which concerns many details that can affect the tracking results. In this paper, we propose an object tracking approach by formulating a new objective function that integrates the learning paradigm of self-paced learning into object tracking such that reliable samples can be automatically selected for model learning. Sample weights and model parameters can be learned by minimizing this single objective function under the framework of kernelized correlation filters. Moreover, a real-valued error-tolerant self-paced function with a constraint vector is proposed to combine prior knowledge, i.e., the characteristics of object tracking, with information learned during tracking. We demonstrate the robustness and efficiency of our object tracking approach on a recent object tracking benchmark data set: OTB 2013.
Wenhui Huang 0002, Jason Gu, Xin Ma 0001, Yibin Li 0001
ICRA3
2014 Intelligent mobility assisted mobile sensor network localization
abstract
The trajectories of mobile seeds have a great influence on localization accuracy and efficiency. This paper presents a novel information-driven intelligent mobility-assisted wireless sensor network localization algorithm. Without requiring any prior knowledge of the sensing field, seeds' or pseudo-seeds' (common sensors which have been positioned) trajectories are scheduled dynamically aiming at position estimates of neighboring non-positioned common sensors. With an information-theoretic utility measure as the objective function, mobile seeds or pseudo-seeds actively determine their motion directions for minimizing the uncertainty in position estimates of neighboring sensors. At the first level, seeds estimate the neighboring sensor nodes' positions with bearing measurements by means of extended Kalman filters and optimize their motion directions by maximizing the mutual information between the position estimates and the motions of seeds. Afterwards the seeds forward the position estimates to the corresponding sensor nodes, which then act as pseudo-seeds. By repeating this process at the following levels, all sensor nodes can obtain position estimates. Compared with heuristic mobility and random mobility-assisted mobile sensor network localization algorithms, the proposed algorithm requires fewer maneuvers of seed or pseudo-seeds for quick convergence to good position estimates. Extensive simulations show that this algorithm can provide more accurate position estimates with fewer maneuvers, especially in the case of limited seeds.
Xin Ma 0001, Mingang Zhou, Yibin Li 0001, Jindong Tan
ICRA1
2014 Depth-Based Human Fall Detection via Shape Features and Improved Extreme Learning Machine
abstract
Falls are one of the major causes leading to injury of elderly people. Using wearable devices for fall detection has a high cost and may cause inconvenience to the daily lives of the elderly. In this paper, we present an automated fall detection approach that requires only a low-cost depth camera. Our approach combines two computer vision techniques-shape-based fall characterization and a learning-based classifier to distinguish falls from other daily actions. Given a fall video clip, we extract curvature scale space (CSS) features of human silhouettes at each frame and represent the action by a bag of CSS words (BoCSS). Then, we utilize the extreme learning machine (ELM) classifier to identify the BoCSS representation of a fall from those of other actions. In order to eliminate the sensitivity of ELM to its hyperparameters, we present a variable-length particle swarm optimization algorithm to optimize the number of hidden neurons, corresponding input weights, and biases of ELM. Using a low-cost Kinect depth camera, we build an action dataset that consists of six types of actions (falling, bending, sitting, squatting, walking, and lying) from ten subjects. Experimenting with the dataset shows that our approach can achieve up to 91.15% sensitivity, 77.14% specificity, and 86.83% accuracy. On a public dataset, our approach performs comparably to state-of-the-art fall detection methods that need multiple cameras.
Xin Ma 0001, Bingxia Xue, Mingang Zhou, Bing Ji 0001, Yibin Li 0001
IEEE J. Biomed. Health Informatics1
2013 State-chain sequential feedback reinforcement learning for path planning of autonomous mobile robots
abstract
This paper deals with a new approach based on Q -learning for solving the problem of mobile robot path planning in complex unknown static environments. As a computational approach to learning through interaction with the environment, reinforcement learning algorithms have been widely used for intelligent robot control, especially in the field of autonomous mobile robots. However, the learning process is slow and cumbersome. For practical applications, rapid rates of convergence are required. Aiming at the problem of slow convergence and long learning time for Q -learning based mobile robot path planning, a state-chain sequential feedback Q -learning algorithm is proposed for quickly searching for the optimal path of mobile robots in complex unknown static environments. The state chain is built during the searching process. After one action is chosen and the reward is received, the Q -values of the state-action pairs on the previously built state chain are sequentially updated with one-step Q -learning. With the increasing number of Q -values updated after one action, the number of actual steps for convergence decreases and thus, the learning time decreases, where a step is a state transition. Extensive simulations validate the efficiency of the newly proposed approach for mobile robot path planning in complex environments. The results show that the new approach has a high convergence speed and that the robot can find the collision-free optimal path in complex unknown static environments with much shorter time, compared with the one-step Q -learning algorithm and the Q ( λ )-learning algorithm.
Xin Ma 0001, Ya Xu, Guo-qiang Sun, Yibin Li 0001
J. Zhejiang Univ. Sci. C1
2007 Immunity-Based Adaptive Genetic Algorithm for Multi-robot Cooperative Exploration
Xin Ma 0001, Yibin Li 0001
ICIC (2)1