Dingsheng Luo

dblp:64/5953 · DBLP profile ↗
← Back
32ranked-venue papers
4as first author
17since 2021 · last 2025
0000-0001-6997-3984ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 4 first-author · 16 since 2021Systems, architecture and hardware · 10 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Human-computer interaction and ubiquitous computing · 7 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 CEMSSL: Conditional Embodied Self-Supervised Learning is All You Need for High-precision Multi-solution Inverse Kinematics of Robot Arms
abstract
In the field of signal processing for robotics, the inverse kinematics of robot arms presents a significant challenge due to multiple solutions caused by redundant degrees of freedom (DOFs). Precision is also a crucial performance indicator for robot arms. Current methods typically rely on conditional deep generative models (CDGMs), which often fall short in precision. In this paper, we propose Conditional Embodied Self-Supervised Learning (CEMSSL) and introduce a unified framework based on CEMSSL for high-precision multi-solution inverse kinematics learning. This framework enhances the precision of existing CDGMs by up to 2-3 orders of magnitude while maintaining their original properties. Furthermore, our method is extendable to other fields of signal processing where obtaining multi-solution data in advance is challenging, as well as to other problems involving multi-solution inverse processes.
Weiming Qu, Tianlin Liu, Dingsheng Luo
ICASSP4
2025 Domain Adaptation-Based Crossmodal Knowledge Distillation for 3D Semantic Segmentation
abstract
Semantic segmentation of 3D LiDAR data plays a pivotal role in autonomous driving. Traditional approaches rely on extensive annotated data for point cloud analysis, incurring high costs and time investments. In contrast, real-world image datasets offer abundant availability and substantial scale. To mitigate the burden of annotating 3D LiDAR point clouds, we propose two crossmodal knowledge distillation methods: Unsupervised Domain Adaptation Knowledge Distillation (UDAKD) and Feature and Semantic-based Knowledge Distillation (FSKD). Leveraging readily available spatio-temporally synchronized data from cameras and LiDARs in autonomous driving scenarios, we directly apply a pretrained 2D image model to unlabeled 2D data. Through crossmodal knowledge distillation with known 2D-3D correspondence, we actively align the output of the 3D network with the corresponding points of the 2D network, thereby obviating the necessity for 3D annotations. Our focus is on preserving modality-general information while filtering out modality-specific details during crossmodal distillation. To achieve this, we deploy self-calibrated convolution on 3D point clouds as the foundation of our domain adaptation module. Rigorous experimentation validates the effectiveness of our proposed methods, consistently surpassing the performance of state-of-the-art approaches in the field. Code is available at https://github.com/KangJialiang/DAKD.
Jialiang Kang, Dingsheng Luo
ICRA3
2025 EHC-MM: Embodied Holistic Control for Mobile Manipulation
abstract
Mobile manipulation typically entails the base for mobility, the arm for accurate manipulation, and the camera for perception. The principle of Distant Mobility, Close Grasping(DMCG) is essential for holistic control. We propose Embod-ied Holistic Control for Mobile Manipulation(EHC-MM) with the embodied function of sig($\omega$): By formulating the DMCG principle as a Quadratic Programming (QP) problem, sig($\omega$) dynamically balances the robot's emphasis between movement and manipulation with the consideration of the robot's state and environment. In addition, we propose the Monitor-Position-Based Servoing (MPBS) with sig($\omega$), enabling the tracking of the target during the operation. This approach enables coordinated control among the robot's base, arm, and camera, enhancing task efficiency. Through extensive simulations and real-world experiments, our approach significantly improves both the success rate and efficiency of mobile manipulation tasks, achieving a 95.6% success rate in real-world scenarios and a 52.8% increase in time efficiency.
Yixiang Jin, Jun Shi 0008, Yong A, Dingzhe Li, Fuchun Sun 0001, Dingsheng Luo, Bin Fang 0003
ICRA7
2025 An Intelligent Tennis Training Robot with Timely Motion Feedback
abstract
Tennis is widely popular across various age groups. However, mastering fluid stroke mechanics remains a significant challenge, requiring substantial time and practice. The lack of scientifically grounded tools for skill acquisition and training tools further hampers the development of tennis proficiency. In this paper, we investigate the application of an intelligent tennis robot equipped with advanced human-machine interaction capabilities, designed to serve as an effective tool for tennis learners, especially for enhancing motion geometry and coordination. First, we detail the design and development of the intelligent tennis robot, highlighting its core functions and operational principles, including ball serving, human motion sensing, and motion feedback mechanisms. Subsequently, we introduce a novel human motion analysis method that integrates motion geometry and kinematic analyses. Following this, we present a real-time motion feedback system that identifies deficiencies in players' movements, thereby facilitating the enhancement of their motion memory. Finally, we conduct experiments over players of varying skill levels, analyzing their motion patterns and providing practical examples. The proposed human-machine interaction framework offers a pioneering solution for intelligent tennis training, enabling players to understand their movements and correct errors immediately after each stroke.
Weiming Qu, Weizheng Chen, Dingsheng Luo
IROS3
2025 DPGP: A Hybrid 2D-3D Dual Path Potential Ghost Probe Zone Prediction Framework for Safe Autonomous Driving
abstract
Modern robots must coexist with humans in dense urban environments. A key challenge is the ghost probe problem, where pedestrians or objects unexpectedly rush into traffic paths. This issue affects both autonomous vehicles and human drivers. Existing works propose vehicle-to-everything (V2X) strategies and non-line-of-sight (NLOS) imaging for ghost probe zone detection. However, most require high computational power or specialized hardware, limiting real-world feasibility. Additionally, many methods do not explicitly address this issue. To tackle this, we propose DPGP, a hybrid 2D-3D fusion framework for ghost probe zone prediction using only a monocular camera during training and inference. With unsupervised depth prediction, we observe ghost probe zones align with depth discontinuities, but different depth representations offer varying robustness. To exploit this, we fuse multiple feature embeddings to improve prediction. To validate our approach, we created a 12K-image dataset annotated with ghost probe zones, carefully sourced and cross-checked for accuracy. Experimental results show our framework outperforms existing methods while remaining cost-effective. To our knowledge, this is the first work extending ghost probe zone prediction beyond vehicles, addressing diverse non-vehicle objects. We will open-source our code and dataset for community benefit.
Weiming Qu, Shenghai Yuan 0001, Shengyi Liu, Yuanhao Zhu, Jiayi Rao, Xihong Wu, Dingsheng Luo
IROS14
2025 Online Iterative Learning with Forward Simulation for Sub-minimum End-effector Displacement Positioning
abstract
Precision is a crucial performance indicator for robot arms. During interacting with human, high precision enables a robot arm to be used effectively and safely, while low precision may lead to safety issues. Traditional methods for improving robot arm precision rely on error compensation. However, these methods are often not robust and lack adaptability. Learning-based methods offer greater flexibility and adaptability, while current researches show that they often fall short in achieving high precision and struggle to handle many scenarios requiring high precision. In this paper, we propose a novel high-precision robot arm manipulation framework based on online iterative learning and forward simulation, which can achieve positioning error (precision) less than end-effector physical minimum displacement. In other words, our proposed method can compensate for the precision-limitation of the hardware structure of the robot arms. Furthermore, we consider the joint angular resolution of the real robot arm, which is usually neglected in related works. A series of experiments on both simulation and real UR3 robot arm platforms demonstrate that our proposed method is effective and promising. The related code will be available soon.
Weiming Qu, Tianlin Liu, Xihong Wu, Dingsheng Luo
IROS5
2025 SILM: A Subjective Intent Based Low-Latency Framework for Multiple Traffic Participants Joint Trajectory Prediction
abstract
Trajectory prediction is a fundamental technology for advanced autonomous driving systems and represents one of the most challenging problems in the field of cognitive intelligence. Accurately predicting the future trajectories of each traffic participant is a prerequisite for building high safety and high reliability decision-making, planning, and control capabilities in autonomous driving. However, existing methods often focus solely on the motion of other traffic participants without considering the underlying intent behind that motion, which increases the uncertainty in trajectory prediction. Autonomous vehicles operate in real-time environments, meaning that trajectory prediction algorithms must be able to process data and generate predictions in real-time. While many existing methods achieve high accuracy, they often struggle to effectively handle heterogeneous traffic scenarios. In this paper, we propose a Subjective Intent-based Low-latency framework for Multiple traffic participants joint trajectory prediction. Our method explicitly incorporates the subjective intent of traffic participants based on their key points, and predicts the future trajectories jointly without map, which ensures promising performance while significantly reducing the prediction latency. Additionally, we introduce a novel dataset designed specifically for trajectory prediction. Related code and dataset will be available soon.
Weiming Qu, Yuanhao Zhu, Xihong Wu, Dingsheng Luo
IROS9
2025 Real-Time Incremental Mapping and Degeneration-Aware Localization for Multi-Floor Parking Lots Based on IPM Image
abstract
In indoor parking lots, the use of RTK/GNSS for vehicle localization is often impractical due to the significantly smaller space compared to outdoor roads, which demands higher precision in both mapping and localization. Although feature point based visual SLAM algorithms have achieved high localization accuracy, they impose significant storage demands on embedded systems, and the visual feature point maps are not time-stable and are sensitive to lighting conditions. In this paper, we propose a real-time mapping and localization system for multi-floor parking lot. For the mapping part, we introduce a map-free SLAM method for precise ego-pose estimation, along with an efficient incremental map update framework that supports loop closure and multi-session mapping tasks. In the localization part, a semantic map is reused for vehicle localization based on bidirectional incentive descriptors. We incorporate degenerate cases into our optimization process, which greatly enhances the localization results. To the best of our knowledge, this is the first comprehensive system proposed for multi-floor parking lots. Experimental results demonstrate that our approach achieves state-of-the-art mapping and localization accuracy in multi-floor environments on embedded platforms.
Feng Youyang, Weiming Qu, Hongyao Wang, He Shizheng, Dingsheng Luo
IROS7
2025 DynaTra: A Dynamic Framework for Realistic and Scalable Trajectory Simulation in AV Platforms
abstract
Realistic and scalable trajectory simulation is a critical component in autonomous vehicle (AV) research. Existing methods trade off between computational efficiency and behavioral realism—either relying on full-stack autonomous driving systems with high overhead or simplified rule-based models lacking fidelity. In this paper, we propose DynaTra, a hybrid trajectory generation framework that dynamically switches between structured-scenario models—the Enhanced Intelligent Driver Model (EIDM) for autonomous vehicles and the Human Driver Model (HDM) for human drivers—and a full-stack system (Apollo 6.0) for unstructured or complex driving situations. A real-time model selection mechanism, informed by unstructured scenario detection, traffic density, and complex agent behavior recognition, enables adaptive and context-aware switching. Experimental evaluations in realistic simulation environments demonstrate that DynaTra significantly outperforms baseline IDM in both accuracy and efficiency and achieves comparable realism to Apollo 6.0 while maintaining substantially higher simulation speed. The proposed framework offers a principled path toward efficient and realistic AV simulation at scale.
Yige Wei, Ziyang Deng, Dingsheng Luo
RO-MAN4
2024 OW3Det: Toward Open-World 3D Object Detection for Autonomous Driving
abstract
Despite their success in LIDAR object detection, modern detectors are vulnerable to uncommon instances and corner cases (e.g., a runaway tire) since they are closed-set and static. Networks under the closed-set setup only predict labels of seen classes, while static models suffer from catastrophic forgetting when gradually learning novel concepts. This motivates us to formulate the open-world 3D object detection task for autonomous driving, which aims to 1) tackle the closed-set issue by identifying unseen instances as unknown and 2) incrementally learn novel classes without forgetting previously obtained knowledge. To achieve the open-world objectives, we propose Open-World 3D Detector (OW3Det), the first framework for open-world 3D object detection. The OW3Det comprises a base detector, a self-supervised unknown identifier, and a knowledge-distillation-restricted incremental learner. Although knowledge distillation facilitates preserving memories, imposing penalties on areas containing unknown objects hinders the incremental learning process. We mitigate this hindrance by employing unknown-driven pivotal mask, which eliminates unnecessary restrictions on regions overlapping with novel instances. Abundant experiments and visualizations demonstrate that the proposed OW3Det attains state-of-the-art performance.
Wenfei Hu, Weikai Lin, Hongyu Fang, Dingsheng Luo
IROS5
2024 Sensorimotor Coordinated Multi-UAV Coverage Path Planning
abstract
We propose a novel on-line multi-UAV coverage path planning method for 3D reconstruction. UAVs play homogeneous roles in most existing multi-UAV path planning methods, where their flexibility is not fully exploited. The fact that UAVs can explore with different missions and strategies at different heights simultaneously is not well considered. Therefore, we give UAVs different identities by presenting a global-local pattern, where global UAV locates the interesting regions while local UAV explores these regions in succession. Leveraging the mechanism of sensorimotor coordination, our strategy adjusts the paths of local UAV on-line according to the input from global UAV in order to gather more information and save time. In addition, our method is compatible with real-time SLAM systems. Simulation shows that our strategy takes less time than other strategies based on area-division, and indicates that it gathers more information for better reconstruction.
Hongyu Fang, Ziyang Deng, Dingsheng Luo
RO-MAN4
2024 Employing feature mixture for active learning of object detection
Licheng Zhang 0004, Siew-Kei Lam, Dingsheng Luo, Xihong Wu
Neurocomputing3
2023 Learning Clear Class Separation for Open-set 3D Detector in Autonomous Vehicle via Selective Forgetting
abstract
A trustworthy 3D detector is essential in the perception system of autonomous vehicles, ensuring accurate detection of their surroundings. However, autonomous vehicles have to operate in ever-changing real-world driving scenes, where unknown objects that do not belong to the training set are commonly encountered. Confusion about known and unknown objects could result in severe and dangerous consequences for road safety. To address this problem, we improve the reliability of autonomous driving systems by formulating open-set 3D object detection task. An Open-set 3D Detector (Open3Det) is proposed to reject unknown instances while maintaining performance on known categories. Distinct from 2D objects, clear space separation exists between each 3D instance. Motivated by this, we propose selective forgetting, a novel method capable of filtering out misleading predictions. Given a close-set teacher model, knowledge distillation is introduced to build a open-set student model. The student model preserves its predictions for known objects, whereas predictions of backgrounds and unknown instances are discarded to minimize misleading results. Extensive experiments and visualizations reveal the efficacy of the proposed method.
Wenfei Hu, Weikai Lin, Hongyu Fang, Dingsheng Luo
RO-MAN5
2023 Learning Clear Class Separation for Open-set 3D Detector in Autonomous Vehicle via Selective Forgetting
abstract
A trustworthy 3D detector is essential in the perception system of autonomous vehicles, ensuring accurate detection of their surroundings. However, autonomous vehicles have to operate in ever-changing real-world driving scenes, where unknown objects that do not belong to the training set are commonly encountered. Confusion about known and unknown objects could result in severe and dangerous consequences for road safety. To address this problem, we improve the reliability of autonomous driving systems by formulating open-set 3D object detection task. An Open-set 3D Detector (Open3Det) is proposed to reject unknown instances while maintaining performance on known categories. Distinct from 2D objects, clear space separation exists between each 3D instance. Motivated by this, we propose selective forgetting, a novel method capable of filtering out misleading predictions. Given a close-set teacher model, knowledge distillation is introduced to build a open-set student model. The student model preserves its predictions for known objects, whereas predictions of backgrounds and unknown instances are discarded to minimize misleading results. Extensive experiments and visualizations reveal the efficacy of the proposed method.
Wenfei Hu, Weikai Lin, Hongyu Fang, Dingsheng Luo
RO-MAN5
2023 A Novel Meta Control Framework for Robot Arm Reaching with Changeable Configuration
abstract
When deploying a robot to real-world environments, it is crucial to execute tasks amidst constantly changing surroundings. The conventional kinematics control of robot arms is primarily reliant on the inverse kinematics model. Unfortunately, due to the lack of adaptability, high-precision control models often falter when the robot utilizes tools of varying lengths or when the robot arm is worn out. This work aims to address this issue by proposing a meta-learning-based control framework. We achieve rapid and seamless online adaptation by updating control models when the robot arm’s configuration changes. The control framework comprises an Adaptive Global Inverse Model (Adaptive GIM) and an Adaptive Local Inverse Model (Adaptive LIM). The Adaptive GIM employs configuration-independent meta-learning, which allows the control model to swiftly adapt to different arm configurations. The Adaptive LIM adopts a meta-learning approach for location-independent training, enabling the robot to adapt to diverse local positions. As the Adaptive GIM suffers from the adverse effects stemming from the multiple solutions of inverse kinematics, utilizing the Adaptive LIM with relative position as input can alleviate this issue and enable more precise reaching towards the target. Extensive validation conducted on PKUHR6.0 demonstrates that the proposed approach significantly enhances online adaptation speed and precision compared to existing methods
Wenfei Hu, Dingsheng Luo
RO-MAN4
2023 Optimizing Robot Arm Reaching Ability with Different Joints Functionality
abstract
During the process of reaching a target, a robotic arm may generate a large number of redundant movements due to the stochastic nature of planning algorithms. We think that each joint of the robotic arm has a unique functionality towards the arm’s end effector. Therefore, during motion planning, we optimized the weights of each joint based on its individual functionality. We introduce the human arm operation mechanism - When a person is reaching for a distant target, the shoulder-near joints take on more actions. We proposed an advanced auxiliary loss function. By utilizing this function, we can filter out multiple solutions from the inverse model, thus reducing the amount of joint change and redundant movements. We conducted comprehensive comparative experiments on the UR3 robotic arm, and the results showed that our approach achieved better efficiency and outcomes.
Dingsheng Luo
RO-MAN4
2022 Advanced Face Anti-Spoofing with Depth Segmentation
abstract
Face anti-spoofing (FAS) plays a vital role in securing face recognition systems. In state-of-the-art FAS methods, face depth is determined for every position in a facial image. However, face depth varies at different positions, which leads to low accuracy when predicting face depth. As we observe, spoof faces have depth values that are 0, while live faces have depths that are equal to or greater than 0. As a result, for faces that have a depth greater than 0, if they are estimated as merely a positive value, instead of an accurate value, the prediction of whether they are real or not will not change. Further, if a range of depths are considered as one category, then there are more samples per category for the network training. Based on the above observation, in this paper, we propose to aggregate simple depth values to the same category and perform classification to optimize the FAS network. To evaluate the performance of the proposed approach, we perform extensive experiments on four benchmark databases, respectively, OULU-NPU, SiW, CASIA-FASD, and Replay-Attack. The results demonstrate that the proposed approach outperforms state-of-the-art methods on intra-database testing. Furthermore, our proposed approach shows advanced performance on cross-database testing.
Nan Sun 0002, Xihong Wu, Dingsheng Luo
IJCNN4
2019 A Hierarchical Model for StarCraft II Mini-Game
abstract
StarCraft II is one of the most challenging real-time strategy games, due to huge action space, large observation space, imperfect information, etc. Therefore, it is hard to learn the full game of StarCraft II. To reduce the learning complexity, DeepMind and Blizzard released several mini-games, in which the BuildMarines mini-game is most challenging, due to long time horizons, partially-observed state, high-dimensional, continuous action space and observation space. In this paper, we propose a hierarchical modeling method to solve those challenges in BuildMarines mini-game. Our approach consists of two levels, combining learning-based (high-level) and rule-based (low-level) method. The learning-based method leverages DQN reinforcement learning algorithm, while the rule-based method leverages script to realize. Experimental results show that the proposed approach is effective for an agent to learn the long planning horizon game, BuildMarines.
Tianlin Liu, Xihong Wu, Dingsheng Luo
ICMLA3
2017 A hierarchical inverse model based on proprioception and DNN for robot reaching
abstract
Robot reaching ability serves as one of essential basis for many other manipulation skills, such as grasping, placing etc., and has been widely focused for decades. Inverse model, which plays a fundamental role within robot reaching ability, aims to produce motor commands to drive the system to the desired state. However, since the inverse model is always an one-to-many mapping, it suffers from the multi-solution issue and the adaptation problem for changing circumstances when traditional kinematic/dynamic models are employed. And thus, learning based approaches are investigated, including the deep learning based models. In this paper, to further improve the performance of deep neural networks (DNN) methods, a novel model is proposed, where both the proprioception and a hierarchical structure are involved. Here, the employed concept of proprioception is aimed to follow human mechanism, while the hierarchical structure is expected to mimic the fact that different joints usually play different effects in a manipulation process. Experiments are performed with respect to PKU-HR6.0 II humanoid robot, and the results illustrate the effectiveness and superiority of the proposed model.
Tao Zhang 0071, Yian Deng, Jun-Hai Zhai, Xihong Wu, Dingsheng Luo
IECON6
2016 Biped robot falling motion control with human-inspired active compliance
abstract
Protecting robot from broken of falling is always a challenge issue for a bipedal humanoid robot in dealing with various locomotion related tasks to serve human society, especially as the assigned tasks turns increasingly complicated and the corresponding real environment gets more and more complex. Unlike several previous successful approaches on humanoid falling control, in this study, a new approach is suggested in the light of how human do when a fall happens. The proposed approach takes a tripod-like falling controller followed by a human-inspired active compliance strategy using joints' active flexion and torque increment. Therefore, robot falling action covers both stages of a fall, i.e. before and after landing impact, so that to reduce the fall damage as far as possible. And the tripod like posture prevents accumulation of kinetic energy, while the active compliance absorbs the impact energy in a tender way. Meanwhile, considering the complexity of robot dynamics, other than taking expert experiences, the proposed human-inspired falling control strategy is parametrically modelled and optimized with policy gradient reinforcement learning. Experiments on both simulation and real robot PKU-HR5.1 are performed, and the results demonstrate this approach is effective and promising.
Dingsheng Luo, Yian Deng, Xiaoqiang Han, Xihong Wu
IROS1
2015 Recognizing Human Activities from Raw Accelerometer Data Using Deep Neural Networks
abstract
Activity recognition from wearable sensor data has been researched for many years. Previous works usually extracted features manually, which were hand-designed by the researchers, and then were fed into the classifiers as the inputs. Due to the blindness of manually extracted features, it was hard to choose suitable features for the specific classification task. Besides, this heuristic method for feature extraction could not generalize across different application domains, because different application domains needed to extract different features for classification. There was also work that used auto-encoders to learn features automatically and then fed the features into the K-nearest neighbor classifier. However, these features were learned in an unsupervised manner without using the information of the labels, thus might not be related to the specific classification task. In this paper, we recommend deep neural networks (DNNs) for activity recognition, which can automatically learn suitable features. DNNs overcome the blindness of hand-designed features and make use of the precious label information to improve activity recognition performance. We did experiments on three publicly available datasets for activity recognition and compared deep neural networks with traditional methods, including those that extracted features manually and auto-encoders followed by a K-nearest neighbor classifier. The results showed that deep neural networks could generalize across different application domains and got higher accuracy than traditional methods.
Xihong Wu, Dingsheng Luo
ICMLA3
2014 Learning the Taxonomy of Function Words for Parsing
Dingsheng Luo, Xihong Wu
COLING3
2014 A Cyclic Contrastive Divergence Learning Algorithm for High-Order RBMs
abstract
The Restricted Boltzmann Machine (RBM), a special case of general Boltzmann Machines and a typical Probabilistic Graphical Models, has attracted much attention in recent years due to its powerful ability in extracting features and representing the distribution underlying the training data. A most commonly used algorithm in learning RBMs is called Contrastive Divergence (CD) proposed by Hinton, which starts a Markov chain at a data point and runs the chain for only a few iterations to get a low variance estimator. However, when referring to a high-order RBM, since there are interactions among its visible layers, the gradient approximation via CD learning usually becomes far from the log-likelihood gradient and even may cause CD learning to fall into an infinite loop with high reconstruction error. In this paper, a new algorithm named Cyclic Contrastive Divergence (CCD) is introduced for learning high-order RBMs. Unlike the standard CD algorithm, CCD updates the parameters according to each visible layer in turn, by borrowing the idea of Cyclic Block Coordinate Descent method. To evaluate the performance of the proposed CCD algorithm, regarding to high-order RBMs learning, both algorithms CCD and standard CD are theoretically analyzed, including convergence, estimate upper bound and both biases comparison, from which the superiority of CCD learning is revealed. Experiments on MNIST dataset for the handwritten digit classification task are performed. The experimental results show that CCD is more applicable and consistently outperforms the standard CD in both convergent speed and performance.
Dingsheng Luo, Xiaoqiang Han, Xihong Wu
ICMLA1
2014 Visual gesture recognition for human robot interaction using dynamic movement primitives
abstract
In this paper a method to address the efficiency and robustness of dynamic hand gesture recognition for human robot interaction is proposed. By using on-board monocular camera and specialized gesture detection algorithms, the humanoid robot is able to detect gestures fast. To model the dynamics of gestures, the dynamic movement primitives (DMP) model is employed, which well characterizes both spatial and temporal evolutions of gestures. The invariance properties of the DMP model against different spatiotemporal scales also offer expected robustness to handle the variances in gestures. To cope with the diversity and noise of gestures, an efficient adaptive DMP learning method is further proposed. Since the learnt weights of the DMP compactly represent the original gestures, they serve as ideal feature vectors for building a classifier to recognize new gestures. To evaluate the proposed method, a nine-class human gestures recognition task on a real humanoid robot is performed and 98.06% accuracy is obtained. Experimental results demonstrate the effectiveness of our method.
Dingsheng Luo, Xihong Wu
SMC3
2013 Discriminative Apprenticeship Learning with Both Preference and Non-preference Behavior
abstract
Considering that expert's demonstrations are usually sub optimal and failed demonstrations often have some useful guidance, in this paper, a Discriminative Apprenticeship Learning algorithm is proposed, where the apprentice is taught with the join of failed attempts to acquire the ability that could discriminate the preference and non-preference cases so that to actively take a corresponding action. Since robot usually encounters changing environments, generalization ability is taken into account in the algorithm through which the reward function is recovered under the evaluation of generalization error. The problem of the representation error is also analyzed and involved in the algorithm. To ensure performance of the algorithm, theoretical guarantee is presented. Experiments on a simple car-driving robot and the comparison with a variety of inverse reinforcement learning methods are performed, which illustrate the proposed method is an effective and promising alternative.
Dingsheng Luo, Xihong Wu
ICMLA (1)1
2012 A non-Gaussian factor analysis approach to transcription Network Component Analysis
abstract
Transcription factor activities (TFAs), rather than expression levels, control gene expression and provide valuable information for investigating TF-gene regulations. Network Component Analysis (NCA) is a model based method to deduce TFAs and TF-gene control strengths from microarray data and a priori TF-gene connectivity data. We modify NCA to model gene expression regulation by non-Gaussian Factor Analysis (NFA), which assumes TFAs independently comes from Gaussian mixture densities. We properly incorporate a priori connectivity and/or sparsity on the mixing matrix of NFA, and derive, under Bayesian Ying-Yang (BYY) learning framework, a BYY-NFA algorithm that can not only uncover the latent TFA profile similar to NCA, but also is capable of automatically shutting off unnecessary connections. Simulation study demonstrates the effectiveness of BYY-NFA, and a preliminary application to two real world data sets shows that BYY-NFA improves NCA for the case when TF-gene connectivity is not available or not reliable, and may provide a preliminary set of candidate TF-gene interactions or double check unreliable connections for experimental verification.
Shikui Tu, Dingsheng Luo, Runsheng Chen, Lei Xu 0001
CIBCB2
2008 Integrating Multi-level Linguistic Knowledge with a Unified Framework for Mandarin Speech Recognition
Jiazhong Nie, Dingsheng Luo, Xihong Wu
EMNLP3
2008 Exploiting prosodic and lexical features for tone modeling in a conditional random field framework
abstract
Tonal cues play an important role in distinguishing ambiguous words in Mandarin speech recognition. This paper explores an innovative tone modeling framework using prosodic and lexical features, as well as syllable context information. A discriminative model, namely a Conditional Random Field (CRF), is adopted, which is sufficiently flexible to handle multiple interacting features and long-range dependencies of observations. After the first pass search of a recognition system, the CRF based tone models are employed to rerank N-best hypotheses according to the tonal scores which can represent the correctness of the tone sequence given each candidate hypothesis and the observed speech signal. Experiments results show that the tonal cues help to achieve 7.8% and 8.6% relative reductions of character error rate on two widely used Mandarin speech recognition tasks, Hub-4 test and 863 test.
Hongxiu Wei, Dingsheng Luo, Xihong Wu
ICASSP4
2008 A Joint Segmenting and Labeling Approach for Chinese Lexical Analysis
Jiazhong Nie, Dingsheng Luo, Xihong Wu
ECML/PKDD (2)3
2007 Refine bigram PLSA model by assigning latent topics unevenly
abstract
As an important component in many speech and language processing applications, statistical language model has been widely investigated. The bigram topic model, which combines advantages of both the traditional n-gram model and the topic model, turns out to be a promising language modeling approach. However, the original bigram topic model assigns the same topic number for each context word but ignores the fact that there are different complexities to the latent semantics of context words. we present a new bigram topic model, the bigram PLSA model, and propose a modified training strategy that unevenly assigns latent topics to context words according to an estimation of their latent semantic complexities. As a consequence, a refined bigram PLSA model is reached. Experiments on HUB4 Mandarin test transcriptions reveal the superiority over existing models and further performance improvements on perplexity are achieved through the use of the refined bigram PLSA model.
Jiazhong Nie, Runxin Li, Dingsheng Luo, Xihong Wu
ASRU3
2003 Refine decision boundaries of a statistical ensemble by active learning
abstract
For pattern classification, the decision boundaries are gradually constructed in a statistical ensemble through a divide-and-conquer procedure based on resampling techniques. Hence a resampling criterion critically governs the process of forming the final decision boundaries. Motivated by active learning ideas, we propose an alternative resampling criterion based on the zero-one loss measure in this paper, where all the patterns in the training set are ranked in terms of their "difficulty" for classification no matter whether a pattern has been incorrectly classified or not. Our resampling criterion incorporated by Adaboost has been applied to benchmark handwritten digit recognition and text-independent speaker identification tasks. Comparative results demonstrate that our method refines decision boundaries and therefore yields the better generalization performance.
Dingsheng Luo
IJCNN1
2003 Biomimetics speaker identification systems for network security gatekeepers
abstract
The perception mechanisms of the human auditory periphery and cochlear nucleus were simulated and the potential application to the voice password gatekeeper was discussed. A biomimetics speaker identification system was implemented based on the auditory processing. Obvious improvement in the robustness was shown under a noisy environment.
Xihong Wu, Dingsheng Luo, Huisheng Chi, Harold Szu
IJCNN2