Lei Shi 0013

dblp:29/563-13 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
2since 2021 · last 2022
0000-0002-5215-025XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 2 since 2021Systems, architecture and hardware · 8 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Autonomous driving · 30% Video understanding and tracking · 30% Image recognition and object detection · 13%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Autonomous driving
trajectory prediction
0.612022
Exact-likelihood User Intention Estimation for Scene-compliant Shared-control Navigation · ICRA 2022
Computer vision › Video understanding and tracking › object tracking › 3d object tracking
3d human tracking
0.312017
Real-time 3D human tracking for mobile robots with multisensors · ICRA 2017
Computer vision › Video understanding and tracking › object tracking
person tracking
0.312017
Real-time 3D human tracking for mobile robots with multisensors · ICRA 2017
Computer vision › Image recognition and object detection
scene recognition
0.212016
Understand scene categories by objects: A semantic regularized scene classifier using Convolutional Neural Networks · ICRA 2016
Computer vision › Segmentation and scene understanding
semantic segmentation
0.212016
Understand scene categories by objects: A semantic regularized scene classifier using Convolutional Neural Networks · ICRA 2016
Computer vision › 3D vision
environmental understanding
0.112016
Understand scene categories by objects: A semantic regularized scene classifier using Convolutional Neural Networks · ICRA 2016

Methods — techniques the papers use, named apart from their topics

normalizing flow · 0.6generative model · 0.6exact likelihood estimation · 0.6visual tracking · 0.3gaussian process regression · 0.3extended kalman filter · 0.3semantic regularization · 0.2convolutional neural network · 0.2
YearPublicationVenuePosition
2022 Exact-likelihood User Intention Estimation for Scene-compliant Shared-control Navigation
abstract
A predictive model for mobility systems capable of understanding the trajectory a user intends to follow in the environment is proposed. Understanding user intention is paramount for any shared-control navigation strategy between a user and an active robotic agent. Equally important however is being able to go beyond simple sample generation to assign probabilistic meaning to the set of possible future trajectories, so most likely scenarios can be assumed. The framework estimates a distribution over possible intentions, proposing a novel generative model predicated on Normalizing Flows which accounts for past behaviours, as traditionally reported in the literature, but also incorporates visual scene information. As the model permits trajectories to be assigned exact likelihoods, tractable density estimates can be readily exploited to finalize an executable intention. Baseline comparisons with the publicly available and widely used KITTI navigational dataset show significant improvements (up to 11.08%) with respect to traditional metrics such as Average and Final Displacement Errors. A novel metric that stands independent of the number of samples is also proposed as a more fitting comparison for future works.
Kavindie Katuwandeniya, Stefan H. Kiss, Lei Shi 0013, Jaime Valls Miró
ICRA3
2021 Multi-modal Scene-compliant User Intention Estimation in Navigation
abstract
A multi-modal framework to generate user intention distributions when operating a mobile vehicle is proposed in this work. The model learns from past observed trajectories and leverages traversability information derived from the visual surroundings to produce a set of future trajectories, suitable to be directly embedded into a perception-action shared control strategy on a mobile agent, or as a safety layer to supervise the prudent operation of the vehicle. We base our solution on a conditional Generative Adversarial Network with Long-Short Term Memory cells to capture trajectory distributions conditioned on past trajectories, further fused with traversability probabilities derived from visual segmentation with a Convolutional Neural Network. The proposed data-driven framework results in a significant reduction in error of the predicted trajectories (versus the ground truth) from comparable strategies in the literature (e.g. Social-GAN) that fail to account for information other than the agent’s past history. Experiments were conducted on a dataset collected with a custom wheelchair model built onto the open-source urban driving simulator CARLA, proving also that the proposed framework can be used with a small, unannotated dataset.
Kavindie Katuwandeniya, Stefan H. Kiss, Lei Shi 0013, Jaime Valls Miró
IROS3
2018 Towards open-set semantic labeling in 3D point clouds : Analysis on the unknown class
Huifang Ma, Rong Xiong, Yue Wang 0020, Sarath Kodagoda, Lei Shi 0013
Neurocomputing5
2017 Gaussian process model enabled particle filter for device-free localization
abstract
Device-free localization (DFL) is an emerging wireless network target localization technique that does not need to attach any electronic device with the target. It is remaining as a challenging research problem due to the weak wireless signals and the uncertain wireless communication environment. In this paper, a novel Gaussian Process (GP) based wireless propagation model is proposed to describe the likelihood relationship between the target location and the changes of the RSS measurement for a wireless link. Sequentially Particle Filter (PF) is applied to the DFL for estimating the location of the target, after the GP model is trained using the experimental measurements of the link. Experimental results demonstrate that the proposed GP-PF algorithm can track the target with much better localization accuracy than the Support Vector Machine (SVM) based PF approach.
Biao Song, Wendong Xiao, Shoudong Huang, Lei Shi 0013
FUSION5
2017 Real-time 3D human tracking for mobile robots with multisensors
abstract
Acquiring the accurate 3-D position of a target person around a robot provides fundamental and valuable information that is applicable to a wide range of robotic tasks, including home service, navigation and entertainment. This paper presents a real-time robotic 3-D human tracking system which combines a monocular camera with an ultrasonic sensor by the extended Kalman filter (EKF). The proposed system consists of three sub-modules: monocular camera sensor tracking model, ultrasonic sensor tracking model and multi-sensor fusion. An improved visual tracking algorithm is presented to provide partial location estimation (2-D). The algorithm is designed to overcome severe occlusions, scale variation, target missing and achieve robust re-detection. The scale accuracy is further enhanced by the estimated 3-D information. An ultrasonic sensor array is employed to provide the range information from the target person to the robot and Gaussian Process Regression is used for partial location estimation (2-D). EKF is adopted to sequentially process multiple, heterogeneous measurements arriving in an asynchronous order from the vision sensor and the ultrasonic sensor separately. In the experiments, the proposed tracking system is tested in both simulation platform and actual mobile robot for various indoor and outdoor scenes. The experimental results show the superior performance of the 3-D tracking system in terms of both the accuracy and robustness.
Mengmeng Wang 0005, Daobilige Su, Lei Shi 0013, Yong Liu 0007, Jaime Valls Miró
ICRA3
2016 Understand scene categories by objects: A semantic regularized scene classifier using Convolutional Neural Networks
abstract
Scene classification is a fundamental perception task for environmental understanding in today's robotics. In this paper, we have attempted to exploit the use of popular machine learning technique of deep learning to enhance scene understanding, particularly in robotics applications. As scene images have larger diversity than the iconic object images, it is more challenging for deep learning methods to automatically learn features from scene images with less samples. Inspired by human scene understanding based on object knowledge, we address the problem of scene classification by encouraging deep neural networks to incorporate object-level information. This is implemented with a regularization of semantic segmentation. With only 5 thousand training images, as opposed to 2.5 million images, we show the proposed deep architecture achieves superior scene classification results to the state-of-the-art on a publicly available SUN RGB-D dataset. In addition, performance of semantic segmentation, the regularizer, also reaches a new record with refinement derived from predicted scene labels. Finally, we apply our model trained on SUN RGB-D dataset to a set of images captured in our university using a mobile robot, demonstrating the generalization ability of the proposed algorithm.
Yiyi Liao, Sarath Kodagoda, Yue Wang 0020, Lei Shi 0013, Yong Liu 0007
ICRA4
2016 Constrained sampling of 2.5D probabilistic maps for augmented inference
abstract
This work exploits modeling spatial correlation in 2.5D data using Gaussian Processes (GPs), and produces constrained sampling realizations on these models to improve certainty in the predictions by means of integrating additional sparse information. Data organized in 2.5D such as elevation and thickness maps has been extensively studied in the fields of robotics and geostatistics. These maps are typically represented as a probabilistic 2D grid that stores an estimated value (height or thickness) for each cell. With the increasing popularity and deployment of robotic devices for infrastructure inspection, 2.5D data becomes a common interpretation of the condition of the target being inspected. Modeling the spatial dependencies and making inferences on new grid locations is a common task that has been addressed using GPs, but inference results on locations which are weakly correlated with the training data are generally not sufficiently informative and distinctly uncertain. The predictive capability of the proposed framework, which is applicable to any 2.5D data, is demonstrated with field inspection data from pipelines. Specifically, sparse and complementary measurements from alternative sensing modalities have been incorporated into the model to predict in more detail local thickness conditions where GP training data is limited. The output of this work aims to probabilistically present variations of the target in the case that both accuracy and reasonable diversity are of significant interest.
Lei Shi 0013, Jaime Valls Miró, Teng Zhang 0003, Teresa Vidal-Calleja, Liye Sun, Gamini Dissanayake
IROS1
2013 Towards simultaneous place classification and object detection based on conditional random field with multiple cues
abstract
Simultaneous place classification and object detection (SPCOD) is an algorithm which is able to categorize the environment (place) and detect the objects presented in the environment. Although both place classification and object detection problems have been in discussion in literature, as a concept SPCOD is still in its early stage of research. Focusing mainly on the discrimination ability of SPCOD, in this paper we have proposed a pairwise conditional random field (CRF) framework to integrate mature techniques on laser data based place classification and vision based off-the-shelf object descriptor. Extensive experimental results on a public data set demonstrate the effectiveness of the proposed method.
Lei Shi 0013, Sarath Kodagoda, Massimo Piccardi
IROS1
2012 Application of CRF and SVM based semi-supervised learning for semantic labeling of environments
abstract
Understanding the environment in both geometric and semantic levels enables a robot to perform high-level tasks in complex environments. Therefore in recent years research towards identifying and semantically labeling the environments based on onboard sensors for mobile robots has been gaining popularity. After the era of heuristic and rule-based approaches, supervised learning algorithms like Support Vector Machines (SVM) and AdaBoost have been extensively used for this purpose showing satisfactory performance. With the introduction of graphical models, approaches like Conditional Random Fields (CRF) which take the advantage of connectivity of samples provide more flexibility to capture complex dependencies. In this paper, we focus on a real-world task which challenges the generalization ability of the model, evaluate some graph based features, propose a semi-supervised learning algorithm by iteratively utilizing the results from SVM and CRF, and suggest a solution for CRF parameter estimation with partially labeled training data. Experiments have been conducted on six real-world indoor environments demonstrating the competence of the algorithm.
Lei Shi 0013, Rami N. Khushaba, Sarath Kodagoda, Gamini Dissanayake
ICARCV1
2012 Application of semi-supervised learning with Voronoi Graph for place classification
abstract
Representation of spaces including both geometric and semantic information enables a robot to perform high-level tasks in complex environments. Therefore, in recent years identifying and semantically labeling the environments based on onboard sensors has become an important competency for mobile robots. Supervised learning algorithms have been extensively used for this purpose with SVM-based solutions showing good generalization properties. The CRF-based approaches take the advantage of connectivity information of samples thereby provide a mechanism to capture complex dependencies. Blending the complementary strengths of Support Vector Machine (SVM) and Conditional Random Field (CRF), there have been algorithms to exploit the advantages of both to enhance the overall accuracy of place classification in indoor environments. However, experiments show that none of the above approaches deal well with diversified testing data. In this paper, we focus mainly on the generalization ability of the model and propose a semi-supervised learning strategy, which essentially improves the performance of the system. Experiments have been carried out on six real-world maps from different universities around the world and the results from rigorous testing demonstrate the feasibility of the approach.
Lei Shi 0013, Sarath Kodagoda, Gamini Dissanayake
IROS1
2010 Multi-class classification for semantic labeling of places
abstract
Human robot interaction is an emerging area of research, where human understandable robotic representations can play a major role. Knowledge of semantic labels of places can be used to effectively communicate with people and to develop efficient navigation solutions in complex environments. In this paper, we propose a new approach that enables a robot to learn and classify observations in an indoor environment using a labeled semantic grid map, which is similar to an Occupancy Grid like representation. Classification of the places based on data collected by laser range finder (LRF) is achieved through a machine learning approach, which implements logistic regression as a multi-class classifier. The classifier output is probabilistically fused using independent opinion pool strategy. Appealing experimental results are presented based on a data set gathered in various indoor scenarios.
Lei Shi 0013, Sarath Kodagoda, Gamini Dissanayake
ICARCV1
2010 Laser range data based semantic labeling of places
abstract
Extending metric space representations of an environment with other high level information, such as semantic and topological representations enable a robotic device to efficiently operate in complex environments. This paper proposes a methodology for a robot to classify indoor environments into semantic categories. Classification task, using data collected from a laser range finder, is achieved by a machine learning approach based on the logistic regression algorithm. The classification is followed by a probabilistic temporal update of the semantic labels of places. The innovation here is that the new algorithm is able to classify parts of a single laser scan into different semantic labels rather than the conventional approach of gross categorization of locations based on the whole laser scan. We demonstrate the effectiveness of the algorithm using a data set available in the public domain.
Lei Shi 0013, Sarath Kodagoda, Gamini Dissanayake
IROS1