VLDB 2026 Research / reviewers in the wild / expert
Guohui Tian
dblp:47/1234
· DBLP profile ↗
39ranked-venue papers
0as first author
29since 2021 · last 2026
0000-0001-8332-3064ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GRHP: Graph-Fused Hierarchical Planning for Embodied Long-Horizon Robotic Task
Guohui Tian, Yongcheng Cui, Xuyang Shao 0002 |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Object search strategy for service robots with knowledge-based viewpoint selection and hierarchical action decisions
Guohui Tian |
Expert Syst. Appl. | 2 |
| 2026 | Approaching for manipulation: Robot termination pose generation via active object detection and task-oriented logical reasoning
Guohui Tian, Xuyang Shao 0002 |
Expert Syst. Appl. | 2 |
| 2026 | Temporal adaptive Bayesian search-driven embodied task sequence optimization based on 3D dynamic scene graphs
Guohui Tian |
Knowl. Based Syst. | 2 |
| 2026 | Semantically Guided Task Planning: Supervised Vision-Language-Action Model by Large Language ModelsabstractEnabling robots to perform everyday tasks has become increasingly important. Task planning, which decomposes task instructions into executable action sequences, is crucial for equipping robots with the ability to handle daily activities. Currently, there are two main effective methods for task planning: one relies on the reasoning capabilities of Large Language Models (LLMs), but it struggles with handling the underlying motion. The other is based on the generative capabilities of Vision-Language-Action (VLA) model, which often lacks essential semantic details. To overcome these limitations, this paper introduces a novel Semantically Supervised Vision-Language-Action (SS-VLA) model. This model addresses the constraints of previous method that relied solely on single-frame image by designing an adaptive visual sequence encoder that integrates continuous visual streams. This encoder efficiently captures and integrates multi-scale spatial and temporal features from the robot’s first-person visual perspective. Furthermore, the model utilizes LLMs to decompose task instructions into subtasks and organize them into graph structure, using Graph Attention Network (GAT) to extract features from subtask sequences and supervise the generation of action sequences. This method not only enhances the alignment of actions with task instructions but also ensures the contextual and semantic accuracy of the robot’s activities, significantly enhancing the task execution capabilities of robots in complex environments. We evaluated our model on the ALFRED and TEACh benchmark, achieving higher performance compared to existing methods, especially in unseen scenes. Additionally, we successfully deployed our model in the AI2-THOR virtual environment and on the TIAGo real robot, demonstrating the effectiveness of our method. Our code is available at: https://github.com/Li-XD-Pro/SS-VLA. Guohui Tian, Yongcheng Cui |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | SenseMax: IoT-Based User-Device Interaction Prediction With Complementary Masked Modeling and Contrastive Learning
Tengfan Fu, Fei Lu 0005, Guohui Tian |
IEEE Internet Things J. | 4 |
| 2025 | Goal-Oriented Visual Semantic Navigation Using Semantic Knowledge Graph and TransformerabstractWhen determining navigation actions, it is important to design effective visual and semantic representations of the observation scenes and robust navigation strategies. The paper proposes a goal-oriented visual semantic navigation method using semantic knowledge graph and transformer. Two kinds of knowledge graphs representing the location relationship between objects are constructed, namely current knowledge graph and prior knowledge graph. The pre-constructed prior knowledge graph is periodically updated by the current knowledge graph obtained in real time, and embedded into the semantic feature vector through graph convolutional network (GCN). The semantic features and extracted scene features are jointly embedded and stored, they are jointly fed into the transformer module to explore the spatio-temporal dependencies between objects in the environment. The navigation strategy is obtained from the Asynchronous Advantage Actor-Critic (A3C) model composed of Long-Short Term Memory (LSTM) and Multi-Layer Perception (MLP). Experiments show that the knowledge graph can significantly improve the navigation performance. More importantly, our experimental results show that our method can improve the generalization ability of navigation across novel scenes and novel objects. Video can be available at https://youtu.be/ZMjNvoK2rbY.Note to Practitioners— The motivation of this work is to develop an efficient visual semantic navigation method. Conventional navigation algorithms lack semantic information and learning ability, and can not adapt to the complex unknown environments. When semantic information is included in navigation, the location relationship between objects can be obtained as a prior knowledge, which can be combined with reinforcement learning to achieve autonomous navigation of agents. In this article, a knowledge graph representing the location relationships between objects has been constructed and regularly updated in real-time. The proposed visual semantic navigation method further improves the generalization ability of navigation. This navigation method can be applied to mobile robots and deployed in many scenarios such as home, restaurant, hospitals, and even factories. Guohui Tian |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Transformer-Driven Semantic-Spatial Adaptive Fusion Representation for Object-Goal NavigationabstractVisual object-goal navigation requires an agent to make decisions to search for specified target objects within an unknown environment. While learning-based approaches have achieved progress, they still face two limitations: (1) visual representation: the lack of a low-dimensional representation that balances object priors with current observations, and the absence of effective modeling of scene spatial structures, which restricts the agent’s ability to perceive its surroundings. (2) policy learning: existing reward functions often fail to consider the visual essence of object goal-driven tasks, leading to inefficient learning and suboptimal performance. To address these issues, this paper introduces a semantic-spatial fusion representation framework that incorporates both object semantics and the intrinsic spatial structure of the scene. Specifically, a new object context matrix captures semantic relationships between objects while providing distinct low-dimensional representations for different observations, and an episode memory graph is also constructed and dynamically updated based on observation similarity and geodesic distance to represent the spatial structure of the agent’s environment in real-time. Then, a transformer-based adaptive multimodal feature fusion module is proposed to integrate these dual representations. Moreover, a sparse reward function is designed based on the target’s bounding box to guide the agent to learn correctly. The proposed method is the first to be evaluated on both the simulation platform AI2-THOR and the real-world dataset AVD, demonstrating its generalization in unseen environments. The method is also deployed on a physical mobile robot and tested in real-world scenarios, further validating its practical effectiveness. Fei Lu 0005, Xiaolei Li 0003, Guohui Tian, Tengfan Fu |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Efficiency-Driven Adaptive Task Planning for Household Robot Based on Hierarchical Item-Environment CognitionabstractTask planning focused on household robots represents a conventional yet complex research domain, necessitating the development of task plans that enable robots to execute unfamiliar household services. This area has garnered significant research interest due to its extensive applications in robotics, particularly concerning household robots. Nevertheless, the majority of task planning methodologies exhibit suboptimal performance regarding the success and efficiency of completing household tasks, primarily due to a lack of cognitive capacity of household items and home environments. To address these challenges, we propose an efficiency-driven adaptive task planning approach based on hierarchical item-environment cognition. Initially, we establish a multiple semantic attribute-based priori knowledge (MSAPK) framework to facilitate the attributive representation of household items. Utilizing MSAPK, we develop a long short-term memory (LSTM) based item cognition model that assigns relevant attributes and substitutes to specified household items, thereby enhancing the cognitive capabilities of household robots at the attribute level. Subsequently, we construct an environment cognition model that delineates the relationships between household items and room types, enabling household robots to locate target items more efficiently. Through hierarchical item-environment cognition, we introduce a strategy for adaptive task planning, empowering household robots to execute household tasks with both flexibility and efficiency. The generated plans are evaluated in both virtual and real-world experiments, with promising results affirming the effectiveness of our proposed methodology. Mengyang Zhang, Guohui Tian, Yongcheng Cui, Hong Liu 0013, Lei Lyu 0001 |
IEEE Trans. Cybern. | 2 |
| 2025 | GMM Enabled by Multimodal Information Fusion Network for Detection and Motion Planning of Robotic Liquid PouringabstractWhen humans perform pouring tasks, they exhibit consistent accuracy, regardless of the liquid type, container, or environmental conditions. This proficiency stems from their ability to effectively utilize both vision and hearing while also considering various factors. However, in the domain of robotic liquid pouring, the combination of multimodal information is effectively rarely leveraged to accomplish automatic control of robotic liquid pouring. To address this limitation, a multimodal information fusion network (MMFNet) is designed for estimating liquid height and pouring state. The MMFNet employs cross-attention networks and motion features to enhance visual features (VFs). Subsequently, multimodal transformers are utilized to fuse audio features with the enhanced VFs, enabling the MMFNet to estimate both liquid height and pouring state accurately. Finally, the detection results are combined with demonstration learning to make robots learn pouring motion trajectory encoded by the Gaussian mixture model (GMM). The experimental results demonstrate the effectiveness of MMFNet in significantly improving the detection accuracy of liquid height and pouring state. Furthermore, by employing the GMM enabled by MMFNet, robots can acquire robust pouring motion planning, enhancing their capabilities in performing pouring tasks. Guohui Tian, Shijie Guo |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Transformer-Based Relationship Inference Model for Household Object Organization by Integrating Graph Topology and OntologyabstractIn domestic environments, the conventional organization of objects by service robots often relies on the inherent properties of each object, such as placing fragile bowls in enclosed cupboards. However, this approach tends to overlook the importance of the orderly arrangement of objects, neglecting the specific placement order of bowls within the cabinet. In practice, effective object organization necessitates consideration of both individual properties and the relationships defined by these properties. In this paper, we have constructed a specialized dataset encompassing the ontological properties of household objects along with their relationships. Furthermore, we have introduced a graph-based model to explicitly represent these relationships and proposed a novel feature extraction technique that integrates the Graph Attention Network (GAT) with the BERT model to predict the relationships among objects. Subsequently, we utilized the Transformer framework to train a model, enabling it to infer relationships between objects. Experimental validation demonstrates the effectiveness of our approach in accurately predicting relationships between household objects, thus facilitating their orderly organization. Our approach significantly augments the object organization capabilities for service robots by accurately predicting the relationships among household objects. Our code is available at: https://github.com/Li-XD-Pro/Household-Object-Organization Guohui Tian, Yongcheng Cui |
IROS | 2 |
| 2024 | An Active Task Cognition Method for Home Service Robot Using Multi-Graph Attention Fusion MechanismabstractActive Task Cognition (ATC) requires the robot to comprehend the current scene using the image within the field of view, enabling them to reason about appropriate and executable tasks, thus allowing the robot to achieve service task scene discovery capability similar to humans. This capability is paramount for robots to provide comfort and intelligent service while performing their tasks. To enhance home service robots’ ATC capability, a multi-graph fusion mechanism based on Graph Attention Network (GAT) is proposed in this paper to model the semantic feature related to the task. First, a multi-graph fusion encoder is proposed to maximally capture the integrated features of objects, tasks, and scenes from the images, thereby obtaining a semantic representation related to the home service task from the robot’s perspective. Next, to enhance the interpretability of the model, we propose a multi-task scene understanding decoder based on the attention mechanism to utilize the integration features of multi-graph fusion efficiently. Lastly, we present a loss function for multi-task scene understanding in the proposed Encoder-Decoder network model for scene comprehension. Furthermore, a new dataset comprising various daily household tasks is constructed in the experiments. Extensive experimental results indicate that the proposed method significantly enhances the robot’s active cognitive abilities in service tasks, empowering it with advanced levels of intelligence. Yongcheng Cui, Guohui Tian, Zhengsong Jiang, Mengyang Zhang, Yifei Wang 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Task-Oriented Robot Cognitive Manipulation Planning Using Affordance Segmentation and Logic ReasoningabstractThe purpose of task-oriented robot cognitive manipulation planning is to enable robots to select appropriate actions to manipulate appropriate parts of an object according to different tasks, so as to complete the human-like task execution. This ability is crucial for robots to understand how to manipulate and grasp objects under given tasks. This article proposes a task-oriented robot cognitive manipulation planning method using affordance segmentation and logic reasoning, which can provide robots with semantic reasoning skills about the most appropriate parts of the object to be manipulated and oriented by tasks. Object affordance can be obtained by constructing a convolutional neural network based on the attention mechanism. In view of the diversity of service tasks and objects in service environments, object/task ontologies are constructed to realize the management of objects and tasks, and the object-task affordances are established through causal probability logic. On this basis, the Dempster-Shafer theory is used to design a robot cognitive manipulation planning framework, which can reason manipulation regions' configuration for the intended task. The experimental results demonstrate that our proposed method can effectively improve the cognitive manipulation ability of robots and make robots preform various tasks more intelligently. Guohui Tian |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | A semantic robotic grasping framework based on multi-task learning in stacking scenes
Shengqi Duan, Guohui Tian, Chenrui Feng |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Sequential Learning for Ingredient Recognition From ImagesabstractTo incorporate the cooking logic into ingredient recognition from food images is beneficial for food cognition. Compared with food categorization, ingredient recognition gives a better understanding on food cognition, by providing crucial information on food compositions. However, there exist situations in which different food are made of different ingredients, thus it is necessary to incorporate cooking logic into ingredient recognition to achieve a better food cognition. Based on this point, our paper proposes a sequential learning method to guide a neural network based (NN-based) model on producing ingredients following the corresponding cooking logic in recipes. Firstly, in order to make a maximum utilization of visual features from images, a double-flow feature fusion module (DFFF) is proposed to obtain features from two image-based, visual tasks (food name proposal and multi-label ingredient proposal). After that, fused features from DFFF, together with original image features, are feed into a bidirectional long short time memory (Bi-LSTM) based ingredient generator to produce sequential ingredients. To guide the sequential ingredient generation process, reinforcement learning is employed by designing a hybrid loss related to both the common and personality traits in ingredients for optimizing the model ability of associating images and sequential ingredients. In addition, sequential ingredients are utilized in a backward flow by reconstructing food images, so that sequential ingredient generation can be further optimized in a complementary manner. In experiments, the results demonstrate the superiority of our method on driving the model to allocate more attention to the correlation between images and sequential ingredients, and produced ingredients are comprehensive and logical. Mengyang Zhang, Guohui Tian, Ying Zhang 0043, Hong Liu 0013 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Behavior Cloning-Based Robot Active Object Detection With Automatically Generated Data and Revision MethodabstractActive object detection (AOD), one of the greatest challenges in the robotics field, is the main focus of this article. Most current AOD methods are developed by reinforcement learning (RL) algorithms while they can be further improved in the aspects of training time, training efficiency, model performance, and model prediction. Therefore, different from the existing works, we propose an AOD method based on behavior cloning trained by automatically generated data. We transform the AOD task into an action classification problem to not only shorten the training time but also improve the training efficiency and model performance. As there is no available expert data for training the presented classification-based AOD model, we design an autonomous method of data generation to avoid the large amounts of manual annotations. We introduce a multiinput network for better obstacle avoidance and AOD performance, where the depth image is added to help the robot to perceive distance information of environments and objects. Moreover, we develop a revision method for model prediction to reduce the accumulation of compounding error, which improves the successful rate of the long path AOD tasks effectively. We extensively evaluate our method on an AOD dataset by the comparable experiments and the ablation study, proving that our approach outperforms other methods in AOD performance and efficiency. In addition, the AOD experiments in the real-world scenario with a TIAGo robot indicate the validity of our method. Guohui Tian, Xuyang Shao 0002 |
IEEE Trans. Robotics | 2 |
| 2022 | A comprehensive survey on 3D face recognition methods
Guohui Tian |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | Service planning oriented efficient object search: A knowledge-based framework for home service robot
Guohui Tian, Ying Zhang 0043, Mengyang Zhang |
Expert Syst. Appl. | 2 |
| 2022 | Safe distance prediction for braking control of bridge cranes considering anti-swingabstractCranes are widely deployed for lifting and moving heavy objects in dynamic environments with human coexistence. Suddenly appeared workers, vehicles, and robots can affect the safety of the cranes. To avoid possible collisions, the cranes must have prediction ability to know how dangerous the situation is. In this paper, we address the safety issues of bridge cranes based on its online physical states and control model. Due to the swing of the payload, the safe braking distance cannot be a constant value. Therefore, we here propose a model prediction control (MPC)-based anti-swing method for non-zero initial states, where a new reference trajectory and a new cost function for optimization are proposed, such that the proposed MPC method can control the crane to follow the proposed reference trajectory and achieve a stable stop state with anti-swing. Furthermore, an offline learning mechanism is introduced to learn a statistical model between the velocity of the crane and the safe braking distance achieved by using the proposed MPC braking control method. In this way, we can predict how far the crane would require to safely stop without swing based on its current velocity, which is the safe distance prediction to evaluate the dangerous level of the dynamic obstacle. Experiments using both a simulated crane and a real crane demonstrate that the proposed safe braking distance prediction method is effective for safe braking control of the bridge cranes. Huili Chen, Guohui Tian, Jianhua Zhang 0010, Ze Ji |
Int. J. Intell. Syst. | 3 |
| 2022 | Hybrid offline and online task planning for service robot using object-level semantic map and probabilistic inference
Guohui Tian |
Inf. Sci. | 2 |
| 2022 | A deep Q-learning network based active object detection model with a novel training algorithm for service robotsabstractThis paper focuses on the problem of active object detection (AOD). AOD is important for service robots to complete tasks in the family environment, and leads robots to approach the target object by taking appropriate moving actions. Most of the current AOD methods are based on reinforcement learning with low training efficiency and testing accuracy. Therefore, an AOD model based on a deep Q-learning network (DQN) with a novel training algorithm is proposed in this paper. The DQN model is designed to fit the Q-values of various actions, and includes state space, feature extraction, and a multilayer perceptron. In contrast to existing research, a novel training algorithm based on memory is designed for the proposed DQN model to improve training efficiency and testing accuracy. In addition, a method of generating the end state is presented to judge when to stop the AOD task during the training process. Sufficient comparison experiments and ablation studies are performed based on an AOD dataset, proving that the presented method has better performance than the comparable methods and that the proposed training algorithm is more effective than the raw training algorithm. Guohui Tian, Yongcheng Cui, Xuyang Shao 0002 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2022 | Efficient 3D CNNs with knowledge transfer for sign language recognition
Xiangzu Han, Fei Lu 0005, Guohui Tian |
Multim. Tools Appl. | 3 |
| 2022 | Autonomous Generation of Service Strategy for Household Tasks: A Progressive Learning Method With A Priori Knowledge and Reinforcement LearningabstractHuman beings tend to learn unknown knowledge in a gradual process, from the basic to the complex. Based on this point, we propose a progressive learning method for producing service strategies according to requests, with a hierarchical priori knowledge and reinforcement learning. Service strategy aims to guide how to perform home services and takes into consideration the relationship between actions and objects in home environment. In this paper, strategy generation is regarded as a text generation problem in question answering (QA). Firstly, a hierarchical priori knowledge with service-object correlation at the bottom and action-object correlation at the top is constructed to assist the understanding on the relationship of objects and actions in service strategies. Service-object correlation guides how to select proper objects with the correct order, while action-object correlation associates actions in strategies according to selected objects. Based on the hierarchical priori knowledge, a progressive learning method is proposed to make the model produce effective strategies with a sequential cognition, from service-object correlation (objects) to action-object correlation (actions). After that, reinforcement learning is employed to enhance the progressive guidance, by designing rewards in terms of the hierarchical priori knowledge. Finally, the proposed method is tested with both comparative experiments and ablation studies, and the experimental results demonstrate the superiority in producing comprehensive and logical strategies, indicating that the progressive learning method in our paper can further improve the QA performance. Mengyang Zhang, Guohui Tian, Huanbing Gao, Ying Zhang 0043 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Sign Language Recognition Based on R(2+1)D With Spatial-Temporal-Channel AttentionabstractPrevious work utilized three-dimensional (3-D) convolutional neural networks (CNNs) tomodel the spatial appearance and temporal evolution concurrently for sign language recognition (SLR) and exhibited impressive performance. However, there are still challenges for 3-D CNN-based methods. First, motion information plays a more significant role than spatial content in sign language. Therefore, it is still questionable whether to treat space and time equally and model them jointly by heavy 3-D convolutions in a unified approach. Second, because of the interference from the highly redundant information in sign videos, it is still nontrivial to effectively extract discriminative spatiotemporal features related to sign language. In this study, deep R(2+1)D was adopted for separate spatial and temporal modeling and demonstrated that decomposing 3-D convolution filters into independent spatial and temporal convolutions facilitates the optimization process in SLR. A lightweight spatial–temporal–channel attention module, including two submodules called channel–temporal attention and spatial–temporal attention, was proposed to make the network concentrate on the significant information along spatial, temporal, and channel dimensions by combining squeeze and excitation attention with self-attention. By embedding this module into R(2+1)D, superior or comparable results to the state-of-the-art methods on the CSL-500, Jester, and EgoGesture datasets were obtained, which demonstrated the effectiveness of the proposed method. Xiangzu Han, Fei Lu 0005, Jianqin Yin, Guohui Tian, Jun Liu 0007 |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2022 | Scene Recognition Mechanism for Service Robot Adapting Various Families: A CNN-Based Approach Using Multi-Type CamerasabstractThe key challenges of scene recognition for service robots in various family environments are the view shortage of holistic scenes and poor adaptation. To address these problems, a family scene recognition mechanism for the service robot is proposed in this paper. A comprehensive application of fish-eye, pinhole, and depth cameras is provided to guarantee the sufficient view of robot. A selective CNN features fusion for the recognition of fish-eye scene images is designed to improve the training efficiency and the recognition accuracy. The mechanism is deployed in a designed hybrid cloud including public and private clouds. The proposed family scene recognition model is trained by large-scale datasets in the public cloud and runs in the private cloud. Besides, the recognition skill can be reinforced and increased by matching human guidance and CNN features to help the robot learn new scenes and improve the adaptation in different family environments. Extensive experiments are implemented to evaluate the proposed method using real scene images from six families. The experiment results show the validity and good performance of our method for the service robot scene recognition in various family environments. Guohui Tian, Ying Zhang 0043, Peng Duan 0002 |
IEEE Trans. Multim. | 2 |
| 2022 | Reinforcement Learning for Logic Recipe Generation: Bridging Gaps From Images to PlansabstractIt is a challenging task to produce recipes from images, due to the difficulty in bridging the gap from intuitive, static images to sequential, dynamic recipes. In this paper, we propose a novel recipe generation system for producing effective recipes from images. As medium steps, ingredient generation is introduced to guide recipe generation in our system. With potential information in ingredient lists, ingredient selection and ingredient sequence, the system is taught to generate effective recipes. For information representation, a hierarchical attention mechanism is designed to extract effective features for ingredient production and recipe generation. In order to guarantee the comprehensiveness and logic in recipes, a specific and explicit criterion around ingredients is designed under the framework of reinforcement learning. In ingredient generation, the system is required to generate ingredients with correct sequence in cooking procedures. And in recipe generation, ingredients in recipes are required to be consistent with produced ingredients. In experiments, the proposed method is compared with state-of-the-art methods to evaluate the feasibility. The results indicate that the proposed system achieves a better performance than other methods on both aspects of producing proper ingredients and effective recipes. Mengyang Zhang, Guohui Tian, Ying Zhang 0043, Peng Duan 0002 |
IEEE Trans. Multim. | 2 |
| 2022 | Effective Safety Strategy for Mobile Robots Based on Laser-Visual Fusion in Home EnvironmentsabstractThe proven efficacy of safety strategies based on 2-D laser rangefinder (LRF) strongly stimulates their application to mobile robots operating in the home environment. However, it remains a challenge for the robot to avoid collisions with all obstacles in the environment. Since LRF can only scan a horizontal slice of the world, some objects cannot be fully observed, such as tables and chairs. In this article, an effective solution based on laser-visual fusion is presented to enhance the safety of the robot. First, a vision sensor is adopted to help detect obstacles that are not fully visible to LRF. Then we propose a method to convert the depth information of the visual image into 2-Dpseudo-laser datarepresentation. With this representation, a strategy for 2-D mapping is developed. On this basis, a novel map fusion algorithm is proposed to generate an improved grid map that amends the incorrect representation of obstacles on the traditional 2-D grid map. We further investigate a robot autonomous navigation strategy that considers LRF data and pseudo-laser data to avoid all obstacles. Experimental results show that the improved grid map together with the presented navigation strategy allows the robot not only to plan a “real” collision-free path, but also to navigate safely in both static and dynamic scenarios, and the proposed strategies can significantly enhance the performance of robot navigation in terms of safety, reliability and robustness. Ying Zhang 0043, Guohui Tian, Xuyang Shao 0002, Jiyu Cheng |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2021 | Online human action recognition with spatial and temporal skeleton features using a distributed camera networkabstractOnline action recognition is an important task for human-centered intelligent services. However, it remains a highly challenging problem due to the high varieties and uncertainties of spatial and temporal scales of human actions. In this paper, the following core ideas are proposed to deal with the online action recognition problem. First, we combine spatial and temporal skeleton features to represent human actions, which include not only geometrical features, but also multiscale motion features, such that both spatial and temporal information of the actions are covered. We use an efficient one-dimensional convolutional neural network to fuse spatial and temporal features and train them for action recognition. Second, we propose a group sampling method to combine the previous action frames and current action frames, which are based on the hypothesis that the neighboring frames are largely redundant, and the sampling mechanism ensures that the long-term contextual information is also considered. Third, the skeletons from multiview cameras are fused in a distributed manner, which can improve the human pose accuracy in the case of occlusions. Finally, we propose a Restful style based client-server service architecture to deploy the proposed online action recognition module on the remote server as a public service, such that camera networks for online action recognition can benefit from this architecture due to the limited onboard computational resources. We evaluated our model on the data sets of JHMDB and UT-Kinect, which achieved highly promising accuracy levels of 80.1% and 96.9%, respectively. Our online experiments show that our memory group sampling mechanism is far superior to the traditional sliding window. Yichao Cao, Guohui Tian, Ze Ji |
Int. J. Intell. Syst. | 4 |
| 2021 | Service skill improvement for home robots: Autonomous generation of action sequence based on reinforcement learning
Mengyang Zhang, Guohui Tian, Ying Zhang 0043, Peng Duan 0002 |
Knowl. Based Syst. | 2 |
| 2020 | Transferring the semantic constraints in human manipulation behaviors to robots
Cici Li, Guohui Tian |
Appl. Intell. | 2 |
| 2020 | Exploring the cognitive process for service task in smart home: A robot service mechanism
Ying Zhang 0043, Guohui Tian, Huanzhao Chen |
Future Gener. Comput. Syst. | 2 |
| 2020 | Integrating manifold ranking with boundary expansion and corners clustering for saliency detection of home scene
Guohui Tian |
Neurocomputing | 2 |
| 2020 | Home service robot task planning using semantic knowledge and probabilistic inference
Guohui Tian, Xuyang Shao 0002 |
Knowl. Based Syst. | 2 |
| 2020 | Facial expression recognition based on deep convolution long short-term memory networks of double-channel weighted mixture
Hepeng Zhang, Guohui Tian |
Pattern Recognit. Lett. | 3 |
| 2019 | Cloud robot: semantic map building for intelligent service task
Hao Wu 0065, Xiaojian Wu, Guohui Tian |
Appl. Intell. | 4 |
| 2019 | A novel scene classification model combining ResNet based transfer learning and data augmentation with a filter
Guohui Tian, Yuan Xu 0003 |
Neurocomputing | 2 |
| 2018 | Distributed RGBD Camera Network for 3D Human Pose Estimation and Action RecognitionabstractSkeleton based human action recognition has recently attracted a lot of attention in the research community. 3D skeleton data is becoming easier to access due to the evolution of new depth sensors like Kinect v2. However, the performance of the depth sensors is subjected to viewpoint variations and occlusions. In this paper, we propose a novel distributed sensor data fusion method to address this problem. The information weighted consensus filter(ICF) is introduced to fuse the skeleton data, so as to get more precise joint positions. To demonstrate the proposed idea, we capture the human action sequences in different views and compare the recognition accuracy between the fused and the raw data, and prove that the fused data can help improve recognition performance. Guohui Tian, Xianglai Zhu, Ziren Wang |
FUSION | 3 |
| 2014 | A method of abnormal habits recognition in intelligent space
Guohui Tian, Hao Wu 0065, Fengyu Zhou 0002 |
Eng. Appl. Artif. Intell. | 2 |
| 2006 | The Application of Robot Formation Approach in the Control of Subway TrainabstractIn order to increase the carrying capacity of subway train, it is necessary to adjust the running of the train for the realization of high operation frequency and safety. After an analysis of the subway train's running under moving block system, a dynamic mathematical model of subway train is constructed. The control among multiple trains is discussed, the method of decreasing the headway and increasing the carrying capacity is discussed by using the event-based technology and formation approach to ensure the safety and to avoid the stop out of station, this approach can reconstruct the system easily and harmonize the cooperation between subsystems. Take account of three trains, the control strategy of the following train is given, the velocity, acceleration of the train can be adjusted according to the safety distance between the successive trains and the distance which the train had traveled Fei Lu 0005, Mumin Song, Guohui Tian, Xiaolei Li 0003 |
IROS | 3 |