VLDB 2026 Research / reviewers in the wild / expert
Yao Guo 0002
dblp:07/6300-2
· DBLP profile ↗
45ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0001-8041-1245ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 3 first-author · 22 since 2021Systems, architecture and hardware · 22 · 1 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Simultaneous surgical stereo depth and motion estimation via brightness-aware self-supervised learning
Yuxuan Liu 0013, Xinyao Zhou, Yating Luo, Yunfei Luan, Zhennan Xiao, Yao Guo 0002, Guang-Zhong Yang |
Pattern Recognit. | 6 |
| 2025 | CLEA: Closed-Loop Embodied Agent for Enhancing Task Execution in Dynamic EnvironmentsabstractLarge Language Models (LLMs) exhibit remarkable capabilities in the hierarchical decomposition of complex tasks through semantic reasoning. However, their application in embodied systems faces challenges in ensuring reliable execution of subtask sequences and achieving one-shot success in long-term task completion. To address these limitations in dynamic environments, we propose Closed-Loop Embodied Agent (CLEA)—a novel architecture incorporating four specialized open-source LLMs with functional decoupling for closed-loop task management. The framework features two core innovations: (1) Interactive task planner that dynamically generates executable subtasks based on the environmental memory, and (2) Multimodal execution critic employing an evaluation framework to conduct a probabilistic assessment of action feasibility, triggering hierarchical re-planning mechanisms when environmental perturbations exceed preset thresholds. To validate CLEA’s effectiveness, we conduct experiments in a real environment with manipulable objects, using two heterogeneous robots for object search, manipulation, and search-manipulation integration tasks. Across 12 task trials, CLEA outperforms the baseline model, achieving a 67.3% improvement in success rate and a 52.8% increase in task completion rate. These results demonstrate that CLEA significantly enhances the robustness of task planning and execution in dynamic environments. Our code is available at https://sp4595.github.io/CLEA/. Mingcong Lei, Ge Wang 0007, Zhixin Mai, Yao Guo 0002, Zhen Li 0026, Shuguang Cui, Yatong Han, Jinke Ren |
IROS | 6 |
| 2025 | Towards Accurate Brain Electrode Implantation via Cross-modality Fusion of White-light and Photoacoustic MicroscopyabstractInvasive flexible neural electrodes are becoming increasingly prevalent in monitoring and modulating brain neural activity, necessitating the precise and minimally invasive implantation of these electrodes to a depth of a few millimeters beneath the cerebral surface. Although Neuralink has pioneered robot-assisted neural electrode implantation guided by microscopy, it currently lacks the ability to detect non-cerebral surface microvessels that are invisible under the white-light microscope, leading to inaccurate implantation planning and a high risk of trauma. To address this limitation, we introduce a vascular-enhanced strategy that fuses intraoperative white-light microscopy and preoperative photoacoustic microscopy and applies the fusion results to our established microsurgical robotic system for brain electrode implantation. Specifically, a multi-modality data preprocessing pipeline is devised to extract representative features, and a 2.5D fusion network that incorporates a depth encoding mechanism is proposed to predict cross-modality correspondence. The enhanced fusion results are utilized for implantation planning and intraoperative guidance during in vivo surgical procedures. Both quantitative and qualitative results are presented to demonstrate the effectiveness of our proposed cross-modality fusion methods. Furthermore, in vivo surgical implementations on mice underscore the potential of the proposed approach for achieving more precise and minimally invasive brain electrode implantation. Yuxuan Liu 0013, Yating Luo, Yunfei Luan, Xinyao Zhou, Jianxin Yang, Yao Guo 0002, Guang-Zhong Yang |
IROS | 6 |
| 2025 | Adaptive Motion Scaling in Teleoperated Robotic Surgery based on Human Intention and AttentionabstractIn teleoperated surgery, the motion scaling factor directly influences both the operator’s control precision of surgical instruments and operational comfort. Previous studies have revealed that the master manipulator state and operator’s gaze information can reflect the complexity of surgical operations and the operator’s intention to some extent. Although enabling real-time adjustment of scaling factors, they were limited by the narrow range of core parameters and the results were significantly influenced by subjective factors. To tackle these challenges, this paper presents a multi-dimensional adaptive motion scaling strategy based on the Bayesian optimization. The prediction of operator’s intention and attention is achieved by integrating multiple dimensional parameters, including master-slave manipulator states, gaze information, as well as pupillary data, all of which have been experimentally validated. Specifically, there exists a significant temporal synchronization between the Index of Pupillary Activity (IPA) and teleoperation tasks, which aligns with research on the correlation between IPA and attention levels. Furthermore, to evaluate the proposed adaptive scaling strategy, we combine subjective questionnaire surveys with objective metric assessments, effectively reducing the excessive influence of operators’ personal conditions and proficiency levels on optimization results. Yiming Zhai, Jingsong Liu, Yating Luo, Yao Guo 0002 |
IROS | 5 |
| 2025 | Hybrid Learning-based Balance Function Assessment of Stroke Patients with a Single Ear-Worn IMUabstractRehabilitation robotics has attracted increasing attention due to its ability to provide continuous, precise, and adaptive treatment programs for stroke patients during their recovery. Accurately assessing lower-limb motor function is crucial in effectively implementing robot-assisted rehabilitation. This study proposes a novel application of a hybrid learning framework that leverages a single-ear-worn inertial measurement unit (IMU) combined with deep learning techniques to predict the Berg Balance Scale (BBS) scores. Participants performed a 3-meter Timed Up and Go (TUG) test while wearing the e-AR sensor. The collected 6-axis IMU data were processed through a CNN-LSTM framework, where we integrated time-domain, frequency-domain, and static features to enhance the model’s regression performance. Experimental results demonstrate that our proposed method achieves a mean absolute error (MAE) of 1.074, surpassing previous studies’ reported results and outperforming traditional machine learning and conventional deep learning algorithms when applied to ear-worn sensor data. The proposed framework is simple to operate yet accurate, making it suitable for patients’ self-assessment even in a home environment. Tianshu Zhao, Zhenye Xu, Yao Guo 0002 |
IROS | 4 |
| 2025 | Deep Coarse-to-Fine Networks for Robust Segmentation and Pose Estimation of Surgical Suturing ThreadsabstractAutonomous suturing is a critical challenge in robot-assisted surgery, where accurate segmentation and pose estimation of suturing threads are essential prerequisites. However, suturing threads are easily occluded by moving instruments and embedded in deformable tissues which make the task much more challenging. To address this, we propose a coarse-to-fine network for detailed segmentation and pose estimation of suturing threads. The coarse stage aims to capture global thread structure, while the fine stage refines the detailed structure through error residual correction. A spatial context fusion module is incorporated to improve the perception of occluded regions, and weighted balanced cross entropy loss as well as hard sample mining strategy is implemented to enhance small target segmentation performance. To deal with severe occlusions, topological constraints are utilized to effectively identify and reconstruct invisible thread segments. Experiments have been conducted on three datasets collected from different surgical scenes including phantom, endoscopy, and microsurgery. Both quantitative and qualitative results have demonstrated that our proposed framework outperforms baseline methods on segmentation and pose estimation of suturing threads, particularly in detecting occluded threads. Our proposed framework generalizes well across different surgical scenarios, showing its potential for automatic suturing. Xinyao Zhou, Yuxuan Liu 0013, Musen Zhang, Yao Guo 0002, Guang-Zhong Yang |
IROS | 5 |
| 2025 | FPM-R2Net: Fused Photoacoustic and operating Microscopic imaging with cross-modality Representation and Registration Network
Yuxuan Liu 0013, Yating Luo, Sung-Liang Chen, Yao Guo 0002, Guang-Zhong Yang |
Medical Image Anal. | 5 |
| 2025 | PoseSDF++: Point Cloud-Based 3-D Human Pose Estimation via Implicit Neural RepresentationabstractPredicting accurate human pose from 3-D visual observation presents a formidable challenge in computer vision, with numerous applications across various industries. However, most existing studies tackled this issue by regressing the 3-D pose from depth maps via 2-D convolutional neural networks or parametric human models, with limited development in point cloud-based methods. To this end, we propose PoseSDF++, i.e., a point cloud-based encoder–decoder network utilizing implicit neural representation to perform 3-D human pose estimation (HPE) and nonparametric shape reconstruction simultaneously. Leveraging the representative capacity of the signed distance function (SDF), we conceptualize the 3-D HPE as a multiple-shape reconstruction task and propose a distance-aware regression method to accurately estimate the 3-D joint positions. In specific, our PoseSDF++ consists of three modules: first,a hierarchical encoderwith vector neuron layers extracts the multiscale rotation equivariant features from the point clouds captured from an arbitrary viewpoint, addressing the degradation issue caused by viewpoint variation of implicit representation; second,a shape decodermaps the extracted feature and the query to its corresponding shape SDF; third,a pose decodercomputes the distance between the query and the target keypoints, namely, the pose SDF. Extensive experiments on four publicly available datasets demonstrate that our PoseSDF++ achieves competitive performance against the state-of-the-art point cloud-based methods and covering the human hand (HANDS 2019), lower limbs (ICL-Gait), and full body (DFAUST, LiDARHuman2.6M) pose estimation. Jianxin Yang, Yuxuan Liu 0013, Xiao Gu 0003, Guang-Zhong Yang, Yao Guo 0002 |
IEEE Trans. Ind. Informatics | 6 |
| 2025 | Toward Fine-Grained 3-D Visual Grounding Through Referring Textual PhrasesabstractRecent progress in 3-D scene understanding has explored visual grounding [3D visual grounding (3DVG)] to localize a target object through a language description. However, existing methods only consider the dependency between the entire sentence and the target object, ignoring fine-grained relationships between contexts and nontarget ones. In this article, we extend 3DVG to a more fine-grained task, called 3D phrase-aware grounding (3DPAG). The 3DPAG task aims to localize the target objects in a 3-D scene by explicitly identifying all phrase-related objects and then conducting the reasoning according to contextual phrases. To tackle this problem, we manually labeled about 227 K phrase-level annotations using a self-developed platform, from 88 K sentences of widely used 3DVG datasets, i.e., Natural Reference in 3-D (Nr3D), Spatial Reference in 3-D (Sr3D), and ScanRefer. By tapping on our datasets, we can extend previous 3DVG methods to the fine-grained phrase-aware scenario. It is achieved through the proposed novel phrase-object alignment (POA) optimization and phrase-specific pretraining (PSP), boosting conventional 3DVG performance as well. Extensive results confirm significant improvements, i.e., previous state-of-the-art method achieves 3.9%, 3.5%, and 4.6% overall accuracy gains on Nr3D, Sr3D, and ScanRefer, respectively. Our datasets and platform are released in https://github.com/CurryYuan/PhraseRefer. Zhihao Yuan, Xu Yan 0005, Xuhao Li, Yao Guo 0002, Shuguang Cui, Zhen Li 0026 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Skill Learning in Robot-Assisted Micro-Manipulation Through Human Demonstrations with Attention GuidanceabstractFor the development of robotic systems for micromanipulation, it is challenging to design appropriate control strategies due to either the lack of sufficient information for feedback or the difficulty in extracting subtle yet critical visual features. With the same system under the teleoperated mode, however, human operators seem to be able to complete the task more successfully with an inherent motion and control strategy. The extraction of implicit human attention during the task and integration of this with robot control could provide crucial guidance in the design of feature extraction and motion control algorithms. In this paper, a micro-assembly task of miniature thin membrane sensors is considered. For human demonstrations, we collected data from repeated tests performed by ten operators following three motion strategies. The human attention during the task is explored according to the coordinates of the eye gaze, and then a neural network with gaze-guided attention is trained to segment the visual Region of Interest (ROI). After quantitative evaluation of operator results in terms of success rate, efficiency, reset time, and the Index of Pupillary Activity (IPA), an optimized motion strategy based on the "palpation" framework was derived. Consequently, we apply this strategy to automated tasks and achieve superior results than human operators, showing an average task completion time of 34.8±5.9s and a success rate of over 90%. Yujian An, Jianxin Yang, Bingze He, Yao Guo 0002, Guang-Zhong Yang |
ICRA | 5 |
| 2024 | Intelligent Disinfection Robot with High-Touch Surface Detection and Dynamic Pedestrian AvoidanceabstractThe increasing awareness of public health issues has highlighted the need for effective disinfection of crowded indoor public areas, leading to the development of automated disinfection robots. However, most of the existing robots spray disinfectant in all areas, and they are still immature to navigate in densely populated environments. Hence, in this paper, we design a new disinfection robotic system consisting of a mobile platform, an RGB-D camera, and a robotic arm with a spray disinfection device. To address the above challenges, we propose a vision-based method for accurately detecting high-touch areas in the surroundings, enabling the disinfection robot to achieve superior disinfection efficiency. In addition, we propose a dynamic pedestrian avoidance method, namely Socially Aware APF (SA-APF), which can predict the movement trend of pedestrians and plan the path in real-time. Both simulated and real-world experiments are conducted to demonstrate the effectiveness of our disinfection robot system, especially highlighting the ability to detect high-touch areas and navigate in the environment while avoiding dynamic pedestrians. Yunfei Luan, Muhang He, Yudong Tian, Chengjie Lin, Yunhan Fang, Zihao Zhao 0005, Jianxin Yang, Yao Guo 0002 |
ICRA | 8 |
| 2024 | Fast Photoacoustic Microscopy with Robot Controlled Microtrajectory OptimizationabstractPhotoacoustic Microscopy (PAM) is a relatively new imaging modality in biomedicine. However, point-by-point raster scanning in PAM suffers from low imaging speed. Sparse sampling has been studied in recent years and with the development of deep learning algorithms, extensive efforts have been devoted to sparse image reconstruction while little attention has been paid to sparse sampling trajectory design required for actual implementation. The use of real-time adaptive robotically controlled sampling with micro-scale accuracy with due consideration of physical constraints can pave the way for using PAM for robot-assisted microsurgery. This work proposes a fast PAM scheme with robot-controlled microtrajectory optimization. The proposed method is adaptive to imaging details of different regions of interest (ROI) and detailed experiments have been conducted on both simulation and in-vivo settings. Results show that our proposed method can achieve faster scanning speed than traditional raster scanning and improved image quality in ROI than the standard spiral trajectory, which demonstrates the effectiveness of our proposed method and its potential to be deployed in other point-by-point scanning systems. Yating Luo, Yuxuan Liu 0013, Sung-Liang Chen, Yao Guo 0002, Guang-Zhong Yang |
ICRA | 5 |
| 2024 | Real-Time Acoustic Holography With Physics-Based Deep Learning for Robotic ManipulationabstractAcoustic holography (AH) is a promising technique for precise noncontact micro-nano robotic manipulation. It encodes a three-dimensional (3D) acoustic field acting as a virtual end-effector into a two-dimensional (2D) hologram, whereby the desired acoustic field reconstruction is made possible. Most traditional methods to implement AH, such as 3D printed holographic lens and phased array of transducers (PAT), have limitations of dynamic and dexterous manipulation. Furthermore, existing iterative optimization algorithms to calculate 2D holograms have inadequate accuracy and real-time performance. To address these issues, this paper proposes a physics-based deep learning method with a novel training framework for phase-only hologram (POH) calculation enabling further pushing forward the PAT-based AH for noncontact robotic manipulation. By implementing independent control of each channel on PAT referring real-time calculated POH by a well-trained network, the desired acoustic field can be reconstructed in real-time with high fidelity. The results both on a simulated dataset and a real dataset demonstrate that our method supports accurate and dynamic reconstruction of desired acoustic field with distinct morphologies, with an average reconstruction error of 0.085 and average POH computing time of 47 milliseconds on GPU. Indeed, this work shows the future potential of AH in the field of noninvasive medical therapy, exogenous material delivery, and miniaturized industrial assembly.Note to Practitioners—This paper addresses the challenge of noncontact micro-nano robotic manipulation by PAT-based AH, an intriguing technique in bioengineering, micro-assembly, and material characterization. However, existing approaches have limited precision and real-time performance. To overcome these limitations, this paper proposes a physics-based deep learning method with a novel training framework. Our method achieves excellent accuracy and real-time performance, enabling efficient reconstruction of various complicated acoustic field morphologies for precise and dynamic acoustic manipulation. Experimental results demonstrate its high manipulation flexibility due to the independent modulation of each channel of PAT and real-time precise control due to the ultrafast calculation of the proposed deep learning method, though the method has not yet been deployed into an acoustic manipulation system and tested in practice. Future research will focus on designing physical experiments for further evaluation. Overall, the proposed method provides a novel and promising basis for desired acoustic field generation. Chengxi Zhong, Jiaqi Li 0029, Zhenhuan Sun, Teng Li 0017, Yao Guo 0002, David C. Jeong, Hu Su, Song Liu 0003 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | Comprehensive Visual Question Answering on Point Clouds through Compositional Scene ManipulationabstractVisual Question Answering on 3D Point Cloud (VQA-3D) is an emerging yet challenging field that aims at answering various types of textual questions given an entire point cloud scene. To tackle this problem, we propose the CLEVR3D, a large-scale VQA-3D dataset consisting of 171K questions from 8,771 3D scenes. Specifically, we develop a question engine leveraging 3D scene graph structures to generate diverse reasoning questions, covering the questions of objects' attributes (i.e., size, color, and material) and their spatial relationships. Through such a manner, we initially generated 44K questions from 1,333 real-world scenes. Moreover, a more challenging setup is proposed to remove the confounding bias and adjust the context from a common-sense layout. Such a setup requires the network to achieve comprehensive visual understanding when the 3D scene is different from the general co-occurrence context (e.g., chairs always exist with tables). To this end, we further introduce the compositional scene manipulation strategy and generate 127K questions from 7,438 augmented 3D scenes, which can improve VQA-3D models for real-world comprehension. Built upon the proposed dataset, we baseline several VQA-3D models, where experimental results verify that the CLEVR3D can significantly boost other 3D scene understanding tasks. Xu Yan 0005, Zhihao Yuan, Yinghong Liao, Yao Guo 0002, Shuguang Cui, Zhen Li 0026 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | EgoHMR: Egocentric Human Mesh Recovery via Hierarchical Latent Diffusion ModelabstractEgocentric vision has gained increasing popularity in social robotics, demonstrating great potentials for personal assistance and human-centric behavior analysis. Holistic per-ception of human body itself is a prerequisite for downstream applications, including action recognition and anticipation. Extensive research has been performed for human mesh recovery from the exocentric images captured from a third-person view, but limited studies are conducted for heavily distorted yet occluded egocentric images. In this paper, we propose Egocentric Human Mesh Recovery (EgoHMR), a novel hierarchical network based on latent diffusion models. Our method takes a single egocentric frame as the input and it can be trained in an end-to-end manner without supervision of 2D pose. The network is built upon the latent diffusion model by incorporating both global and local features in a hierarchical structure. To train the proposed network, we generate weak labels from synchronized exocentric images. The proposed method can perform human mesh recovery directly from egocentric images and detailed quantitative and qualitative experiments have been conducted to demonstrate the effectiveness of the proposed EgoHMR method. Yuxuan Liu 0013, Jianxin Yang, Xiao Gu 0003, Yao Guo 0002, Guang-Zhong Yang |
ICRA | 4 |
| 2023 | EasyGaze3D: Towards Effective and Flexible 3D Gaze Estimation from a Single RGB CameraabstractEye gaze can convey rich information of human intentions, which enables the social robots to comprehend the cognition and behavior of human targets. However, the existing 3D gaze estimation methods generally have high requirements either on the dedicated hardware or the quantity and quality of training databases, which largely limits their practical application values. This paper proposes EasyGaze3D, an effective 3D gaze estimation framework using a single RGB camera. First, the framework detects the 2D facial landmarks and recovers the 3D facial shape from the input image, and derives the required camera parameters with these features. Then, without loss of generality, the gaze direction can be regarded as the vector pointing from the eyeball center to the pupil center, which are derived respectively from the detected facial landmarks and the spherical fitting performed on the recovered 3D facial shape. Besides, we propose a flexible yet efficient calibration module, namely Easy-Cali, for deriving the subject-specific 3D facial shape and eyeball centers. The features calibrated by Easy-Cali can further boost the performance of EasyGaze3D. Experimental results show that our proposed method, being plug-and-play and without the need of training on large-scale dataset, can achieve superior performance against the existing methods based on deep models. Jianxin Yang, Yuxuan Liu 0013, Zhen Li 0026, Guang-Zhong Yang, Yao Guo 0002 |
IROS | 6 |
| 2023 | EgoFish3D: Egocentric 3D Pose Estimation From a Fisheye Camera via Self-Supervised LearningabstractEgocentric vision has gained increasing popularity recently, opening new avenues for human-centric applications. However, the use of the egocentric fisheye cameras allows wide angle coverage but image distortion is introduced along with strong human body self-occlusion imposing significant challenges in data processing and model reconstruction. Unlike previous work only leveraging synthetic data for model training, this paper presents a new real-world EgoCentric Human Pose (ECHP) dataset. To tackle the difficulty of collecting 3D ground truth using motion capture systems, we simultaneously collect images from a head-mounted egocentric fisheye camera as well as from two third-person-view cameras, circumventing the environmental restrictions. By using self-supervised learning under multi-view constraints, we propose a simple yet effective framework, namely EgoFish3D, for egocentric 3D pose estimation from a single image in different real-world scenarios. The proposed EgoFish3D incorporates three main modules. 1)The third-person-view moduletakes two exocentric images as input and estimates the 3D pose represented in the third-person camera frame; 2)the egocentric modulepredicts the 3D pose in the egocentric camera frame; and 3)the interactive moduleestimates the rotation matrix between the third-person and the egocentric views. Experimental results on our ECHP dataset and existing benchmark datasets demonstrate the effectiveness of the proposed EgoFish3D, which can achieve superior performance to existing methods. Yuxuan Liu 0013, Jianxin Yang, Xiao Gu 0003, Yao Guo 0002, Guang-Zhong Yang |
IEEE Trans. Multim. | 5 |
| 2023 | An Intelligent Vision-Based Nutritional Assessment Method for Handheld Food ItemsabstractDietary assessment has proven to be effective to evaluate the dietary intake of patients with diabetes and obesity. The traditional approach of accessing the dietary intake is to conduct a 24-hour dietary recall, a structured interview designed to obtain information on food categories and volume consumed by the participants. Due to unconscious biases in this kind of self-reporting approaches, many research studies have explored the use of vision-based approaches to provide accurate and objective assessments. Despite the promising results of food recognition by deep neural networks, there still exist several hurdles in deep learning-based food volume estimation ranging from domain shift between synthetic and raw 3D models, shape completion ambiguity and lack of large-scale paired training dataset. Therefore, this paper proposed an intelligent nutritional assessment approach via weakly-supervised point cloud completion, which aims to close the reality gap in 3D point cloud completion tasks and address the targeted challenges. Then the volume can be easily estimated from the completed representation of the food. Another major merit of our system is that it can be used to estimate the volume of handheld food items without requiring the constraints including placing the food items on a table or next to fiducial markers, which facilitates the implementation on both wearable and handheld cameras. Comprehensive experiments have been carried out on major benchmark datasets and self-constructed volume-annotated dataset respectively, in which the proposed method demonstrates comparable results with several strong fully-supervised baseline methods and shows superior completion ability in handling food volume estimation. Frank P.-W. Lo, Yao Guo 0002, Yingnan Sun, Jianing Qiu, Benny P. L. Lo |
IEEE Trans. Multim. | 2 |
| 2022 | APAUNet: Axis Projection Attention UNet for Small Target in 3D Medical Segmentation
Yuncheng Jiang 0002, Zixun Zhang, Shixi Qin, Yao Guo 0002, Zhen Li 0026, Shuguang Cui |
ACCV (6) | 4 |
| 2022 | X -Trans2Cap: Cross-Modal Knowledge Transfer using Transformer for 3D Dense Captioningabstract3D dense captioning aims to describe individual objects in 3D scenes by natural language, where 3D scenes are usually represented as RGB-D scans or point clouds. However, only exploiting single modal information, e.g., point cloud, previous approaches fail to produce faithful descriptions. Though aggregating 2D features into point clouds may be beneficial, it introduces an extra computational burden, especially in the inference phase. In this study, we investigate a cross-modal knowledge transfer using Transformer for 3D dense captioning, namely X-Trans2Cap. Our proposed X-Trans2Cap effectively boost the performance of single-modal 3D captioning through the knowledge distillation enabled by a teacher-student framework. In practice, during the training phase, the teacher network exploits auxiliary 2D modality and guides the student network that only takes point clouds as input through the feature consistency constraints. Owing to the well-designed cross-modal feature fusion module and the feature alignment in the training phase, X-Trans2Cap acquires rich appearance information embedded in 2D images with ease. Thus, a more faithful caption can be generated only using point clouds during the inference. Qualitative and quantitative results confirm that X-Trans2Cap outperforms previous state-of-the-art by a large margin, i.e., about +21 and +16 CIDEr points on ScanRefer and Nr3D datasets, respectively. Zhihao Yuan, Xu Yan 0005, Yinghong Liao, Yao Guo 0002, Guanbin Li, Shuguang Cui, Zhen Li 0026 |
CVPR | 4 |
| 2022 | Tackling Long-Tailed Category Distribution Under Domain Shifts
Xiao Gu 0003, Yao Guo 0002, Zeju Li, Jianing Qiu, Qi Dou 0001, Yuxuan Liu 0013, Benny P. L. Lo, Guang-Zhong Yang |
ECCV (23) | 2 |
| 2022 | PoseSDF: Simultaneous 3D Human Shape Reconstruction and Gait Pose Estimation Using Signed Distance FunctionsabstractVision-based 3D human pose estimation and shape reconstruction play important roles in robot-assisted healthcare monitoring and personal assistance. However, 3D data captured from a single viewpoint always encounter occlusions and exhibit substantial heterogeneity across different views, resulting in significant challenges for both tasks. Extensive approaches have been proposed to perform each task separately, but few of them present a unified solution. In this paper, we propose a novel network based on signed distance functions, namely PoseSDF, to simultaneously reconstruct 3D lower limb shape and estimate gait pose by two dedicated branches. To promote multi-task learning, several strategies are developed to ensure that these two branches leverage the same latent shape code while exchanging information between them. More importantly, an auxiliary RotNet is incorporated into the inference phase, overcoming the inherent limitations of implicit neural functions under cross-view scenarios. Experimental results demonstrate that our proposed PoseSDF can achieve both high-quality shape reconstruction and precise pose estimation, generalizing well on the data from novel views, gait patterns, as well as real-world. Jianxin Yang, Yuxuan Liu 0013, Xiao Gu 0003, Guang-Zhong Yang, Yao Guo 0002 |
ICRA | 5 |
| 2022 | Human-Robot Shared Control for Surgical Robot Based on Context-Aware Sim-to-Real AdaptationabstractHuman-robot shared control, which integrates the advantages of both humans and robots, is an effective approach to facilitate efficient surgical operation. Learning from demonstration (LfD) techniques can be used to automate some of the surgical sub tasks for the construction of the shared control mechanism. However, a sufficient amount of data is required for the robot to learn the manoeuvres. Using a surgical simulator to collect data is a less resource-demanding approach. With sim-to-real adaptation, the manoeuvres learned from a simulator can be transferred to a physical robot. To this end, we propose a sim-to-real adaptation method to construct a human-robot shared control framework for robotic surgery. In this paper, a desired trajectory is generated from a simulator using LfD method, while dynamic motion primitives (DMP) is used to transfer the desired trajectory from the simulator to the physical robotic platform. Moreover, a role adaptation mechanism is developed such that the robot can adjust its role according to the surgical operation contexts predicted by a neural network model. The effectiveness of the proposed framework is validated on the da Vinci Research Kit (dVRK). Results of the user studies indicated that with the adaptive human-robot shared control framework, the path length of the remote controller, the total clutching number and the task completion time can be reduced significantly. The proposed method outperformed the traditional manual control via teleoperation. Dandan Zhang 0001, Zicong Wu, Adnan Munawar, Bo Xiao 0002, Yuan Guan, Wuzhou Hong, Yao Guo 0002, Gregory S. Fischer, Benny P. L. Lo, Guang-Zhong Yang |
ICRA | 10 |
| 2022 | Ego+X: An Egocentric Vision System for Global 3D Human Pose Estimation and Social Interaction CharacterizationabstractEgocentric vision is an emerging topic, which has demonstrated great potential in assistive healthcare scenarios, ranging from human-centric behavior analysis to personal social assistance. Within this field, due to the heterogeneity of visual perception from first-person views, egocentric pose estimation is one of the most significant prerequisites for enabling various downstream applications. However, existing methods for egocentric pose estimation mainly focus on predicting the pose represented in the camera coordinates from a single image, which ignores the latent cues in the temporal domain and results in less accuracy. In this paper, we propose Ego+X, an egocentric vision based system for 3D canonical pose estimation and human-centric social interaction characterization. Our system is composed of two head-mounted egocentric cameras, where one is faced downwards and the other looks outwards. By leveraging the global context provided by visual SLAM, we first propose Ego-Glo for spatial-accurate and temporal-consistent egocentric 3D pose estimation in the canonical coordinate system. With the help of an egocentric camera looking outwards, we then propose Ego-Soc by extending Ego-Glo to various social interaction tasks, e.g., object detection and human-human interaction. Quantitative and qualitative experiments have been conducted to demonstrate the effectiveness of our proposed Ego+X. Yuxuan Liu 0013, Jianxin Yang, Xiao Gu 0003, Yao Guo 0002, Guang-Zhong Yang |
IROS | 4 |
| 2022 | Real-time Acoustic Holography with Physics-based Deep Learning for Acoustic Robotic ManipulationabstractAcoustic holography is a newly emerging and promising technique to dynamically generate arbitrary desired holographic acoustic field in 3D space for contactless robotic manipulation. The latest technology supporting complex dynamic holographic acoustic field reconstruction is through phased transducer array (PTA), where the phase profile of emitted acoustic wave from discrete transducers is controlled independently by sophisticated circuits to modulate the acoustic interference field. While the forward kinematics of a phased array based robotic manipulation system is simple and straightforward, the inverse kinematics of the required holographic acoustic field is mathematically non-linear and unsolvable, which substantially limits the application of dynamic holographic acoustic field for robot manipulation. In this work, we propose a physics-based deep learning framework for this phase retrieval inverse kinematics problem so that the target complex hologram could be reconstructed precisely with average MAE of 0.022 and in real time with prediction time of 47 milliseconds on GPU. The accuracy and real time of the proposed method for dynamic holographic acoustic field reconstruction from PTA are demonstrated experimentally. Chengxi Zhong, Zhenhuan Sun, Kunyong Lyu, Yao Guo 0002, Song Liu 0003 |
IROS | 4 |
| 2022 | Eye-Tracking for Performance Evaluation and Workload Estimation in Space Telerobotic TrainingabstractMonitoring the mental workload of operators is of paramount importance in space telerobotic training and other teleoperation tasks. Instead of the estimation of task-specific workload, this article aims at investigating the impact of two significant confounding factors (time-pressure and latency) on space teleoperation and explored the use of eye-tracking technology for factor-induced mental workload estimation and performance evaluation. Ten subjects teleoperated a Canadarm2 robot to complete a complex on-orbit assembly task in our photo-realistic training simulator while wearing a head-mounted eye-tracker. To understand how time-pressure and latency influence eye-tracking features works, we first performed the statistical analysis on various features with respect to a single factor and across multiple groups. Next, eye-tracking features extracted from segment data and trial data is used to identify the mental workload induced by confounding factors, which can be used for developing personalized training programs and guaranteeing safe teleoperation. Furthermore, to improve the recognition performance using segment data, we propose the activity ratio and time ratio to characterize the informative segments. Finally, the relationship between simulator-defined performance measures and eye-tracking features is examined. Results show that fixation duration, saccade frequency and duration, pupil diameter, and index of pupillary activity are significant features that can be used in both factor-induced mental workload estimation and task performance evaluation. Yao Guo 0002, Daniel R. Freer, Fani Deligianni, Guang-Zhong Yang |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2022 | Cross-Domain Self-Supervised Complete Geometric Representation Learning for Real-Scanned Point Cloud Based Pathological Gait AnalysisabstractAccurate lower-limb pose estimation is aprerequisite of skeleton based pathological gait analysis. To achieve this goal in free-living environments for long-term monitoring, single depth sensor has been proposed in research. However, the depth map acquired from a single viewpoint encodes only partial geometric information of the lower limbs and exhibits large variations across different viewpoints. Existing off-the-shelf 3D pose tracking algorithms and public datasets for depth based human pose estimation are mainly targeted at activity recognition applications. They are relatively insensitive to skeleton estimation accuracy, especially at the foot segments. Furthermore, acquiring ground truth skeleton data for detailed biomechanics analysis also requires considerable efforts. To address these issues, we propose a novel cross-domain self-supervised complete geometric representation learning framework, with knowledge transfer from the unlabelled synthetic point clouds of full lower-limb surfaces. The proposed method can significantly reduce the number of ground truth skeletons (with only 1%) in the training phase, meanwhile ensuring accurate and precise pose estimation and capturing discriminative features across different pathological gait patterns compared to other methods. Xiao Gu 0003, Yao Guo 0002, Guang-Zhong Yang, Benny P. L. Lo |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Simultaneous Precision Assembly of Multiple Objects through Coordinated Micro-robot ManipulationabstractSimultaneous assembly of multiple objects is a key technology to form solid connections among objects to get compact structures in precision assembly and micro-assembly. Dramatically different from traditional assembly of two objects, the interaction among multiple objects is more complicated on analysis and control. During simultaneous assembly of multiple objects, there are multiple mutually effected contact surfaces, and multiple force sensors are needed to perceive the interaction status. In this paper, a coordinated micro-robot manipulation strategy is proposed for simultaneous assembly problem, which is based on microscopic vision and force information. Taking simultaneous assembly of three objects as an instance, the proposed method is well articulated, including calibration of assembly system, force analysis for each contacting surface, and insertion control strategy for assembly process. The proposed method is applicable also to case with more objects. Experiment results demonstrate effectiveness of the proposed method. Song Liu 0003, Yuyu Jia, Youfu Li 0001, Yao Guo 0002, Haojian Lu |
ICRA | 4 |
| 2021 | Deep3DRanker: A Novel Framework for Learning to Rank 3D Models with Self-Attention in Robotic VisionabstractResearch on generating or processing point clouds has become an increasingly popular domain in robotic research due to its extensive applications, such as robotic grasping, augmented reality and autonomous vehicle navigation. In this paper, we explore a new research area on point clouds - Learning to rank 3D models captured from a single depth image. In the Learning To Rank (LTR) task, we aim at optimizing the order of a list of 3D models according to the given query. Inspired by the recent advances in Natural Language Processing (NLP), we propose a novel framework, namely Deep3DRanker, for ranking 3D models by leveraging graph-based encoding and self-attention mechanisms. Comprehensive experiments are conducted to validate our methods on publicly available YCB synthetic and YCB video datasets. The promising results have shown that our proposed framework is generic enough to be applicable with any combinations of randomly positioned, oriented, and unseen object items with accuracy ranging from 59.2% to 94.9%, which shows great potentials of the proposed framework for robotic applications, in particular, for making decisions under different circumstances. Frank P.-W. Lo, Yao Guo 0002, Yingnan Sun, Jianing Qiu, Benny P. L. Lo |
ICRA | 2 |
| 2021 | MCDCD: Multi-Source Unsupervised Domain Adaptation for Abnormal Human Gait DetectionabstractFor gait analysis, especially for the detection of subtle gait abnormalities, the collected datasets involve high variability across subjects due to inherent biometric traits and movement behaviors, leading to limited detection accuracy and poor generalizability. To address this, we propose a novel deep multi-source Unsupervised Domain Adaptation (UDA) approach, namely Maximum Cross-Domain Classifier Discrepancy (MCDCD), which aims to improve the classification performance on the test subject (target domain) by leveraging the information from multiple labelled training subjects (source domains). Specifically, the proposed model consists of a feature extractor and a domain-specific category classifier per source domain. The former feature extractor learns to generate discriminative gait features. For the latter classifiers, we minimize the cross-entropy loss to accurately classify source samples, and simultaneously maximize a novel cross-domain discrepancy loss between any two category classifiers to minimize domain shift between multiple sources and the target domain. To validate the proposed MCDCD for detecting gait abnormalities on novel subjects, we collected both high-quality Motion capture (Mocap) and noisy Electromyography (EMG) data from eighteen subjects with both normal and imitated abnormal gaits. Experiment results using both data modalities demonstrate that the proposed approach can achieve superior performance in abnormal gait classification compared to baseline deep models and state-of-the-art UDA methods. Yao Guo 0002, Xiao Gu 0003, Guang-Zhong Yang |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Cross-Subject and Cross-Modal Transfer for Generalized Abnormal Gait Pattern RecognitionabstractFor abnormal gait recognition, pattern-specific features indicating abnormalities are interleaved with the subject-specific differences representing biometric traits. Deep representations are, therefore, prone to overfitting, and the models derived cannot generalize well to new subjects. Furthermore, there is limited availability of abnormal gait data obtained from precise Motion Capture (Mocap) systems because of regulatory issues and slow adaptation of new technologies in health care. On the other hand, data captured from markerless vision sensors or wearable sensors can be obtained in home environments, but noises from such devices may prevent the effective extraction of relevant features. To address these challenges, we propose a cascade of deep architectures that can encode cross-modal and cross-subject transfer for abnormal gait recognition. Cross-modal transfer maps noisy data obtained from RGBD and wearable sensors to accurate 4-D representations of the lower limb and joints obtained from the Mocap system. Subsequently, cross-subject transfer allows disentangling subject-specific from abnormal pattern-specific gait features based on a multiencoder autoencoder architecture. To validate the proposed methodology, we obtained multimodal gait data based on a multicamera motion capture system along with synchronized recordings of electromyography (EMG) data and 4-D skeleton data extracted from a single RGBD camera. Classification accuracy was improved significantly in both Mocap and noisy modalities. Xiao Gu 0003, Yao Guo 0002, Fani Deligianni, Benny P. L. Lo, Guang-Zhong Yang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | End-to-End Real-time Catheter Segmentation with Optical Flow-Guided Warping during Endovascular InterventionabstractAccurate real-time catheter segmentation is an important pre-requisite for robot-assisted endovascular intervention. Most of the existing learning-based methods for catheter segmentation and tracking are only trained on smallscale datasets or synthetic data due to the difficulties of ground-truth annotation. Furthermore, the temporal continuity in intraoperative imaging sequences is not fully utilised. In this paper, we present FW-Net, an end-to-end and real-time deep learning framework for endovascular intervention. The proposed FW-Net has three modules: a segmentation network with encoder-decoder architecture, a flow network to extract optical flow information, and a novel flow-guided warping function to learn the frame-to-frame temporal continuity. We show that by effectively learning temporal continuity, the network can successfully segment and track the catheters in real-time sequences using only raw ground-truth for training. Detailed validation results confirm that our FW-Net outperforms stateof-the-art techniques while achieving real-time performance. Anh Nguyen 0003, Dennis Kundrat, Giulio Dagnino, Wenqiang Chi, Mohamed E. M. K. Abdelaziz, Yao Guo 0002, YingLiang Ma, Trevor M. Y. Kwok, Celia V. Riga, Guang-Zhong Yang |
ICRA | 6 |
| 2020 | Coupled Real-Synthetic Domain Adaptation for Real-World Deep Depth EnhancementabstractAdvances in depth sensing technologies have allowed simultaneous acquisition of both color and depth data under different environments. However, most depth sensors have lower resolution than that of the associated color channels and such a mismatch can affect applications that require accurate depth recovery. Existing depth enhancement methods use simplistic noise models and cannot generalize well under real-world conditions. In this paper, a coupled real-synthetic domain adaptation method is proposed, which enables domain transfer between high-quality depth simulators and real depth camera information for super-resolution depth recovery. The method first enables the realistic degradation from synthetic images, and then enhances degraded depth data to high quality with a color-guided sub-network. The key advantage of the work is that it generalizes well to real-world datasets without further training or fine-tuning. Detailed quantitative and qualitative results are presented, and it is demonstrated that the proposed method achieves improved performance compared to previous methods fine-tuned on the specific datasets. Xiao Gu 0003, Yao Guo 0002, Fani Deligianni, Guang-Zhong Yang |
IEEE Trans. Image Process. | 2 |
| 2019 | Transfer Learning for Surgical Task SegmentationabstractIn this paper, we present a novel approach for surgical task segmentation. A segmentation policy learns the correlations between features and segmentation points from manually labeled data. The most correlated features and rules for segmenting them are identified and learned. These form a complete set of segmentation policy. The proposed approach is developed to segment new but similar tasks through transfer learning. It is verified through applying the segmentation rule learned from the labeled data to segment other tasks. The performance of the proposed algorithm was evaluated by comparing the results against the ground truths. Experimental results demonstrate that our approach can achieve high segmentation rates with an accuracy of between 68.8% - 81.8%. Ya-Yen Tsai, Bidan Huang, Yao Guo 0002, Guang-Zhong Yang |
ICRA | 3 |
| 2019 | Unsupervised Task Segmentation Approach for Bimanual Surgical Tasks using Spatiotemporal and Variance PropertiesabstractIn surgical workflow analysis and training in robot-assisted surgery, automatic task segmentation could significantly reduce the manual labeling time and enhance robot learning efficiency. This paper presents an unsupervised segmentation approach to automatically segment a given surgical task without manual intervention. A new segmentation method is presented, which relies only on bimanual kinematic trajectories without the need for prior information about the data. Specifically, surgical tasks are segmented by fusing trajectories' spatiotemporal and variance properties. To demonstrate the effectiveness of the proposed method, detailed experiments were first conducted on our dataset. We segmented trajectories of three different surgical stitches and observed an average F1score of 77.9% against the ground truths. The same trajectories were then added with different levels of noises and the segmentation comparison was made with four other methods. The proposed algorithm had demonstrated its robustness against the noises. Finally, to assess its generalization ability, the method was evaluated on publicly available JIGSAWS dataset and an average F1score of 75.5% was achieved. Ya-Yen Tsai, Yao Guo 0002, Guang-Zhong Yang |
IROS | 2 |
| 2019 | A Handheld Master Controller for Robot-Assisted MicrosurgeryabstractAccurate master-slave control is important for Robot-Assisted Microsurgery (RAMS). This paper presents a handheld master controller for the operation and training of RAMS. A 9-axis Inertial Measure Unit (IMU) and a micro camera are utilized to form the sensing system for the handheld controller. A new hybrid marker pattern is designed to achieve reliable visual tracking, which integrated QR codes, Aruco markers, and chessboard vertices. Real-time multi-sensor fusion is implemented to further improve the tracking accuracy. The proposed handheld controller has been verified on an in-house microsurgical robot to assess its usability and robustness. User studies were conducted based on a trajectory following task, which indicated that the proposed handheld controller had comparable performance with the Phantom Omni, demonstrating its potential applications in microsurgical robot control and training. Dandan Zhang 0001, Yao Guo 0002, Guang-Zhong Yang |
IROS | 2 |
| 2019 | A Hierarchical Model for Human Action Recognition From Body-PartsabstractAs increasing attention is paid to human action recognition from skeleton data, this paper focuses on such tasks by proposing a hierarchical model to discover the structure information of body-parts involved in actions for better analysis of human actions in the skeleton data. Considering human actions as simultaneous motions of body-parts of the human skeleton, we propose a hierarchical model to simultaneously apply discriminative body-parts selection at a same scale and group coupling of bundles of body-parts at different scales, while we decompose the human skeleton into a hierarchy of body-parts of varying scales. To represent such hierarchy of body-parts, we accordingly build a hierarchical rotation and relative velocity (HRRV) descriptor. The hierarchical representations encoded by Fisher vectors of the HRRV descriptors are properly formulated into the hierarchical model via the proposed mixed norm, to apply the sparse selection of body-parts and regularize the structure of such hierarchy of body-parts. The extensive evaluations on three challenging datasets demonstrate the effectiveness of our proposed approach, which achieves superior performance compared to the state-of-the-art algorithms on datasets with various sizes, showing it is more widely applicable than existing approaches. Zhanpeng Shao, Youfu Li 0001, Yao Guo 0002, Xiaolong Zhou 0001, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | From Emotions to Mood Disorders: A Survey on Gait Analysis MethodologyabstractMood disorders affect more than 300 million people worldwide and can cause devastating consequences. Elderly people and patients with neurological conditions are particularly susceptible to depression. Gait and body movements can be affected by mood disorders, and thus they can be used as a surrogate sign, as well as an objective index for pervasive monitoring of emotion and mood disorders in daily life. Here we review evidence that demonstrates the relationship between gait, emotions and mood disorders, highlighting the potential of a multimodal approach that couples gait data with physiological signals and home-based monitoring for early detection and management of mood disorders. This could enhance self-awareness, enable the development of objective biomarkers that identify high risk subjects and promote subject-specific treatment. Fani Deligianni, Yao Guo 0002, Guang-Zhong Yang |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | A Hierarchical Model for Action Recognition Based on Body PartsabstractAs increasing attention is paid on human action recognition from skeleton data, this paper focuses on such tasks by proposing a hierarchical model to discover the structure information of body-parts involved in human actions. Considering human actions as simultaneous motions of different body-parts of the human skeleton, we propose a hierarchical model to simultaneously apply discriminative body-parts selection at a same scale and group coupling of bundles of body-parts at different scales, while we decompose the human skeleton into a hierarchy of body-parts of varying scales. To represent such hierarchy of body-parts, we accordingly build a hierarchical RRV (Rotation and Relative Velocity) descriptors. The hierarchical representations encoded by Fisher vectors of the hierarchical RRV descriptors are properly formulated into the hierarchical model via the proposed hierarchical mixed norm, to apply sparse selection of body-parts and regularize the structure of such hierarchy of body-parts. The extensive evaluations on three challenging datasets demonstrate the effectiveness of our proposed approach, which achieves superior performance compared to state-of-the-art results on different sizes of datasets, showing it is more widely applicable than existing approaches. Zhanpeng Shao, Youfu Li 0001, Yao Guo 0002, Jianyu Yang 0002, Zhenhua Wang 0003 |
ICRA | 3 |
| 2018 | DSRF: A flexible trajectory descriptor for articulated human action recognition
Yao Guo 0002, Youfu Li 0001, Zhanpeng Shao |
Pattern Recognit. | 1 |
| 2018 | RRV: A Spatiotemporal Descriptor for Rigid Body Motion RecognitionabstractThe motion behaviors of a rigid body can be characterized by a six degrees of freedom motion trajectory, which contains the 3-D position vectors of a reference point on the rigid body and 3-D rotations of this rigid body over time. This paper devises a rotation and relative velocity (RRV) descriptor by exploring the local translational and rotational invariants of rigid body motion trajectories, which is insensitive to noise, invariant to rigid transformation and scale. The RRV descriptor is then applied to characterize motions of a human body skeleton modeled as articulated interconnections of multiple rigid bodies. To show the descriptive ability of our RRV descriptor, we explore its potentials and applications in different rigid body motion recognition tasks. The experimental results on benchmark datasets demonstrate that our RRV descriptor learning discriminative motion patterns can achieve superior results for various recognition tasks. Yao Guo 0002, Youfu Li 0001, Zhanpeng Shao |
IEEE Trans. Cybern. | 1 |
| 2017 | MSM-HOG: A flexible trajectory descriptor for rigid body motion recognitionabstractThis paper proposes a flexible descriptor for representing 6-D rigid body motion trajectories, which not only shows strong invariances and descriptive ability but also achieves satisfactory results in both recognition accuracy and efficiency. 6-D rigid body motion trajectories are first transformed into the Multi-layer Self-similarity Matrices (MSM) representation. The MSM is the combination of the square similarity matrices in three layers, which captures both local and global spatiotemporal features of the trajectories. Next, the Histogram of Oriented Gradients (HOG) features extracted from the MSM representation are concatenated as the final MSM-HOG trajectory descriptor. Then we train the Support Vector Machine (SVM) classifier with the linear kernel for multicalss motion recognition. Finally, rigid body motion recognition experiments on two public datasets are conducted to verify the effectiveness and efficiency of the proposed method. Yao Guo 0002, Youfu Li 0001, Zhanpeng Shao |
IROS | 1 |
| 2017 | On Multiscale Self-Similarities Description for Effective Three-Dimensional/Six-Dimensional Motion Trajectory RecognitionabstractMotion trajectories provide compact informative clues in characterizing motion behaviors of human bodies, robots, and moving objects. This paper devises an invariant and unified descriptor for three-dimensional/six-dimensional (3-D/6-D) motion trajectories recognition by exploring the latent motion patterns in the multiscale self-similarity matrices (MSM) within a motion trajectory and its components. The MSM approach transforms a motion trajectory in Euclidean space into a set of similarity matrices and exhibits strong invariances, in which each matrix can be regarded as a grayscale image. Next, the histograms of oriented gradients features extracted from the MSM representation are concatenated as the final trajectory descriptor. In addition, an improved kernel MSM is raised by calculating the pairwise kernel distances. Finally, extensive 3-D/6-D motion trajectory recognition experiments on three public datasets with a linear support vector machine classifier are conducted to verify the effectiveness and efficiency of the proposed approach. Yao Guo 0002, Youfu Li 0001, Zhanpeng Shao |
IEEE Trans. Ind. Informatics | 1 |
| 2015 | An Exponential-Rayleigh Model for RSS-Based Device-Free Localization and TrackingabstractA common technical difficulty in device-free localization and tracking (DFLT) with a wireless sensor network is that the change of the received signal strength (RSS) of the link often becomes more unpredictable due to the multipath interferences. This challenge can lead to unsatisfactory or even unstable DFLT performance. This work focuses on developing a new RSS model, called Exponential-Rayleigh (ER) model, for addressing this challenge. Based on data from our extensive experiments, we first develop the ER model of the received signal strength. This model consists of two parts: the large-scale exponential attenuation part and the small-scale Rayleigh enhancement part. The new consideration on using the Rayleigh model is to depict the target-induced multipath components. We then explore the use of the ER model with a particle filter in the context of multi-target localization and tracking. Finally, we experimentally demonstrate that our ER model outperforms the existing models. The experimental results highlight the advantages of using the Rayleigh model in mitigating the multipath interferences thus improving the DFLT performance. Yao Guo 0002, Kaide Huang, Nanyong Jiang, Xuemei Guo, Youfu Li 0001, Guoli Wang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2014 | Heterogeneous Bayesian compressive sensing for sparse signal recoveryabstractThis study focuses on the issue of sparse signal recovery with sparse Bayesian learning in the context of a heterogeneous noise model, called by the heterogeneous Bayesian compressive sensing. The main contribution is to exploit the capability of noise variance learning in performance improvement and applicability enhancement. Experimental results on synthetic and real‐world data demonstrate that heterogeneous Bayesian compressive sensing has superior performance in terms of accuracy and sparsity for both homogeneous and heterogeneous noise scenarios. Kaide Huang, Yao Guo 0002, Xuemei Guo, Guoli Wang 0001 |
IET Signal Process. | 2 |