Xiao Gu 0003

dblp:41/2698-3 · DBLP profile ↗
← Back
25ranked-venue papers
6as first author
22since 2021 · last 2025
0000-0002-3015-5818ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 7 since 2021Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Information Transfer Across Clinical Tasks via Adaptive Parameter Optimisation
abstract
This paper presents Adaptive Parameter Optimisation (APO), a novel framework for optimising shared models across multiple clinical tasks, addressing the challenges of balancing strict parameter sharing—often leading to task conflicts—and soft parameter sharing, which may limit effective cross-task information exchange. The proposed APO framework leverages insights from the lazy behaviour observed in over-parameterised neural networks, where only a small subset of parameters undergo any substantial updates during training. APO dynamically identifies and updates task-specific parameters while treating parameters associated with other tasks as protected, limiting their modification to prevent interference. The remaining unassigned parameters remain unchanged, embodying the lazy training phenomenon. This dynamic management of task-specific, protected, and unclaimed parameters across tasks enables effective information sharing, preserves task-specific adaptability, and mitigates gradient conflicts without enforcing a uniform representation. Experimental results across diverse healthcare datasets demonstrate that APO surpasses traditional information-sharing approaches, such as multi-task learning and model-agnostic meta-learning, in improving task performance.
Anshul Thakur, Elena Gal, Soheila Molaei, Xiao Gu 0003, Patrick Schwab, Danielle Belgrave, Kim Branson 0001, David A. Clifton
AISTATS4
2025 An EEG Conformer Model for Error Feedback During Human-Robot Interaction
abstract
Identifying a brain signal that enables the detection of incorrect execution in human-robot interaction (HRI) is considered a holy grail for real-time systems. A major challenge in achieving this is the inherent imbalance caused by the sparsity of error-related potential (ErrP) events in streaming electroencephalogram (EEG) data, which often leads models to learn irrelevant features and perform poorly. Thus, while deep learning-based ErrP detection has seen considerable advancements, the variability in individual user reaction times introduces labeling errors, complicating model adaptation to new subjects. Moreover, most deep learning methods are developed and validated on discrete, offline experiments using pre-defined windows, which fail to translate effectively to continuous, real-time HRI. Addressing these challenges is crucial to improving the robustness and adaptability of real-time ErrP detection in practical HRI applications. Here, we develop a causal EEG conformer framework, combining a convolutional neural network (CNN) encoder and a transformer with causal attention for real-time prediction of ErrP signals during HRI. We evaluated our ErrP model in a pseudo-online environment in both inter-session and inter-subject cross-validation settings for exoskeleton assistive robotics. Our model demonstrated superior performance in decoding accuracy and efficiency, showcasing better generalization for real-world dynamic HRI applications.
Jinpei Han, Yinxuan Li, Xiao Gu 0003, A. Aldo Faisal
ICRA3
2025 Learning Semi-Supervised Medical Image Segmentation from Spatial Registration
abstract
Semi-supervised medical image segmentation has shown promise in training models with limited labeled data and abundant unlabeled data. However, state-of-the-art methods ignore a potentially valuable source of unsupervised semantic information-spatial registration transforms between image volumes. To address this, we propose CCT-R, a contrastive cross-teaching framework incorporating registration information. To leverage the semantic information available in registrations between volume pairs, CCT-R incorporates two proposed modules: Registration Supervision Loss (RSL) and Registration-Enhanced Positive Sampling (REPS). The RSL leverages segmentation knowledge derived from transforms between labeled and unlabeled volume pairs, providing an additional source of pseudo-labels. REPS enhances contrastive learning by identifying anatomically-corresponding positives across volumes using registration transforms. Experimental results on two challenging medical segmentation benchmarks demonstrate the effectiveness and superiority of CCT-R across various semi-supervised settings, with as few as one labeled case. Our code is available at https://github.com/kathyliu579/ContrastiveCross-teachingWithRegistration.
Qianying Liu, Paul Henderson, Xiao Gu 0003, Hang Dai, Fani Deligianni
WACV3
2025 PoseSDF++: Point Cloud-Based 3-D Human Pose Estimation via Implicit Neural Representation
abstract
Predicting accurate human pose from 3-D visual observation presents a formidable challenge in computer vision, with numerous applications across various industries. However, most existing studies tackled this issue by regressing the 3-D pose from depth maps via 2-D convolutional neural networks or parametric human models, with limited development in point cloud-based methods. To this end, we propose PoseSDF++, i.e., a point cloud-based encoder–decoder network utilizing implicit neural representation to perform 3-D human pose estimation (HPE) and nonparametric shape reconstruction simultaneously. Leveraging the representative capacity of the signed distance function (SDF), we conceptualize the 3-D HPE as a multiple-shape reconstruction task and propose a distance-aware regression method to accurately estimate the 3-D joint positions. In specific, our PoseSDF++ consists of three modules: first,a hierarchical encoderwith vector neuron layers extracts the multiscale rotation equivariant features from the point clouds captured from an arbitrary viewpoint, addressing the degradation issue caused by viewpoint variation of implicit representation; second,a shape decodermaps the extracted feature and the query to its corresponding shape SDF; third,a pose decodercomputes the distance between the query and the target keypoints, namely, the pose SDF. Extensive experiments on four publicly available datasets demonstrate that our PoseSDF++ achieves competitive performance against the state-of-the-art point cloud-based methods and covering the human hand (HANDS 2019), lower limbs (ICL-Gait), and full body (DFAUST, LiDARHuman2.6M) pose estimation.
Jianxin Yang, Yuxuan Liu 0013, Xiao Gu 0003, Guang-Zhong Yang, Yao Guo 0002
IEEE Trans. Ind. Informatics4
2025 Unsupervised Domain Adaptation With Synchronized Self-Training for Cross- Domain Motor Imagery Recognition
abstract
Robust decoding performance is essential for the practical deployment of brain-computer interface (BCI) systems. Existing EEG decoding models often rely on large amounts of annotated data collected through specific experimental setups, which fail to address the heterogeneity of data distributions across different domains. This limitation hinders BCI systems from effectively managing the complexity and variability of real-world data. To overcome these challenges, we propose Synchronized Self-Training Domain Adaptation (SSTDA) for cross-domain motor imagery classification. Specifically, SSTDA leverages labeled signals from a source domain and applies self-training to unlabeled signals from a target domain, enabling the simultaneous training of a more robust classifier. The raw EEG signals are mapped into a latent space by a feature extractor for discriminative representation learning. A domain-shared latent space is then learned by optimizing the feature extractor with both source and target samples, using an easy-tohard self-training process. We validate the method with extensive experiments on two public motor imagery datasets: Dataset IIa of BCI Competition IV and the High Gamma dataset. In the inter-subject task, our method achieves classification accuracies of 64.43% and 80.40%, respectively. It also outperforms existing methods in the inter-session task. Moreover, we develope a new six-class motor imagery dataset and achieve test accuracies of 77.09% and 80.18% across different datasets. All experimental results demonstrate that our SSTDA outperforms existing algorithms in inter-session, inter-subject, and inter-dataset validation protocols, highlighting its capability to learn discriminative, domain-invariant representations that enhance EEG decoding performance.
Peiyin Chen, Xiaofeng Liu 0006, Chao Ma 0015, He Wang 0049, Xiong Yang 0001, Celso Grebogi, Xiao Gu 0003, Zhongke Gao
IEEE J. Biomed. Health Informatics7
2024 Egocentric Image Captioning for Privacy-Preserved Passive Dietary Intake Monitoring
abstract
Camera-based passive dietary intake monitoring is able to continuously capture the eating episodes of a subject, recording rich visual information, such as the type and volume of food being consumed, as well as the eating behaviors of the subject. However, there currently is no method that is able to incorporate these visual clues and provide a comprehensive context of dietary intake from passive recording (e.g., is the subject sharing food with others, what food the subject is eating, and how much food is left in the bowl). On the other hand, privacy is a major concern while egocentric wearable cameras are used for capturing. In this article, we propose a privacy-preserved secure solution (i.e., egocentric image captioning) for dietary assessment with passive monitoring, which unifies food recognition, volume estimation, and scene understanding. By converting images into rich text descriptions, nutritionists can assess individual dietary intake based on the captions instead of the original images, reducing the risk of privacy leakage from images. To this end, an egocentric dietary image captioning dataset has been built, which consists of in-the-wild images captured by head-worn and chest-worn cameras in field studies in Ghana. A novel transformer-based architecture is designed to caption egocentric dietary images. Comprehensive experiments have been conducted to evaluate the effectiveness and to justify the design of the proposed architecture for egocentric dietary image captioning. To the best of our knowledge, this is the first work that applies image captioning for dietary intake assessment in real-life settings.
Jianing Qiu, Frank P.-W. Lo, Xiao Gu 0003, Modou L. Jobarteh, Wenyan Jia, Thomas Baranowski, Matilda Steiner-Asiedu, Alex K. Anderson, Megan A. McCrory, Edward Sazonov, Mingui Sun, Gary S. Frost, Benny P. L. Lo
IEEE Trans. Cybern.3
2024 Noise-Factorized Disentangled Representation Learning for Generalizable Motor Imagery EEG Classification
abstract
Motor Imagery (MI) Electroencephalography (EEG) is one of the most common Brain-Computer Interface (BCI) paradigms that has been widely used in neural rehabilitation and gaming. Although considerable research efforts have been dedicated to developing MI EEG classification algorithms, they are mostly limited in handling scenarios where the training and testing data are not from the same subject or session. Such poor generalization capability significantly limits the realization of BCI in real-world applications. In this paper, we proposed a novel framework to disentangle the representation of raw EEG data into three components, subject/session-specific, MI-task-specific, and random noises, so that the subject/session-specific feature extends the generalization capability of the system. This is realized by a joint discriminative and generative framework, supported by a series of fundamental training losses and training strategies. We evaluated our framework on three public MI EEG datasets, and detailed experimental results show that our method can achieve superior performance by a large margin compared to current state-of-the-art benchmark algorithms.
Jinpei Han, Xiao Gu 0003, Guang-Zhong Yang, Benny P. L. Lo
IEEE J. Biomed. Health Informatics2
2024 A Survey of Wearable Lower Extremity Neurorehabilitation Exoskeleton: Sensing, Gait Dynamics, and Human-Robot Collaboration
abstract
The lower extremity exoskeleton, which can sense the neural motion state of the human body and then provide motion assistance, is gradually replacing the traditional wheelchairs and assistive devices, making many patients with disabilities or movement disorders able to regain the walking function. This survey provides a comprehensive review on recent technological advances in lower extremity neurorehabilitation exoskeleton from the perspectives of sensing, gait dynamics, and human–robot collaboration. For each technology category, a detailed comparison among state-of-the-art solutions is provided. The results show that the exoskeleton has been greatly improved in mechanical and learning ability. However, some issues, such as adaptability, safety, and efficiency still restrict the development of exoskeleton technology. To address these problems, the remaining open challenges and future directions to improve intelligence, sensing, gait analysis, trust, efficiency, generalization, and power consumption of exoskeleton are also presented and discussed.
Jie Li 0009, Xiao Gu 0003, Sen Qiu, Xu Zhou 0002, Angelo Cangelosi, Chu Kiong Loo, Xiaofeng Liu 0006
IEEE Trans. Syst. Man Cybern. Syst.2
2023 Multi-Scale Cross Contrastive Learning for Semi-Supervised Medical Image Segmentation
Qianying Liu, Xiao Gu 0003, Paul Henderson, Fani Deligianni
BMVC2
2023 Generalizable Movement Intention Recognition with Multiple Heterogeneous EEG Datasets
abstract
Human movement intention recognition is important for human-robot interaction. Existing work based on motor imagery electroencephalogram (EEG) provides a non-invasive and portable solution for intention detection. However, the data-driven methods may suffer from the limited scale and diversity of the training datasets, which result in poor generalization performance on new test subjects. It is practically difficult to directly aggregate data from multiple datasets for training, since they often employ different channels and collected data suffers from significant domain shifts caused by different devices, experiment setup, etc. On the other hand, the inter-subject heterogeneity is also substantial due to individual differences in EEG representations. In this work, we developed two networks to learn from both the shared and the complete channels across datasets, handling inter-subject and inter-dataset heterogeneity respectively. Based on both networks, we further developed an online knowledge co-distillation framework to collaboratively learn from both networks, achieving coherent performance boosts. Experimental results have shown that our proposed method can effectively aggregate knowledge from multiple datasets, demonstrating better generalization in the context of cross-subject validation.
Xiao Gu 0003, Jinpei Han, Guang-Zhong Yang, Benny P. L. Lo
ICRA1
2023 EgoHMR: Egocentric Human Mesh Recovery via Hierarchical Latent Diffusion Model
abstract
Egocentric vision has gained increasing popularity in social robotics, demonstrating great potentials for personal assistance and human-centric behavior analysis. Holistic per-ception of human body itself is a prerequisite for downstream applications, including action recognition and anticipation. Extensive research has been performed for human mesh recovery from the exocentric images captured from a third-person view, but limited studies are conducted for heavily distorted yet occluded egocentric images. In this paper, we propose Egocentric Human Mesh Recovery (EgoHMR), a novel hierarchical network based on latent diffusion models. Our method takes a single egocentric frame as the input and it can be trained in an end-to-end manner without supervision of 2D pose. The network is built upon the latent diffusion model by incorporating both global and local features in a hierarchical structure. To train the proposed network, we generate weak labels from synchronized exocentric images. The proposed method can perform human mesh recovery directly from egocentric images and detailed quantitative and qualitative experiments have been conducted to demonstrate the effectiveness of the proposed EgoHMR method.
Yuxuan Liu 0013, Jianxin Yang, Xiao Gu 0003, Yao Guo 0002, Guang-Zhong Yang
ICRA3
2023 Trustworthy learning with (un)sure annotation for lung nodule diagnosis with CT
Liang Chen 0023, Xiao Gu 0003, Yulei Qin, Zhexin Wang, Yun Gu, Guang-Zhong Yang
Medical Image Anal.3
2023 EgoFish3D: Egocentric 3D Pose Estimation From a Fisheye Camera via Self-Supervised Learning
abstract
Egocentric vision has gained increasing popularity recently, opening new avenues for human-centric applications. However, the use of the egocentric fisheye cameras allows wide angle coverage but image distortion is introduced along with strong human body self-occlusion imposing significant challenges in data processing and model reconstruction. Unlike previous work only leveraging synthetic data for model training, this paper presents a new real-world EgoCentric Human Pose (ECHP) dataset. To tackle the difficulty of collecting 3D ground truth using motion capture systems, we simultaneously collect images from a head-mounted egocentric fisheye camera as well as from two third-person-view cameras, circumventing the environmental restrictions. By using self-supervised learning under multi-view constraints, we propose a simple yet effective framework, namely EgoFish3D, for egocentric 3D pose estimation from a single image in different real-world scenarios. The proposed EgoFish3D incorporates three main modules. 1)The third-person-view moduletakes two exocentric images as input and estimates the 3D pose represented in the third-person camera frame; 2)the egocentric modulepredicts the 3D pose in the egocentric camera frame; and 3)the interactive moduleestimates the rotation matrix between the third-person and the egocentric views. Experimental results on our ECHP dataset and existing benchmark datasets demonstrate the effectiveness of the proposed EgoFish3D, which can achieve superior performance to existing methods.
Yuxuan Liu 0013, Jianxin Yang, Xiao Gu 0003, Yao Guo 0002, Guang-Zhong Yang
IEEE Trans. Multim.3
2022 Revisiting Self-Supervised Contrastive Learning for Facial Expression Recognition
Yuxuan Shu, Xiao Gu 0003, Guang-Zhong Yang, Benny P. L. Lo
BMVC2
2022 Tackling Long-Tailed Category Distribution Under Domain Shifts
Xiao Gu 0003, Yao Guo 0002, Zeju Li, Jianing Qiu, Qi Dou 0001, Yuxuan Liu 0013, Benny P. L. Lo, Guang-Zhong Yang
ECCV (23)1
2022 PoseSDF: Simultaneous 3D Human Shape Reconstruction and Gait Pose Estimation Using Signed Distance Functions
abstract
Vision-based 3D human pose estimation and shape reconstruction play important roles in robot-assisted healthcare monitoring and personal assistance. However, 3D data captured from a single viewpoint always encounter occlusions and exhibit substantial heterogeneity across different views, resulting in significant challenges for both tasks. Extensive approaches have been proposed to perform each task separately, but few of them present a unified solution. In this paper, we propose a novel network based on signed distance functions, namely PoseSDF, to simultaneously reconstruct 3D lower limb shape and estimate gait pose by two dedicated branches. To promote multi-task learning, several strategies are developed to ensure that these two branches leverage the same latent shape code while exchanging information between them. More importantly, an auxiliary RotNet is incorporated into the inference phase, overcoming the inherent limitations of implicit neural functions under cross-view scenarios. Experimental results demonstrate that our proposed PoseSDF can achieve both high-quality shape reconstruction and precise pose estimation, generalizing well on the data from novel views, gait patterns, as well as real-world.
Jianxin Yang, Yuxuan Liu 0013, Xiao Gu 0003, Guang-Zhong Yang, Yao Guo 0002
ICRA3
2022 Ego+X: An Egocentric Vision System for Global 3D Human Pose Estimation and Social Interaction Characterization
abstract
Egocentric vision is an emerging topic, which has demonstrated great potential in assistive healthcare scenarios, ranging from human-centric behavior analysis to personal social assistance. Within this field, due to the heterogeneity of visual perception from first-person views, egocentric pose estimation is one of the most significant prerequisites for enabling various downstream applications. However, existing methods for egocentric pose estimation mainly focus on predicting the pose represented in the camera coordinates from a single image, which ignores the latent cues in the temporal domain and results in less accuracy. In this paper, we propose Ego+X, an egocentric vision based system for 3D canonical pose estimation and human-centric social interaction characterization. Our system is composed of two head-mounted egocentric cameras, where one is faced downwards and the other looks outwards. By leveraging the global context provided by visual SLAM, we first propose Ego-Glo for spatial-accurate and temporal-consistent egocentric 3D pose estimation in the canonical coordinate system. With the help of an egocentric camera looking outwards, we then propose Ego-Soc by extending Ego-Glo to various social interaction tasks, e.g., object detection and human-human interaction. Quantitative and qualitative experiments have been conducted to demonstrate the effectiveness of our proposed Ego+X.
Yuxuan Liu 0013, Jianxin Yang, Xiao Gu 0003, Yao Guo 0002, Guang-Zhong Yang
IROS3
2022 Cross-Domain Self-Supervised Complete Geometric Representation Learning for Real-Scanned Point Cloud Based Pathological Gait Analysis
abstract
Accurate lower-limb pose estimation is aprerequisite of skeleton based pathological gait analysis. To achieve this goal in free-living environments for long-term monitoring, single depth sensor has been proposed in research. However, the depth map acquired from a single viewpoint encodes only partial geometric information of the lower limbs and exhibits large variations across different viewpoints. Existing off-the-shelf 3D pose tracking algorithms and public datasets for depth based human pose estimation are mainly targeted at activity recognition applications. They are relatively insensitive to skeleton estimation accuracy, especially at the foot segments. Furthermore, acquiring ground truth skeleton data for detailed biomechanics analysis also requires considerable efforts. To address these issues, we propose a novel cross-domain self-supervised complete geometric representation learning framework, with knowledge transfer from the unlabelled synthetic point clouds of full lower-limb surfaces. The proposed method can significantly reduce the number of ground truth skeletons (with only 1%) in the training phase, meanwhile ensuring accurate and precise pose estimation and capturing discriminative features across different pathological gait patterns compared to other methods.
Xiao Gu 0003, Yao Guo 0002, Guang-Zhong Yang, Benny P. L. Lo
IEEE J. Biomed. Health Informatics1
2021 Semi-Supervised Contrastive Learning for Generalizable Motor Imagery EEG Classification
abstract
Electroencephalography (EEG) is one of the most widely used brain-activity recording methods in non-invasive brain-machine interfaces (BCIs). However, EEG data is highly nonlinear, and its datasets often suffer from issues such as data heterogeneity, label uncertainty and data/label scarcity. To address these, we propose a domain independent, end-to-end semi-supervised learning framework with contrastive learning and adversarial training strategies. Our method was evaluated in experiments with different amounts of labels and an ablation study in a motor imagery EEG dataset. The experiments demonstrate that the proposed framework with two different backbone deep neural networks show improved performance over their supervised counterparts under the same condition.
Jinpei Han, Xiao Gu 0003, Benny P. L. Lo
BSN2
2021 Indoor Future Person Localization from an Egocentric Wearable Camera
abstract
Accurate prediction of future person location and movement trajectory from an egocentric wearable camera can benefit a wide range of applications, such as assisting visually impaired people in navigation, and the development of mobility assistance for people with disability. In this work, a new egocentric dataset was constructed using a wearable camera, with 8,250 short clips of a targeted person either walking 1) toward, 2) away, or 3) across the camera wearer in indoor environments, or 4) staying still in the scene, and 13,817 person bounding boxes were manually labelled. Apart from the bounding boxes, the dataset also contains the estimated pose of the targeted person as well as the IMU signal of the wearable camera at each time point. An LSTM-based encoder-decoder framework was designed to predict the future location and movement trajectory of the targeted person in this egocentric setting. Extensive experiments have been conducted on the new dataset, and have shown that the proposed method is able to reliably and better predict future person location and trajectory in egocentric videos captured by the wearable camera compared to three baselines.
Jianing Qiu, Frank P.-W. Lo, Xiao Gu 0003, Yingnan Sun, Benny P. L. Lo
IROS3
2021 MCDCD: Multi-Source Unsupervised Domain Adaptation for Abnormal Human Gait Detection
abstract
For gait analysis, especially for the detection of subtle gait abnormalities, the collected datasets involve high variability across subjects due to inherent biometric traits and movement behaviors, leading to limited detection accuracy and poor generalizability. To address this, we propose a novel deep multi-source Unsupervised Domain Adaptation (UDA) approach, namely Maximum Cross-Domain Classifier Discrepancy (MCDCD), which aims to improve the classification performance on the test subject (target domain) by leveraging the information from multiple labelled training subjects (source domains). Specifically, the proposed model consists of a feature extractor and a domain-specific category classifier per source domain. The former feature extractor learns to generate discriminative gait features. For the latter classifiers, we minimize the cross-entropy loss to accurately classify source samples, and simultaneously maximize a novel cross-domain discrepancy loss between any two category classifiers to minimize domain shift between multiple sources and the target domain. To validate the proposed MCDCD for detecting gait abnormalities on novel subjects, we collected both high-quality Motion capture (Mocap) and noisy Electromyography (EMG) data from eighteen subjects with both normal and imitated abnormal gaits. Experiment results using both data modalities demonstrate that the proposed approach can achieve superior performance in abnormal gait classification compared to baseline deep models and state-of-the-art UDA methods.
Yao Guo 0002, Xiao Gu 0003, Guang-Zhong Yang
IEEE J. Biomed. Health Informatics2
2021 Cross-Subject and Cross-Modal Transfer for Generalized Abnormal Gait Pattern Recognition
abstract
For abnormal gait recognition, pattern-specific features indicating abnormalities are interleaved with the subject-specific differences representing biometric traits. Deep representations are, therefore, prone to overfitting, and the models derived cannot generalize well to new subjects. Furthermore, there is limited availability of abnormal gait data obtained from precise Motion Capture (Mocap) systems because of regulatory issues and slow adaptation of new technologies in health care. On the other hand, data captured from markerless vision sensors or wearable sensors can be obtained in home environments, but noises from such devices may prevent the effective extraction of relevant features. To address these challenges, we propose a cascade of deep architectures that can encode cross-modal and cross-subject transfer for abnormal gait recognition. Cross-modal transfer maps noisy data obtained from RGBD and wearable sensors to accurate 4-D representations of the lower limb and joints obtained from the Mocap system. Subsequently, cross-subject transfer allows disentangling subject-specific from abnormal pattern-specific gait features based on a multiencoder autoencoder architecture. To validate the proposed methodology, we obtained multimodal gait data based on a multicamera motion capture system along with synchronized recordings of electromyography (EMG) data and 4-D skeleton data extracted from a single RGBD camera. Classification accuracy was improved significantly in both Mocap and noisy modalities.
Xiao Gu 0003, Yao Guo 0002, Fani Deligianni, Benny P. L. Lo, Guang-Zhong Yang
IEEE Trans. Neural Networks Learn. Syst.1
2020 Coupled Real-Synthetic Domain Adaptation for Real-World Deep Depth Enhancement
abstract
Advances in depth sensing technologies have allowed simultaneous acquisition of both color and depth data under different environments. However, most depth sensors have lower resolution than that of the associated color channels and such a mismatch can affect applications that require accurate depth recovery. Existing depth enhancement methods use simplistic noise models and cannot generalize well under real-world conditions. In this paper, a coupled real-synthetic domain adaptation method is proposed, which enables domain transfer between high-quality depth simulators and real depth camera information for super-resolution depth recovery. The method first enables the realistic degradation from synthetic images, and then enhances degraded depth data to high quality with a color-guided sub-network. The key advantage of the work is that it generalizes well to real-world datasets without further training or fine-tuning. Detailed quantitative and qualitative results are presented, and it is demonstrated that the proposed method achieves improved performance compared to previous methods fine-tuned on the specific datasets.
Xiao Gu 0003, Yao Guo 0002, Fani Deligianni, Guang-Zhong Yang
IEEE Trans. Image Process.1
2018 Markerless gait analysis based on a single RGB camera
abstract
Gait analysis is an important tool for monitoring and preventing injuries as well as to quantify functional decline in neurological diseases and elderly people. In most cases, it is more meaningful to monitor patients in natural living environments with low-end equipment such as cameras and wearable sensors. However, inertial sensors cannot provide enough details on angular dynamics. This paper presents a method that uses a single RGB camera to track the 2D joint coordinates with state-of-the-art vision algorithms. Reconstruction of the 3D trajectories uses sparse representation of an active shape model. Subsequently, we extract gait features and validate our results in comparison with a state-of-the-art commercial multi-camera tracking system. Our results are comparable to those from the current literature based on depth cameras and optical markers to extract gait characteristics.
Xiao Gu 0003, Fani Deligianni, Benny P. L. Lo, Wei Chen 0015, Guang-Zhong Yang
BSN1
2017 A wearable sensor system for neonatal seizure monitoring
abstract
A novel wearable sensor system for seizure monitoring of neonates comprised of smart clothing, video recording and cloud platform is presented. Textile electrodes and Inertial Measurement Unit (IMU) are embedded in the smart clothing to obtain ECG signal and motion signal whereby epileptic seizure detection algorithm is performed. Moreover, a video monitoring module provides real-time information about patients. The cloud platform receives the pre-processed data and enables remote monitoring, centralized signal processing and data management. Comparison with commercial instruments shows that the smart clothing is capable of acquiring high-quality signals. Pilot tests under disinfection operations at Children's Hospital of Fudan University confirm clinical feasibility of the proposed system. The scalability and modularity of the unobtrusive wearable front end and the design of system architecture based on cloud enable the whole system with great potential in clinical practice and home monitoring scenarios.
Hongyu Chen 0002, Xiao Gu 0003, Zhenning Mei, Ke Xu 0006, Chunmei Lu, Laishuan Wang, Feng Shu 0001, Qixin Xu, Sidarto Bambang-Oetomo, Wei Chen 0015
BSN2