VLDB 2026 Research / reviewers in the wild / expert
Hong Cheng 0002
dblp:85/5637-2
· DBLP profile ↗
114ranked-venue papers
18as first author
54since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 69 · 8 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 12 first-author · 13 since 2021Systems, architecture and hardware · 18 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3Computer networks · 2 · 2 since 2021Theory of computation · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Optimal Tracking Control of Uncertain Robotic Systems With Global Predefined-Time Stability: A Unified Observer-Identifier-Learning FrameworkabstractThis paper proposes a unified observer–identifier–learning framework (OILF) for predefined-time optimal tracking control with prescribed performance for robotic systems subject to unmeasurable states and uncertain dynamics. Most existing optimal control approaches rely on full-state information or accurate dynamic models, which are often unavailable in practice. To overcome this issue, a predefined-time dynamic regression extension and mixing (PTDREM) method is proposed to realize the co-design of state observer and parameter identifier, enabling synchronous predefined-time estimation of unknown states and dynamic parameters. Subsequently, to achieve optimal tracking control for robotic systems, a prescribed-performance-based critic–actor (PPCA) structure is developed via reinforcement learning (RL), in which all hierarchical tracking errors are driven into prescribed neighborhoods of the origin within a predefined time. In contrast to most existing works that solely ensure uniform ultimate boundedness (UUB) of the closed-loop system, the proposed scheme enables the upper bounds of the convergence time of the state observer, system identifier, and optimal controller to be preset through independent parameter design, thereby establishing global predefined-time stability (G-PTS) for the overall closed-loop system. Numerical simulations on a two-degree-of-freedom (DOF) robotic manipulator verify the effectiveness of the proposed OILF. Lin Hao, Rui Luo 0003, Zhinan Peng, Linpu He, Zhipeng Du, Rui Huang 0008, Hong Cheng 0002, Bijoy K. Ghosh |
IEEE Internet Things J. | 8 |
| 2026 | Fixed-time learning-based optimal tracking control for robotic systems with prescribed performance constraints
Zhinan Peng, Zhuo Xia, Lin Hao, Linpu He, Hong Cheng 0002 |
Neural Networks | 6 |
| 2026 | SegMIC: A universal model for medical image segmentation through in-context learning
Fan Yang 0054, Xin Li 0079, Zhicheng Jiao, Qiang Zhai, Xiaomeng Li 0001, De Wu, Huazhu Fu, Hong Cheng 0002 |
Pattern Recognit. | 9 |
| 2026 | Dynamics-Based Collaborative Control for an Exoskeleton-Walker System via Deterministic Learning
Weitian He, Chaobin Zou, Fukai Zhang, Hong Cheng 0002, Cong Wang 0007 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | Adaptive Hierarchical Event-Triggered H∞ Output Tracking of IT2 Fuzzy Heterogeneous Multiagent Systems Under Multiple-Channel DoS AttacksabstractOutput tracking control has been extensively applied in the cooperative control of multiagent systems (MASs), including mobile robot obstacle avoidance and autonomous aerial vehicle formation. This article investigates the $H_{\infty }$ output consensus tracking problem of interval type-2 (IT2) fuzzy heterogeneous MASs subject to multiple-channel denial-of-service (DoS) attacks. First, in the presence of DoS attacks on both leader-follower and follower-follower communication channels, a fully distributed adaptive compensator is designed to approximate the convex hull of the leader's state. Then, to address DoS attacks occurring in the compensator-controller and observer-controller channels, an observer-based fuzzy switched controller is developed for each follower agent to accomplish the tracking objective. Furthermore, a two-layer hierarchical hybrid event-triggered mechanism (ETM) is established to significantly reduce the communication burden of the MAS network. In the proposed ETM, asynchronous communication and Zeno behavior are rigorously excluded, while the triggering frequency is effectively decreased. Moreover, a sufficient condition is derived to guarantee the exponential stability of the tracking error with a prescribed $H_{\infty }$ performance. Finally, simulation results are provided to demonstrate the feasibility and superiority of the proposed approach. Sheng Han 0002, Hong Zhu 0001, Lanfeng Hua, Kaibo Shi, Zhinan Peng, Hong Cheng 0002, Yeng Chai Soh |
IEEE Trans. Cybern. | 6 |
| 2026 | Spatiotemporal Dynamics Modeling of Brain Activity for Human-Robot Cognitive Interaction: A Distributed-Lumped Parameter System FrameworkabstractThis article investigates the system modeling problem for the dynamical process of human brain activity in human-robot cognitive interaction (HRCI). An important novelty of the proposed approaches is to build a computational model of a human-distributed robot-lumped parameter system (HDRLPS) that describes the inherent dynamical principle of human brain activity (with spatiotemporal-varying characteristic) undergoing the interaction between the intrinsic cognitive dynamics and extrinsic robot stimuli. A deterministic learning (DL)-based spatiotemporal dynamics identification scheme is proposed to accurately identify the spatiotemporal dynamics of HDRLS and obtain the associated knowledge as a constant radial basis functional neural network (RBF NN) model. A spatiotemporal dynamics estimator is designed with this model, which can accurately evaluate and monitor the dynamical process of human brain activity in real-time HRCI by the generated dynamics-synchronized state. The effectiveness and practicability of the approaches in the dynamics identification and evaluation for the human brain activity in HRCI are validated by the thorough analysis, including the mathematical proof, the simulation study, and the brain-computer interface (BCI) experiment using publicly available datasets. Our method is compared with state-of-the-art (SOTA) methods, such as LGGNet, EEGNet, Tsception, EEG-Deformer, EEG-Transformer, and EEGViT. The results show that our method can outperform these methods with better recognition accuracy and macro- $F1$ scores. The source code can be found at: https://github.com/alonexing/source_code/tree/master. Jingting Zhang, Lianchi Zhang, Fengjun Mu, Zonghai Huang, Chaobin Zou, Rui Huang 0008, Cong Wang 0007, Hong Cheng 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2025 | MExD: An Expert-Infused Diffusion Model for Whole-Slide Image ClassificationabstractWhole Slide Image (WSI) classification poses unique challenges due to the vast image size and numerous non-informative regions, which introduce noise and cause data imbalance during feature aggregation. To address these issues, we propose MExD, an Expert-Infused Diffusion Model that combines the strengths of a Mixture-of-Experts (MoE) mechanism with a diffusion model for enhanced classification. MExD balances patch feature distribution through a novel MoE-based aggregator that selectively emphasizes relevant information, effectively filtering noise, addressing data imbalance, and extracting essential features. These features are then integrated via a diffusion-based generative process to directly yield the class distribution for the WSI. Moving beyond conventional discriminative approaches, MExD represents the first generative strategy in WSI classification, capturing fine-grained details for robust and precise results. Our MExD is validated on three widely-used benchmarks—Camelyon16, TCGANSCLC, and BRACS—consistently achieving state-of-the-art performance in both binary and multi-class tasks. Our code and model are available at https://github.com/JWZhao-uestc/MExD. Xin Li 0079, Fan Yang 0054, Qiang Zhai, Ao Luo, Yang Zhao 0024, Hong Cheng 0002, Huazhu Fu |
CVPR | 7 |
| 2025 | Plug-and-Play Multi-Domain Fusion Adaptation for Cross-Subject EEG-Based Motor Imagery ClassificationabstractMotor imagery (MI) classification in rehabilitation brain-computer interfaces (RBCIs) faces significant challenges due to the variability of electroencephalography (EEG) signals across subjects. Existing methods typically require extensive EEG data collection from each new subject, which is time-consuming and results in poor user experience. To address this issue, this paper decompose MI-EEG into subject-specific private components and shared components common across all subjects, and propose a plug-and-play domain fusion adaptive method (PPMDFA) to handle variability between subjects. In the training phase, PPMDFA introduces a Multi-Domain Fusion Graph Convolutional Network (MDFGCN) module to extract shared and private features from the MI processes of source domain subjects. In the calibration phase, the method constructs private classifiers for the target new subject using the extracted shared features combined with a small amount of labeled data. During testing, PPMDFA leverages the similarity of private components to utilize knowledge from source subjects, thereby enhancing classification accuracy for target subjects' MI. We validated the proposed method on the PhysioNet and LLMBCImotion datasets. Experimental results show that PPMDFA achieves state-of-the-art classification accuracy on both datasets, with rapid adaptation to new subjects using only 20% of the data, reaching accuracies of 73.33% and 61.62%, demonstrating strong generalization ability and robustness. Rui Huang 0008, Jianzhi Lyu, Yang Zhao 0024, Guangkui Song, Hong Cheng 0002, Jianwei Zhang 0001 |
ICRA | 7 |
| 2025 | KneeMamba: A Multi-level Feature Extraction Model for Knee MRI Assisted Diagnosis with MambaabstractMagnetic resonance imaging (MRI) is highly important for diagnosing knee injuries because of its ability to provide detailed images. However, the key problem in intelligent MRI diagnosis is how to identify and extract the most relevant features from a large amount of data. In this paper, we present a single training stage method. This method can extract features from images, sequences, and anatomical planes at different levels simultaneously. First, we propose a nested Mamba module. This module allows parallel extraction of image-level and sequence-level features from the same anatomical plane in MRI. Second, an anatomical-level Mamba module is designed. It integrates features extracted from different anatomical planes. Finally, self-distillation is applied to improve the model’s feature extraction efficiency. Experimental results show that our proposed model outperforms sub-optimal models by 3.7% in accuracy and 2.4% in F1-score. Moreover, compared with the multi-stage model, the single-stage multi-class classification model greatly improves classification efficiency. In conclusion, the model developed in this paper can effectively and accurately conduct intelligent knee MRI diagnosis. Zonghai Huang, Jiatong Si, Jingting Zhang, Fengjun Mu, Rui Huang 0008, Hong Cheng 0002 |
IJCNN | 6 |
| 2025 | A VisuoMotor Human-Robot Interaction Framework for Attention-Motion-Integrated TrainingabstractFocus of attention is one of the most influential factors facilitating motor training performance. Most of robotic training methods have not well solved the negative effect of divided-attention on motor execution performance, resulting in limited rehabilitation efficiency for motor-cognitive dysfunction. In this study, we propose a novel visuomotor human-robot interaction framework by integrating a gaze-visual game and force-movement robot, to realize more efficient training for both attentional and motor function. An important novelty of this framework is to design a dynamical pattern recognition scheme for the hierarchical-coupled behavior of attentional and motor execution, to facilitate efficient human-robot interaction in both cognitive and motor perspectives. Specifically, an attentional-motor dynamical system modeling method is first developed by using the gaze, force and movement data collected from the human under different attentional-motor behavior. Then, an online dynamical pattern recognition scheme can be design with these models to online recognizing the human’s attentional and motor behavior states. The training robot system can dynamically adjust the parameters according to the recognition results, to guide the collaboration of both attentional and motor training. Experimental study are conducted to demonstrate the desired accuracy and efficiency of our designed approaches in attentional-motor behavior recognition and training. Chen Chen 0137, Shuhe Yuan, Jingting Zhang, Fengjun Mu, Chaobin Zou, Hong Cheng 0002 |
IROS | 6 |
| 2025 | Force-Sensor-free Contact Estimation for Lower Limb Exoskeleton Robots Based on Probabilistic Modeling and FusionabstractLower limb exoskeletons (LLEs) play a crucial role in assisting paraplegic patients with walking in outdoor environments characterized by complex terrains, including various stairs, slopes, and uneven grounds. However, most existing control methods for LLEs rely on predefined joint angles, lacking the flexibility to adapt to diverse terrains. This deficiency often leads to unexpected contacts between the feet of the LLEs and the ground, thereby disrupting the walking balance of the LLEs. In this paper, a novel force-sensor-free contact estimation method is proposed to tackle this problem. This method utilizes only the sensors already present on the LLEs, eliminating the need for any additional force sensors. The proposed approach is founded on the probabilistic modeling of gait phases, knee joint torques, foot heights, and the displacement of the center of mass. Moreover, Kalman filtering is employed to enhance the contact estimation accuracy by integrating multiple probabilistic models. Experiments were carried out on both robot simulation platforms and real exoskeleton robots. The experimental results demonstrate that the proposed approach can accurately estimate contacts during walking on flat ground and stairs. Specifically, it achieves an accuracy of 99% with a time deviation of 8 ms on the flat ground and an accuracy of 95% with a time deviation of 10 ms on stairs. Weigen Ye, Chaobin Zou, Jingting Zhang, Guangkui Song, Hong Cheng 0002 |
IROS | 7 |
| 2025 | Adaptive Coordinated Motion Planning for lower limb exoskeleton robots with a robotic walker
Chaobin Zou, Rui Huang 0008, Jingting Zhang, Zhinan Peng, Hong Cheng 0002 |
Eng. Appl. Artif. Intell. | 7 |
| 2025 | Dynamic event-driven composite optimal impedance control for lower limb exoskeletons using finite-time reinforcement learning
Linpu He, Zhinan Peng, Yongxiang Liu, Rui Luo 0003, Yiqun Kuang, Hong Cheng 0002 |
Expert Syst. Appl. | 6 |
| 2025 | From honeybee to helicopter: A low-cost dual bionic model for flight control under hazardous situations
Shiwen Pan, Fengshuo Yan, Hong Cheng 0002, Kun Guo 0004, Jiawei Xu 0004, Zhao-Hui Sun, Xiaoru Wanyan, Edmond Q. Wu |
Neurocomputing | 5 |
| 2025 | Reinforcement Learning-Based Fixed-Time Optimal Impedance Control for Human-Robot Collaboration With Input Disturbances
Linpu He, Zhinan Peng, Yongxiang Liu, Yiqun Kuang, Hong Cheng 0002, Bijoy K. Ghosh |
IEEE Internet Things J. | 6 |
| 2025 | Multi-Level Skeleton Self-Supervised Learning: Enhancing 3D action representation learning with Large Multimodal Models
Yang Chen 0039, Ling Wang 0013, Rui Huang 0008, Hong Cheng 0002 |
Knowl. Based Syst. | 6 |
| 2025 | Memory-Guided Transformer with group attention for knee MRI diagnosis
Rui Huang 0008, Zonghai Huang, Hantang Zhou, Qiang Zhai, Fengjun Mu, Huayi Zhan, Hong Cheng 0002 |
Pattern Recognit. | 7 |
| 2025 | EEG-Based Motor Imagery Classification With Tuned Heuristic Fusion Graph Convolutional Network for Rehabilitation TrainingabstractMotor imagery-based brain–computer interfaces (MI-BCIs) hold significant promise for rehabilitation training in individuals with neurological impairments such as stroke and spinal cord injury (SCI). Achieving precise and robust lower limb movement prediction for each patient is crucial. However, the variability in MI response frequencies and brain activation patterns among subjects presents a great challenge to the generalizability of MI-BCIs. This paper proposes a Tuned Heuristic Fusion Graph Convolutional Network (THFGCN) for limb movement prediction in rehabilitation scenarios. THFGCN innovatively designs a learnable EEG frequency band tuned module and a heuristic space topology module. These two modules allow for the intricate extraction of both frequency and spatial topological features, utilizing graph adjacency matrices that encapsulate channel correlations and spatial relationships, hence fostering individualized analysis and enhanced generalizability across subjects. Furthermore, a spatio-temporal convolution module paired with a feature map attention mechanism is proposed to extract the critical spatio-temporal features of electroencephalogram (EEG) data. Validation experiments on the PhysioNet and LLM-BCImotion datasets against six mainstream methods demonstrate that THFGCN outperforms state-of-theart methods, achieving 88.41% and 82.82% accuracy in the within-subject case, and 65.93% and 60.56% accuracy in the cross-subject case, respectively. Detailed frequency band weight and T-distributed Stochastic Neighbor Embedding visualization validate the effectiveness of proposed modules. Furthermore, feature interpretability analysis proves the extracted features’ profound MI task relevance, underlining THFGCN’s exceptional interpretability. Rui Huang 0008, Jianzhi Lyu, Fengjun Mu, Zhinan Peng, Chaobin Zou, Hong Cheng 0002, Jianwei Zhang 0001, Bijoy K. Ghosh |
IEEE Trans Autom. Sci. Eng. | 8 |
| 2025 | Optimization-Based Adaptive Assistance for Lower Limb Exoskeleton Robots With a Robotic Walker via Spatially Quantized GaitabstractGait training with human-like gait patterns can be provided by lower limb exoskeletons (LLEs) for patients with gait impairments. For patients with little effort to keep balance, using a mobile robotic walker to assist gait training with LLEs is an effective way. Since gait patterns are varying with walking speeds, it is a critical issue to coordinated control the robotic walker and the exoskeleton to obtain a natural and human-like walking posture. In this paper, a novel adaptive assistance approach named SQG-OPT is proposed to tackle the problem, which comprises of two parts: the Spatially Quantized Gait (SQG) and the optimization. The SQG generates reference joint angles and reference trajectory of the Center Of Mass (COM) for the human-exoskeleton system in space domain. The optimization part is constructed to convert the reference joint angles from the space domain to the time domain, which is based on the dynamics model of the human-exoskeleton-walker system and adaptive to different walking speeds. The proposed approach has been tested on the robot simulation platform CoppeliaSim, the experimental results indicate that the proposed approach can generate human-like gait patterns for different walking speeds from 0 to 0.8 m/s. Additionally, in comparison with other methods, the proposed approach has a better performance on the movement tracking of the COM for a natural walking posture.Note to Practitioners—The coordinated control is important for the exoskeleton robot with a mobile robotic walker, one of the potential challenge is the adaptive coordinated motion planning for these robotic devices. The proposed approach is for the coordinated control of the exoskeleton robots with a robotic walker and adaptive to different walking speeds, which may inspires more extended coordinated control strategies for human-exoskeleton systems in more gait training applications. The proposed approach is also potential for the coordinated motion planning of the other human-centered assistance robots, such as the wheeled walking assistance robots for the elderly. Chaobin Zou, Zhinan Peng, Fengjun Mu, Rui Huang 0008, Hong Cheng 0002 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | Enhancing Skeleton-Based Action Recognition With Language Descriptions From Pre-Trained Large Multimodal ModelsabstractSkeleton data has become popular in human action recognition because of its efficacy in capturing human motion patterns while mitigating the influence of environmental noise. However, overlooking critical action-related environmental descriptors presents challenges in distinguishing actions characterized by similar body movements. To address this limitation, we propose a novel framework that integrates skeleton data with language descriptions to easily capture essential environmental information for fine-grained action recognition while maintaining the robustness of skeleton-based methods. We first develop a Language Environment Description Generation (LEDG) module that utilizes the open-world understanding ability of Large Multimodal Models to generate instance-level action-related language environment descriptions without the need to train additional modules. Then, we introduce a Skeleton-supported Environment Feature Extraction (SEFE) module that leverages the temporal dependency inherent in skeleton data to extract key semantic environmental features. Additionally, we propose an Entropy-based Feature Fusion (EFF) module to dynamically amalgamate complementary features from both skeleton and language domains. Experimental results demonstrate the superiority of our framework, which can improve the accuracy of existing skeleton-based action recognition methods and achieve state-of-the-art performance on four well-established skeleton-based action recognition benchmarks. Yang Chen 0039, Ling Wang 0013, Hong Cheng 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Human-Factors-in-Aviation-Loop: Multimodal Deep Learning for Pilot Situation Awareness Analysis Using Gaze Position and Flight Control DataabstractSituation awareness (SA) is a crucial factor affecting flight safety for pilots, yet few studies have focused specifically on modeling SA for pilots, resulting in limited success. In this paper, we propose a novel multimodal deep learning approach to monitor pilots’ SA. The approach combines handcrafted and deep features obtained from eye movement and flight control data collected from 27 novice pilots across different training phases using a flight simulator. Ground truth SA measurements were obtained using the Situation Awareness Global Assessment Technique (SAGAT). The handcrafted features included 13 eye movements and 22 flight control features, while deep features were extracted from time-series of gaze positions using a deep extractor based on Transformer. By fusing the handcrafted features of eye movement and flight control, along with one deep feature of eye movement, we predicted the final SA level. Through leave-one-flight-out cross-validation, our model achieved a higher accuracy of 92.04%. The results indicate that the multimodal model outperforms the unimodal models, with the eye movement modality demonstrating superiority over the flight control modality in predicting SA. This suggests our method provides an objective means of predicting pilot’s SA and offers new insights for SA assessment in aviation and other fields. Overall, our multimodal deep learning approach holds promise for enhancing pilot training and flight safety by facilitating a more comprehensive understanding of pilots’ SA during critical flight scenarios. Jiawei Xu 0004, Sicheng Pan, Zhao-Hui Sun, Kun Guo 0004, Seop Hyeong Park, Fengshuo Yan, Xiaoru Wanyan, Hong Cheng 0002, Qi Wu 0003 |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2025 | Vision-Language Meets the Skeleton: Progressively Distillation With Cross-Modal Knowledge for 3D Action Representation LearningabstractSkeleton-based action representation learning aims to interpret and understand human behaviors by encoding the skeleton sequences, which can be categorized into two primary training paradigms: supervised learning and self-supervised learning. However, the former one-hot classification requires labor-intensive predefined action categories annotations, while the latter involves skeleton transformations (e.g., cropping) in the pretext tasks that may impair the skeleton structure. To address these challenges, we introduce a novel skeleton-based training framework (C$^{2}$VL) based onCross-modalContrastive learning that uses the progressive distillation to learn task-agnostic human skeleton action representation from theVision-Language knowledge prompts. Specifically, we establish the vision-language action concept space through vision-language knowledge prompts generated by pre-trained large multimodal models (LMMs), which enrich the fine-grained details that the skeleton action space lacks. Moreover, we propose the intra-modal self-similarity and inter-modal cross-consistency softened targets in the cross-modal representation learning process to progressively control and guide the degree of pulling vision-language knowledge prompts and corresponding skeletons closer. These soft instance discrimination and self-knowledge distillation strategies contribute to the learning of better skeleton-based action representations from the noisy skeleton-vision-language pairs. During the inference phase, our method requires only the skeleton data as the input for action recognition and no longer for vision-language prompts. Extensive experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD datasets demonstrate that our method outperforms the previous methods and achieves state-of-the-art results. Yang Chen 0039, Junfeng Fu, Ling Wang 0013, Jingcai Guo, Hong Cheng 0002 |
IEEE Trans. Multim. | 7 |
| 2024 | FocusDiffuser: Perceiving Local Disparities for Camouflaged Object Detection
Xin Li 0079, Fan Yang 0054, Qiang Zhai, Ao Luo, Zicheng Jiao, Hong Cheng 0002 |
ECCV (53) | 7 |
| 2024 | Joint-Loss Enhanced Self-Supervised Learning for Refinement-Coupled Object 6D Pose Estimationabstract6D object pose estimation plays a crucial role in robot grasping and manipulation. However, the prevalent methods for 6D object pose estimation heavily rely on 6D annotated data to train deep neural networks, which poses challenges due to the difficulty in obtaining sufficient pose annotations. To address this limitation, this paper presents a self-supervised pose estimation method based on a novel pixelwise weighted dense fusion architecture. This method allows for direct learning from unannotated RGB-D data facilitated by an Iterative Annotation Resolver. Furthermore, a self-supervised pose refinement method based on joint loss is proposed to enhance the pose estimation accuracy. This refinement method employs a differentiable renderer to construct joint optimization constraints. The experimental results demonstrate that our approach achieves a level of pose estimation accuracy that closely rivals that of supervised methods. Fengjun Mu, Shixiang Sun, Rui Huang 0008, Chaobin Zou, Wenjiang Li, Huayi Zhan, Hong Cheng 0002 |
ICRA | 7 |
| 2024 | Target-point Attention Transformer: A novel trajectory predict network for end-to-end autonomous drivingabstractThe network of end-to-end automatic driving algorithms can be divided into perception network part and planning network part. Most of the research on the end-to-end automatic driving algorithm focuses on the part of the perception network, while the improvement of the planning network is less. However, the existing planning network can not effectively use the perceptual features, which may lead to traffic accidents. In this paper, we propose a Transformer-based trajectory prediction network for end-to-end autonomous driving without rules called Target-point Attention Transformer network (TAT). Leveraging the attention mechanism, our proposed model facilitates interaction between the predicted trajectory and perception features, along with target-points. Comparative evaluations with existing conditional imitation learning and GRU-based methods show the superior performance of our approach, particularly in reducing accident occurrences and improving route completion. Extensive assessments conducted in complex closed-loop driving scenarios within urban settings, utilizing the CARLA simulator, affirm the state-of-the-art proficiency of our proposed method. Yang Zhao 0024, Jingyu Du, Ruoyu Deng, Hong Cheng 0002 |
IV | 4 |
| 2024 | Spatio-temporal features for fast early warning of unplanned self-extubation in ICU
Yang Chen 0039, Ling Wang 0013, Guorong Wang, MingFang Xiang, Dekun Hu, Hong Cheng 0002 |
Eng. Appl. Artif. Intell. | 10 |
| 2024 | Event-triggered learning-based robust tracking control for robotic manipulators with uncertain dynamics and non-zero equilibrium
Chen Chen 0137, Zhinan Peng, Chaobin Zou, Rui Huang 0008, Kaibo Shi, Hong Cheng 0002 |
Expert Syst. Appl. | 6 |
| 2024 | SS-Pose: Self-Supervised 6-D Object Pose Representation Learning Without RenderingabstractObject pose estimation has extensive applications in various industrial scenarios. However, the heavy reliance on dense 6-D annotation and textured object models has become a significant obstacle to the widespread industrial application of 6-D object pose estimation methods. In this work, we presentSS-Pose, a self-supervised learning framework for estimating 6-D object poses without annotated 6-D data and textured model.SS-Poseproposes thecoordinate system datum reinitializerstage to dynamically establish a sequence-level pose representation datum, and thetemporal–spatial constraint resolvermodule to obtain the self-supervised learning target through interframe constraints. We introduce a one-shotcross-coordinate transformationthat establishes the relationship between the 6-D representation and the object poses, which can be further utilized in real-world tasks. We evaluated the proposedSS-Poseon the challenging YCB-Video dataset and texture-less T-LESS dataset. Our approach achieves competitive performance with significantly lower data dependency, making it suitable for visual perception in industrial applications. Fengjun Mu, Rui Huang 0008, Jingting Zhang, Chaobin Zou, Shixiang Sun, Huayi Zhan, Pengbo Zhao, Jing Qiu 0004, Hong Cheng 0002 |
IEEE Trans. Ind. Informatics | 10 |
| 2023 | PLFormer: Prompt Learning for Early Warning of Unplanned Extubation in ICUabstractPatients’ Unplanned Extubation (UEX) behaviors in ICU have adverse effects on their postoperative recovery. Therefore, it is necessary to design a early warning systems to detect UEX tendency. However, the fineness and rapidity of UEX behaviors, coupled with the complexity of the ICU environment, renders the utilization of RGB monitory videos for early warning extremely challenging. To address the aforementioned challenges, we propose a PLFormer to make early warning of UEX behaviors in ICU by using the prompt learning approach. Specifically, we provide click prompts to the Track Anything model (TAM) with the ability to segment and track patient regions, producing a mask sequence. Then we introduce the Prompt-Guided Adaptive Fusion (PAF) module, which utilizes mask prompts to guide the model’s attention towards the dynamically active region at both global and local levels. Subsequently, we incorporate the ST-Transformer to delve deeply into the spatial fine-grained representation and long-term temporal dependency properties of UEX behaviors. Experimental results demonstrate that our PLFormer achieves state-of-the-art performance on an ICU monitory dataset. Yang Chen 0039, Hong Cheng 0002, Ling Wang 0013 |
BIBM | 4 |
| 2023 | A Dual-Arm Participated Human-Robot Collaboration Method for Upper Limb Rehabilitation of Hemiplegic PatientsabstractUpper limb rehabilitation robots are mainly used as a physical therapy method to passively or actively train the affected side. However, they are rarely implemented in accordance with the occupational therapy theory, which is dedicated to improving the sensorimotor coordination of hemiplegic patients by considering both healthy and affected limbs. To realize the occupational therapy concept in robot-assisted upper limb rehabilitation, we propose a new human-robot collaboration framework for hemiplegic patients that integrates healthy/affected limbs and robot. The strategy aims at achieving patient-specific movement capabilities and improving the participation of the affected limb during rehabilitation. To accomplish this task, we have addressed two essential issues: accurate motion estimation of the healthy limb and the rehabilitation trajectory learning technique. The posture estimation is achieved by introducing the calibration model to reduce static and time dependent errors during the measurement. We also introduce a force term to the conventional imitation learning method to improve the adaptability in integrating the affected side in cooperation with the robot. Various experiments have been conducted to validate the feasibility and effectiveness of our proposed dual-arm collaboration strategy. Lufeng Chen, Jing Qiu 0004, Xuan Zou, Hong Cheng 0002 |
ICRA | 4 |
| 2023 | Emotional Voice Conversion with Semi-Supervised Generative Modeling
Huayi Zhan, Hong Cheng 0002, Ying Wu 0001 |
INTERSPEECH | 3 |
| 2023 | Optimal tracking control for motion constrained robot systems via event-sampled critic learning
Zhinan Peng, Hong Cheng 0002, Kaibo Shi, Chaobin Zou, Rui Huang 0008, Xiaoqing Li 0003, Bijoy K. Ghosh |
Expert Syst. Appl. | 2 |
| 2023 | Multi-view graph convolution network for the recognition of human action with spatial and temporal occlusion problems
Yang Chen 0039, Ling Wang 0013, Dekun Hu, Hong Cheng 0002 |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | Optimal H∞ tracking control of nonlinear systems with zero-equilibrium-free via novel adaptive critic designs
Zhinan Peng, Hanqi Ji, Chaobin Zou, Yiqun Kuang, Hong Cheng 0002, Kaibo Shi, Bijoy K. Ghosh |
Neural Networks | 5 |
| 2023 | MGL: Mutual Graph Learning for Camouflaged Object DetectionabstractCamouflaged object detection, which aims to detect/segment the object(s) that blend in with their surrounding, remains challenging for deep models due to the intrinsic similarities between foreground objects and background surroundings. Ideally, an effective model should be capable of finding valuable clues from the given scene and integrating them into a joint learning framework to co-enhance the representation. Inspired by this observation, we propose a novel Mutual Graph Learning (MGL) model by shifting the conventional perspective of mutual learning from regular grids to graph domain. Specifically, an image is decoupled by MGL into two task-specific feature maps - one for finding the rough location of the target and the other for capturing its accurate boundary details. Then, the mutual benefits can be fully exploited by reasoning their high-order relations through graphs recurrently. It should be noted that our method is different from most mutual learning models that model all between-task interactions with the use of a shared function. To increase information interactions, MGL is built with typed functions for dealing with different complementary relations. To overcome the accuracy loss caused by interpolation to higher resolution and the computational redundancy resulting from recurrent learning, the S-MGL is equipped with a multi-source attention contextual recovery module, called R-MGL_v2, which uses the pixel feature information iteratively. Experiments on challenging datasets, including CHAMELEON, CAMO, COD10K, and NC4K demonstrate the effectiveness of our MGL with superior performance to existing state-of-the-art methods. The code can be found at https://github.com/fanyang587/MGL. Qiang Zhai, Xin Li 0079, Fan Yang 0054, Zhicheng Jiao, Ping Luo 0002, Hong Cheng 0002, Zicheng Liu 0001 |
IEEE Trans. Image Process. | 6 |
| 2023 | Co-Communication Graph Convolutional Network for Multi-View Crowd CountingabstractWe study and address the multi-view crowd counting (MVCC) problem which poses more realistic challenges than single-view crowd counting for better facilitating crowd management/public safety systems. Its major challenge lies in how to fully distill and aggregate useful, complementary information among multiple camera views to create powerful ground-plane representations for wide-area crowd analysis. In this paper, we present a graph-based, multi-view learning model called Co-Communication Graph Convolutional Network (CoCo-GCN) to jointly investigate intra-view contextual dependencies and inter-view complementary relations. More specifically, CoCo-GCN builds a view-agnostic graph interaction space for each camera view to conduct efficient contextual reasoning, and extends the intra-view reasoning by using a novel Graph Communication Layer (GCL) to also take between-graph (cross-view), complementary information into account. Moreover, CoCo-GCN uses a new Co-Memory Layer (CoML) to jointly coarsen the graphs and close the ‘representational gap’ among them for further exploiting the compositional nature of graphs and learning more consistent representations. Finally, these jointly learned features of multiple views can be easily fused to create ground-plane representations for wide-area crowd counting. Experiments show that the proposed CoCo-GCN achieves state-of-the-art results on three MVCC datasets, i.e., PETS2009, DukeMTMC, and City Street, significantly improving the scene-level accuracy over previous models. Qiang Zhai, Fan Yang 0054, Xin Li 0079, Guosen Xie, Hong Cheng 0002, Zicheng Liu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Robust Scene Parsing by Mining Supportive Knowledge From DatasetabstractScene parsing, or semantic segmentation, aims at labeling all pixels in an image with the predefined categories of things and stuff. Learning a robust representation for each pixel is crucial for this task. Existing state-of-the-art (SOTA) algorithms employ deep neural networks to learn (discover) the representations needed for parsing from raw data. Nevertheless, these networks discover desired features or representations only from the given image (content), ignoring more generic knowledge contained in the dataset. To overcome this deficiency, we make the first attempt to explore the meaningful supportive knowledge, including general visual concepts (i.e., the generic representations for objects and stuff) and their relations from the whole dataset to enhance the underlying representations of a specific scene for better scene parsing. Specifically, we propose a novel supportive knowledge mining module (SKMM) and a knowledge augmentation operator (KAO), which can be easily plugged into modern scene parsing networks. By taking image-specific content and dataset-level supportive knowledge into full consideration, the resulting model, called knowledge augmented neural network (KANN), can better understand the given scene and provide greater representational power. Experiments are conducted on three challenging scene parsing and semantic segmentation datasets: Cityscapes, Pascal-Context, and ADE20K. The results show that our KANN is effective and achieves better results than all existing SOTA methods. Ao Luo, Fan Yang 0054, Xin Li 0079, Yuezun Li, Zhicheng Jiao, Hong Cheng 0002, Siwei Lyu |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Depth Estimation via Sparse Radar Prior and Driving Scene Semantics
Shuguang Li 0004, Kongjian Qin, Zhenxu Li, Yang Zhao 0024, Zhinan Peng, Hong Cheng 0002 |
ACCV (2) | 7 |
| 2022 | A Novel Multimodal Human-Exoskeleton Interface Based on EEG and sEMG Activity for Rehabilitation TrainingabstractDespite the advances in the field of human-robot interface (HRI) based on biological neural signal, the use of the sole electroencephalography (EEG) signal to help robotic exoskeleton predict the limb movement is currently no mature in rehabilitation training, due to its unreliability. Multimodal HRI represents a very recent solution to enhance the performance of single modal HRI. These HRI normally include the EEG signal with surface electromyography (sEMG) signal. However, their use for the lower limb movement prediction in hemiplegia is still limited, and the deep fusion feature of sEMG and EEG signal is ignored. This paper proposes a Dense co-attention mechanism-based Multimodal Enhance fusion Network (DMEFNet) for the lower limb movement prediction in hemiplegia. The DMEFNet can realize the mapping and deep fusion between the sEMG and EEG signal features and get a high accuracy movement prediction of the lower limbs. A sEMG and EEG data acquisition experiment and an incomplete asynchronous data collection paradigm are designed to verify the effectiveness of DMEFNet. The experimental results show that DMEFNet has a good movement prediction performance in both within-subject and cross-subject situations, reaching an accuracy of 82.96% and 88.44% respectively. Rui Huang 0008, Fengjun Mu, Zhinan Peng, Yizhe Qin, Hong Cheng 0002 |
ICRA | 8 |
| 2022 | Human-exoskeleton Cooperative Balance Strategy for a Human-powered Augmentation Lower ExoskeletonabstractLower Limb Exoskeletons (LLE) have received considerable interest in strength augmentation, rehabilitation, and walking assistance scenarios. For strength augmentation, LLE is expected to have the capability of reducing metabolic energy. However, the energy for adjusting Center of Gravity (CoG) is a main part of the total energy consumed during walking. This paper proposes a novel Human-exoskeleton Cooperative Balance (HCB) strategy which gives assistive torques balance ability and combine with the direction selected by the pilot to achieve balance walking of human-exoskeleton systems. In which, a Dynamic Torque Primitive Model (DTPM) is designed to plan a bionic assistive torque, and the balance parameters obtained by an Inverted Pendulum Model (IPM) is superimposed on it. Finally, the performance improved by the HCB strategy can break the limitation of traditional strategies and substantially increase the efficiency of assistance. We demonstrated the effectiveness of the proposed HCB strategy on the HUman-powered Augmentation Lower EXoskeleton (HUALEX) system. Experimental results indicate that the proposed HCB is more efficient than traditional strategies. Guangkui Song, Rui Huang 0008, Zhinan Peng, Jing Qiu 0004, Huayi Zhan, Hong Cheng 0002 |
IROS | 9 |
| 2022 | Attention-Based Deep Driving Model for Autonomous Vehicles with Surround-View CamerasabstractExperienced human drivers always make safe driving decisions by selectively observing the front, rear and side- view mirrors. Several end - to-end methods have been pro-posed to learn driving models with multi-view visual infor-mation. However, these benchmark methods lack semantic understanding of multi-view image contents, where human drivers usually reason these information for decision making with different visual region of interests. In this paper, we propose an attention-based deep learning method to learn a driving model with input of surround-view visual information and the route planner, in which a multi-view attention module is designed for obtaining region of interests from human drivers. We evaluate our model on the Drive360 dataset with comparison of benchmarking deep driving models. Results demonstrate that our model achieves a competitive accuracy in both steering angle and speed prediction than benchmarking methods. Code is available at https://githuh.com/jet-uestc/MVA-Net. Yang Zhao 0024, Rui Huang 0008, Boqi Li 0001, Ao Luo, Yaochen Li, Hong Cheng 0002 |
IROS | 7 |
| 2022 | EFRNet: Efficient Feature Reconstructing Network for Real-Time Scene ParsingabstractIn this paper, we introduce a light-weight and powerful convolutional neural network, termed asefficient feature reconstructing network(EFRNet), for real-time scene parsing. Our key idea is to decompose the process of learning high-resolution representations into two stages: i) bottom-up codebook/coding matrix learning and ii) top-down feature reconstructing. Specifically, the bottom-up process focuses on learningimage-specificcodewords (codebook) using deep-layer features and generating a coding matrix with the shallow-layer feature map. In the top-down process, the learned codebook and coding matrix are used to rebuild high-resolution features via a lightweightfeature reconstructing operator(FRO). In addition, our EFRNet is constructed on a new building block, named efficient adaptive abstraction (EAA) block, to further reduce the overall network parameters and achieve a significant speed up. Extensive experiments are conducted on challenging benchmarks, such as CamVid and Cityscapes. The results show that EFRNet demonstrates state-of-the-art performance with an optimal balance between accuracy and speed. Xin Li 0079, Fan Yang 0054, Ao Luo, Zhicheng Jiao, Hong Cheng 0002, Zicheng Liu 0001 |
IEEE Trans. Multim. | 5 |
| 2021 | Probabilistic Model Distillation for Semantic CorrespondenceabstractSemantic correspondence is a fundamental problem in computer vision, which aims at establishing dense correspondences across images depicting different instances under the same category. This task is challenging due to large intra-class variations and a severe lack of ground truth. A popular solution is to learn correspondences from synthetic data. However, because of the limited intra-class appearance and background variations within synthetically generated training data, the model’s capability for handling “real” image pairs using such strategy is intrinsically constrained. We address this problem with the use of a novel Probabilistic Model Distillation (PMD) approach which transfers knowledge learned by a probabilistic teacher model on synthetic data to a static student model with the use of unlabeled real image pairs. A probabilistic supervision reweighting (PSR) module together with a confidence-aware loss (CAL) is used to mine the useful knowledge and alleviate the impact of errors. Experimental results on a variety of benchmarks show that our PMD achieves state-of-the-art performance. To demonstrate the generalizability of our approach, we extend PMD to incorporate stronger supervision for better accuracy – the probabilistic teacher is trained with stronger key-point supervision. Again, we observe the superiority of our PMD. The extensive experiments verify that PMD is able to infer more reliable supervision signals from the probabilistic teacher for representation learning and largely alleviate the influence of errors in pseudo labels. Code is available at https://github.com/fanyang587/PMD. Xin Li 0079, Deng-Ping Fan, Fan Yang 0054, Ao Luo, Hong Cheng 0002, Zicheng Liu 0001 |
CVPR | 5 |
| 2021 | Mutual Graph Learning for Camouflaged Object DetectionabstractAutomatically detecting/segmenting object(s) that blend in with their surroundings is difficult for current models. A major challenge is that the intrinsic similarities between such foreground objects and background surroundings make the features extracted by deep model indistinguishable. To overcome this challenge, an ideal model should be able to seek valuable, extra clues from the given scene and incorporate them into a joint learning framework for representation co-enhancement. With this inspiration, we design a novel Mutual Graph Learning (MGL) model, which generalizes the idea of conventional mutual learning from regular grids to the graph domain. Specifically, MGL decouples an image into two task-specific feature maps — one for roughly locating the target and the other for accurately capturing its boundary details — and fully exploits the mutual benefits by recurrently reasoning their high-order relations through graphs. Importantly, in contrast to most mutual learning approaches that use a shared function to model all between-task interactions, MGL is equipped with typed functions for handling different complementary relations to maximize information interactions. Experiments on challenging datasets, including CHAMELEON, CAMO and COD10K, demonstrate the effectiveness of our MGL with superior performance to existing state-of-the-art methods. Code is available at https://github.com/fanyang587/MGL. Qiang Zhai, Xin Li 0079, Fan Yang 0054, Chenglizhao Chen, Hong Cheng 0002, Deng-Ping Fan |
CVPR | 5 |
| 2021 | Uncertainty-Guided Transformer Reasoning for Camouflaged Object DetectionabstractSpotting objects that are visually adapted to their surroundings is challenging for both humans and AI. Conventional generic / salient object detection techniques are suboptimal for this task because they tend to only discover easy and clear objects, while overlooking the difficult-to-detect ones with inherent uncertainties derived from indistinguishable textures. In this work, we contribute a novel approach using a probabilistic representational model in combination with transformers to explicitly reason under uncertainties, namely uncertainty-guided transformer reasoning (UGTR), for camouflaged object detection. The core idea is to first learn a conditional distribution over the backbone's output to obtain initial estimates and associated uncertainties, and then reason over these uncertain regions with attention mechanism to produce final predictions. Our approach combines the benefits of both Bayesian learning and Transformer-based reasoning, allowing the model to handle camouflaged object detection by leveraging both deterministic and probabilistic information. We empirically demonstrate that our proposed approach can achieve higher accuracy than existing state-of-the-art models on CHAMELEON, CAMO and COD10K datasets. Code is available at https://github.com/fanyang587/UGTR. Fan Yang 0054, Qiang Zhai, Xin Li 0079, Rui Huang 0008, Ao Luo, Hong Cheng 0002, Deng-Ping Fan |
ICCV | 6 |
| 2021 | Estimating the Center of Mass of Human-Exoskeleton Systems with Physically Coupled Serial ChainabstractEstimating the center of mass (CoM) is essential for both gait planning and controlling of lower limb exoskeletons. Different from CoM estimation in human and humanoid robots, a critical issue in human-exoskeleton systems pis how to describe the effect of physical human-exoskeleton interactions in estimating the CoM of lower limb exoskeletons. This paper presents a novel center of mass estimation method Physically Coupled Serial Chain (PCSC) for human-coupled lower limb exoskeleton systems. Different from traditional serial chain methods, the proposed PCSC involves physical human-exoskeleton models to describe physical interactions between the pilot and the lower limb exoskeleton. We demonstrated the effectiveness of proposed PCSC model in the AIDER lower limb exoskeleton system. Experimental results indicate that the proposed PCSC model is more accuracy than traditional serial chain methods. Rui Huang 0008, Zhinan Peng, Siying Guo, Chaobin Zou, Jing Qiu 0004, Hong Cheng 0002 |
IROS | 7 |
| 2021 | TemporalFusion: Temporal Motion Reasoning with Multi-Frame Fusion for 6D Object Pose Estimationabstract6D object pose estimation is an essential task in vision-based robotic grasping and manipulation. Prior works extract spatial features by fusing the RGB image and depth without considering the temporal motion information, limiting their performance in heavy occlusion robotic grasping scenarios. In this paper, we present an end-to-end model named TemporalFusion, which integrates the temporal motion information from RGB-D images for 6D object pose estimation. The core of proposed TemporalFusion model is to embed and fuse the temporal motion information from multi-frame RGB-D sequences, which could handle heavy occlusion in robotic grasping tasks. Furthermore, the proposed deep model can also obtain stable pose sequences, which is essential for real-time robotic grasping tasks. We evaluated the proposed method in the YCB-Video dataset, and experimental results show our model outperforms state-of-the-art approaches. Our code is available at https://github.com/mufengjun260/TemporalFusion21. Fengjun Mu, Rui Huang 0008, Ao Luo, Xin Li 0079, Jing Qiu 0004, Hong Cheng 0002 |
IROS | 6 |
| 2021 | Synergetic Gait Prediction for Stroke Rehabilitation with Varying Walking SpeedsabstractLower Limb Exoskeletons (LLEs) are promising in gait rehabilitation for stroke survivors. In gait training of post-stroke patients with LLEs, one of the main challenges is how to generate appropriate gait patterns from the sound leg to the paretic leg for different patients with varying walking speeds. In this paper, we proposed a Synergetic Gait Prediction (SGP) model for rehabilitation LLEs with post-stroke patients, which can generate adaptive synergetic gait patterns for different patients with varying walking speeds. The proposed SGP model is based on Sequence-to-Sequence (Seq2Seq) neural networks with temporal attention mechanisms. In the training procedure of the proposed SGP model, a gait database with collected gait patterns from healthy subjects is employed to learn the parameters of SGP model. The SGP model takes current joint angles from the sound leg and a segment of observed history joint angles from both legs as input and predicts the future joint angles for the paretic leg. We compared the effectiveness of the SGP model with the Long Short Term Memory (LSTM) model, experimental results indicate that SGP model can generate synergetic gait patterns for different subjects via varying walking speeds with less prediction error. Chaobin Zou, Rui Huang 0008, Zhinan Peng, Jing Qiu 0004, Hong Cheng 0002 |
IROS | 5 |
| 2021 | Learning continuous coupled multi-controller coefficients based on actor-critic algorithm for lower-limb exoskeleton
Guangkui Song, Rui Huang 0008, Hong Cheng 0002, Jing Qiu 0004, Qiming Cheng, Shuai Fan 0002 |
Sci. China Inf. Sci. | 3 |
| 2021 | Adaptive compensation for time-varying uncertainties in model-based control of lower-limb exoskeleton systems
Guangkui Song, Rui Huang 0008, Hong Cheng 0002, Jing Qiu 0004, Shuai Fan 0002 |
Sci. China Inf. Sci. | 3 |
| 2021 | The AIDER system and its clinical applications
Yilin Wang 0036, Hong Cheng 0002, Jing Qiu 0004, Anren Zhang, Hongchen He |
Sci. China Inf. Sci. | 2 |
| 2021 | EKENet: Efficient knowledge enhanced network for real-time scene parsing
Ao Luo, Fan Yang 0054, Xin Li 0079, Rui Huang 0008, Hong Cheng 0002 |
Pattern Recognit. | 5 |
| 2021 | Slope Gradient Adaptive Gait Planning for Walking Assistance Lower Limb ExoskeletonsabstractIn recent years, lower limb exoskeletons have gained considerable interest in applications of walking assistance for paraplegic patients. In daily lives, the exoskeleton should have the ability to help the patients to walk over different terrains. For sloped terrains, how to plan the stepping locations on slopes with different gradients and generate stable human-like gaits for patients is a critical issue. In this article, we proposed a slope gradient estimator (SGE) based on the sensor data fusion of the exoskeleton and combined SGE with the capture point theory and dynamic movement primitives (DMP) to construct an adaptive gait planning approach for slopes. After learning from demonstrated gaits sampled from healthy subjects, adaptive gait trajectories can be reproduced online to adapt to slopes with different gradients. The efficiency of the proposed approach was demonstrated on an exoskeleton system named AIDER. Experimental results indicate that the proposed approach can endow exoskeletons with the ability to generate appropriate gaits for different slopes. Note to Practitioners-For lower limb exoskeletons, it is a vital problem to plan the gait for sloped terrains. Considering different gradients among slopes, fixed predefined gait planning cannot cover all cases; thus, a slope gradient adaptive gait planning approach is necessary. The slope gradient estimator proposed in this article provides a possible slope gradient estimation method for exoskeletons or humanoid bipedal robots; it is easy to estimate the slope gradient only based on the local sensor data of the robot. The proposed dynamic gait generator provides lower limb exoskeletons and humanoid bipedal robots a possible adaptive gait planning framework and some flexibility for different slopes. The proposed approach may inspire more extended gait planning strategies for other terrains, such as stairs. Chaobin Zou, Rui Huang 0008, Jing Qiu 0004, Hong Cheng 0002 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2021 | Simultaneously Encoding Movement and sEMG-Based Stiffness for Robotic Skill LearningabstractTransferring human stiffness regulation strategies to robots enables them to effectively and efficiently acquire adaptive impedance control policies to deal with uncertainties during the accomplishment of physical contact tasks in an unstructured environment. In this article, we develop such a physical human-robot interaction system which allows robots to learn variable impedance skills from human demonstrations. Specifically, the biological signals, i.e., surface electromyography are utilized for the extraction of human arm stiffness features during the task demonstration. The estimated human arm stiffness is then mapped into a robot impedance controller. The dynamics of both movement and stiffness are simultaneously modeled by using a model combining the hidden semi-Markov model and the Gaussian mixture regression. More importantly, the correlation between the movement information and the stiffness information is encoded in a systematic manner. This approach enables capturing uncertainties over time and space and allows the robot to satisfy both position and stiffness requirements in a task with modulation of the impedance controller. The experimental study validated the proposed approach. Chao Zeng 0002, Chenguang Yang 0001, Hong Cheng 0002, Yanan Li 0001, Shi-Lu Dai |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Hybrid Graph Neural Networks for Crowd CountingabstractCrowd counting is an important yet challenging task due to the large scale and density variation. Recent investigations have shown that distilling rich relations among multi-scale features and exploiting useful information from the auxiliary task, i.e., localization, are vital for this task. Nevertheless, how to comprehensively leverage these relations within a unified network architecture is still a challenging problem. In this paper, we present a novel network structure called Hybrid Graph Neural Network (HyGnn) which targets to relieve the problem by interweaving the multi-scale features for crowd density as well as its auxiliary task (localization) together and performing joint reasoning over a graph. Specifically, HyGnn integrates a hybrid graph to jointly represent the task-specific feature maps of different scales as nodes, and two types of relations as edges: (i) multi-scale relations capturing the feature dependencies across scales and (ii) mutual beneficial relations building bridges for the cooperation between counting and localization. Thus, through message passing, HyGnn can capture and distill richer relations between nodes to obtain more powerful representations, providing robust and accurate results. Our HyGnn performs significantly well on four challenging datasets: ShanghaiTech Part A, ShanghaiTech Part B, UCF_CC_50 and UCF_QNRF, outperforming the state-of-the-art algorithms by a large margin. Ao Luo, Fan Yang 0054, Xin Li 0079, Dong Nie, Zhicheng Jiao, Shangchen Zhou, Hong Cheng 0002 |
AAAI | 7 |
| 2020 | Cascade Graph Neural Networks for RGB-D Salient Object Detection
Ao Luo, Xin Li 0079, Fan Yang 0054, Zhicheng Jiao, Hong Cheng 0002, Siwei Lyu |
ECCV (12) | 5 |
| 2020 | Data-Driven Reinforcement Learning for Walking Assistance Control of a Lower Limb Exoskeleton with Hemiplegic PatientsabstractLower limb exoskeleton (LLE) has received considerable interests in strength augmentation, rehabilitation and walking assistance scenarios. For walking assistance, the LLE is expected to have the capability of controlling the affected leg to track the unaffected leg’s motion naturally. An important issue in this scenario is that the exoskeleton system needs to deal with unpredictable disturbance from the patient, which requires the controller of exoskeleton system to have the ability to adapt to different wearers. This paper proposes a novel Data-Driven Reinforcement Learning (DDRL) control strategy to adapt different hemiplegic patients with unpredictable disturbances. In the proposed DDRL strategy, the interaction between two lower limbs of LLE and the legs of hemiplegic patient are modeled in the context of leader-follower framework. The walking assistance control problem is transformed into a optimal control problem. Then, a policy iteration (PI) algorithm is introduced to learn optimal controller. To achieve online adaptation control for different patients, based on PI algorithm, an Actor-Critic Neural Network (ACNN) technology of the reinforcement learning (RL) is employed in the proposed DDRL. We conduct experiments both on a simulation environment and a real LLE system. Experimental results demonstrate that the proposed control strategy has strong robustness against disturbances and adaptability to different pilots. Zhinan Peng, Rui Luo 0003, Rui Huang 0008, Jiangping Hu, Hong Cheng 0002, Bijoy K. Ghosh |
ICRA | 6 |
| 2020 | Webly-supervised learning for salient object detection
Ao Luo, Xin Li 0079, Fan Yang 0054, Zhicheng Jiao, Hong Cheng 0002 |
Pattern Recognit. | 5 |
| 2019 | Adaptive Gait Planning for Walking Assistance Lower Limb Exoskeletons in Slope ScenariosabstractLower-limb exoskeleton has gained considerable interests in walking assistance applications for paraplegic patients. In walking assistance of paraplegic patients, the exoskeleton should have the ability to help patients to walk over different terrains in the daily life, such as slope terrains. One critical issue is how to plan the stepping locations on slopes with different gradients, and generate stable and human-like gaits for patients. This paper proposed an adaptive gait planning approach which can generate gait trajectories adapt to slopes with different gradients for lower-limb walking assistance exoskeletons. We modeled the human-exoskeleton system as a 2D Linear Inverted Pendulum Model (2D-LIPM) with an external force in the two-dimensional sagittal plane, and proposed a Dynamic Gait Generator (DGG) based on an extension of the conventional Capture Point (CP) theory and Dynamic Movement Primitives (DMPs). The proposed approach can dynamically generate reference foot locations for each step on slopes, and human-like adaptive gait trajectories can be reproduced after the learning from demonstrated trajectories that sampled from level ground walking of normal healthy human. We demonstrated the efficiency of the proposed approach on both the Gazebo simulation platform and an exoskeleton named AIDER. Experimental results indicate that the proposed approach is able to provide the ability for exoskeletons to generate appropriate gaits adapt to slopes with different gradients. Chaobin Zou, Rui Huang 0008, Hong Cheng 0002, Jing Qiu 0004 |
ICRA | 3 |
| 2019 | End-to-End Driving Model for Steering Control of Autonomous Vehicles with Future Spatiotemporal FeaturesabstractEnd-to-end deep learning has gained considerable interests in autonomous driving vehicles in both academic and industrial fields, especially in decision making process. One critical issue in decision making process of autonomous driving vehicles is steering control. Researchers has already trained different artificial neural networks to predict steering angle with front-facing camera data stream. However, existing end-to-end methods only consider the spatiotemporal relation on a single layer and lack the ability of extracting future spatiotemporal information. In this paper, we propose an end-to-end driving model based on Convolutional Long Short-Term Memory (Conv-LSTM) neural network with a Multi-scale Spatiotemporal Integration (MSI) module, which aiming to encode the spatiotemporal information from different scales for steering angle prediction. Moreover, we employ future sequential information to enhance spatiotemporal features of the end-to-end driving model. We demonstrate the efficiency of proposed end-to-end driving model on the public Udacity dataset with comparison of some existing methods. Experimental results show that the proposed model has better performances than other existing methods, especially in some complex scenarios. Furthermore, we evaluate the proposed driving model on a real-time autonomous vehicle, and results show that the proposed driving model is able to predict the steering angle with high accuracy compared to skilled human driver. Ao Luo, Rui Huang 0008, Hong Cheng 0002, Yang Zhao 0024 |
IROS | 4 |
| 2019 | Robust Quadratic Programming for MDPs with uncertain observation noise
Jianmei Su, Hong Cheng 0002, Hongliang Guo 0001, Zhinan Peng |
Neurocomputing | 2 |
| 2019 | Learning Physical Human-Robot Interaction With Coupled Cooperative Primitives for a Lower ExoskeletonabstractHuman-powered lower exoskeletons have received considerable interests from both academia and industry over the past decades, and encountered increasing applications in human locomotion assistance and strength augmentation. One of the most important aspects in those applications is to achieve robust control of lower exoskeletons, which, in the first place, requires the proactive modeling of human movement trajectories through physical human-robot interaction (pHRI). As a powerful representative tool for motion trajectories, dynamic movement primitives (DMP) have been used to model human movement trajectories. However, canonical DMP only offers a general representation of human movement trajectory and may neglects the interactive term, therefore it cannot be directly applied to lower exoskeletons which need to track human joint trajectories online, because different pilots have different trajectories and even same pilot might change his/her motion during walking. This paper presents a novel coupled cooperative primitive (CCP) strategy, which aims at modeling the motion trajectories online. Besides maintaining canonical motion primitives, we model the interaction term between the pilot and exoskeletons through impedance models, and propose a reinforcement learning method based on policy improvement and path integrals (PI2) to learn the parameters online. Experimental results on both a single degree-of-freedom platform and a HUman-powered Augmentation Lower EXoskeleton (HUALEX) system demonstrate the advantages of our proposed CCP scheme. Rui Huang 0008, Hong Cheng 0002, Jing Qiu 0004, Jianwei Zhang 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2019 | Robust High-Order Manifold Constrained Sparse Principal Component Analysis for Image RepresentationabstractIn order to efficiently utilize the information in the data and eliminate the negative effects of outliers in the principal component analysis (PCA) method, in this paper, we propose a novel robust sparse PCA method based on maximum correntropy criterion (MCC) with high-order manifold constraints called the RHSPCA. Compared with the traditional PCA methods, the proposed RHSPCA has the following benefits: 1) the MCC regression term is more robust to outliers than the MSE-based regression term; 2) thanks to the high-order manifold constraints, the low-dimensional representations can preserve the local relations of the data and greatly improve the clustering and classification performance for image processing tasks; and 3) in order to further counteract the adverse effects of outliers, the MCC-based samples' mean is proposed to better centralize the data. We also propose a new solver based on the half-quadratic technique and accelerated block coordinate update strategy to solve the RHSPCA model. Extensive experimental results show that the proposed method can outperform the state-of-the-art robust PCA methods on a variety of image processing tasks, including reconstruction, clustering, and classification, on outliers contaminated datasets. Nan Zhou 0010, Hong Cheng 0002, Harry Qin, Yuanhua Du, Badong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Maximum Correntropy Criterion-Based Sparse Subspace Learning for Unsupervised Feature SelectionabstractHigh-dimensional data contain not only redundancy but also noises produced by the sensors. These noises are usually non-Gaussian distributed. The metrics based on Euclidean distance are not suitable for these situations in general. In order to select the useful features and combat the adverse effects of the noises simultaneously, a robust sparse subspace learning method in unsupervised scenario is proposed in this paper based on the maximum correntropy criterion that shows strong robustness against outliers. Furthermore, an iterative strategy based on half quadratic and an accelerated block coordinate update is proposed. The convergence analysis of the proposed method is also carried out to ensure the convergence to a reliable solution. Extensive experiments are conducted on real-world data sets to show that the new method can filter out the outliers and outperform several state-of-the-art unsupervised feature selection methods. Nan Zhou 0010, Yangyang Xu 0005, Hong Cheng 0002, Zejian Yuan, Badong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Multi-Scale Bidirectional FCN for Object Skeleton ExtractionabstractObject skeleton detection is a challenging problem with wide application. Recently, deep Convolutional Neural Networks (CNNs) have substantially improved the performance of the state-of-the-art in this task. However, most of the existing CNN-Based methods are based on a skip-layer structure where low-level and high-level features are combined and learned so as to gather multi-level contextual information. As shallow features are too messy and lack semantic knowledge, they may cause errors and inaccuracy. Therefore, we propose a novel network architecture, Multi-Scale Bidirectional Fully Convolutional Network (MSB-FCN), to better capture and consolidate multi-scale high-level context information for object skeleton detection. Our network uses only deep features to build multi-scale feature representations, and employs a bidirectional structure to collect contextual knowledge. Hence the proposed MSB-FCN has the ability to learn the semantic-level information from different sub-regions. Furthermore, we introduce dense connections into the bidirectional structure of our MSB-FCN to ensure that the learning process at each scale can directly encode information from all other scales. Extensive experiments on various commonly used benchmarks demonstrate that the proposed MSB-FCN has achieved significant improvements over the state-of-the-art algorithms. Fan Yang 0054, Xin Li 0079, Hong Cheng 0002, Yuxiao Guo 0001, Leiting Chen |
AAAI | 3 |
| 2018 | Contour Knowledge Transfer for Salient Object Detection
Xin Li 0079, Fan Yang 0054, Hong Cheng 0002, Dinggang Shen |
ECCV (15) | 3 |
| 2018 | Learning-based Walking Assistance Control Strategy for a Lower Limb Exoskeleton with Hemiplegia PatientsabstractLower exoskeleton has gained considerable interests in walking assistance applications for both paraplegia and hemiplegia patients. In walking assistance of hemiplegia patients, the exoskeleton should have the ability to control the affected leg to follow the unaffected leg's motion naturally. One critical issue of walking assistance for hemiplegia patients is how to adapt the controller of both lower limbs with different patients. This paper presents a novel learning-based walking assistance control strategy for lower exoskeleton with hemiplegia patients. In the proposed control strategy, we modeled the control system of lower exoskeleton with hemiplegia patient as a Leader-Follower Multi-Agent System (LF -MAS). In order to adapt different patients with different conditions, reinforcement learning framework is utilized to adapt controllers online. In reinforcement learning framework with LF-MAS, we employed a Policy Iteration Adaptive Dynamic Programming (PI-ADP) algorithm, which aims to achieve better tracking control performance for lower exoskeleton with hemiplegia patient. We demonstrate the efficiency of proposed learning-based walking assistance control strategy in an exoskeleton system with healthy subjects who simulate hemiplegia patients. Experimental results indicate that the proposed control strategy can adapt different pilots with good tracking performance. Rui Huang 0008, Zhinan Peng, Hong Cheng 0002, Jiangping Hu, Jing Qiu 0004, Chaobin Zou |
IROS | 3 |
| 2018 | Chaos Theory in Urban Traffic Flow: Is Crowd Sensed Data Driving the Macro-traffic Behavior to Oscillation or EquilibriumabstractStability theory tells us that a dynamic system will eventually converge to its stable state, in which the system's overall energy is at its minimum. On the other hand, chaos theory states that small perturbations of the system are able to drive itself from previously-stable state to another state. This phenomenon has been observed in many fields like cosmetol- ogy, physics, biology and chemistry. Our research question is whether chaos theory also applies to the transportation domain. Specifically, when we are given imperfect or delayed crowd- sensed data, will we observe the cyclic/oscillatory transition between different traffic states? This paper aims at investigating this chaotic phenomenon (oscillatory traffic behavior in this paper) on urban transportation with imperfect or delayed crowd-sensed information and delivering recommendations for crowdsensing-based traffic applications to avoid the undesirable oscillations. Huanghuang Liang, Lu Yang 0002, Jiacheng Wei, Hong Cheng 0002 |
Intelligent Vehicles Symposium | 4 |
| 2018 | Hierarchical learning control with physical human-exoskeleton interaction
Rui Huang 0008, Hong Cheng 0002, Hongliang Guo 0001, XiChuan Lin, Jianwei Zhang 0001 |
Inf. Sci. | 2 |
| 2018 | A set-to-set nearest neighbor approach for robust and efficient face recognition with image sets
Ling Wang 0013, Hong Cheng 0002, Zicheng Liu 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Robust spatial-temporal Bayesian view synthesis for video stitching with occlusion handling
Jianmei Su, Hong Cheng 0002, Lu Yang 0002, Ao Luo |
Mach. Vis. Appl. | 2 |
| 2018 | Structured dynamic time warping for continuous hand trajectory gesture recognition
Jingren Tang, Hong Cheng 0002, Yang Zhao 0024, Hongliang Guo 0001 |
Pattern Recognit. | 2 |
| 2018 | Deep Background Modeling Using Fully Convolutional NetworkabstractBackground modeling plays an important role for video surveillance, object tracking, and object counting. In this paper, we propose a novel deep background modeling approach utilizing fully convolutional network. In the network block constructing the deep background model, three atrous convolution branches with different dilate are used to extract spatial information from different neighborhoods of pixels, which breaks the limitation that extracting spatial information of the pixel from fixed pixel neighborhood. Furthermore, we sample multiple frames from original sequential images with increasing interval, in order to capture more temporal information and reduce the computation. Compared with classical background modeling approaches, our approach outperforms the state-of-art approaches both in indoor and outdoor scenes. Lu Yang 0002, Jing Li 0013, Yuansheng Luo, Yang Zhao 0024, Hong Cheng 0002 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2017 | Object-Aware Dense Semantic Correspondence
Fan Yang 0054, Xin Li 0079, Hong Cheng 0002, Leiting Chen |
CVPR | 3 |
| 2017 | Multi-Scale Cascade Network for Salient Object DetectionabstractIn this paper we present a novel network architecture, called Multi-Scale Cascade Network (MSC-Net), to identify the most visually conspicuous objects in an image. Our network consists of several stages (sub-networks) for handling saliency detection across different scales. All these sub-networks form a cascade structure (in a coarse-to-fine manner) where the same underlying convolutional feature representations are fully shared. Compared with existing CNN-based saliency models, the MSC-Net can naturally enable the learning process in the finer cascade stages to encode more global contextual information while progressively incorporating the saliency prior knowledge obtained from coarser stages and thus lead to better detection accuracy. We also design a novel refinement module to further filter out errors by considering the intermediate feedback information. Our MSC-Net is highly integrated, end-to-end trainable, and very powerful. The proposed method achieves state-of-the-art performance on five widely-used salient object detection benchmarks, outperforming existing methods and also maintaining high efficiency. Code and pre-trained models are available at https://github.com/lixin666/MSC-NET. Xin Li 0079, Fan Yang 0054, Hong Cheng 0002, Junyu Chen 0002, Yuxiao Guo 0001, Leiting Chen |
ACM Multimedia | 3 |
| 2017 | Gazing point dependent eye gaze estimation
Hong Cheng 0002, Yanli Ji, Lu Yang 0002, Yang Zhao 0024, Jie Yang 0001 |
Pattern Recognit. | 1 |
| 2017 | Sparse Bayesian dictionary learning with a Gaussian hierarchical model
Linxiao Yang, Jun Fang 0001, Hong Cheng 0002, Hongbin Li 0001 |
Signal Process. | 3 |
| 2017 | Haptic Identification by ELM-Controlled Uncertain ManipulatorabstractThis paper presents an extreme learning machine (ELM)-based control scheme for uncertain robot manipulators to perform haptic identification. ELM is used to compensate for the unknown nonlinearity in the manipulator dynamics. The ELM enhanced controller ensures that the closed-loop controlled manipulator follows a specified reference model, in which the reference point as well as the feedforward force is adjusted after each trial for haptic identification of geometry and stiffness of an unknown object. A neural learning law is designed to ensure finite-time convergence of the neural weight learning, such that exact matching with the reference model can be achieved after the initial iteration. The usefulness of the proposed method is tested and demonstrated by extensive simulation studies. Chenguang Yang 0001, Kunxia Huang, Hong Cheng 0002, Yanan Li 0001, Chun-Yi Su |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2016 | Hierarchical Interactive Learning for a HUman-Powered Augmentation Lower EXoskeletonabstractLearning by demonstration methods have gained considerable interest in human-coupled robot control. It aims at modeling the goal motion trajectories through human demonstration. However, in lower exoskeleton control, the physical human-robot interaction is changing from pilot to pilot or even for one pilot in different walking patterns. This characteristic requires that the exoskeletons should have the ability to learn and adapt the motion trajectories as well as controllers online. This paper presents a novel Hierarchical Interactive Learning (HIL) strategy which reduces the complexity of the exoskeleton sensory system and is able to handle varying interaction dynamics. The proposed HIL strategy is composed of two learning hierarchies, namely, high-level motion learning and low-level controller learning. The Dynamic Movement Primitives (DMPs) combined with Locally Weighted Regression (LWR) are employed to model and learn the motion trajectories, while reinforcement learning (RL) is used to learn the model-based controller. We demonstrate the efficiency of proposed HIL strategy on a single degree-of-freedom (DOF) platform as well as a HUman-powered Augmentation Lower EXoskeleton (HUALEX) system. Experimental results indicate that the proposed HIL strategy is able to handle the varying interaction dynamics with less interaction force between the pilot and the exoskeleton when compared to traditional model-based control algorithms. Rui Huang 0008, Hong Cheng 0002, Hongliang Guo 0001, XiChuan Lin |
ICRA | 2 |
| 2016 | Learning Cooperative Primitives with physical Human-Robot Interaction for a HUman-powered Lower EXoskeletonabstractHuman-powered lower exoskeletons have gained considerable interests from both academia and industry over the past few decades, and thus have seen increasing applications in areas of human locomotion assistance and strength augmentation. One of the most important aspects in those applications is to achieve robust control of lower exoskeletons, which, in the first place, requires the proactive modeling of human movement trajectories through physical Human-Robot Interaction (pHRI). As a powerful representation tool for motion trajectories, Dynamic Movement Primitive (DMP) has been used extensively to model human movement trajectories. However, canonical DMPs only offers a general offline representation of human movement trajectory and neglects the real-time interaction term, therefore it cannot be directly applied to lower exoskeletons which need to model human motion trajectories online since different pilots have different trajectories and even one pilot might change his/her intended trajectory during walking. This paper presents a novel Coupled Cooperative Primitives (CCPs) scheme, which models the motion trajectories online. Besides maintaining canonical motion primitives, we also model the interaction term between the pilot and exoskeletons through impedance models and apply a reinforcement learning method based on Policy Improvement and Path Integrals (PI2) to learn the parameters online. Experimental results on both a single Degree-Of-Freedom (DOF) platform and a HUman-powered Augmentation Lower EXoskeleton (HUALEX) system demonstrate the advantages of our proposed CCP scheme. Rui Huang 0008, Hong Cheng 0002, Hongliang Guo 0001, XiChuan Lin, Fuchun Sun 0001 |
IROS | 2 |
| 2016 | An image-to-class dynamic time warping approach for both 3D static and trajectory hand gesture recognition
Hong Cheng 0002, Zhongjun Dai, Zicheng Liu 0001, Yang Zhao 0024 |
Pattern Recognit. | 1 |
| 2016 | Global and local structure preserving sparse subspace learning: An iterative approach to unsupervised feature selection
Nan Zhou 0010, Yangyang Xu 0005, Hong Cheng 0002, Jun Fang 0001, Witold Pedrycz |
Pattern Recognit. | 3 |
| 2016 | Sparsity-Induced Similarity Measure and Its ApplicationsabstractThe structures of feature vectors-based semisupervised/supervised learning have gained considerable interest in recent years due to their effectiveness for better object modeling and classification. In many machine learning and computer vision tasks, a critical issue is the similarity between two feature vectors. In this paper, we present a novel technique to measure similarities among feature vectors by decomposing each feature vector as an ℓ1sparse linear combination of the rest of the feature vectors. The main idea is that the coefficients in such sparse decomposition reflect the features' neighborhood structure, thus providing better similarity measures among the decomposed feature vector and the rest of the feature vectors. The proposed approach is applied to label propagation and action recognition, and is evaluated on several commonly used datasets. The experimental results show that the proposed sparsity-induced similarity measure significantly improves the performance of both label propagation and action recognition. Hong Cheng 0002, Zicheng Liu 0001, Lei Hou 0019, Jie Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Survey on 3D Hand Gesture RecognitionabstractThree-dimensional hand gesture recognition has attracted increasing research interests in computer vision, pattern recognition, and human-computer interaction. The emerging depth sensors greatly inspired various hand gesture recognition approaches and applications, which were severely limited in the 2D domain with conventional cameras. This paper presents a survey of some recent works on hand gesture recognition using 3D depth sensors. We first review the commercial depth sensors and public data sets that are widely used in this field. Then, we review the state-of-the-art research for 3D hand gesture recognition in four aspects: 1) 3D hand modeling; 2) static hand gesture recognition; 3) hand trajectory gesture recognition; and 4) continuous hand gesture recognition. While the emphasis is on 3D hand gesture recognition approaches, the related applications and typical systems are also briefly summarized for practitioners. Hong Cheng 0002, Lu Yang 0002, Zicheng Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Pixel-to-Model Distance for Robust Background ReconstructionabstractBackground information is crucial for many video surveillance applications such as object detection and scene understanding. In this paper, we present a novel pixel-to-model (P2M) paradigm for background modeling and restoration in surveillance scenes. In particular, the proposed approach models the background with a set of context features for each pixel, which are compressively sensed from local patches. We determine whether a pixel belongs to the background according to the minimum P2M distance, which measures the similarity between the pixel and its background model in the space of compressive local descriptors. The pixel feature descriptors of the background model are properly updated with respect to the minimum P2M distance. Meanwhile, the neighboring background model will be renewed according to the maximum P2M distance to handle ghost holes. The P2M distance plays an important role of background reliability in the 3-D spatial-temporal domain of surveillance videos, leading to the robust background model and recovered background videos. We applied the proposed P2M distance for foreground detection and background restoration on synthetic and real-world surveillance videos. Experimental results show that the proposed P2M approach outperforms the state-of-the-art approaches both in indoor and outdoor surveillance scenes. Lu Yang 0002, Hong Cheng 0002, Jianan Su, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Fast Online Upper Body Pose Estimation from VideoabstractEstimation of human body poses from video is an important problem in computer vision with many applications. Most existing methods for video pose estimation are offline in nature, where all frames in the video are used in the process to estimate the body pose in each frame. In this work, we describe a fast online video upper body pose estimation method (CDBN-MODEC) that is based on a conditional dynamic Bayesian network model, which predicts upper body pose in a frame without using information from future frames. Our method combines fast single image based pose estimation methods with the temporal correlation of poses between frames. We collect a new high frame rate upper body pose dataset that better reflects practical scenarios calling for fast online video pose estimation. When evaluated on this dataset and the VideoPose2 benchmark dataset, CDBN-MODEC achieves improvements in both performance and running efficiency over several state-of-art online video pose estimation methods. Ming-Ching Chang, Honggang Qi, Xin Wang 0045, Hong Cheng 0002, Siwei Lyu |
BMVC | 4 |
| 2015 | Robust Kernel Dictionary Learning Using a Whole Sequence Convergent Algorithm
Huaping Liu 0001, Hong Cheng 0002, Fuchun Sun 0001 |
IJCAI | 3 |
| 2015 | Interactive learning for sensitivity factors of a human-powered augmentation lower exoskeletonabstractSensitivity Amplification Control (SAC) algorithm was first proposed in the augmentation applications of Berkeley Lower Extremity Exoskeleton (BLEEX). The SAC algorithm is widely used in human augmentation applications since it just need the information from the exoskeleton robot, so that the complexity of exoskeleton system can be reduced greatly. However, the SAC algorithm has two main drawbacks: 1) requiring accurate dynamic models of the exoskeleton, 2) can not manage the variation of interaction dynamics from different walking speed. This paper presents a novel developed learning control strategy based on SAC algorithm. In the proposed Adaptive Sensitivity Amplification Control (ASAC) strategy, the reinforcement learning method is utilized to learn the sensitivity factors online for the sake of handling the variation of interaction dynamics. We demonstrate the control efficiency of ASAC on an one degree-of-freedom (DOF) platform with swing movements first, and then extend it into a HUman-powered Augmentation Lower EXoskeleton (HUALEX). The experimental results show that the proposed ASAC strategy can handle the changing interaction dynamics with less interaction force between the pilot and the exoskeleton as compared with traditional SAC algorithm. Rui Huang 0008, Hong Cheng 0002, Huu-Toan Tran, XiChuan Lin |
IROS | 2 |
| 2015 | Learning contrastive feature distribution model for interaction recognition
Yanli Ji, Hong Cheng 0002, Haoxin Li |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Learning sparse and scale-free networksabstractGaussian networks study undirected interactions between random variables, through the estimation of the precision matrices. Recently, it has been demonstrated that some of the important networks display features similar to scale-free graphs. There have been few works on the learning of the sparse Gaussian graphical models aiming to preserve properties of networks which are believed to be scale-free or have dominating hubs. We prefer to name both networks as ‘scale-free’ networks for simplicity in this paper. We propose a new log-likelihood formulation, which promotes the sparseness of the precision matrix and features of scale-free graphical topology. We used the alternating direction method of multipliers (ADMM) form, which is used for the convex optimization, to solve the general L1regularized loss optimization. Our proposed method exhibits better estimation performance on various data sets and various number of samples, N. Also, the proposed method and some of the state of the arts methods are tested under various penalty constants to validate the robustness. Melih S. Aslan, Xue-wen Chen 0001, Hong Cheng 0002 |
DSAA | 3 |
| 2014 | Pseudo labels for imbalanced multi-label learningabstractThe classification with instances which can be tagged with any of the 2Lpossible subsets from the predefined L labels is called multi-label classification. Multi-label classification is commonly applied in domains, such as multimedia, text, web and biological data analysis. The main challenge lying in multi-label classification is the dilemma of optimising label correlations over exponentially large label powerset and the ignorance of label correlations using binary relevance strategy (1-vs-all heuristic). The classification with label powerset usually encounters with highly skewed data distribution, called imbalanced problem. While binary relevance strategy reduces the problem from exponential to linear, it totally neglects the label correlations. In this artical, we propose a novel strategy of introducing Balanced Pseudo-Labels (BPL) which build more robust classifiers for imbalanced multi-label classification, which embeds imbalanced data in the problems innately. By incorporating the new balanced labels we aim to increase the average distances among the distinct label vectors. In this way, we also code the label correlation implicitly in the algorithm. Another advantage of the proposed method is that it can combined with any classifier and it is proportional to linear label transformation. In the experiment, we choose five multi-label benchmark data sets and compare our algorithm with the most state-of-art algorithms. Our algorithm outperforms them in standard multi-label evaluation in most scenarios. Wenrong Zeng, Xue-wen Chen 0001, Hong Cheng 0002 |
DSAA | 3 |
| 2014 | A windowed dynamic time warping approach for 3D continuous hand gesture recognitionabstractDetecting the beginning and end of a specific gesture from an infinite trajectory gesture sequence has gained considerable interests in the past several years. Traditional begin-end dynamic time warping approach for gesture recognition could provide multiple different gesture labels for one trajectory segment. This paper presents a Windowed Dynamic Time Warping (WDTW) approach for 3D continuous hand trajectory gesture recognition. The main contribution is that we introduce a parameterized searching window in the cost matrix of traditional DTW approach to detect the beginning and end of the specific gesture from an infinite trajectory gesture sequence. By doing so, we formulate continuous gesture recognition into online parameter estimation of the searching window. Moreover, the proposed gesture recognition can handle the multilabel issue. We evaluate the proposed windowed dynamic time warping approach in our gesture dataset. The experimental results show that the proposed WDTW can significantly improve the begin-end gesture recognition performance. Hong Cheng 0002, Xue-wen Chen 0001 |
ICME | 1 |
| 2014 | Pixel-to-Model background modeling in crowded scenesabstractBackground modeling is an important step for many video surveillance applications such as object detection and scene understanding. In this paper, we present a novel Pixel-to-Model (P2M) paradigm for background modeling in crowded scenes. In particular, the proposed method models the background with a set of context features for each pixel, which are compressively sensed from local patches. We determine whether a pixel belongs to the background according to the minimum P2M distance, which measures the similarity between the pixel and its background model in the space of compressive local descriptors. Moreover, the background updating utilizes minimum and maximum P2M distances to update the pixel feature descriptors in local and neighboring background models, respectively. We evaluate the proposed approach with foreground detection tasks on real crowded surveillance videos. Experiments results show that the proposed P2M approach outperforms the state-of-the-art methods both in indoor and outdoor crowded scenes. Lu Yang 0002, Hong Cheng 0002, Jianan Su, Xue-wen Chen 0001 |
ICME | 2 |
| 2014 | The relationship between physical human-exoskeleton interaction and dynamic factors: using a learning approach for control applications
Huu-Toan Tran, Hong Cheng 0002, XiChuan Lin, Mien-Ka Duong, Rui Huang 0008 |
Sci. China Inf. Sci. | 2 |
| 2014 | A robust elastic net approach for feature learning
Ling Wang 0013, Hong Cheng 0002, Zicheng Liu 0001, Ce Zhu |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Real world activity summary for senior home monitoring
Hong Cheng 0002, Zicheng Liu 0001, Yang Zhao 0024, Guo Ye, Xinghai Sun |
Multim. Tools Appl. | 1 |
| 2014 | Kernelized pyramid nearest-neighbor search for object categorization
Hong Cheng 0002, Rongchao Yu, Zicheng Liu 0001, Lu Yang 0002, Xue-wen Chen 0001 |
Mach. Vis. Appl. | 1 |
| 2013 | Image-to-Class Dynamic Time Warping for 3D hand gesture recognitionabstract3D Human Computer Interaction (HCI) becomes more and more popular thanks to the emergence of commercial depth cameras. Moreover, hand gestures provide a natural and attractive alternative to cumbersome interface devices for HCI. In this paper, we present an Image-to-Class Dynamic Time Warping (I2C-DTW) approach for 3D hand gesture recognition. Themain idea is that we divide the time-series curve of a 3D hand gesture into various finger combinations, called `fingerlets', which can either be learned or be set manually to represent each gesture and to capture inter-class variations. Furthermore, the I2C-DTW approach searches for the minimal path to warp two fingerlets, which are fromone test image and the specific class, respectively. Then the gesture recognition is to use the ensemble of multiple image-to-class DTW distance of fingerlets to obtain better performance. The proposed approach is evaluated on two 3D hand gesture datasets and the experiment results show that the proposed I2C-DTW approach significantly improves the recognizing performance. Hong Cheng 0002, Zhongjun Dai, Zicheng Liu 0001 |
ICME | 1 |
| 2013 | Sparse representation and learning in visual recognition: Theory and applications
Hong Cheng 0002, Zicheng Liu 0001, Lu Yang 0002, Xue-wen Chen 0001 |
Signal Process. | 1 |
| 2013 | Recovering shape and motion by a dynamic system for low-rank matrix approximation in L 1 norm
Yiguang Liu, Liping Cao, Yi-Fei Pu, Hong Cheng 0002 |
Vis. Comput. | 5 |
| 2012 | A Pyramid Nearest Neighbor Search Kernel for object categorization
Hong Cheng 0002, Rongchao Yu, Zicheng Liu 0001, Yiguang Liu |
ICPR | 1 |
| 2012 | Low-rank matrix decomposition in L1-norm by dynamic systems
Yiguang Liu, Yi-Fei Pu, Hong Cheng 0002 |
Image Vis. Comput. | 5 |
| 2012 | Multi-support-region image descriptors and its application to street landmark localization
Hong Cheng 0002, Zicheng Liu 0001, Jie Yang 0001 |
Mach. Vis. Appl. | 1 |
| 2011 | Real world activity summary for senior home monitoringabstractFrom a senior person's daily activities, one can tell a lot about the health condition of the senior person. Thus we believe that senior home activity analysis will play an important role in the health care of senior people. Toward this goal, we propose a senior home activity summary system. One challenging problem in such a real world application is that senior's activities are usually accompanied by nurse's walking. It is impractical to predefine and label all the potential activities of all the potential visitors. To address this problem, we propose a novel feature filtering technique to reduce or eliminate the effects of the interest points that belong to other people. To evaluate the proposed activity summary system, we have collected a senior home activity dataset (SAR), and performed activity recognition for eating and walking classes. The experimental results show that the proposed system provides quite accurate activity summaries for a real world application scenario. Hong Cheng 0002, Zicheng Liu 0001, Yang Zhao 0024, Guo Ye |
ICME | 1 |
| 2010 | Learning feature transforms for object detection from panoramic imagesabstractWe present a novel technique to detect objects from panoramic images using existing object detectors trained from perspective images. By leveraging existing object detectors, we save the cost of training a new detector which requires tedious and time consuming training data collection and labeling. The core of our technique is learning a feature transform which is represented by Gaussian Process Regression (GPR). Feature vectors computed directly from panoramic images are transformed into new feature vectors in such a way that the existing classifier has much better detection rate on the transformed feature vectors. Our feature transform has the interesting property that it not only corrects for the geometric distortions resulted from panoramic imaging process, but also corrects for the pose mismatches between the objects on the panoramic images and those on the training images. Our experiments show that we are able to successfully apply an existing car detector trained on perspective images to panoramic images which have both geometric distortions and larger pose variations. Hong Cheng 0002, Zicheng Liu 0001, Jie Yang 0001 |
ICME | 1 |
| 2009 | Sparsity induced similarity measure for label propagationabstractGraph-based semi-supervised learning has gained considerable interests in the past several years thanks to its effectiveness in combining labeled and unlabeled data through label propagation for better object modeling and classification. A critical issue in constructing a graph is the weight assignment where the weight of an edge specifies the similarity between two data points. In this paper, we present a novel technique to measure the similarities among data points by decomposing each data point as an L1sparse linear combination of the rest of the data points. The main idea is that the coefficients in such a sparse decomposition reflect the point's neighborhood structure thus providing better similarity measures among the decomposed data point and the rest of the data points. The proposed approach is evaluated on four commonly-used data sets and the experimental results show that the proposed Sparsity Induced Similarity (SIS) measure significantly improves label propagation performance. As an application of the SIS-based label propagation, we show that the SIS measure can be used to improve the Bag-of-Words approach for scene classification. Hong Cheng 0002, Zicheng Liu 0001, Jie Yang 0001 |
ICCV | 1 |
| 2008 | A deformable local image descriptorabstractThis paper presents a novel local image descriptor that is robust to general image deformations. A limitation with traditional image descriptors is that they use a single support region for each interest point. For general image deformations, the amount of deformation for each location varies and is unpredictable such that it is difficult to choose the best scale of the support region. To overcome this difficulty, we propose to use multiple support regions of different sizes surrounding an interest point. A feature vector is computed for each support region, and the concatenation of these feature vectors forms the descriptor for this interest point. Furthermore, we propose a new similarity measure model, Local-to-Global Similarity (LGS) model, for point matching that takes advantage of the multi-size support regions. Each support region acts as a ‘weak’ classifier and the weights of these classifiers are learned in an unsupervised manner. The proposed approach is evaluated on a number of images with real and synthetic deformations. The experiment results show that our method outperforms existing techniques under different deformations. Hong Cheng 0002, Zicheng Liu 0001, Nanning Zheng 0001, Jie Yang 0001 |
CVPR | 1 |
| 2008 | Layered object categorizationabstractIn this paper, we propose a novel framework of object categorization, namely layered object categorization, which takes advantage of hierarchical category information and performs object categorization at different levels. The proposed hierarchical structure of object categories is built bottom-up and top-down simultaneously accordingly to cognitive rules. First, part-based models are learnt to evaluate structure similarities at the basic level and objects are divided into basic categories. Then the decision cues for object categorization at different layers are optimally selected. Prior knowledge about inter-category relationships is utilized to infer objectspsila higher inclusive concept labels, while the most discriminative visual details of each category at the lower specific levels are selected automatically. We evaluate the proposed method with a hierarchical database and show promising results. The layered object categorization provides an efficient way for dynamically adapting the object categorization results to different applications. Lei Yang 0063, Jie Yang 0001, Nanning Zheng 0001, Hong Cheng 0002 |
ICPR | 4 |
| 2007 | Enhancing a Driver's Situation Awareness using a Global View MapabstractThis paper proposes a novel method to enhance a driver's situation awareness by dynamically providing a global view of surroundings for the driver. The surroundings of a vehicle are captured by an omni-directional vision system mounted on the top of the vehicle. The video stream from the camera is processed to detect nearby vehicles. Positions of these detected objects are overlaid on a global view of a local map (e.g., an aerial imagery or satellite imagery map). We establish the relationship between the omni-directional vision system and the global view map. The global view map dynamically provides a realistic perspective view of the driving environment. This map can be projected onto an HUD on the windshield. By looking at the display, a driver can have a global picture of the situation and potentially produce a good driving strategy. We illustrate the proposed method by dynamically mapping a video stream onto Google Earth map. Hong Cheng 0002, Zicheng Liu 0001, Nanning Zheng 0001, Jie Yang 0001 |
ICME | 1 |
| 2007 | Interactive Road Situation Analysis for Driver Assistance and Safety Warning Systems: Framework and AlgorithmsabstractRoad situation analysis in Interactive Intelligent Driver-Assistance and Safety Warning (I2DASW) systems involves estimation and prediction of the position and size of various on-road obstacles. Real-time processing, given incomplete and uncertain information, is a challenge for current object detection and tracking technologies. This paper proposed a development framework and novel algorithms for road situation analysis based on driving action behavior, where the safety situation is analyzed by simulating real driving action behaviors. First, we review recent development and trends in road situation analysis to provide perspective for the related research. Second, we introduce a road situation analysis framework, where onboard sensors provide information about drivers, traffic environment, and vehicles. Finally, on the basis of the previous frameworks, we proposed multiple-obstacle detection and tracking algorithms using multiple sensors including radar, lidar, and a camera, where a decentralized track-to-track fusion approach is introduced to fuse these sensors. In order to reduce the effect of obstacle shape and appearance, we cluster lidar data and then classify obstacles into two categories: static and moving objects. Future collisions are assessed by computation of local tracks of moving obstacles using extended Kalman filter, maximum likelihood estimation to fuse distributed local tracks into global tracks, and finally, computation of future collision distribution from the global tracks. Our experimental results show that our approach is efficient for road situation evaluation and prediction Hong Cheng 0002, Nanning Zheng 0001, Xuetao Zhang 0001, Junjie Qin, Huub van de Wetering |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2006 | Vanishing Point and Gabor Feature Based Multi-resolution On-Road Vehicle Detection
Hong Cheng 0002, Nanning Zheng 0001, Huub van de Wetering |
ISNN (2) | 1 |
| 2005 | Neural Network Based Online Feature Selection for Vehicle Tracking
Nanning Zheng 0001, Hong Cheng 0002 |
ISNN (2) | 3 |
| 2004 | Principal Component Analysis Neural Network Based Probabilistic Tracking of Unpaved Road
Nanning Zheng 0001, Hong Cheng 0002 |
ISNN (1) | 4 |
| 2004 | Springrobot: a prototype autonomous vehicle and its algorithms for lane detectionabstractThis work presents the current status of the Springrobot autonomous vehicle project, whose main objective is to develop a safety-warning and driver-assistance system and an automatic pilot for rural and urban traffic environments. This system uses a high precise digital map and a combination of various sensors. The architecture and strategy for the system are briefly described and the details of lane-marking detection algorithms are presented. The R and G channels of the color image are used to form graylevel images. The size of the resulting gray image is reduced and the Sobel operator with a very low threshold is used to get a grayscale edge image. In the adaptive randomized Hough transform, pixels of the gray-edge image are sampled randomly according to their weights corresponding to their gradient magnitudes. The three-dimensional (3-D) parametric space of the curve is reduced to the two-dimensional (2-D) and the one-dimensional (1-D) space. The paired parameters in two dimensions are estimated by gradient directions and the last parameter in one dimension is used to verify the estimated parameters by histogram. The parameters are determined coarsely and quantization accuracy is increased relatively by a multiresolution strategy. Experimental results in different road scene and a comparison with other methods have proven the validity of the proposed method. Nanning Zheng 0001, Hong Cheng 0002 |
IEEE Trans. Intell. Transp. Syst. | 3 |