Jason Gu

dblp:60/955 · also Jason Jianjun Gu, Jianjun Gu 0001 · DBLP profile ↗
← Back
57ranked-venue papers
4as first author
27since 2021 · last 2026
0000-0002-7626-1077ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 2 first-author · 13 since 2021Systems, architecture and hardware · 23 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Guest Editorial: Artificial Intelligence Generated Content (AIGC) for Industrial Manufacturing
Huaping Liu 0001, Weiwei Wan, Jason Gu, Valeria Villani, Giulia Pedrielli, Yiannis Aloimonos
IEEE Trans Autom. Sci. Eng.3
2025 SX-Stitch: An Efficient VMS-UNet Based Framework for Intraoperative Scoliosis X-Ray Image Stitching
abstract
In scoliosis surgery, the limited field of view of the C-arm Xray machine restricts the surgeons’ holistic analysis of spinal structures. This paper presents an end-to-end efficient and robust intraoperative X-ray image stitching method for scoliosis surgery, named SX-Stitch. The method is divided into two stages: segmentation and stitching. In the segmentation stage, we propose a medical image segmentation model named Vision Mamba of Spine-UNet (VMS-UNet), which utilizes the state space Mamba to capture long-distance contextual information while maintaining linear computational complexity. Meanwhile, the proposed model incorporates the SimAM attention mechanism, significantly improving the segmentation performance. In the stitching stage, we simplify the alignment process between images to minimize a registration energy function. The total energy function is then optimized to order unordered images and a hybrid energy function is introduced to optimize the best seam, effectively eliminating parallax artifacts. On the clinical dataset, Sx-Stitch demonstrates superiority over SOTA schemes both qualitatively and quantitatively.
Heting Gao, Mingde He, Jinqian Liang, Jason Gu
ICASSP5
2025 Surface Roughness Estimation for Terrain Perception
abstract
Ground terrain perception has become the primary visual task for the robust navigation of intelligent systems in unstructured outdoor environments. However, complex ter-rain poses a significant challenge to vision-based perception. This work introduces a novel estimation task using RGB images to facilitate low-cost terrain perception in extracting surface roughness information. The proposed task presents both semantic-aware and edge-aware roughness descriptors at the pixel level instead of a single value for a given image. To promote the research on the proposed novel terrain roughness estimation task, we introduce a multimodal synthetic dataset for terrain perception in outdoor scenes, containing multiple terrain categories, diverse viewpoints, different lighting and weather conditions, as well as semantic and roughness annotations. Additionally, inspired by computer graphics, we introduce TRENet, a roughness estimation architecture to model the intrinsic correlation of depth-normal-roughness. We also perform ablation studies on the effect of each component and diverse types of inputs. Extensive evaluations and comparisons demonstrate that our method can effectively predict pixel-wise terrain surface roughness with high accuracy.
Minxiang Ye, Jason Gu, Senwei Xiang, Lingyu Kong, Anhuan Xie
ICRA3
2025 Referring Expression Comprehension in semi-structured human-robot interaction
Tianlei Jin, Qiwei Meng, Qiulan Huang, Fangtai Guo, Shu Kong, Wei Song 0008, Jiakai Zhu, Jason Gu
Expert Syst. Appl.9
2025 Survey of federated learning in intrusion detection
Hao Zhang 0078, Junwei Ye, Wei Huang 0037, Ximeng Liu, Jason Gu
J. Parallel Distributed Comput.5
2025 A Lighter and Faster One-Stage Algorithm for Object Detection in Remote Sensing Images
abstract
Remote sensing images processing and analysis face significant challenges due to varying object scales and complex backgrounds. Existing detection algorithms often suffer from high computational complexity and suboptimal performance. A lightweight algorithm SCC-YOLO was proposed for remote sensing objects detection. It incorporates three key innovations: (1) Slimneck-V feature fusion architecture to enhance multi-scale adaptability while reducing computational load. (2) Cross Stage Partial with Context Anchor Attention (C2CAA) module to improve feature representation of key object regions. (3) Cross Stage Partial with Ghost (CSPGhost) module that optimizes feature extraction efficiency. The algorithm is validated on DOTA and RSOD datasets. Experimental results demonstrate that, compared to baseline algorithms, SCC-YOLO reduces model parameters by 15.3% and computational complexity by 26%. On the DOTA dataset, detection accuracy and inference speed are improved by 3.9% and 6.5%, respectively.
Yifeng Du, Yang Shengqi, Junmei Guo, Dehao Dong, Jason Gu, Chaoqun Wang 0009, Lida Liu
IEEE Geosci. Remote. Sens. Lett.6
2025 GPD: Learning Geometric Primitive Deformation for Unseen Object Pose Estimation
abstract
Witnessing the rapid progress and development in instance-level object pose estimation, increasing attention has shifted to the more challenging problem for unseen objects, which is in great demand for various robotic applications. In this paper, we propose the GPD, a novel framework for unseen object pose estimation, including both category-level and cross-category objects. The key innovation of the GPD model is the effective utilization of geometric primitives in target reconstruction and pose estimation, as it can generalize the learned primitive deformation across intra-class and inter-class instances. Additionally, we also design an advanced scheme for representative object feature extraction, including attention-aware excitation, multi-scale fusion, and semantic feature encoding. Extensive evaluations validate the effectiveness of individual innovation modules and the overall superior performance of the GPD. It not only achieves the SOTA results on category-level benchmarks CAMERA25 and REAL275, but also demonstrates impressive generalization ability across novel objects on the GraspNet-1Billion dataset. Furthermore, we deploy the trained GPD model for vision-guided robotic grasping experiments in simulation and real-world settings, again exhibiting its outstanding robustness and practicability in robotic manipulations. Note to Practitioners—This paper is motivated by the problem of unseen object pose estimation and robotic manipulation in unstructured environments. For intelligent robots expected to interact with their surroundings, rather than just passively perceiving them like surveillance cameras, 6Dof pose estimation is a critical capability. However, existing approaches generally face two key challenges. On the one hand, robots are likely to encounter unseen objects in real-world applications. Without the availability of prior models or specific training data for these unseen objects, instance-level and category-level methods may become ineffective or even fail to work. On the other hand, the error tolerance of precise tabletop robotic manipulation is very tight, and the varying lighting conditions and background noise impose higher robustness requirements on pose estimation algorithms. To address these difficulties, we propose a novel network that learns geometric primitive deformation for pose estimation. This model is less dependent on object prior information, thereby enhancing the generalization ability. Additionally, by incorporating cross-modal excitation and multi-scale fusion during feature extraction, our model can capture representative appearance and geometric information of objects for accurate pose estimation. Extensive experimental results on benchmark datasets quantitatively validate the superior performance of our approach. We also demonstrate its effectiveness in robotic applications through unseen object grasping experiments on Kinova and Franka Emika robot platforms. In the future, we plan to explore primitive combination schemes for compound object representation, enabling pose estimation for more complex-shaped objects.
Qiwei Meng, Jason Gu, Yun-Hui Liu 0001
IEEE Trans Autom. Sci. Eng.2
2025 Learning 6-DoF Fine-Grained Grasp Detection Based on Part Affordance Grounding
abstract
Robotic grasping is a fundamental ability for a robot to interact with the environment. Current methods focus on how to obtain a stable and reliable grasping pose in object level, while little work has been studied on part (shape)-wise grasping which is related to fine-grained grasping and robotic affordance. Parts can be seen as atomic elements to compose an object, which contains rich semantic knowledge and a strong correlation with affordance. However, lacking a large part-wise 3D robotic dataset limits the development of part representation learning and downstream applications. In this paper, we propose a new large Language-guided SHape grAsPing datasEt (named LangSHAPE) to promote 3D part-level affordance and grasping ability learning. From the perspective of robotic cognition, we design a two-stage fine-grained robotic grasping framework (named LangPartGPD), including a novel 3D part language grounding model and a partaware grasp pose detection model, in which explicit language input from human or large language models (LLMs) could guide a robot to generate part-level 6-DoF grasping pose with textual explanation. Our method combines the advantages of humanrobot collaboration and LLMs’ planning ability using explicit language as a symbolic intermediate. To evaluate the effectiveness of our proposed method, we perform 3D part grounding and fine-grained grasp detection experiments on both simulation and physical robot settings, following language instructions across different degrees of textual complexity. Results show our method achieves competitive performance in 3D geometry fine-grained grounding, object affordance inference, and 3D part-aware grasping tasks. Our dataset and code are available on our project website https://sites.google.com/view/lang-shape.
Yaoxian Song, Penglei Sun, Piaopiao Jin, Yu Zheng 0001, Zhixu Li, Xiaowen Chu 0001, Yue Zhang 0004, Tiefeng Li, Jason Gu
IEEE Trans Autom. Sci. Eng.10
2025 TransCLIP: Transferring Vision-Language Models for Efficient Video Action Recognition
abstract
Transferring contrastive vision-language pretrained models, such as contrastive language–image pretraining (CLIP), to video recognition task has attracted much attention. Recent studies in this area utilized prompt learning within either text or vision branches, or employed an end-to-end CLIP fine-tuning approach. However, these methods do not fully leverage the learning potential of two branches and may compromise zero-shot generalization. In this work, we present a multimodal framework TransCLIP, aiming to adapt vision–language models by integrating both adapter and prompt tuning techniques for the vision and text encoders. Specifically, we incorporate learnable prompt tokens into each transformer encoder layer’s input of vision and text branches, and integrate lightweight adapters into the key and value matrices of the multi-head self-attention modules, enhancing the model’s capability to capture more related video-specific features. To effectively leverage temporal information in videos, we implement a temporal difference attention module (TemDiff attn) that explicitly computes differences between adjacent frame embeddings and conducts difference-level attention to encode motion-related temporal dependency in videos. In addition, a coarse-and-fine contrastive leaning strategy is employed to better align the video and text branches, enhancing the learning capability of the whole framework. Across different evaluation settings, our model consistently outperforms previous State-of-the-Art methods on several video action recognition benchmarks.
Wen Wang 0012, Yanzhou Su, Jason Gu
IEEE Trans. Ind. Informatics3
2025 Network intrusion detection based on feature fusion of attack dimension
Xiaolong Sun, Zhengyao Gu, Hao Zhang 0078, Jason Gu, Chen Dong 0002, Junwei Ye
J. Supercomput.4
2024 Leveraging the efficiency of multi-task robot manipulation via task-evoked planner and reinforcement learning
abstract
Multi-task learning has expanded the boundaries of robotic manipulation, enabling the execution of increasingly complex tasks. However, policies learned through reinforcement learning exhibit limited generalization and narrow distributions, which restrict their effectiveness in multi-task training. Addressing the challenge of obtaining policies with generalization and stability represents a non-trivial problem. To tackle this issue, we propose a planning-guided reinforcement learning method. It leverages a task-evoked planner(TEP) and a reinforcement learning approach with planner’s guidance. TEP utilizes reusable samples as the source, with the aim of learning reachability information across different task scenarios. Then in reinforcement learning, TEP assesses and guides the Actor towards better outputs and smoothly enhances the performance in multi-task benchmarks. We evaluate this approach within the Meta-World framework and compare it with prior works in terms of learning efficiency and effectiveness. Depending on experimental results, our method has more efficiency, higher success rates, and demonstrates more realistic behavior.
Haofu Qian, Jiatao Zhang, Jason Gu, Wei Song 0008, Shiqiang Zhu
ICRA5
2024 Aligning Knowledge Graph with Visual Perception for Object-goal Navigation
abstract
Object-goal navigation is a challenging task that requires guiding an agent to specific objects based on first-person visual observations. The ability of agent to comprehend its surroundings plays a crucial role in achieving successful object finding. However, existing knowledge-graph-based navigators often rely on discrete categorical one-hot vectors and vote counting strategy to construct graph representation of the scenes, which results in misalignment with visual images. To provide more accurate and coherent scene descriptions and address this misalignment issue, we propose the Aligning Knowledge Graph with Visual Perception (AKGVP) method for object-goal navigation. Technically, our approach introduces continuous modeling of the hierarchical scene architecture and leverages visual-language pre-training to align natural language description with visual perception. The integration of a continuous knowledge graph architecture and multimodal feature alignment empowers the navigator with a remarkable zero-shot navigation capability. We extensively evaluate our method using the AI2-THOR simulator and conduct a series of experiments to demonstrate the effectiveness and efficiency of our navigator.
Nuo Xu 0006, Wen Wang 0017, Zheyuan Lin, Wei Song 0008, Chunlong Zhang, Jason Gu, Chao Li 0028
ICRA8
2024 FLTRNN: Faithful Long-Horizon Task Planning for Robotics with Large Language Models
abstract
Recent planning methods based on Large Language Models typically employ the In-Context Learning paradigm. Complex long-horizon planning tasks require more context(including instructions and demonstrations) to guarantee that the generated plan can be executed correctly. However, in such conditions, LLMs may overlook(unfaithful) the rules in the given context, resulting in the generated plans being invalid or even leading to dangerous actions. In this paper, we investigate the faithfulness of LLMs for complex long-horizon tasks. Inspired by human intelligence, we introduce a novel framework named FLTRNN. FLTRNN employs a language-based RNN structure to integrate task decomposition and memory management into LLM planning inference, which could effectively improve the faithfulness of LLMs and make the planner more reliable. We conducted experiments in VirtualHome household tasks. Results show that our model significantly improves faithfulness and success rates for complex long-horizon tasks. Website at https://tannl.github.io/FLTRNN.github.io/
Jiatao Zhang, Lanling Tang, Qiwei Meng, Haofu Qian, Wei Song 0008, Shiqiang Zhu, Jason Gu
ICRA9
2024 Alzheimer's disease diagnosis from single and multimodal data using machine and deep learning models: Achievements and future directions
Ahmed El-Azab, Changmiao Wang, Mohammed Abdelaziz, Jason Gu, Juan Manuel Górriz, Yudong Zhang 0001, Chunqi Chang
Expert Syst. Appl.5
2024 Prototype learning based generic multiple object tracking via point-to-box supervision
Wenxi Liu, Qi Li 0038, Yinhua She, Yuanlong Yu 0001, Jia Pan 0001, Jason Gu
Pattern Recognit.7
2024 Proximal Policy Optimization With Time-Varying Muscle Synergy for the Control of an Upper Limb Musculoskeletal System
abstract
Because of their unique adaptability, flexibility, and robustness, musculoskeletal robotic systems are regarded potentially as next-generation robots. However, motion learning and generation of such a robotic system are still challenging. This paper presents a neuromuscular control method, namely, TMS-PPO, based on time-varying muscle synergy (TMS) and proximal policy optimization (PPO). The electromyogram (EMG) activation signals of actual human motions are decomposed to obtain TMSs based on the temporal properties of the TMS. The weights of networks are trained to generate the scale and phase coefficients through the PPO. The coefficients modulate the TMSs to generate appropriate activation patterns to optimize motion learning of the musculoskeletal system. To verify the effectiveness of the proposed method, the TMSs are extracted from human upper limb muscle activation signals, and we compare TMS-PPO with PPO in the motion learning and generation process of an upper limb musculoskeletal system. The results show that TMS-PPO can complete the control tasks because the average errors of the joints are less than 0.05 rad. In the meantime, TMSs are used as motion primitives of the musculoskeletal system to simulate the process of the human CNS controlling muscles. It shows that TMS-PPO reduces the energy consumption and improves the learning rate significantly compared with the PPO. The learning episodes reduce from$10^{4}$to$10^{3}$, which indicates that TMS-PPO has a stronger learning ability and better physiological explanation.Note to Practitioners—Due to the superiorities of the musculoskeletal system, humanoid robots that imitate human driven mechanisms are vigorously carried out worldwide. Taking advantages of human-like characteristics, the musculoskeletal robot provides new opportunities to understand and validate the human mechanisms of muscle control and motion learning, to compare the performance of the robot to that of humans as well as work in real world, e.g., human interactive robots, amusement robots and medical training robots in the future. However, strong redundancy, coupling, and nonlinearity of the system also raises many challenges for the investigation of the control problem. Inspired by how the human CNS controls a musculoskeletal system and realize motion generalization, a novel muscle-synergies-based neuromuscular control that combines time-varying muscle synergy (TMS) and Proximal Policy Optimization (PPO), namely, TMS-PPO is proposed in this paper. The learning efficiency of PPO and the physiological interpretation of the control process are improved during the motion learning and generation processes of the musculoskeletal system. Preliminary simulation experiments suggest that this method is feasible in terms of control accuracy and efficiency. Moreover, the performance of the TMS-PPO is comparable to the PPO without significant improvement. To solve this problem, in future work, we will introduce the cerebellar model into the control method which plays the role of adjusting and correcting the motions of the limbs to achieve accurate and stable control in the actions process of humans.
Yaru Chen 0002, Yongxuan Wang, Jason Gu
IEEE Trans Autom. Sci. Eng.6
2024 Whole-Body Inverse Kinematics and Operation-Oriented Motion Planning for Robot Mobile Manipulation
abstract
High DoF mobile manipulation of robots is a nonlinear, nonchain redundant problem. In this article, we focus on two subissues of robot mobile manipulation: whole-body inverse kinematics (whole-body IK) and operation-oriented motion planning (OOMP). Whole-body IK solves the robot arm joint configuration and the mobile base position configuration according to the target pose. OOMP generates a feasible trajectory from the current pose to the target pose. The trajectory can avoid obstacles and touch operated objects. We introduce neural network optimization (NNO) methods with two variations to solve whole-body IK and OOMP, respectively. For whole-body IK, we design a fully connected network (FCN) to predict ten DoF of position and joint configurations based on the target pose. We use these ten DoF configurations to derive the predicted pose for online optimization. For OOMP, we design a GRU-based network to generate trajectories based on the initial and goal states. We mainly adopt sphere masks to modify the point cloud properties of the target object dynamically. During optimization, the trajectory keeps away from point clouds but approaches sphere masks. Finally, we conduct extensive experiments both on a Franka Panda robot and a mobile dual-arm robot. The results demonstrate the superior performance of our NNO method on whole body IK and OOMP, and implement mobile manipulation in different environments successfully.
Tianlei Jin, Jiakai Zhu, Shiqiang Zhu, Zaixing He, Shuyou Zhang 0001, Wei Song 0008, Jason Gu
IEEE Trans. Ind. Informatics8
2023 CHAN: Cross-Modal Hybrid Attention Network for Temporal Language Grounding in Videos
abstract
The goal of temporal language grounding (TLG) task is to temporally localize the most semantically matched video segment with respect to a given sentence query in an untrimmed video. How to effectively incorporate the cross-modal interactions between video and language is the key to improve grounding performance. Previous approaches focus on learning correlations by computing the attention matrix between each frame-word pair, while ignoring the global semantics conditioned on one modality for better associating the complex video contents and sentence query of the target modality. In this paper, we propose a novel Cross-modal Hybrid Attention Network, which integrates two parallel attention fusion modules to exploit the semantics of each modality and interactions in cross modalities. One is Intra-Modal Attention Fusion, which utilizes gated self-attention to capture the frame-by-frame and word-by-word relations conditioned on the other modality. The other is Inter-Modal Attention Fusion, which utilizes query and key features derived from different modalities to calculate the co-attention weights and further promote inter-modal fusion. Experimental results show that our CHAN significantly outperforms several existing state-of-the-arts on three challenging datasets (ActivityNet Captions, Charades-STA and TACOS), demonstrating the effectiveness of our proposed method.
Wen Wang 0017, Minhong Wan, Jason Gu
ICME5
2023 KGNet: Knowledge-Guided Networks for Category-Level 6D Object Pose and Size Estimation
abstract
Despite the giant leap made in object 6D pose estimation and robotic grasping under structured scenarios, most approaches depend heavily on the exact CAD models of target objects beforehand, thereby limiting their wide applications. To address this, we propose a novel knowledge-guided network - KGNet to estimate the pose and size of category-level unseen objects. This network includes three primary innovations: knowledge-guided categorical model generation, pointwise deformation probability matrix and synergetic RGBD feature fusion, with the former two leveraging categorical object knowledge for unseen object reconstruction and the latter one facilitating pose-sensitive feature extraction. Exten-sive experiments on CAMERA25 and REAL275 verify their effectiveness, and KGNet achieves the SOTA performance on these two acknowledged benchmarks. Additionally, a real-world robotic grasping experiment is conducted, and its results further qualitatively prove the practicability and robustness of KGNet.
Qiwei Meng, Jason Gu, Shiqiang Zhu, Jianfeng Liao, Tianlei Jin, Fangtai Guo, Wen Wang 0017, Wei Song 0008
ICRA2
2023 RFFCE: Residual Feature Fusion and Confidence Evaluation Network for 6DoF Pose Estimation
abstract
In this paper, we propose a novel RGBD-based object 6DoF pose estimation network - RFFCE. It is a two-stage method that firstly leverages deep neural networks for feature extraction and object points matching, and then the geometric principles are utilized for final pose computation. Our approach consists of three primary innovations: residual feature fusion for representative RGBD feature extraction; confidence evaluation and confidence-based paired points offsets regression for self-evaluation and self-optimization respectively. Their effectiveness is verified through an ablation study, and our RFFCE achieves the SOTA performance on LineMOD, Occlusion-LineMOD and YCB-Video datasets. Additionally, we also conduct a real-world object grasping experiment for visualization and qualitative evaluation of the RFFCE.
Qiwei Meng, Shanshan Ji, Shiqiang Zhu, Tianlei Jin, Jason Gu, Wei Song 0008
ICRA6
2023 Event-Triggered Fixed-Time Practical Tracking Control for Flexible-Joint Robot
abstract
This article studies the adaptive fuzzy event-triggered fixed-time practical tracking control problem for flexible-joint robot system. Since the nonlinearities of the system are difficult to obtain, fuzzy logic systems are utilized. Second-order command filters are used to avoid the “explosion of complexity” problem. Moreover, a novel compensation system is proposed. The new error compensation system cannot only compensate for the error of the filter band, but also make the error converge in fixed time. By using backstepping technique, the virtual control laws and the adaptive law are designed. Notice that compared to the reporting achievements, our proposed virtual control laws are second-order derivable by using the novel switch function, which avoids “singularity hindrance” problem. To reduce communication pressure, the event-triggered controller is designed and Zeno behavior is avoided. Our proposed control strategy ensures that the tracking error can be arbitrarily small in fixed time and all variables of the closed-loop system remain bounded. Finally, simulation results are given to show the effectiveness of our control strategy.
Yingkang Xie, Qian Ma 0001, Jason Gu, Guopeng Zhou
IEEE Trans. Fuzzy Syst.3
2023 A network anomaly detection algorithm based on semi-supervised learning and adaptive multiclass balancing
Hao Zhang 0078, Zude Xiao, Jason Gu
J. Supercomput.3
2022 BCOT: A Markerless High-Precision 3D Object Tracking Benchmark
abstract
Template-based 3D object tracking still lacks a high-precision benchmark of real scenes due to the difficulty of annotating the accurate 3D poses of real moving video objects without using markers. In this paper, we present a multi-view approach to estimate the accurate 3D poses of real moving objects, and then use binocular data to construct a new benchmark for monocular textureless 3D object tracking. The proposed method requires no markers, and the cameras only need to be synchronous, relatively fixed as cross-view and calibrated. Based on our object-centered model, we jointly optimize the object pose by minimizing shape reprojection constraints in all views, which greatly improves the accuracy compared with the single-view approach, and is even more accurate than the depth-based method. Our new benchmark dataset contains 20 textureless objects, 22 scenes, 404 video sequences and 126K images captured in real scenes. The annotation error is guaranteed to be less than 2mm, according to both theoretical analysis and validation experiments. We reevaluate the state-of-the-art 3D object tracking methods with our dataset, reporting their performance ranking in real scenes. Our BCOT benchmark and code can be found at https://ar3dv.github.io/BCOT-Benchmark/.
Bin Wang 0035, Shiqiang Zhu, Xin Cao 0010, Fan Zhong 0001, Wenxuan Chen, Jason Gu, Xueying Qin
CVPR8
2021 Exploiting Probabilistic Siamese Visual Tracking with a Conditional Variational Autoencoder
abstract
Visual tracking is a fundamental capability for robots tasked with humans and environment interaction. However, state-of-the-art visual tracking methods are still prone to failures and are imprecise when applied to challenging stereos, and their results are generally confidence agonistic. These methods depend on an embedded deep learning model to provide deterministic features or regression maps. A deterministic output with low confidence can result in disastrous consequences and lacks evidence needed for subsequent operations. Moreover, training data ambiguities or noise in the observations (so-called data uncertainty) can also lead to inherent uncertainty. In this paper, we focus on exploiting probabilistic Siamese visual tracking with a conditional variational autoencoder (CVAE). First, we build a bridge between the Siamese architecture and the CVAE and propose a novel Bayesian visual tracking method. Second, the proposed method generates a complete probability distribution that enables the production of multiple plausible tracking outputs. Third, CVAE conditioned by ground truth data encodes a low-dimensional latent space and conducts noise-injection training to prevent overfitting. Our proposed tracking method outperformed the state-of-the-art trackers on the VOT2016, VOT2018 and TColor-128 datasets.
Wenhui Huang 0002, Jason Gu, Peiyong Duan, Sujuan Hou, Yuanjie Zheng
ICRA2
2021 A Capturability-based Control Framework for the Underactuated Bipedal Walking *
abstract
This work considers the control of underactuated bipedal walking, and a novel capturability-based control framework is presented. Compared with traditional approaches, the presented control method does not rely on the use of the Poincaré map, which may take significant computational cost. Firstly, a new definition of stable walking is presented, and a novel foot-placement based control method is proposed. Then, a controller design method is presented based on this control method. For the controller design, the foot placement adjustment is achieved by updating the virtual constraints using a heuristic method, and an improved virtual constraint control method is proposed to enforce the virtual constraints. Finally, the effectiveness of the presented control framework is illustrated on a five-link underactuated planar biped by numerical simulations.
Haihui Yuan, Sumian Song, Ruilong Du, Shiqiang Zhu, Jason Gu, Mingguo Zhao, Jianxin Pang
ICRA5
2021 From Edge to Keypoint: An End-to-End Framework For Indoor Layout Estimation
abstract
The task of spatial layout estimation of monocular image is to segment an RGB image of indoor scenes with semantic surface labels (i.e., ceiling, floor, front wall, left wall, and right wall). Most recent methods have to produce layout hypotheses based on the estimated edge map or semantic labels, and then rank the layout hypotheses. In this paper, we present an end-to-end framework that can directly output the layout type and keypoint coordinates (defined in the LSUN challenge). The proposed method takes advantage of transfer learning via learning on the fake samples, i.e., plenty of artificial {type, keypoints, edge map} triplets are generated to learn the mapping from edge maps to keypoint coordinates. Generative adversarial network (GAN) is implemented in this work for domain adaptation of the edge maps. Experimental results show that the proposed method can achieve state-of-the-art layout estimation performance on benchmark datasets.
Weidong Zhang 0005, Qian Zhang 0076, Wei Zhang 0021, Jason Gu, Yibin Li 0001
IEEE Trans. Multim.4
2021 Robust Exact Predictive Scheme for Output-Feedback Control of Input-Delay Systems With Unmatched Sinusoidal Disturbances
abstract
Predictor-based control has been widely used to compensate input delays. However, the conventional predictors such as Smith predictor, usually have poor robustness with respect to system disturbances. In this article, a novel robust predictive scheme, that can achieve exact state prediction in finite-time, is proposed for linear time-invariant (LTI) systems subject to input delay and unmatched sinusoidal disturbances. By combining the super-twisting technique with the proposed predictive scheme, we also propose a new predictor-based output-feedback super-twisting controller which can compensate arbitrary long input delay and eliminate the effect of unmatched disturbances from the corresponding state completely. The simulation comparison shows the effectiveness of the theoretical results.
Shengyuan Xu 0001, Jason Gu, Zhengqiang Zhang
IEEE Trans. Syst. Man Cybern. Syst.3
2020 Long Range Underwater Localization and Navigation using Gravity-Based Measurements
abstract
This paper reports on work to assess the feasibility of gravity-based long range underwater navigation and localization. As a first step, this is explored in simulations with RAO-Blackwellized particle filter simultaneous localization and mapping (SLAM). When implemented on an autonomous underwater vehicle it can operate submerged for extended periods without the use of an active sensor, thus widening the variety of AUV missions. Additionally, this work applies information theory to navigate through a region such that the SLAM data association, and thus the localization, performance is improved. The results also indicate that characteristic values for a region can be used as a SLAM metric for the region. Combining the characteristic value with information theory techniques improves the localization performance at extended ranges and is a first step towards long range underwater localization using gravimeters. Future work will optimize the particle filter, explore more sophisticated loop closures as well as hardware-in-the loop tests.
P. Pasnani, Mae Seto, Jason Gu
SMC3
2020 End-to-end multitask Siamese network with residual hierarchical attention for real-time object tracking
Wenhui Huang 0002, Jason Gu, Xin Ma 0001, Yibin Li 0001
Appl. Intell.2
2020 Edge-Semantic Learning Strategy for Layout Estimation in Indoor Environment
abstract
Visual cognition of the indoor environment can benefit from the spatial layout estimation, which is to represent an indoor scene with a 2-D box on a monocular image. In this paper, we propose to fully exploit the edge and semantic information of a room image for layout estimation. More specifically, we present an encoder-decoder network with shared encoder and two separate decoders, which are composed of multiple deconvolution (transposed convolution) layers, to jointly learn the edge maps and semantic labels of a room image. We combine these two network predictions in a scoring function to evaluate the quality of the layouts, which are generated by ray sampling and from a predefined layout pool. Guided by the scoring function, we apply a novel refinement strategy to further optimize the layout hypotheses. Experimental results show that the proposed network can yield accurate estimates of edge maps and semantic labels. By fully utilizing the two different types of labels, the proposed method achieves the state-of-the-art layout estimation performance on the benchmark datasets.
Weidong Zhang 0005, Wei Zhang 0021, Jason Gu
IEEE Trans. Cybern.3
2020 RBFNN-Based Adaptive Sliding Mode Control Design for Delayed Nonlinear Multilateral Telerobotic System With Cooperative Manipulation
abstract
Multilateral telerobotic system has potential applications in the industry environments with the advantages of cooperative manipulation for the remote and hazardous tasks, and its control design is quite challenging due to several coupling issues such as stability, position tracking, force feedback, and cooperative manipulation under time delays, various uncertainties, and external disturbance. In this paper, a novel radial basis function neural network (RBFNN) based adaptive sliding mode control design is proposed for nonlinear multilateral telerobotic system with n-master-n-slave manipulators. The environment force is modeled with a general form via the RBFNN-based environment parameters estimation in the slave side. The estimated environment parameters (nonpower signals) are transmitted to rebuild the environment dynamics in the master side and provide the good force feedback for the human operators. The RBFNN-based adaptive sliding mode controllers are designed separately for master and slave manipulators to achieve good position tracking under parameter variations and external disturbance. The coordinated force distribution algorithm is designed to achieve cooperative manipulation with the balance of force acting on the target object. The theoretical analysis is given and the comparative experiment for a nonlinear multilateral telerobotic system with 2-master-2-slave manipulators is implemented. The results show the good performance of our design.
Zheng Chen 0004, Fanghao Huang, Weichao Sun, Jason Gu, Shiqiang Zhu
IEEE Trans. Ind. Informatics7
2020 Further Results on Adaptive Stabilization of High-Order Stochastic Nonlinear Systems Subject to Uncertainties
abstract
This paper concerns the adaptive state-feedback control for a class of high-order stochastic nonlinear systems with uncertainties including time-varying delay, unknown control gain, and parameter perturbation. The commonly used growth assumptions on system nonlinearities are removed, and the adaptive control technique is combined with the sign function to deal with the unknown control gain. Then, with the help of the radial basis function neural network approximation approach and Lyapunov-Krasovskii functional, an adaptive state-feedback controller is obtained through the backstepping design procedure. It is verified that the constructed controller can render the closed-loop system semiglobally uniformly ultimately bounded. Finally, both the practical and numerical examples are presented to validate the effectiveness of the proposed scheme.
Huifang Min, Shengyuan Xu 0001, Jason Gu, Baoyong Zhang, Zhengqiang Zhang
IEEE Trans. Neural Networks Learn. Syst.3
2018 A Feature Descriptor Based on Local Normalized Difference for Real-World Texture Classification
abstract
In this paper, we propose a normalized difference vector (NDV) for texture representation. Compared to local-binary-pattern-based descriptors, the proposed NDV takes full advantage of the local difference, and the size can be extended flexibly to cover a large local region. We further employ the bag-of-words model to integrate the local descriptors into a global feature representation of an image. In addition, two strategies are introduced for the proposed NDV to achieve rotation invariance. We test the proposed texture descriptor on benchmark datasets, such as AniTex, VehApp, KTH-TIPS2a, OpenSurface, and Kylberg. Classification results demonstrate the superiority of the proposed descriptor over state-of-the-art methods.
Wei Zhang 0021, Weidong Zhang 0005, Kan Liu 0001, Jason Gu
IEEE Trans. Multim.4
2017 Correlation filter-based self-paced object tracking
abstract
Object tracking is an important capability for robots tasked with interacting with humans and the environment, and it enables robots to manipulate objects. In object tracking, selecting samples to learn a robust and efficient appearance model is a challenging task. Model learning determines both the strategy and frequency of model updating, which concerns many details that can affect the tracking results. In this paper, we propose an object tracking approach by formulating a new objective function that integrates the learning paradigm of self-paced learning into object tracking such that reliable samples can be automatically selected for model learning. Sample weights and model parameters can be learned by minimizing this single objective function under the framework of kernelized correlation filters. Moreover, a real-valued error-tolerant self-paced function with a constraint vector is proposed to combine prior knowledge, i.e., the characteristics of object tracking, with information learned during tracking. We demonstrate the robustness and efficiency of our object tracking approach on a recent object tracking benchmark data set: OTB 2013.
Wenhui Huang 0002, Jason Gu, Xin Ma 0001, Yibin Li 0001
ICRA2
2017 Visual-Tactile Fusion for Object Recognition
abstract
The camera provides rich visual information regarding objects and becomes one of the most mainstream sensors in the automation community. However, it is often difficult to be applicable when the objects are not visually distinguished. On the other hand, tactile sensors can be used to capture multiple object properties, such as textures, roughness, spatial features, compliance, and friction, and therefore provide another important modality for the perception. Nevertheless, effective combination of the visual and tactile modalities is still a challenging problem. In this paper, we develop a visual–tactile fusion framework for object recognition tasks. This paper uses the multivariate-time-series model to represent the tactile sequence and the covariance descriptor to characterize the image. Further, we design a joint group kernel sparse coding (JGKSC) method to tackle the intrinsically weak pairing problem in visual–tactile data samples. Finally, we develop a visual–tactile data set, composed of 18 household objects for validation. The experimental results show that considering both visual and tactile inputs is beneficial and the proposed method indeed provides an effective strategy for fusion.
Huaping Liu 0001, Yuanlong Yu 0001, Fuchun Sun 0001, Jason Gu
IEEE Trans Autom. Sci. Eng.4
2017 An Efficient Method for Traffic Sign Recognition Based on Extreme Learning Machine
abstract
This paper proposes a computationally efficient method for traffic sign recognition (TSR). This proposed method consists of two modules: 1) extraction of histogram of oriented gradient variant (HOGv) feature and 2) a single classifier trained by extreme learning machine (ELM) algorithm. The presented HOGv feature keeps a good balance between redundancy and local details such that it can represent distinctive shapes better. The classifier is a single-hidden-layer feedforward network. Based on ELM algorithm, the connection between input and hidden layers realizes the random feature mapping while only the weights between hidden and output layers are trained. As a result, layer-by-layer tuning is not required. Meanwhile, the norm of output weights is included in the cost function. Therefore, the ELM-based classifier can achieve an optimal and generalized solution for multiclass TSR. Furthermore, it can balance the recognition accuracy and computational cost. Three datasets, including the German TSR benchmark dataset, the Belgium traffic sign classification dataset and the revised mapping and assessing the state of traffic infrastructure (revised MASTIF) dataset, are used to evaluate this proposed method. Experimental results have shown that this proposed method obtains not only high recognition accuracy but also extremely high computational efficiency in both training and recognition processes in these three datasets.
Zhiyong Huang 0005, Yuanlong Yu 0001, Jason Gu, Huaping Liu 0001
IEEE Trans. Cybern.3
2017 Learning to Predict High-Quality Edge Maps for Room Layout Estimation
abstract
The goal of room layout estimation is to predict the three-dimensional box that represents the room spatial structure from a monocular image. In this paper, a deconvolution network is trained first to predict the edge map of a room image. Compared to the previous fully convolutional networks, the proposed deconvolution network has a multilayer deconvolution process that can refine the edge map estimate layer by layer. The deconvolution network also has fully connected layers to aggregate the information of every region throughout the entire image. During the layout generation process, an adaptive sampling strategy is introduced based on the obtained high-quality edge maps. Experimental results prove that the learned edge maps are highly reliable and can produce accurate layouts of room images.
Weidong Zhang 0005, Wei Zhang 0021, Kan Liu 0001, Jason Gu
IEEE Trans. Multim.4
2016 Deep Neural Networks for wireless localization in indoor and outdoor environments
Wei Zhang 0021, Kan Liu 0001, Weidong Zhang 0005, Youmei Zhang, Jason Gu
Neurocomputing5
2015 Experimental study of optimal Takagi Sugeno fuzzy controller for rotary inverted pendulum
abstract
This paper presents an experimental study of optimal Takagi Sugeno (TS) fuzzy controller for a rotary inverted pendulum. A TS fuzzy model of the simplified plant is first constructed using sector nonlinearity approach. A guaranteed cost TS fuzzy PDC optimal controller is then designed using LMI toolbox of MATLAB. The designed controller is experimentally evaluated on a ‘Quanser Qube Servo’ platform where it is compared with linear optimal controller which is also designed using LMI toolbox of MATLAB under the same design conditions. It is shown that TS fuzzy optimal controller shows better performance even though it is designed based on a simplified plant model.
Umar Farooq 0001, Jason Gu, Mohamed E. El-Hawary, Valentina Emilia Balas, Muhammad Usman 0011
FUZZ-IEEE2
2015 No-reference blur assessment based on edge modeling
Jingwei Guan, Wei Zhang 0021, Jason Gu, Hongliang Ren 0001
J. Vis. Commun. Image Represent.3
2014 Bhattacharyya distance-based irregular pyramid method for image segmentation
abstract
This paper proposes a new unsupervised image segmentation method by using Bhattacharyya distance‐based irregular pyramid, termed as ‘BDIP’ algorithm. The proposed BDIP algorithm obtains a suboptimal labelling solution under the condition that the number of segments is not manually given. It hierarchically builds each level of the irregular pyramid, with the result that the final segments emerge as they are represented by single nodes at certain levels. The BDIP algorithm employs Bhattacharyya distance to estimate the intra‐level similarity at higher pyramidal levels so as to improve the accuracy and robustness to noise. Furthermore, an adaptive neighbour search method is proposed such that the BDIP algorithm can self‐determine the number of segments. This method considers not only the graphic constraint, but also the similarity constraint in the sense that a candidate node is selected as a neighbour of the centre node if there is no boundary evidence between these two nodes. With the pyramidal accumulation, this evaluation is aggregated into the approximately global evidence, based on which the number of segments can be self‐determined. Experimental results have shown that this proposed BDIP algorithm outperforms other benchmark segmentation algorithms in terms of segmentation accuracy, labelling cost and robustness to noise.
Yuanlong Yu 0001, Jason Gu
IET Comput. Vis.2
2013 Development and Evaluation of Object-Based Visual Attention for Automatic Perception of Robots
abstract
Bottom-up visual attention is an automatic behavior to guide visual perception to a conspicuous object in a scene. This paper develops a new object-based bottom-up attention (OBA) model for robots. This model includes four modules: Extraction of preattentive features, preattentive segmentation, estimation of space-based saliency, and estimation of proto-object-based saliency. In terms of computation, preattentive segmentation serves as a bridge to connect the space-based saliency and object-based saliency. This paper therefore proposes a preattentive segmentation algorithm, which is able to self-determine the number of proto-objects, has low computational cost, and is robust in a variety of conditions such as noise and spatial transformations. Experimental results have shown that the proposed OBA model outperforms space-based attention model and other object-based attention methods in terms of accuracy of attentional selection, consistency under a series of noise settings and object completion.
Yuanlong Yu 0001, Jason Gu, George K. I. Mann, Ray G. Gosine
IEEE Trans Autom. Sci. Eng.2
2010 Multi-focus image fusion using PCNN
Zhaobin Wang, Yide Ma, Jason Gu
Pattern Recognit.3
2009 Regulation control of underactuated mechanical systems based on a new matching equation of port-controlled hamiltonian systems
abstract
We consider the control of Port-Controlled Hamiltonian (PCH) systems, which are a generalization of Euler-Lagrange Systems. A new matching equation for PCH systems is developed so that interconnection damping assignment passivity-based control (IDA-PBC) can be extended to the regulation of some underactuated PCH systems whosekinetic energy must be modified. A simple underactuated mechanical system (the inertial wheel pendulum) is used to demonstrate the effectiveness of the proposed method.
Peter B. Goldsmith, Jason Gu
ICRA3
2007 Bilateral Teleoperation of Robotic Systems with Predictive Control
abstract
This paper presents a new control approach with prediction to minimize the effects of time delays while ensuring stability and system performance. Two predictors at the slave and master sides are constructed assuming that the time delays in both transmission channels are measurable. Simulation and experimental results are compared with the scheme without prediction to show the effectiveness of this approach. The influence of data dropout to the proposed teleoperation system is studied in the experiment.
Ya-Jun Pan 0001, Jason Gu, Max Q.-H. Meng, Jayaprashanth Jayachandran
ICRA2
2007 Robust design for bilateral teleoperation system with Markov jumping parameters
abstract
This paper presents a robust design for the general bilateral teleoperation system with Markov jumping parameters. Although passivity-based approach and robust design-based approach can stabilize the bilateral teleoperation system, sufficient conditions are too conservative. In this paper, requirements for a general bilateral teleoperation system are given first. Next, time delays on the internet are modeled by a Markov process with nine states, three for the forward communication branch and three for the backward communication branch. Based on the requirements of the master and the slave controllers, the standard robust design are carried out. Then, with the model of internet time delays, the bilateral teleoperation system is reconstructed by a Markov jumping system. Based on the mode-dependent stability conditions for the stochastic time delay system, the closed-loop controller is designed to obtain the less conservative stable system and to therefore improve the performance of tracking and transparency. Simulation has been carried out and the results verify the feasibility and efficiency of the proposed approach.
Weimin Shen, Jason Gu, Evangelos E. Milios
IROS2
2007 Robust mixed H2/H∞ control of time-varying delay systems with extended LMI
abstract
This paper studies robust performance analysis of H2/Hinfincontrol problem in time-varying systems. In the case where the state-space matrices of the system depend affinely on the uncertain parameters, it is known that recently developed extended or dilated linear matrix inequalities (LMIs) are effective to assess the robust performance in a less conservative fashion. This paper further probes into those preceding results and proposes a new form of extended LMIs for time-varying H2/Hinfincontrollers synthesis. The new method enables us to parameterize controllers without involving the Lyapunov variables in the parameterization. The numerical simulations prove the validity of this framework.
Weimin Shen, Jason Gu, Max Q.-H. Meng
IROS3
2007 Control gain design for bilateral teleoperation systems using linear matrix inequalities
abstract
The application of LMI methods to teleoperation with a bounded communication delay is considered. A method to design a state and force feedback controller that guarantees the stability of the system with bounded error related to the rate of change of the operator’s and environment’s exerted force is derived. A numerical example is considered and the means of choosing the design parameters introduced in the derivation are demonstrated. A trade-off between position and force fidelity is outlined, whereby the controller gain could be computed offline and used to adjust the controllers in real-time to suit the current tast. The performance is demonstrated, showing the stability of the system.
Kevin Walker, Ya-Jun Pan 0001, Jason Gu
SMC3
2005 Fuzzy approach for mobile robot positioning
abstract
This paper presents a fuzzy approach for positioning an iRobot B21r mobile robot in an indoor environment. A novel error model for the laser rangefinder is built with consideration of the detection distance and the detection angle, and a new concept, the virtual angular point, is introduced as the feature for positioning the mobile robot in this paper. Such points as break points, real angular points, and virtual angular points are employed for positioning a mobile robot. Positions obtained by two arbitrary pairs of feature points are fused together by the weighted mean technique, and the weights are determined by the fuzzy accuracy of the feature points. Experimental study has been carried out to verify the effectiveness of the algorithms.
Weimin Shen, Jason Gu
IROS2
2005 Vision data registration for robot self-localization in 3D
abstract
We address the problem of globally consistent estimation of the trajectory of a robot arm moving in three dimensional space based on a sequence of binocular stereo images from a stereo camera mounted on the tip of the arm. Correspondence between 3D points from successive stereo camera positions is established through matching of 2D SIFT features in the images. We compare three different methods for solving this estimation problem, based on three distance measures between 3D points, Euclidean distance, Mahalanobis distance and a distance measure defined by a maximum likelihood formulation. Theoretical analysis and experimental results demonstrate that the maximum likelihood formulation is the most accurate. If the measurement error is guaranteed to be small, then Euclidean distance is the fastest, without significantly compromising accuracy, and therefore it is best for on-line robot navigation.
Pifu Zhang, Evangelos E. Milios, Jason Gu
IROS3
2004 An HTK-developed hidden Markov model (HMM) for a voice-controlled robotic system
abstract
This paper presents a voice-controlled mobile robotic system capable of recognizing voice commands and relaying them to a mobile robot. First, the specifications of the robot used will be presented. Second, a description of the HTK toolkit, the HMM, and the VCR software developed will be discussed. Finally, comparing our HTK-developed HMM against a commercially available Microsoft speech recognition engine (SDK 5.1) in terms of accuracy. The experimental evaluation and accuracy tests will show the ability of the VCR software to control a robot with simple human voice commands. Conclusions and future work are presented towards the end of the paper.
Osama Majdalawieh, Jason Gu, Max Q.-H. Meng
IROS2
2003 Control and data transmission for internet robots
abstract
For Internet-based tele-robotic systems (Internet robots), the most challenging and distinct difficulties are associated with Internet transmission delays, delay jitter and not-guaranteed bandwidth availability, which might lead to dramatic performance degradation or even instability. In this paper, a new approach to dealing with these problems is explored and implemented. Specifically, a rate-based end-to-end transport protocol is developed for real-time data transmission and an adaptive control scheme is developed to control the robot remotely. A mobile robot teleoperation system, ArtBot-I, is developed to verify and test the solutions. In the experiments, the users successfully guided a Pioneer-2 mobile robot through a laboratory environment remotely via the Internet using a web browser.
Peter Xiaoping Liu, Max Q.-H. Meng, Jason Gu, Simon X. Yang
ICRA3
2002 End-to-End Delay Boundary Prediction using Maximum Entropy Principle (MEP) for Internet-Based Teleoperation
abstract
Since data packets may get lost somewhere in the Internet connections, for real-time applications such as Internet-based teleoperation, delay boundary prediction plays an important role in determining properly whether a packet is lost or not. The predictors currently employed are lowpass filters based on the autoregressive and moving average (ARMA) models. However, recent studies and the results of the experiments in this paper show that the traditional ARMA model is not suitable because sometimes delays develop with quick and evident variation. In this paper, we present a novel adaptive algorithm for delay boundary prediction based on the maximum entropy principle (MEP). The results of our 3 successive working day experiments on 9 links which consists of academic, commercial and governmental ones among Northern America, Asia and Europe show that the MEP algorithm proposed has a better performance than the traditional ARMA method.
Peter Xiaoping Liu, Max Q.-H. Meng, Xiufen Ye, Jason Gu
ICRA4
2001 A study of Natural Eye Movement Detection and Ocular Implant Movement Control Using processed Electrooculograph Signals
abstract
This paper describes an intelligent sensor and control system, robotic prosthetic eye system, with artificial eye model, biomedical electrodes and a micro controller. The system intends to provide a rehabilitation ocular implant device that can be used by people with ocular implant. The proposed system can acquire the dynamic natural eye orientation signal, which is sent to the micro controller to control the artificial eye to have the same orientation. This paper starts with a brief review of various eye movement detection methods and then a proposed system is set up to carry out experimental study. Pilot study has demonstrated its potential for clinical applications.
Jason Gu, Max Q.-H. Meng, Albert Cook, M. Gary Faulkner
ICRA1
2001 Sensing and control of a robotic prosthetic eye for ocular implant
abstract
Describes two robotic prosthetic eye prototype models. The first model uses an external infrared sensor array mounted on a frame of a pair of eyeglasses to detect natural eye movement and to feed the control system to drive the artificial eye to move with the natural eye. The second model uses human brain EOG (electrooculography) signals picked up by electrodes placed on both sides of a person's head to carry out the same eye movement detection and control tasks as mentioned above. Theoretical issues on sensor failure detection and recovery, and signal processing techniques used in sensor data fusion are studied using statistical methods and artificial neural network based techniques. In addition, practical control system design and implementation using micro controllers are studied and implemented to carry out the natural eye movement detection and artificial robotic eye control tasks.
Jason Gu, Max Q.-H. Meng, A. Cook, M. Gary Faulkner, Peter Xiaoping Liu
IROS1
2000 Analysis of eye tracking movements using FIR median hybrid filters
abstract
This paper presents an approach of using FIR Median Hybrid Filters for analysis of eye tracking movements. The proposed filter can remove the eye blink artifact from the eye movement signal. The background of the project is described first. The whole idea is to put movements into eyes, which are used as static prosthesis, so that the ocular implant will have the same natural movement as the real eye. First step is to obtain the movement of the real eye. From the review of the eye movement methods, the electro-oculogram (EOG) is used to determine the eye position. Because the eye blink artifact is always corrupted in the EOG signal, it must be filtered out for the purpose of our project. The FIR Median Hybrid Filter is studied in the paper; its properties are explored with examples. Finally the filter is used to deal the real eye blink corrupted EOG signal. Examples are given of analysis procedure for eye tracking or a random moving target. The method is proved to be highly reliable.
Jason Gu, Max Q.-H. Meng, Albert Cook, M. Gary Faulkner
ETRA1
1998 Movement control system design for an artificial eye implant
abstract
The main objective of the research project reported in this paper is to design an assistive device that will help patients with eye-implant to have natural eye movement. The patients lose their eye for various reasons. The loss of an eye can be solved by the ocular implant. The artificial eye can be made like a real eye cosmetically. But the problem is that it is static and does not have the natural movement of an eye. We design an ocular assistive system to enable the artificial eye have the natural movement of a real eye. This paper starts with the literature review of the eye movement detection methods and followed by the description of the experimental system we designed and constructed. The paper is concluded with further considerations.
Jason Gu, Max Q.-H. Meng, M. Gary Faulkner, A. Cook
SMC1