VLDB 2026 Research / reviewers in the wild / expert
Jianwei Zhang 0001
dblp:z/JianweiZhang1
· DBLP profile ↗
218ranked-venue papers
12as first author
71since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 169 · 10 first-author · 54 since 2021Systems, architecture and hardware · 118 · 7 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Force-guided multimodal property estimation for robotic scooping of deformable objectsabstractRobotic feeding and serving require accurate reasoning about deformable foods, including recognizing food type and estimating the scooped weight. This is challenging because visually similar foods (e.g., millet vs. rice) provide ambiguous appearance cues, while force responses during scooping vary nonlinearly with portion size and utensil dynamics. We propose the Force-guided Text-prompt Multimodal Transformer (FTMT), a framework that integrates vision, force, and language supervision for robust food property estimation. Text prompts are used during training to semantically regularize force features, guiding multimodal tokenization and bi-directional cross-attention between visual and force representations. Experiments on eight deformable food types demonstrate that FTMT achieves improved recognition accuracy and lower weight estimation error compared to force-based and multimodal fusion baselines. Ablation studies highlight the importance of text–force alignment, force preprocessing, and cross-attention fusion. By enabling robots to reliably identify food and predict scooped weight, FTMT represents a step toward adaptive and preference-aware robot-assisted feeding in real-world settings. Zongdao Li, Qingdu Li, Fuchun Sun 0001, Jianwei Zhang 0001 |
Neurocomputing | 5 |
| 2026 | Language-guided cross-modal collaborative learning for unsupervised domain adaptation
Ai Peng, Zhouli Shen, Mao Ye 0001, Jianwei Zhang 0001 |
Neurocomputing | 6 |
| 2026 | Visibility-Guaranteed Tracking Control for Robotic Ureteroscopy Using Robust Control Barrier Functions and Neurodynamic Optimization
Zhen Deng, Bingwei He, Jianwei Zhang 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | Lie Algebra-Based Probabilistic Shape Estimation of Flexible Continuum Robots From Sparse Pose MeasurementsabstractEstimating the three-dimensional (3D) shape of tendon-driven continuum robots (TDCRs) under sparse pose measurements and unknown tip forces remains a significant challenge. This paper presents a novel Lie algebra-based probabilistic framework for continuous-arclength shape reconstruction of variable-curvature TDCRs subjected to unknown tip loads. By unifying Cosserat-rod and Cosserat-string theories, a Lie algebra-coupled static model is established that explicitly represents tendon actuation as distributed wrenches along the robot’s backbone. A sparse Gaussian Process prior in se(3) is derived via linearization to encode actuation-coupled shape distributions. This prior is fused with sparse, noisy pose measurements from electromagnetic sensors in a continuous-arclength batch estimation framework, enabling uncertainty-aware reconstruction along the backbone. Simulations and experiments on single-segment prototypes under diverse tip-loading conditions validate the method’s efficacy. Results demonstrate that the proposed approach achieves accurate shape estimation, with average shape (tip) errors of 1.6% (2.5%) of the total length, surpassing state-of-the-art methods. Comparable precision on helical-routing TDCR further confirms the framework’s generalizability. Zhouhui Yan, Zhen Deng, Jianwei Zhang 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2026 | Nonlinear Shape Control of Flexible Continuum Robots Using Offline-Online Learning With Neurodynamic OptimizationabstractPrecise shape control of tendon-driven continuum robots (TDCRs) remains challenging due to their inherent nonlinearity and environmental uncertainties. Analytical kinematic models fail to accurately characterize nonlinear deformation behavior. This paper proposes a model-less optimal shape control (MLOSC) framework using an offline-online learning strategy to achieve precise shape tracking under unknown loading conditions. The method approximates the TDCR’s shape Jacobian by a Radial Basis Function Neural Network (RBFNN), pretrained offline. Online adaptation of network parameters is performed via neurodynamic optimization to enhance prediction accuracy in the presence of uncertainties. The shape control problem is subsequently formulated as a constrained quadratic program (QP) that incorporates the robot’s physical limits. A neurodynamic solver is designed to compute optimal control inputs in real time. Control stability is analyzed via Lyapunov’s method. Experimental evaluations under static and dynamic loading conditions achieve average shape tracking errors of 2.04 mm (1.4% of robot total length) and 5.09 mm (3.4%), respectively. Comparative results confirm improved convergence speed and tracking accuracy over state-of-the-art methods. Ziqi Zhuang, Zhen Deng, Chuanchuan Pan, Ying Hu 0001, Jianwei Zhang 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | Zero-Shot Visual Grounding via Cascade Semantic Prototype LearningabstractCompared to standard visual grounding (VG), zero-shot visual grounding (ZSVG) requires robust localization of novel image–query pairs unseen during training, creating a clear gap from supervised settings. Existing two-stage approaches rely too heavily on external knowledge from pre-trained vision-language models and underexplore in-context semantics, while current one-stage models fail to leverage region-level cues to enhance discriminative representation learning. In this paper, inspired by the information processing procedure of human beings, we propose a novel one-stage network, dubbed Cascade Semantic Prototype Learning (CaSePro), to tackle these limitations. Specifically, CaSePro builds two types of prototypes that capture prototypical multimodal representations for global contexts and local semantics, respectively. With global context prototypes, CaSePro explores context-related cues to enhance the coarsely fused global multimodal features, enabling the model to learn more comprehensive representations. With local semantic prototypes, CaSePro searches for region-level multimodal correspondences to augment the prototypical global patterns, thereby strengthening the cross-modal relationship between the query and visual representations. Moreover, to enhance the generalization ability of the proposed approach on sample-imbalanced benchmark datasets, CaSePro further incorporates an Adaptive Focal Loss that dynamically reweights hard and easy samples by accounting for the model’s grounding uncertainty. Extensive experiments on ten benchmarks demonstrate that CaSePro outperforms all existing one-stage approaches and achieves performance comparable to state-of-the-art two-stage methods. Jinpeng Mi, Jinbo Yang, Xian Wei, Jianwei Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | StereoAnything: Advanced Zero-Shot Stereo Imaging for Robotic Grasp Detection With Transparent ObjectsabstractGrasping transparent objects remains challenging for robotic systems due to their reflective and refractive properties, which distort depth perception and introduce background noise. Unlike humans, who leverage life experience to perceive depth intuitively, robotic algorithms often fail to generalize across different object types. To address this, we propose a novel framework inspired by human perception for grasping transparent objects. Our approach extends features extracted by foundation models to implicitly learn reconstruction strategies for transparent objects without requiring segmentation priors. Crucially, our framework maintains strong performance across all types of objects and scenes, preventing catastrophic forgetting of opaque objects while learning to perceive transparent ones. By integrating affordance information, our method dynamically guides a five-finger dexterous hand to execute diverse grasping strategies based on human intent. To tackle the challenge of annotating transparent objects, we constructed a large-scale synthetic dataset with depth information, affordance data, and automated annotations. Our framework demonstrates strong generalization, achieving a 96% grasp success rate in real-world robotic experiments and proving its broad applicability across varied environments. Kaixin Bai, Lei Zhang 0198, Zhaopeng Chen, Jianwei Zhang 0001 |
IEEE Trans. Cybern. | 5 |
| 2025 | Proxy Denoising for Source-Free Domain AdaptationabstractSource-Free Domain Adaptation (SFDA) aims to adapt a pre-trained source model to an unlabeled target domain with no access to the source data. Inspired by the success of large Vision-Language (ViL) models in many applications, the latest research has validated ViL's benefit for SFDA by using their predictions as pseudo supervision. However, we observe that ViL's supervision could be noisy and inaccurate at an unknown rate, potentially introducing additional negative effects during adaption. To address this thus-far ignored challenge, we introduce a novel Proxy Denoising (__ProDe__) approach. The key idea is to leverage the ViL model as a proxy to facilitate the adaptation process towards the latent domain-invariant space. Concretely, we design a proxy denoising mechanism to correct ViL's predictions. This is grounded on a proxy confidence theory that models the dynamic effect of proxy's divergence against the domain-invariant space during adaptation. To capitalize the corrected proxy, we further derive a mutual knowledge distilling regularization. Extensive experiments show that ProDe significantly outperforms the current state-of-the-art alternatives under both conventional closed-set setting and the more challenging open-set, partial-set, generalized SFDA, multi-target, multi-source, and test-time settings. Our code and data are available at https://github.com/tntek/source-free-domain-adaptation. Song Tang 0001, Wenxin Su, Yan Gan, Mao Ye 0001, Jianwei Zhang 0001, Xiatian Zhu |
ICLR | 5 |
| 2025 | Multi-Segment Soft Robot Control Via Deep Koopman-Based Model Predictive ControlabstractSoft robots, compared to regular rigid robots, as their multiple segments with soft materials bring flexibility and compliance, have the advantages of safe interaction and dexterous operation in the environment. However, due to its characteristics of high dimensional, nonlinearity, time-varying nature, and infinite degree of freedom, it has been challenges in achieving precise and dynamic control such as trajectory tracking and position reaching. To address these challenges, we propose a framework of Deep Koopman-based Model Predictive Control (DK-MPC) for handling multi-segment soft robots. We first employ a deep learning approach with sampling data to approximate the Koopman operator, which therefore linearizes the high-dimensional nonlinear dynamics of the soft robots into a finite-dimensional linear representation. Secondly, this linearized model is utilized within a model predictive control framework to compute optimal control inputs that minimize the tracking error between the desired and actual state trajectories. The real-world experiments on the soft robot “Chordata” demonstrate that DK-MPC could achieve highprecision control, showing the potential of DK-MPC for future applications to soft robots. More visualization results can be found at https://pinkmoon-io.github.io/DKMPC/. Lei Lv, Lei Liu 0076, Fuchun Sun 0001, Jiahong Dong, Jianwei Zhang 0001, Xuemei Shan, Kai Sun 0014, Hao Huang 0009, Yu Luo 0021 |
ICRA | 6 |
| 2025 | 3D Dense Captioning via Prototypical Momentum Distillationabstract3D dense captioning aims to describe the crucial regions in 3D visual scenes in the form of natural language. Recent prevailing approaches achieve promising results by leveraging complicated structures incorporated with large-scale models, which necessitate abundant parameters and pose challenges regarding its practical applications. Besides, with limited training data, 3D dense captioners are often susceptible to overfitting, directly degrading caption generation performance. Drawing inspiration from the recent advancements in knowledge distillation, we propose a novel approach termed Prototypical Momentum Distillation (PMD) to prompt the model to generate more detailed captions. PMD incorporates Momentum Distillation (MD) with an Uncertainty-aware Prototype-anchored Clustering (UPC) strategy to transfer knowledge by considering the uncertainty of the teacher knowledge. Specifically, we employ the original captioner as the student model and maintain an Exponential Moving Average (EMA) copy of the captioner as the teacher model to impart knowledge as the auxiliary supervision of the student. To abate the misleading caused by uncertain knowledge, we present an Uncertainty-aware Prototype-anchored Clustering (UPC) strategy to cluster the distilled knowledge according to its confidence. We then transfer the rearranged knowledge from the teacher to guide the training route of the student. We conduct extensive experiments and ablation studies on two widely used benchmark datasets, ScanRefer and Nr3D. Experimental results demonstrate that PMD outperforms all state-of-the-art approaches on the benchmarks with MLE training, highlighting its effectiveness. Jinpeng Mi, Shaofei Jin, Xian Wei, Jianwei Zhang 0001 |
ICRA | 6 |
| 2025 | Plug-and-Play Multi-Domain Fusion Adaptation for Cross-Subject EEG-Based Motor Imagery ClassificationabstractMotor imagery (MI) classification in rehabilitation brain-computer interfaces (RBCIs) faces significant challenges due to the variability of electroencephalography (EEG) signals across subjects. Existing methods typically require extensive EEG data collection from each new subject, which is time-consuming and results in poor user experience. To address this issue, this paper decompose MI-EEG into subject-specific private components and shared components common across all subjects, and propose a plug-and-play domain fusion adaptive method (PPMDFA) to handle variability between subjects. In the training phase, PPMDFA introduces a Multi-Domain Fusion Graph Convolutional Network (MDFGCN) module to extract shared and private features from the MI processes of source domain subjects. In the calibration phase, the method constructs private classifiers for the target new subject using the extracted shared features combined with a small amount of labeled data. During testing, PPMDFA leverages the similarity of private components to utilize knowledge from source subjects, thereby enhancing classification accuracy for target subjects' MI. We validated the proposed method on the PhysioNet and LLMBCImotion datasets. Experimental results show that PPMDFA achieves state-of-the-art classification accuracy on both datasets, with rapid adaptation to new subjects using only 20% of the data, reaching accuracies of 73.33% and 61.62%, demonstrating strong generalization ability and robustness. Rui Huang 0008, Jianzhi Lyu, Yang Zhao 0024, Guangkui Song, Hong Cheng 0002, Jianwei Zhang 0001 |
ICRA | 8 |
| 2025 | Estimating Continuum Robot Shape under External Loading using spatiotemporal Neural NetworksabstractThis paper presents a learning-based approach for accurately estimating the 3D shape of flexible continuum robots subjected to external loads. The proposed method introduces a spatiotemporal neural network architecture that fuses multimodal inputs, including current and historical tendon displacement data and RGB images, to generate point clouds representing the robot’s deformed configuration. The network integrates a recurrent neural module for temporal feature extraction, an encoding module for spatial feature extraction, and a multimodal fusion module to combine spatial features extracted from visual data with temporal dependencies from historical actuator inputs. Continuous 3D shape reconstruction is achieved by fitting Bézier curves to the predicted point clouds. Experimental validation demonstrates that our approach achieves high precision, with mean shape estimation errors of 0.08 mm (unloaded) and 0.22 mm (loaded), outperforming state-of-the-art methods in shape sensing for TDCRs. The results validate the efficacy of deep learning-based spatiotemporal data fusion for precise shape estimation under loading conditions. Enyi Wang, Zhen Deng, Chuanchuan Pan, Bingwei He, Jianwei Zhang 0001 |
IROS | 5 |
| 2025 | ContactDexNet: Multi-fingered Robotic Hand Grasping in Cluttered Environments through Hand-Object Contact Semantic MappingabstractThe deep learning models has significantly advanced dexterous manipulation techniques for multi-fingered hand grasping. However, the contact information-guided grasping in cluttered environments remains largely underexplored. To address this gap, we have developed ContactDexNet, a method for generating multi-fingered hand grasp samples in cluttered settings through contact semantic map. We introduce a contact semantic conditional variational autoencoder network (CoSe-CVAE) for creating comprehensive contact semantic map from object point cloud. We utilize grasp detection method to estimate hand grasp poses from the contact semantic map. Finally, an unified grasp evaluation model PointNetGPD++ is designed to assess grasp quality and collision probability, substantially improving the reliability of identifying optimal grasps in cluttered scenarios. Our grasp generation method has demonstrated remarkable success, outperforming state-of-the-art (SOTA) methods by at least 4.7%, with 81.0% average grasping success rate in real-world single-object grasping using a known hand, and by at least 9.0% when using an unknown hand. Moreover, in cluttered scenes, our method attains a 76.7% success rate, outperforming the SOTA method by 6.3%. We also proposed the multi-modal multi-fingered grasping dataset generation method. Our multi-fingered hand grasping dataset outperforms previous datasets in scene diversity, modality diversity. More details and supplementary materials can be found at https://sites.google.com/view/contact-dexnet. Lei Zhang 0035, Kaixin Bai, Guowen Huang, Zhenshan Bing, Zhaopeng Chen, Alois C. Knoll, Jianwei Zhang 0001 |
IROS | 7 |
| 2025 | Awakening Facial Emotional Expressions in Human-RobotabstractThe facial expression generation capability of humanoid social robots is critical for achieving natural and human-like interactions, playing a vital role in enhancing the fluidity of human-robot interactions and the accuracy of emotional expression. Currently, facial expression generation in humanoid social robots still relies on pre-programmed be-havioral patterns, which are manually coded at high human and time costs. To enable humanoid robots to autonomously acquire generalized expressive capabilities, they need to develop the ability to learn human-like expressions through self-training. To address this challenge, we have designed a highly biomimetic robotic face with physical-electronic animated facial units and developed an end-to-end learning framework based on KAN (Kolmogorov-Arnold Network) and attention mechanisms. Unlike previous humanoid social robots, we have also meticulously designed an automated data collection system based on expert strategies of facial motion primitives to construct the dataset. Notably, to the best of our knowledge, this is the first open-source facial dataset for humanoid social robots. Comprehensive evaluations indicate that our approach achieves accurate and diverse facial mimicry across different test subjects. Yongtong Zhu, Iggy Qian, Qingdu Li, Na Liu 0007, Jianwei Zhang 0001 |
IROS | 8 |
| 2025 | Catching spinning table tennis balls in simulation with end-to-end Curriculum Reinforcement Learning
Yue Mao, Gang Wang 0024, Qingdu Li, Jianwei Zhang 0001, Yunfeng Ji |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | A lightweight network-based sign language robot with facial mirroring and speech system
Na Liu 0007, Xinchao Li, Baolei Wu, Lihong Wan, Jianwei Zhang 0001, Qingdu Li |
Expert Syst. Appl. | 7 |
| 2025 | Multi-level semantic-assisted prototype learning for Few-Shot Action Recognition
Mao Ye 0001, Jianwei Zhang 0001 |
Neurocomputing | 5 |
| 2025 | Boosting few-shot action recognition via time-enhanced multimodal adaptation learning
Ai Peng, Zhouli Shen, Feng Zhang 0052, Mao Ye 0001, Jianwei Zhang 0001 |
Neurocomputing | 7 |
| 2025 | Few-shot medical image segmentation with high-fidelity prototypesabstractFew-shot Semantic Segmentation (FSS) aims to adapt a pretrained model to new classes with as few as a single labeled training sample per class. Despite the prototype based approaches have achieved substantial success, existing models are limited to the imaging scenarios with considerably distinct objects and not highly complex background, e.g., natural images. This makes such models suboptimal for medical imaging with both conditions invalid. To address this problem, we propose a novel D etail S elf-refined P rototype Net work ( DSPNet ) to construct high-fidelity prototypes representing the object foreground and the background more comprehensively. Specifically, to construct global semantics while maintaining the captured detail semantics, we learn the foreground prototypes by modeling the multimodal structures with clustering and then fusing each in a channel-wise manner. Considering that the background often has no apparent semantic relation in the spatial dimensions, we integrate channel-specific structural information under sparse channel-aware regulation. Extensive experiments on three challenging medical image benchmarks show the superiority of DSPNet over previous state-of-the-art methods. The code and data are available at https://github.com/tntek/DSPNet . • A novel prototypical FSS approach DSPNet that enhances prototypes’ self-representation. • A class prototype self-refining method FSPA integrating the cluster prototypes. • A background prototype self-refining method BCMA coding channel-specific structure. Song Tang 0001, Shaxu Yan, Xiaozhi Qi, Mao Ye 0001, Jianwei Zhang 0001, Xiatian Zhu |
Medical Image Anal. | 6 |
| 2025 | SAM-Net: Semantic-assisted multimodal network for action recognition in RGB-D videos
Jinpeng Mi, Mao Ye 0001, Qingdu Li, Jianwei Zhang 0001 |
Pattern Recognit. | 6 |
| 2025 | Dynamic Grasping Based on Reinforcement Learning and Differentiable MPCabstractDynamic object grasping in human-robot shared workspaces presents significant challenges in ensuring grasping success, efficiency, and safety. To address these challenges, we propose a novel Actor-Critic reinforcement learning (RL) framework integrated with a differentiable nonlinear least squares (DNLS) optimization module. The RL component generates target poses and adaptively tunes cost function weights for DNLS, while the DNLS module refines motion trajectories to satisfy kinodynamic and safety constraints. The framework supports end-to-end training and enhances grasping performance and training speed, thanks to the GPU-accelerated DNLS optimization. We evaluate the proposed method in both simulation and real-world environments, demonstrating superior performance over baseline method in terms of success rate, efficiency, safety, and trajectory smoothness. These results highlight the effectiveness of our approach and its potential for enabling safe and reliable human-robot collaboration in dynamic, safety-critical environments. Jianzhi Lyu, Hui Zhang 0092, Youshuang Ding, Fuchun Sun 0001, Jianwei Zhang 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | EEG-Based Motor Imagery Classification With Tuned Heuristic Fusion Graph Convolutional Network for Rehabilitation TrainingabstractMotor imagery-based brain–computer interfaces (MI-BCIs) hold significant promise for rehabilitation training in individuals with neurological impairments such as stroke and spinal cord injury (SCI). Achieving precise and robust lower limb movement prediction for each patient is crucial. However, the variability in MI response frequencies and brain activation patterns among subjects presents a great challenge to the generalizability of MI-BCIs. This paper proposes a Tuned Heuristic Fusion Graph Convolutional Network (THFGCN) for limb movement prediction in rehabilitation scenarios. THFGCN innovatively designs a learnable EEG frequency band tuned module and a heuristic space topology module. These two modules allow for the intricate extraction of both frequency and spatial topological features, utilizing graph adjacency matrices that encapsulate channel correlations and spatial relationships, hence fostering individualized analysis and enhanced generalizability across subjects. Furthermore, a spatio-temporal convolution module paired with a feature map attention mechanism is proposed to extract the critical spatio-temporal features of electroencephalogram (EEG) data. Validation experiments on the PhysioNet and LLM-BCImotion datasets against six mainstream methods demonstrate that THFGCN outperforms state-of-theart methods, achieving 88.41% and 82.82% accuracy in the within-subject case, and 65.93% and 60.56% accuracy in the cross-subject case, respectively. Detailed frequency band weight and T-distributed Stochastic Neighbor Embedding visualization validate the effectiveness of proposed modules. Furthermore, feature interpretability analysis proves the extracted features’ profound MI task relevance, underlining THFGCN’s exceptional interpretability. Rui Huang 0008, Jianzhi Lyu, Fengjun Mu, Zhinan Peng, Chaobin Zou, Hong Cheng 0002, Jianwei Zhang 0001, Bijoy K. Ghosh |
IEEE Trans Autom. Sci. Eng. | 9 |
| 2025 | Soft Contact Simulation and Manipulation Learning of Deformable Objects With Vision-Based Tactile SensorabstractDeformable object manipulation is a challenging problem due to its complex deformable properties. With the development of artificial intelligence, learning-based methods have shown outstanding performance in robotic manipulation. Previous works have investigated the manipulation of deformable objects via Reinforcement Learning (RL) in simulation. However, they approximate object deformation with particles, using particle states as observations, which are unavailable in reality. To address these issues, we utilize Vision-Based Tactile Sensors (VBTSs) as the end-effector to manipulate and observe the deformable objects. In this work, we develop a new contact simulation environment for deformable objects, including elastic, plastic, and elastoplastic. We utilize RL strategies and expert demonstrations to train agents in the simulation. Finally, we build a real experimental platform to complete the sim-to-real tasks and robustness testing. Our work introduces an innovative strategy that utilizes high-resolution VBTSs for contact simulation and manipulation of deformable objects. The experimental results show superior performances of deformable object manipulation with the proposed method. Shixin Zhang, Zixi Chen 0002, Zirong Shen, Fuchun Sun 0001, Cesare Stefanini, Di Guo 0002, Shan Luo 0001, Jianwei Zhang 0001, Jianhua Shan, Bin Fang 0003 |
IEEE Trans Autom. Sci. Eng. | 9 |
| 2025 | ADG-Net: A Sim2Real Multimodal Learning Framework for Adaptive Dexterous GraspingabstractIn this article, a novel simulation-to-real (sim2real) multimodal learning framework is proposed for adaptive dexterous grasping and grasp status prediction. A two-stage approach is built upon the Isaac Gym and several proposed pluggable modules, which can effectively simulate dexterous grasps with multimodal sensing data, including RGB-D images of grasping scenarios, joint angles, 3-D tactile forces of soft fingertips, etc. Over 500K multimodal synthetic grasping scenarios are collected for neural network training. An adaptive dexterous grasping neural network (ADG-Net) is trained to learn dexterous grasp principles and predict grasp parameters, employing an attention mechanism and a graph convolutional neural network module to fuse multimodal information. The proposed adaptive dexterous grasping method can detect feasible grasp parameters from an RGB-D image of a grasp scene and then optimize grasp parameters based on multimodal sensing data when the dexterous hand touches a target object. Various experiments in both simulation and physical grasps indicate that our ADG-Net grasping method outperforms state-of-the-art grasping methods, achieving an average success rate of 92% for grasping isolated unseen objects and 83% for stacked objects. Code and video demos are available at https://github.com/huikul/adgnet. Hui Zhang 0092, Jianzhi Lyu, Chuangchuang Zhou, Hongzhuo Liang, Yuyang Tu, Fuchun Sun 0001, Jianwei Zhang 0001 |
IEEE Trans. Cybern. | 7 |
| 2025 | Goal-Conditioned Hierarchical Reinforcement Learning With High-Level Model ApproximationabstractHierarchical reinforcement learning (HRL) exhibits remarkable potential in addressing large-scale and long-horizon complex tasks. However, a fundamental challenge, which arises from the inherently entangled nature of hierarchical policies, has not been understood well, consequently compromising the training stability and exploration efficiency of HRL. In this article, we propose a novel HRL algorithm, high-level model approximation (HLMA), presenting both theoretical foundations and practical implementations. In HLMA, a Planner constructs an innovative high-level dynamic model to predict the -step transition of the Controller in a subtask. This allows for the estimation of the evolving performance of the Controller. At low level, we leverage the initial state of each subtask, transforming absolute states into relative deviations by a designed operator as Controller input. This approach facilitates the reuse of subtask domain knowledge, enhancing data efficiency. With this designed structure, we establish the local convergence of each component within HLMA and subsequently derive regret bounds to ensure global convergence. Abundant experiments conducted on complex locomotion and navigation tasks demonstrate that HLMA surpasses other state-of-the-art single-level RL and HRL algorithms in terms of sample efficiency and asymptotic performance. In addition, thorough ablation studies validate the effectiveness of each component of HLMA. Yu Luo 0021, Tianying Ji, Fuchun Sun 0001, Huaping Liu 0001, Jianwei Zhang 0001, Mingxuan Jing, Wenbing Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | DreamArrangement: Learning Language-Conditioned Robotic Rearrangement of Objects via Denoising Diffusion and VLM PlannerabstractThe capability for robotic systems to rearrange objects based on human instructions represents a critical step toward realizing embodied intelligence. Recently, diffusion-based learning has shown significant advancements in the field of data generation while prompt-based learning has proven effective in formulating robot manipulation strategies. However, prior solutions for robotic rearrangement have overlooked the significance of integrating human preferences and optimizing for rearrangement efficiency. Additionally, traditional prompt-based approaches struggle with complex, semantically meaningful rearrangement tasks without predefined target states for objects. To address these challenges, our work first introduces a comprehensive two dimensional (2-D) tabletop rearrangement dataset, utilizing a physical simulator to capture interobject relationships and semantic configurations. Then, we present DreamArrangement, a novel language-conditioned object rearrangement scheme, consisting of two primary processes: employing a transformer-based multimodal denoising diffusion model to envisage the desired arrangement of objects, and leveraging a vision–language foundational model to derive actionable policies from text, alongside initial and target visual information. In particular, we introduce an efficiency-oriented learning strategy to minimize the average motion distance of objects. Given few-shot instruction examples, the learned policy from our synthetic dataset can be transferred to the real world without extra human intervention. Extensive simulations validate DreamArrangement’s superior rearrangement quality and efficiency. Moreover, real-world robotic experiments confirm that our method can adeptly execute a range of challenging, language-conditioned, and long-horizon tasks with a singular model. The demonstration video can be found at https://youtu.be/fq25-DjrbQE Changming Xiao, Fuchun Sun 0001, Changshui Zhang, Jianwei Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2024 | Subequivariant Reinforcement Learning in 3D Multi-Entity Physical EnvironmentsabstractLearning policies for multi-entity systems in 3D environments is far more complicated against single-entity scenarios, due to the exponential expansion of the global state space as the number of entities increases. One potential solution of alleviating the exponential complexity is dividing the global space into independent local views that are invariant to transformations including translations and rotations. To this end, this paper proposes Subequivariant Hierarchical Neural Networks (SHNN) to facilitate multi-entity policy learning. In particular, SHNN first dynamically decouples the global space into local entity-level graphs via task assignment. Second, it leverages subequivariant message passing over the local entity-level graphs to devise local reference frames, remarkably compressing the representation redundancy, particularly in gravity-affected environments. Furthermore, to overcome the limitations of existing benchmarks in capturing the subtleties of multi-entity systems under the Euclidean symmetry, we propose the Multi-entity Benchmark (MEBEN), a new suite of environments tailored for exploring a wide range of multi-entity reinforcement learning. Extensive experiments demonstrate significant advancements of SHNN on the proposed benchmarks compared to existing methods. Comprehensive ablations are conducted to verify the indispensability of task assignment and subequivariance. Runfa Chen, Tianrui Xue, Fuchun Sun 0001, Jianwei Zhang 0001, Wenbing Huang 0001 |
ICML | 6 |
| 2024 | Seizing Serendipity: Exploiting the Value of Past Success in Off-Policy Actor-CriticabstractLearning high-quality $Q$-value functions plays a key role in the success of many modern off-policy deep reinforcement learning (RL) algorithms. Previous works primarily focus on addressing the value overestimation issue, an outcome of adopting function approximators and off-policy learning. Deviating from the common viewpoint, we observe that $Q$-values are often underestimated in the latter stage of the RL training process, potentially hindering policy learning and reducing sample efficiency. We find that such a long-neglected phenomenon is often related to the use of inferior actions from the current policy in Bellman updates as compared to the more optimal action samples in the replay buffer. We propose the Blended Exploitation and Exploration (BEE) operator, a simple yet effective approach that updates $Q$-value using both historical best-performing actions and the current policy. Based on BEE, the resulting practical algorithm BAC outperforms state-of-the-art methods in over 50 continuous control tasks and achieves strong performance in failure-prone scenarios and real-world robot tasks. Benchmark results and videos are available at https://jity16.github.io/BEE/. Tianying Ji, Yu Luo 0021, Fuchun Sun 0001, Xianyuan Zhan, Jianwei Zhang 0001, Huazhe Xu |
ICML | 5 |
| 2024 | OMPO: A Unified Framework for RL under Policy and Dynamics ShiftsabstractTraining reinforcement learning policies using environment interaction data collected from varying policies or dynamics presents a fundamental challenge. Existing works often overlook the distribution discrepancies induced by policy or dynamics shifts, or rely on specialized algorithms with task priors, thus often resulting in suboptimal policy performances and high learning variances. In this paper, we identify a unified strategy for online RL policy learning under diverse settings of policy and dynamics shifts: transition occupancy matching. In light of this, we introduce a surrogate policy learning objective by considering the transition occupancy discrepancies and then cast it into a tractable min-max optimization problem through dual reformulation. Our method, dubbed Occupancy-Matching Policy Optimization (OMPO), features a specialized actor-critic structure equipped with a distribution discriminator and a small-size local buffer. We conduct extensive experiments based on the OpenAI Gym, Meta-World, and Panda Robots environments, encompassing policy shifts under stationary and non-stationary dynamics, as well as domain adaption. The results demonstrate that OMPO outperforms the specialized baselines from different categories in all settings. We also find that OMPO exhibits particularly strong performance when combined with domain randomization, highlighting its potential in RL-based robotics applications. Yu Luo 0021, Tianying Ji, Fuchun Sun 0001, Jianwei Zhang 0001, Huazhe Xu, Xianyuan Zhan |
ICML | 4 |
| 2024 | Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RLabstractOff-policy reinforcement learning (RL) has achieved notable success in tackling many complex real-world tasks, by leveraging previously collected data for policy learning. However, most existing off-policy RL algorithms fail to maximally exploit the information in the replay buffer, limiting sample efficiency and policy performance. In this work, we discover that concurrently training an offline RL policy based on the shared online replay buffer can sometimes outperform the original online learning policy, though the occurrence of such performance gains remains uncertain. This motivates a new possibility of harnessing the emergent outperforming offline optimal policy to improve online policy learning. Based on this insight, we present Offline-Boosted Actor-Critic (OBAC), a model-free online RL framework that elegantly identifies the outperforming offline policy through value comparison, and uses it as an adaptive constraint to guarantee stronger policy learning performance. Our experiments demonstrate that OBAC outperforms other popular model-free RL baselines and rivals advanced model-based RL methods in terms of sample efficiency and asymptotic performance across **53** tasks spanning **6** task suites. Yu Luo 0021, Tianying Ji, Fuchun Sun 0001, Jianwei Zhang 0001, Huazhe Xu, Xianyuan Zhan |
ICML | 4 |
| 2024 | A Collision-Aware Cable Grasping Method in Cluttered EnvironmentabstractWe introduce a Cable Grasping-Convolutional Neural Network (CG-CNN) designed to facilitate robust cable grasping in cluttered environments. Utilizing physics simulations, we generate an extensive dataset that mimics the intricacies of cable grasping, factoring in potential collisions between cables and robotic grippers. We employ the Approximate Convex Decomposition technique to dissect the non-convex cable model, with grasp quality autonomously labeled based on simulated grasping attempts. The CG-CNN is refined using this simulated dataset and enhanced through domain randomization techniques. Subsequently, the trained model predicts grasp quality, guiding the optimal grasp pose to the robot’s controller for execution. Grasping efficacy is assessed across both synthetic and real-world settings. Given our model’s implicit collision sensitivity, we achieved commendable success rates of 92.3% for known cables and 88.4% for unknown cables, surpassing contemporary state-of-the-art approaches. Supplementary materials can be found at https://leizhang-public.github.io/cg-cnn/. Lei Zhang 0198, Kaixin Bai, Qiang Li 0001, Zhaopeng Chen, Jianwei Zhang 0001 |
ICRA | 5 |
| 2024 | Close the Sim2real Gap via Physically-based Structured Light Synthetic Data SimulationabstractDespite the substantial progress in deep learning, its adoption in industrial robotics projects remains limited, primarily due to challenges in data acquisition and labeling. Previous sim2real approaches using domain randomization require extensive scene and model optimization. To address these issues, we introduce an innovative physically-based structured light simulation system, generating both RGB and physically realistic depth images, surpassing previous dataset generation tools. We create an RGBD dataset tailored for robotic industrial grasping scenarios and evaluate it across various tasks, including object detection, instance segmentation, and embedding sim2real visual perception in industrial robotic grasping. By reducing the sim2real gap and enhancing deep learning training, we facilitate the application of deep learning models in industrial settings. Project details are available at https://baikaixin-public.github.io/structured_light_3D_synthesizer/ Kaixin Bai, Lei Zhang 0198, Zhaopeng Chen, Jianwei Zhang 0001 |
ICRA | 5 |
| 2024 | Pluck and Play: Self-supervised Exploration of Chordophones for Robotic PlayingabstractExisting robotic musicians utilize detailed handcrafted instrument models to generate or learn policies for playing because model-free or inaccurate policy rollouts might easily damage or wear out fragile instruments. We introduce an approach to characterize geometric models of chordophones and their audio onset responses directly through audio-tactile exploration with a physical robot arm. Initially, the system refines prior estimates of string positions, provided by kinesthetic teaching or visual estimation, through repeated attempts to pluck individual strings. A subsequent stage implements a Safe Active Exploration paradigm based on Gaussian Processes to explore and characterize the audio onset response of feasible plucking motions while minimizing invalid attempts. The resulting models can be used to actuate an imprecise robotic arm to play sequences of notes with varying loudness on a Chinese Guzheng. Michael Görner, Norman Hendrich, Jianwei Zhang 0001 |
ICRA | 3 |
| 2024 | Smooth Computation without Input Delay: Robust Tube-Based Model Predictive Control for Robot Manipulator PlanningabstractModel Predictive Control (MPC) has exhibited remarkable capabilities in optimizing objectives and meeting constraints. However, the substantial computational burden associated with solving the Optimal Control Problem (OCP) at each triggering instant introduces significant delays between state sampling and control application. These delays limit the practicality of MPC in resource-constrained systems when engaging in complex tasks. The intuition to address this issue in this paper is that by predicting the successor state, the controller can solve the OCP one time step ahead of time thus avoiding the delay of the next action. To this end, we compute deviations between real and nominal system states, predicting forthcoming real states as initial conditions for the imminent OCP solution. Anticipatory computation stores optimal control based on current nominal states, thus mitigating the delay effects. Additionally, we establish an upper bound for linearization error, effectively linearizing the nonlinear system, reducing OCP complexity, and enhancing response speed. We provide empirical validation through two numerical simulations and corresponding real-world robot tasks, demonstrating significant performance improvements and augmented response speed (up to 90%) resulting from the seamless integration of our proposed approach compared to conventional time-triggered MPC strategies. Yu Luo 0021, Qie Sima, Tianying Ji, Fuchun Sun 0001, Huaping Liu 0001, Jianwei Zhang 0001 |
ICRA | 6 |
| 2024 | Sensor-agnostic Visuo-Tactile Robot Calibration Exploiting Assembly-Precision Model GeometriesabstractVisual sensor modalities dominate traditional robot calibration, but when environment contacts are relevant, the tactile modality can provide another natural, accurate, and highly relevant modality. Most existing tactile sensing methods for robot calibration are constrained to specific sensor-object pairs, limiting their applicability. This paper pioneers a general approach to exploit contacts in robot calibration, supporting self-touch throughout the entire system kinematics by generalizing touchable surfaces to any accurately represented mesh surface. The approach supports different contact sensors as long as a simple single-contact interface can be provided. Integrated into the Atomic Transformation Optimization Method (ATOM) calibration methodology, our work facilitates seamless integration of both modalities in a single approach. Our results demonstrate comparable performance to single-modality calibration but can trade off accuracy between both modalities, thus increasing overall robustness. Furthermore, we observe that utilizing a touch point at the end of a kinematic chain slightly improves calibration over touching the chain links with an external sensor but find no significant advantage of restricting touch to end-effector contacts when calibrating a dual-arm system with our method. Manuel Gomes, Michael Görner, Miguel Armando Riem de Oliveira, Jianwei Zhang 0001 |
IROS | 4 |
| 2024 | State Estimation of an Adaptive 3-Finger Gripper using Recurrent Neural NetworksabstractAdaptive grippers enable easy and robust grasping of diverse objects by adapting to their shapes and enclosing them. However, determining the exact state of the hand remains challenging. This is not always straightforward but is often necessary to assess grip success, quality, or the pose of the object. In this work, we present two deep learning approaches using recurrent neural networks to successfully estimate the joint states of the Robotiq 3-Finger Adaptive Gripper. The models are compared with an existing analytical approach, which does not distinguish between the fingers of the hand and calculates the three angles of their respective joints using joint limits, contact information, and motor position in a transition model. We test the differences in accuracy with our networks by not distinguishing between the fingers as the analytical approach for the first model and by looking at the entire hand in the second model. Our experiments demonstrate that the model considering the entire hand outperforms the other two approaches, is more robust against object movements and achieves an average joint position accuracy of 2.29 degrees. Yannick Jonetzko, Theresa Alexandra Aurelia Naß, Niklas Fiedler, Jianwei Zhang 0001 |
IROS | 4 |
| 2024 | ToolEENet: Tool Affordance 6D Pose EstimationabstractThe exploration of robotic dexterous hands utilizing tools has recently attracted considerable attention. A significant challenge in this field is the precise awareness of a tool’s pose when grasped, as occlusion by the hand often degrades the quality of the estimation. Additionally, the tool’s overall pose often fails to accurately represent the contact interaction, thereby limiting the effectiveness of vision-guided, contact-dependent activities. To overcome this limitation, we present the innovative TOOLEE dataset, which, to the best of our knowledge, is the first to feature affordance segmentation of a tool’s end-effector (EE) along with its defined 6D pose based on its usage. Furthermore, we propose the ToolEENet framework for accurate 6D pose estimation of the tool’s EE. This framework begins by segmenting the tool’s EE from raw RGB-D data, then uses a diffusion model-based pose estimator for 6D pose estimation at a category-specific level. Addressing the issue of symmetry in pose estimation, we introduce a symmetry-aware pose representation that enhances the consistency of pose estimation. Our approach excels in this field, demonstrating high levels of precision and generalization. Furthermore, it shows great promise for application in contact-based manipulation scenarios. All data and codes are available on the project website: https://tooleenet-iros2024.github.io/ Lei Zhang 0198, Yuyang Tu, Hui Zhang 0070, Kaixin Bai, Zhaopeng Chen, Jianwei Zhang 0001 |
IROS | 7 |
| 2024 | Deep Learning Based Measurement Model for Monte Carlo Localization in the RoboCup Humanoid League
Jasper Güldenstein, Niklas Fiedler, Jianwei Zhang 0001 |
RoboCup | 3 |
| 2024 | Robust tube-based MPC with smooth computation for dexterous robot manipulation
Yu Luo 0021, Tianying Ji, Fuchun Sun 0001, Qie Sima, Huaping Liu 0001, Mingxuan Jing, Jianwei Zhang 0001 |
Sci. China Inf. Sci. | 7 |
| 2024 | Temporal cues enhanced multimodal learning for action recognition in RGB-D videos
Zhiyuan Ma 0001, Jinpeng Mi, Yan Gan, Mao Ye 0001, Jianwei Zhang 0001 |
Neurocomputing | 8 |
| 2024 | Zero-shot visual grounding via coarse-to-fine representation learning
Jinpeng Mi, Shaofei Jin, Zhiqian Chen, Xian Wei, Jianwei Zhang 0001 |
Neurocomputing | 6 |
| 2024 | Adaptive knowledge distillation and integration for weakly supervised referring expression comprehension
Jinpeng Mi, Stefan Wermter, Jianwei Zhang 0001 |
Knowl. Based Syst. | 3 |
| 2024 | Multimodality Driven Impedance-Based Sim2Real Transfer Learning for Robotic Multiple Peg-in-Hole AssemblyabstractRobotic rigid contact-rich manipulation in an unstructured dynamic environment requires an effective resolution for smart manufacturing. As the most common use case for the intelligence industry, a lot of studies based on reinforcement learning (RL) algorithms have been conducted to improve the performances of single peg-in-hole assembly. However, existing RL methods are difficult to apply to multiple peg-in-hole issues due to more complicated geometric and physical constraints. In addition, previously limited solutions for multiple peg-in-hole assembly are hard to transfer into real industrial scenarios flexibly. To effectively address these issues, this work designs a novel and more challenging multiple peg-in-hole assembly setup by using the advantage of the Industrial Metaverse. We propose a detailed solution scheme to solve this task. Specifically, multiple modalities, including vision, proprioception, and force/torque, are learned as compact representations to account for the complexity and uncertainties and improve the sample efficiency. Furthermore, RL is used in the simulation to train the policy, and the learned policy is transferred to the real world without extra exploration. Domain randomization and impedance control are embedded into the policy to narrow the gap between simulation and reality. Evaluation results demonstrate the effectiveness of the proposed solution, showcasing successful multiple peg-in-hole assembly and generalization across different object shapes in real-world scenarios. Chao Zeng 0002, Hongzhuo Liang, Fuchun Sun 0001, Jianwei Zhang 0001 |
IEEE Trans. Cybern. | 5 |
| 2024 | A Dexterous Hand-Arm Teleoperation System Based on Hand Pose Estimation and Active VisionabstractMarkerless vision-based teleoperation that leverages innovations in computer vision offers the advantages of allowing natural and noninvasive finger motions for multifingered robot hands. However, current pose estimation methods still face inaccuracy issues due to the self-occlusion of the fingers. Herein, we develop a novel vision-based hand-arm teleoperation system that captures the human hands from the best viewpoint and at a suitable distance. This teleoperation system consists of an end-to-end hand pose regression network and a controlled active vision system. The end-to-end pose regression network (Transteleop), combined with an auxiliary reconstruction loss function, captures the human hand through a low-cost depth camera and predicts joint commands of the robot based on the image-to-image translation method. To obtain the optimal observation of the human hand, an active vision system is implemented by a robot arm at the local site that ensures the high accuracy of the proposed neural network. Human arm motions are simultaneously mapped to the slave robot arm under relative control. Quantitative network evaluation and a variety of complex manipulation tasks, for example, tower building, pouring, and multitable cup stacking, demonstrate the practicality and stability of the proposed teleoperation system. Shuang Li 0014, Norman Hendrich, Hongzhuo Liang, Philipp Ruppel, Changshui Zhang, Jianwei Zhang 0001 |
IEEE Trans. Cybern. | 6 |
| 2024 | EHoA: A Benchmark for Task-Oriented Hand-Object Action Recognition via Event VisionabstractThe event-based camera is a novel neuromorphic vision sensor that can perceive different dynamic behaviors due to its low latency, asynchronous data stream, and high dynamic range characteristics. There has been much work based on event cameras to solve problems, such as object tracking, visual odometry, and gesture recognition. However, the adoption of event vision to analyze hand-object action in a dynamic environment, a problem that regular CMOS cameras cannot handle, is still lacking in relevant research. This work presents a richly annotated task-oriented hand-object action dataset consisting of asynchronous event streams, captured by the event-based camera system on different application scenarios. In addition, we design an attention-based residual spiking neural network by learning temporal-wise and spatial-wise attention simultaneously and introducing a particular residual connection structure to achieve dynamic hand-object action recognition. Extensive experiments are validated by comparing with existing baseline methods to form a vision benchmark. We also show that the learned recognition model can be transferred to classify a real robot hand-object action. Shang-Ching Liu, Jianwei Zhang 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Progressive Source-Aware Transformer for Generalized Source-Free Domain AdaptationabstractSource-free domain adaptation (SFDA) tends to forget the source domain, suffering from limitations in real-world scenarios. Recently, generalized source-free domain adaptation (GSFDA) problem naturally emerges, aiming for good performance on both target and source domains. The existing methods attempt to retain model parameters associated with the source domain to prevent such forgetting. However, this strategy is not conducive to improving cross-domain performance on the target domain, prioritizing mitigating forgetting on the source domain. This article introduces a Progressive Source-Aware Transformer approach for GSFDA, dubbed PSAT-GDA. Our core idea is to enforce the domain adaptation process to remember the source domain by imposing source guidance, offering a target domain-centric anti-forgetting mechanism. Specifically, for each epoch, a Transformer-based deep network is adapted to do domain alignment like the traditional SFDA method, because the transformer working on the image patch sequence helps to reduce image noise caused by domain shift. Meanwhile, another Transformer is designed to generate source guidance supervising domain alignment. By augmenting target sample and mining the source information from the historical models before current epoch, source injected feature group is constructed. Based on the Transformer mechanism, the attention block can select useful source information for each target sample. From it, we devise neighbour-based and augmentation-based regularizations to shape the source guidance. Experiments on three challenging datasets show that our method can achieve evident cross-domain improvement on the target domains. Also, it can mitigate forgetting on all domains after adapting to single or multiple target domains. Song Tang 0001, Yuji Shi, Mao Ye 0001, Changshui Zhang, Jianwei Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | Weakly Supervised Referring Expression Grounding via Target-Guided Knowledge DistillationabstractWeakly supervised referring expression grounding aims to train a model without the manual labels between image regions and referring expressions during the training phase. Current predominant models often adopt deep structures to reconstruct the region-expression correspondence. A crucial deficiency of the existing approaches lies in that these models neglect to exploit potential valuable information to further improve their grounding performance. To address this issue, we leverage knowledge distillation as a unique scheme to excavate and transfer helpful information for acquiring a better model. Specifically, we propose a target-guided knowledge distillation framework that accounts for region-expression pairs reconstruction and matching. We reactivate the target-related prediction information learned by a pre-trained teacher model and transfer the target-related prediction knowledge from the teacher to guide the training process and boost the performance of the student model. We conduct extensive experiments on three benchmark datasets, i.e., RefCOCO, RefCOCO+, and RefCOCOg. Without bells and whistles, our approach achieves state-of-the-art results on several splits of benchmark datasets. The implementation codes and trained models are available at: https://github.com/dami23/WREG_KD. Jinpeng Mi, Song Tang 0001, Zhiyuan Ma 0001, Qingdu Li, Jianwei Zhang 0001 |
ICRA | 6 |
| 2023 | Reinforcement Learning Based Pushing and Grasping Objects from Ungraspable PosesabstractGrasping an object when it is in an ungraspable pose is a challenging task, such as books or other large flat objects placed horizontally on a table. Inspired by human manipulation, we address this problem by pushing the object to the edge of the table and then grasping it from the hanging part. In this paper, we develop a model-free Deep Reinforcement Learning framework to synergize pushing and grasping actions. We first pre-train a Variational Autoencoder to extract high-dimensional features of input scenario images. One Proximal Policy Optimization algorithm with the common reward and sharing layers of Actor-Critic is employed to learn both pushing and grasping actions with high data efficiency. Experiments show that our one network policy can converge 2.5 times faster than the policy using two parallel networks. Moreover, the experiments on unseen objects show that our policy can generalize to the challenging case of objects with curved surfaces and off-center irregularly shaped objects. Lastly, our policy can be transferred to a real robot without fine-tuning by using CycleGAN for domain adaption and outperforms the push-to-wall baseline. Hongzhuo Liang, Jianzhi Lyu, Long Zeng 0001, Pingfa Feng, Jianwei Zhang 0001 |
ICRA | 7 |
| 2023 | Weakly Supervised Referring Expression Grounding via Dynamic Self-Knowledge DistillationabstractWeakly supervised referring expression grounding (WREG) is an attractive and challenging task for grounding target regions in images by understanding given referring expressions. WREG learns to ground target objects without the manual annotations between image regions and referring expressions during the model training phase. Different from the predominant grounding pattern of existing models, which locates target objects by reconstructing the region-expression correspondence, we investigate WREG from a novel perspective and enrich the prevailing pattern with self-knowledge distillation. Specifically, we propose a target-guided self-knowledge distillation approach that adopts the target prediction knowledge learned from the previous training iterations as the teacher to guide the subsequent training procedure. In order to avoid the misleading caused by the teacher knowledge with low prediction confidence, we present an uncertaintyaware knowledge refinement strategy to adaptively rectify the teacher knowledge by learning dynamic threshold values based on the model prediction uncertainty. To validate the proposed approach, we implement extensive experiments on three benchmark datasets, i.e., Ref Coco, RefCOCO+, and RefCOCOg. Our approach achieves new state-of-the-art results on several splits of the benchmark datasets, showcasing the advantage of the proposed framework for WREG. The implementation codes and trained models are available at: https://github.com/dami23IWREG.sar_KD. Jinpeng Mi, Zhiqian Chen, Jianwei Zhang 0001 |
IROS | 3 |
| 2023 | PoseFusion: Robust Object-in-Hand Pose Estimation with SelectLSTMabstractAccurate estimation of the relative pose between an object and a robot hand is critical for many manipulation tasks. However, most of the existing object-in-hand pose datasets use two-finger grippers and also assume that the object remains fixed in the hand without any relative movements, which is not representative of real-world scenarios. To address this issue, a 6D object-in-hand pose dataset is proposed using a teleoperation method with an anthropomorphic Shadow Dexterous hand. Our dataset comprises RGB-D images, proprioception and tactile data, covering diverse grasping poses, finger contact states, and object occlusions. To overcome the significant hand occlusion and limited tactile sensor contact in real-world scenarios, we propose PoseFusion, a hybrid multi-modal fusion approach that integrates the information from visual and tactile perception channels. PoseFusion generates three candidate object poses from three estimators (tactile only, visual only, and visuo-tactile fusion), which are then filtered by a SelectLSTM network to select the optimal pose, avoiding inferior fusion poses resulting from modality collapse. Extensive experiments demonstrate the robustness and advantages of our framework. All data and codes are available on the project website: https://elevenjiang1.github.io/ObjectlnHand-Dataset/. Yuyang Tu, Junnan Jiang, Shuang Li 0014, Norman Hendrich, Miao Li 0002, Jianwei Zhang 0001 |
IROS | 6 |
| 2023 | Efficient Human Motion Reconstruction from Monocular Videos with Physical Consistency LossabstractVision-only motion reconstruction from monocular videos often produces artifacts such as foot sliding and jittering. Existing physics-based methods typically either simplify the problem to focus solely on foot-ground contacts, or they reconstruct full-body contacts within a physics simulator, necessitating the solution of a time-consuming bilevel optimization problem. To overcome these limitations, we present an efficient gradient-based method for reconstructing complex human motions (including highly dynamic and acrobatic movements) with physical constraints. Our approach reformulates human motion dynamics through a differentiable physical consistency loss within an augmented search space that accounts both for contacts and camera alignment. This enables us to transform the motion reconstruction task into a single-level trajectory optimization problem. Experimental results demonstrate that our method can reconstruct complex human motions from real-world videos in minutes, which is substantially faster than previous approaches. Additionally, the reconstructed results show enhanced physical realism compared to existing methods. Philipp Ruppel, Yizhou Wang 0001, Norman Hendrich, Jianwei Zhang 0001 |
SIGGRAPH Asia | 6 |
| 2023 | Cross-domain video action recognition via adaptive gradual learning
Zhenwei Bao, Jinpeng Mi, Yan Gan, Mao Ye 0001, Jianwei Zhang 0001 |
Neurocomputing | 6 |
| 2023 | A survey on machine learning from few samples
Jiang Lu, Pinghua Gong, Jieping Ye, Jianwei Zhang 0001, Changshui Zhang |
Pattern Recognit. | 4 |
| 2023 | Where you edit is what you get: Text-guided image editing with region-based attentionabstractLeveraging the abundant knowledge learned from pre-trained multi-modal models like CLIP has recently proved to be effective for text-guided image editing. Though convincing results have been made when combining the image generator StyleGAN with CLIP, most methods need to train separate models for different prompts, and irrelevant regions are often changed after editing due to the lack of spatial disentanglement. We propose a novel framework that can edit different images according to different prompts in one model. Besides, an innovative region-based spatial attention mechanism is adopted to explicitly guarantee the locality of editing. Experiments mainly in the face domain verify the feasibility of our framework and show that when multi-text editing and local editing are accomplishable, our method can complete practical applications like sequential editing and regional style transfer. Changming Xiao, Xiaoqiang Xu, Jianwei Zhang 0001, Changshui Zhang |
Pattern Recognit. | 4 |
| 2023 | Multifingered Robot Hand Compliant Manipulation Based on Vision-Based Demonstration and Adaptive Force ControlabstractMultifingered hand dexterous manipulation is quite challenging in the domain of robotics. One remaining issue is how to achieve compliant behaviors. In this work, we propose a human-in-the-loop learning-control approach for acquiring compliant grasping and manipulation skills of a multifinger robot hand. This approach takes the depth image of the human hand as input and generates the desired force commands for the robot. The markerless vision-based teleoperation system is used for the task demonstration, and an end-to-end neural network model (i.e., TeachNet) is trained to map the pose of the human hand to the joint angles of the robot hand in real-time. To endow the robot hand with compliant human-like behaviors, an adaptive force control strategy is designed to predict the desired force control commands based on the pose difference between the robot hand and the human hand during the demonstration. The force controller is derived from a computational model of the biomimetic control strategy in human motor learning, which allows adapting the control variables (impedance and feedforward force) online during the execution of the reference joint angles. The simultaneous adaptation of the impedance and feedforward profiles enables the robot to interact with the environment compliantly. Our approach has been verified in both simulation and real-world task scenarios based on a multifingered robot hand, that is, the Shadow Hand, and has shown more reliable performances than the current widely used position control mode for obtaining compliant grasping and manipulation behaviors. Chao Zeng 0002, Shuang Li 0014, Zhaopeng Chen, Chenguang Yang 0001, Fuchun Sun 0001, Jianwei Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Event-Triggered Tracking Control Scheme for Quadrotors with External Disturbances: Theory and ValidationsabstractThis article studies the tracking control of a quadrotor unmanned aerial vehicle (UAV) under time-varying external disturbances. An event-triggered sliding mode control (SMC) strategy is proposed by introducing a new triggering condition form of desired trajectory, quadrotor position, and velocity. In the sense of Lyapunov theory, the stability of the entire closed-loop control system is analyzed, and it is proved that the tracking error is adjusted to an adjustable set around zero. We show that the Zeno phenomenon can be avoided; that is, a positive minimum inter-event time is assured. One of the salient features of the proposed strategy is that it can reduce the update frequency of the control efforts, thereby ensuring desirable tracking performance under limited communication bandwidth. Comparative simulation and experimental results are provided to show the efficacy of our framework. Gang Wang 0024, Yunfeng Ji, Qingdu Li, Jianwei Zhang 0001, Yantao Shen 0001, Peng Li 0019 |
ICRA | 5 |
| 2022 | Learning Friction Model for Magnet-Actuated Tethered Capsule RobotabstractThe potential diagnostic applications of magnet-actuated capsules have been greatly increased in recent years. For most of these potential applications, accurate position control of the capsule have been highly demanding. However, the friction between the robot and the environment as well as the drag force from the tether play a significant role during the motion control of the capsule. Moreover, these forces especially the friction force are typically hard to model beforehand. In this paper, we first designed a magnet-actuated tethered capsule robot, where the driving magnet is mounted on the end of a robotic arm. Then, we proposed a learning-based approach to model the friction force between the capsule and the environment, with the goal of increasing the control accuracy of the whole system. Finally, several real robot experiments are demonstrated to showcase the effectiveness of our proposed approach. Yuyang Tu, Yuchen He 0004, Xutian Deng, Ziwei Lei, Jianwei Zhang 0001, Miao Li 0002 |
ICRA | 6 |
| 2022 | SOF-UNet: SAR and Optical Fusion Unet for Land Cover ClassificationabstractWe propose a SAR and Optical Fusion Network based on the UNet framework (SOF-UNet) for multi-modal land cover classification. The two-stream SOF-UNet consists of three parts: two encoders to extract features, a sharing decoder to upsample the feature maps and specially designed skip connections to fuse multi-modal features. The qualitative and quantitative experimental results show that SOF-UNet has a promising capability to identify different land cover classes and can retain fine details in the prediction maps. Symmetric Cross Entropy (SCE) loss is also verified useful in this framework. Di Zhang 0021, Martin Gade, Jianwei Zhang 0001 |
IGARSS | 3 |
| 2022 | Learning 6-DoF Task-oriented Grasp Detection via Implicit Estimation and Visual AffordanceabstractCurrently, task-oriented grasp detection approaches are mostly based on pixel-level affordance detection and semantic segmentation. These pixel-level approaches heavily rely on the accuracy of a 2D affordance mask, and the generated grasp candidates are restricted to a small workspace. To mitigate these limitations, we firstly construct a novel affordance-based grasp dataset and propose a 6-DoF task-oriented grasp detection framework, which takes the observed object point cloud as input and predicts diverse 6-DoF grasp poses for different tasks. Specifically, our implicit estimation network and visual affordance network in this framework could directly predict coarse grasp candidates, and corresponding 3D affordance heatmap for each potential task, respectively. Furthermore, the grasping scores from coarse grasps are combined with heatmap values to generate more accurate and finer candidates. Our proposed framework shows significant improvements compared to baselines for existing and novel objects on our simulation dataset. Although our framework is trained based on the simulated objects and environment, the final generated grasp candidates can be accurately and stably executed in the real robot experiments when the object is randomly placed on a support surface. Hongzhuo Liang, Zhaopeng Chen, Fuchun Sun 0001, Jianwei Zhang 0001 |
IROS | 5 |
| 2022 | Bipedal Walking on Humanoid Robots Through Parameter Optimization
Marc Bestmann, Jianwei Zhang 0001 |
RoboCup | 2 |
| 2022 | Semantic consistency learning on manifold for source data-free unsupervised domain adaptation
Song Tang 0001, Yan Zou, Jianzhi Lyu, Mao Ye 0001, Shouming Zhong, Jianwei Zhang 0001 |
Neural Networks | 8 |
| 2021 | CloudAAE: Learning 6D Object Pose Regression with On-line Data Synthesis on Point CloudsabstractIt is often desired to train 6D pose estimation systems on synthetic data because manual annotation is expensive. However, due to the large domain gap between the synthetic and real images, synthesizing color images is expensive. In contrast, this domain gap is considerably smaller and easier to fill for depth information. In this work, we present a system that regresses 6D object pose from depth information represented by point clouds, and a lightweight data synthesis pipeline that creates synthetic point cloud segments for training. We use an augmented autoencoder (AAE) for learning a latent code that encodes 6D object pose information for pose regression. The data synthesis pipeline only requires texture-less 3D object models and desired viewpoints, and it is cheap in terms of both time and hardware storage. Our data synthesis process is up to three orders of magnitude faster than commonly applied approaches that render RGB image data. We show the effectiveness of our system on the LineMOD, LineMOD Occlusion, and YCB Video datasets. The implementation of our system is available at: https://github.com/GeeeG/CloudAAE. Mikko Lauri, Xiaolin Hu 0001, Jianwei Zhang 0001, Simone Frintrop |
ICRA | 4 |
| 2021 | Proactive Action Visual Residual Reinforcement Learning for Contact-Rich Tasks Using a Torque-Controlled RobotabstractContact-rich manipulation tasks are commonly found in modern manufacturing settings. However, manually designing a robot controller is considered hard for traditional control methods as the controller requires an effective combination of modalities and vastly different characteristics. In this paper, we first consider incorporating operational space visual and haptic information into a reinforcement learning (RL) method to solve the target uncertainty problems in unstructured environments. Moreover, we propose a novel idea of introducing a proactive action to solve a partially observable Markov decision process (POMDP) problem. With these two ideas, our method can either adapt to reasonable variations in unstructured environments or improve the sample efficiency of policy learning. We evaluated our method on a task that involved inserting a random-access memory (RAM) using a torque-controlled robot and tested the success rates of different baselines used in the traditional methods. We proved that our method is robust and can tolerate environmental variations. Yunlei Shi, Zhaopeng Chen, Sebastian Riedel 0002, Chunhui Gao, Jianwei Zhang 0001 |
ICRA | 8 |
| 2021 | SOFNet: SAR-Optical Fusion Network for Land Cover ClassificationabstractThe objective of this research is to realize automatic land cover classification from synthetic aperture radar (SAR) and multispectral remote sensing imagery. We develop a SAR-optical fusion network (SOFNet) with the symmetric cross entropy (SCE) loss to utilize both the SAR and optical information in a novel deep neural network. The proposed framework has been trained on the public SEN12MS dataset and tested on the 2020 IEEE-GRSS Data Fusion Contest (DFC2020) dataset. Experimental results show that our approach takes full advantage of multimodal information and outperforms the state-of-the-art convolutional architectures. Di Zhang 0021, Martin Gade, Jianwei Zhang 0001 |
IGARSS | 3 |
| 2021 | Model Adaptation through Hypothesis Transfer with Gradual Knowledge DistillationabstractThe ability to adapt their perception to changing environments is a core characterization of intelligent robots. At present, Unsupervised Domain Adaptation (UDA) methods are used to address this problem where the adaptation task is formulated as a transfer problem from a well-described scenario (source domain) to a new scenario (target domain). In order to implement the domain adaptation, these methods require access to the source data for achieving the distribution matching between both domains. However, in many real-world applications, the source data is inaccessible and only a source model pre-trained on the source domain is available during the transfer process. Therefore, the traditional UDA methods cannot support the challenging setting. This paper developed a new hypothesis transfer method to achieve model adaptation with gradual knowledge distillation. Specifically, we first prepare a source model through training a deep network on the labeled source domain by supervised learning. Then, we transfer the source model to the unlabeled target domain by self-training. To implement gradual knowledge distillation, we sliced the self-training into several epochs and then used the soft pseudo-labels from the latest epoch to guide the current epoch. In this process, the soft labels were generated by a semantic fusion on a proposed geometry of the neighborhood. To regulate the self-training, we developed a new objective constructed on the neighborhood. Experiments on three benchmarks have confirmed the state-of-the-art results of our method. Song Tang 0001, Yuji Shi, Zhiyuan Ma 0001, Jianzhi Lyu, Qingdu Li, Jianwei Zhang 0001 |
IROS | 7 |
| 2021 | A Low-Cost Modular System of Customizable, Versatile, and Flexible Tactile Sensor ArraysabstractThe key role of tactile sensing for human grasping and manipulation is widely acknowledged, but most industrial robot grippers and even multi-fingered hands are still designed and used without any tactile sensors. While the basic design principles for resistive or capacitive sensors are well known, several factors keep tactile sensing from large-scale deployment — high sensor costs, short lifespan, poor reliability, difficult production processes, a lack of suitable software and tools for system integration, and the unique requirement for tactile sensors to conform to application-specific shapes.In this work, we describe a very simple but efficient approach to design low-cost resistive matrix sensors, where sensor layout and geometry, taxel-size, and measurement sensitivity can be customized over a wide range. Sensor assembly needs nothing more than a hobby cutting plotter for precise cutting of aluminum tape and Velostat foils, as well as adhesive plastic tape. Our electronics combines transimpedance amplifiers with common Arduino microcontrollers, supporting standard communication protocols, and using either cabled or wireless data transfer to the host. We present three different application examples and sketch our ROS software for sensor calibration and visualization. All parts of our project, including detailed building instructions, bill-of-materials, electronics, and firmware are available open-source. Niklas Fiedler, Philipp Ruppel, Yannick Jonetzko, Norman Hendrich, Jianwei Zhang 0001 |
IROS | 5 |
| 2021 | Model-Based Trajectory Prediction and Hitting Velocity Control for a New Table Tennis RobotabstractCurrently, most table tennis robots concentrate on the canonical position control problem while ignoring the actual velocity control requirements. In this paper, we consider these requirements and propose a new table tennis robot framework. First, a tailor-made mechanical structure is designed such that the robot can reach large workspaces. Thereafter, in the table tennis trajectory prediction process, a clustering algorithm is introduced to screen the heterogeneous predicted hitting points and filter the invalid ones, thereby significantly improving the prediction accuracy. By using quintic polynomial trajectory planning, smooth and stable high-speed control of the robot hitting motion can be obtained. Finally, a position-based strategy and a velocity-based strategy are devised for returning the table tennis. Extensive experiments demonstrate that the accuracy of the ball's trajectory prediction algorithm is more than 92%. The success rate of returning the ball exceeds 95% at the ball velocity of 3-7 m/s, and the velocity-based strategy performs better compared with the position-based approach at the ball velocity of 7-9 m/s. Yunfeng Ji, Yue Mao, Gang Wang 0024, Qingdu Li, Jianwei Zhang 0001 |
IROS | 7 |
| 2021 | Combining Learning from Demonstration with Learning by Exploration to Facilitate Contact-Rich TasksabstractCollaborative robots are expected to work alongside humans and directly replace human workers in some cases, thus effectively responding to rapid changes in assembly lines. Current methods for programming contact-rich tasks, particularly in heavily constrained spaces, tend to be fairly inefficient. Therefore, faster and more intuitive approaches are urgently required for robot teaching. This study focuses on combining visual servoing-based learning from demonstration (LfD) and force-based learning by exploration (LbE) to enable the fast and intuitive programming of contact-rich tasks with minimal user efforts. Two learning approaches were developed and integrated into a framework, one relying on human-to-robot motion mapping (visual servoing approach) and the other relying on force-based reinforcement learning. The developed framework implements the noncontact demonstration teaching method based on the visual servoing approach and optimizes the demonstrated robot target positions according to the detected contact state. The developed framework is compared with two most commonly used baseline techniques, i.e., teach pendant-based teaching and hand-guiding teaching. Furthermore, the efficiency and reliability of the framework are validated via comparison experiments involving the teaching and execution of contact-rich tasks. The proposed framework shows the best performance in terms of the teaching time, execution success rate, risk of damage, and ease of use. Yunlei Shi, Zhaopeng Chen, Yansong Wu, Dimitri Henkel, Sebastian Riedel 0002, Jianwei Zhang 0001 |
IROS | 8 |
| 2021 | Learning compliant grasping and manipulation by teleoperation with adaptive force controlabstractIn this work, we focus on improving the robot’s dexterous capability by exploiting visual sensing and adaptive force control. TeachNet, a vision-based teleoperation learning framework, is exploited to map human hand postures to a multi-fingered robot hand. We augment TeachNet, which is originally based on an imprecise kinematic mapping and position-only servoing, with a biomimetic learning-based compliance control algorithm for dexterous manipulation tasks. This compliance controller takes the mapped robotic joint angles from TeachNet as the desired goal, computes the desired joint torques. It is derived from a computational model of the biomimetic control strategy in human motor learning, which allows adapting the control variables (impedance and feedforward force) online during the execution of the reference joint angle trajectories. The simultaneous adaptation of the impedance and feedforward profiles enables the robot to interact with the environment in a compliant manner. Our approach has been verified in multiple tasks in physics simulation, i.e., grasping, opening-a-door, turning-a-cap, and touching-a-mouse, and has shown more reliable performances than the existing position control and the fixed-gain-based force control approaches. Chao Zeng 0002, Shuang Li 0014, Yiming Jiang 0001, Qiang Li 0001, Zhaopeng Chen, Chenguang Yang 0001, Jianwei Zhang 0001 |
IROS | 7 |
| 2021 | Encode-decode network with fully connected CRF for dynamic objects detection and static maps reconstruction
Bingwei He, Mingzhu Zhu, Jianwei Zhang 0001 |
Signal Process. Image Commun. | 5 |
| 2021 | Detection and Segmentation of Unlearned Objects in Unknown EnvironmentabstractDetecting and segmenting unlearned objects in unknown environment is a very important visual perception ability to enhance industrial intelligence. In this article, we present a novel conditional random field model integrating unimodal and cross-modal terms for detecting and segmenting object instances without knowing their categories and without sampling extra proposals. This model takes a paired image and point cloud as input, from which we first develop a set of novel category-independent features to distinguish objects. Then, a set of unary, pairwise, and higher order potentials are designed according to these category-independent features, and the cross-modal potential is introduced as a novel global constraints to keep the spatial consistency in both 2-D and 3-D modalities. In this novel model, the unlearned object detection and segmentation is treated as the process of pixel labeling. Thus, adjacent or occlusion object instances can also be separated efficiently from a labeled map. By comparison with the baseline methods, experimental results on a public RGB+D dataset show that the proposed model can obtain better performance with improved precision and recall rate. Moreover, we use the proposed method in a real industrial scene and achieve satisfactory performance. Jianhua Zhang 0002, Jingbo Chen, Shengyong Chen, Zhenhua Wang 0003, Jianwei Zhang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2020 | Domain Optimization for Hierarchical Planning based on Set-Theory
Bernd Kast, Vincent Dietrich, Sebastian Albrecht 0001, Wendelin Feiten, Jianwei Zhang 0001 |
ICINCO | 5 |
| 2020 | Center-of-Mass-based Robust Grasp Planning for Unknown Objects Using Tactile-Visual SensorsabstractAn unstable grasp pose can lead to slip, thus an unstable grasp pose can be predicted by slip detection. A regrasp is required afterwards to correct the grasp pose in order to finish the task. In this work, we propose a novel regrasp planner with multi-sensor modules to plan grasp adjustments with the feedback from a slip detector. Then a regrasp planner is trained to estimate the location of center of mass, which helps robots find an optimal grasp pose. The dataset in this work consists of 1 025 slip experiments and 1 347 regrasps collected by one pair of tactile sensors, an RGB-D camera and one Franka Emika robot arm equipped with joint force/torque sensors. We show that our algorithm can successfully detect and classify the slip for 5 unknown test objects with an accuracy of 76.88% and a regrasp planner increases the grasp success rate by 31.0% compared to the state-of-the-art vision-based grasping algorithm. Zhaopeng Chen, Chunhui Gao, Jianwei Zhang 0001, Alois C. Knoll |
ICRA | 5 |
| 2020 | 6D Object Pose Regression via Supervised Learning on Point CloudsabstractThis paper addresses the task of estimating the 6 degrees of freedom pose of a known 3D object from depth information represented by a point cloud. Deep features learned by convolutional neural networks from color information have been the dominant features to be used for inferring object poses, while depth information receives much less attention. However, depth information contains rich geometric information of the object shape, which is important for inferring the object pose. We use depth information represented by point clouds as the input to both deep networks and geometry-based pose refinement and use separate networks for rotation and translation regression. We argue that the axis-angle representation is a suitable rotation representation for deep learning, and use a geodesic loss function for rotation regression. Ablation studies show that these design choices outperform alternatives such as the quaternion representation and L2 loss, or regressing translation and rotation with the same network. Our simple yet effective approach clearly outperforms state-of-the-art methods on the YCB-video dataset. Mikko Lauri, Xiaolin Hu 0001, Jianwei Zhang 0001, Simone Frintrop |
ICRA | 5 |
| 2020 | SAR Eddy Detection Using Mask-RCNN and Edge EnhancementabstractThe objective of this research is to detect ocean eddies automatically on Synthetic Aperture Radar (SAR) images. We develop a new approach using Mask Region-based Convolutional Neural Networks (Mask R-CNN) and edge enhancement. First, we use Canny edge detector to extract a wide range of edges in SAR images. Then we put both the edge detection results and the corresponding original images into a Mask R-CNN based model for learning, thereby strengthening edge information. The proposed framework has been trained on a sample dataset of Sentinel-1A SAR-C imagery of the Western Mediterranean Sea. Experimental results revealed that the proposed method improved the performance by 2.3% on the MS COCO metrics compared to the method without edge enhancement. Di Zhang 0021, Martin Gade, Jianwei Zhang 0001 |
IGARSS | 3 |
| 2020 | Self-Adapting Recurrent Models for Object Pushing from Learning in SimulationabstractPlanar pushing remains a challenging research topic, where building the dynamic model of the interaction is the core issue. Even an accurate analytical dynamic model is inherently unstable because physics parameters such as inertia and friction can only be approximated. Data-driven models usually rely on large amounts of training data, but data collection is time consuming when working with real robots.In this paper, we collect all training data in a physics simulator and build an LSTM-based model to fit the pushing dynamics. Domain Randomization is applied to capture the pushing trajectories of a generalized class of objects. When executed on the real robot, the trained recursive model adapts to the tracked object's real dynamics within a few steps. We propose the algorithm Recurrent Model Predictive Path Integral (RMPPI) as a variation of the original MPPI approach, employing state-dependent recurrent models. As a comparison, we also train a Deep Deterministic Policy Gradient (DDPG) network as a model-free baseline, which is also used as the action generator in the data collection phase. During policy training, Hindsight Experience Replay is used to improve exploration efficiency. Pushing experiments on our UR5 platform demonstrate the model's adaptability and the effectiveness of the proposed framework. Michael Görner, Philipp Ruppel, Hongzhuo Liang, Norman Hendrich, Jianwei Zhang 0001 |
IROS | 6 |
| 2020 | Learning Local Planners for Human-aware Navigation in Indoor EnvironmentsabstractEstablished indoor robot navigation frameworks build on the separation between global and local planners. Whereas global planners rely on traditional graph search algorithms, local planners are expected to handle driving dynamics and resolve minor conflicts. We present a system to train neural-network policies for such a local planner component, explicitly accounting for humans navigating the space. DRL-agents are trained in randomized virtual 2D environments with simulated human interaction. The trained agents can be deployed as a drop-in replacement for other local planners and significantly improve on traditional implementations. Performance is demonstrated on a MiR-100 transport robot. Ronja Güldenring, Michael Görner, Norman Hendrich, Niels Jul Jacobsen, Jianwei Zhang 0001 |
IROS | 5 |
| 2020 | A Mobile Robot Hand-Arm Teleoperation System by Vision and IMUabstractIn this paper, we present a multimodal mobile teleoperation system that consists of a novel vision-based hand pose regression network (Transteleop) and an IMU (inertial measurement units)-based arm tracking method. Transteleop observes the human hand through a low-cost depth camera and generates not only joint angles but also depth images of paired robot hand poses through an image-to-image translation process. A keypoint-based reconstruction loss explores the resemblance in appearance and anatomy between human and robotic hands and enriches the local features of reconstructed images. A wearable camera holder enables simultaneous hand-arm control and facilitates the mobility of the whole teleoperation system. Network evaluation results on a test dataset and a variety of complex manipulation tasks that go beyond simple pick-and-place operations show the efficiency and stability of our multimodal teleoperation system. Shuang Li 0014, Jiaxi Jiang, Philipp Ruppel, Hongzhuo Liang, Xiaojian Ma 0001, Norman Hendrich, Fuchun Sun 0001, Jianwei Zhang 0001 |
IROS | 8 |
| 2020 | Robust Robotic Pouring using Audition and HapticsabstractRobust and accurate estimation of liquid height lies as an essential part of pouring tasks for service robots. However, vision-based methods often fail in occluded conditions while audio-based methods cannot work well in a noisy environment. We instead propose a multimodal pouring network (MP-Net) that is able to robustly predict liquid height by conditioning on both audition and haptics input. MP-Net is trained on a self-collected multimodal pouring dataset. This dataset contains 300 robot pouring recordings with audio and force/torque measurements for three types of target containers. We also augment the audio data by inserting robot noise. We evaluated MP-Net on our collected dataset and a wide variety of robot experiments. Both network training results and robot experiments demonstrate that MP-Net is robust against noise and changes to the task and environment. Moreover, we further combine the predicted height and force data to estimate the shape of the target container. Hongzhuo Liang, Chuangchuang Zhou, Shuang Li 0014, Xiaojian Ma 0001, Norman Hendrich, Timo Gerkmann, Fuchun Sun 0001, Marcus Stoffel, Jianwei Zhang 0001 |
IROS | 9 |
| 2020 | Learning Object Manipulation with Dexterous Hand-Arm Systems from Human DemonstrationabstractWe present a novel learning and control framework that combines artificial neural networks with online trajectory optimization to learn dexterous manipulation skills from human demonstration and to transfer the learned behaviors to real robots. Humans can perform the demonstrations with their own hands and with real objects. An instrumented glove is used to record motions and tactile data. Our system learns neural control policies that generalize to modified object poses directly from limited amounts of demonstration data. Outputs from the neural policy network are combined at runtime with kinematic and dynamic safety and feasibility constraints as well as a learned regularizer to obtain commands for a real robot through online trajectory optimization. We test our approach on multiple tasks and robots. Philipp Ruppel, Jianwei Zhang 0001 |
IROS | 2 |
| 2020 | Improving Action Recognition Using Sequence Prediction LearningabstractSkeleton-based action recognition distinguishes human actions using the trajectories of skeleton joints, which can be a good representation of human behaviors. Conventional methods usually construct classifiers with hand-crafted or the learned features to recognize human actions. Different from constructing a direct action classifier for action recognition task, this paper attempts to identify human actions based on the development trends of behavior sequences. Specifically, we first utilize the memory neural network to construct action predictors for each kind of activity. These action predictors can then output the action trends at the next time step. According to the predictions of these action predictors at each time step and the removal rule, the poor predictors can be eliminated step by step, and the IDentity(ID) number of the last predictor left is considered as the label of the action sequence to be categorized. We compare the proposed action recognition algorithm using sequence prediction learning with other methods on two publicly available datasets. Our experimental results consistently demonstrate the feasibility and effectiveness of the suggested method. It also proves the importance of prediction learning for action recognition. Mao Ye 0001, Jianwei Zhang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2020 | Cutting Depth Monitoring Based on Milling Force for Robot-Assisted LaminectomyabstractGoal: In the context of robot-assisted laminectomy surgery, an analytical force model is introduced to guarantee procedural safety. The aim of the method is to intraoperatively monitor the cutting depth via modeling the milling status. Methods: The theoretical dynamic model for the surgical milling process is based on the flute geometry of the ball-end milling tool. A particle swarm optimization algorithm is exploited to calibrate the model using the local average force, and to validate it using the denoised dynamic force. A wear detection method based on the fast Fourier transform is proposed to determine the quality of the tool geometry and to avoid using worn tools, which may lead to imprecise and unsafe operations. Results: Milling experiments were performed on machined fresh bovine femur bones. The experimental results thus obtained from the mechanical model are in good accordance with the numerical model. The proposed method can monitor the current cutting depth with an accuracy of ±0.1 mm in regions located within the depth [0.8-1.2 mm], and ±0.2 mm within [1.2-1.6 mm]. Conclusion: The proposed model can successfully estimate the milling force and the cutting depth intraoperatively in experimental conditions. Significance: This approach has the potential to improve the safety of laminectomy operations in humans, and make it more accessible to younger surgeons by lowering the required manual skills threshold. Zhongliang Jiang, Xiaozhi Qi, Yu Sun 0018, Ying Hu 0001, Guillaume Zahnd, Jianwei Zhang 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2019 | A Hierarchical Planner based on Set-theoretic Models: Towards Automating the Automation for Autonomous Systems
Bernd Kast, Vincent Dietrich, Sebastian Albrecht 0001, Wendelin Feiten, Jianwei Zhang 0001 |
ICINCO (1) | 5 |
| 2019 | MoveIt! Task Constructor for Task-Level Motion PlanningabstractA lot of motion planning research in robotics focuses on efficient means to find trajectories between individual start and goal regions, but it remains challenging to specify and plan robotic manipulation actions which consist of multiple interdependent subtasks. The Task Constructor framework we present in this work provides a flexible and transparent way to define and plan such actions, enhancing the capabilities of the popular robotic manipulation framework MoveIt!.11The Task Constructor framework is publicly available at https://github.com/ros-planning/moveit_task_constructor Subproblems are solved in isolation in black-box planning stages and a common interface is used to pass solution hypotheses between stages. The framework enables the hierarchical organization of basic stages using containers, allowing for sequential as well as parallel compositions. The flexibility of the framework is illustrated in multiple scenarios performed on various robot platforms, including bimanual ones. Michael Görner, Robert Haschke, Helge J. Ritter, Jianwei Zhang 0001 |
ICRA | 4 |
| 2019 | Vision-based Teleoperation of Shadow Dexterous Hand using End-to-End Deep Neural NetworkabstractIn this paper, we present TeachNet, a novel neural network architecture for intuitive and markerless vision-based teleoperation of dexterous robotic hands. Robot joint angles are directly generated from depth images of the human hand that produce visually similar robot hand poses in an end-to-end fashion. The special structure of TeachNet, combined with a consistency loss function, handles the differences in appearance and anatomy between human and robotic hands. A synchronized human-robot training set is generated from an existing dataset of labeled depth images of the human hand and simulated depth images of a robotic hand. The final training set includes 400K pairwise depth images and joint angles of a Shadow C6 robotic hand. The network evaluation results verify the superiority of TeachNet, especially regarding the high-precision condition. Imitation experiments and grasp tasks teleoperated by novice users demonstrate that TeachNet is more reliable and faster than the state-of-the-art vision-based teleoperation method. Shuang Li 0014, Xiaojian Ma 0001, Hongzhuo Liang, Michael Görner, Philipp Ruppel, Bin Fang 0003, Fuchun Sun 0001, Jianwei Zhang 0001 |
ICRA | 8 |
| 2019 | PointNetGPD: Detecting Grasp Configurations from Point SetsabstractIn this paper, we propose an end-to-end grasp evaluation model to address the challenging problem of localizing robot grasp configurations directly from the point cloud. Compared to recent grasp evaluation metrics that are based on handcrafted depth features and a convolutional neural network (CNN), our proposed PointNetGPD is lightweight and can directly process the 3D point cloud that locates within the gripper for grasp evaluation. Taking the raw point cloud as input, our proposed grasp evaluation network can capture the complex geometric structure of the contact area between the gripper and the object even if the point cloud is very sparse. To further improve our proposed model, we generate a large-scale grasp dataset with 350k real point cloud and grasps with the YCB object set for training. The performance of the proposed model is quantitatively measured both in simulation and on robotic hardware. Experiments on object grasping and clutter removal show that our proposed model generalizes well to novel objects and outperforms state-of-the-art methods. Code and video are available at https://lianghongzhuo.github.io/PointNetGPD. Hongzhuo Liang, Xiaojian Ma 0001, Shuang Li 0014, Michael Görner, Song Tang 0001, Bin Fang 0003, Fuchun Sun 0001, Jianwei Zhang 0001 |
ICRA | 8 |
| 2019 | Image Captioning with Partially Rewarded Imitation LearningabstractCurrent state-of-the-art image captioning algorithms have achieved great progress via reinforcement learning or generative adversarial nets, with hand-craft metrics such as CIDEr as the reward for the former and signals from adversarial discriminative networks for the latter. Despite the high scores on metrics or improvement in diversity gained from the application of these methods, they suffer from distinction with human-written sentences and drop of ratings on metrics respectively.In this paper, we propose a novel training objective for image captioning that consists of two parts representing explicit and implicit knowledge respectively. Optimizing the new reward partially with imitation learning, we devise an algorithm in which the caption generator is trained to maximize the combination of CIDEr and predictions from adversarial discriminator. Experiments on MSCOCO dataset demonstrate that the proposed method can integrate the strengths of state-of-the-arts, producing more human-like captions while maintaining comparable performance on traditional metrics. Xintong Yu 0002, Tszhang Guo, Kun Fu 0002, Changshui Zhang, Jianwei Zhang 0001 |
IJCNN | 6 |
| 2019 | Making Sense of Audio Vibration for Liquid Height Estimation in Robotic PouringabstractIn this paper, we focus on the challenging perception problem in robotic pouring. Most of the existing approaches either leverage visual or haptic information. However, these techniques may suffer from poor generalization performances on opaque containers or concerning measuring precision. To tackle these drawbacks, we propose to make use of audio vibration sensing and design a deep neural network PouringNet to predict the liquid height from the audio fragment during the robotic pouring task. PouringNet is trained on our collected real-world pouring dataset with multimodal sensing data, which contains more than 3000 recordings of audio, force feedback, video and trajectory data of the human hand that performs the pouring task. Each record represents a complete pouring procedure. We conduct several evaluations on PouringNet with our dataset and robotic hardware. The results demonstrate that our PouringNet generalizes well across different liquid containers, positions of the audio receiver, initial liquid heights and types of liquid, and facilitates a more robust and accurate audio-based perception for robotic pouring. Hongzhuo Liang, Shuang Li 0014, Xiaojian Ma 0001, Norman Hendrich, Timo Gerkmann, Fuchun Sun 0001, Jianwei Zhang 0001 |
IROS | 7 |
| 2019 | Visual Domain Adaptation Exploiting Confidence-SamplesabstractDomain adaptation methods are used to address a problem, in which train scenario (source domain) and test scenario (target domain) are different. The existing methods mainly perform adaptation via reducing domain discrepancy from the view of a probability distribution. However, the idea of probability distribution matching always leads to a complex optimization process. Thereby these methods are difficult to apply in some scenario like online application or fast perception in dynamic environments. In this paper, we propose a new and simple domain adaptation method that utilizes confidence- samples to facilitate the classifier training on the target domain. Here, the confidence-samples are a subset of the target samples, and they have very credibly predicted labels. In order to detect the samples, a Category Similarity Collaborative Representation (CSCR) is first developed, by which the raw labels of all target samples are predicted using the smallest projection error according to the law of category. After this, the confidence score of the raw predicted labels is evaluated by the energy context information of CSCR. Finally, the target samples with a high confidence score are selected. Because of the linearity of CSCR, our method avoids complex optimization for matching the probability distribution. Empirical studies on a standard dataset demonstrate the advantages of our method. Song Tang 0001, Yunfeng Ji, Jianzhi Lyu, Jinpeng Mi, Qingdu Li, Jianwei Zhang 0001 |
IROS | 6 |
| 2019 | Robust High Accuracy Visual-Inertial-Laser SLAM SystemabstractIn recent years, many excellent works on visual-inertial SLAM and laser-based SLAM have been proposed. Although inertial measurement unit (IMU) significantly improve the motion estimate performance by reducing the impact of illumination variation or texture-less region on visual tracking, tracking failures occur when in such an environment for a long time. Similarly, when in structure-less environments, laser module will fail since lack of sufficient geometric features. Besides, motion estimation by moving lidar has the problem of distortion since range measurements are received continuously. To solve these problems, we propose a robust and high-accuracy visual-inertial-laser SLAM system. The system starts with a visual-inertial tightly-coupled method for motion estimation, followed by scan matching to further optimize the estimation and register point cloud on the map. Furthermore, we enable modules to be adjusted automatically and flexibly. That is, when one of these modules fails, the remaining modules will undertake the motion-tracking task. For further improving the accuracy, loop closure and proximity detection are implemented to eliminate drift accumulation. When loop or proximity is detected, we perform six degree-of-freedom (6-DOF) pose graph optimization to achieve the global consistency. The performance of our system is verified on public dataset, and the experimental results show that the proposed method achieves superior accuracy against other state-of-the-art algorithms. Zengyuan Wang, Jianhua Zhang 0002, Shengyong Chen, Conger Yuan, Jingqian Zhang, Jianwei Zhang 0001 |
IROS | 6 |
| 2019 | High-Frequency Multi Bus Servo and Sensor Communication Using the Dynamixel Protocol
Marc Bestmann, Jasper Güldenstein, Jianwei Zhang 0001 |
RoboCup | 3 |
| 2019 | Network Intrusion Feature Map Node Equalization Algorithm Based on Modified Variable Step-Size Constant ModulusabstractWhen the network is subject to intrusion and attack, the node output channel equalization will be affected, resulting in bit error and distortion in the output of network transmission symbols. In order to improve the anti-attack ability and equalization of network node, a network intrusion feature map node equalization algorithm based on modified variable step-size constant modulus blind equalization algorithm (MISO-VSS-MCMA) is proposed. In this algorithm, the node transmission channel model after network intrusion is constructed, and sequential processing is performed to intruded nodes with the variable structure feedback link control method. With diversity spread spectrum technology, the channel loss after network intrusion is compensated and the network intrusion map feature is extracted. According to the extracted feature amount, channel equalization processing is performed for the cost function with the MISO-VSS-MCMA method to reduce the damage of network intrusion to the channel. Simulation results show that in node transmission channel equalization after network intrusion, this algorithm can reduce the error bit rate of signal transmission in network, and provide a good ability of correcting phase deflection in the output constellation, thus avoiding the error bit distortion and channel damage caused by network intrusion to the signal with a good equalization effect. This algorithm provides stronger convergence and map concentration, which demonstrates that its anti-interference and signal recovery capabilities are better, so it improves the anti-attack ability of the network. Xiaolei Liu 0001, Jianwei Zhang 0001, Xiaosong Zhang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2019 | Detecting adversarial examples via prediction difference for deep neural networks
Qingjie Zhao, Xiaohui Kuang, Jianwei Zhang 0001, Yahong Han, Yu-an Tan 0001 |
Inf. Sci. | 5 |
| 2019 | Scene flow estimation by depth map upsampling and layer assignment for camera-LiDAR system
Bingwei He, Mingzhu Zhu, Jianwei Zhang 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | Boosting VLAD with weighted fusion of local descriptors for image retrieval
Qingjie Zhao, Jimmy T. Mbelwa, Song Tang 0001, Jianwei Zhang 0001 |
Multim. Tools Appl. | 6 |
| 2019 | Learning motion field of LiDAR point cloud with convolutional networks
Bingwei He, Mingzhu Zhu, Jianwei Zhang 0001 |
Pattern Recognit. Lett. | 5 |
| 2019 | Learning Physical Human-Robot Interaction With Coupled Cooperative Primitives for a Lower ExoskeletonabstractHuman-powered lower exoskeletons have received considerable interests from both academia and industry over the past decades, and encountered increasing applications in human locomotion assistance and strength augmentation. One of the most important aspects in those applications is to achieve robust control of lower exoskeletons, which, in the first place, requires the proactive modeling of human movement trajectories through physical human-robot interaction (pHRI). As a powerful representative tool for motion trajectories, dynamic movement primitives (DMP) have been used to model human movement trajectories. However, canonical DMP only offers a general representation of human movement trajectory and may neglects the interactive term, therefore it cannot be directly applied to lower exoskeletons which need to track human joint trajectories online, because different pilots have different trajectories and even same pilot might change his/her motion during walking. This paper presents a novel coupled cooperative primitive (CCP) strategy, which aims at modeling the motion trajectories online. Besides maintaining canonical motion primitives, we model the interaction term between the pilot and exoskeletons through impedance models, and propose a reinforcement learning method based on policy improvement and path integrals (PI2) to learn the parameters online. Experimental results on both a single degree-of-freedom platform and a HUman-powered Augmentation Lower EXoskeleton (HUALEX) system demonstrate the advantages of our proposed CCP scheme. Rui Huang 0008, Hong Cheng 0002, Jing Qiu 0004, Jianwei Zhang 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2019 | Memetic Evolution for Generic Full-Body Inverse Kinematics in Robotics and AnimationabstractIn this paper, a novel and fast memetic evolutionary algorithm is presented which can solve fully constrained generic inverse kinematics with multiple end effectors and goal objectives, leaving high flexibility for the design of custom cost functions. The algorithm utilizes a hybridization of evolutionary and swarm optimization, combined with the limited-memory-Broyden-Fletcher-Goldfarb-Shanno with bound constraints algorithm for gradient-based optimization. Accurate solutions can be found in real-time and suboptimal extrema are robustly avoided, scaling well even for greatly higher degree of freedom. The algorithm provides a general framework for bounded continuous optimization which only requires two parameters for the number of individuals and elites to be set, and supports adding additional goals and constraints for inverse kinematics, such as minimal displacement between solutions, collision avoidance, or functional joint relations. Experimental results on several industrial and anthropomorphic robots as well as on virtual characters demonstrate the algorithm to be applicable for solving complex kinematic postures for different challenging tasks in robotics, human-robot interaction and character animation, including dexterous object manipulation, collision-free full-body motion, as well as animation post-processing for video games and films. Implementations are made available for Unity3D and robot operating system. Sebastian Starke, Norman Hendrich, Jianwei Zhang 0001 |
IEEE Trans. Evol. Comput. | 3 |
| 2019 | Topic-Oriented Image Captioning Based on Order-EmbeddingabstractWe present an image captioning framework that generates captions under a given topic. The topic candidates are extracted from the caption corpus. A given image's topics are then selected from these candidates by a CNN-based multi-label classifier. The input to the caption generation model is an image-topic pair, and the output is a caption of the image. For this purpose, a cross-modal embedding method is learned for the images, topics, and captions. In the proposed framework, the topic, caption, and image are organized in a hierarchical structure, which is preserved in the embedding space by using the order-embedding method. The caption embedding is upper bounded by the corresponding image embedding and lower bounded by the topic embedding. The lower bound pushes the images and captions about the same topic closer together in the embedding space. A bidirectional caption-image retrieval task is conducted on the learned embedding space and achieves the state-of-the-art performance on the MS-COCO and Flickr30K datasets, demonstrating the effectiveness of the embedding method. To generate a caption for an image, an embedding vector is sampled from the region bounded by the embeddings of the image and the topic, then a language model decodes it to a sentence as the output. The lower bound set by the topic shrinks the output space of the language model, which may help the model to learn to match images and captions better. Experiments on the image captioning task on the MS-COCO and Flickr30K datasets validate the usefulness of this framework by showing that the different given topics can lead to different captions describing specific aspects of the given image and that the quality of generated captions is higher than the control model without a topic as input. In addition, the proposed method is competitive with many state-of-the-art methods in terms of standard evaluation metrics. Niange Yu, Xiaolin Hu 0001, Binheng Song, Jian Yang 0003, Jianwei Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Weighted two-step aggregated VLAD for image retrieval
Qingjie Zhao, Jimmy T. Mbelwa, Song Tang 0001, Jianwei Zhang 0001 |
Vis. Comput. | 5 |
| 2018 | Cost Functions to Specify Full-Body Motion and Multi-Goal Manipulation TasksabstractWhile the problem of inverse kinematics on serial kinematic chains is well researched, solving motion tasks quickly on more complex robots remains an open problem. Examples include dual-arm manipulation, grasping with multi-finger hands, and full-body motion generation for humanoids. In this paper, we introduce an open-source software package for ROS and MoveIt! that solves inverse kinematics and motion tasks on robots with arbitrary kinematic trees. The underlying memetic algorithm integrates evolutionary optimization, particle swarm optimization, and gradient methods. The optimization respects joint limits, effectively avoids local minima, and achieves fast convergence to accurate solutions. More importantly, the overall motion goal is specified using a set of weighted sub-goals, providing great flexibility and control of secondary objectives. Several application examples demonstrate how to combine the predefined sub-goals to achieve complex motion tasks. Philipp Ruppel, Norman Hendrich, Sebastian Starke, Jianwei Zhang 0001 |
ICRA | 4 |
| 2018 | Comparison of Multimodal Heading and Pointing Gestures for Co-Located Mixed Reality Human-Robot InteractionabstractMixed reality (MR)opens up new vistas for human-robot interaction (HRI)scenarios in which a human operator can control and collaborate with co-located robots. For instance, when using a see-through head-mounted-display (HMD)such as the Microsoft HoloLens, the operator can see the real robots and additional virtual information can be superimposed over the real-world view to improve security, acceptability and predictability in HRI situations. In particular, previewing potential robot actions in-situ before they are executed has enormous potential to reduce the risks of damaging the system or injuring the human operator. In this paper, we introduce the concept and implementation of such an MR human-robot collaboration system in which a human can intuitively and naturally control a co-located industrial robot arm for pick-and-place tasks. In addition, we compared two different, multimodal HRI techniques to select the pick location on a target object using (i)head orientation (aka heading)or (ii)pointing, both in combination with speech. The results show that heading-based interaction techniques are more precise, require less time and are perceived as less physically, temporally and mentally demanding for MR-based pick-and-place scenarios. We confirmed these results in an additional usability study in a delivery-service task with a multi-robot system. The developed MR interface shows a preview of the current robot programming to the operator, e. g., pick selection or trajectory. The findings provide important implications for the design of future MR setups. Dennis Krupke, Frank Steinicke, Paul Lubos, Yannick Jonetzko, Michael Görner, Jianwei Zhang 0001 |
IROS | 6 |
| 2018 | Static map reconstruction and dynamic object tracking for a camera and laser scanner systemabstractThe vision‐based mobile robot's simultaneous localisation and mapping and navigation capability in dynamic environments are highly problematic elements of robot vision applications. The goal of this study is to reconstruct a static map and track the dynamic object for a camera and laser scanner system. An improved automatic calibration is designed to merge image and laser point clouds. Then, the fusion data is exploited to detect the slowly moved object and reconstruct static map. Tracking‐by‐detection requires the correct assignment of noisy detection results to object trajectories. In the proposed method, occluded regions are combined 3D motion models with object appearance to manage difficulties in crowded scenes. The proposed method was validated by experimental results gathered in a real environment and on publicly available data. Bingwei He, Jianwei Zhang 0001 |
IET Comput. Vis. | 4 |
| 2018 | Scene flow for 3D laser scanner and camera systemabstractEstimating the three‐dimensional (3D) motion from sparse laser point clouds is a highly challenging endeavour facing computer and robotic vision engineers. In this study, a novel method is proposed for robustly estimating the scene flow from a laser scanner assisted by a camera. Conditional random field (CRF) is constructed by a spatial structure of point clouds, the energy of which is minimised by a synchronous calibrated image. With the high frame rate of a laser scanner, the authors’ method allows for estimating the potential motion field as the CRF label. The authors ran an experiment on a public dataset to demonstrate that their method can accurately estimate rigid motion in outdoor scenes. They also tested the method on a laser scanner and omni‐directional camera system to find that it also accurately estimates the rigid and semi‐rigid motion of objects in a controlled indoor environment. Bingwei He, Jianwei Zhang 0001 |
IET Image Process. | 4 |
| 2018 | Hierarchical learning control with physical human-exoskeleton interaction
Rui Huang 0008, Hong Cheng 0002, Hongliang Guo 0001, XiChuan Lin, Jianwei Zhang 0001 |
Inf. Sci. | 5 |
| 2017 | A memetic evolutionary algorithm for real-time articulated kinematic motionabstractSolving kinematic motion is a challenging field of research which is relevant for various applications in character animation and robotics. This paper presents a novel and fast hybrid evolutionary algorithm for inverse kinematics which can handle fully constrained and highly articulated geometries with multiple end effectors and individual objectives. Several experiments on the 42 DoF human body mannequin and other kinematic models demonstrate a robust multimodal and multi-objective optimisation, and the ability to evolve accurate solutions in real-time while offering maximum flexibility for the design of custom cost functions. Sebastian Starke, Norman Hendrich, Jianwei Zhang 0001 |
CEC | 3 |
| 2017 | Saliency-guided adaptive seeding for supervoxel segmentationabstractWe propose a new saliency-guided method for generating supervoxels in 3D space. Rather than using an evenly distributed spatial seeding procedure, our method uses visual saliency to guide the process of supervoxel generation. This results in densely distributed, small, and precise supervoxels in salient regions which often contain objects, and larger supervoxels in less salient regions that often correspond to background. Our approach largely improves the quality of the resulting supervoxel segmentation in terms of boundary recall and under-segmentation error on publicly available benchmarks. Mikko Lauri, Jianwei Zhang 0001, Simone Frintrop |
IROS | 3 |
| 2017 | A model of vertebral motion and key point recognition of drilling with force in robot-assisted spinal surgeryabstractPedicle drilling is a crucial and high-risk process in spinal surgery. Due to the respiration and cardiac cycle, the position of spine would fluctuate during operations, which result in an increase of the difficulty in state recognition of pedicle drilling. To guarantee the safety and validity, a model-based compensation method is proposed in this paper. To build the empirical model of vertebral motion, vertebral displacement and tidal volume (Tv) signals are collected from volunteers. To rule out disturbances in original signal, FFT and wavelet transform (DWT) are used to process the experimental signals. In order to select the apt basis for different signals, the root mean square error decision-making method (RMSE-DMM) is introduced. When the filtered vertebral displacement signal is obtained, the particle swarm optimization (PSO) algorithm is used to figure out the empirical model. The robot assisted systems (RAS) can easily compensate vertebral fluctuation based on the empirical model. Due to the goal of pedicle drilling is to drill a hole from surface of first cortical layer to the interior of second cortical layer, a new key point recognition algorithm, based on force, proposed in this paper. To verify the effectiveness of compensation and the recognition algorithm, 3 sets of comparison experiments are carried out. And the results of experiment show the compensation method and new key point recognition algorithm perform effectively. Zhongliang Jiang, Yu Sun 0018, Shijia Zhao, Ying Hu 0001, Jianwei Zhang 0001 |
IROS | 5 |
| 2017 | Evolutionary multi-objective inverse kinematics on highly articulated and humanoid robotsabstractWhile solving inverse kinematics on serial kinematic chains is well researched, many methods still seem rather limited in jointly handling more complex geometries, including dexterous multi-finger hands or humanoid robots. In particular, object manipulation and motion tasks would benefit from the ability to define intermediate goals along the kinematic chains, such as an elbow position or wrist orientation. In this paper, we propose a fast hybrid evolutionary approach that is capable of solving inverse kinematics for multiple end effectors simultaneously, leaving high flexibility for specifying full-body postures with different objectives. Accurate solutions can be found in real-time and suboptimal extrema are robustly avoided. Our experimental results on the NASA Valkyrie and Shadow Dexterous Hand demonstrate that the algorithm is fast and can be efficiently applied for different robotic tasks which require flexible control of fully-constrained geometries. Sebastian Starke, Norman Hendrich, Dennis Krupke, Jianwei Zhang 0001 |
IROS | 4 |
| 2016 | Multimodal Object Recognition and Categorisation by Interactive Behaviours
Haojun Guan, Jianwei Zhang 0001 |
CogSci | 2 |
| 2016 | Experience Representation of Artificial Cognitive System in Interaction with Real World
Zhen Deng, Jianwei Zhang 0001, Ying Hu 0001 |
CogSci | 3 |
| 2016 | A simple 2D straight-leg passive dynamic walking model without foot-scuffing problemabstractThis paper presents a simple 2D passive dynamic walking model with straight legs based on a novel hip joint, called T-joint. The model directly solves the common foot-scuffing problem in straight-legged walkers without introducing any new degree of freedom or additional motion phase, which is unavoidable in kneed robots. The mathematic model of the new walker is a 4D hybrid dynamical system, which can be reduced to a 3D Poincaré map. A stable period-1 walking gait is found by searching the map with a proposed algorithm. Numerical studies show that the model performs well for disturbances on initial conditions and also on a large slope. The stable walking gait is verified by a virtual prototype in ADAMS. This study has been successfully used to make an efficient walking robot called Xingzhe, which has set a new world record for the longest distance after 134km non-stoping walking. Qingdu Li, Jianwei Zhang 0001 |
IROS | 4 |
| 2016 | Vision-based real-time 3D mapping for UAV with laser sensorabstractReal-time 3D mapping with MAV (Micro Aerial Vehicle) in GPS-denied environment is a challenging problem. In this paper, we present an effective vision-based 3D mapping system with 2D laser-scanner. All algorithms necessary for this system are on-board. In this system, two cameras work together with the laser-scanner for motion estimation. The distance of the points detected by laser-scanner are transformed and treated as the depth of image features, which improves the robustness and accuracy of the pose estimation. The output of visual odometry is used as an initial pose in the Iterative Closest Point (ICP) algorithm and the motion trajectory is optimized by the registration result. We finally get the MAV's state by fusing IMU with the pose estimation from mapping process. This method maximizes the utility of the point clouds information and overcomes the scale problem of lacking depth information in the monocular visual odometry. The results of the experiments prove that this method has good characteristics in real-time and accuracy. Jinqiao Shi, Bingwei He, Jianwei Zhang 0001 |
IROS | 4 |
| 2016 | Immersive remote grasping: realtime gripper control by a heterogenous robot control systemabstractCurrent developments in the field of user interface (UI) technologies as well as robotic systems provide enormous potential to reshape the future of human-robot interaction (HRI) and collaboration. However, the design of reliable, intuitive and comfortable user interfaces is a challenging task. In this paper, we focus on one important aspect of such interfaces, i.e., teleoperation. We explain how to setup a heterogeneous, extendible and immersive system for controlling a distant robotic system via the network. Therefore, we exploit current technologies from the area of virtual reality (VR) and the Unity3D game engine in order to provide natural user interfaces for teleoperation. Regarding robot control, we use the well-known robot operating system (ROS) and apply its freely available modular components. The contribution of this work lies in the implementation of a flexible immersive grasping control system using a network layer (ROSbridge) between Unity3D and ROS for arbitary robotic hardware. Dennis Krupke, Lasse Einig, Eike Langbehn, Jianwei Zhang 0001, Frank Steinicke |
VRST | 4 |
| 2016 | Objectness ranking by uniform Bayesian model with multimodal and global cuesabstractCategory‐independent object detection and localisation plays an important role in many computer vision tasks. In this study, an efficient method is proposed for generic objectness ranking by fusing two dimension (2D) or 3D information. A novel Bayesian model is designed to integrate multimodal cues and global cues to estimate object location, scale and number. In the pure trichannel colour space, the authors employ global spatial information as new global cues. From the colour+depth (red, green and blue+D) aspect, the authors compute multimodal saliency and oversegments to find two new multimodal cues. Local and regional depth cues are also explored and combined with them together so that a reliable objectness ranking scheme can be implemented. The proposed method is evaluated on web‐public common 2D and RGB+D datasets. In RGB+D cases, the experimental results show that the proposed method achieves an average 5% improvement over state‐of‐the‐art methods. Furthermore, for achieving the similar recall rates, the authors’ method only needs 30% amounts of sampled windows with respect of other available methods. Jianhua Zhang 0002, Junhao Xiao 0001, Shengyong Chen, Jianwei Zhang 0001 |
IET Comput. Vis. | 5 |
| 2016 | Kinematic Analysis and Motion Control of Wheeled Mobile Robots in Cylindrical WorkspacesabstractWheeled mobile robots (WMRs) are often used for maintenance of round pipes or ducts, which can typically be represented as a cylindrical workspace. Working in round pipes or ducts, kinematic models of WMRs are different from those applying on a plane and thus pose significant challenges in terms of kinematic analysis and motion control. To address these challenges, the kinematic properties of WMRs in a cylindrical workspace are analyzed in this paper. First, we discuss the kinematic properties of a single wheel in a cylindrical workspace. Then, we analyze the geometric constraints of WMRs in round pipes or ducts with analytical geometry. Based on these analyses, kinematic properties of WMRs in cylindrical workspaces are discussed with screw theory. A control law based on biaxial clinometer information is proposed, and it enables the robot to move horizontally in round pipes or ducts. Finally, the motion of a single wheel purely rolling in a cylindrical workspace is simulated. Experiments using a car-like mobile robot moving in round ducts are carried out to show the feasibility of the proposed algorithm. Zhangjun Song, Hongliang Ren 0001, Jianwei Zhang 0001, Shuzhi Sam Ge |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2016 | Object Classification and Grasp Planning Using Visual and Tactile SensingabstractThe perception of the tactile and vision modalities plays a crucial role in object classification and robot precise operations. On one hand, the visual perception helps acquire the objects apparent characteristics (e.g., shape and color) that are paramount for grasp planning. On the other hand, when grasping the object, the tactile sensors can detect the softness or stiffness and the surface texture of the object, thus can be used to classify the objects. In this paper, two different tactile models are first developed to model the tactile sequences, namely bag-of-system (BoS) and deep dynamical system (DDS). To be specific, BoS applies the extreme learning machine method as the classifier, and DDS employs deep neural networks as the feature mapping functions in order to deal with the highly nonlinear data. As for visual sensing, we formulate a signature of histograms of orientations descriptor to model the shape of objects. Moreover, by using human experience, we identify the graspable components of objects on the category level instead of the object level. Finally, a fast planning method considering both contact points exaction and hand kinematics is proposed to accomplish the grasping manipulation. Comparative experiments demonstrate the effectiveness of our proposed methods on the tasks including the robot grasping and the object classification. Fuchun Sun 0001, Chunfang Liu, Wenbing Huang 0001, Jianwei Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2015 | Motion planning and control of a robotic system for orthodontic archwire bendingabstractIn clinics, customized archwires are demanded for lingual orthodontic treatment. However, only very experienced orthodontists can handle the manual appliance preparation. This pattern not only occupies lots of the orthodontist's labor time, but also can not ensure the accuracy of the appliances. Therefore, a robotic system was developed for automatic and accurate orthodontic archwire bending in our study. First, a method for customized archwire parameterization was developed. Second, an adaptive sampling-based bending planner with collision checker in time-varying environment was designed. Finally, a bending control strategy was used to eliminate the springback effect of the archwires and bending point shift during the bending process. A self-developed simulation platform based on Robot Operating System with MoveIt was used for preliminary validation of the proposed method. Physical experiments for multi-functional orthodontic bends on the robotic system were conducted as well. The results have shown that the developed robotic system using the proposed planning and control method was able to accomplish the automatic and accurate orthodontic archwire bending. Hao Deng 0005, Zeyang Xia, Shaokui Weng, Yangzhou Gan, Jing Xiong 0001, Yongsheng Ou, Jianwei Zhang 0001 |
IROS | 7 |
| 2015 | A novel optical tracking based tele-control system for tabletop object manipulation tasksabstractFor a robot serving in a complex environment such as in a restaurant, it is difficult to perform a task like tabletop object manipulation completely by itself, in that some information may be missing. An approach to deal with this is to use a tele-control system and method to control the robot or demonstrate. In this paper, a LeapMotion sensor based non-contact tele-control method is developed for a robot to perform tabletop object manipulation tasks. A coordinate system for mapping from the operation space of the LeapMotion sensor to the workspace of the robot is established. A gesture recognition and action generating algorithm is proposed for control or to demonstrate the motion to the robot. To evaluate the performance of the LeapMotion sensor and proposed method for tele-control of a robot, a comprehensive assessment index based on entropy weighting is proposed. Three common tele-control modes, including demonstration mode, teleoperation mode and semi-teleoperation mode, are developed on a PR2 robot. The experimental results show that the proposed tele-control system is more appropriate for use in task demonstration. Haiyang Jin, Sebastian Rockel, Ying Hu 0001, Jianwei Zhang 0001 |
IROS | 6 |
| 2015 | Grasp planning by human experience on a variety of objects with complex geometryabstractWe present an effective method of identifying the graspable components of a variety of complex objects for grasp planning based on human experience. Instead of focusing on individual objects, our method identifies graspable components on the category level under the assumption that geometrically alike objects share similar graspable components. Employing a modified SHOT descriptor, we propose a fast KNN-based method for object categorization. Then the graspable components are identified by adopting a learning framework based on human experience. Finally, a fast grasp planning method comprised of contact points exaction and hand kinematics calculation accomplishes the grasp on the identified graspable component. Our experiments demonstrate the effectiveness of this method by realizing grasps on the graspable components of human choice for a variety of unseen objects. Chunfang Liu, Fuchun Sun 0001, Jianwei Zhang 0001 |
IROS | 4 |
| 2015 | Integrating physics-based prediction with Semantic plan Execution MonitoringabstractReal-world robotic systems have to perform reliably in uncertain and dynamic environments. State-of-the-art cognitive robotic systems use an abstract symbolic representation of the real world for high-level reasoning. Some aspects of the world, such as object dynamics, are inherently difficult to capture in an abstract symbolic form, yet they influence whether the executed action will succeed or fail. This paper presents an integrated system that uses a physics-based simulation to predict robot action results and durations, combined with a Hierarchical Task Network (HTN) planner and semantic execution monitoring. We describe a fully integrated system in which a Semantic Execution Monitor (SEM) uses information from the planning domain to perform functional imagination. Based on information obtained from functional imagination, the robot control system decides whether it is necessary to adapt the plan currently being executed. As a proof of concept, we demonstrate a PR2 able to carry tall objects on a tray without the objects toppling. Our approach achieves this by simulating robot and object dynamics. A validation shows that robot action results in simulation can be transferred to the real world. The system improves on state-of-the-art AI plan-based systems by feeding simulated prediction results back into the execution system. Sebastian Rockel, Stefan Konecny, Sebastian Stock 0001, Joachim Hertzberg, Federico Pecora, Jianwei Zhang 0001 |
IROS | 6 |
| 2015 | Caterpillar-Like Climbing Method Incorporating a Dual-Mode Optimal ControllerabstractThis paper presents a bio-inspired caterpillar-like climbing method. Natural caterpillars climb relatively slow, but their multisegmented body trunk strongly enhances climbing versatility and stability, which thus motivates us to design an imitative robot to study the caterpillar-like climbing locomotion. Based on observations from natural caterpillars, we propose a three-stage climbing method to imitate the caterpillar-like straight line climbing on a planar wall. Due to the redundancy property of caterpillars' multisegmented body trunk, we formulate the caterpillar-like climbing problem as an end-effector tracking problem using the redundant robotics terminology. A dual-mode optimal controller, which effectively resolves the end-effector tracking problem even when the robot configuration is ill-conditioned, is incorporated for realizing the caterpillar-like climbing locomotion. Both simulation and experiment results are presented to demonstrate the effectiveness of the proposed caterpillar-like climbing method. Guoyuan Li, Jianwei Zhang 0001, Junzhi Yu 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2015 | An Assistive Navigation Framework for the Visually ImpairedabstractThis paper provides a framework for context-aware navigation services for vision impaired people. Integrating advanced intelligence into navigation requires knowledge of the semantic properties of the objects around the user's environment. This interaction is required to enhance communication about objects and places to improve travel decisions. Our intelligent system is a human-in-the-loop cyber-physical system that interprets ubiquitous semantic entities by interacting with the physical world and the cyber domain, viz., 1) visual cues and distance sensing of material objects as line-of-sight interaction to interpret location-context information, and 2) data (tweets) from social media as event-based interaction to interpret situational vibes. The case study elaborates our proposed localization methods (viz., topological, landmark, metric, crowdsourced, and sound localization) for applications in way finding, way confirmation, user tracking, socialization, and situation alerts. Our pilot evaluation provides a proof of concept for an assistive navigation system. Jizhong Xiao, Samleo L. Joseph, Bing Li 0008, Xiaohai Li, Jianwei Zhang 0001 |
IEEE Trans. Hum. Mach. Syst. | 6 |
| 2014 | Model-based state recognition of bone drilling with robotic orthopedic surgery systemabstractScrew path drilling is an important process among many orthopedic surgeries. To guarantee the safety and correctness of this process, a model-based drilling state recognition method is proposed in this paper. The thrust force in the drilling process is modeled based on an accurate 3D bone model restructured by means of Micro-CT images. In theoretical modeling of the thrust force, the resistance and the elasticity of the bone tissues are considered. The cutting energy and elastic modulus are defined as the material parameters in the theoretical model, which are identified via a least square method. Some key parameters are proposed to support the state recognition: the peak forces in the first and the second cortical layers, the average force in the cancellous layer and the thickness of each layer. Based on these key parameters in the model, a state recognition strategy with a robotic orthopedic surgery system is proposed to recognize the switch position of each layer. Experiments are performed to demonstrate the effectiveness of the modeling approach and the state recognition method. Haiyang Jin, Ying Hu 0001, Zhen Deng, Peng Zhang 0012, Zhangjun Song, Jianwei Zhang 0001 |
ICRA | 6 |
| 2014 | An hyperreality imagination based reasoning and evaluation system (HIRES)abstractIn this work we ask whether an integrated system based on the concept of human imagination and realized as a hyperreal setup can improve system robustness and autonomy. In particular we focus on how non-nominal failures in a planning-based system can be detected before actual failure. To investigate, we integrated a system combining an accurate physics-based simulation, robust object recognition and a symbolic planner to achieve realistic prediction of robot actions. A Gazebo simulation was used to reason about and evaluate situations before and during plan execution. The simulation enabled re-planning to take place in advance of actual plan failure. We present a restaurant scenario in which our system prevents plan failure and successfully lets the robot serve a drink on a table cluttered with objects. The results give us confidence in our approach to improving situations where unavoidable abstractions of robot action planning meet the real world. Sebastian Rockel, Denis Klimentjew, Jianwei Zhang 0001 |
ICRA | 4 |
| 2014 | A hough transform based scan registration strategy for Mobile Robotic MappingabstractThe scan registration is the cornerstone to Mobile Robotic Mapping, and the majority of existing global registration methods are dependent on specific features. This paper presents a global feature-less scan registration strategy based on the ground surface, which is extremely common in Mobile Robotic Mapping scenarios. The 3D rotation is decoupled from 3D translation by transforming the input scans into the Hough domain, wherein Phase Only Matched Filtering (POMF) is adopted for the partially overlapped signal registration. No particular features in the input data are prerequisite to our algorithm. The algorithm is validated by the challenging scans captured by our custom-built platform and a public dataset. The result illustrates the reliability of this algorithm to align feature-less, partially overlapped and noisy scans. Bo Sun 0009, Junhao Xiao 0001, Jianwei Zhang 0001 |
ICRA | 4 |
| 2014 | Push resistance in in-hand manipulationabstractHaptic perception plays an important role in human life. It provides comprehensive information about the world, especially in in-hand manipulation tasks. For human the in-hand manipulation is an easy work; for robots it is still a challenge. One of its issues lies in the uncertainty of the interaction state. This paper researches robot object interaction from a novel angle, the method is called haptic exploration, which helps robots acquire the ability to explore the object in hand. For in-hand manipulation task, the haptic exploration is a process where the robot hand perceives contact force feedback from slightly push along different directions. In order to describe the force feedback, two definitions are given: pushable direction and push resistance. Additionally, a single finger push model and spatial multi-finger push model are proposed to illustrate push resistance. Furthermore an object push strategy is presented to reduce the number of push directions. At last real robot experiments are conducted to verify the proposed models. And its result also shows the feasibility of haptic exploration method. Junhu He, Jianwei Zhang 0001 |
IROS | 2 |
| 2014 | A ground-based optical system for autonomous landing of a fixed wing UAVabstractThis paper presents a new ground-based visual approach for guidance and safe landing of an unmanned aerial vehicle (UAV) in Global Navigation Satellite System(GNSS)-denied environments. In our previous work, the old system consists of one pan-tilt unit(PTU) with two cameras, whose detection range is limited by the baseline. To achieve long-range detection and cover wide field of regard, we mounted two separate sets of PTU integrated with visible light camera on both sides of the runway instead of our previous assembled stereo vision system. Then, the well-known AdaBoost method was evaluated with regard to detecting and tracking the target. To achieve the relative position between the UAV and landing area, we used triangulation to calculate the 3D coordinates of the UAV. By combining the estimated position in the closed loop control, we obtain the autonomous landing strategy. Finally, we present several real flights in outdoor environments, and compare its accuracy with ground truth provided by GNSS. The results support the validity and accuracy of the presented system. Dianle Zhou, Yu Zhang 0082, Daibing Zhang, Xun Wang 0003, Boxin Zhao, Chengping Yan, Lincheng Shen, Jianwei Zhang 0001 |
IROS | 9 |
| 2014 | State recognition of bone drilling with audio signal in Robotic Orthopedics Surgery SystemabstractBone drilling is an important and difficult process in orthopedic surgeries. To detect the drilling state of a Robotic Orthopedic Surgery System (ROSS) in real-time, a state recognition method based on audio signals, the Acoustic Emission (AE) signals generated in drilling process, is proposed in this paper. By an analysis via power spectral density of the AE signals, an appropriate frequency band is selected for state recognition. The Exponential Mean Amplitude (EMA) and the Hurst Exponent (HE) are used to illustrate the energy characteristics and stability of the AE signals in the chosen frequency band, respectively. The recognition algorithm combines the two different features is performed on a embedded device in the experiments. Finally, the experiments are carried out to demonstrate the effectiveness of the proposed drilling state recognition method. Yu Sun 0018, Haiyang Jin, Ying Hu 0001, Peng Zhang 0012, Jianwei Zhang 0001 |
IROS | 5 |
| 2014 | In-hand haptic perception in dexterous manipulations
Junhu He, Jianwei Zhang 0001 |
Sci. China Inf. Sci. | 2 |
| 2014 | Modeling response properties of V2 neurons using a hierarchical K-means model
Xiaolin Hu 0001, Jianwei Zhang 0001, Peng Qi 0004, Bo Zhang 0010 |
Neurocomputing | 2 |
| 2014 | Spatial Neighborhood-Constrained Linear Coding for Visual Object TrackingabstractIn this paper, a new spatial neighborhood-constrained linear coding strategy which realizes sparse representation is proposed for visual object tracking. Unlike conventional sparse and locality-constrained linear coding approaches that need an extra post-processing stage to incorporate the spatial layout information, the proposed coding strategy intrinsically embeds the spatial layout information into the coding stage. The proposed coding strategy can also be used to effectively realize joint sparse representation for different feature descriptors. In addition, based on the distance to the “ideal point” in the reconstruction error space, a new multicue integration approach for robust tracking is proposed and a co-learning approach is developed to update the dictionaries. Finally, the proposed tracking algorithm is compared with other state-of-the-art trackers on some challenging video sequences and shows promising results. Huaping Liu 0001, Mingyi Yuan, Fuchun Sun 0001, Jianwei Zhang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2014 | A Survey on CPG-Inspired Control Models and System ImplementationabstractThis paper surveys the developments of the last 20 years in the field of central pattern generator (CPG) inspired locomotion control, with particular emphasis on the fast emerging robotics-related applications. Functioning as a biological neural network, CPGs can be considered as a group of coupled neurons that generate rhythmic signals without sensory feedback; however, sensory feedback is needed to shape the CPG signals. The basic idea in engineering endeavors is to replicate this intrinsic, computationally efficient, distributed control mechanism for multiple articulated joints, or multi-DOF control cases. In terms of various abstraction levels, existing CPG control models and their extensions are reviewed with a focus on the relative advantages and disadvantages of the models, including ease of design and implementation. The main issues arising from design, optimization, and implementation of the CPG-based control as well as possible alternatives are further discussed, with an attempt to shed more light on locomotion control-oriented theories and applications. The design challenges and trends associated with the further advancement of this area are also summarized. Junzhi Yu 0001, Min Tan 0001, Jianwei Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2013 | Finding next best views for autonomous UAV mapping through GPU-accelerated particle simulationabstractThis paper presents a novel algorithm capable of generating multiple next best views (NBVs), sorted by achievable information gain. Although being designed for way-point generation in autonomous airborne mapping of outdoor environments, it works directly on raw point clouds and thus can be used with any sensor generating spatial occupancy information (e.g. LIDAR, kinect or Time-of-Flight cameras). To satisfy time-constraints introduced by operation on UAVs, the algorithm is implemented on a highly parallel architecture and benchmarked against the previous, CPU-based proof of concept. As the underlying hardware imposes limitations with regards to memory access and concurrency, necessary data structures and further performance considerations are explained in detail. Open-source code for this paper is available at http://www.github.com/benadler/. Benjamin Adler, Junhao Xiao 0001, Jianwei Zhang 0001 |
IROS | 3 |
| 2013 | Autonomous landing of an UAV with a ground-based actuated infrared stereo vision systemabstractIn this study, we focus on the problem of landing an unmanned aerial vehicle (UAV) in unknown and Global Navigation Satellite System(GNSS)-denied environments based on an infrared stereo vision system. This system is fixed on the ground and used to track the UAV's position during the landing process. In order to enlarge the search field of view (FOV), a pan-tilt unit (PTU) is employed to actuate the vision system. The infrared camera is chosen as the exteroceptive sensor for two main reasons: first, it can be used under all weather conditions and around the clock; second, infrared targets can be tracked based on infrared spectrum features at a lower computational cost compared to tracking texture features in visible spectrum. State-of-the-art active contour based algorithms and the mean shift algorithm have been evaluated with regard to detecting and tracking an infrared target. Field experiments have been carried out using an unmanned quadrotor and a fixed-wing unmanned aircraft, with both qualitative and quantitative evaluations. The results demonstrate that our system can track UAVs without artificial markers and is sufficient to enhance or replace the GNSS-based localization in GNSS-denied environment or where its information is inaccurate. Daibing Zhang, Xun Wang 0003, Zhiwen Xian, Jianwei Zhang 0001 |
IROS | 5 |
| 2013 | Discover Novel Visual Categories From Dynamic Hierarchies Using Multimodal AttributesabstractLearning novel visual categories from observations and experiences in unexplored environment is a vitally important cognitive ability for human beings. A dynamic category hierarchy that is an inherent structure in a human mind is a key component for this ability. This paper develops a framework to build dynamic category hierarchy based on object attributes and a topic model. Since humans trend to utilize multimodal information to learn novel categories, we also develop an algorithm to learn multimodal object attributes from multimodal data. The new multimodal attributes can describe objects efficiently and can generalize from learned categories to novel ones. By comparison with a state-of-the-art unimodal attribute, the multimodal attributes can achieve 4%-19% improvements on average. We also develop a constrained topic model, which can accurately construct category hierarchies for large-scale categories. Based on them, the novel framework can effectively detect novel categories and relate them with known categories for further category learning. Extensive experiments are conducted using a public multimodal dataset, i.e., color and point cloud data, to evaluate the multimodal attributes and the dynamic category hierarchy. The experimental results show the effectiveness of multimodal attributes to describe objects and the satisfactory performance of the dynamic category hierarchy to discover novel categories. By comparison with state-of-the-art methods, the dynamic category hierarchy achieves 7% improvements. Jianhua Zhang 0002, Jianwei Zhang 0001, Shengyong Chen |
IEEE Trans. Ind. Informatics | 2 |
| 2012 | Flexible Modular Robotic Simulation Environment For Research And EducationabstractIn this paper a novel GUI for a modular robots simulation environment is introduced. The GUI is intended to be used by unexperienced users that take part in an educational workshop as well as by experienced researchers who want to work on the topic of control algorithms of modular robots with the help of a framework. It offers two modes for the two kinds of users. Each mode makes it possible to configure everything needed with a graphical interface and stores configurations in XML files. Furthermore, the GUI not only supports importing the user’s control algorithms, but also provides online modulation for these algorithms. Some learning techniques such as genetic algorithms and reinforcement learning are also integrated into the GUI for locomotion optimization. Thus, its easy to use, and its scalability makes it suitable for research and education. Dennis Krupke, Guoyuan Li, Jianwei Zhang 0001, Houxiang Zhang, Hans Petter Hildre |
ECMS | 3 |
| 2012 | Continual HTN Planning and Acting in Open-ended Domains - Considering Knowledge Acquisition Opportunities
Dominik Off, Jianwei Zhang 0001 |
ICAART (1) | 2 |
| 2012 | Hybrid physics simulation of multi-fingered hands for dexterous in-hand manipulationabstractDextrous object manipulation with multi-fingered robot hands remains one of the key challenges of service robotics. So far, most theoretical approaches and simulators have concentrated on the search for and evaluation of static stable grasps, but with neither a model of the full hand-arm system nor the system dynamics. GraspIt! is probably the best-known simulator of this kind. In this work we present a simulator that uses the JBullet physics engine to realistically model grasps with multi-fingered hands. It supports manipulation tasks based on a complete arm and hand system, with full calculation of hand and object dynamics. A hybrid dynamics and kinematics approach avoids the oscillations introduced by the different size scale of the arm and hand, so that force-closure grasps are possible in addition to form-closure grasps. The software includes detailed models of our 24-DOF Shadow Dextrous hand and the 6-DOF Mitsubishi PA-10 robot arm. A real-time interface allows us to prepare or to replay and analyze grasp experiments performed on our real robots. Hanno Scharfe, Norman Hendrich, Jianwei Zhang 0001 |
ICRA | 3 |
| 2012 | Action gist based automatic segmentation for periodic in-hand manipulation movement learningabstractWe consider in-hand manipulation tasks that consists of periodic movements. In order to improve the manipulation learning ability of a robot with a human-like hand, this paper introduces a segmentation method based on the techniques of action gist. Action gist is the key motion information in manipulation with the property of semantics. In the techniques of in-hand manipulation action gist, there is a Meta Motion Occurrence Histogram describing the motion information in the demonstration set. This paper proposes an algorithm related to the Meta Motion Occurrence Histogram to maximize the common motions in each segment, so as to figure out the best segmentation solution in the in-hand manipulation sequence. The experiments illustrate the performance of the proposed method, and discuss the possibility of segmentation fusing with the information from tactile sensor. Gang Cheng 0001, Norman Hendrich, Jianwei Zhang 0001 |
IROS | 3 |
| 2012 | Constructing dynamic category hierarchies for novel visual category discoveryabstractCategory hierarchies are commonly used to compactly represent large numbers of categories and reduce the complexity of the classification problem. In this paper we introduce a novel and extended application of category hierarchies which is a powerful novel framework developed to construct dynamic category hierarchies and automatically discover novel visual categories. The dynamic is a characteristic of category hierarchies which can facilitate an important cognitive ability, the discovering of novel categories. We develop a constrained hierarchical latent Dirichlet allocation to build accurate category hierarchies. We employ object attributes as features to describe objects, which can transfer knowledge across categories and can efficiently describe novel categories. By combining them in the novel framework, novel visual object categories can be efficiently discovered and described. Extensive experiments based on PASCAL VOC 2008 and the LabelMe image database show the satisfactory performance of the proposed framework. Jianhua Zhang 0002, Jianwei Zhang 0001, Shengyong Chen, Ying Hu 0001, Haojun Guan |
IROS | 2 |
| 2012 | A Hierarchical Model Incorporating Segmented Regions and Pixel Descriptors for Video Background SubtractionabstractBackground subtraction is important for detecting moving objects in videos. Currently, there are many approaches to performing background subtraction. However, they usually neglect the fact that the background images consist of different objects whose conditions may change frequently. In this paper, a novel hierarchical background model is proposed based on segmented background images. It first segments the background images into several regions by the mean-shift algorithm. Then, a hierarchical model, which consists of the region models and pixel models, is created. The region model is a kind of approximate Gaussian mixture model extracted from the histogram of a specific region. The pixel model is based on the cooccurrence of image variations described by histograms of oriented gradients of pixels in each region. Benefiting from the background segmentation, the region models and pixel models corresponding to different regions can be set to different parameters. The pixel descriptors are calculated only from neighboring pixels belonging to the same object. The experimental results are carried out with a video database to demonstrate the effectiveness, which is applied to both static and dynamic scenes by comparing it with some well-known background subtraction methods. Shengyong Chen, Jianhua Zhang 0002, Youfu Li 0001, Jianwei Zhang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2012 | Analyzing Image Deblurring Through Three ParadigmsabstractTo recover a sharp version from a blurred image is a long-standing inverse problem. In this paper, we analyze the research on this topic both theoretically and experimentally through three paradigms: 1) the deterministic filter; 2) Bayesian estimation; and 3) the conjunctive deblurring algorithm (CODA), which performs the deterministic filter and Bayesian estimation in a conjunctive manner. We point out the weaknesses of the deterministic filter and unify the limitation latent in two kinds of Bayesian estimators. We further explain why the CODA is able to handle quite large blurs beyond Bayesian estimation. Finally, we propose a novel method to overcome several unreported limitations of the CODA. Although extensive experiments demonstrate that our method outperforms state-of-the-art methods with a large margin, some common problems of image deblurring still remain unsolved and should attract further research efforts. Chao Wang 0063, Lifeng Sun, Peng Cui 0001, Jianwei Zhang 0001, Shiqiang Yang |
IEEE Trans. Image Process. | 4 |
| 2012 | A Matrix-Based Approach to Unsupervised Human Action CategorizationabstractHuman action, as the basic unit of most human-relevant video content, bridges the gap between low-level visual features and high-level semantics. Human action recognition is of great significance in the applications of human-computer interaction, intelligent video surveillance, video retrieval and search. In this paper, we propose a novel unsupervised approach to mining categories from action video sequences, which consists of two modules: action representation for video data structurization and learning model for unsupervised categorization. In action representation, a novel view of video decomposition is presented. Videos are regarded as spatially distributed dynamic pixel time series, and these dynamic pixels are first quantized into pixel prototypes. After replacing the pixel time series with their corresponding prototype labels, the video sequences are compressed into two-dimensional action matrices. In the learning model, we put these matrices together to form an multi-action tensor, and propose the joint matrix factorization method to simultaneously cluster the pixel prototypes into pixel signatures, and matrices into action classes with the consideration of the duality between pixel clustering and action clustering. The approach is tested on public and popular Weizmann, and KTH datasets, and promising results are achieved. Peng Cui 0001, Fei Wang 0001, Lifeng Sun, Jianwei Zhang 0001, Shiqiang Yang |
IEEE Trans. Multim. | 4 |
| 2012 | Control of Yaw and Pitch Maneuvers of a Multilink Dolphin RobotabstractThis paper is devoted to the active turn control of a free-swimming multilink dolphin-like robot, with emphasis on yaw and pitch controls. With full consideration of both mechanical configuration and propulsive principle of the robot consisting of a yaw joint and multiple pitch joints, a viable approach to perform yaw maneuvers via laterally directed biases is formed, providing an advantage in qualitative and quantitative assessment. Meanwhile, based on the feedback of the pitch angle measured by an onboard gyroscope, a closed-loop control strategy in dorsoventral motions is proposed to achieve agile and swift pitch maneuvers. More remarkably, two hybrid acrobatic stunts, i.e., frontflip and backflip, are first implemented on the physical robot. The latest results obtained demonstrate the effectiveness of the proposed methods. It is also confirmed that the dolphin robot achieves better performance for pitch maneuvers than it does for yaw maneuvers, agreeing well with the biological observations. Junzhi Yu 0001, Zongshuai Su, Ming Wang 0001, Min Tan 0001, Jianwei Zhang 0001 |
IEEE Trans. Robotics | 5 |
| 2011 | CPG-based behavior design and implementation for a biomimetic amphibious robotabstractThis paper presents the behavior design and multimodal locomotion control of a biomimetic amphibious robot based on a bio-inspired CPG (central pattern generator). A set of four key parameters are introduced serving as external stimuli to shape the CPG rhythmic activities where necessary speed and orientation modulation as well as 3-D locomotion can be obtained. In terms of the built parameter set, a library of movement primitives based on finite state machine is established to facilitate rapid and smooth gait transitions. To enhance adaptive behaviors, well-integrated sensory feedback by means of two liquid-level detectors enables the gait transition between ground and water autonomously. Simulations and experiments are also conducted to demonstrate the feasibility of a behavior based control architecture governed by CPGs. Rui Ding 0006, Junzhi Yu 0001, Qinghai Yang, Min Tan 0001, Jianwei Zhang 0001 |
ICRA | 5 |
| 2011 | Analysis and control of a biped line-walking robot for inspection of power transmission linesabstractThe subject of this paper is the analysis and control of a biped line walking robot for inspection of power transmission lines. The novel mechanism enables the center of the robot to concentrate on the hip joint to minimize the drive torque of the hip joint The mechanical structure of the robot is discussed, as well as forward kinematics. A dynamic model is established in this paper. Locomotion of the line-walking robot is discussed in detail and the line-walking cycle of the designed mechanism is composed of a single-support phase and a double-support phase. Finally, a PD+ gravity compensation method is introduced for the controller design in the single-support phase and a complaint controller with estimated force is implemented during the double-support phase. The feasibility of this concept is then confirmed by performing experiments with a simulated line environment. Ludan Wang, Fei Liu 0029, Shaoqiang Xu, Jianwei Zhang 0001 |
ICRA | 5 |
| 2011 | Dynamic modeling and its application for a CPG-coupled robotic fishabstractIn this paper, we present the formulation of a dynamic model of a free-swimming multi-joint robotic fish with a pair of wing-like pectoral fins, in which the whole robot is regarded as a moving multilink rigid body in fluids. Considering that the thrust of fish mainly results from the force of trailing vortex, added lateral pressure, and leading-edge suction force, the dynamic equations of the swimming fish have been derived by summing up the longitudinal force, lateral force, and yaw moment on each propulsive component in the framework of Lagrangian mechanics. Furthermore, using the bio-inspired Central Pattern Generators (CPGs) as the swimming data generator, the overall dynamic propulsive characteristics of the swimming robot are estimated in a mathematical environment (i.e., Mathematica). As a case study, the created dynamic model offers a good guide to seeking pragmatic backward swimming patterns for a carangiform robotic fish, which exemplifies the validity of the CPG-coupled dynamic model. Junzhi Yu 0001, Ming Wang 0001, Zongshuai Su, Min Tan 0001, Jianwei Zhang 0001 |
ICRA | 5 |
| 2011 | Design and control of a fish-inspired multimodal swimming robotabstractPresented in this paper is our effort to create a multifunctional swimming robot, i.e., robotic fish, inspired by the well-integrated, configurable multiple control surfaces existing in real fish. By virtue of the hybrid propulsion capability in the tail plus the caudal fin and the maneuverability in accessory fins, a novel, synthesized propulsion scheme composed of multiple artificial control surfaces is proposed, involving the tail plus the caudal fin, pectoral fins, pelvic fin, and dorsal fin. Multimodal locomotion is then accomplished by manipulation of control surfaces, separately or cooperatively, allowing the robot to maneuver more diversely and agilely. In particular, bio inspired Central Pattern Generators (CPGs) based locomotion control is adopted for online swimming gait generation. Aquatic testing has been carried out to demonstrate the improved maneuverability and stability of the robotic fish underwater as well as the effectiveness of the conceived multi-fin mechatronic design. Junzhi Yu 0001, Ming Wang 0001, Weibing Wang, Min Tan 0001, Jianwei Zhang 0001 |
ICRA | 5 |
| 2011 | Integrate multi-modal cues for category-independent object detection and localizationabstractTo detect and localize objects is an indispensable step for many computer vision tasks. Most of the state-of-the-art methods of object detection and localization are category-dependent. These methods can achieve a significant performance. However, they are useless for detecting and localizing objects belonging to an unknown category when applying them to an unknown environment. In this paper, a method is proposed for detecting and localizing generic objects without specifying their categories. The proposed method combines diverse cues, including multi-scale saliency, superpixels straddling, intensity, depth and global information, into a uniform Bayesian framework to obtain accurate detection and localization. By comparison to state-of-the-art methods, our experiments show the promising performance of the proposed method based on the PASCAL VOC 08 dataset and our indoor scene dataset. Jianhua Zhang 0002, Junhao Xiao 0001, Jianwei Zhang 0001, Houxiang Zhang, Shengyong Chen |
IROS | 3 |
| 2010 | Internal force compensating method for wall-climbing caterpillar robotabstractThe redundant driving problem is an inherent phenomenon existing in a modular caterpillar robot. To limit the internal forces arising from the redundant driving, this paper proposes a joint torque control method, which is based on an assumption that there is only one active joint in the four-link mechanism driving the climbing gait. Except the active joint, the other three joints are all considered as passive joints, whose torques tend to be zero, although they are driven by motors in reality. According to the analyses of static forces in the closed chain state, the ideal torque of named active joint is calculated, and will be followed by the joint in real climbing locomotion. The experiments reveal the reasonability and feasibility of the proposed joint control method, as well as the limitations of current prototype and control algorithm. Wei Wang 0034, Houxiang Zhang, Jianwei Zhang 0001 |
ICRA | 4 |
| 2010 | HTN robot planning in partially observable dynamic environmentsabstractExperiments showed that Hierarchical Task Network (HTN) planners are suitable to find solutions for nontrivial tasks in complex scenarios. Mobile service robots are able to execute actions which may constitute the basic building blocks to achieve high-level goals. However, only few experiments demonstrate the application of a general purpose deliberative planner in the domain of mobile service robots. One challenging problem arises from the fact that adaptive AI-based planners presume the closed-world assumption (CWA) and are therefore unable to deal with incomplete information. Unknown objects which are not represented in the planning domain, for example, cannot be integrated into the planning process. Since mobile service robots act in a real dynamic environment and construct or adapt their world model autonomously based on sensory data, they are inevitably confronted with uncertain and incomplete information about the world. This conflict between simplified assumptions for planning on the one hand and the complexity of the real world on the other constitutes a major problem of modern robotics. This paper describes two approaches to dealing with incomplete world knowledge in the context of HTN robot planning. Several experiments demonstrate that the approaches can successfully be applied in a dynamic and unstructured environment. Martin Weser, Dominik Off, Jianwei Zhang 0001 |
ICRA | 3 |
| 2010 | A cloud computing approach to complex robot vision tasks using smart camera systemsabstractIn this paper we show our work on enabling service robot systems to distribute parts of the image processing functions to different off-board computer systems in the working environment of the robot. Thus complex algorithms can be carried out on high performance systems circumventing the restrictions considering space and power consumption that a mobile platform imposes. As high resolution cameras provide a huge amount of image data and the bandwidth of a wireless network connection is strongly limited, we are using intelligent camera systems on the mobile robot platform to execute parts of the image processing functions directly on the robot. This way only preprocessed image information will be transmitted instead of raw image data. We are using a flexible modular software framework that allows us to split image processing tasks into a pipeline of modular functions that can run on different systems. We show how our approach can be used to enable a service robot system to speed up high resolution SIFT-based object detection. Hannes Bistry, Jianwei Zhang 0001 |
IROS | 2 |
| 2010 | Robust gait control in biomimetic amphibious robot using central pattern generatorabstractThis paper presents a control architecture for the underwater locomotion control of a biomimetic amphibious robot with multi-mobility mechanism. In view of both hydrodynamic problem and engineering approach, we develop a robotic prototype capable of multi-mode motion. A robust gait control for steady swimming using the central pattern generator (CPG) is proposed and has been successfully applied to the robot. The CPG can produce coordinated patterns of rhythmic activity while being simply modulated by control parameters including input drive, frequency, amplitude, threshold, etc., which will be suitable for manually interactive modulation. Using the CPG model, the robot is capable of performing and switching between various locomotion modes such as swimming forwards and backwards, turning and pitching, with the speed, direction and gait types modulated accordingly. A test-bed is provided and results are presented demonstrating interesting properties of the CPG-based control approach and feasibility of the CPG control for efficient propulsion. Rui Ding 0006, Junzhi Yu 0001, Qinghai Yang, Min Tan 0001, Jianwei Zhang 0001 |
IROS | 5 |
| 2010 | Closed-loop precise turning control for a BCF-mode robotic fishabstractThis paper deals with a novel closed-loop maneuvering control method to enhance the turning precision and turning response speed of a robotic fish propelled via the body and/or caudal fin (BCF) mode. Although the BCF propulsion is favorable for the cases requiring greater thrust and accelerations, its maneuverability can be compensated by effective turning control. In our method, the turning maneuver is divided into three phases: the bending, holding, and unbending phases. After much consideration on turning details, the functions of each phase and the basic control laws are further identified. Results of experiments on in-situ direction tracking and direction maintaining verify the effectiveness of the proposed turning control. Zongshuai Su, Junzhi Yu 0001, Min Tan 0001, Jianwei Zhang 0001 |
IROS | 4 |
| 2010 | Development of a practical power transmission line inspection robot based on a novel line walking mechanismabstractA mobile robot based on novel line-walking mechanism is proposed for inspecting power transmission lines. The novel mechanism enables the centroid of the robot to concentrate on the hip joint to minimize the drive torque of the hip joint and keep the robot stable when only one leg is hung on line. After reviewing of the line-walking mechanism, power line inspection robot is described in detail. The pose adjustment analysis is carried out to make sure that the robot will keep in stable state when rolling on the power transmission line which forms a catenary curve. The obstacle-navigation cycle of the designed inspection robot is composed of a single-support phase and a double-support phase. The centroid of the robot will be adjusted to shift to the other leg to start a new single-support phase. The feasibility of this concept is then confirmed by performing experiments with a simulated line environment. Ludan Wang, Fei Liu 0029, Shaoqiang Xu, Jianwei Zhang 0001 |
IROS | 6 |
| 2010 | EpistemeBase: A semantic memory system for task planning under uncertaintiesabstractTasks planning under uncertainties is one of fundamental skills for enabling autonomous robots to make proper manipulations in the complex environment. But owing to inexpressive representations, autonomous robots hardly conduct efficient tasks planning, especially in unknown conditions. The application of semantic knowledge in task planning is critically required in artificial intelligence research. In this paper, we focus on two topics: semantic knowledge representations and parallel planning for uncertainties. Firstly, a semantic memory system which is called EpistemeBase is proposed for indoor tasks planning, it includes five parallel agents: Assertion, Plan, Anticipation, Behaviour and Effect. Its framework is an evolving process, which consists of Datum, Information, Knowledge and Intelligence. Secondly, the same task planning is synchronously represented by five paralleled agents. This paralleled structure can well accelerate the process of tasks planning as well as better handle it under uncertainties. Finally, the experiment of tasks planning is conducted for measuring the reaction time of planning and uncertainties by using the EpistemeBase and the Open Mind Common Sense (OMCS) respectively. Xiaofeng Xiong, Ying Hu 0001, Jianwei Zhang 0001 |
IROS | 3 |
| 2010 | Efficient kinematic solution to a multi-robot with serial and parallel mechanismsabstractThis paper presents an efficient kinematical solution to a multi-robot system with serial and parallel mechanisms. JL-I is a reconfigurable robot featuring active spherical joints formed by serial and parallel mechanisms endowing the robotic system with the ability of changing shapes in three dimensions. The active joint here can combine the advantages of the high rigidity of a parallel mechanism and the extended workspace of a serial mechanism. However, the kinematic analysis of the serial and parallel mechanism is always the bottleneck in designing a robot and control realization. In order to deal with this problem, the whole kinematical analysis is organized in the sequence from the direct mechanical analysis related to the serial and parallel mechanism over the numerical solutions to the simplified kinematics expression. The latest results obtained demonstrate that the deduced closure-form solution is time efficient and easy to implement while offering a satisfactory motion performance in on-site experiments. Houxiang Zhang, Gionata Salvietti, Wei Wang 0034, Guoyuan Li, Junzhi Yu 0001, Jianwei Zhang 0001 |
IROS | 6 |
| 2009 | Robust inter-scale non-blind image motion deblurringabstractKernel estimate errors and image noise are major causes of visual artifacts in image motion deblurring. We propose an inter-scale non-blind image motion deblurring approach that significantly reduces those artifacts. We use Gaussian Scale Mixture Field of Experts (GSM FOE) model as image prior. The inter-scale smoothness constraint is adopted to suppress the ringing artifacts. In each scale, image details are recovered by the residual deconvolution and the cross bilateral filter (CBF). We further propose a std-controlled CBF to denoise the result. The experimental results are much better than those of previous methods. Chao Wang 0063, Lifeng Sun, Zhuoyuan Chen, Shiqiang Yang, Jianwei Zhang 0001 |
ICIP | 5 |
| 2009 | High-quality non-blind motion deblurringabstractTraditional non-blind motion deblurring methods are sensitive to kernel estimate errors and image noise, thus suffering from either ringing artifacts, enlarged image noise, or over-smoothed image details. We introduce a robust non-blind deblurring algorithm that produces high quality results even from many challenging images with noisy kernels. We adopt the Gaussian Scale Mixture Fields of Experts (GSM FOE) model and the smoothness constraint as image prior, and use the iterative re-weight least-square (IRLS) algorithm to produce the temporal result. The residual deconvolution suite is used to restore the lost image details. We denoise the result using our std-controlled cross bilateral filter. The experimental results are much better than those of previous approaches. Chao Wang 0063, Lifeng Sun, Zhuoyuan Chen, Shiqiang Yang, Jianwei Zhang 0001 |
ICIP | 5 |
| 2009 | Task oriented control of smart camera systems in the context of mobile service robotsabstractIn this paper we present our work on enabling smart cameras to act as autonomous vision systems for mobile service robots. The smart cameras will adapt their functionality to the current task of the robot system. Instead of transferring raw image data, higher level image information or regions-of-interest are transmitted. This way, the amount of data is reduced and computationally intensive image interpretation will not affect other tasks of the robot system. We are developing a flexible software solution to integrate smart camera systems to the architecture of our robot system. Each image processing task is constructed by a composition of modular functions that can be distributed over different systems. Timing aspects of the flow of data can be analyzed with different tools to evaluate the performance. For this work we are using a commercially available smart camera, but due to the use of a modular architecture the porting to other camera models is easy. We show the advantages of our approach in a setup where image regions containing a face are detected and extracted for further processing steps. This task is accomplished using different setups, where more or fewer subtasks are assigned to the camera system. The performance of the overall system is evaluated with respect to processor load, network load and latency of image data. Hannes Bistry, Jianwei Zhang 0001 |
IROS | 2 |
| 2009 | Online motion planning for HOAP-2 humanoid robot navigationabstractAutonomous robot navigation is becoming an increasingly important research topic for mobile robots. In the last few years, significant progress has been made towards stable robotic bipedal walking. This is creating an increased research interest in developing autonomous navigation strategies which are tailored specifically to humanoid robots. Efficient approaches to perception and motion planning, which are suited to the unique characteristics of biped humanoid robots and their typical operating environments, are receiving special interest. In this paper, we present a time-efficient motion planning system for a Fujitsu HOAP-2 humanoid robot. The sampling based algorithm is used to provide the robot with minimal free configuration space which is sampled to extract the robot path. For collision detection, a cylinder model is used to approximate the trajectory for the body center of the humanoid robot during navigation. It calculates the actual distances required to execute different actions of the robot and compares them with the distances to the nearest obstacles. The A* search algorithm is then implemented to find smooth and low-cost footstep placements of the humanoid robot within the resulting configuration space. The proposed hybrid algorithm reduces searching time and produces a smoother path for the humanoid robot at a low cost. Mohammed M. Elmogy, Christopher Habel, Jianwei Zhang 0001 |
IROS | 3 |
| 2009 | Development of a line-walking mechanism for power transmission line inspection purposeabstractA mobile mechanism with biped configuration is proposed for power transmission line inspection purpose. The wire-walking cycle of the designed mechanism composed of single-support phase and double-support phase. During the process of single-support phase, one foot hang on line and the other foot swing from rear to front to overcome obstacles on line and realize wire-walking locomotion. The novel mechanism designing enable the centroid of the robot concentrate on the hip joint to minimize the drive toque of hip joint and keep the robot stable during the single-support phase. And the centroid of the robot will be adjusted to concentrate to the other leg to start a new single-support phase. Both the forward kinematics model and inverse kinematics model are established in this paper for motion control. The feasibility of this concept is then confirmed by designing a real wire-walking robot and by performing experiment with a simulated line environment. Ludan Wang, Jianwei Zhang 0001 |
IROS | 3 |
| 2009 | Docking manipulator for a reconfigurable mobile robot systemabstractJL-2, as a new version of the JL reconfigurable mobile robot system, features not only a docking and 3D posture adjusting capability between its robots, but also a multi-functional docking gripper. The basic concept of JL is that the robots in the system can simultaneously perform basic tasks in flat terrains, and in the case of rugged terrains, the robots can interconnect to enhance their locomotion capabilities. This paper introduces new designs for JL-2 by which the docking mechanism can be used as a simple gripper with 3 DOFs. Then the technologies of the docking mechanism are discussed in detail, including the workspace of the docking gripper, the docking procedure and analyses of the self-aligning ability. Then the workspaces of the posture adjusting mechanisms between two docked robots are analyzed to clarify the reconfiguration ability of JL-2. At last, a series of real experiments are proposed to test the designs and analyses and the basic performance of JL-2. Wei Wang 0034, Wenpeng Yu, Houxiang Zhang, Jianwei Zhang 0001 |
IROS | 4 |
| 2009 | Crawling Locomotion of Modular Climbing Caterpillar Robot with Changing Kinematic ChainabstractBased on the modular concept, this paper presents two caterpillar robot prototypes which are inspired by two typical caterpillars: Inchworm and Pine Caterpillar. The inchworm robot prototype features simplest kinematics and open chain architecture. Due to the fact that there is only one attachment module supporting the inchworm robot during crawling, we apply an Unsymmetrical Phase Method (UPM) to realize a stable crawling gait for it. A pine caterpillar robot is derived from combining two inchworm robots together. The crawling gait of it features a repetitive changing chain: Open-Closed-Open. Besides the UPM in open chain states, a four-links kinematic model is applied to control the corresponding joints to transfer the crawling wave along the robot body in the closed chain state. These two prototypes are all constructed and, and their crawling locomotion abilities have been tested on vertical glasses respectively. Wei Wang 0034, Houxiang Zhang, Jianwei Zhang 0001 |
IROS | 3 |
| 2009 | Autonomous planning for mobile manipulation services based on multi-level robot skillsabstractGeneral purpose service robots are expected to deal with many different tasks in unknown environments. The number of possible tasks and changing situations prevent developers from writing control programs for all tasks and possible situations. Complex robot tasks are thus accomplished by sequential execution of less complex robot actions that are triggered and configured by a task planner. The question of the appropriate abstraction level of robot actions is still being researched and not discussed conclusively. In this paper, we address the problem of atomicity of robot actions and provide some key properties that have to be considered while designing plan-based robot control systems. Based on these properties, we define and implement atomic skills of different abstraction level for the service robot TASER. The HTN planner JShop2 is used to complete the plan-based control architecture which is evaluated in a set of experiments. Martin Weser, Jianwei Zhang 0001 |
IROS | 2 |
| 2009 | Probabilistic Cluster Signature for Modeling Motion ClassesabstractIn this paper, a novel 3-D motion trajectory signature is introduced to serve as an effective description to the raw trajectory. More importantly, based on the trajectory signature, a probabilistic model-based cluster signature is further developed for modeling a motion class. The cluster signature is a mixture model-based motion description that is useful for motion class perception, recognition and to benefit a generalized robot task representation. The signature modeling process is supported by integrating the EM and IPRA algorithms. The conducted experiments verified the cluster signature's effectiveness. Shandong Wu, Youfu Li 0001, Jianwei Zhang 0001 |
IROS | 3 |
| 2009 | The Bilateral Wavelet Pyramid (BWP): A Novel Image RepresentationabstractIn this paper, we study the properties of the bilateral pyramid, based on which a novel image representation-the bilateral wavelet pyramid (BWP) is proposed. The BWP provides a multiresolution technique to manipulate image variances simultaneously in radiometric, spatial and frequency domains. Application to texture retrieval is discussed to demonstrate the effectiveness of the BWP. Chao Wang 0063, Lifeng Sun, Zhuoyuan Chen, Jianwei Zhang 0001, Shiqiang Yang |
ISCAS | 4 |
| 2009 | Statistics in Bilateral Domain: Novel Statistics of Natural ImagesabstractRecently, there has been a great deal of interest in statistics of natural images in linear domain. However, a detail study of image statistics in non-linear domain on a very large dataset is still missing. In this paper, we analyze natural images' statistics in ldquobilateral domainrdquo on different sub bands which are generated by the classical bilateral filter. Compared with those in linear domain, the statistics in bilateral domain are significantly superior in generalization, de-correlation, and local feature extraction. We also observe some entirely new but interesting features. Chao Wang 0063, Lifeng Sun, Zhuoyuan Chen, Jianwei Zhang 0001, Shiqiang Yang |
ISCAS | 4 |
| 2009 | Function of EEG Temporal Complexity Analysis in Neural Activities Measurement
Xiuquan Li, Zhidong Deng, Jianwei Zhang 0001 |
ISNN (1) | 3 |
| 2008 | Fuzzy color extractor based algorithm for segmenting an odor source in near shore ocean conditionsabstractA mission of chemical plume tracing (CPT) in near-shore and ocean environments is to find out an odor source via an autonomous underwater vehicle (AUV). It is necessary to confirm the detected odor source using a visual system interactively or automatically, when a chemical sensor identifies the odor source. However, color images taken in near-shore ocean environments are very vague due to dim illumination conditions and fluid advection effects. This paper presents a fuzzy algorithm for recognizing the chemical plume and its source in near-shore and ocean environments. This algorithm iteratively generates color patterns based on a defined reference color and extracts color components of the chemical plume and its source from fuzzy images using a fuzzy color extractor (FCE). The proposed approach to color image segmentation might be of general interest to robot vision. Wei Li 0006, Yunyi Li, Jianwei Zhang 0001 |
FUZZ-IEEE | 3 |
| 2008 | Moth-inspired chemical plume tracing by integration of fuzzy following-obstacle behaviorabstractThis paper presents a strategy for chemical plume tracing (CPT) with soft obstacle avoidance, such as kelp forests or seaweed in near-shore, oceanic environments. The basic idea is to integrate a low-level fuzzy following obstacle behavior into a subsumption architecture for moth-inspired plume tracing. By on-line generation of dynamic reference targets (DRTs) on a course of the most recent CPT activity, an autonomous underwater vehicle (AUV) can quickly and correctly resume CPT maneuvering after passing obstacles. Simulation studies of CPT with obstacle avoidance are performed using a simulated plume with significant meander and filament intermittency in an environment with length scales of 100 m. Wei Li 0006, Jianwei Zhang 0001 |
FUZZ-IEEE | 2 |
| 2008 | Fuzzy multisensor fusion for autonomous proactive robot perceptionabstractRobot perception still lacks reliability in complex natural environments. A commonly used method to improve perception is to incorporate more sensors with different modalities. This leads to increased computational requirements due to the parallel processing of huge amounts of sensor data. Appropriate sensor fusion methods are needed if contradictory information is provided by different sensors. We propose a feature-based technique to fuse multimodal sensor data using fuzzy rules. Probabilistic methods are avoided by applying fuzzyfication at the feature level. We propose a higher information gain of the available sensors by utilizing robot actions to focus sensors on objects of interest. Therefore sensor readings, algorithms and robot actions are combined into feature detectors. A goal-directed activation of these feature detectors renders parallel processing of all sensor data unnecessary. Martin Weser, Sascha Jockel, Jianwei Zhang 0001 |
FUZZ-IEEE | 3 |
| 2008 | Topic mining on web-shared videosabstractInternet videos have grown exponentially with the help from video sharing Websites. Automatic topic mining is therefore increasingly important for organizing and navigating such large video databases. Most of current solutions of topic detection and mining were done on news videos and cannot be directly applied on Web videos, because of their limited and noisy semantic information. In this paper, we will try to address this problem and propose an automatic topic mining framework on Web videos. We develop an iterative weight-updated co-clustering scheme to filter "noisy" tags and mine the "hot" topics. We then propose a visual-based clustering approach to further group the videos with similar content, and rank the visual-similar groups by their similarity to the topic center. Experiments on a large Web video database demonstrate the superior performance of our weight-updated co-clustering to both of the traditional co-clustering and k-means. The experiments also demonstrate significant improvement of users' experience by our visual-based clustering and ranking. Lu Liu 0005, Yong Rui, Lifeng Sun, Bo Yang 0008, Jianwei Zhang 0001, Shiqiang Yang |
ICASSP | 5 |
| 2008 | A hierarchical motion trajectory signature descriptorabstractMotion trajectory is a compact clue for motion characterization. However, it is normally used directly in its raw data form in most work and effective trajectory description is lacking. In this paper, we propose a novel hierarchical motion trajectory signature descriptor, which can not only fully capture motion features for detailed perception, but also can be used for probabilistic fast recognition. The hierarchy enables the signature to exhibit high functional adaptability meeting different application requirements. At the first-level, differential invariants are employed to describe trajectory features and a nonlinear signature warping method is developed to perceive and recognize trajectories. The second-level signature is the condensation of the first-level signature by applying PCA based dimension optimization. It behaves more efficiently in recognition based on the Gaussian Mixture modeling and Bayesian classifier. The conducted experiments verified the signature’s effectiveness. Shandong Wu, Youfu Li 0001, Jianwei Zhang 0001 |
ICRA | 3 |
| 2008 | Invariant signature description and trajectory reproduction for robot Learning by DemonstrationabstractIn most reported works about robot learning by demonstration (LbD), the demonstration is normally limited to simple gestures or grasp actions. In this paper, motion trajectory oriented LbD is studied in which free form 3-D motion trajectory is extracted to characterize certain human demonstrations. We propose to build effective description to motion trajectories to be learned by a robot instead of learning the raw trajectory data. A novel signature descriptor is formulated which serves as a generic and invariant description for motion trajectories. More importantly, a trajectory reproduction algorithm based on the learned signature is investigated to enable a robot to repeat/follow the reproduced trajectory instance. Experiments are reported to show the signature description and the reproduction algorithm for further application to the LbD. Shandong Wu, Youfu Li 0001, Jianwei Zhang 0001 |
IROS | 3 |
| 2008 | Using the Tandem Approach for AF Classification in an AVSR System
Wolfgang Menzel, Jianwei Zhang 0001 |
ISNN (2) | 3 |
| 2008 | Coal and Gas Outburst Prediction Combining a Neural Network with the Dempster-Shafter Evidence
Yanzi Miao, Jianwei Zhang 0001, Houxiang Zhang, Zhongxiang Zhao |
ISNN (2) | 2 |
| 2008 | Vision Processing for Realtime 3-D Data Acquisition Based on Coded Structured LightabstractStructured light vision systems have been successfully used for accurate measurement of 3-D surfaces in computer vision. However, their applications are mainly limited to scanning stationary objects so far since tens of images have to be captured for recovering one 3-D scene. This paper presents an idea for real-time acquisition of 3-D surface data by a specially coded vision system. To achieve 3-D measurement for a dynamic scene, the data acquisition must be performed with only a single image. A principle of uniquely color-encoded pattern projection is proposed to design a color matrix for improving the reconstruction efficiency. The matrix is produced by a special code sequence and a number of state transitions. A color projector is controlled by a computer to generate the desired color patterns in the scene. The unique indexing of the light codes is crucial here for color projection since it is essential that each light grid be uniquely identified by incorporating local neighborhoods so that 3-D reconstruction can be performed with only local analysis of a single image. A scheme is presented to describe such a vision processing method for fast 3-D data acquisition. Practical experimental performance is provided to analyze the efficiency of the proposed methods. Shengyong Chen, Youfu Li 0001, Jianwei Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2007 | Realtime Structured Light Vision with the Principle of Unique Color CodesabstractTo date, several successful structured light vision systems for accurate 3D measurement in machine vision have been set up. However, these are usually limited to scanning stationary objects or static environments since tens of images have to be captured for recovering one 3D scene, which results in the industry largely avoiding this technology. This paper presents a method of grid-pattern design based on the principles of uniquely color-encoded structured light, to improve the reconstruction efficiency for real-time processing. For a live scene, the 3D measurement is desired to only capture a single image. To realize this, an important problem for the color-encoded projection is the unique indexing of the color codes in the image. It is essential that each light grid be uniquely identified by incorporating the local neighborhoods in the pattern so that 3D reconstruction can be performed with only local analysis of a single image. This paper describes such a method in the design of the special grid patterns and its corresponding 3D reconstruction method for fast vision perception. Shengyong Chen, Youfu Li 0001, Jianwei Zhang 0001 |
ICRA | 3 |
| 2007 | Active Illumination for Robot VisionabstractA vision sensor is the robot's eye to perceive its environment, but the perception performance can be significantly affected by illumination conditions. This paper presents strategies of adaptive illumination control for robot vision to achieve the best scene interpretation. It investigates how to obtain the most comfortable illumination conditions for a vision sensor. In a "comfort" condition the image reflects the natural properties of the concerned object. "Discomfort" may occur if some scene information is lost. Strategies are proposed to optimize the pose and optical parameters of the luminaire and the sensor, with emphasis on controlling the intensity and avoiding glare. Shengyong Chen, Jianwei Zhang 0001, Houxiang Zhang, Wanliang Wang, Youfu Li 0001 |
ICRA | 2 |
| 2007 | Learning to grasp everyday objects using reinforcement-learning with automatic value cut-offabstractAlthough grasping of everyday objects has been a research topic over the last decades, it still is a crucial task for service robots. Several methods have been proposed to generate suitable grasps for objects. Many of them are restricted to a certain type of grasp or limited to a fixed number of contacts. In this paper we propose an algorithm based on reinforcement learning, to enable a service robot to grasp every kind of object with as many contacts as needed. The proposed method will be evaluated using a simulation with a three-fingered robotic hand. Tim Baier-Löwenstein, Jianwei Zhang 0001 |
IROS | 2 |
| 2007 | A smart interface-unit for the integration of pre-processed laser range measurements into robotic systems and sensor networksabstractIn this paper we propose a novel method of connecting laser range finders to a service robot or to an arbitrary sensor network. The developed unit can be connected to two laser range finders and provides a standard Ethernet port. Autonomous computing power enables the unit to carry out data manipulation tasks. Every device in the network will be able to access both the raw measurement data and preprocessed measurements calculated by the smart interface unit. The paper will discuss the technical aspects of the developed unit as well as the implemented novel processing algorithms. These algorithms are specific to tasks many service robots have to carry out. We present a use case where the developed unit has been applied successfully. Hannes Bistry, Daniel Westhoff, Jianwei Zhang 0001 |
IROS | 3 |
| 2007 | F&A compensating variable bang-bang control algorithm for pneumatic driving glass-wall cleaning robotabstractTo overcome the overshoot and oscillation in the normal Bang-Bang controller applied in the pneumatic driving glass-wall cleaning robot, this paper presents a F&A compensating variable Bang-Bang control algorithm for the pneumatic position servo. Besides a new cylinder driving plan is designed, the control algorithm is also improved by adding a fiction force compensating item in the chamber force evaluation equation. To adjust the controlling parameters dynamically, certain increasing functions, such as step function, linear function and arc-tangent function, are used to define the piston's expectative acceleration along with the reduction of the piston's position error. To testify the feasibility of the control algorithms with the above three expectative acceleration setting functions respectively, a pneumatic glass-wall cleaning robot "Sky Cleaner 4" is used as a testing platform to implement a series of on-load position servo experiments. And the different results of them are compared. Wei Wang 0034, Houxiang Zhang, Wenpeng Yu, Guanghua Zong, Jianwei Zhang 0001 |
IROS | 5 |
| 2007 | Runtime reconfiguration of a modular mobile robot with serial and parallel mechanismsabstractThis paper presents a novel field robot JL-I based on a reconfigurable concept for urban search and rescue applications. The robot consists of three identical modules; each module is an entire robotic system that can perform distributed activities. It features three-degrees-of-freedom (DOF) active joints actuated by serial and parallel mechanisms for changing shape and flexible docking mechanism. The docking mechanism enables adjacent modules to connect or disconnect flexibly and automatically. DOF analysis, working space analysis and the kinematics of the 3D active joint between connected modules are studied thoroughly. In the end a series of successful tests confirm the principles and the robot's capabilities. Houxiang Zhang, Shengyong Chen, Wanliang Wang, Jianwei Zhang 0001, Guanghua Zong |
IROS | 4 |
| 2006 | Learning of Demonstrated grasping Skills by Stereoscopic Tracking of Human Head ConfigurationabstractIn this paper a novel approach to learning by demonstration (LbD) is presented. A multimodal service robot is taught grasping skills by a human instructor who demonstrates a grasping action. Our approach contributes novel solutions to the aspects of robustly tracking the demonstrator's hands in real time as well as to the transformation of tracking results into grasping skills. To track the demonstrator's hands in stereoscopic images a mean-shift-like algorithm is adapted. For the very first time this algorithm is applied to local binary patterns (LBP) and color histograms. To retrieve the hand configuration we use view-based principal component analysis (PCA). To develop grasping skills from tracking results the robot repetitively tracks the demonstrator's grasping actions and transforms the results into three-dimensional self organizing maps (SOMs). The SOMs give a spatial description of the collected data and serve as data structures for a reinforcement learning (RL) algorithm which optimizes trajectories for use by the robot. The approach is applied to a multimodal service robot. Experiments show the effectiveness of the LBP-enhanced mean-shift-like tracking and the robustness of LbD based on SOMs and RL Markus Hüser, Tim Baier-Löwenstein, Jianwei Zhang 0001 |
ICRA | 3 |
| 2006 | A Focal Cue for Metric Measurement of 3D SurfacesabstractThis paper finds a method for computing the best-focused location from an image and using it as a dimensional cue for acquisition of a 3D scene surface. In some situations in 3D vision, an object cannot be reconstructed into a 3D model with metric dimensions. Rather, it can only be reconstructed into a 3D structure up to a similarity transformation. To upgrade the 3D model from a similarity transformation to a Euclidean transformation, we propose a method based on the best-focused locations. By analyzing the blur distribution in an image, this method finds the best-focused locations from an image, which provides an additional cue for upgrading the reconstructed 3D structure. Hence, we can obtain not only the object's shape, but also the dimensions and sizes of surface features Shengyong Chen, Youfu Li 0001, Jianwei Zhang 0001 |
IROS | 3 |
| 2006 | Stable Symmetric Feature Detection and Classification in Panoramic Robot Vision SystemsabstractWe propose a novel approach to detect sparse and stable image features by symmetric properties extracted from the visual data. The regional features are formed by a fast qualitative symmetry operator in combination with quantitative symmetry range information. We apply a simple color histogram descriptor to match pre-selected features to those features acquired by our omnidirectional vision system at run time. The complete algorithm produces regional symmetry-based features that are sparse and highly robust to scale change and panoramic image warp, in particular. In this video, we present the feature processing in an object classification experiment using our platform, the Bremen autonomous wheelchair "Rolland III". Kai Huebner, Jianwei Zhang 0001 |
IROS | 2 |
| 2006 | Stable Symmetry Feature Detection and Classification in Panoramic Robot Vision SystemsabstractWe propose a novel approach to detect sparse and stable image features by symmetric properties extracted from the visual data. The regional features are formed by a fast qualitative symmetry operator in combination with quantitative symmetry range information. We apply a simple color histogram descriptor to match pre-selected features to those features acquired by our omnidirectional vision system at run time. The complete algorithm produces regional symmetry-based features that are sparse and highly robust to scale change and panoramic image warp, in particular. We present the algorithms of symmetry and feature processing and show their application in an object classification experiment using our platform, the Bremen autonomous wheelchair "Holland III" Kai Huebner, Jianwei Zhang 0001 |
IROS | 2 |
| 2006 | Online Approaches to Camera Pose RecalibrationabstractThis paper proposes some practical methods for online pose recalibration for cameras with known internal parameters. These approaches work only with reference points, which may be any artificial marks, or any parts of objects fixed in the working environment. Regardless of whether the camera pose is changed by mistake or because of system re-design, the camera will observe the reference points after the change, the recalibration for the camera pose will be completed in a few seconds, and the camera can immediately be reintegrated into the vision system. Compared to the results from a measuring device with a laser tracker and a famous industrial vision system, the accuracy of the recalibration methods with real data completely satisfies the practical demands. Fangwu Shu, Werner Neddermeyer, Jianwei Zhang 0001 |
IROS | 3 |
| 2006 | Visual Object Recognition in Diverse Scenes with Multiple Instance LearningabstractVisual object recognition is important to the robot industry and is a prerequisite for other robot functionalities, such as grasping and manipulation. Object representation and a learning technique are two indispensable parts for this demanding task while arbitrary object appearance and diverse scenes with cluttered background are two great challenges. However, compared with object representation, the learning technique is less developed to deal with these challenges. This paper extends the multiple instance learning (MIL) technique to the multi-class classification scenario and introduces this multi-class MIL framework to the object recognition domain for the first time. This framework is independent of object representation and is useful for object/background discrimination in unseen scenes. Preliminary experiments show that it compares favorably with the supervised learning approach which takes whole images as the classifier training input Dong Wang 0022, Bo Zhang 0010, Jianwei Zhang 0001 |
IROS | 3 |
| 2006 | Locomotion Capabilities of a Novel Reconfigurable Robot with 3 DOF Active Joints for Rugged TerrainabstractReconfigurable robots consist of many modules which are able to change the way they are connected. As a result, the robots have the capability of adopting different configurations to match various tasks and suit complex environments. This paper presents a novel field robot named JL-I which consists of three uniform modules. With the docking mechanisms, the modules can connect or disconnect flexibly and automatically. Furthermore the active joints formed by serial and parallel mechanisms endow the robot with the ability of changing shape in three dimensions. Consequently useful locomotion capabilities, such as crossing high vertical obstacles, getting self-recovery when the robot is upside-down are achieved. After describing the structural principle of the robot, the related kinematics analysis follows. In the end, the successful on-site tests confirm the principles described above and the robot's ability. Houxiang Zhang, Zhicheng Deng, Wei Wang 0034, Jianwei Zhang 0001, Guanghua Zong |
IROS | 4 |
| 2006 | A semi-supervised incremental learning framework for sports video view classificationabstractSports videos have special characteristics such as well-defined video structure, specialized sports syntax, and typically having some canonical view types. In this paper, we propose a semi-supervised incremental learning framework for sports video view classification. Baseball is selected as an example to explain the main ideas. In order to obtain an optimal model based on a small number of pre-labeled training samples, the semi-supervised incremental learning framework explores the local distributed properties of the video sequences and sufficiently utilizes the information of a positive model pool and a negative model pool. After each round of online optimization process for the under-investigating video, a locally-optimized positive model and a set of negative models are added into the positive model pool and the negative model pool according to some heuristic criteria, respectively. Experiments results on real sports video data show that the proposed system is effective and promising Jun Wu 0022, Bo Zhang 0010, Xian-Sheng Hua 0001, Jianwei Zhang 0001 |
MMM | 4 |
| 2004 | Neural network plus fuzzy PD control of tip vibration for flexible-link manipulatorsabstractA neural network (NN) plus fuzzy PD controller is presented in this paper for the trajectory tracking of a flexible-link manipulator, where the robot tip vibration is measured by an optical measurement device with position-sensitive detectors (PSD). Based on the singular perturbation method and two time-scale decompositions, the flexible-link robot model is decomposed into a slow subsystem of an equivalent rigid-link manipulator and a fast subsystem of flexible mode. An NN-based adaptive controller is used to implement the angle position control of the slow dynamics, while the fuzzy PD control with PSD tip deflection measurement is designed to suppress the vibration along with the robot control process. Finally, the control performance of the proposed control approach is illustrated by experimental research. Fuchun Sun 0001, Lingbo Zhang, Yuangang Tang, Jianwei Zhang 0001 |
IROS | 4 |
| 2004 | A novel approach to pneumatic position servo control of a glass wall cleaning robotabstractTaking Shanghai Science and Technology Museum as the operation target, the robot named skycleaner which is totally actuated by pneumatic cylinders and is sucked to the glass walls with vacuum suckers is presented. In order to solve the problems of lower stiffness and the nonlinear movement characteristic of the pneumatic system, a method of segment and variable bang-bang controller is proposed to implement the accurate control of the position servo system using the principle of pneumatic pulse width-modulation (PWM). Testing results show that the controller can effectively improve the control quality. This implies that the method can meet the requirements of realization. Houxiang Zhang, Jianwei Zhang 0001, Guanghua Zong |
IROS | 2 |
| 2004 | Improvements to Bennett?s Nearest Point Algorithm for Support Vector Machines
Jianmin Li 0001, Jianwei Zhang 0001, Bo Zhang 0010 |
ISNN (1) | 2 |
| 2003 | A service robot for automating the sample management in biotechnological cell cultivationsabstractIn this paper we present a mobile robot system that is capable of automating the sample management in a biotechnological laboratory. The system consists of a mobile platform and a robot arm. It can navigate freely in the laboratory and operate standard devices needed for the sample management. The platform uses an extended Kalman filter for localization and A* algorithm for path planning on a tangent graph computed from the laboratory's map. Motion execution has been designed to be as predictable as possible to not irritate, disturb or harm human personnel. The robot arm uses color vision to detect devices and compensate for positioning errors. The parameters and tasks needed to operate the devices are specified in simple scripts to allow quick and easy adaptations to other situations. Torsten Scherer, Iris Poggendorf, Axel Schneider, Daniel Westhoff, Jianwei Zhang 0001, Dirk Lütkemeyer, Jürgen Lehmann, Alois C. Knoll |
ETFA (2) | 5 |
| 2003 | Extraction and transfer of fuzzy control rules for sensor-based robotic operations
Jianwei Zhang 0001, Markus Ferch |
Fuzzy Sets Syst. | 1 |
| 2002 | Visual Guided Grasping of Aggregates using Self-Valuing LearningabstractWe present a self-valuing learning technique which is capable of learning how to grasp unfamiliar objects and generalize the learned abilities. The learning system consists of two learners which distinguish between local and global grasping criteria. The local criteria are not object specific while the global criteria cover physical properties of each object. The system is self-valuing, i.e. it rates its actions by evaluating sensory information and the usage of image processing techniques. An experimental setup consisting of a PUMA-260 manipulator, equipped with a hand-camera and a force/torque sensor, was used to test this scheme. The system has shown the ability to grasp a wide range of objects and to apply previously learned knowledge to new objects. Bernd Rössler, Jianwei Zhang 0001, Alois C. Knoll |
ICRA | 2 |
| 2002 | Control Architecture and Experiment of a Situated Robot System for Interactive AssemblyabstractWe present the development of and experiment with a robot system showing cognitive capabilities of children of three to four years. We focus on two topics: assembly by two hands and understanding human instructions in natural language as a precondition for assembly systems being perceived by humans as "intelligent". A typical application of such a system is interactive assembly. A human communicator sharing a view of the assembly scenario with the robot instructs the latter by speaking to it in the same way that he would communicate with a child. His instructions can be under-specified, incomplete and/or context-dependent. After introducing the general purpose of our project, we present the hardware and software components of our robots necessary for interactive assembly tasks. The control architecture of the robot system with two stationary robot arms is discussed. We then describe the functionalities of the instruction understanding, planning and execution levels. The implementations of a layered-learning methodology, memories and monitoring functions are briefly introduced. Finally, we outline a list of future research topics for extending our system. Jianwei Zhang 0001, Alois C. Knoll |
ICRA | 1 |
| 2002 | Learning cooperative assembly with the graph representation of a state-action spaceabstractIn this paper, we present a method for two robot manipulators to learn cooperative assembly tasks. A learning algorithm based on trial end error is used to find a sequence for each robot to assemble the goal aggregate. It is shown that a distributed learning method based on a Markov decision process is able to learn the sequences for the involved robots. A novel state-action graph is used to store the reinforcement values of the learning process. The approach is designed in a way that not only exact matches but also similar aggregates are accepted by the system. Markus Ferch, Matthias Höchsmann, Jianwei Zhang 0001 |
IROS | 3 |
| 2002 | Visual guided grasping and generalization using self-valuing learningabstractWe present a self-valuing learning technique which is capable of learning how to grasp unfamiliar objects and generalize the learned abilities. The learning system consists of two learners which distinguish between local and global grasping criteria. The local criteria are not object specific while the global criteria cover physical properties of each object. In this case we present a generalization method of the learning parameters based on a tree distance model for the medial axis transformations. The system is self-valuing, i.e. it rates its actions by evaluating sensory information and the usage of image processing techniques. An experimental setup consisting of a PUMA-260 manipulator, equipped with a hand-camera and a force/torque sensor was used to test this scheme. The system has shown the ability to grasp a wide range of objects and to apply previously learned knowledge to new objects. Bernd Rössler, Jianwei Zhang 0001, Matthias Höchsmann |
IROS | 2 |
| 2002 | Extracting compact fuzzy rules based on adaptive data approximation using B-splines
Jianwei Zhang 0001, S. Köper, Alois C. Knoll |
Inf. Sci. | 1 |
| 2001 | Computation of Fingertip Positions for a Form-Closure GraspabstractThis paper proposes a simple and efficient algorithm for computing a form-closure grasp on a 3D polyhedral object. This algorithm searches for a form-closure grasp from a "good" initial grasp in a promising search direction that pulls the convex hull of the primitive contact wrenches towards the origin of the wrench space. The "good" initial grasp is a set of contact points that minimizes the distance between the origin and the centroid of the primitive contact wrenches, and can be calculated by the quadratic programming. The local promising search direction at every step is readily determined by the ray-shooting based qualitative test algorithm developed in our early work. By using the "good" initial grasp, the iteration times of search can be significantly reduced so that a form-closure grasp can be found more efficiently. Since the algorithm adopts a local search strategy, its computational cost is less dependent on the complexity of the object surfacer. Finally, the algorithm was implemented and its efficiency ascertained by three examples. Yun-Hui Liu 0001, Jianwei Zhang 0001, Alois C. Knoll |
ICRA | 3 |
| 2001 | Asymptotic Motion Control of Robot Manipulators Using Uncalibrated Visual FeedbackabstractTo implement a visual feedback controller, it is necessary to calibrate the homogeneous transformation matrix between the robot base frame and the vision frame besides the intrinsic parameters of the vision system. The calibration accuracy greatly affects the control performance. In this paper, we address the problem of controlling a robot manipulator using visual feedback without calibrating the transformation matrix. We propose an adaptive algorithm to estimate the unknown matrix online. It is proved by the Lyapunov method that the robot motion approaches asymptotically to the desired one and the estimated matrix is bounded under the control of the proposed visual feedback controller. The performance was confirmed by simulations and experiments. Yantao Shen 0001, Yun-Hui Liu 0001, Kejie Li, Jianwei Zhang 0001, Alois C. Knoll |
ICRA | 4 |
| 2000 | A General Learning Approach to Multisensor Based Control using Statistic IndicesabstractWe propose a concept for integrating multiple sensors in real-time robot control. To increase the controller robustness under diverse uncertainties, the robot systematically generates series of sensor data (as robot state) while memorising the corresponding motion parameters. From the collection of (multi-) sensor trajectories, statistical indices for each sensor type can be extracted. If the sensor data are preselected as output relevant, the principal components can be used very efficiently to approximately represent the original perception scenarios. After this dimension reduction procedure, a nonlinear fuzzy controller can be trained to map the subspace projection into the robot control parameters. We apply the approach to a real robot system with two arms and multiple vision and force/torque sensors. These external sensors are used simultaneously to control the robot arm performing insertion and screwing operations. The experiments show that the robustness as well as the precision of robot control can be enhanced by integrating multiple additional sensors using this concept. Yorck von Collani, Markus Ferch, Jianwei Zhang 0001, Alois C. Knoll |
ICRA | 3 |
| 1999 | Modelling multivariate data by neuro-fuzzy systemsabstractThe paper proposes an approach for solving multivariate modelling problems with neuro-fuzzy systems. Instead of using selected input variables, statistical indices are extracted to feed the fuzzy controller. The original input space is transformed into an eigenspace. If a sequence of training data are sampled in a local context, a small number of eigenvectors which possess larger eigenvalues provide a good summary of all the original variables. Fuzzy controllers can be trained for mapping the input projection in the eigenspace to the outputs. Implementations with the prediction of time series validate the concept. Jianwei Zhang 0001, Alois C. Knoll |
CIFEr | 1 |
| 1999 | Robot Skill Transfer Based on B-Spline Fuzzy Controllers for Force-Control TasksabstractHuman-beings can easily describe their behaviour by IF-THEN rules, which can be transferred from one task to another with slight local changes. However, standard techniques for function approximation like neural networks or associative memories are unable to work with rules. We introduce a method for extracting and importing human readable rules from and to a B-spline fuzzy controller. Rule import is used to initialise a B-spline fuzzy controller with a priori knowledge to decrease the learning time and overcome the problem of partially trained B-spline controllers. In the experimental section we show how a set of rules for a two arm cooperation task are generated through "learning-by-doing" and transferred to a robot screwing operation. The successful experiment shows how rule-based knowledge can be used for skill transfer in similar tasks. Markus Ferch, Jianwei Zhang 0001, Alois C. Knoll |
ICRA | 2 |
| 1999 | Appearance-Based Visual Learning in a Neuro-Fuzzy Model for Fine-Positioning of ManipulatorsabstractThis paper presents an implementation of visual learning by appearance in conjunction with an adaptive nonlinear controller for fine-positioning a manipulator onto a grasping position. We use principal component analysis to reduce the dimension of raw camera images (about 10,000 pixels) to lower-dimension vectors that can be used as inputs of our neuro-fuzzy controllers. It is shown that this approach leads to a very robust system that is stable under variable environment conditions. The approach needs no camera calibration and is applied to tasks of three degrees of freedom, e.g., translating the gripper in the x-y-plane and rotating it about the z-axis. Jianwei Zhang 0001, Alois C. Knoll |
ICRA | 1 |
| 1999 | Designing fuzzy controllers by rapid learning
Jianwei Zhang 0001, Alois C. Knoll |
Fuzzy Sets Syst. | 1 |
| 1998 | A Neuro-Fuzzy Solution for Fine-Motion Control Based on Vision and Force SensorsabstractIn this paper the use of a B-spline neuro-fuzzy model for different tasks such as vision-based fine-positioning and force control is presented. It is shown that neuro-fuzzy controllers can be used not only for low-dimensional problems like force control but also for high-dimensional problems like vision-based sensorimotor control. Controllers of this type can be modularly combined to solve a given assembly problem. Yorck von Collani, Jianwei Zhang 0001, Alois C. Knoll |
ICRA | 2 |
| 1998 | Rapid On-Line Learning of Compliant Motion for Two-Arm CoordinationabstractTo enable a two-arm manipulator system to perform complex cooperating tasks, such as carrying a rigid object, establishing a contact to a surface, grinding, etc., a nonlinear multiple input system should be solved. We apply an approach to online automatic learning of a B-spline fuzzy controller. This controller model directly connects the sensor inputs to the compensation motion. By using the adaptation of control actions in all possible situations through practising of the robots in the real environment, uncertainties of the robot-object model can be taken into account. The compliant motion controller of the robot system can be adapted to new situations in a short time due to the online learning approach. Jianwei Zhang 0001, Markus Ferch |
ICRA | 1 |
| 1998 | Constructing fuzzy controllers with B-spline models - Principles and applicationsabstractIn this paper we present an approach to designing a novel type of fuzzy controller. B-spline basis functions are used for input variables and fuzzy singletons for output variables to specify linguistic terms. “Product” is chosen as the fuzzy conjunction, and “centroid” as the defuzzification method. By appropriately designing the rule base, a fuzzy controller can be interpreted as a B-spline interpolator. Such a fuzzy controller may learn to approximate any known data sequences and to minimize a certain cost function. By choosing such a function appropriately, the learning process can be made to converge rapidly. We applied this approach to the problems of function approximation and both supervised and unsupervised learning of mobile robots. Experiments validate the advantages of this approach. © 1998 John Wiley & Sons, Inc.13: 257–285, 1998 Jianwei Zhang 0001, Alois C. Knoll |
Int. J. Intell. Syst. | 1 |
| 1997 | Instructing cooperating assembly robots through situated dialogues in natural languageabstractWe present an assembly cell consisting of two cooperating robots and a variety of sensors. It offers a number of complex skills necessary for constructing aggregates from elements of a toy construction set. A high degree of flexibility was achieved because the skills were realised only through sensory feedback not by resorting to fixtures or specialised tools. The operation of the cell is completely controlled through natural language. Results from experiments in cognitive sciences and computer linguistics were incorporated to integrate natural language with vision as well as to control the construction dialogue between a human instructor and the robotic system. The experimental setup is described, and a sample dialogue demonstrates the capabilities of the cell. A brief discussion of issues for further research concludes the paper. Alois C. Knoll, Bernd Hildebrandt, Jianwei Zhang 0001 |
ICRA | 3 |
| 1997 | Online learning of B-spline fuzzy controller to acquire sensor-based assembly skillsabstractWe present an online learning approach for developing sensor-based controllers and show how it can make the programming of assembly tasks easier. We suggest combining the qualitative modelling of human expert skills and the self-tuning of control parameters so that a controller for a complex assembly task can be efficiently developed. It is then discussed how to construct a fuzzy controller with B-splines and why rapid convergence of its learning can be achieved. To apply the concept in a screwing operation, we propose several gradual steps of active online learning. Experiments were carried out with two independently controlled robot arms. Although general-purpose jaw-grippers are used and diverse uncertainties exist, the "elevator control" of a toy aircraft can be robustly built. Jianwei Zhang 0001, Yorck von Collani, Alois C. Knoll |
ICRA | 1 |
| 1996 | Modular design of fuzzy controller integrating deliberative and reactive strategiesabstractThis work presents the concept and realisation of the integration of deliberative and reactive strategies for controlling mobile robot systems. Programming at the task-level in a partially-known environment is divided into two consecutive steps: subgoal planning and subgoal-guided plan execution. A modular fuzzy control scheme is proposed, which allows independent development and flexible integration of different rule bases, each for fulfilling a certain subtask. The design process of the fuzzy controller is demonstrated with the examples of three rule bases. Jianwei Zhang 0001, Frank Wille, Alois C. Knoll |
ICRA | 1 |
| 1994 | Design of a Fuzzy Controller for Optimal Execution of Subgoal-Guided Robot MotionsabstractIn this paper, an approach is presented for designing a robot controller to integrate path planning and motion execution. Instead of following an exactly planned geometric path, which cannot be easily modified to deal with uncertainties, a robot realizes its motion under the guidance of a sequence of grossly selected crucial points called subgoals. A fuzzy model is used for on-line controlling the robot to pass through the subgoals smoothly. With the help of heuristics for optimal motion control, the rule base can be developed. Such a concept provides a control method in real time as well as a possibility of integrating sensor data for the subgoal-guided motion.> Jianwei Zhang 0001, Andreas Herp |
ICRA | 1 |
| 1994 | Robust subgoal planning and motion execution for robots in fuzzy environmentsabstractIn this paper, a fuzzy environment includes obstacles in the vicinity of robot motion, which cannot be modelled exactly or are entirely unknown, and a vague description of the relationship between the robot and the obstacles. An integrated concept for solving the problem of motion planning and control of robots in fuzzy environments is proposed. First, critical points for the desired motion are found in a connectivity network on the boundary of static obstacles. By considering the degree of uncertainty of the obstacles, the local geometry, the available free space and the path shape, these points are further adjusted to form subgoals which reliably decide the global motion direction. During on-line motion execution, a fuzzy controller evaluates sensor data, realizes the motion towards subgoals and avoids local unexpected collisions. Simulations with mobile robots show some control results using this concept.> Jianwei Zhang 0001, J. Raozkowsky |
IROS | 1 |