Bin Fang 0003

dblp:94/4033-3 · DBLP profile ↗
← Back
43ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0002-9149-7336ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 2 first-author · 9 since 2021Systems, architecture and hardware · 13 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 STOLA: Self-Adaptive Touch-Language Framework for Tactile Commonsense Reasoning in Open-Ended Scenarios
abstract
This paper explores the challenges of integrating tactile sensing into intelligent systems for multimodal reasoning, particularly in enabling commonsense reasoning about the open-ended physical world. We identify two key challenges: modality discrepancy, where existing touch-language models often treat touch as a mere sub-modality of language without further addressing the semantic differences, and open-ended tactile data scarcity, where current datasets lack the diversity, open-endedness, and complexity needed for reasoning. To overcome these challenges, we introduce SToLa, a Self-Adaptive Touch-Language framework. SToLa utilizes Mixture of Experts (MoE) to dynamically process, unify, and manage tactile and language modalities, capturing their unique characteristics. Crucially, we also present a comprehensive tactile commonsense reasoning dataset and benchmark featuring free-form questions and responses, 8 physical properties, 4 interactive characteristics, and diverse commonsense knowledge. Experiments show SToLa exhibits competitive performance compared to existing models on the PHYSICLEAR benchmark and self-constructed datasets, proving the effectiveness of the Mixture of Experts architecture in multimodal management and the performance advantages for open-scenario tactile commonsense reasoning tasks.
Jin An Xu, Jialing Chen, Bin Fang 0003, Wenjuan Han
AAAI4
2026 G-Anomaly: A Pyramid Graph Transformer-Based Vision-Language Model for General Industrial Anomaly Detection
abstract
Visual-language alignment is crucial for enhancing the domain adaptability of industrial anomaly detection models. However, the existing methods overlook the importance of structured image representation, fail to further distinguish topological differences between anomalies and the inherent textures of products, which reduces the accuracy of semantic matching. To address this problem, we propose a novel industrial anomaly detection model G-Anomaly, to preserve the topological structure of the sample images and further enhance the model’s domain adaptability. We designed Pyramid Graph Transformer as a visual encoder to extract multi-scale visual features, which can directly preserve the structural relationships between different regions of the image, and also optimize the over-smoothing issue present in deep graph networks, thereby retaining the distinguishability of anomalous nodes. Additionally, we design a Multi-level Domain Adapter that ensures semantic consistency of anomalous features across different scales and contexts by performing visual-language matching at various resolutions and levels of abstraction. This enhances the model’s domain adaptability for anomaly detection for a wide range of industrial products. We collect and craft an actual solar panel dataset PV_actual AD, and conduct extensive experiments on the public dataset MVTec AD as well as the actual solar panel dataset PV_actual AD. This has demonstrated that G-Anomaly not only performs well in standard testing environments but also exhibits robustness and domain adaptability for anomaly detection tasks in real-world scenarios.
Jiaqi Li 0016, Shuhuan Wen, Bin Fang 0003
IEEE Trans Autom. Sci. Eng.3
2026 Personalized Lumbar Vertebrae Modeling for Dynamic Assessment of Idiopathic Scoliosis
abstract
Clinical assessment of idiopathic scoliosis (IS) patients primarily relies on static imaging techniques. Dynamic digital human (DDH) can provide comprehensive spatio-temporal information for dynamic assessment of the scoliotic spine in IS patients comparing with static imaging techniques, such as X-ray for general assessment and computed tomography (CT) for surgical planning. The lumbar vertebrae exhibit greater morphological variability than the thoracic region when subjected to different postures and mechanical loads, making them particularly important for dynamic assessment. Therefore, a personalized lumbar vertebrae model (PLVM) is proposed in this work to simulate lumbar vertebrae motion for IS patients; furthermore, an individualized DDH (i-DDH) is proposed by embedding PLVM into DDH to capture the spatio-temporal information. First, we use a bone primitive generation method to construct the DDH by incorporating Neural Radiance Fields (NeRF) and three-dimensional (3D) Gaussian splatting methods. Next, we develop the PLVM generation method to simulate lumbar vertebrae motion under different loads and postures. Finally, the bone primitives and PLVM are merged to generate the i-DDH for dynamic assessment. We validated i-DDH using multi-posture radiographs from eight IS patients awaiting surgery. The results demonstrate high accuracy compared to state-of-the-art (SOTA) models, with a mean angular error of 0.96$^\circ$ and a maximum error of 3.6$^\circ$ relative to radiographs. The proposed i-DDH framework is able to capture the spinal posture and conduct the dynamic assessment of IS patients rather than fixed positions. It overcomes the soft tissue artifact (STA) problem from motion capture systems and the failure to generate 3D spinal curvature of IS patients by training healthy subjects from computer vision methods. It also shows great clinical significance for preoperative planning and clinical assessment by providing dynamic spinal posture that cannot be achieved with static imaging.
Chengyin Wang, Jianfeng Li 0007, Shuo Wang 0028, Mingjie Dong, Bin Fang 0003, Qianyu Zhuang
IEEE J. Biomed. Health Informatics7
2026 A Multigranularity Fuzzy Inference Approach for Out-of-Distribution Detection in Fault Diagnosis
abstract
The intelligent fault diagnosis has achieved notable success in identifying known mechanical failures; however, reliably detecting out-of-distribution (OOD) faults remains a key challenge to achieve the diagnostic robustness. In industrial applications, vibration signals are typically collected as time-series data whose dynamic characteristics vary with load, speed, and environmental interference, with weak early fault patterns that blur class boundaries. As a result, models trained under limited laboratory conditions inevitably encounter unseen OOD inputs after deployment, requiring the ability to recognize and reject them reliably. Existing representation- and similarity-based OOD methods have shown promise but typically rely on single-granularity prototypes, capturing only coarse similarity structures and overlooking latent subclass relations—thus limiting the generalization under complex degradation modes. To address these limitations, we propose a multigranularity fuzzy inference (MgFI) framework for enhanced uncertainty quantification in fault diagnosis. MgFI models fine-grained subclass memberships on a hyperspherical manifold, aggregates them into class-level fuzzy sets, and infers coarse-grained In-distribution (ID) confidence through the hierarchical fuzzy reasoning. Extensive experiments demonstrate that MgFI substantially improves the OOD detection accuracy and provides a principled, interpretable framework for trustworthy open-set industrial diagnostics.
Fir Dunkin, Xinde Li, Bin Fang 0003, Guoliang Wu, Tao Shen 0004, Bing Li 0033, Shuzhi Sam Ge
IEEE Trans. Syst. Man Cybern. Syst.3
2025 AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors
abstract
Visuo-tactile sensors aim to emulate human tactile perception, enabling robots to precisely understand and manipulate objects. Over time, numerous meticulously designed visuo-tactile sensors have been integrated into robotic systems, aiding in completing various tasks. However, the distinct data characteristics of these low-standardized visuo-tactile sensors hinder the establishment of a powerful tactile perception system. We consider that the key to addressing this issue lies in learning unified multi-sensor representations, thereby integrating the sensors and promoting tactile knowledge transfer between them. To achieve unified representation of this nature, we introduce TacQuad, an aligned multi-modal multi-sensor tactile dataset from four different visuo-tactile sensors, which enables the explicit integration of various sensors. Recognizing that humans perceive the physical environment by acquiring diverse tactile information such as texture and pressure changes, we further propose to learn unified multi-sensor representations from both static and dynamic perspectives. By integrating tactile images and videos, we present AnyTouch, a unified static-dynamic multi-sensor representation learning framework with a multi-level structure, aimed at both enhancing comprehensive perceptual abilities and enabling effective cross-sensor transfer. This multi-level architecture captures pixel-level details from tactile data via masked modeling and enhances perception and transferability by learning semantic-level sensor-agnostic features through multi-modal alignment and cross-sensor matching. We provide a comprehensive analysis of multi-sensor transferability, and validate our method on various offline datasets and in the real-world pouring task. Experimental results show that our method outperforms existing methods, exhibits outstanding static and dynamic perception capabilities across various sensors. The code, TacQuad dataset and AnyTouch model are fully available at gewu-lab.github.io/AnyTouch/.
Ruoxuan Feng, Jiangyu Hu, Wenke Xia, Tianci Gao, Bin Fang 0003, Di Hu 0001
ICLR7
2025 EHC-MM: Embodied Holistic Control for Mobile Manipulation
abstract
Mobile manipulation typically entails the base for mobility, the arm for accurate manipulation, and the camera for perception. The principle of Distant Mobility, Close Grasping(DMCG) is essential for holistic control. We propose Embod-ied Holistic Control for Mobile Manipulation(EHC-MM) with the embodied function of sig($\omega$): By formulating the DMCG principle as a Quadratic Programming (QP) problem, sig($\omega$) dynamically balances the robot's emphasis between movement and manipulation with the consideration of the robot's state and environment. In addition, we propose the Monitor-Position-Based Servoing (MPBS) with sig($\omega$), enabling the tracking of the target during the operation. This approach enables coordinated control among the robot's base, arm, and camera, enhancing task efficiency. Through extensive simulations and real-world experiments, our approach significantly improves both the success rate and efficiency of mobile manipulation tasks, achieving a 95.6% success rate in real-world scenarios and a 52.8% increase in time efficiency.
Yixiang Jin, Jun Shi 0008, Yong A, Dingzhe Li, Fuchun Sun 0001, Dingsheng Luo, Bin Fang 0003
ICRA8
2025 MagicGel: A Novel Visual-Based Tactile Sensor Design with Magnetic Gel
abstract
Force estimation is the core indicator for evaluating the performance of tactile sensors, and it is also the key technical path to achieving precise force feedback mechanisms. This study proposes a design method for a visual tactile sensor (VBTS) that integrates a magnetic perception mechanism, and develops a new tactile sensor called MagicGel. The sensor uses strong magnetic particles as markers and captures magnetic field changes in real time through Hall sensors. On this basis, MagicGel achieves the coordinated optimization of multimodal perception capabilities: it not only has fast response characteristics, but also can perceive non-contact status information of home electronic products. Specifically, MagicGel simultaneously analyzes the visual characteristics of magnetic particles and the multimodal data of changes in magnetic field intensity, ultimately improving force estimation capabilities.
Jianhua Shan, Jiangduo Liu, Xiangbo Wang, Ziwei Xia, Guangzeng Chen, Guangyuan Xu, Bin Fang 0003
IROS9
2025 Soft Contact Simulation and Manipulation Learning of Deformable Objects With Vision-Based Tactile Sensor
abstract
Deformable object manipulation is a challenging problem due to its complex deformable properties. With the development of artificial intelligence, learning-based methods have shown outstanding performance in robotic manipulation. Previous works have investigated the manipulation of deformable objects via Reinforcement Learning (RL) in simulation. However, they approximate object deformation with particles, using particle states as observations, which are unavailable in reality. To address these issues, we utilize Vision-Based Tactile Sensors (VBTSs) as the end-effector to manipulate and observe the deformable objects. In this work, we develop a new contact simulation environment for deformable objects, including elastic, plastic, and elastoplastic. We utilize RL strategies and expert demonstrations to train agents in the simulation. Finally, we build a real experimental platform to complete the sim-to-real tasks and robustness testing. Our work introduces an innovative strategy that utilizes high-resolution VBTSs for contact simulation and manipulation of deformable objects. The experimental results show superior performances of deformable object manipulation with the proposed method.
Shixin Zhang, Zixi Chen 0002, Zirong Shen, Fuchun Sun 0001, Cesare Stefanini, Di Guo 0002, Shan Luo 0001, Jianwei Zhang 0001, Jianhua Shan, Bin Fang 0003
IEEE Trans Autom. Sci. Eng.11
2024 Bionic Soft Fingers with Hybrid Variable Stiffness Mechanisms for Multimode Grasping
abstract
This paper presents a novel Bionic Soft Finger (BSF) that aims to overcome the limitations of conventional rigid manipulators in terms of adaptability and safety, as well as the challenges faced by soft hands regarding carrying capacity and stability. The BSF design uses a hybrid variable stiffness mechanism combining memory alloy actuators with particle jamming to achieve the desired bending angle and actuator stiffness. Our innovative approach utilizes a bionic finger design that incorporates a memory alloy skeleton and a water-cooled recirculation system, leading to a substantial reduction in the time required for each operation. Through the integration of particle jamming, we have enhanced the overall stiffness and performance of the manipulator, enabling load capacities of up to 3N per finger and more than twice the stiffness of a normal condition. Additionally, our design enables multimode grasping and incorporates a liquid metal strain sensor (METT) for real-time monitoring of finger bending angles. Comparative analyses demonstrate that our design exhibits superior stiffness and enables five-mode grasping in comparison to pneumatic actuators. We believe that bionic soft fingers present a promising solution for enhancing adaptability, safety, and performance in human-robot interaction applications.
Xiangbo Wang, Hongze Yu, Zhenwei Wen, Lide Fang, Huaping Liu 0001, Fuchun Sun 0001, Lixue Tang, Bin Fang 0003
ICRA9
2024 Skill enhancement learning with knowledge distillation
Naijun Liu, Fuchun Sun 0001, Bin Fang 0003, Huaping Liu 0001
Sci. China Inf. Sci.3
2024 Digital-Twin-Assisted Skill Learning for 3C Assembly Tasks
abstract
The utilization of robots in computer, communication, and consumer electronics (3C) assembly has the potential to significantly reduce labor costs and enhance assembly efficiency. However, many typical scenarios in 3C assembly, such as the assembly of flexible printed circuits (FPCs), involve complex manipulations with long-horizon steps and high-precision requirements that cannot be effectively accomplished through manual programming or conventional skill-learning methods. To address this challenge, this article proposes a learning-based framework for the acquisition of complex 3C assembly skills assisted by a multimodal digital-twin environment. First, we construct a fully equivalent digital-twin environment based on the real-world counterpart, equipped with visual, tactile force, and proprioception information, and then collect multimodal demonstration data using virtual reality (VR) devices. Next, we construct a skill knowledge base through multimodal skill parsing of demonstration data, resulting in primitive policy sequences for achieving 3C assembly tasks. Finally, we train primitive policies via a combination of curriculum learning, residual reinforcement learning, and domain randomization methods and transfer the learned skill from the digital-twin environment to the real-world environment. The experiments are conducted to verify the effectiveness of our proposed method.
Fuchun Sun 0001, Naijun Liu, Ruize Sun, Shengyi Miao, Zengxin Kang, Bin Fang 0003, Huaping Liu 0001, Yongjia Zhao, Haiming Huang
IEEE Trans. Cybern.7
2023 Implementation and Optimization of Grasping Learning with Dual-modal Soft Gripper
abstract
Robust and efficient grasping of different objects is still an open problem due to the difficulty of integrating multidisciplinary knowledge such as gripper ontology design, perception, control, and learning. In recent years, learning-based methods have achieved excellent results in grasping various novel objects. However, current methods are usually limited to a single grasping mode or rely on different end effectors to grasp objects of different shapes. For human beings, our hands are capable of grasping various objects with changes in grasping methods and form of hands. In light of this, developing a gripper with similar performance could possibly improve the robot's gripping ability. In this paper, we design a dual-modal soft gripper (DSG) and propose a deep reinforcement learning (DRL) framework to implement the operations. Both of our grasping modes, namely enveloping and pinching, are achieved through the tendon drive system and the deformation of the spring steel plate, which enables the gripper to switch between the two grasping modes in real time. We also combined the cutting-edge achievements of deep learning and reinforcement learning to design an autonomous grasping algorithm based on Q-learning and a deep Q network. Moreover, to fully utilize the visual input from the sensor, we added semantic embeddings of target objects to facilitate the learning, which is especially useful in deciding the grasping method for objects previously unseen. We also evaluate our DRL framework in different scenarios, offering a detailed comparison of each grasping mode and the mixed method (with or without semantic information). Our design has proved efficient in reducing the number of failing grasping actions and improving the success rate when facing novel and tricky objects.
Feihan Li, Xingyu Ding, Fuchun Sun 0001, Jianhua Shan, Lincheng Li, Bin Fang 0003
ICRA10
2023 Autonomous Oropharyngeal-Swab Robot System for COVID-19 Pandemic
abstract
The outbreak of COVID-19 has led to the shortage of medical personnel and the increasing need for nucleic acid testing. Manual oropharyngeal sampling is susceptible to inconsistency caused by fatigue and close contact could also cause healthcare personnel exposure and cross infection. The innate deficiency calls for a safer and more consistent way to collect the oropharyngeal samples. Therefore a fully autonomous oropharyngeal-swab robot system is proposed in this paper. The system is installed in a negative pressure chamber and carrying out a standardized sampling process to minimize individual sampling differences. A hierarchical throat detection algorithm is presented and multiple modality sensory information are fused to safely and accurately localize the optimum sampling location. Also, a force/position hybrid control method is adopted to ensure both accurate sampling and subject comfort. The robot system described in this paper can safely and efficiently collect the oropharyngeal sample, providing a scalable solution for large-scale Polymerase Chain Reaction (PCR) Molecular sample collection for various respiratory diseases. Note to Practitioners—During the COVID-19 pandemic, pre-diagnostic is essential for both prevention and treatment. Existing approaches, including nasal swab and oropharyngeal-swab, require extensive medical worker training and increase the chance of cross-infection. The robot system introduced in this paper can take oropharyngeal-swab samples from subjects with minimum human intervention, reducing medical worker exposure, alleviating the work pressure of medical staff, and speed up large quantity of sampling plan. The robot will first guide the subject into position with vocal commands, and automatically detect the optimum sampling location with a real-time machine learning algorithm. A dedicated control strategy aiming at minimizing discomfort and uniforming sample quantity is then applied to safely collect nucleic samples from the throat. Eventually, while the swab is being stored in the culture medium, a disinfection process is carried out simultaneously to prepare the robot for the next subject. Preliminary clinical trials show that our robot system can safely and accurately collect samples from subjects.
Fuchun Sun 0001, Huaping Liu 0001, Bin Fang 0003
IEEE Trans Autom. Sci. Eng.5
2023 Hybrid Robotic Grasping With a Soft Multimodal Gripper and a Deep Multistage Learning Scheme
abstract
Grasping has long been considered an important and practical task in robotic manipulation. Yet achieving robust and efficient grasps of diverse objects is challenging, since it involves gripper design, perception, control, and learning, etc. Recent learning-based approaches have shown excellent performance in grasping a variety of novel objects. However, these methods either are typically limited to one single grasping mode or else more end effectors are needed to grasp various objects. In addition, gripper design and learning methods are commonly developed separately, which may not adequately explore the ability of a multimodal gripper. In this article, we present a deep reinforcement learning (DRL) framework to achieve multistage hybrid robotic grasping with a new soft multimodal gripper. A soft gripper with three grasping modes (i.e.,enveloping,sucking, andenveloping_then_sucking) can both deal with objects of different shapes and grasp more than one object simultaneously. We propose a novel hybrid grasping method integrated with the multimodal gripper to optimize the number of grasping actions. We evaluate the DRL framework under different scenarios (i.e., with different ratios of objects of two grasp types). The proposed algorithm is shown to reduce the number of grasping actions (i.e., enlarge the grasping efficiency, with maximum values of 161.0% in simulations, and 153.5% in real-world experiments) compared to single grasping modes.
Fukang Liu, Fuchun Sun 0001, Bin Fang 0003, Xiang Li 0009, Songyu Sun, Huaping Liu 0001
IEEE Trans. Robotics3
2022 Non-destructive Fruit Firmness Evaluation Using Vision-Based Tactile Information
abstract
During postharvest storage, fruit firmness usually decreases due to respiration and bruise, the former of which indicates the fruit ripeness while the latter negatively influence consumers' taste preference. This paper presents a portable and low-cost device using vision-based tactile information to evaluate fruit firmness in a non-destructive manner. The device consists of a camera, LED lights, and a soft sensing layer with small bumps to capture detailed tactile information of the fruit. Two working modes are designed and a CNN-LSTM architecture is developed to relate the tactile information to fruit overall firmness or detect local firmness distortion. According to the experimental results, an R2 up to 92.9% was achieved for the evaluation of the overall firmness of Cuixiang kiwifruit, and accuracy of 98.0% was obtained for the detection of local firmness distortion of Fuji apples. These results demonstrate the efficacy of the proposed solution to evaluate fruit firmness featuring high precision, and its non-destructive and potable nature is also anticipated to be favorable by the fresh fruit market.
Yaohui Chen 0002, Bin Fang 0003, Fuchun Sun 0001, Shanjun Li
ICRA4
2022 Visual Affordance Guided Tactile Material Recognition for Waste Recycling
abstract
Because more and more solid waste is generated, in particular, in cities, the management of solid waste disposal has become a global challenge. A solution is to find an effective way to sort solid waste materials and recycle them into reusable products. In this article, we propose to use a material recognition method with vision-guided tactile to form a robotic system for waste sorting. The vision guidance module integrates an object detector and an affordance network together. It allows the robot to not only detect the desired containers and packaging from an assortment of the waste but also obtain a configuration to grasp the target and actively collect its tactile data. By classifying the object with the tactile data, the robot can sort containers and packaging into their respective categories according to the type of material. Our experimental results demonstrate the effectiveness of the proposed robotic waste sorting system in sorting containers and packaging various types of materials. Note to Practitioners—The management of waste has become a great challenge in environmental protection in the world. We propose a robotic waste sorting system, which utilizes a vision-guided tactile sensing approach to find target waste and sort the waste according to materials. A visual module is used to find the waste of interest. A class-specific affordance map is generated to guide a robotic hand to actively grasp waste and collect the tactile data for material recognition. With the recorded tactile data, the robot can recognize the material of the target waste and sort the waste according to the type of material. The proposed system demonstrates good performance in waste sorting, and it can be easily implemented in practical scenarios. The target material can be generalized to a broader group of objects.
Di Guo 0002, Huaping Liu 0001, Bin Fang 0003, Fuchun Sun 0001, Wuqiang Yang
IEEE Trans Autom. Sci. Eng.3
2021 Fabric defect detection using tactile information
abstract
This paper proposes a method of fabric structure defect detection based on tactile information. Different from traditional visual-based detection methods, the proposed method uses a tactile sensor to attain the information of fabric. The advantage of using tactile information is to avoid different irregular dyeing patterns and reduce the influence of ambient light. Therefore, the proposed method can be more concise and universal, which makes the defect detection system more robust. Experiments are conducted to verify the performance of the developed tactile method. The results demonstrate that the proposed method is much better than the visual method in the detection of structural defects. In addition, we propose a design of a tactile sensing device that can greatly improve the actual detection efficiency of the tactile method. It is showed that the proposed method is efficient enough to meet our requirements and it provides a possible solution for the application of the tactile method in actual fabric production.
Xingming Long, Bin Fang 0003, GuoYi Luo, Fuchun Sun 0001
ICRA2
2021 Elastic Tactile Simulation Towards Tactile-Visual Perception
abstract
Tactile sensing plays an important role in robotic perception and manipulation tasks. To overcome the real-world limitations of data collection, simulating tactile response in a virtual environment comes as a desirable direction of robotic research. In this paper, we propose Elastic Interaction of Particles (EIP) for tactile simulation, which is capable of reflecting the elastic property of the tactile sensor as well as characterizing the fine-grained physical interaction during contact. Specifically, EIP models the tactile sensor as a group of coordinated particles, and the elastic property is applied to regulate the deformation of particles during contact. With the tactile simulation by EIP, we further propose a tactile-visual perception network that enables information fusion between tactile data and visual images. The perception network is based on a global-to-local fusion mechanism where multi-scale tactile features are aggregated to the corresponding local region of the visual modality with the guidance of tactile positions and directions. The fusion method exhibits superiority regarding the 3D geometric reconstruction task. Our code for EIP is available at https://github.com/yikaiw/EIP.
Yikai Wang 0001, Wenbing Huang 0001, Bin Fang 0003, Fuchun Sun 0001
ACM Multimedia3
2021 Toward Image-to-Tactile Cross-Modal Perception for Visually Impaired People
abstract
It is still a great challenge for the visually impaired people to perceive their surroundings from a global perspective, which makes it difficult for them to interact with unfamiliar environments. The reason is that these conventional assisting devices only address the obstacle avoidance problem. They do not provide visually impaired people with a global perception of the surrounding environment. In this article, a new generative adversarial network (GAN) model is developed to effectively transform the ground images into the tactile signal, which can be displayed by an off-the-shelf vibration device. The algorithm module and the hardware are integrated into a portable device, which provides visually impaired people with effective surrounding perception capability. In addition, a visual-tactile cross-modal data set is constructed to train the proposed deep-learning architecture. Experimental results show that the proposed system can help visually impaired people sense the ground and bring a better traveling experience for them. Note to Practitioners-This article presents a portable device that provides tactile recognition assistance for visually impaired people. Such a technology can be extensively used in tactile mouse and white cane. The developed technology can be extensively used for various industrial applications, such as surrounding monitoring and manipulation. The proposed work demonstrates the promising ability of artificial intelligence in healthcare applications. The generated tactile signals are expected to be used in many human-centered systems, and we believe that our contribution is an important step toward the development of a more comprehensive assisting technology for visually impaired people.
Huaping Liu 0001, Di Guo 0002, Xinyu Zhang 0001, Wenlin Zhu, Bin Fang 0003, Fuchun Sun 0001
IEEE Trans Autom. Sci. Eng.5
2021 An Interactive Perception Method for Warehouse Automation in Smart Cities
abstract
The smart city is an integrated environment that heavily relies on intelligent robots, which provides the basis for the warehouse automation. However, a warehouse is a typical unstructured environment, and robotic grasp and manipulation are extremely important for the package, transfer, search, and so on. Currently, the most usual method is to detect the picking or grasping points for some specific end-effector including suction cup, gripper, or robotic hand. The manipulation performance is, therefore, strongly influenced by the visual detector. To tackle this problem, the affordance map has recently been developed. It characterizes the operation possibilities afforded by the operation scene and has been used for several grasp tasks. Nevertheless, the conventional affordance method often fails in complicated environments due to the mistake calculation results. In this article, we develop a novel framework to integrate the interactive exploration with a composite robotic hand for robotic grasping in a complicated environment. The exploration strategy is obtained by a deep reinforcement learning procedure. The developed new composite hand, which integrates the suction cup and grippers, is used to test the merits of the proposed interactive perception method. Experimental results show the proposed method significantly increases the manipulation efficiency and may bring great economic and social and benefits for smart cities.
Huaping Liu 0001, Yuhong Deng, Di Guo 0002, Bin Fang 0003, Fuchun Sun 0001, Wuqiang Yang
IEEE Trans. Ind. Informatics4
2020 Reinforcement Learning from Imperfect Demonstrations under Soft Expert Guidance
abstract
In this paper, we study Reinforcement Learning from Demonstrations (RLfD) that improves the exploration efficiency of Reinforcement Learning (RL) by providing expert demonstrations. Most of existing RLfD methods require demonstrations to be perfect and sufficient, which yet is unrealistic to meet in practice. To work on imperfect demonstrations, we first define an imperfect expert setting for RLfD in a formal way, and then point out that previous methods suffer from two issues in terms of optimality and convergence, respectively. Upon the theoretical findings we have derived, we tackle these two issues by regarding the expert guidance as a soft constraint on regulating the policy exploration of the agent, which eventually leads to a constrained optimization problem. We further demonstrate that such problem is able to be addressed efficiently by performing a local linear search on its dual form. Considerable empirical evaluations on a comprehensive collection of benchmarks indicate our method attains consistent improvement over other RLfD counterparts.
Mingxuan Jing, Xiaojian Ma 0001, Wenbing Huang 0001, Fuchun Sun 0001, Chao Yang 0026, Bin Fang 0003, Huaping Liu 0001
AAAI6
2020 Reusing Discriminators for Encoding: Towards Unsupervised Image-to-Image Translation
abstract
Unsupervised image-to-image translation is a central task in computer vision. Current translation frameworks will abandon the discriminator once the training process is completed. This paper contends a novel role of the discriminator by reusing it for encoding the images of the target domain. The proposed architecture, termed as NICE-GAN, exhibits two advantageous patterns over previous approaches: First, it is more compact since no independent encoding component is required; Second, this plug-in encoder is directly trained by the adversary loss, making it more informative and trained more effectively if a multi-scale discriminator is applied. The main issue in NICE-GAN is the coupling of translation with discrimination along the encoder, which could incur training inconsistency when we play the min-max game via GAN. To tackle this issue, we develop a decoupled training strategy by which the encoder is only trained when maximizing the adversary loss while keeping frozen otherwise. Extensive experiments on four popular benchmarks demonstrate the superior performance of NICE-GAN over state-of-the-art methods in terms of FID, KID, and also human preference. Comprehensive ablation studies are also carried out to isolate the validity of each proposed component. Our codes are available at https://github.com/alpc91/NICE-GAN-pytorch.
Runfa Chen, Wenbing Huang 0001, Binghui Huang, Fuchun Sun 0001, Bin Fang 0003
CVPR5
2020 Cross-Modal Zero-Shot-Learning for Tactile Object Recognition
abstract
In this paper, we address the learning problem of classifying untouched tactile instance with the help of visual modality. The proposed method is based on dictionary learning and we impose different penalty terms on coding vectors between visual and tactile modalities. Using such structured coding vectors, the visual-tactile cross-modal transfer can be achieved. A set of optimization algorithms are developed to obtain the solutions of the proposed optimization problems. After then, we can use the obtained dictionary to predict the coding vectors of the new untouched tactile samples and further determine its label. Finally, we perform extensive experimental evaluations on publicly available datasets to show the effectiveness of the proposed method.
Huaping Liu 0001, Fuchun Sun 0001, Bin Fang 0003, Di Guo 0002
IEEE Trans. Syst. Man Cybern. Syst.3
2019 Attention-based Transfer Learning for Brain-computer Interface
abstract
Different functional areas of the human brain play different roles in brain activity, which has not been paid sufficient research attention in the brain-computer interface (BCI) field. This paper presents a new approach for electroencephalography (EEG) classification that applies attention-based transfer learning. Our approach considers the importance of different brain functional areas to improve the accuracy of EEG classification, and provides an additional way to automatically identify brain functional areas associated with new activities without the involvement of a medical professional. We demonstrate empirically that our approach out-performs state-of-the-art approaches in the task of EEG classification, and the results of visualization indicate that our approach can detect brain functional areas related to a certain task.
Chuanqi Tan, Fuchun Sun 0001, Tao Kong, Bin Fang 0003
ICASSP4
2019 Lifelong Learning for Heterogeneous Multi-Modal Tasks
abstract
In this work, we investigate the lifelong learning problem from the viewpoint of heterogeneous multi-modal fusion. The main challenges come from the fact that the common representation between heterogeneous modalities should be persistently learned and the learned classifier for each multi-modal task should be persistently updated. To address this problem, we construct a multi-modal lifelong learning framework which deals with the consecutive multi-modal learning tasks and develop an efficient online dictionary learning algorithm to solve the multi-modal lifelong learning problem. Finally, we perform experimental validation on a complicated material recognition task and show the promising results.
Huaping Liu 0001, Fuchun Sun 0001, Bin Fang 0003
ICRA3
2019 Vision-based Teleoperation of Shadow Dexterous Hand using End-to-End Deep Neural Network
abstract
In this paper, we present TeachNet, a novel neural network architecture for intuitive and markerless vision-based teleoperation of dexterous robotic hands. Robot joint angles are directly generated from depth images of the human hand that produce visually similar robot hand poses in an end-to-end fashion. The special structure of TeachNet, combined with a consistency loss function, handles the differences in appearance and anatomy between human and robotic hands. A synchronized human-robot training set is generated from an existing dataset of labeled depth images of the human hand and simulated depth images of a robotic hand. The final training set includes 400K pairwise depth images and joint angles of a Shadow C6 robotic hand. The network evaluation results verify the superiority of TeachNet, especially regarding the high-precision condition. Imitation experiments and grasp tasks teleoperated by novice users demonstrate that TeachNet is more reliable and faster than the state-of-the-art vision-based teleoperation method.
Shuang Li 0014, Xiaojian Ma 0001, Hongzhuo Liang, Michael Görner, Philipp Ruppel, Bin Fang 0003, Fuchun Sun 0001, Jianwei Zhang 0001
ICRA6
2019 PointNetGPD: Detecting Grasp Configurations from Point Sets
abstract
In this paper, we propose an end-to-end grasp evaluation model to address the challenging problem of localizing robot grasp configurations directly from the point cloud. Compared to recent grasp evaluation metrics that are based on handcrafted depth features and a convolutional neural network (CNN), our proposed PointNetGPD is lightweight and can directly process the 3D point cloud that locates within the gripper for grasp evaluation. Taking the raw point cloud as input, our proposed grasp evaluation network can capture the complex geometric structure of the contact area between the gripper and the object even if the point cloud is very sparse. To further improve our proposed model, we generate a large-scale grasp dataset with 350k real point cloud and grasps with the YCB object set for training. The performance of the proposed model is quantitatively measured both in simulation and on robotic hardware. Experiments on object grasping and clutter removal show that our proposed model generalizes well to novel objects and outperforms state-of-the-art methods. Code and video are available at https://lianghongzhuo.github.io/PointNetGPD.
Hongzhuo Liang, Xiaojian Ma 0001, Shuang Li 0014, Michael Görner, Song Tang 0001, Bin Fang 0003, Fuchun Sun 0001, Jianwei Zhang 0001
ICRA6
2019 Deep Reinforcement Learning for Robotic Pushing and Picking in Cluttered Environment
abstract
In this paper, a novel robotic grasping system is established to automatically pick up objects in cluttered scenes. A composite robotic hand composed of a suction cup and a gripper is designed for grasping the object stably. The suction cup is used for lifting the object from the clutter first and the gripper for grasping the object accordingly. We utilize the affordance map to provide pixel-wise lifting point candidates for the suction cup. To obtain a good affordance map, the active exploration mechanism is introduced to the system. An effective metric is designed to calculate the reward for the current affordance map, and a deep Q-Network (DQN) is employed to guide the robotic hand to actively explore the environment until the generated affordance map is suitable for grasping. Experimental results have demonstrated that the proposed robotic grasping system is able to greatly increase the success rate of the robotic grasping in cluttered scenes.
Yuhong Deng, Yixuan Wei, Kai Lu 0003, Bin Fang 0003, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001
IROS5
2019 A glove-based system for object recognition via visual-tactile fusion
Bin Fang 0003, Fuchun Sun 0001, Huaping Liu 0001, Chuanqi Tan, Di Guo 0002
Sci. China Inf. Sci.1
2019 A novel multi-modal tactile sensor design using thermochromic material
Fuchun Sun 0001, Bin Fang 0003, Hongxiang Xue, Huaping Liu 0001, Haiming Huang
Sci. China Inf. Sci.2
2019 Interactive video summarization with human intentions
Huaping Liu 0001, Fuchun Sun 0001, Xinyu Zhang 0001, Bin Fang 0003
Multim. Tools Appl.4
2019 Surface Material Retrieval Using Weakly Paired Cross-Modal Learning
abstract
In this paper, we investigate the cross-modal material retrieval problem, which permits the user to submit a multimodal query including tactile and auditory modalities, and retrieve the image results of visual modalities. Since multiple significantly different modalities are involved in this process, we encounter more challenges compared with the existing cross-modal retrieval tasks. Our focus is to learn cross-modal representations when the modalities are significantly different and with minimal supervision. A novelty is that we establish a framework that deals with weakly paired multimodal fusion method for heterogenous tactile and auditory modalities and weakly paired cross-modal transfer for visual modality. A structured dictionary learning method with a low rank and common classifier is developed to obtain the modal-invariant representation. Finally, some cross-modal validations on publicly available data sets are performed to show the advantages of the proposed method.Note to Practitioners—Cross-modal retrieval is an important task for industrial intelligence. In this paper, we establish a framework to effectively solve the cross-modal material retrieval problem. In the developed framework, the user may submit a multimodal query including acceleration and sound about an object, and the system may return the most relevant retrieved images. Such a framework may find extensive applications in many fields, because it can be flexible to deal with a multiple-modal query and uses the minimal category label supervision without the need of strong sample pairing information between modalities. Compared with the previous material analysis systems, this paper goes beyond previously proposed surface material classification approaches as it returns an ordered list of perceptually similar surface materials for a query.
Huaping Liu 0001, Feng Wang 0034, Fuchun Sun 0001, Bin Fang 0003
IEEE Trans Autom. Sci. Eng.4
2019 Feature Pyramid Reconfiguration With Consistent Loss for Object Detection
abstract
Taking the feature pyramids into account has become a crucial way to boost the object detection performance. While various pyramid representations have been developed, previous works are still inefficient to integrate the semantical information over different scales. Moreover, recent object detectors are suffering from accurate object location applications, mainly due to the coarse definition of the "positive" examples at training and predicting phases. In this paper, we begin by analyzing current pyramid solutions, and then propose a novel architecture by reconfiguring the feature hierarchy in a flexible yet effective way. In particular, our architecture consists of two lightweight and trainable processes: global attention and local reconfiguration. The global attention is to emphasize the global information of each feature scale, while the local reconfiguration is to capture the local correlations across different scales. Both the global attention and local reconfiguration are non-linear and thus exhibit more expressive ability. Then, we discover that the loss function for object detectors during training is the central cause of the inaccurate location problem. We propose to address this issue by reshaping the standard cross entropy loss such that it focuses more on accurate predictions. Both the feature reconfiguration and the consistent loss could be utilized in popular one-stage (SSD, RetinaNet) and two-stage (Faster R-CNN) detection frameworks. Extensive experimental evaluations on PASCAL VOC 2007, PASCAL VOC 2012 and MS COCO datasets demonstrate that, our models achieve consistent and significant boosts compared with other state-of-the-art methods.
Fuchun Sun 0001, Tao Kong, Wenbing Huang 0001, Chuanqi Tan, Bin Fang 0003, Huaping Liu 0001
IEEE Trans. Image Process.5
2019 Learning to Grasp Familiar Objects Based on Experience and Objects' Shape Affordance
abstract
Stably grasping objects for a specific task is a hot research topic in robotics due to multiple degrees of freedom of hand kinematics, various shapes of objects, and incomplete visual sensing of objects (partial point clouds). This paper proposes an effective grasp planning method by integrating the crucial grasp cues (positions and orientations of thumb fingertips and the wrist) from humans’ grasp experience. This approach has multiple advantages: greatly reducing the search space of the hand kinematics; no reconstruction or registration; being able to directly perform on the partial point cloud of objects. Meanwhile, for various shapes of objects which are partially observable in the single-view visual sensing, the presented approach learns the “thumb” grasp point employing a signature of histograms of orientations shape descriptor based on objects’ category level. This method recognizes the grasp point according to the shape affordance at each point on the object, which performs the grasp point generalization on the familiar objects. Finally, we verify the developed methods via both simulations and experiments by grasping various shapes of objects.
Chunfang Liu, Bin Fang 0003, Fuchun Sun 0001, Xiaoli Li 0011, Wenbing Huang 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2019 Kernel Regularized Nonlinear Dictionary Learning for Sparse Coding
abstract
For most sparse coding methods, data samples are first encoded as hand-crafted features, followed by another separate learning step that generates dictionary and sparse codes. However, such feature representations may not be optimally compatible with the learning process, thus producing suboptimal results. In this paper, we propose a new architecture for nonlinear dictionary learning with sparse coding, in which samples are mapped into sparse codes via carefully designed stacked auto-encoder (SAE) networks. We jointly learn a low-dimensional embedding of the data samples by means of an SAE and a dictionary in the low-dimensional space. Further, to leverage the prior knowledge, we develop a kernel regularized nonlinear dictionary learning method, which effectively incorporates the knowledge provided by the hand-crafted kernel. An iterative algorithm is developed to jointly search the solutions of the associated optimization problem and extensive experimental validations are performed to show that the proposed kernel regularized dictionary learning method achieves satisfactory performance.
Huaping Liu 0001, Fuchun Sun 0001, Bin Fang 0003
IEEE Trans. Syst. Man Cybern. Syst.4
2018 A Dual-Modal Vision-Based Tactile Sensor for Robotic Hand Grasping
abstract
Humans' fingertips can perceive not only the magnitude and the direction of force but also the texture of object. When we grasp an object, the surface texture sensing of the fingertip helps us recognize the object and the force feeling that is parallel to the skin helps us grasp stably. Focusing on these points, we have developed a dual-modal vision-based tactile sensor that can measure the texture of object and a distribution of force vectors. The tactile sensor consists of a transparent elastomer, a camera, a piece of transparent acrylic board, LEDs and supporting structures. A reflective membrane and markers array are on the surface of the elastomer. An applied force on the elastic body results in movements of the markers, which are acquired by the CCD camera. In addition, the shape and texture of the object's contact surface can be reflected by the membrane deformations. The distribution of force vectors is determined by the BP neural network. The local binary pattern algorithm using captured images calculates the texture information. This paper reports experimental evaluation results concerning accuracy of determination of magnitude, direction of force, and texture recognition rate.
Bin Fang 0003, Fuchun Sun 0001, Chao Yang 0026, Hongxiang Xue, Wendan Chen, Chun Zhang 0001, Di Guo 0002, Huaping Liu 0001
ICRA1
2018 3D human gesture capturing and recognition by the IMMU-based data glove
Bin Fang 0003, Fuchun Sun 0001, Huaping Liu 0001, Chunfang Liu
Neurocomputing1
2018 Weakly Paired Multimodal Fusion for Object Recognition
abstract
The ever-growing development of sensor technology has led to the use of multimodal sensors to develop robotics and automation systems. It is therefore highly expected to develop methodologies capable of integrating information from multimodal sensors with the goal of improving the performance of surveillance, diagnosis, prediction, and so on. However, real multimodal data often suffer from significant weak-pairing characteristics, i.e., the full pairing between data samples may not be known, while pairing of a group of samples from one modality to a group of samples in another modality is known. In this paper, we establish a novel projective dictionary learning framework for weakly paired multimodal data fusion. By introducing a latent pairing matrix, we realize the simultaneous dictionary learning and the pairing matrix estimation, and therefore improve the fusion effect. In addition, the kernelized version and the optimization algorithms are also addressed. Extensive experimental validations on some existing data sets are performed to show the advantages of the proposed method.Note to Practitioners—In many industrial environments, we usually use multiple heterogeneous sensors, which provide multimodal information. Such multimodal data usually lead to two technical challenges. First, different sensors may provide different patterns of data. Second, the full-pairing information between modalities may not be known. In this paper, we develop a unified model to tackle such problems. This model is based on a projective dictionary learning method, which efficiently produces the representation vector for the original data by an explicit form. In addition, the latent pairing relation between samples can be learned automatically and be used to improve the classification performance. Such a method can be flexibly used for multimodal fusion with full-pairing, partial-pairing and weak-pairing cases.
Huaping Liu 0001, Yupei Wu, Fuchun Sun 0001, Bin Fang 0003, Di Guo 0002
IEEE Trans Autom. Sci. Eng.4
2017 A hybrid deep architecture for robotic grasp detection
abstract
The robotic grasp detection is a great challenge in the area of robotics. Previous work mainly employs the visual approaches to solve this problem. In this paper, a hybrid deep architecture combining the visual and tactile sensing for robotic grasp detection is proposed. We have demonstrated that the visual sensing and tactile sensing are complementary to each other and important for the robotic grasping. A new THU grasp dataset has also been collected which contains the visual, tactile and grasp configuration information. The experiments conducted on a public grasp dataset and our collected dataset show that the performance of the proposed model is superior to state of the art methods. The results also indicate that the tactile data could help to enable the network to learn better visual features for the robotic grasp detection task.
Di Guo 0002, Fuchun Sun 0001, Huaping Liu 0001, Tao Kong, Bin Fang 0003, Ning Xi 0001
ICRA5
2017 Multi-label tactile property analysis
abstract
In this paper, we exploit the intrinsic relation between different adjective labels and develop a novel multilabel dictionary learning and sparse coding method which is improved by introducing the structured output association information. Such a method makes use of the label correlation information and is more suitable for the multi-label tactile understanding task. In addition, we develop a globally-convergent iterative algorithms to solve the dictionary learning problem. Finally, we perform extensive experimental validations on the public available tactile sequence dataset PHAC-2 and show the advantages of the proposed method.
Huaping Liu 0001, Yupei Wu, Fuchun Sun 0001, Di Guo 0002, Bin Fang 0003
ICRA5
2017 Robotic grasping using visual and tactile sensing
Di Guo 0002, Fuchun Sun 0001, Bin Fang 0003, Chao Yang 0026, Ning Xi 0001
Inf. Sci.3
2017 Structured Output-Associated Dictionary Learning for Haptic Understanding
abstract
Haptic sensing and feedback play extremely important roles for humans and robots to perceive, understand, and manipulate the world. Since many properties perceived by the haptic sensors can be characterized by adjectives, it is reasonable to develop a set of haptic adjectives for the haptic understanding. This formulates the haptic understanding as a multilabel classification problem. In this paper, we exploit the intrinsic relation between different adjective labels and develop a novel dictionary learning method which is improved by introducing the structured output association information. Such a method makes use of the label correlation information and is more suitable for the multilabel haptic understanding task. In addition, we develop two iterative algorithms to solve the dictionary learning and classifier design problems, respectively. Finally, we perform extensive experimental validations on the public available haptic sequence dataset Penn Haptic Adjective Corpus 2 and show the advantages of the proposed method.
Huaping Liu 0001, Fuchun Sun 0001, Di Guo 0002, Bin Fang 0003, Zhengchun Peng
IEEE Trans. Syst. Man Cybern. Syst.4
2016 Tactile sequence based object categorization: A Bag of features modeled by Linear Dynamic System with Symmetric Transition Matrix
abstract
In this paper, we propose a novel categorization framework to recognize tactile sequences based on two particular properties of the tactile data. For the first one, tactile sequences are spatio-temporal data which is sequential and dynamic, depicting the process of grasping an object in different grasping stages; therefore, it is reasonable to discover the dynamical pattern by modeling tactile data as integral sequences rather than individual frames. For the second one, a tactile sequence contains various dynamical patterns in different stages of the grasping process; therefore, we decompose the whole sequence into multiple mini-sequences so as to enhance feature resolution. To address both properties in our framework, we take advantage of a Bag-of-System model using parameters of the Linear Dynamic System (LDS) as feature descriptors. Moreover, we employ the LDS with Symmetric Transition matrix (LDSST) rather than the original LDS as the building-block in order to obtain accurate codewords of the codebook of the Bag-of-System. The performance of our framework is evaluated on six real-world databases of three groups. Our experiments show that classification using LDSST is better than the original LDS, and the decomposition of tactile sequences does improve the accuracy of classification. The experiment results also show the superiority of our framework in comparison with other state-of-the-art sequence classifiers.
Fuchun Sun 0001, Wenbing Huang 0001, Le-le Cao, Bin Fang 0003
IJCNN5