VLDB 2026 Research / reviewers in the wild / expert
Di Guo 0002
dblp:29/7028-2
· DBLP profile ↗
44ranked-venue papers
6as first author
25since 2021 · last 2026
0000-0002-9816-0103ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 3 first-author · 16 since 2021Systems, architecture and hardware · 20 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RoboCleaner: Robotic Tabletop Cleaning via VLM-Powered Multi-Agent CollaborationabstractRobotic tabletop cleaning is applicable in various environments ranging from domestic to industrial settings, yet it still faces challenges in the cluttered scenario where wastes of diverse types exist. Moreover, to deal with different waste, such as liquid spills or fine crumbs, different cleaning tools might be required, which further complicates the task. Inspired by significant advancements in vision language models (VLMs), in this paper, we proposeRoboCleaner, a VLM-powered multi-agent framework for tabletop cleaning that leverages the collaborative intelligence of multiple agents. Specifically, the proposed framework consists of three VLM-powered agents: aPlanning Agentfor decision-making, anExecution Agentfor precise operation and aReflection Agentfor outcome evaluation and providing feedback for iterative improvement. Through the collaboration of these agents, our framework is capable of handling a wide range of challenging cleaning tasks. Extensive experiments are conducted demonstrating that the proposedRoboCleanercan achieve high task success rates and operational efficiency within cluttered environments. Also, we have observed the emergent problem-solving capabilities of the proposed framework, which further validates the robustness and adaptability of the framework. Di Guo 0002 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | AssistantX: An LLM-Powered Proactive Assistant in Collaborative Human-Populated EnvironmentsabstractCurrent service robots suffer from limited natural language communication abilities, heavy reliance on predefined commands, ongoing human intervention, and, most notably, a lack of proactive collaboration awareness in human-populated environments. This results in narrow applicability and low utility. In this paper, we introduce AssistantX, an LLM-powered proactive assistant designed for autonomous operation in real-world scenarios with high accuracy. AssistantX employs a multi-agent framework consisting of 4 specialized LLM agents, each dedicated to perception, planning, decision-making, and reflective review, facilitating advanced inference capabilities and comprehensive collaboration awareness, much like a human assistant by your side. We built a dataset of 210 real-world tasks to validate AssistantX, which includes instruction content and status information on whether relevant personnel are available. Extensive experiments were conducted in both text-based simulations and a real office environment over the course of a month and a half. Our experiments demonstrate the effectiveness of the proposed framework, showing that AssistantX can reactively respond to user instructions, actively adjust strategies to adapt to contingencies, and proactively seek assistance from humans to ensure successful task completion. More details and videos can be found at https://assistantx-agent.github.io/AssistantX/. Yongchang Li, Di Guo 0002, Huaping Liu 0001 |
IROS | 4 |
| 2025 | A Patch-Based Transformer Method for Electrical Capacitance Tomography Image ReconstructionabstractElectrical capacitance tomography (ECT) is a contactless and non-invasive imaging technique, which visualizes the internal permittivity distribution around a region utilizing boundary capacitance measurements. It has been widely used in the fields of object classification, tactile sensing and multiphase flows monitoring. However, due to the inherent nonlinearity and ill-conditioned nature of the ECT inverse problem, its practical implementation remains limited by challenges in the image reconstruction accuracy. To tackle the above problems, we propose a patch-based transformer method (PT) for an accurate reconstruction of ECT images. Specifically, the complex capacitance-to-image mapping is systematically decoupled into the capacitance-to-patch feature extraction and patch-to-image reconstruction, enabling more efficient and accurate permittivity distribution recovery through localized feature learning and global context integration. Additionally, a simulation ECT dataset for objects with varying sizes and positions is established. Duanpeng Shi, Huaping Liu 0001, Di Guo 0002 |
IROS | 4 |
| 2025 | A Novel Terrain Classification System with Planar ECT SensorabstractTerrain classification is crucial for robotic navigation especially in unknown environment. Existing terrain classification methods usually have high requirements for environment conditions and robot motions, making them challenging to apply to real-world scenarios. In this paper, we develop a novel terrain classification system with the planar electrical capacitance tomography (ECT) sensor, which provides a non-contact, real-time, and cost-effective way for terrain classification. Specifically, we design a planar ECT sensor and integrate it at the bottom of a mobile robot. The proposed system leverages the collected capacitance measurements to reflect the inherent differences in dielectric permittivity across various terrain types. And a multilayer perception networks is used to fuse the collected capacitance and IMU measurements for classification. Additionally, a large scale ECT dataset including 10 different types of terrains is collected with the proposed system. Extensive experiments are conducted demonstrating the effectiveness and robustness of the proposed system. Wenju Yang, Duanpeng Shi, Wuqiang Yang, Tengchen Sun, Huaping Liu 0001, Di Guo 0002 |
IROS | 6 |
| 2025 | Trust-Aware Human-Robot Fusion Decision-Making for Emergency Indoor PatrollingabstractTrust plays a crucial role in decision-making during human-robot collaboration, particularly in emergency scenarios where it becomes more susceptible due to dynamic factors. Misalignment of human-robot trust significantly hampers the efficiency of the collaboration. Therefore, it is imperative to establish effective human-robot interaction and decision support mechanisms that mitigate biases in human confidence levels regarding the robot’s capabilities in dynamic environments. Additionally, online correction of human-robot trust based on human behavioral feedback is vital. This paper focuses on a specific type of emergency task scenario, specifically, indoor human-robot collaborative patrolling in the event of sudden power outages. We propose a trust model based on linear Gaussian and sparse Gaussian processes (sparse GP). We also employ Monte Carlo Tree Search (MCTS) method to determine the robot’s optimal fusion decision-making. Through VR-based human-robot collaborative experiments, we ascertain that the robot prioritizes enhancing human-robot trust in emergency scenarios to mitigate the long-term costs of human-robot collaboration.Note to Practitioners—The primary motivation of this paper lies in tackling the issue of decreased collaboration efficiency in human-robot cooperation, stemming from irrational human decision-making, a problem exacerbated in emergency scenarios where comprehending all available information proves challenging. Grounded in the realm of human-robot trust, this study evaluates the temporal and substantive aspects of human-robot interactions contingent upon the degree of trust. Moreover, it adapts the fusion decision-making process within the human-robot team in accordance with the decision-making strategies adopted by humans at varying trust levels. The optimized decisions of the human-robot ensemble are conveyed to the human participant via decision support. This approach ensures the maintenance of trust levels, facilitating the acceptance of decision support that may seem intuitively irrational but holds objective superiority. Furthermore, it mitigates redundant human-robot interactions once a sufficient level of trust is established. Yang Li 0029, Jiaxin Xu, Di Guo 0002, Huaping Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Soft Contact Simulation and Manipulation Learning of Deformable Objects With Vision-Based Tactile SensorabstractDeformable object manipulation is a challenging problem due to its complex deformable properties. With the development of artificial intelligence, learning-based methods have shown outstanding performance in robotic manipulation. Previous works have investigated the manipulation of deformable objects via Reinforcement Learning (RL) in simulation. However, they approximate object deformation with particles, using particle states as observations, which are unavailable in reality. To address these issues, we utilize Vision-Based Tactile Sensors (VBTSs) as the end-effector to manipulate and observe the deformable objects. In this work, we develop a new contact simulation environment for deformable objects, including elastic, plastic, and elastoplastic. We utilize RL strategies and expert demonstrations to train agents in the simulation. Finally, we build a real experimental platform to complete the sim-to-real tasks and robustness testing. Our work introduces an innovative strategy that utilizes high-resolution VBTSs for contact simulation and manipulation of deformable objects. The experimental results show superior performances of deformable object manipulation with the proposed method. Shixin Zhang, Zixi Chen 0002, Zirong Shen, Fuchun Sun 0001, Cesare Stefanini, Di Guo 0002, Shan Luo 0001, Jianwei Zhang 0001, Jianhua Shan, Bin Fang 0003 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2025 | Semantics-Aware Hierarchical Decision Framework for Embodied Visual Room RearrangementabstractIn embodied visual room rearrangement, the agent needs to recover the scene state to the goal state through interacting with the environment based on the egocentric visual observations after the locations and states of some objects are changed. It has important application potential in the field of robotics. This task is challenging in visual perception, scene understanding, and action execution. Existing methods do not take full advantage of the semantic information and spatial relationship of objects in the scene perception and understanding process. To tackle the challenges and shortcomings of the current methods, we build a hierarchical decision framework based on the pretrained semantic scene representation and transformer-based scene memory to solve this task. The results in the unseen scenes demonstrate the effectiveness of the proposed model compared with other methods. Xinzhu Liu, Di Guo 0002, Huaping Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Self-Supervised 3-D Semantic Representation Learning for Vision-and-Language NavigationabstractIn vision-and-language navigation (VLN) tasks, most current methods primarily utilize RGB images, overlooking the rich 3-D semantic data inherent to environments. To rectify this, we introduce a novel VLN framework that integrates 3-D semantic information into the navigation process. Our approach features a self-supervised training scheme that incorporates voxel-level 3-D semantic reconstruction to create a detailed 3-D semantic representation. A key component of this framework is a pretext task focused on region queries, which determines the presence of objects in specific 3-D areas. Following this, we devise an long short-term memory (LSTM)-based navigation model that is trained using our 3-D semantic representations. To maximize the utility of these 3-D semantic representations, we implement a cross-modal distillation strategy. This strategy encourages the RGB model's outputs to emulate those from the 3-D semantic feature network, enabling the concurrent training of both branches to merge RGB and 3-D semantic data effectively. Comprehensive evaluations on both the R2R and R4R datasets reveal that our method significantly enhances performance in VLN tasks. Sinan Tan, Kuankuan Sima, Dunzheng Wang, Mengmeng Ge 0002, Di Guo 0002, Huaping Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Stimulate the Potential of Robots via CompetitionabstractIt is common for us to feel pressure in a competition environment, which arises from the desire to obtain success comparing with other individuals or opponents. Although we might get anxious under the pressure, it could also be a drive for us to stimulate our potentials to the best in order to keep up with others. Inspired by this, we propose a competitive learning framework which is able to help individual robot to acquire knowledge from the competition, fully stimulating its dynamics potential in the race. Specifically, the competition information among competitors is introduced as the additional auxiliary signal to learn advantaged actions. We further build a Multiagent-Race environment, and extensive experiments are conducted, demonstrating that robots trained in competitive environments outperform ones that are trained with SoTA algorithms in single robot environment. Kangyao Huang, Di Guo 0002, Xinyu Zhang 0001, Xiangyang Ji, Huaping Liu 0001 |
ICRA | 2 |
| 2024 | A Large-area Tactile Sensor for Distributed Force Sensing Using Highly Sensitive Piezoresistive SpongeabstractTactile sensing plays a critical role in enabling robots to interact safely with target objects in dynamic and unstructured environments. While various tactile sensors based on different sensing principles or different sensitive materials have been proposed, the development of flexible large-area tactile sensors for robots is still challenging. In this paper, a novel highly sensitive piezoresistive sponge based on multi-walled carbon nanotubes (MWCNTs) and polyurethane (PU) sponge is fabricated for pressure sensing. The sensing behavior of the piezoresistive sponge was experimentally evaluated, showing high sensitivity and fast response. Based on the piezoresistive sponge, a flexible large-area tactile sensor is designed for distributed force detection with electrical resistance tomography technology. The sensing performance of the sensor is validated by touch location, sensitivity analysis, real-time touch discrimination, and touch modality recognition. The experimental results indicate that the sensor performs well in detecting the position and force of contact in a large area. The sensor’s performance shows promise in embodied tactile sensing and human–robot interaction. Wendong Zheng, Di Guo 0002, Wuqiang Yang, Huaping Liu 0001 |
ICRA | 3 |
| 2024 | CompetEvo: Towards Morphological Evolution from Competition
Kangyao Huang, Di Guo 0002, Xinyu Zhang 0001, Xiangyang Ji, Huaping Liu 0001 |
IJCAI | 2 |
| 2024 | Continual Residual Reservoir Computing for Remaining Useful Life PredictionabstractIn practical engineering applications, it is inevitable that production is often faced with different working conditions. Therefore, it is necessary to have a continual learning system, which can adapt to a sequence of tasks and keep learning over time. In this work, we propose a continual deep residual reservoir computing framework for the practical remaining useful life prediction task. Specifically, we propose a novel deep echo state network structure with residual blocks to effectively mitigate the performance degradation of the deep reservoir computing framework and reduce the difficulty of model training. Furthermore, the proposed framework is trained with the elastic weight consolidation method to alleviate the impact of catastrophic forgetting in continual learning systems. Extensive experiments are conducted with the FEMTO-ST bearing and high-intensity-radiated fields battery dataset. And the proposed framework is proven to be effective in multiple continual learning tasks compared with other state-of-the-art methods. Huaping Liu 0001, Di Guo 0002, Xi-Ming Sun |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Rich Action-Semantic Consistent Knowledge for Early Action PredictionabstractEarly action prediction (EAP) aims to recognize human actions from a part of action execution in ongoing videos, which is an important task for many practical applications. Most prior works treat partial or full videos as a whole, ignoring rich action knowledge hidden in videos, i.e., semantic consistencies among different partial videos. In contrast, we partition original partial or full videos to form a new series of partial videos and mine the Action-Semantic Consistent Knowledge (ASCK) among these new partial videos evolving in arbitrary progress levels. Moreover, a novel Rich Action-semantic Consistent Knowledge network (RACK) under the teacher-student framework is proposed for EAP. Firstly, we use a two-stream pre-trained model to extract features of videos. Secondly, we treat the RGB or flow features of the partial videos as nodes and their action semantic consistencies as edges. Next, we build a bi-directional semantic graph for the teacher network and a single-directional semantic graph for the student network to model rich ASCK among partial videos. The MSE and MMD losses are incorporated as our distillation loss to enrich the ASCK of partial videos from the teacher to the student network. Finally, we obtain the final prediction by summering the logits of different subnetworks and applying a softmax layer. Extensive experiments and ablative studies have been conducted, demonstrating the effectiveness of modeling rich ASCK for EAP. With the proposed RACK, we have achieved state-of-the-art performance on three benchmarks. The code is available at https://github.com/lily2lab/RACK.git. Jianqin Yin, Di Guo 0002, Huaping Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Natural Language Instruction Understanding for Robotic Manipulation: a Multisensory Perception ApproachabstractIt has always been expected that the robot can understand the natural language instruction and thus a more natural human-robot interaction is achieved. Currently, the robot usually interprets the instruction by visually grounding the textual information to its surroundings, while it may be not enough for some complex situations with only visual perception. So it is reasonable for the robot to leverage its multisensory perception ability to better understand the instruction. In this paper, we propose a multisensory perception approach to tackle the task of natural language instruction understanding for robotic manipulation, in which the robot coordinates its visual, tactile and auditory perception to fully understand the instruction and then executes the manipulation task. Extensive experiments have been conducted demonstrating the superiority of the multisensory perception compared with single sensory perception for instruction understanding. Moreover, we establish a user-friendly human-robot interaction interface where the human sends instruction to the robot via a mobile APP. Yanzhi Dong, Di Guo 0002, Huaping Liu 0001 |
ICRA | 5 |
| 2023 | Adaptive Optimal Electrical Resistance Tomography for Large-Area Tactile SensingabstractIt is critical to perceive physical contact for intelligent robots to safely interact in dynamic, unstructured environments. As physical contacts can occur at any location, a well-performing tactile sensing system should be able to deploy a large area on robotic surface. Some researchers have implemented large-area tactile sensors by using sensing arrays, but it is challenging to deploy many sensing elements. Electrical resistance tomography (ERT) has recently been introduced into tactile sensing to overcome some of the limitations with conventional tactile sensing arrays, and good results have been achieved for some robotic applications. However, a particular challenge is that spatial resolution is low. Although various attempts have been made to improve the performance of ERT-based tactile sensors, the intrinsic resolution issue remains unsolved. In this paper, we propose a novel adaptive optimal drive strategy for efficient ERT-based large-area tactile sensing for robotic applications, which can adaptively select the current injection and voltage measurement pattern for optimal tactile stimulus. In particular, regions of tactile contacts are preliminarily detected and localized by a base scanning pattern with only a few measurement data. According to this detected region, the adaptive strategy can select the optimal current injection and voltage measurement pattern to improve the sensing performance by maximizing the current density. To verify the effectiveness of the proposed strategy, the proposed method is comprehensively evaluated by simulation and experiments. The results revealed that the optimal strategy can effectively improve both spatial and temporal resolution. Wendong Zheng, Huaping Liu 0001, Di Guo 0002, Wuqiang Yang |
ICRA | 3 |
| 2023 | Knowledge-Based Embodied Question AnsweringabstractIn this paper, we propose a novel Knowledge-based Embodied Question Answering (K-EQA) task, in which the agent intelligently explores the environment to answer various questions with the knowledge. Different from explicitly specifying the target object in the question as existing EQA work, the agent can resort to external knowledge to understand more complicated question such as "Please tell me what are objects used to cut food in the room?", in which the agent must know the knowledge such as "knife is used for cutting food". To address this K-EQA problem, a novel framework based on neural program synthesis reasoning is proposed, where the joint reasoning of the external knowledge and 3D scene graph is performed to realize navigation and question answering. Especially, the 3D scene graph can provide the memory to store the visual information of visited scenes, which significantly improves the efficiency for the multi-turn question answering. Experimental results have demonstrated that the proposed framework is capable of answering more complicated and realistic questions in the embodied environment. The proposed method is also applicable to multi-agent scenarios. Sinan Tan, Mengmeng Ge 0002, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Depth-Aware Vision-and-Language Navigation using Scene Query Attention NetworkabstractVision-and-language navigation (VLN) has been an important task in the field of Robotics and Computer Vision. However, most existing vision-and-language navigation models only use features extracted from RGB observation as input, while robots can utilize depth sensors in the real world. Existing research has also shown that simply adding a depth stream to neural models could only provide a marginal improvement to the performance of the VLN task. Therefore, in our work, we develop a novel method for the VLN task using semantic map observations built from RGB-D input. We use vision-pretraining to efficiently encode the semantic map with CNN and scene query attention network by answering queries about semantic information of specific regions of a scene. The proposed method could be used with a simple model and does not require large-scale vision-language transformer pretraining, bringing a more than 10% increase in the success rate compared with a baseline model. When used together with the Speaker-Follower training technique, it achieves a success rate of 58 % on the test set for the R2R dataset in single-run setting, outperforming the previous RGB-D method and most existing RGB-only models that do not use large-scale vision-language transformers pretraining. Sinan Tan, Mengmeng Ge 0002, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001 |
ICRA | 3 |
| 2022 | Audio-Visual Grounding Referring Expression for Robotic ManipulationabstractReferring expressions are commonly used when referring to a specific target in people's daily dialogue. In this paper, we develop a novel task of audio-visual grounding referring expression for robotic manipulation. The robot leverages both the audio and visual information to understand the referring expression in the given manipulation instruction and the corresponding manipulations are implemented. To solve the proposed task, an audio-visual framework is proposed for visual localization and sound recognition. We have also established a dataset which contains visual data, auditory data and manipulation instructions for evaluation. Finally, extensive experiments are conducted both offline and online to verify the effectiveness of the proposed audio-visual framework. And it is demonstrated that the robot performs better with the audio-visual data than with only the visual data. Yefei Wang, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001 |
ICRA | 4 |
| 2022 | Remaining useful life predictions for turbofan engine degradation based on concurrent semi-supervised model
Di Guo 0002, Xi-Ming Sun |
Neural Comput. Appl. | 2 |
| 2022 | Visual Affordance Guided Tactile Material Recognition for Waste RecyclingabstractBecause more and more solid waste is generated, in particular, in cities, the management of solid waste disposal has become a global challenge. A solution is to find an effective way to sort solid waste materials and recycle them into reusable products. In this article, we propose to use a material recognition method with vision-guided tactile to form a robotic system for waste sorting. The vision guidance module integrates an object detector and an affordance network together. It allows the robot to not only detect the desired containers and packaging from an assortment of the waste but also obtain a configuration to grasp the target and actively collect its tactile data. By classifying the object with the tactile data, the robot can sort containers and packaging into their respective categories according to the type of material. Our experimental results demonstrate the effectiveness of the proposed robotic waste sorting system in sorting containers and packaging various types of materials. Note to Practitioners—The management of waste has become a great challenge in environmental protection in the world. We propose a robotic waste sorting system, which utilizes a vision-guided tactile sensing approach to find target waste and sort the waste according to materials. A visual module is used to find the waste of interest. A class-specific affordance map is generated to guide a robotic hand to actively grasp waste and collect the tactile data for material recognition. With the recorded tactile data, the robot can recognize the material of the target waste and sort the waste according to the type of material. The proposed system demonstrates good performance in waste sorting, and it can be easily implemented in practical scenarios. The target material can be generalized to a broader group of objects. Di Guo 0002, Huaping Liu 0001, Bin Fang 0003, Fuchun Sun 0001, Wuqiang Yang |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2021 | Adversarial Skill Learning for Robust ManipulationabstractDeep reinforcement learning has made significant progress in robotic manipulation tasks and it works well in the ideal disturbance-free environment. However, in a real-world environment, both internal and external disturbances are inevitable, thus the performance of the trained policy will dramatically drop. To improve the robustness of the policy, we introduce the adversarial training mechanism to the robotic manipulation tasks in this paper, and an adversarial skill learning algorithm based on soft actor-critic (SAC) is proposed for robust manipulation. Extensive experiments are conducted to demonstrate that the learned policy is robust to internal and external disturbances. Additionally, the proposed algorithm is evaluated in both the simulation environment and on the real robotic platform. Pingcheng Jian, Chao Yang 0026, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001 |
ICRA | 3 |
| 2021 | Robotic Indoor Scene Captioning from Streaming VideoabstractRobots are usually equipped with cameras to explore the indoor scene and it is expected that the robot can well describe the scene with natural language. Although some great success has been achieved in image and video captioning technology, especially on many public datasets, the caption generated from indoor scene video is still not informative and coherent enough. In this paper, we propose the problem of Indoor Scene Captioning from Streaming Video, which aims at generating a more accurate and informative caption from streaming video. To solve this problem, we firstly design an algorithm to organize the visual information of the indoor scene into a scene graph, and then implement a scene graph guided captioning method, which takes the scene graph and video frames as input to generate the caption from the video streaming. The proposed framework is evaluated both on the AI2THOR dataset and a real-world robotic platform, demonstrating the effectiveness of the framework. Xinghang Li, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001 |
ICRA | 2 |
| 2021 | Toward Image-to-Tactile Cross-Modal Perception for Visually Impaired PeopleabstractIt is still a great challenge for the visually impaired people to perceive their surroundings from a global perspective, which makes it difficult for them to interact with unfamiliar environments. The reason is that these conventional assisting devices only address the obstacle avoidance problem. They do not provide visually impaired people with a global perception of the surrounding environment. In this article, a new generative adversarial network (GAN) model is developed to effectively transform the ground images into the tactile signal, which can be displayed by an off-the-shelf vibration device. The algorithm module and the hardware are integrated into a portable device, which provides visually impaired people with effective surrounding perception capability. In addition, a visual-tactile cross-modal data set is constructed to train the proposed deep-learning architecture. Experimental results show that the proposed system can help visually impaired people sense the ground and bring a better traveling experience for them. Note to Practitioners-This article presents a portable device that provides tactile recognition assistance for visually impaired people. Such a technology can be extensively used in tactile mouse and white cane. The developed technology can be extensively used for various industrial applications, such as surrounding monitoring and manipulation. The proposed work demonstrates the promising ability of artificial intelligence in healthcare applications. The generated tactile signals are expected to be used in many human-centered systems, and we believe that our contribution is an important step toward the development of a more comprehensive assisting technology for visually impaired people. Huaping Liu 0001, Di Guo 0002, Xinyu Zhang 0001, Wenlin Zhu, Bin Fang 0003, Fuchun Sun 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2021 | An Interactive Perception Method for Warehouse Automation in Smart CitiesabstractThe smart city is an integrated environment that heavily relies on intelligent robots, which provides the basis for the warehouse automation. However, a warehouse is a typical unstructured environment, and robotic grasp and manipulation are extremely important for the package, transfer, search, and so on. Currently, the most usual method is to detect the picking or grasping points for some specific end-effector including suction cup, gripper, or robotic hand. The manipulation performance is, therefore, strongly influenced by the visual detector. To tackle this problem, the affordance map has recently been developed. It characterizes the operation possibilities afforded by the operation scene and has been used for several grasp tasks. Nevertheless, the conventional affordance method often fails in complicated environments due to the mistake calculation results. In this article, we develop a novel framework to integrate the interactive exploration with a composite robotic hand for robotic grasping in a complicated environment. The exploration strategy is obtained by a deep reinforcement learning procedure. The developed new composite hand, which integrates the suction cup and grippers, is used to test the merits of the proposed interactive perception method. Experimental results show the proposed method significantly increases the manipulation efficiency and may bring great economic and social and benefits for smart cities. Huaping Liu 0001, Yuhong Deng, Di Guo 0002, Bin Fang 0003, Fuchun Sun 0001, Wuqiang Yang |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | Active Object Discovery and Localization Using Sound-Induced AttentionabstractIndustrial intelligent devices are usually equipped with both microphones and cameras to perceive and understand the physical world. Though visual object detection technology has achieved a great success, its combination with other sensing modalities remains unsolved. In this article, we establish a novel sound-induced attention framework for the visual object detection, and develop a two-stream weakly supervised deep learning architecture to combine the visual and audio modalities for localizing the sounding object. A dataset is constructed from the Audio Set to validate the proposed method and some realistic experiments are conducted to demonstrate the effectiveness of the proposed system. Huaping Liu 0001, Feng Wang 0034, Di Guo 0002, Xinzhu Liu, Xinyu Zhang 0001, Fuchun Sun 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Multi-agent Embodied Question Answering in Interactive Environments
Sinan Tan, Weilai Xiang, Huaping Liu 0001, Di Guo 0002, Fuchun Sun 0001 |
ECCV (13) | 4 |
| 2020 | Self-Supervised Learning for Alignment of Objects and SoundabstractThe sound source separation problem has many useful applications in the field of robotics, such as human-robot interaction, scene understanding, etc. However, it remains a very challenging problem. In this paper, we utilize both visual and audio information of videos to perform the sound source separation task. A self-supervised learning framework is proposed to implement the object detection and sound separation modules simultaneously. Such an approach is designed to better find the alignment between the detected objects and separated sound components. Our experiments, conducted on both the synthetic and real datasets, validate this approach and demonstrate the effectiveness of the proposed model in the task of object and sound alignment. Xinzhu Liu, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001, Haibo Min |
ICRA | 3 |
| 2020 | Unsupervised Representation Learning by Invariance PropagationabstractUnsupervised learning methods based on contrastive learning have drawn increasing attention and achieved promising results. Most of them aim to learn representations invariant to instance-level variations, which are provided by different views of the same instance. In this paper, we propose Invariance Propagation to focus on learning representations invariant to category-level variations, which are provided by different instances from the same category. Our method recursively discovers semantically consistent samples residing in the same high-density regions in representation space. We demonstrate a hard sampling strategy to concentrate on maximizing the agreement between the anchor sample and its hard positive samples, which provide more intra-class variations to help capture more abstract invariance. As a result, with a ResNet-50 as the backbone, our method achieves 71.3% top-1 accuracy on ImageNet linear classification and 78.2% top-5 accuracy fine-tuning on only 1% labels, surpassing previous results. We also achieve state-of-the-art performance on other downstream tasks, including linear classification on Places205 and Pascal VOC, and transfer learning on small scale datasets. Feng Wang 0034, Huaping Liu 0001, Di Guo 0002, Fuchun Sun 0001 |
NeurIPS | 3 |
| 2020 | Cross-Modal Zero-Shot-Learning for Tactile Object RecognitionabstractIn this paper, we address the learning problem of classifying untouched tactile instance with the help of visual modality. The proposed method is based on dictionary learning and we impose different penalty terms on coding vectors between visual and tactile modalities. Using such structured coding vectors, the visual-tactile cross-modal transfer can be achieved. A set of optimization algorithms are developed to obtain the solutions of the proposed optimization problems. After then, we can use the obtained dictionary to predict the coding vectors of the new untouched tactile samples and further determine its label. Finally, we perform extensive experimental evaluations on publicly available datasets to show the effectiveness of the proposed method. Huaping Liu 0001, Fuchun Sun 0001, Bin Fang 0003, Di Guo 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2019 | Sound-Indicated Visual Object Detection for Robotic ExplorationabstractRobots are usually equipped with microphones and cameras to perceive and understand the physical world. Though visual object detection technology has achieved great success, the detection in other modalities remains unsolved. In this paper, we establish a novel robotic sound-indicated visual object detection framework, and develop a two-stream weakly-supervised deep learning architecture to connect the visual and audio modalities for localizing the sounding object. A dataset is constructed from the AudioSet to validate the proposed method and some promising applications are demonstrated on robotic platforms. Feng Wang 0034, Di Guo 0002, Huaping Liu 0001, Junfeng Zhou, Fuchun Sun 0001 |
ICRA | 2 |
| 2019 | Deep Reinforcement Learning for Robotic Pushing and Picking in Cluttered EnvironmentabstractIn this paper, a novel robotic grasping system is established to automatically pick up objects in cluttered scenes. A composite robotic hand composed of a suction cup and a gripper is designed for grasping the object stably. The suction cup is used for lifting the object from the clutter first and the gripper for grasping the object accordingly. We utilize the affordance map to provide pixel-wise lifting point candidates for the suction cup. To obtain a good affordance map, the active exploration mechanism is introduced to the system. An effective metric is designed to calculate the reward for the current affordance map, and a deep Q-Network (DQN) is employed to guide the robotic hand to actively explore the environment until the generated affordance map is suitable for grasping. Experimental results have demonstrated that the proposed robotic grasping system is able to greatly increase the success rate of the robotic grasping in cluttered scenes. Yuhong Deng, Yixuan Wei, Kai Lu 0003, Bin Fang 0003, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001 |
IROS | 6 |
| 2019 | A glove-based system for object recognition via visual-tactile fusion
Bin Fang 0003, Fuchun Sun 0001, Huaping Liu 0001, Chuanqi Tan, Di Guo 0002 |
Sci. China Inf. Sci. | 5 |
| 2018 | A Dual-Modal Vision-Based Tactile Sensor for Robotic Hand GraspingabstractHumans' fingertips can perceive not only the magnitude and the direction of force but also the texture of object. When we grasp an object, the surface texture sensing of the fingertip helps us recognize the object and the force feeling that is parallel to the skin helps us grasp stably. Focusing on these points, we have developed a dual-modal vision-based tactile sensor that can measure the texture of object and a distribution of force vectors. The tactile sensor consists of a transparent elastomer, a camera, a piece of transparent acrylic board, LEDs and supporting structures. A reflective membrane and markers array are on the surface of the elastomer. An applied force on the elastic body results in movements of the markers, which are acquired by the CCD camera. In addition, the shape and texture of the object's contact surface can be reflected by the membrane deformations. The distribution of force vectors is determined by the BP neural network. The local binary pattern algorithm using captured images calculates the texture information. This paper reports experimental evaluation results concerning accuracy of determination of magnitude, direction of force, and texture recognition rate. Bin Fang 0003, Fuchun Sun 0001, Chao Yang 0026, Hongxiang Xue, Wendan Chen, Chun Zhang 0001, Di Guo 0002, Huaping Liu 0001 |
ICRA | 7 |
| 2018 | Weakly Paired Multimodal Fusion for Object RecognitionabstractThe ever-growing development of sensor technology has led to the use of multimodal sensors to develop robotics and automation systems. It is therefore highly expected to develop methodologies capable of integrating information from multimodal sensors with the goal of improving the performance of surveillance, diagnosis, prediction, and so on. However, real multimodal data often suffer from significant weak-pairing characteristics, i.e., the full pairing between data samples may not be known, while pairing of a group of samples from one modality to a group of samples in another modality is known. In this paper, we establish a novel projective dictionary learning framework for weakly paired multimodal data fusion. By introducing a latent pairing matrix, we realize the simultaneous dictionary learning and the pairing matrix estimation, and therefore improve the fusion effect. In addition, the kernelized version and the optimization algorithms are also addressed. Extensive experimental validations on some existing data sets are performed to show the advantages of the proposed method.Note to Practitioners—In many industrial environments, we usually use multiple heterogeneous sensors, which provide multimodal information. Such multimodal data usually lead to two technical challenges. First, different sensors may provide different patterns of data. Second, the full-pairing information between modalities may not be known. In this paper, we develop a unified model to tackle such problems. This model is based on a projective dictionary learning method, which efficiently produces the representation vector for the original data by an explicit form. In addition, the latent pairing relation between samples can be learned automatically and be used to improve the classification performance. Such a method can be flexibly used for multimodal fusion with full-pairing, partial-pairing and weak-pairing cases. Huaping Liu 0001, Yupei Wu, Fuchun Sun 0001, Bin Fang 0003, Di Guo 0002 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2017 | From foot to head: Active face finding using deep Q-learningabstractIn the existing work on active face detection and tracking, it is usually required that the face has to appear in the field-of-view. However, this may not be practical in some challenging scenarios. In this paper, we formulate the problem of active face finding as a Markov Decision Process and resort to the deep Q-learning to solve it in an end-to-end manner. Under the proposed framework, the agent is able to learn how to adjust the control parameters of a camera in order to find the face. Even if the captured image contains only some parts of the person, the PTZ camera can still adjust its pose until the face is found. Extensive experimental validations are performed to show the effectiveness of the developed system. Hui Zhang 0092, Huaping Liu 0001, Di Guo 0002, Fuchun Sun 0001 |
ICIP | 3 |
| 2017 | A hybrid deep architecture for robotic grasp detectionabstractThe robotic grasp detection is a great challenge in the area of robotics. Previous work mainly employs the visual approaches to solve this problem. In this paper, a hybrid deep architecture combining the visual and tactile sensing for robotic grasp detection is proposed. We have demonstrated that the visual sensing and tactile sensing are complementary to each other and important for the robotic grasping. A new THU grasp dataset has also been collected which contains the visual, tactile and grasp configuration information. The experiments conducted on a public grasp dataset and our collected dataset show that the performance of the proposed model is superior to state of the art methods. The results also indicate that the tactile data could help to enable the network to learn better visual features for the robotic grasp detection task. Di Guo 0002, Fuchun Sun 0001, Huaping Liu 0001, Tao Kong, Bin Fang 0003, Ning Xi 0001 |
ICRA | 1 |
| 2017 | Multi-label tactile property analysisabstractIn this paper, we exploit the intrinsic relation between different adjective labels and develop a novel multilabel dictionary learning and sparse coding method which is improved by introducing the structured output association information. Such a method makes use of the label correlation information and is more suitable for the multi-label tactile understanding task. In addition, we develop a globally-convergent iterative algorithms to solve the dictionary learning problem. Finally, we perform extensive experimental validations on the public available tactile sequence dataset PHAC-2 and show the advantages of the proposed method. Huaping Liu 0001, Yupei Wu, Fuchun Sun 0001, Di Guo 0002, Bin Fang 0003 |
ICRA | 4 |
| 2017 | Robotic grasping using visual and tactile sensing
Di Guo 0002, Fuchun Sun 0001, Bin Fang 0003, Chao Yang 0026, Ning Xi 0001 |
Inf. Sci. | 1 |
| 2017 | Extreme Kernel Sparse Learning for Tactile Object RecognitionabstractTactile sensors play very important role for robot perception in the dynamic or unknown environment. However, the tactile object recognition exhibits great challenges in practical scenarios. In this paper, we address this problem by developing an extreme kernel sparse learning methodology. This method combines the advantages of extreme learning machine and kernel sparse learning by simultaneously addressing the dictionary learning and the classifier design problems. Furthermore, to tackle the intrinsic difficulties which are introduced by the representer theorem, we develop a reduced kernel dictionary learning method by introducing row-sparsity constraint. A globally convergent algorithm is developed to solve the optimization problem and the theoretical proof is provided. Finally, we perform extensive experimental validations on some public available tactile sequence datasets and show the advantages of the proposed method. Huaping Liu 0001, Fuchun Sun 0001, Di Guo 0002 |
IEEE Trans. Cybern. | 4 |
| 2017 | Structured Output-Associated Dictionary Learning for Haptic UnderstandingabstractHaptic sensing and feedback play extremely important roles for humans and robots to perceive, understand, and manipulate the world. Since many properties perceived by the haptic sensors can be characterized by adjectives, it is reasonable to develop a set of haptic adjectives for the haptic understanding. This formulates the haptic understanding as a multilabel classification problem. In this paper, we exploit the intrinsic relation between different adjective labels and develop a novel dictionary learning method which is improved by introducing the structured output association information. Such a method makes use of the label correlation information and is more suitable for the multilabel haptic understanding task. In addition, we develop two iterative algorithms to solve the dictionary learning and classifier design problems, respectively. Finally, we perform extensive experimental validations on the public available haptic sequence dataset Penn Haptic Adjective Corpus 2 and show the advantages of the proposed method. Huaping Liu 0001, Fuchun Sun 0001, Di Guo 0002, Bin Fang 0003, Zhengchun Peng |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2016 | Object discovery and grasp detection with a shared convolutional neural networkabstractGrasp an object from a stack of objects in real-time is still a challenge in robotics. This requires the robot to have the ability of both fast object discovery and grasp detection: a target object should be picked out from the stack first and then a proper grasp configuration is applied to grasp the object. In this paper, we propose a shared convolutional neural network (CNN) which can simultaneously implement these two tasks in real-time. The processing speed of the model is about 100 frames per second on a GPU which largely satisfies the requirement. Meanwhile, we also establish a labeled RGBD dataset which contains scenes of stacked objects for robotic grasping. At last, we demonstrate the implementation of our shared CNN model on a real robotic platform and show that the robot can accurately discover a target object from the stack and successfully grasp it. Di Guo 0002, Tao Kong, Fuchun Sun 0001, Huaping Liu 0001 |
ICRA | 1 |
| 2015 | Discovery of topical object in image collectionsabstractAutomatic discovery of topical objects from a set of image collections provides more strong cognitive capability of robot to understand the unstructured environment. In this paper, we propose a novel framework based on dictionary learning for such a task. Different from existing work which utilizes multiple segmentations to coarsely obtain the object regions, we adopt the most recently developed objectness operator to extract candidate objects. Such a method admits a great advantage that the interested objects can be more reliably segmented. A dictionary learning method is proposed to discover the topical objects. Such an optimization model exploits the observation that any image only includes a few topical objects and therefore sparsity is encouraged. Further, a globally convergent algorithm is developed to solve the dictionary learning problem and extensive experiments show that the proposed method outperforms the state-of-the-arts. Huaping Liu 0001, Yunhui Liu 0003, Liming Huang, Fuchun Sun 0001, Di Guo 0002 |
ICRA | 5 |
| 2015 | Transmissive optical pretouch sensing for robotic graspingabstractRobotic grasping has been hindered by the inability of robots to perceive unstructured environments. Because these environments can be complex or dynamic, it is important to obtain additional and precise sensing information just before grasping. This paper expands upon the pretouch modality by introducing a transmissive optical sensor. It can unambiguously indicate the presence or lack of objects in close proximity. A wide variety of items that other sensors fail to sense, such as extremely soft or shiny objects, can be detected by the proposed sensor. The sensor is also fully integrated into the fingertips of the PR2 robotic platform and manufactured with inexpensive, commercially-available components. Several experiments are conducted to verify its utility in both environment perception and robotic grasping. It is shown that the perception information supplied by the sensor facilitates effective robotic grasping. Di Guo 0002, Patrick Lancaster, Liang-Ting Jiang, Fuchun Sun 0001, Joshua R. Smith 0001 |
IROS | 1 |
| 2014 | A system of robotic grasping with experience acquisition
Di Guo 0002, Fuchun Sun 0001, Chunfang Liu |
Sci. China Inf. Sci. | 1 |