EDBT 2026 Demo / reviewers in the wild / expert
Tao Lu 0006
dblp:03/5189-6
· DBLP profile ↗
18ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0003-3374-5845ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 10 since 2021Systems, architecture and hardware · 9 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HOCOpt: Hand-Object Contact Optimization to Improve Pose Estimation in Physical InteractionsabstractReconstructing hand–object physical interaction through visual sensing is crucial to understanding human intentions and guaranteeing the safety of human–robot collaboration in industrial applications. Due to heavy occlusion and cluttered backgrounds, existing methods generate inaccurate pose estimation, leading to unrealistic physical interactions between hands and objects. In this article, we present a novel contact-driven pose optimization framework called hand–object contact optimization (HOCOpt) to achieve accurate hand–object pose estimation. The HOCOpt includes two parts: contact estimation and pose optimization. In the contact estimation, we propose a contact region estimation network (CREN) to predict the potential contact across the hand–object meshes with inaccurate poses. A novel contact entropy weight and an auxiliary network are introduced to the training process of CREN to accelerate the model learning and improve the prediction accuracy. For pose optimization, a two-stage hand–object pose optimization method is utilized to refine inaccurate poses by considering both contact distribution and contact stability. During optimization, an orientation-aware differentiable contact model is introduced to account for hand deformation and contact forces to achieve accurate contact modeling. Extensive experiments on ContactPose, HO3D, and DexYCB datasets show that our approach outperforms the existing baselines. Besides, experiments on physical interaction tasks for human–robot collaboration are conducted to demonstrate the practical significance of HOCOpt in industrial scenarios. Xiaoge Cao, Tao Lu 0006, Wenhao Yu 0011, Yinghao Cai, Shuo Wang 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | NeuGrasp: Generalizable Neural Surface Reconstruction with Background Priors for Material-Agnostic Object Grasp DetectionabstractRobotic grasping in scenes with transparent and specular objects presents great challenges for methods relying on accurate depth information. In this paper, we introduce NeuGrasp, a neural surface reconstruction method that leverages background priors for material-agnostic grasp detection. NeuGrasp integrates transformers and global prior volumes to aggregate multi-view features with spatial encoding, enabling robust surface reconstruction in narrow and sparse viewing conditions. By focusing on foreground objects through residual feature enhancement and refining spatial perception with an occupancy-prior volume, NeuGrasp excels in handling objects with transparent and specular surfaces. Extensive experiments in both simulated and real-world scenarios show that NeuGrasp outperforms state-of-the-art methods in grasping while maintaining comparable reconstruction quality. More details are available at https://neugrasp.github.io/. Qingyu Fan, Yinghao Cai, Wenzhe He, Tao Lu 0006, Shuo Wang 0001 |
ICRA | 6 |
| 2025 | MISCGrasp: Leveraging Multiple Integrated Scales and Contrastive Learning for Enhanced Volumetric GraspingabstractRobotic grasping faces challenges in adapting to objects with varying shapes and sizes. In this paper, we introduce MISCGrasp, a volumetric grasping method that integrates multi-scale feature extraction with contrastive feature enhancement for self-adaptive grasping. We propose a query-based interaction between high-level and low-level features through the Insight Transformer, while the Empower Transformer selectively attends to the highest-level features, which synergistically strikes a balance between focusing on fine geometric details and overall geometric structures. Furthermore, MISCGrasp utilizes multi-scale contrastive learning to exploit similarities among positive grasp samples, ensuring consistency across multi-scale features. Extensive experiments in both simulated and real-world environments demonstrate that MISCGrasp outperforms baseline and variant methods in tabletop decluttering tasks. More details are available at https://miscgrasp.github.io/. Qingyu Fan, Yinghao Cai, Chunting Jiao, Tao Lu 0006, Shuo Wang 0001 |
IROS | 6 |
| 2025 | SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement LearningabstractPreference-based Reinforcement Learning (PbRL) methods provide a solution to avoid reward engineering by learning reward models based on human preferences. However, poor feedback- and sample- efficiency still remain the problems that hinder the application of PbRL. In this paper, we present a novel efficient query selection and preference-guided exploration method, called SENIOR, which could select the meaningful and easy-to-comparison behavior segment pairs to improve human feedback-efficiency and accelerate policy learning with the designed preference-guided intrinsic rewards. Our key idea is twofold: (1) We designed a Motion-Distinction-based Selection scheme (MDS). It selects segment pairs with apparent motion and different directions through kernel density estimation of states, which is more task-related and easy for human preference labeling; (2) We proposed a novel preference-guided exploration method (PGE). It encourages the exploration towards the states with high preference and low visits and continuously guides the agent achieving the valuable samples. The synergy between the two mechanisms could significantly accelerate the progress of reward and policy learning. Our experiments show that SENIOR outperforms other five existing methods in both human feedback-efficiency and policy convergence speed on six complex robot manipulation tasks from simulation and four real-worlds. Videos can be found on our project website: https://2025senior.github.io/ Hexian Ni, Tao Lu 0006, Haoyuan Hu, Yinghao Cai, Shuo Wang 0001 |
IROS | 2 |
| 2025 | Learn-Gen-Plan: Bridging the Gap Between Vision Language Models and Real-World Long-Horizon Dexterous ManipulationsabstractLong-horizon dexterous tasks have been a long-standing problem in robotic manipulation. Previous studies have developed task and motion planning, imitation learning, and reinforcement learning methods for long-horizon manipulations. However, these methods are hard to achieve efficient planning for new tasks. Empowered with the Vision Language Model (VLM), recent studies significantly improve the generalization of robot systems. However, these works are only verified in simple pick-and-place tasks due to limited skills. To this end, we propose the Learn-Gen-Plan (LGP), which combines the VLM and learning-based primitives to endow robots with the ability to efficiently plan and complete various long-horizon dexterous tasks. LGP contains two key phases: skill generation and task planning. In skill generation, the Skill Generator is proposed to utilize the learned key primitives and hand-crafted trivial primitives to generate adaptive robot skills. In task planning, the Multimodal Planner generates the robot plan based on image observation, generated skills, and text prompts. We set up a series of dexterous tasks (e.g., cable routing, peg-in-hole assembly) in a real-world lighting circuit wiring scenario to evaluate LGP. The experimental results show that LGP efficiently generates robot plans with learned skills, controlling the robot to complete various multi-step cable wiring tasks. Peng Hao 0003, Shaowei Cui, Junhang Wei, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2022 | Meta-Imitation Learning by Watching Video Demonstrations
Tao Lu 0006, Xiaoge Cao, Yinghao Cai, Shuo Wang 0001 |
ICLR | 2 |
| 2022 | Joint Self-Supervised Monocular Depth Estimation and SLAMabstractClassical monocular Simultaneous Localization and Mapping (SLAM) and convolutional neural networks (CNNs) based monocular depth estimation represent two different methods towards reconstructing the 3D geometry of the scene. In this paper, we leverage SLAM and depth estimation for their respective advantages to further improve the performance of both tasks. For SLAM, running pseudo RGBD-SLAM with CNN-predicted depths improves the accuracy of visual odometry and mapping compared with the monocular SLAM baseline. For depth estimation, we use 3D scene structures from geometric SLAM to refine the pre-trained monocular depth estimation network to update the model which did not reach the optimum due to the photometric inconsistency. Moreover, the proposed method incorporates an optional Sparse Auxiliary Network [1] into the original depth estimation network, from which the sparse depth features are dynamically combined with RGB features for predicting the depth map. Experimental results on KITTI and TUM RGB-D datasets show that our method achieves state-of-the-art performances on both depth prediction and pose estimation tasks. Xiaoxia Xing, Yinghao Cai, Tao Lu 0006, Dayong Wen |
ICPR | 3 |
| 2022 | VGPN: 6-DoF Grasp Pose Detection Network Based on Hough VotingabstractIn this paper, we propose a novel Voting based Grasp Pose Network (VGPN) to detect 6-DoF grasps in cluttered scenes. The motivation of this paper is that local object geometry can provide useful clues about where the object can be grasped. Generated by the sampled seed points from raw point cloud, the votes allow seed points in different object regions to contribute to locations where the object can be grasped. Geometric features from various local regions are aggregated to generate grasps in a more confident and dense space, which enables grasp prediction utilizing more global context features. The search space of grasp pose detection is also greatly reduced. Experimental results on both simulation and real-world environments show that our proposed method outperforms state-of-the-art approaches in terms of both success rate and coverage of the ground truth grasps. The objects can be grasped with fewer attempts which is critical in real-world applications. Yinghao Cai, Tao Lu 0006, Shuo Wang 0001 |
IROS | 3 |
| 2022 | Manipulation skill learning on multi-step complex task based on explicit and implicit curriculum learning
Naijun Liu, Tao Lu 0006, Yinghao Cai, Rui Wang 0031, Shuo Wang 0001 |
Sci. China Inf. Sci. | 2 |
| 2021 | Hierarchical Learning from Demonstrations for Long-Horizon TasksabstractAlthough reinforcement learning (RL) has achieved great success in robotic manipulation skills learning, it is still challenging for long-horizon tasks. Combining RL with demonstrations is an effective solution. In this paper, we propose a novel hierarchical learning from demonstrations method for long-horizon tasks, which leverages (i) object-centered segmentation of demonstrations to automatically segment the teaching trajectories into episodes. (ii) a bi-level hierarchical imitation learning method with a parallel training mechanism to train the two-level policies simultaneously. Experimental results on three challenging long-horizon tasks with sparse rewards show that our proposed method significantly outperforms state-of-art approaches in terms of both sample-efficiency and success rate. Moreover, our method is the only one which achieves satisfactory performance in tasks of multi-object stack and multi-object push&stack. Boyao Li, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001 |
ICRA | 3 |
| 2021 | DIMSAN: Fast Exploration with the Synergy between Density-based Intrinsic Motivation and Self-adaptive Action NoiseabstractExploration in environments with sparse rewards remains a challenging problem in Deep Reinforcement Learning (DRL). For the off-policy method, it usually needs a large number of training samples. With the growing dimensions of state and action space, this method becomes more and more sample-inefficient. In this paper, we propose a novel fast exploration method for off-policy reinforcement learning, called Density-based Intrinsic Motivation and Self-adaptive Action Noise (DIMSAN). Our main contribution is twofold: (1) We propose a Density-based Intrinsic Motivation (DIM) method. It introduces a new intrinsic-reward generation mechanism based on samples’ density estimation during experience replay and encourages the agent to seek novel and unfamiliar states. (2) We propose a Self-adaptive Action Noise (SAN) to deal with the exploration-exploitation tradeoffs, which could automatically change the exploration step through adding adaptive action space noise. The synergy between DIM and SAN could guide the agent to search the state and action space with high efficiency. We evaluate our method on the benchmark manipulation tasks and the designed challenging ones. Empirical results show that our method outperforms the existing methods in terms of convergence speed and sample efficiency, especially in challenging tasks. Boyao Li, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001 |
ICRA | 3 |
| 2021 | 3DTDesc: learning local features using 2D and 3D cues
Xiaoxia Xing, Yinghao Cai, Tao Lu 0006, Dayong Wen |
Mach. Vis. Appl. | 3 |
| 2021 | Correction to: 3DTDesc: learning local features using 2D and 3D cues
Xiaoxia Xing, Yinghao Cai, Tao Lu 0006, Dayong Wen |
Mach. Vis. Appl. | 3 |
| 2020 | Dynamic Guided Network for Monocular Depth EstimationabstractSelf-attention and encoder-decoder have been widely used in the deep neural network for monocular depth estimation. The self-attention mechanism is capable of capturing long-range dependencies by computing the representation of each image position by a weighted sum of the features at all positions, while the encoder-decoder can capture detailed structural information by gradually recovering spatial information. In this work, we combine the advantages of both methods. Specifically, our proposed model, DGNet, extends EMANet [1] by adding an effective decoder module to progressively refine the coarse depth map. In the decoder stage, we design a dynamic guided upsampling module that employs dynamically generated kernel conditioned on low-level features to guide the upsampling of the coarse depth map. Experimental results demonstrate that our method obtains higher accuracy and generates visually pleasant depth maps. Xiaoxia Xing, Yinghao Cai, Tao Lu 0006, Dayong Wen |
ICPR | 4 |
| 2020 | ACDER: Augmented Curiosity-Driven Experience ReplayabstractExploration in environments with sparse feed-back remains a challenging research problem in reinforcement learning (RL). When the RL agent explores the environment randomly, it results in low exploration efficiency, especially in robotic manipulation tasks with high dimensional continuous state and action space. In this paper, we propose a novel method, called Augmented Curiosity-Driven Experience Replay (ACDER), which leverages (i) a new goal-oriented curiosity-driven exploration to encourage the agent to pursue novel and task-relevant states more purposefully and (ii) the dynamic initial states selection as an automatic exploratory curriculum to further improve the sample-efficiency. Our approach complements Hindsight Experience Replay (HER) by introducing a new way to pursue valuable states. Experiments conducted on four challenging robotic manipulation tasks with binary rewards, including Reach, Push, Pick&Place and Multi-step Push. The empirical results show that our proposed method significantly outperforms existing methods in the first three basic tasks and also achieves satisfactory performance in multi-step robotic task learning. Boyao Li, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001 |
ICRA | 2 |
| 2019 | Localizing Discriminative Visual Landmarks for Place RecognitionabstractWe address the problem of visual place recognition with perceptual changes. The fundamental problem of visual place recognition is generating robust image representations which are not only insensitive to environmental changes but also distinguishable to different places. Taking advantage of the feature extraction ability of Convolutional Neural Networks (CNNs), we further investigate how to localize discriminative visual landmarks that positively contribute to the similarity measurement, such as buildings and vegetations. In particular, a Landmark Localization Network (LLN) is designed to indicate which regions of an image are used for discrimination. Detailed experiments are conducted on open source datasets with varied appearance and viewpoint changes. The proposed approach achieves superior performance against state-of-the-art methods. Zhe Xin, Yinghao Cai, Tao Lu 0006, Xiaoxia Xing, Shaojun Cai, Jixiang Zhang 0001 |
ICRA | 3 |
| 2019 | Self-modeling Tracking Control of Crawler Fire Fighting Robot Based on Causal Network*abstractIn this paper, a self-modeling method based on a causal network is proposed for the tracking control of the Crawler Fire Fighting Robot (CFFR). The method mainly consists of two parts, one is a motion model, based on data driving, learning to establish the correspondence between control signal sequence and vehicle motion, estimating the motion state of the next moment from historical data, eliminating complex CFFR modeling. The other is the tracking network. Based on the simulation data of the motion model, the relationship between the target trajectory and the current control command is learned, which simplifies the design and cumbersome tuning of the complex controller. The effectiveness of the proposed method is verified in both simulated and real-world environments. Qualitative and quantitative experimental results verify the accuracy of the tracking. Wenkai Chang, Caiyun Yang, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001 |
IROS | 4 |
| 2018 | 3DTNet: Learning Local Features Using 2D and 3D CuesabstractWe present an approach to learn 3D local descriptor by combining both 2D texture and 3D geometric information, which can be used to register partial 3D data for a variety of vision applications. Unlike previous approaches which simply concatenate features learned from multiple sources into one feature descriptor, we learn 2D and 3D feature representations jointly. We design a network, 3DTNet with an architecture particularly designed for learning robust local feature representation leveraging both texture and geometric information. Two types of information are interacted with each other which results in more robust and stable feature representation. Finally, feature representations of multi-scale neighborhoods are aggregated to further improve the performance of feature matching. Extensive experimental results show that our method outperforms state-of-art 2D or 3D descriptors in terms of both accuracy and efficiency. Xiaoxia Xing, Yinghao Cai, Tao Lu 0006, Shaojun Cai, Dayong Wen |
3DV | 3 |