EDBT 2026 Demo / reviewers in the wild / expert
Kaixin Bai
dblp:338/9572
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0001-8579-0547ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StereoAnything: Advanced Zero-Shot Stereo Imaging for Robotic Grasp Detection With Transparent ObjectsabstractGrasping transparent objects remains challenging for robotic systems due to their reflective and refractive properties, which distort depth perception and introduce background noise. Unlike humans, who leverage life experience to perceive depth intuitively, robotic algorithms often fail to generalize across different object types. To address this, we propose a novel framework inspired by human perception for grasping transparent objects. Our approach extends features extracted by foundation models to implicitly learn reconstruction strategies for transparent objects without requiring segmentation priors. Crucially, our framework maintains strong performance across all types of objects and scenes, preventing catastrophic forgetting of opaque objects while learning to perceive transparent ones. By integrating affordance information, our method dynamically guides a five-finger dexterous hand to execute diverse grasping strategies based on human intent. To tackle the challenge of annotating transparent objects, we constructed a large-scale synthetic dataset with depth information, affordance data, and automated annotations. Our framework demonstrates strong generalization, achieving a 96% grasp success rate in real-world robotic experiments and proving its broad applicability across varied environments. Kaixin Bai, Lei Zhang 0198, Zhaopeng Chen, Jianwei Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2025 | ContactDexNet: Multi-fingered Robotic Hand Grasping in Cluttered Environments through Hand-Object Contact Semantic MappingabstractThe deep learning models has significantly advanced dexterous manipulation techniques for multi-fingered hand grasping. However, the contact information-guided grasping in cluttered environments remains largely underexplored. To address this gap, we have developed ContactDexNet, a method for generating multi-fingered hand grasp samples in cluttered settings through contact semantic map. We introduce a contact semantic conditional variational autoencoder network (CoSe-CVAE) for creating comprehensive contact semantic map from object point cloud. We utilize grasp detection method to estimate hand grasp poses from the contact semantic map. Finally, an unified grasp evaluation model PointNetGPD++ is designed to assess grasp quality and collision probability, substantially improving the reliability of identifying optimal grasps in cluttered scenarios. Our grasp generation method has demonstrated remarkable success, outperforming state-of-the-art (SOTA) methods by at least 4.7%, with 81.0% average grasping success rate in real-world single-object grasping using a known hand, and by at least 9.0% when using an unknown hand. Moreover, in cluttered scenes, our method attains a 76.7% success rate, outperforming the SOTA method by 6.3%. We also proposed the multi-modal multi-fingered grasping dataset generation method. Our multi-fingered hand grasping dataset outperforms previous datasets in scene diversity, modality diversity. More details and supplementary materials can be found at https://sites.google.com/view/contact-dexnet. Lei Zhang 0035, Kaixin Bai, Guowen Huang, Zhenshan Bing, Zhaopeng Chen, Alois C. Knoll, Jianwei Zhang 0001 |
IROS | 2 |
| 2024 | A Collision-Aware Cable Grasping Method in Cluttered EnvironmentabstractWe introduce a Cable Grasping-Convolutional Neural Network (CG-CNN) designed to facilitate robust cable grasping in cluttered environments. Utilizing physics simulations, we generate an extensive dataset that mimics the intricacies of cable grasping, factoring in potential collisions between cables and robotic grippers. We employ the Approximate Convex Decomposition technique to dissect the non-convex cable model, with grasp quality autonomously labeled based on simulated grasping attempts. The CG-CNN is refined using this simulated dataset and enhanced through domain randomization techniques. Subsequently, the trained model predicts grasp quality, guiding the optimal grasp pose to the robot’s controller for execution. Grasping efficacy is assessed across both synthetic and real-world settings. Given our model’s implicit collision sensitivity, we achieved commendable success rates of 92.3% for known cables and 88.4% for unknown cables, surpassing contemporary state-of-the-art approaches. Supplementary materials can be found at https://leizhang-public.github.io/cg-cnn/. Lei Zhang 0198, Kaixin Bai, Qiang Li 0001, Zhaopeng Chen, Jianwei Zhang 0001 |
ICRA | 2 |
| 2024 | Close the Sim2real Gap via Physically-based Structured Light Synthetic Data SimulationabstractDespite the substantial progress in deep learning, its adoption in industrial robotics projects remains limited, primarily due to challenges in data acquisition and labeling. Previous sim2real approaches using domain randomization require extensive scene and model optimization. To address these issues, we introduce an innovative physically-based structured light simulation system, generating both RGB and physically realistic depth images, surpassing previous dataset generation tools. We create an RGBD dataset tailored for robotic industrial grasping scenarios and evaluate it across various tasks, including object detection, instance segmentation, and embedding sim2real visual perception in industrial robotic grasping. By reducing the sim2real gap and enhancing deep learning training, we facilitate the application of deep learning models in industrial settings. Project details are available at https://baikaixin-public.github.io/structured_light_3D_synthesizer/ Kaixin Bai, Lei Zhang 0198, Zhaopeng Chen, Jianwei Zhang 0001 |
ICRA | 1 |
| 2024 | ToolEENet: Tool Affordance 6D Pose EstimationabstractThe exploration of robotic dexterous hands utilizing tools has recently attracted considerable attention. A significant challenge in this field is the precise awareness of a tool’s pose when grasped, as occlusion by the hand often degrades the quality of the estimation. Additionally, the tool’s overall pose often fails to accurately represent the contact interaction, thereby limiting the effectiveness of vision-guided, contact-dependent activities. To overcome this limitation, we present the innovative TOOLEE dataset, which, to the best of our knowledge, is the first to feature affordance segmentation of a tool’s end-effector (EE) along with its defined 6D pose based on its usage. Furthermore, we propose the ToolEENet framework for accurate 6D pose estimation of the tool’s EE. This framework begins by segmenting the tool’s EE from raw RGB-D data, then uses a diffusion model-based pose estimator for 6D pose estimation at a category-specific level. Addressing the issue of symmetry in pose estimation, we introduce a symmetry-aware pose representation that enhances the consistency of pose estimation. Our approach excels in this field, demonstrating high levels of precision and generalization. Furthermore, it shows great promise for application in contact-based manipulation scenarios. All data and codes are available on the project website: https://tooleenet-iros2024.github.io/ Lei Zhang 0198, Yuyang Tu, Hui Zhang 0070, Kaixin Bai, Zhaopeng Chen, Jianwei Zhang 0001 |
IROS | 5 |