EDBT 2026 Demo / reviewers in the wild / expert
Jiayi Chen 0003
dblp:42/1159-3
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-5817-8317ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Robot manipulation · 66% 3D vision · 18% Transfer learning and domain adaptation · 6% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 13 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot manipulation › grasping
multifingered grasping |
3.1 | 4 | 2025 | BODex: Scalable and Efficient Robotic Dexterous Grasp Synthesis Using Bilevel Optimization · ICRA 2025 DexVLG: Dexterous Vision-Language-Grasp Model at Scale · ICCV 2025 DexGraspNet: A Large-Scale Robotic Dexterous Grasp Dataset for General Objects Based on Simulation · ICRA 2023 |
Robotics › Robot manipulation
grasping |
1.5 | 2 | 2025 | DexVLG: Dexterous Vision-Language-Grasp Model at Scale · ICCV 2025 UniDexGrasp: Universal Robotic Dexterous Grasping via Learning Diverse Proposal Generation and Goal-Conditioned Policy · CVPR 2023 |
Robotics › Robot manipulation › grasping › grasp planning
grasp synthesis |
1.5 | 2 | 2025 | BODex: Scalable and Efficient Robotic Dexterous Grasp Synthesis Using Bilevel Optimization · ICRA 2025 DexGraspNet: A Large-Scale Robotic Dexterous Grasp Dataset for General Objects Based on Simulation · ICRA 2023 |
Machine learning › Transfer learning and domain adaptation › generalization to unseen classes
cross-category generalization |
0.7 | 1 | 2023 | PartManip: Learning Cross-Category Generalizable Part Manipulation Policy from Point Cloud Observations · CVPR 2023 |
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
goal-conditioned policy |
0.7 | 1 | 2023 | UniDexGrasp: Universal Robotic Dexterous Grasping via Learning Diverse Proposal Generation and Goal-Conditioned Policy · CVPR 2023 |
Robotics › Robot manipulation › grasping › grasp planning
grasp pose generation |
0.7 | 1 | 2023 | UniDexGrasp: Universal Robotic Dexterous Grasping via Learning Diverse Proposal Generation and Goal-Conditioned Policy · CVPR 2023 |
Robotics › Robot manipulation › object manipulation
part manipulation |
0.7 | 1 | 2023 | PartManip: Learning Cross-Category Generalizable Part Manipulation Policy from Point Cloud Observations · CVPR 2023 |
Computer vision › 3D vision › pose estimation
rotation estimation |
0.6 | 1 | 2022 | Projective Manifold Gradient Layer for Deep Rotation Regression · CVPR 2022 |
Computer vision › 3D vision
pose estimation |
0.4 | 2 | 2025 | DexVLG: Dexterous Vision-Language-Grasp Model at Scale · ICCV 2025 Projective Manifold Gradient Layer for Deep Rotation Regression · CVPR 2022 |
Robotics › Robot manipulation › grasping › grasp detection
grasp pose estimation |
0.3 | 1 | 2025 | DexVLG: Dexterous Vision-Language-Grasp Model at Scale · ICCV 2025 |
Mathematical optimization
bilevel optimization |
0.3 | 1 | 2025 | BODex: Scalable and Efficient Robotic Dexterous Grasp Synthesis Using Bilevel Optimization · ICRA 2025 |
Computer vision › 3D vision › 3d shape reconstruction
object shape reconstruction |
0.2 | 1 | 2023 | Tracking and Reconstructing Hand Object Interactions from Point Cloud Sequences in the Wild · AAAI 2023 |
Computer vision › 3D vision
point cloud |
0.2 | 1 | 2023 | PartManip: Learning Cross-Category Generalizable Part Manipulation Policy from Point Cloud Observations · CVPR 2023 |
Methods — techniques the papers use, named apart from their topics
quadratic programming · 1.7gradient descent · 1.7CUDA · 1.7vision-language model · 0.9flow matching · 0.9bilevel optimization · 0.9bi-level optimization · 0.9optimization-based tracking · 0.7joint optimization · 0.7handtracknet · 0.7MANO hand model · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DexVLG: Dexterous Vision-Language-Grasp Model at ScaleabstractAs large models gain traction, vision-language-action (VLA) systems are enabling robots to tackle increasingly complex tasks. However, limited by the difficulty of data collection, progress has mainly focused on controlling simple gripper end-effectors. There is little research on functional grasping with large models for human-like dexterous hands. In this paper, we introduce DexVLG, a large Vision-Language-Grasp model for Dexterous grasp pose prediction aligned with language instructions using single-view RGBD input. To accomplish this, we generate a dataset of 170 million dexterous grasp poses mapped to semantic parts across 174,000 objects in simulation, paired with detailed part-level captions. This large-scale dataset, named DexGraspNet 3.0, is used to train a VLM and flow-matching-based pose head capable of producing instruction-aligned grasp poses for tabletop objects. To assess DexVLG's performance, we create benchmarks in physics-based simulations and conduct real-world experiments. Extensive testing demonstrates DexVLG's strong zero-shot generalization capabilities-achieving over 76% zero-shot execution success rate and state-of-the-art part-grasp accuracy in simulation-and successful part-aligned grasps on physical objects in real-world scenarios. Jiawei He 0002, Danshi Li, Xinqiang Yu, Zekun Qi, Jiayi Chen 0003, Zhaoxiang Zhang 0001, Zhizheng Zhang 0011, Li Yi 0001, He Wang 0010 |
ICCV | 6 |
| 2025 | BODex: Scalable and Efficient Robotic Dexterous Grasp Synthesis Using Bilevel OptimizationabstractRobotic dexterous grasping is important for interacting with the environment. To unleash the potential of data-driven models for dexterous grasping, a large-scale, highquality dataset is essential. While gradient-based optimization offers a promising way for constructing such datasets, previous works suffer from limitations, such as inefficiency, strong assumptions in the grasp quality energy, or limited object sets for experiments. Moreover, the lack of a standard benchmark for comparing different methods and datasets hinders progress in this field. To address these challenges, we develop a highly efficient synthesis system and a comprehensive benchmark with MuJoCo for dexterous grasping. We formulate grasp synthesis as a bilevel optimization problem, combining a novel lowerlevel quadratic programming (QP) with an upper-level gradient descent process. By leveraging recent advances in CUDAaccelerated robotic libraries and GPU-based QP solvers, our system can parallelize thousands of grasps and synthesize over 49 grasps per second on a single 3090 GPU. Our synthesized grasps for Shadow, Allegro, and Leap hands all achieve a success rate above 75 % in simulation, with a penetration depth under 1 mm, outperforming existing baselines on nearly all metrics. Compared to the previous large-scale dataset, DexGraspNet, our dataset significantly improves the performance of learning models, with a success rate from around 40 % to 80 % in simulation. Real-world testing of the trained model on the Shadow Hand achieves an 81 % success rate across 20 diverse objects. The codes and datasets are released on our project page: https://pku-epic.github.io/BODex. Jiayi Chen 0003, Yubin Ke, He Wang 0010 |
ICRA | 1 |
| 2024 | Task-Oriented Dexterous Hand Pose Synthesis Using Differentiable Grasp Wrench Boundary EstimatorabstractThis work tackles the problem of task-oriented dexterous hand pose synthesis, which involves generating a static hand pose capable of applying a task-specific set of wrenches to manipulate objects. Unlike previous approaches that focus solely on force-closure grasps, which are unsuitable for non-prehensile manipulation tasks (e.g., turning a knob or pressing a button), we introduce a unified framework covering force-closure grasps, non-force-closure grasps, and a variety of non-prehensile poses. Our key idea is a novel optimization objective quantifying the disparity between the Task Wrench Space (TWS, the desired wrenches predefined as a task prior) and the Grasp Wrench Space (GWS, the achievable wrenches computed from the current hand pose). By minimizing this objective, gradient-based optimization algorithms can synthe-size task-oriented hand poses without additional human demonstrations. Our specific contributions include 1) a fast, accurate, and differentiable technique for estimating the GWS boundary; 2) a task-oriented objective function based on the disparity between the estimated GWS boundary and the provided TWS boundary; and 3) an efficient implementation of the synthesis pipeline that leverages CUDA accelerations and supports large-scale parallelization. Experimental results on 10 diverse tasks demonstrate a 72.6% success rate in simulation. Furthermore, real-world validation for 4 tasks confirms the effectiveness of synthesized poses for manipulation. Notably, despite being primarily tailored for task-oriented hand pose synthesis, our pipeline can generate force-closure grasps 50 times faster than DexGraspNet while maintaining comparable grasp quality. Project page: https://pku-epic.github.io/TaskDexGrasp/. Jiayi Chen 0003, He Wang 0010 |
IROS | 1 |
| 2023 | Tracking and Reconstructing Hand Object Interactions from Point Cloud Sequences in the WildabstractIn this work, we tackle the challenging task of jointly tracking hand object poses and reconstructing their shapes from depth point cloud sequences in the wild, given the initial poses at frame 0. We for the first time propose a point cloud-based hand joint tracking network, HandTrackNet, to estimate the inter-frame hand joint motion. Our HandTrackNet proposes a novel hand pose canonicalization module to ease the tracking task, yielding accurate and robust hand joint tracking. Our pipeline then reconstructs the full hand via converting the predicted hand joints into a MANO hand. For object tracking, we devise a simple yet effective module that estimates the object SDF from the first frame and performs optimization-based tracking. Finally, a joint optimization step is adopted to perform joint hand and object reasoning, which alleviates the occlusion-induced ambiguity and further refines the hand pose. During training, the whole pipeline only sees purely synthetic data, which are synthesized with sufficient variations and by depth simulation for the ease of generalization. The whole pipeline is pertinent to the generalization gaps and thus directly transferable to real in-the-wild data. We evaluate our method on two real hand object interaction datasets, e.g. HO3D and DexYCB, without any fine-tuning. Our experiments demonstrate that the proposed method significantly outperforms the previous state-of-the-art depth-based hand and object pose estimation and tracking methods, running at a frame rate of 9 FPS. We have released our code on https://github.com/PKU-EPIC/HOTrack. Jiayi Chen 0003, Mi Yan, Jiazhao Zhang, Yinzhen Xu, Yijia Weng, Li Yi 0001, Shuran Song, He Wang 0010 |
AAAI | 1 |
| 2023 | PartManip: Learning Cross-Category Generalizable Part Manipulation Policy from Point Cloud ObservationsabstractLearning a generalizable object manipulation policy is vital for an embodied agent to work in complex real-world scenes. Parts, as the shared components in different object categories, have the potential to increase the generalization ability of the manipulation policy and achieve cross-category object manipulation. In this work, we build the first large-scale, part-based cross-category object manipulation benchmark, PartManip, which is composed of 11 object categories, 494 objects, and 1432 tasks in 6 task classes. Compared to previous work, our benchmark is also more diverse and realistic, i.e., having more objects and using sparse-view point cloud as input without oracle information like part segmentation. To tackle the difficulties of vision-based policy learning, we first train a statebased expert with our proposed part-based canonicalization and part-aware rewards, and then distill the knowledge to a vision-based student. We also find an expressive backbone is essential to overcome the large diversity of different objects. For cross-category generalization, we introduce domain adversarial learning for domain-invariant feature extraction. Extensive experiments in simulation show that our learned policy can outperform other methods by a large margin, especially on unseen object categories. We also demonstrate our method can successfully manipulate novel objects in the real world. Our benchmark has been released in https://pku-epic.github.io/PartManip. Yiran Geng, Jiayi Chen 0003, Hao Dong 0003, He Wang 0010 |
CVPR | 4 |
| 2023 | UniDexGrasp: Universal Robotic Dexterous Grasping via Learning Diverse Proposal Generation and Goal-Conditioned PolicyabstractIn this work, we tackle the problem of learning universal robotic dexterous grasping from a point cloud observation under a table-top setting. The goal is to grasp and lift up objects in high-quality and diverse ways and generalize across hundreds of categories and even the unseen. Inspired by successful pipelines used in parallel gripper grasping, we split the task into two stages: 1) grasp proposal (pose) generation and 2) goal-conditioned grasp execution. For the first stage, we propose a novel probabilistic model of grasp pose conditioned on the point cloud observation that factorizes rotation from translation and articulation. Trained on our synthesized large-scale dexterous grasp dataset, this model enables us to sample diverse and high-quality dexterous grasp poses for the object point cloud. For the second stage, we propose to replace the motion planning used in parallel gripper grasping with a goal-conditioned grasp policy, due to the complexity involved in dexterous grasping execution. Note that it is very challenging to learn this highly generalizable grasp policy that only takes realistic inputs without oracle states. We thus propose several important innovations, including state canonicalization, object curriculum, and teacher-student distillation. In-tegrating the two stages, our final pipeline becomes the first to achieve universal generalization for dexterous grasping, demonstrating an average success rate of more than 60% on thousands of object instances, which significantly out-performs all baselines, meanwhile showing only a minimal generalization gap. Yinzhen Xu, Weikang Wan, Zikang Shan, Hao Shen 0015, Ruicheng Wang, Yijia Weng, Jiayi Chen 0003, Tengyu Liu, Li Yi 0001, He Wang 0010 |
CVPR | 10 |
| 2023 | DexGraspNet: A Large-Scale Robotic Dexterous Grasp Dataset for General Objects Based on SimulationabstractRobotic dexterous grasping is the first step to enable human-like dexterous object manipulation and thus a crucial robotic technology. However, dexterous grasping is much more under-explored than object grasping with parallel grippers, partially due to the lack of a large-scale dataset. In this work, we present a large-scale robotic dexterous grasp dataset, DexGraspNet, generated by our proposed highly efficient synthesis method that can be generally applied to any dexterous hand. Our method leverages a deeply accelerated differentiable force closure estimator and thus can efficiently and robustly synthesize stable and diverse grasps on a large scale. We choose ShadowHand and generate 1.32 million grasps for 5355 objects, covering more than 133 object categories and containing more than 200 diverse grasps for each object instance, with all grasps having been validated by the Isaac Gym simulator. Compared to the previous dataset from Liu et al. generated by GraspIt!, our dataset has not only more objects and grasps, but also higher diversity and quality. Via performing cross-dataset experiments, we show that training several algorithms of dexterous grasp synthesis on our dataset significantly outperforms training on the previous one. To access our data and code, including code for human and Allegro grasp synthesis, please visit our project page: https://pku-epic.github.io/DexGraspNet/. Ruicheng Wang, Jiayi Chen 0003, Yinzhen Xu, Puhao Li, Tengyu Liu, He Wang 0010 |
ICRA | 3 |
| 2022 | Projective Manifold Gradient Layer for Deep Rotation RegressionabstractRegressing rotations on SO(3) manifold using deep neural networks is an important yet unsolved problem. The gap between the Euclidean network output space and the non-Euclidean SO(3) manifold imposes a severe challenge for neural network learning in both forward and backward passes. While several works have proposed different regression-friendly rotation representations, very few works have been devoted to improving the gradient back-propagating in the backward pass. In this paper, we propose a manifold-aware gradient that directly backpropagates into deep network weights. Leveraging Riemannian optimization to construct a novel projective gradient, our proposed regularized projective manifold gradient (RPMG) method helps networks achieve new state-of-the-art performance in a variety of rotation estimation tasks. Our proposed gradient layer can also be applied to other smooth manifolds such as the unit sphere. Our project page is at https://jychen18.github.io/RPMG. Jiayi Chen 0003, Yingda Yin, Tolga Birdal, Baoquan Chen, Leonidas J. Guibas, He Wang 0010 |
CVPR | 1 |