EDBT 2026 Demo / reviewers in the wild / expert
Xinyu Zhan 0001
dblp:257/1454-1
· DBLP profile ↗
10ranked-venue papers
1as first author
10since 2021 · last 2025
0009-0004-7859-2592ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dense Policy: Bidirectional Autoregressive Learning of Actions
Xinyu Zhan 0001, Hongjie Fang, Hao-Shu Fang, Yong-Lu Li 0001, Cewu Lu, Lixin Yang 0001 |
ICCV | 2 |
| 2025 | Multi-View Hand Reconstruction With a Point-Embedded TransformerabstractThis work introduces a novel and generalizable multi-view Hand Mesh Reconstruction (HMR) model, named POEM, designed for practical use in real-world hand motion capture scenarios. The advances of the POEM model consist of two main aspects. First, concerning the modeling of the problem, we propose embedding a static basis point within the multi-view stereo space. A point represents a natural form of 3D information and serves as an ideal medium for fusing features across different views, given its varied projections across these views. Consequently, our method harnesses a simple yet effective idea: a complex 3D hand mesh can be represented by a set of 3D basis points that 1) are embedded in the multi-view stereo, 2) carry features from the multi-view images, and 3) encompass the hand in it. The second advance lies in the training strategy. We utilize a combination of five large-scale multi-view datasets and employ randomization in the number, order, and poses of the cameras. By processing such a vast amount of data and a diverse array of camera configurations, our model demonstrates notable generalizability in the real-world applications. As a result, POEM presents a highly practical, plug-and-play solution that enables user-friendly, cost-effective multi-view motion capture for both left and right hands. Lixin Yang 0001, Licheng Zhong, Pengxiang Zhu, Xinyu Zhan 0001, Junxiao Kong, Jian Xu 0027, Cewu Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | FAVOR: Full-Body AR-Driven Virtual Object Rearrangement Guided by Instruction TextabstractRearrangement operations form the crux of interactions between humans and their environment. The ability to generate natural, fluid sequences of this operation is of essential value in AR/VR and CG. Bridging a gap in the field, our study introduces FAVOR: a novel dataset for Full-body AR-driven Virtual Object Rearrangement that uniquely employs motion capture systems and AR eyeglasses. Comprising 3k diverse motion rearrangement sequences and 7.17 million interaction data frames, this dataset breaks new ground in research data. We also present a pipeline FAVORITE for producing digital human rearrangement motion sequences guided by instructions. Experimental results, both qualitative and quantitative, suggest that this dataset and pipeline deliver high-quality motion sequences. Our dataset, code, and appendix are available at https://kailinli.github.io/FAVOR. Kailin Li 0001, Lixin Yang 0001, Zenan Lin, Jian Xu 0027, Xinyu Zhan 0001, Yifei Zhao 0003, Pengxiang Zhu, Wenxiong Kang, Kejian Wu, Cewu Lu |
AAAI | 5 |
| 2024 | OakInk2 : A Dataset of Bimanual Hands-Object Manipulation in Complex Task CompletionabstractWe present Oakink2, a dataset of bimanual object manipulation tasks for complex daily activities. In pursuit of constructing the complex tasks into a structured representation, Oakink2 introduces three level of abstraction to organize the manipulation tasks: Affordance, Primitive Task, and Complex Task. OAKINK2 features on an object-centric perspective for decoding the complex tasks, treating them as a sequence of object affordance fulfillment. The first level, Affordance, outlines the functionalities that objects in the scene can afford, the second level, Primitive Task, describes the minimal interaction units that humans interact with the object to achieve its affordance, and the third level, Complex Task, illustrates how Primitive Tasks are composed and interdependent. Oakink2 dataset provides multi-view image streams and precise pose annotations for the human body, hands and various interacting objects. This extensive collection supports applications such as interaction reconstruction and motion synthesis. Based on the 3-level abstraction of Oakink2, we explore a task-oriented framework for Complex Task Completion (CTC). CTC aims to generate a sequence of bimanual manipulation to achieve task objectives. Within the CTC framework, we employ Large Language Models (LLMs) to decompose the complex task objectives into sequences of Primitive Tasks and have developed a Motion Fulfillment Model that generates bimanual hand motion for each Primitive Task. Oakink2 datasets and models are available at https://oakink.net/v2. Xinyu Zhan 0001, Lixin Yang 0001, Yifei Zhao 0003, Kangrui Mao, Hanlin Xu, Zenan Lin, Kailin Li 0001, Cewu Lu |
CVPR | 1 |
| 2024 | Learning a Contact Potential Field for Modeling the Hand-Object InteractionabstractEstimating and synthesizing the hand's manipulation of objects is central to understanding human behaviour. To accurately model the interaction between the hand and object (referred to as the "hand-object"), we must not only focus on the pose of the hand and object, but also consider the contact between them. This contact provides valuable information for generating semantically and physically plausible grasps. In this paper, we propose an explicit contact representation called Contact Potential Field (CPF). In CPF, we model the contact between a pair of hand-object vertices as a spring-mass system. This system encodes the distance of the pair, as well as a likelihood of that contact being stable. Therefore, the system of multiple extended and compressed springs forms an elastic potential field with minimal energy at the optimal grasp position. We apply CPF to two relevant tasks, namely, hand-object pose estimation and grasping pose generation. Extensive experiments on the two challenging tasks and three commonly used datasets have demonstrated that our method can achieve state-of-the-art in several reconstruction metrics, allowing us to produce more physically plausible hand-object poses even when the ground-truth exhibits severe interpenetration or disjointedness. Lixin Yang 0001, Xinyu Zhan 0001, Kailin Li 0001, Cewu Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | POEM: Reconstructing Hand in a Point Embedded Multi-view StereoabstractEnable neural networks to capture 3D geometrical-aware features is essential in multi-view based vision tasks. Previous methods usually encode the 3D information of multi-view stereo into the 2D features. In contrast, we present a novel method, named POEM, that directly operates on the 3D POints Embedded in the Multi-view stereo for reconstructing hand mesh in it. Point is a natural form of 3D information and an ideal medium for fusing features across views, as it has different projections on different views. Our method is thus in light of a simple yet effective idea, that a complex 3D hand mesh can be represented by a set of 3D points that 1) are embedded in the multi-view stereo, 2) carry features from the multi-view images, and 3) encircle the hand. To leverage the power of points, we design two operations: point-based feature fusion and cross-set point attention mechanism. Evaluation on three challenging multi-view datasets shows that POEM outperforms the state-of-the-art in hand mesh reconstruction. Code and models are available for research at github.com/lixiny/POEM Lixin Yang 0001, Jian Xu 0027, Licheng Zhong, Xinyu Zhan 0001, Zhicheng Wang 0007, Kejian Wu, Cewu Lu |
CVPR | 4 |
| 2023 | Chord: Category-level Hand-held Object Reconstruction via Shape DeformationabstractIn daily life, humans utilize hands to manipulate objects. Modeling the shape of objects that are manipulated by the hand is essential for AI to comprehend daily tasks and to learn manipulation skills. However, previous approaches have encountered difficulties in reconstructing the precise shapes of hand-held objects, primarily owing to a deficiency in prior shape knowledge and inadequate data for training. As illustrated, given a particular type of tool, such as a mug, despite its infinite variations in shape and appearance, humans have a limited number of ‘effective’ modes and poses for its manipulation. This can be attributed to the fact that humans have mastered the shape prior of the ‘mug’ category, and can quickly establish the corresponding relations between different mug instances and the prior, such as where the rim and handle are located. In light of this, we propose a new method, Chord, for Category-level Hand-held Object Reconstruction via shape Deformation. Chord deforms a categorical shape prior for reconstructing the intra-class objects. To ensure accurate reconstruction, we empower Chord with three types of awareness: appearance, shape, and interacting pose. In addition, we have constructed a new dataset, Comic, of category-level hand-object interaction. Comic contains a rich array of object instances, materials, hand interactions, and viewing directions. Extensive evaluation shows that Chord outperforms state-of-the-art approaches in both quantitative and qualitative measures. Code, model, and datasets are available at https://kailinli.github.io/CHORD Kailin Li 0001, Lixin Yang 0001, Haoyu Zhen, Zenan Lin, Xinyu Zhan 0001, Licheng Zhong, Jian Xu 0027, Kejian Wu, Cewu Lu |
ICCV | 5 |
| 2022 | ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and SynthesisabstractEstimating the articulated 3D hand-object pose from a single RGB image is a highly ambiguous and challenging problem, requiring large-scale datasets that contain diverse hand poses, object types, and camera viewpoints. Most real-world datasets lack these diversities. In contrast, data synthesis can easily ensure those diversities separately. However, constructing both valid and diverse hand-object interactions and efficiently learning from the vast synthetic data is still challenging. To address the above issues, we propose ArtiBoost, a lightweight online data enhancement method. ArtiBoost can cover diverse hand-object poses and camera viewpoints through sampling in a Composited hand-object Configuration and View-point space (CCV-space) and can adaptively enrich the current hard-discernable items by loss-feedback and sample re-weighting. ArtiBoost alternatively performs data exploration and synthesis within a learning pipeline, and those synthetic data are blended into real-world source data for training. We apply ArtiBoost on a simple learning baseline network and witness the performance boost on several hand-object benchmarks. Our models and code are available at https://github.com/lixiny/ArtiBoost. Lixin Yang 0001, Kailin Li 0001, Xinyu Zhan 0001, Cewu Lu |
CVPR | 3 |
| 2022 | OakInk: A Large-scale Knowledge Repository for Understanding Hand-Object InteractionabstractLearning how humans manipulate objects requires machines to acquire knowledge from two perspectives: one for understanding object affordances and the other for learning human's interactions based on the affordances. Even though these two knowledge bases are crucial, we find that current databases lack a comprehensive awareness of them. In this work, we propose a multi-modal and rich-annotated knowledge repository, OakInk, for visual and cognitive understanding of hand-object interactions. We start to collect 1,800 common household objects and annotate their affordances to construct the first knowledge base: Oak. Given the affordance, we record rich human interactions with 100 selected objects in Oak. Finally, we transfer the interactions on the 100 recorded objects to their virtual counterparts through a novel method: Tink. The recorded and transferred hand-object interactions constitute the second knowledge base: Ink. As a result, OakInk contains 50,000 distinct affordance-aware and intent-oriented hand-object interactions. We benchmark OakInk on pose estimation and grasp generation tasks. Moreover, we propose two practical applications of OakInk: intent-based interaction generation and handover generation. Our dataset and source code are publicly available at www.oakink.net. Lixin Yang 0001, Kailin Li 0001, Xinyu Zhan 0001, Fei Wu 0001, Anran Xu 0003, Liu Liu 0012, Cewu Lu |
CVPR | 3 |
| 2021 | CPF: Learning a Contact Potential Field to Model the Hand-Object InteractionabstractModeling the hand-object (HO) interaction not only requires estimation of the HO pose, but also pays attention to the contact due to their interaction. Significant progress has been made in estimating hand and object separately with deep learning methods, simultaneous HO pose estimation and contact modeling has not yet been fully explored. In this paper, we present an explicit contact representation namely Contact Potential Field (CPF), and a learning-fitting hybrid framework namely MIHO to Modeling the Interaction of Hand and Object. In CPF, we treat each contacting HO vertex pair as a spring-mass system. Hence the whole system forms a potential field with minimal elastic energy at the grasp position. Extensive experiments on the two commonly used benchmarks have demonstrated that our method can achieve state-of-the-art in several reconstruction metrics, and allow us to produce more physically plausible HO pose even when the ground-truth exhibits severe interpenetration or disjointedness. Our code is available at https://github.com/lixiny/CPF. Lixin Yang 0001, Xinyu Zhan 0001, Kailin Li 0001, Cewu Lu |
ICCV | 2 |