VLDB 2026 Research / reviewers in the wild / expert
Lixin Yang 0001
dblp:59/4517-1
· DBLP profile ↗
19ranked-venue papers
7as first author
18since 2021 · last 2025
0000-0001-6366-3192ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GaPT-DAR: Category-level Garments Pose Tracking via Integrated 2D Deformation and 3D ReconstructionabstractGarments are common in daily life and are important for embodied intelligence community. Current category-level garments pose tracking works focus on predicting point-wise canonical correspondence and learning a shape deformation in point cloud sequences. In this paper, motivated by the 2D warping space and shape prior, we propose GaPT-DAR, a novel category-level Garments Pose Tracking framework with integrated 2D Deformation And 3D Reconstruction function, which fully utilize 3D-2D projection and 2D-3D reconstruction to transform the 3D point-wise learning into 2D warping deformation learning. Specifically, GaPT-DAR firstly builds a Voting-based Project module that learns the optimal 3D-2D projection plane for maintaining the maximum orthogonal entropy during point projection. Next, a Garments Deformation module is designed in 2D space to explicitly model the garments warping procedure with deformation parameters. Finally, we build a Depth Reconstruction module to recover the 2D images into 3D warp field. We provide extensive experiments on VR-Folding dataset to evaluate our GaPT-DAR and the results show obvious improvements on most of the metrics compared to state-of-the-arts (i.e. Garment-Nets [8] and GarmentTracking [32]). More details are available at https://sites.google.com/view/gapt-dar. Li Zhang 0104, Qiaojun Yu, Lixin Yang 0001, Yong-Lu Li 0001, Cewu Lu, Rujing Wang, Liu Liu 0012 |
CVPR | 5 |
| 2025 | Dense Policy: Bidirectional Autoregressive Learning of Actions
Xinyu Zhan 0001, Hongjie Fang, Hao-Shu Fang, Yong-Lu Li 0001, Cewu Lu, Lixin Yang 0001 |
ICCV | 8 |
| 2025 | HybrIK-X: Hybrid Analytical-Neural Inverse Kinematics for Whole-Body Mesh RecoveryabstractRecovering whole-body mesh by inferring the abstract pose and shape parameters from visual content can obtain 3D bodies with realistic structures. However, the inferring process is highly non-linear and suffers from image-mesh misalignment, resulting in inaccurate reconstruction. In contrast, 3D keypoint estimation methods utilize the volumetric representation to achieve pixel-level accuracy but may predict unrealistic body structures. To address these issues, this paper presents a novel hybrid inverse kinematics solution, HybrIK, that integrates the merits of 3D keypoint estimation and body mesh recovery in a unified framework. HybrIK directly transforms accurate 3D joints to body-part rotations via twist-and-swing decomposition. The swing rotations are analytically solved with 3D joints, while the twist rotations are derived from visual cues through neural networks. To capture comprehensive whole-body details, we further develop a holistic framework, HybrIK-X, which enhances HybrIK with articulated hands and an expressive face. HybrIK-X is fast and accurate by solving the whole-body pose with a one-stage model. Experiments demonstrate that HybrIK and HybrIK-X preserve both the accuracy of 3D joints and the realistic structure of the parametric human model, leading to pixel-aligned whole-body mesh recovery. The proposed method significantly surpasses the state-of-the-art methods on various benchmarks for body-only, hand-only, and whole-body scenarios. Siyuan Bian, Chao Xu 0023, Zhicun Chen, Lixin Yang 0001, Cewu Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Multi-View Hand Reconstruction With a Point-Embedded TransformerabstractThis work introduces a novel and generalizable multi-view Hand Mesh Reconstruction (HMR) model, named POEM, designed for practical use in real-world hand motion capture scenarios. The advances of the POEM model consist of two main aspects. First, concerning the modeling of the problem, we propose embedding a static basis point within the multi-view stereo space. A point represents a natural form of 3D information and serves as an ideal medium for fusing features across different views, given its varied projections across these views. Consequently, our method harnesses a simple yet effective idea: a complex 3D hand mesh can be represented by a set of 3D basis points that 1) are embedded in the multi-view stereo, 2) carry features from the multi-view images, and 3) encompass the hand in it. The second advance lies in the training strategy. We utilize a combination of five large-scale multi-view datasets and employ randomization in the number, order, and poses of the cameras. By processing such a vast amount of data and a diverse array of camera configurations, our model demonstrates notable generalizability in the real-world applications. As a result, POEM presents a highly practical, plug-and-play solution that enables user-friendly, cost-effective multi-view motion capture for both left and right hands. Lixin Yang 0001, Licheng Zhong, Pengxiang Zhu, Xinyu Zhan 0001, Junxiao Kong, Jian Xu 0027, Cewu Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Color-NeuS: Reconstructing Neural Implicit Surfaces with ColorabstractThe reconstruction of object surfaces from multi-view images or monocular video is a fundamental issue in computer vision. However, much of the recent research concentrates on reconstructing geometry through implicit or explicit methods. In this paper, we shift our focus towards reconstructing mesh in conjunction with color. We remove the view-dependent color from neural volume rendering while retaining volume rendering performance through a relighting network. Mesh is extracted from the signed distance function (SDF) network for the surface, and color for each surface vertex is drawn from the global color network. To evaluate our approach, we conceived a in hand object scanning task featuring numerous occlusions and dramatic shifts in lighting conditions. We’ve gathered several videos for this task, and the results surpass those of any existing methods capable of reconstructing mesh alongside color. Additionally, our method’s performance was assessed using public datasets, including DTU, BlendedMVS, and OmniObject3D. The results indicated that our method performs well across all these datasets. Project page: colmar-zlicheng.github.io/color_neus. Licheng Zhong, Lixin Yang 0001, Kailin Li 0001, Haoyu Zhen, Cewu Lu |
3DV | 2 |
| 2024 | FAVOR: Full-Body AR-Driven Virtual Object Rearrangement Guided by Instruction TextabstractRearrangement operations form the crux of interactions between humans and their environment. The ability to generate natural, fluid sequences of this operation is of essential value in AR/VR and CG. Bridging a gap in the field, our study introduces FAVOR: a novel dataset for Full-body AR-driven Virtual Object Rearrangement that uniquely employs motion capture systems and AR eyeglasses. Comprising 3k diverse motion rearrangement sequences and 7.17 million interaction data frames, this dataset breaks new ground in research data. We also present a pipeline FAVORITE for producing digital human rearrangement motion sequences guided by instructions. Experimental results, both qualitative and quantitative, suggest that this dataset and pipeline deliver high-quality motion sequences. Our dataset, code, and appendix are available at https://kailinli.github.io/FAVOR. Kailin Li 0001, Lixin Yang 0001, Zenan Lin, Jian Xu 0027, Xinyu Zhan 0001, Yifei Zhao 0003, Pengxiang Zhu, Wenxiong Kang, Kejian Wu, Cewu Lu |
AAAI | 2 |
| 2024 | OakInk2 : A Dataset of Bimanual Hands-Object Manipulation in Complex Task CompletionabstractWe present Oakink2, a dataset of bimanual object manipulation tasks for complex daily activities. In pursuit of constructing the complex tasks into a structured representation, Oakink2 introduces three level of abstraction to organize the manipulation tasks: Affordance, Primitive Task, and Complex Task. OAKINK2 features on an object-centric perspective for decoding the complex tasks, treating them as a sequence of object affordance fulfillment. The first level, Affordance, outlines the functionalities that objects in the scene can afford, the second level, Primitive Task, describes the minimal interaction units that humans interact with the object to achieve its affordance, and the third level, Complex Task, illustrates how Primitive Tasks are composed and interdependent. Oakink2 dataset provides multi-view image streams and precise pose annotations for the human body, hands and various interacting objects. This extensive collection supports applications such as interaction reconstruction and motion synthesis. Based on the 3-level abstraction of Oakink2, we explore a task-oriented framework for Complex Task Completion (CTC). CTC aims to generate a sequence of bimanual manipulation to achieve task objectives. Within the CTC framework, we employ Large Language Models (LLMs) to decompose the complex task objectives into sequences of Primitive Tasks and have developed a Motion Fulfillment Model that generates bimanual hand motion for each Primitive Task. Oakink2 datasets and models are available at https://oakink.net/v2. Xinyu Zhan 0001, Lixin Yang 0001, Yifei Zhao 0003, Kangrui Mao, Hanlin Xu, Zenan Lin, Kailin Li 0001, Cewu Lu |
CVPR | 2 |
| 2024 | SemGrasp : Semantic Grasp Generation via Language Aligned Discretization
Kailin Li 0001, Jingbo Wang 0003, Lixin Yang 0001, Cewu Lu, Bo Dai 0002 |
ECCV (2) | 3 |
| 2024 | General Articulated Objects Manipulation in Real Images via Part-Aware Diffusion ProcessabstractArticulated object manipulation in real images is a fundamental step in computer and robotic vision tasks. Recently, several image editing methods based on diffusion models have been proposed to manipulate articulated objects according to text prompts. However, these methods often generate weird artifacts or even fail in real images. To this end, we introduce the Part-Aware Diffusion Model to approach the manipulation of articulated objects in real images. First, we develop Abstract 3D Models to represent and manipulate articulated objects efficiently. Then we propose dynamic feature maps to transfer the appearance of objects from input images to edited ones, meanwhile generating the novel-appearing parts reasonably. Extensive experiments are provided to illustrate the advanced manipulation capabilities of our method concerning state-of-the-art editing works. Additionally, we verify our method on 3D articulated object understanding for
embodied robot scenarios and the promising results prove that our method supports this task strongly. The project page is https://mvig-rhos.com/pa_diffusion. Yong-Lu Li 0001, Lixin Yang 0001, Cewu Lu |
NeurIPS | 3 |
| 2024 | Learning a Contact Potential Field for Modeling the Hand-Object InteractionabstractEstimating and synthesizing the hand's manipulation of objects is central to understanding human behaviour. To accurately model the interaction between the hand and object (referred to as the "hand-object"), we must not only focus on the pose of the hand and object, but also consider the contact between them. This contact provides valuable information for generating semantically and physically plausible grasps. In this paper, we propose an explicit contact representation called Contact Potential Field (CPF). In CPF, we model the contact between a pair of hand-object vertices as a spring-mass system. This system encodes the distance of the pair, as well as a likelihood of that contact being stable. Therefore, the system of multiple extended and compressed springs forms an elastic potential field with minimal energy at the optimal grasp position. We apply CPF to two relevant tasks, namely, hand-object pose estimation and grasping pose generation. Extensive experiments on the two challenging tasks and three commonly used datasets have demonstrated that our method can achieve state-of-the-art in several reconstruction metrics, allowing us to produce more physically plausible hand-object poses even when the ground-truth exhibits severe interpenetration or disjointedness. Lixin Yang 0001, Xinyu Zhan 0001, Kailin Li 0001, Cewu Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | POEM: Reconstructing Hand in a Point Embedded Multi-view StereoabstractEnable neural networks to capture 3D geometrical-aware features is essential in multi-view based vision tasks. Previous methods usually encode the 3D information of multi-view stereo into the 2D features. In contrast, we present a novel method, named POEM, that directly operates on the 3D POints Embedded in the Multi-view stereo for reconstructing hand mesh in it. Point is a natural form of 3D information and an ideal medium for fusing features across views, as it has different projections on different views. Our method is thus in light of a simple yet effective idea, that a complex 3D hand mesh can be represented by a set of 3D points that 1) are embedded in the multi-view stereo, 2) carry features from the multi-view images, and 3) encircle the hand. To leverage the power of points, we design two operations: point-based feature fusion and cross-set point attention mechanism. Evaluation on three challenging multi-view datasets shows that POEM outperforms the state-of-the-art in hand mesh reconstruction. Code and models are available for research at github.com/lixiny/POEM Lixin Yang 0001, Jian Xu 0027, Licheng Zhong, Xinyu Zhan 0001, Zhicheng Wang 0007, Kejian Wu, Cewu Lu |
CVPR | 1 |
| 2023 | Chord: Category-level Hand-held Object Reconstruction via Shape DeformationabstractIn daily life, humans utilize hands to manipulate objects. Modeling the shape of objects that are manipulated by the hand is essential for AI to comprehend daily tasks and to learn manipulation skills. However, previous approaches have encountered difficulties in reconstructing the precise shapes of hand-held objects, primarily owing to a deficiency in prior shape knowledge and inadequate data for training. As illustrated, given a particular type of tool, such as a mug, despite its infinite variations in shape and appearance, humans have a limited number of ‘effective’ modes and poses for its manipulation. This can be attributed to the fact that humans have mastered the shape prior of the ‘mug’ category, and can quickly establish the corresponding relations between different mug instances and the prior, such as where the rim and handle are located. In light of this, we propose a new method, Chord, for Category-level Hand-held Object Reconstruction via shape Deformation. Chord deforms a categorical shape prior for reconstructing the intra-class objects. To ensure accurate reconstruction, we empower Chord with three types of awareness: appearance, shape, and interacting pose. In addition, we have constructed a new dataset, Comic, of category-level hand-object interaction. Comic contains a rich array of object instances, materials, hand interactions, and viewing directions. Extensive evaluation shows that Chord outperforms state-of-the-art approaches in both quantitative and qualitative measures. Code, model, and datasets are available at https://kailinli.github.io/CHORD Kailin Li 0001, Lixin Yang 0001, Haoyu Zhen, Zenan Lin, Xinyu Zhan 0001, Licheng Zhong, Jian Xu 0027, Kejian Wu, Cewu Lu |
ICCV | 2 |
| 2022 | ArtiBoost: Boosting Articulated 3D Hand-Object Pose Estimation via Online Exploration and SynthesisabstractEstimating the articulated 3D hand-object pose from a single RGB image is a highly ambiguous and challenging problem, requiring large-scale datasets that contain diverse hand poses, object types, and camera viewpoints. Most real-world datasets lack these diversities. In contrast, data synthesis can easily ensure those diversities separately. However, constructing both valid and diverse hand-object interactions and efficiently learning from the vast synthetic data is still challenging. To address the above issues, we propose ArtiBoost, a lightweight online data enhancement method. ArtiBoost can cover diverse hand-object poses and camera viewpoints through sampling in a Composited hand-object Configuration and View-point space (CCV-space) and can adaptively enrich the current hard-discernable items by loss-feedback and sample re-weighting. ArtiBoost alternatively performs data exploration and synthesis within a learning pipeline, and those synthetic data are blended into real-world source data for training. We apply ArtiBoost on a simple learning baseline network and witness the performance boost on several hand-object benchmarks. Our models and code are available at https://github.com/lixiny/ArtiBoost. Lixin Yang 0001, Kailin Li 0001, Xinyu Zhan 0001, Cewu Lu |
CVPR | 1 |
| 2022 | OakInk: A Large-scale Knowledge Repository for Understanding Hand-Object InteractionabstractLearning how humans manipulate objects requires machines to acquire knowledge from two perspectives: one for understanding object affordances and the other for learning human's interactions based on the affordances. Even though these two knowledge bases are crucial, we find that current databases lack a comprehensive awareness of them. In this work, we propose a multi-modal and rich-annotated knowledge repository, OakInk, for visual and cognitive understanding of hand-object interactions. We start to collect 1,800 common household objects and annotate their affordances to construct the first knowledge base: Oak. Given the affordance, we record rich human interactions with 100 selected objects in Oak. Finally, we transfer the interactions on the 100 recorded objects to their virtual counterparts through a novel method: Tink. The recorded and transferred hand-object interactions constitute the second knowledge base: Ink. As a result, OakInk contains 50,000 distinct affordance-aware and intent-oriented hand-object interactions. We benchmark OakInk on pose estimation and grasp generation tasks. Moreover, we propose two practical applications of OakInk: intent-based interaction generation and handover generation. Our dataset and source code are publicly available at www.oakink.net. Lixin Yang 0001, Kailin Li 0001, Xinyu Zhan 0001, Fei Wu 0001, Anran Xu 0003, Liu Liu 0012, Cewu Lu |
CVPR | 1 |
| 2022 | DART: Articulated Hand Model with Diverse Accessories and Rich TexturesabstractHand, the bearer of human productivity and intelligence, is receiving much attention due to the recent fever of digital twins. Among different hand morphable models, MANO has been widely used in vision and graphics community. However, MANO disregards textures and accessories, which largely limits its power to synthesize photorealistic hand data. In this paper, we extend MANO with Diverse Accessories and Rich Textures, namely DART. DART is composed of 50 daily 3D accessories which varies in appearance and shape, and 325 hand-crafted 2D texture maps covers different kinds of blemishes or make-ups. Unity GUI is also provided to generate synthetic hand data with user-defined settings, e.g., pose, camera, background, lighting, textures, and accessories. Finally, we release DARTset, which contains large-scale (800K), high-fidelity synthetic hand images, paired with perfect-aligned 3D labels. Experiments demonstrate its superiority in diversity. As a complement to existing hand datasets, DARTset boosts the generalization in both hand pose estimation and mesh recovery tasks. Raw ingredients (textures, accessories), Unity GUI, source code and DARTset are publicly available at dart2022.github.io. Daiheng Gao, Yuliang Xiu, Kailin Li 0001, Lixin Yang 0001, Feng Wang 0072, Peng Zhang 0080, Bang Zhang, Cewu Lu |
NeurIPS | 4 |
| 2021 | HandTailor: Towards High-Precision Monocular 3D Hand Recovery
Lixin Yang 0001, Sucheng Qian, Chongzhao Mao, Cewu Lu |
BMVC | 3 |
| 2021 | HybrIK: A Hybrid Analytical-Neural Inverse Kinematics Solution for 3D Human Pose and Shape EstimationabstractModel-based 3D pose and shape estimation methods reconstruct a full 3D mesh for the human body by estimating several parameters. However, learning the abstract parameters is a highly non-linear process and suffers from image-model misalignment, leading to mediocre model performance. In contrast, 3D keypoint estimation methods combine deep CNN network with the volumetric representation to achieve pixel-level localization accuracy but may predict unrealistic body structure. In this paper, we address the above issues by bridging the gap between body mesh estimation and 3D keypoint estimation. We propose a novel hybrid inverse kinematics solution (HybrIK). HybrIK directly transforms accurate 3D joints to relative body-part rotations for 3D body mesh reconstruction, via the twistand-swing decomposition. The swing rotation is analytically solved with 3D joints, and the twist rotation is derived from the visual cues through the neural network. We show that HybrIK preserves both the accuracy of 3D pose and the realistic body structure of the parametric human model, leading to a pixel-aligned 3D body mesh and a more accurate 3D pose than the pure 3D keypoint estimation methods. Without bells and whistles, the proposed method surpasses the state-of-the-art methods by a large margin on various 3D human pose and shape benchmarks. As an illustrative example, HybrIK outperforms all the previous methods by 13.2 mm MPJPE and 21.9 mm PVE on 3DPW dataset. Our code is available at https://github.com/Jeff-sjtu/HybrIK. Chao Xu 0023, Zhicun Chen, Siyuan Bian, Lixin Yang 0001, Cewu Lu |
CVPR | 5 |
| 2021 | CPF: Learning a Contact Potential Field to Model the Hand-Object InteractionabstractModeling the hand-object (HO) interaction not only requires estimation of the HO pose, but also pays attention to the contact due to their interaction. Significant progress has been made in estimating hand and object separately with deep learning methods, simultaneous HO pose estimation and contact modeling has not yet been fully explored. In this paper, we present an explicit contact representation namely Contact Potential Field (CPF), and a learning-fitting hybrid framework namely MIHO to Modeling the Interaction of Hand and Object. In CPF, we treat each contacting HO vertex pair as a spring-mass system. Hence the whole system forms a potential field with minimal elastic energy at the grasp position. Extensive experiments on the two commonly used benchmarks have demonstrated that our method can achieve state-of-the-art in several reconstruction metrics, and allow us to produce more physically plausible HO pose even when the ground-truth exhibits severe interpenetration or disjointedness. Our code is available at https://github.com/lixiny/CPF. Lixin Yang 0001, Xinyu Zhan 0001, Kailin Li 0001, Cewu Lu |
ICCV | 1 |
| 2020 | BiHand: Recovering Hand Mesh with Multi-stage Bisected Hourglass Networks
Lixin Yang 0001, Jiasen Li, Yiqun Diao, Cewu Lu |
BMVC | 1 |