EDBT 2026 Demo / reviewers in the wild / expert
Xinyu Yi
dblp:292/4128
· DBLP profile ↗
15ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LRC-PSO: LLM-Driven Reasoning and Closed-Loop Optimization for High-Dimensional VLSI Placement
Xinyu Yi, Feng-Feng Wei |
ICIC (13) | 1 |
| 2025 | MagShield: Towards Better Robustness in Sparse Inertial Motion Capture Under Magnetic DisturbancesabstractThis paper proposes a novel method called MagShield, designed to address the issue of magnetic interference in sparse inertial motion capture (MoCap) systems. Existing Inertial Measurement Unit (IMU) systems are prone to orientation estimation errors in magnetically disturbed environments, limiting their practical application in real-world scenarios. To address this problem, MagShield employs a "detect-then-correct" strategy, first detecting magnetic disturbances through multi-IMU joint analysis, and then correcting orientation errors using human motion priors. MagShield can be integrated with most existing sparse inertial MoCap systems, improving their performance in magnetically disturbed environments. Experimental results demonstrate that MagShield significantly enhances the accuracy of motion capture under magnetic interference and exhibits good compatibility across different sparse inertial MoCap systems. Yunzhe Shao, Xinyu Yi, Shihui Guo, Jun-Hai Yong, Feng Xu 0005 |
ICCV | 2 |
| 2025 | BaroPoser: Real-time Human Motion Tracking from IMUs and Barometers in Everyday Devices
Xinyu Yi, Feng Xu 0005 |
UIST | 2 |
| 2025 | Improving Global Motion Estimation in Sparse IMU-based Motion Capture with PhysicsabstractBy learning human motion priors, motion capture can be achieved by 6 inertial measurement units (IMUs) in recent years with the development of deep learning techniques, even though the sensor inputs are sparse and noisy. However, human global motions are still challenging to be reconstructed by IMUs. This paper aims to solve this problem by involving physics. It proposes a physical optimization scheme based on multiple contacts to enable physically plausible translation estimation in the full 3D space where the z-directional motion is usually challenging for previous works. It also considers gravity in local pose estimation which well constrains human global orientations and refines local pose estimation in a joint estimation manner. Experiments demonstrate that our method achieves more accurate motion capture for both local poses and global motions. Furthermore, by deeply integrating physics, we can also estimate 3D contact, contact forces, joint torques, and interacting proxy surfaces. Code is available at https://xinyu-yi.github.io/GlobalPose/. Xinyu Yi, Shaohua Pan 0002, Feng Xu 0005 |
ACM Trans. Graph. | 1 |
| 2025 | Shape-aware Inertial Poser: Motion Tracking for Humans with Diverse Shapes Using Sparse Inertial SensorsabstractHuman motion capture with sparse inertial sensors has gained significant attention recently. However, existing methods almost exclusively rely on a template adult body shape to model the training data, which poses challenges when generalizing to individuals with largely different body shapes (such as a child). This is primarily due to the variation in IMU-measured acceleration caused by changes in body shape. To fill this gap, we propose Shape-aware Inertial Poser (SAIP), the first solution considering body shape differences in sparse inertial-based motion capture. Specifically, we decompose the sensor measurements related to shape and pose in order to effectively model their joint correlations. Firstly, we train a regression model to transfer the IMU-measured accelerations of a real body to match the template adult body model, compensating for the shape-related sensor measurements. Then, we can easily follow the state-of-the-art methods to estimate the full body motions of the template-shaped body. Finally, we utilize a second regression model to map the joint velocities back to the real body, combined with a shape-aware physical optimization strategy to calculate global motions on the subject. Furthermore, our method relies on body shape awareness, introducing the first inertial shape estimation scheme. This is accomplished by modeling the shape-conditioned IMU-pose correlation using an MLP-based network. To validate the effectiveness of SAIP, we also present the first IMU motion capture dataset containing individuals of different body sizes. This dataset features 10 children and 10 adults, with heights ranging from 110 cm to 190 cm, and a total of 400 minutes of paired IMU-Motion samples. Extensive experimental results demonstrate that SAIP can effectively handle motion capture tasks for diverse body shapes. The code and dataset are available at https://github.com/yinlu5942/SAIP . Ziying Shi, Yinghao Wu, Xinyu Yi, Feng Xu 0005, Shihui Guo |
ACM Trans. Graph. | 4 |
| 2025 | Transformer IMU Calibrator: Dynamic On-body IMU Calibration for Inertial Motion CaptureabstractIn this paper, we propose a novel dynamic calibration method for sparse inertial motion capture systems, which is the first to break the restrictive absolute static assumption in IMU calibration, i.e., the coordinate drift R G′ G and measurement offset R BS remain constant during the entire motion, thereby significantly expanding their application scenarios. Specifically, we achieve real-time estimation of R G′ G and R BS under two relaxed assumptions: i) the matrices change negligibly in a short time window; ii) the human movements/IMU readings are diverse in such a time window. Intuitively, the first assumption reduces the number of candidate matrices, and the second assumption provides diverse constraints, which greatly reduces the solution space and allows for accurate estimation of R G′ G and R BS from a short history of IMU readings in real time. To achieve this, we created synthetic datasets of paired R G′ G , R BS matrices and IMU readings, and learned their mappings using a Transformer-based model. We also designed a calibration trigger based on the diversity of IMU readings to ensure that assumption ii) is met before applying our method. To our knowledge, we are the first to achieve implicit IMU calibration (i.e., seamlessly putting IMUs into use without the need for an explicit calibration process), as well as the first to enable long-term and accurate motion capture using sparse IMUs. The code and dataset are available at https://github.com/ZuoCX1996/TIC. Chengxu Zuo, Xiangren Shi, Xinyu Yi, Feng Xu 0005, Shihui Guo, Yipeng Qin |
ACM Trans. Graph. | 7 |
| 2025 | DiffCap: Diffusion-Based Real-Time Human Motion Capture Using Sparse IMUs and a Monocular CameraabstractCombining sparse IMUs and a monocular camera is a new promising setting to perform real-time human motion capture. This paper proposes a diffusion-based solution to learn human motion priors and fuse the two modalities of signals together seamlessly in a unified framework. By delicately considering the characteristics of the two signals, the sequential visual information is considered as a whole and transformed into a condition embedding, while the inertial measurement is concatenated with the noisy body pose frame by frame to construct a sequential input for the diffusion model. Firstly, we observe that the visual information may be unavailable in some frames due to occlusions or subjects moving out of the camera view. Thus incorporating the sequential visual features as a whole to get a single feature embedding is robust to the occasional degenerations of visual information in those frames. On the other hand, the IMU measurements are robust to occlusions and always stable when signal transmission has no problem. So incorporating them frame-wisely could better explore the temporal information for the system. Experiments have demonstrated the effectiveness of the system design and its state-of-the-art performance in pose estimation compared with the previous works. The code will be released. Shaohua Pan 0002, Xinyu Yi, Yan Zhou 0003, Weihua Jian, Yuan Zhang 0020, Pengfei Wan 0001, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Loose Inertial Poser: Motion Capture with IMU-attached Loose-Wear JacketabstractExisting wearable motion capture methods typically demand tight on-body fixation (often using straps) for reliable sensing, limiting their application in everyday life. In this paper, we introduce Loose Inertial Poser, a novel motion capture solution with high wearing comfortableness, by integrating four Inertial Measurement Units (IMUs) into a loose-wear jacket. Specifically, we address the challenge of scarce loose-wear IMU training data by proposing a Secondary Motion AutoEncoder (SeMo-AE) that learns to model and synthesize the effects of secondary motion between the skin and loose clothing on IMU data. SeMo-AE is leveraged to generate a diverse synthetic dataset of loose-wear IMU data to augment training for the pose estimation network and significantly improve its accuracy. For validation, we collected a dataset with various subjects and 2 wearing styles (zipped and unzipped). Experimental results demonstrate that our approach maintains high-quality real-time posture estimation even in loose-wear scenarios. Our dataset and code are available at: https://github.com/ZuoCX1966/Loose-Inertial-Poser Chengxu Zuo, Lishuang Zhan, Shihui Guo, Xinyu Yi, Feng Xu 0005, Yipeng Qin |
CVPR | 5 |
| 2023 | LI-FPN: Depression and Anxiety Detection from Learning and ImitationabstractWith the rise in societal pressures, depression and anxiety have increasingly become prominent mental health conditions impacting people’s lives. To enhance the efficacy of automatic detection for these disorders, we have developed an experimental framework called the Voluntary Facial Expression Mimicry(VFEM). This framework led to the creation of the VFEM Dataset, which supports related research endeavors. Subsequently, we introduce the LI-FPN designed specifically for the automatic identification of depression and anxiety disorders. The LI-FPN comprises two core components: the Learning and Imitation Module(LIM) and the Spatio-temporal Feature Pyramid Network(STFPN). Within the LIM, we leverage sequence features to facilitate comprehensive feature extraction through learning and imitation steps. The STFPN is designed to focus on outliers in multi-scale features for further screening. Compared with traditional attention methods, LI-FPN is more suitable for processing sequence data features and small sample datasets. Upon training using the VFEM Dataset, the LI-FPN achieves impressive accuracies: 0.850 for depression detection, 0.835 for anxiety detection, and 0.786 for co-occurrence detection of depression and anxiety. Meanwhile, LI-FPN also achieves SOAT results on AVEC2014 dataset. The source code for LI-FPN is accessible at https://github.com/muzixingyun/LI-FPN Xingyun Li, Xinyu Yi, Yunshao Zheng, Yanhong Yu, Qingxiang Wang |
BIBM | 3 |
| 2023 | Fusing Monocular Images and Sparse IMU Signals for Real-time Human Motion CaptureabstractEither RGB images or inertial signals have been used for the task of motion capture (mocap), but combining them together is a new and interesting topic. We believe that the combination is complementary and able to solve the inherent difficulties of using one modality input, including occlusions, extreme lighting/texture, and out-of-view for visual mocap and global drifts for inertial mocap. To this end, we propose a method that fuses monocular images and sparse IMUs for real-time human motion capture. Our method contains a dual coordinate strategy to fully explore the IMU signals with different goals in motion capture. To be specific, besides one branch transforming the IMU signals to the camera coordinate system to combine with the image information, there is another branch to learn from the IMU signals in the body root coordinate system to better estimate body poses. Furthermore, a hidden state feedback mechanism is proposed for both two branches to compensate for their own drawbacks in extreme input cases. Thus our method can easily switch between the two kinds of signals or combine them in different cases to achieve a robust mocap. Quantitative and qualitative results demonstrate that by delicately designing the fusion method, our technique significantly outperforms the state-of-the-art vision, IMU, and combined methods on both global orientation and local pose estimation. Our codes are available for research at https://shaohua-pan.github.io/robustcap-page/. Shaohua Pan 0002, Xinyu Yi, Xingkang Zhou, Jijunnan Li, Feng Xu 0005 |
SIGGRAPH Asia | 3 |
| 2023 | Digital rights management scheme based on redactable blockchain and perceptual hash
Xinyu Yi, Yuping Zhou, Yuqian Lin, Ben Xie, Chenye Wang |
Peer Peer Netw. Appl. | 1 |
| 2023 | EgoLocate: Real-time Motion Capture, Localization, and Mapping with Sparse Body-mounted SensorsabstractHuman and environment sensing are two important topics in Computer Vision and Graphics. Human motion is often captured by inertial sensors, while the environment is mostly reconstructed using cameras. We integrate the two techniques together in EgoLocate, a system that simultaneously performs human motion capture (mocap), localization, and mapping in real time from sparse body-mounted sensors, including 6 inertial measurement units (IMUs) and a monocular phone camera. On one hand, inertial mocap suffers from large translation drift due to the lack of the global positioning signal. EgoLo-cate leverages image-based simultaneous localization and mapping (SLAM) techniquesto locate the human in the reconstructed scene. Onthe other hand, SLAM often fails when the visual feature is poor. EgoLocate involves inertial mocap to provide a strong prior for the camera motion. Experiments show that localization, a key challenge for both two fields, is largely improved by our technique, compared with the state of the art of the two fields. Our codes are available for research at https://xinyu-yi.github.io/EgoLocate/. Xinyu Yi, Yuxiao Zhou 0001, Marc Habermann, Vladislav Golyanik, Shaohua Pan 0002, Christian Theobalt, Feng Xu 0005 |
ACM Trans. Graph. | 1 |
| 2022 | Physical Inertial Poser (PIP): Physics-aware Real-time Human Motion Tracking from Sparse Inertial SensorsabstractMotion capture from sparse inertial sensors has shown great potential compared to image-based approaches since occlusions do not lead to a reduced tracking quality and the recording space is not restricted to be within the viewing frustum of the camera. However, capturing the motion and global position only from a sparse set of inertial sensors is inherently ambiguous and challenging. In consequence, recent state-of-the-art methods can barely handle very long period motions, and unrealistic artifacts are common due to the unawareness of physical constraints. To this end, we present the first method which combines a neural kinematics estimator and a physics-aware motion optimizer to track body motions with only 6 inertial sensors. The kinematics module first regresses the motion status as a reference, and then the physics module refines the motion to satisfy the physical constraints. Experiments demonstrate a clear improvement over the state of the art in terms of capture accuracy, temporal stability, and physical correctness. Xinyu Yi, Yuxiao Zhou 0001, Marc Habermann, Soshi Shimada, Vladislav Golyanik, Christian Theobalt, Feng Xu 0005 |
CVPR | 1 |
| 2022 | Physical Interaction: Reconstructing Hand-object Interactions with PhysicsabstractSingle view-based reconstruction of hand-object interaction is challenging due to the severe observation missing caused by occlusions. This paper proposes a physics-based method to better solve the ambiguities in the reconstruction. It first proposes a force-based dynamic model of the in-hand object, which not only recovers the unobserved contacts but also solves for plausible contact forces. Next, a confidence-based slide prevention scheme is proposed, which combines both the kinematic confidences and the contact forces to jointly model static and sliding contact motion. Qualitative and quantitative experiments show that the proposed technique reconstructs both physically plausible and more accurate hand-object interaction and estimates plausible contact forces in real-time with a single RGBD sensor. Xinyu Yi, Hao Zhang 0042, Jun-Hai Yong, Feng Xu 0005 |
SIGGRAPH Asia | 2 |
| 2021 | TransPose: real-time 3D human translation and pose estimation with six inertial sensorsabstractMotion capture is facing some new possibilities brought by the inertial sensing technologies which do not suffer from occlusion or wide-range recordings as vision-based solutions do. However, as the recorded signals are sparse and quite noisy, online performance and global translation estimation turn out to be two key difficulties. In this paper, we present TransPose, a DNN-based approach to perform full motion capture (with both global translations and body poses) from only 6 Inertial Measurement Units (IMUs) at over 90 fps. For body pose estimation, we propose a multi-stage network that estimates leaf-to-full joint positions as intermediate results. This design makes the pose estimation much easier, and thus achieves both better accuracy and lower computation cost. For global translation estimation, we propose a supporting-foot-based method and an RNN-based method to robustly solve for the global translations with a confidence-based fusion technique. Quantitative and qualitative comparisons show that our method outperforms the state-of-the-art learning- and optimization-based methods with a large margin in both accuracy and efficiency. As a purely inertial sensor-based approach, our method is not limited by environmental settings (e.g., fixed cameras), making the capture free from common difficulties such as wide-range motion space and strong occlusion. Xinyu Yi, Yuxiao Zhou 0001, Feng Xu 0005 |
ACM Trans. Graph. | 1 |