VLDB 2026 Research / reviewers in the wild / expert
Songpengcheng Xia
dblp:274/2787
· DBLP profile ↗
17ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-9452-7967ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Computer networks · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RadarLLM: Empowering Large Language Models to Understand Human Motion from Millimeter-wave Point Cloud SequenceabstractMillimeter-wave radar offers a privacy-preserving and environment-robust alternative to vision-based sensing, enabling human motion analysis in challenging conditions such as low light, occlusions, rain, or smoke. However, its sparse point clouds pose significant challenges for semantic understanding. We present RadarLLM, the first framework that leverages large language models (LLMs) for human motion understanding from radar signals. RadarLLM introduces two key innovations: (1) a motion-guided radar tokenizer based on our Aggregate VQ-VAE architecture, integrating deformable body templates and masked trajectory modeling to convert spatial-temporal radar sequences into compact semantic tokens; and (2) a radar-aware language model that establishes cross-modal alignment between radar and text in a shared embedding space. To overcome the scarcity of paired radar-text data, we generate a realistic radar-text dataset from motion-text datasets with a physics-aware synthesis pipeline. Extensive experiments on both synthetic and real-world benchmarks show that RadarLLM achieves state-of-the-art performance, enabling robust and interpretable motion understanding under privacy and visibility constraints, even in adverse environments. Zengyuan Lai, Songpengcheng Xia, Lizhou Lin, Renwen Wang, Jianran Liu, Qi Wu 0007, Ling Pei |
AAAI | 3 |
| 2025 | EnvPoser: Environment-aware Realistic Human Motion Estimation from Sparse Observations with Uncertainty ModelingabstractEstimating full-body motion using the tracking signals of head and hands from VR devices holds great potential for various applications. However, the sparsity and unique distribution of observations present a significant challenge, resulting in an ill-posed problem with multiple feasible solutions (i.e., hypotheses). This amplifies uncertainty and ambiguity in full-body motion estimation, especially for the lower-body joints. Therefore, we propose a new method, EnvPoser, that employs a two-stage framework to perform full-body motion estimation using sparse tracking signals and pre-scanned environment from VR devices. EnvPoser models the multi-hypothesis nature of human motion through an uncertainty-aware estimation module in the first stage. In the second stage, we refine these multi-hypothesis estimates by integrating semantic and geometric environmental constraints, ensuring that the final motion estimation aligns realistically with both the environmental context and physical interactions. Qualitative and quantitative experiments on two public datasets demonstrate that our method achieves state-of-the-art performance, highlighting significant improvements in human motion estimation within motion-environment interaction scenarios. Project page: https://xspc.github.io/EnvPoser/. Songpengcheng Xia, Zhuo Su 0006, Xiaozheng Zheng, Guidong Wang, Qi Wu 0007, Ling Pei |
CVPR | 1 |
| 2025 | Suite-IN: Aggregating Motion Features from Apple Suite for Robust Inertial NavigationabstractWith the rapid development of wearable technology, devices like smartphones, smartwatches, and headphones equipped with IMUs have become essential for applications such as pedestrian positioning. However, traditional pedestrian dead reckoning (PDR) methods struggle with diverse motion patterns, while recent data-driven approaches, though improving accuracy, often lack robustness due to reliance on a single device. In our work, we attempt to enhance the positioning performance using the low-cost commodity IMUs embedded in the wearable devices. We propose a multi-device deep learning framework named Suite-IN, aggregating motion data from Apple Suite for inertial navigation. Motion data captured by sensors on different body parts contains both local and global motion information, making it essential to reduce the negative effects of localized movements and extract global motion representations from multiple devices. Our model innovatively introduces a contrastive learning module to disentangle motionshared and motion-private latent representations, enhancing positioning accuracy. We validate our method on a self-collected dataset consisting of Apple Suite: iPhone, Apple Watch and Airpods, which supports a variety of movement patterns and flexible device configurations. Experimental results demonstrate that our approach outperforms state-of-the-art models while maintaining robustness across diverse sensor configurations. Songpengcheng Xia, Junyuan Deng, Zengyuan Lai, Qi Wu 0007, Ling Pei |
ICRA | 2 |
| 2025 | mmDEAR: mmWave Point Cloud Density Enhancement for Accurate Human Body ReconstructionabstractMillimeter-wave (mmWave) radar offers robust sensing capabilities in diverse environments, making it a highly promising solution for human body reconstruction due to its privacy-friendly and non-intrusive nature. However, the significant sparsity of mm Wave point clouds limits the estimation accuracy. To overcome this challenge, we propose a two-stage deep learning framework that enhances mm Wave point clouds and improves human body reconstruction accuracy. Our method includes a mm Wave point cloud enhancement module that densifies the raw data by leveraging temporal features and a multi-stage completion network, followed by a 2D-3D fusion module that extracts both 2D and 3D motion features to refine SMPL parameters. The mm Wave point cloud enhancement module learns the detailed shape and posture information from 2D human masks in single-view images. However, image-based supervision is involved only during the training phase, and the inference relies solely on sparse point clouds to maintain privacy. Experiments on multiple datasets demonstrate that our approach outperforms state-of-the-art methods, with the enhanced point clouds further improving performance when integrated into existing models. Songpengcheng Xia, Zengyuan Lai, Qi Wu 0007, Wenxian Yu, Ling Pei |
ICRA | 2 |
| 2025 | 360Recon: An Accurate Reconstruction Method based on Depth Fusion from 360 ImagesabstractAccurate 3D reconstruction is crucial for AR and VR applications. Compared with traditional pinhole camera-based methods, 360° image-based reconstruction can achieve higher precision with fewer input images, making it especially effective in low-texture environments. However, the severe distortion resulting from the wide field of view complicates feature extraction and matching, leading to geometric inconsistencies in multi-view reconstruction. To address these challenges, we propose 360Recon, a novel multi-view stereo (MVS) algorithm specifically designed for equirectangular projection (ERP) images. With the proposed spherical feature extraction module mitigating distortion, 360Recon integrates a 3D cost volume with multi-scale ERP features to deliver high-precision scene reconstruction while preserving local geometric consistency. Experimental results demonstrate that 360Recon outperforms existing methods in terms of accuracy, computational efficiency, and generalization capability. The source code will be released at https://github.com/LeonATP/360Recon. Zhongmiao Yan, Qi Wu 0007, Songpengcheng Xia, Junyuan Deng, Xiang Mu, Renbiao Jin, Changchun Ye, Ling Pei |
IROS | 3 |
| 2025 | SMART: Scene-Motion-Aware Human Action Recognition Framework for Mental Disorder GroupabstractPatients with mental disorders often exhibit risky abnormal actions, such as climbing walls or hitting windows, necessitating intelligent video behavior monitoring for smart healthcare with the rising Internet of Things (IoT) technology. However, the development of vision-based human action recognition (HAR) for these actions is hindered by the lack of specialized algorithms and datasets. In this article, we innovatively propose to build a vision-based HAR dataset, including abnormal actions often occurring in the mental disorder group and then introduce a novel scene-motion-aware action recognition technology framework, named SMART, consisting of two technical modules. First, we propose a scene perception module to extract human motion trajectory and human-scene interaction features, which introduces additional scene information for a supplementary semantic representation of the above actions. Second, the multistage fusion module fuses the skeleton motion, motion trajectory, and human-scene interaction features, enhancing the semantic association between the skeleton motion and the above supplementary representation, thus generating a comprehensive representation with both human motion and scene information. The effectiveness of our proposed method has been validated on our self-collected HAR dataset (MentalHAD), achieving 94.9% and 93.1% accuracy in un-seen subjects and scenes and outperforming state-of-the-art approaches by 6.5% and 13.2%, respectively. The demonstrated subject- and scene- generalizability makes it possible for SMART’s migration to practical deployment in smart healthcare systems for mental disorder patients in medical settings. The code and dataset will be released publicly for further research:https://github.com/Inowlzy/SMART.git. Zengyuan Lai, Songpengcheng Xia, Qi Wu 0007, Wenxian Yu, Ling Pei |
IEEE Internet Things J. | 3 |
| 2025 | Suite-IN++: A FlexiWear BodyNet Integrating Global and Local Motion Features From Apple Suite for Robust Inertial NavigationabstractThe proliferation of wearable technology has established multi-device ecosystems comprising smartphones, smartwatches, and headphones as critical enablers for ubiquitous pedestrian localization. However, traditional pedestrian dead reckoning (PDR) struggles with diverse motion modes, while data-driven methods, despite improving accuracy, often lack robustness due to their reliance on a single-device setup. Therefore, a promising solution is to fully leverage existing wearable devices to form a flexiwear bodynet for robust and accurate pedestrian localization. This paper presents Suite-IN++, a deep learning framework for flexiwear bodynet-based pedestrian localization. Suite-IN++ integrates motion data from wearable devices on different body parts, using contrastive learning to separate global and local motion features. It fuses global features based on the data reliability of each device to capture overall motion trends and employs an attention mechanism to uncover cross-device correlations in local features, extracting motion details helpful for accurate localization. To evaluate our method, we construct a real-life flexiwear bodynet dataset, incorporating Apple Suite (iPhone, Apple Watch, and AirPods) across diverse walking modes and device configurations. Experimental results demonstrate that Suite-IN++ achieves superior localization accuracy and robustness, significantly outperforming state-of-the-art models in real-life pedestrian tracking scenarios. Songpengcheng Xia, Ling Pei |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Dynamic Inertial Poser (DynaIP): Part-Based Motion Dynamics Learning for Enhanced Human Pose Estimation with Sparse Inertial SensorsabstractThis paper introduces a novel human pose estimation approach using sparse inertial sensors, addressing the short-comings of previous methods reliant on synthetic data. It leverages a diverse array of real inertial motion capture data from different skeleton formats to improve motion di-versity and model generalization. This method features two innovative components: a pseudo-velocity regression model for dynamic motion capture with inertial sensors, and a part-based model dividing the body and sensor data into three regions, each focusing on their unique characteristics. The approach demonstrates superior performance over state-of-the-art models across five public datasets, notably reducing pose error by 19% on the DIP-IMU dataset, thus representing a significant improvement in inertial sensor-based human pose estimation. Our codes are available at https://github.com/dx118/dynaip Songpengcheng Xia, Qi Wu 0007, Ling Pei |
CVPR | 2 |
| 2024 | A Learning-Based Multi-Node Fusion Positioning Method Using Wearable Inertial SensorsabstractThis study presents a novel approach to enhance the accuracy and adaptability of pedestrian positioning by fusing data from multiple Inertial Measurement Units (IMUs) attached to the human body. Leveraging the temporal and spatial richness of IMU data, our proposed multi-node sensors fusion strategy accommodates multiple sensor placements and various motion modes, a challenge that has posed difficulties for previous methods. The methodology is based on data-driven principles, directly estimating position information from IMU data. The core of our method is the strategic allocation of weights to optimize the contribution of each sensor based on reliability and relevance. Two restraint strategies, Gumbel Softmax Resampling and Empirical Risk Minimization under Fairness Constraints, are introduced to enhance fusion accuracy and robustness. Experimental results show improvements in positioning accuracy and robust performance in various situations compared to existing methods. Our approach is expected to be applied in domains such as indoor navigation and motion analysis, highlighting the key role of wearable sensors in advancing positioning accuracy and addressing complex scenarios. Songpengcheng Xia, Ling Pei |
ICASSP | 2 |
| 2024 | mmBaT: A Multi-Task Framework for Mmwave-Based Human Body Reconstruction and Translation PredictionabstractHuman body reconstruction with Millimeter Wave (mmWave) radar point clouds has gained significant interest due to its ability to work in adverse environments and its capacity to mitigate privacy concerns associated with traditional camera-based solutions. Despite pioneering efforts in this field, two challenges persist. Firstly, raw point clouds contain massive noise points, usually caused by the ambient objects and multi-path effects of Radio Frequency (RF) signals. Recent approaches typically rely on prior knowledge or elaborate preprocessing methods, limiting their applicability. Secondly, even after noise removal, the sparse and inconsistent body-related points pose an obstacle to accurate human body reconstruction. To address these challenges, we introduce mmBaT, a novel multi-task deep learning framework that concurrently estimates the human body and predicts body translations in subsequent frames to extract body-related point clouds. Our method is evaluated on two public datasets that are collected with different radar devices and noise levels. A comprehensive comparison against other state-of-the-art methods demonstrates that our method has a superior reconstruction performance and generalization ability from noisy raw data, even when compared to methods provided with body-related point clouds. Songpengcheng Xia, Qi Wu 0007, Ling Pei |
ICASSP | 2 |
| 2024 | Thermal-NeRF: Neural Radiance Fields from an Infrared CameraabstractIn recent years, Neural Radiance Fields (NeRFs) have demonstrated significant potential in encoding highly-detailed 3D geometry and environmental appearance, positioning themselves as a promising alternative to traditional explicit representation for 3D scene reconstruction. However, the predominant reliance on RGB imaging presupposes ideal lighting conditions—a premise frequently unmet in robotic applications plagued by poor lighting or visual obstructions. This limitation overlooks the capabilities of infrared (IR) cameras, which excel in low-light detection and present a robust alternative under such adverse scenarios. To tackle these issues, we introduce Thermal-NeRF, the first method that estimates a volumetric scene representation in the form of a NeRF solely from IR imaging. By leveraging a thermal mapping and structural thermal constraint derived from the thermal characteristics of IR imaging, our method showcases unparalleled proficiency in recovering NeRFs in visually degraded scenes where RGB-based methods fall short. We conduct extensive experiments to demonstrate that Thermal-NeRF can achieve superior quality compared to existing methods. Furthermore, we contribute a dataset for IR-based NeRF applications, paving the way for future research in IR NeRF reconstruction, see https://github.com/Cerf-Volant425/Thermal-NeRF. Tianxiang Ye, Qi Wu 0007, Junyuan Deng, Liu Liu 0012, Songpengcheng Xia, Wenxian Yu, Ling Pei |
IROS | 6 |
| 2024 | Timestamp-Supervised Wearable-Based Activity Segmentation and Recognition With Contrastive Learning and Order-Preserving Optimal TransportabstractHuman activity recognition (HAR) with wearables is one of the serviceable technologies in ubiquitous and mobile computing applications. The sliding-window scheme is widely adopted while suffering from the multi-class windows problem. As a result, there is a growing focus on joint segmentation and recognition with deep-learning methods, aiming at simultaneously dealing with HAR and time-series segmentation issues. However, obtaining the full activity annotations of wearable data sequences is resource-intensive or time-consuming, while unsupervised methods yield poor performance. To address these challenges, we propose a novel method for joint activity segmentation and recognition with timestamp supervision, in which only a single annotated sample is needed in each activity segment. However, the limited information of sparse annotations exacerbates the gap between recognition and segmentation tasks, leading to sub-optimal model performance. Therefore, the prototypes are estimated by class-activation maps to form a sample-to-prototype contrast module for well-structured embeddings. Moreover, with the optimal transport theory, our approach generates the sample-level pseudo-labels that take advantage of unlabeled data between timestamp annotations for further performance improvement. Comprehensive experiments on four public HAR datasets demonstrate that our model trained with timestamp supervision is superior to the state-of-the-art weakly-supervised methods and achieves comparable performance to the fully-supervised approaches. Songpengcheng Xia, Ling Pei, Wenxian Yu, Robert C. Qiu |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | NeRF-LOAM: Neural Implicit Representation for Large-Scale Incremental LiDAR Odometry and MappingabstractSimultaneously odometry and mapping using LiDAR data is an important task for mobile systems to achieve full autonomy in large-scale environments. However, most existing LiDAR-based methods prioritize tracking quality over reconstruction quality. Although the recently developed neural radiance fields (NeRF) have shown promising advances in implicit reconstruction for indoor environments, the problem of simultaneous odometry and mapping for large-scale scenarios using incremental LiDAR data remains unexplored. To bridge this gap, in this paper, we propose a novel NeRF-based LiDAR odometry and mapping approach, NeRF-LOAM, consisting of three modules neural odometry, neural mapping, and mesh reconstruction. All these modules utilize our proposed neural signed distance function, which separates LiDAR points into ground and non-ground points to reduce Z-axis drift, optimizes odometry and voxel embeddings concurrently, and in the end generates dense smooth mesh maps of the environment. Moreover, this joint optimization allows our NeRF-LOAM to be pre-trained free and exhibit strong generalization abilities when applied to different environments. Extensive evaluations on three publicly available datasets demonstrate that our approach achieves state-of-the-art odometry and mapping performance, as well as a strong generalization in large-scale environments utilizing LiDAR data. Furthermore, we perform multiple ablation studies to validate the effectiveness of our network design. The implementation of our approach will be made available at https://github.com/JunyuanDeng/NeRF-LOAM. Junyuan Deng, Qi Wu 0007, Xieyuanli Chen, Songpengcheng Xia, Wenxian Yu, Ling Pei |
ICCV | 4 |
| 2023 | A Boundary Consistency-Aware Multitask Learning Framework for Joint Activity Segmentation and Recognition With Wearable SensorsabstractWith the development of industrial and sensing technology, sensor-based activity recognition has become a promising technology for informatics applications. However, in a typical activity recognition procedure, sensory data segmentation, usually considered a preprocess with sliding windows, rarely has been investigated and significantly affected the recognition performance. In this article, we propose a novel deep-learning method to jointly segment and recognize activities with wearable sensors. Our contributions are three-fold: First, we introduce a multistage temporal convolutional network for sample-level activity prediction to overcome the multiclass windows problem. Second, for alleviating oversegmentation errors, our model forms a multitask learning framework with a boundary prediction module to adjust the entire model’s gradients. Third, we innovatively propose a boundary consistency loss to enforce the consistency of the activity and boundary prediction. Our method shows impressive performance on three public datasets, especially achieving 16% improvement over very recently advanced competing methods with class-average F1-score on the Hospital dataset. The code of this work will be open source onhttps://github.com/xspc/Segmentation-Sensor-based-HAR. Songpengcheng Xia, Ling Pei, Wenxian Yu, Robert C. Qiu |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Multi-level Contrast Network for Wearables-based Joint Activity Segmentation and RecognitionabstractHuman activity recognition (HAR) with wearables is promising research that can be widely adopted in many smart healthcare applications. In recent years, the deep learning-based HAR models have achieved impressive recognition performance. However, most HAR algorithms are susceptible to the multi-class windows problem that is essential yet rarely exploited. In this paper, we propose to relieve this challenging problem by introducing the segmentation technology into HAR, yielding joint activity segmentation and recognition. Especially, we introduce the Multi-Stage Temporal Convolutional Network (MS-TCN) architecture for sample-level activity prediction to joint segment and recognize the activity sequence. Furthermore, to enhance the robustness of HAR against the inter-class similarity and intra-class heterogeneity, a multi-level contrastive loss, containing the sample-level and segment-level contrast, has been proposed to learn a well-structured embedding space for better activity segmentation and recognition performance. Finally, with comprehensive experiments, we verify the effectiveness of the proposed method on two public HAR datasets, achieving significant improvements in the various evaluation metrics. Songpengcheng Xia, Ling Pei, Wenxian Yu, Robert C. Qiu |
GLOBECOM | 1 |
| 2021 | Open Set Mixed-Reality Human Activity RecognitionabstractSensor-based human activity recognition (HAR) is a fundamental problem that can have a broad impact on many research/industrial fields. The deep learning methods pave the way for extracting robust and informative features from the non-stationary HAR data, achieving high-accuracy HAR. Most of the works in the literature consider the closed set activity recognition, which assumes all classes of activities are known in both the training and test stages. However, it is a challenging way to apply deep learning models to real-world applications, which can contain unseen activities. These activities will sharply decrease the performance of the deep learning models. To this end, we introduce the problem of Open Set Mixed-Reality (OSM) HAR, which aims to recognize unseen activities while classify seen samples. Furthermore, we propose a novel balanced open set backpropagation method to realize accurate and robust OSM-HAR. Lastly, we verify the effectiveness of the proposed method with a publicly available and our newly collected dataset. Songpengcheng Xia, Ling Pei |
GLOBECOM | 3 |
| 2021 | MARS: Mixed Virtual and Real Wearable Sensors for Human Activity Recognition With Multidomain Deep Learning ModelabstractTogether with the rapid development of the Internet of Things, human activity recognition (HAR) using wearable inertial measurement units (IMUs) becomes a promising technology for many research areas. Recently, deep-learning-based methods pave a new way of understanding and performing analysis of the complex data in the HAR system. However, the performance of these methods is mostly based on the quality and quantity of the collected data. In this article, we innovatively propose to build a large data set based on virtual IMUs and then address technical issues by introducing a multiple-domain deep learning framework consisting of three technical parts. In the first part, we propose to learn the single-frame human activity from the noisy IMU data with hybrid convolutional neural networks in the semisupervised form. For the second part, the extracted data features are fused according to the principle of uncertainty-aware consistency, which reduces the uncertainty by weighting the importance of the features. The transfer learning is performed in the last part based on the newly released archive of motion capture as surface shapes data set, containing abundant synthetic human poses, which enhances the variety and diversity of the training data set and is beneficial for the process of training and feature transfer in the proposed method. The efficiency and effectiveness of the proposed method have been demonstrated in the real deep inertial poser data set. The experimental results show that the proposed methods can surprisingly converge within a few iterations and outperform all competing methods. Ling Pei, Songpengcheng Xia, Fanyi Xiao, Qi Wu 0007, Wenxian Yu, Robert C. Qiu |
IEEE Internet Things J. | 2 |