EDBT 2026 Demo / reviewers in the wild / expert
Qi Wu 0007
dblp:96/3446-7
· DBLP profile ↗
17ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0001-9369-394XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Systems, architecture and hardware · 5 · 5 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RadarLLM: Empowering Large Language Models to Understand Human Motion from Millimeter-wave Point Cloud SequenceabstractMillimeter-wave radar offers a privacy-preserving and environment-robust alternative to vision-based sensing, enabling human motion analysis in challenging conditions such as low light, occlusions, rain, or smoke. However, its sparse point clouds pose significant challenges for semantic understanding. We present RadarLLM, the first framework that leverages large language models (LLMs) for human motion understanding from radar signals. RadarLLM introduces two key innovations: (1) a motion-guided radar tokenizer based on our Aggregate VQ-VAE architecture, integrating deformable body templates and masked trajectory modeling to convert spatial-temporal radar sequences into compact semantic tokens; and (2) a radar-aware language model that establishes cross-modal alignment between radar and text in a shared embedding space. To overcome the scarcity of paired radar-text data, we generate a realistic radar-text dataset from motion-text datasets with a physics-aware synthesis pipeline. Extensive experiments on both synthetic and real-world benchmarks show that RadarLLM achieves state-of-the-art performance, enabling robust and interpretable motion understanding under privacy and visibility constraints, even in adverse environments. Zengyuan Lai, Songpengcheng Xia, Lizhou Lin, Renwen Wang, Jianran Liu, Qi Wu 0007, Ling Pei |
AAAI | 8 |
| 2025 | EnvPoser: Environment-aware Realistic Human Motion Estimation from Sparse Observations with Uncertainty ModelingabstractEstimating full-body motion using the tracking signals of head and hands from VR devices holds great potential for various applications. However, the sparsity and unique distribution of observations present a significant challenge, resulting in an ill-posed problem with multiple feasible solutions (i.e., hypotheses). This amplifies uncertainty and ambiguity in full-body motion estimation, especially for the lower-body joints. Therefore, we propose a new method, EnvPoser, that employs a two-stage framework to perform full-body motion estimation using sparse tracking signals and pre-scanned environment from VR devices. EnvPoser models the multi-hypothesis nature of human motion through an uncertainty-aware estimation module in the first stage. In the second stage, we refine these multi-hypothesis estimates by integrating semantic and geometric environmental constraints, ensuring that the final motion estimation aligns realistically with both the environmental context and physical interactions. Qualitative and quantitative experiments on two public datasets demonstrate that our method achieves state-of-the-art performance, highlighting significant improvements in human motion estimation within motion-environment interaction scenarios. Project page: https://xspc.github.io/EnvPoser/. Songpengcheng Xia, Zhuo Su 0006, Xiaozheng Zheng, Guidong Wang, Qi Wu 0007, Ling Pei |
CVPR | 8 |
| 2025 | GASEM: Boosting Generalized and Actionable Parts Segmentation and Pose Estimation via Object Motion PerceptionabstractCategory-level object understanding has progressed, but generalized part perception remains underexplored. This paper introduces GASEM, a framework for Generalizable and Actionable Parts GAPart Segmentation and pose Estimation via object Motion perception. GASEM utilizes point-wise motion data from observed point clouds and cross-perspective alignment to learn object motion using a scene flow model. It features a segmentation proposal architecture for GAPart segmentation and an Normalized Object Coordinate Space(NPCS) branch for pose estimation. Additionally, a reinforcement learning agent is trained for robust GAPart manipulation in both simulations and real-world environments. Experiments on the GAPartNet dataset show GASEM outperforms state-of-the-art methods. This work promises advancements in embodied intelligence applications like robot-object interaction and generalizable manipulation. Codes are available at https://github.com/Zirconium233/GASEM. Liu Liu 0012, Li Zhang 0104, Yiming Tang 0001, Qi Wu 0007, Hao Wu 0040 |
ICME | 6 |
| 2025 | Suite-IN: Aggregating Motion Features from Apple Suite for Robust Inertial NavigationabstractWith the rapid development of wearable technology, devices like smartphones, smartwatches, and headphones equipped with IMUs have become essential for applications such as pedestrian positioning. However, traditional pedestrian dead reckoning (PDR) methods struggle with diverse motion patterns, while recent data-driven approaches, though improving accuracy, often lack robustness due to reliance on a single device. In our work, we attempt to enhance the positioning performance using the low-cost commodity IMUs embedded in the wearable devices. We propose a multi-device deep learning framework named Suite-IN, aggregating motion data from Apple Suite for inertial navigation. Motion data captured by sensors on different body parts contains both local and global motion information, making it essential to reduce the negative effects of localized movements and extract global motion representations from multiple devices. Our model innovatively introduces a contrastive learning module to disentangle motionshared and motion-private latent representations, enhancing positioning accuracy. We validate our method on a self-collected dataset consisting of Apple Suite: iPhone, Apple Watch and Airpods, which supports a variety of movement patterns and flexible device configurations. Experimental results demonstrate that our approach outperforms state-of-the-art models while maintaining robustness across diverse sensor configurations. Songpengcheng Xia, Junyuan Deng, Zengyuan Lai, Qi Wu 0007, Ling Pei |
ICRA | 6 |
| 2025 | mmDEAR: mmWave Point Cloud Density Enhancement for Accurate Human Body ReconstructionabstractMillimeter-wave (mmWave) radar offers robust sensing capabilities in diverse environments, making it a highly promising solution for human body reconstruction due to its privacy-friendly and non-intrusive nature. However, the significant sparsity of mm Wave point clouds limits the estimation accuracy. To overcome this challenge, we propose a two-stage deep learning framework that enhances mm Wave point clouds and improves human body reconstruction accuracy. Our method includes a mm Wave point cloud enhancement module that densifies the raw data by leveraging temporal features and a multi-stage completion network, followed by a 2D-3D fusion module that extracts both 2D and 3D motion features to refine SMPL parameters. The mm Wave point cloud enhancement module learns the detailed shape and posture information from 2D human masks in single-view images. However, image-based supervision is involved only during the training phase, and the inference relies solely on sparse point clouds to maintain privacy. Experiments on multiple datasets demonstrate that our approach outperforms state-of-the-art methods, with the enhanced point clouds further improving performance when integrated into existing models. Songpengcheng Xia, Zengyuan Lai, Qi Wu 0007, Wenxian Yu, Ling Pei |
ICRA | 5 |
| 2025 | 360Recon: An Accurate Reconstruction Method based on Depth Fusion from 360 ImagesabstractAccurate 3D reconstruction is crucial for AR and VR applications. Compared with traditional pinhole camera-based methods, 360° image-based reconstruction can achieve higher precision with fewer input images, making it especially effective in low-texture environments. However, the severe distortion resulting from the wide field of view complicates feature extraction and matching, leading to geometric inconsistencies in multi-view reconstruction. To address these challenges, we propose 360Recon, a novel multi-view stereo (MVS) algorithm specifically designed for equirectangular projection (ERP) images. With the proposed spherical feature extraction module mitigating distortion, 360Recon integrates a 3D cost volume with multi-scale ERP features to deliver high-precision scene reconstruction while preserving local geometric consistency. Experimental results demonstrate that 360Recon outperforms existing methods in terms of accuracy, computational efficiency, and generalization capability. The source code will be released at https://github.com/LeonATP/360Recon. Zhongmiao Yan, Qi Wu 0007, Songpengcheng Xia, Junyuan Deng, Xiang Mu, Renbiao Jin, Changchun Ye, Ling Pei |
IROS | 2 |
| 2025 | REArtGS: Reconstructing and Generating Articulated Objects via 3D Gaussian Splatting with Geometric and Motion ConstraintsabstractArticulated objects, as prevalent entities in human life, their 3D representations play crucial roles across various applications. However, achieving both high-fidelity textured surface reconstruction and dynamic generation for articulated objects remains challenging for existing methods. In this paper, we present REArtGS, a novel framework that introduces additional geometric and motion constraints to 3D Gaussian primitives, enabling realistic surface reconstruction and generation for articulated objects. Specifically, given multi-view RGB images of arbitrary two states of articulated objects, we first introduce an unbiased Signed Distance Field (SDF) guidance to regularize Gaussian opacity fields, enhancing geometry constraints and improving surface reconstruction quality. Then we establish deformable fields for 3D Gaussians constrained by the kinematic structures of articulated objects, achieving unsupervised generation of surface meshes in unseen states. Extensive experiments on both synthetic and real datasets demonstrate our approach achieves high-quality textured surface reconstruction for given states, and enables high-fidelity surface generation for unseen states. Project site: https://sites.google.com/view/reartgs/home. Liu Liu 0012, Zhou Linli, Anran Huang, Liangtu Song, Qiaojun Yu, Qi Wu 0007, Cewu Lu |
NeurIPS | 7 |
| 2025 | SMART: Scene-Motion-Aware Human Action Recognition Framework for Mental Disorder GroupabstractPatients with mental disorders often exhibit risky abnormal actions, such as climbing walls or hitting windows, necessitating intelligent video behavior monitoring for smart healthcare with the rising Internet of Things (IoT) technology. However, the development of vision-based human action recognition (HAR) for these actions is hindered by the lack of specialized algorithms and datasets. In this article, we innovatively propose to build a vision-based HAR dataset, including abnormal actions often occurring in the mental disorder group and then introduce a novel scene-motion-aware action recognition technology framework, named SMART, consisting of two technical modules. First, we propose a scene perception module to extract human motion trajectory and human-scene interaction features, which introduces additional scene information for a supplementary semantic representation of the above actions. Second, the multistage fusion module fuses the skeleton motion, motion trajectory, and human-scene interaction features, enhancing the semantic association between the skeleton motion and the above supplementary representation, thus generating a comprehensive representation with both human motion and scene information. The effectiveness of our proposed method has been validated on our self-collected HAR dataset (MentalHAD), achieving 94.9% and 93.1% accuracy in un-seen subjects and scenes and outperforming state-of-the-art approaches by 6.5% and 13.2%, respectively. The demonstrated subject- and scene- generalizability makes it possible for SMART’s migration to practical deployment in smart healthcare systems for mental disorder patients in medical settings. The code and dataset will be released publicly for further research:https://github.com/Inowlzy/SMART.git. Zengyuan Lai, Songpengcheng Xia, Qi Wu 0007, Wenxian Yu, Ling Pei |
IEEE Internet Things J. | 4 |
| 2024 | KPA-Tracker: Towards Robust and Real-Time Category-Level Articulated Object 6D Pose TrackingabstractOur life is populated with articulated objects. Current category-level articulation estimation works largely focus on predicting part-level 6D poses on static point cloud observations. In this paper, we tackle the problem of category-level online robust and real-time 6D pose tracking of articulated objects, where we propose KPA-Tracker, a novel 3D KeyPoint based Articulated object pose Tracker. Given an RGB-D image or a partial point cloud at the current frame as well as the estimated per-part 6D poses from the last frame, our KPA-Tracker can effectively update the poses with learned 3D keypoints between the adjacent frames. Specifically, we first canonicalize the input point cloud and formulate the pose tracking as an inter-frame pose increment estimation task. To learn consistent and separate 3D keypoints for every rigid part, we build KPA-Gen that outputs the high-quality ordered 3D keypoints in an unsupervised manner. During pose tracking on the whole video, we further propose a keypoint-based articulation tracking algorithm that mines keyframes as reference for accurate pose updating. We provide extensive experiments on validating our KPA-Tracker on various datasets ranging from synthetic point cloud observation to real-world scenarios, which demonstrates the superior performance and robustness of the KPA-Tracker. We believe that our work has the potential to be applied in many fields including robotics, embodied intelligence and augmented reality. All the datasets and codes are available at https://github.com/hhhhhar/KPA-Tracker. Liu Liu 0012, Anran Huang, Qi Wu 0007, Dan Guo 0001, Xun Yang 0001, Meng Wang 0001 |
AAAI | 3 |
| 2024 | ICAF-4: An Integrated Framework of Category-level Articulated Object Perception and Manipulation for Embodied Intelligence
Li Zhang 0104, Qiankun Li 0004, Qi Wu 0007, Lin Wu 0001, Liu Liu 0012 |
BMVC | 4 |
| 2024 | Dynamic Inertial Poser (DynaIP): Part-Based Motion Dynamics Learning for Enhanced Human Pose Estimation with Sparse Inertial SensorsabstractThis paper introduces a novel human pose estimation approach using sparse inertial sensors, addressing the short-comings of previous methods reliant on synthetic data. It leverages a diverse array of real inertial motion capture data from different skeleton formats to improve motion di-versity and model generalization. This method features two innovative components: a pseudo-velocity regression model for dynamic motion capture with inertial sensors, and a part-based model dividing the body and sensor data into three regions, each focusing on their unique characteristics. The approach demonstrates superior performance over state-of-the-art models across five public datasets, notably reducing pose error by 19% on the DIP-IMU dataset, thus representing a significant improvement in inertial sensor-based human pose estimation. Our codes are available at https://github.com/dx118/dynaip Songpengcheng Xia, Qi Wu 0007, Ling Pei |
CVPR | 5 |
| 2024 | mmBaT: A Multi-Task Framework for Mmwave-Based Human Body Reconstruction and Translation PredictionabstractHuman body reconstruction with Millimeter Wave (mmWave) radar point clouds has gained significant interest due to its ability to work in adverse environments and its capacity to mitigate privacy concerns associated with traditional camera-based solutions. Despite pioneering efforts in this field, two challenges persist. Firstly, raw point clouds contain massive noise points, usually caused by the ambient objects and multi-path effects of Radio Frequency (RF) signals. Recent approaches typically rely on prior knowledge or elaborate preprocessing methods, limiting their applicability. Secondly, even after noise removal, the sparse and inconsistent body-related points pose an obstacle to accurate human body reconstruction. To address these challenges, we introduce mmBaT, a novel multi-task deep learning framework that concurrently estimates the human body and predicts body translations in subsequent frames to extract body-related point clouds. Our method is evaluated on two public datasets that are collected with different radar devices and noise levels. A comprehensive comparison against other state-of-the-art methods demonstrates that our method has a superior reconstruction performance and generalization ability from noisy raw data, even when compared to methods provided with body-related point clouds. Songpengcheng Xia, Qi Wu 0007, Ling Pei |
ICASSP | 4 |
| 2024 | Explicit Interaction for Fusion-Based Place RecognitionabstractFusion-based place recognition is an emerging technique jointly utilizing multi-modal perception data, to recognize previously visited places in GPS-denied scenarios for robots and autonomous vehicles. Recent fusion-based place recognition methods combine multi-modal features in implicit manners. While achieving remarkable results, they do not explicitly consider what the individual modality affords in the fusion system. Therefore, the benefit of multi-modal feature fusion may not be fully explored. In this paper, we propose a novel fusion-based network, dubbed EINet, to achieve explicit interaction of the two modalities. EINet uses LiDAR ranges to supervise more robust vision features for long time spans, and simultaneously uses camera RGB data to improve the discrimination of LiDAR point clouds. In addition, we develop a new benchmark for the place recognition task based on the nuScenes dataset. To establish this benchmark for future research with comprehensive comparisons, we introduce both supervised and self-supervised training schemes alongside evaluation protocols. We conduct extensive experiments on the proposed benchmark, and the experimental results show that our EINet exhibits better recognition performance as well as solid generalization ability compared to the state-of-the-art fusion-based place recognition approaches. Our open-source code and benchmark are released at: https://github.com/BIT-XJY/EINet. Junyi Ma, Qi Wu 0007, Yue Wang 0020, Xieyuanli Chen, Wenxian Yu, Ling Pei |
IROS | 3 |
| 2024 | Thermal-NeRF: Neural Radiance Fields from an Infrared CameraabstractIn recent years, Neural Radiance Fields (NeRFs) have demonstrated significant potential in encoding highly-detailed 3D geometry and environmental appearance, positioning themselves as a promising alternative to traditional explicit representation for 3D scene reconstruction. However, the predominant reliance on RGB imaging presupposes ideal lighting conditions—a premise frequently unmet in robotic applications plagued by poor lighting or visual obstructions. This limitation overlooks the capabilities of infrared (IR) cameras, which excel in low-light detection and present a robust alternative under such adverse scenarios. To tackle these issues, we introduce Thermal-NeRF, the first method that estimates a volumetric scene representation in the form of a NeRF solely from IR imaging. By leveraging a thermal mapping and structural thermal constraint derived from the thermal characteristics of IR imaging, our method showcases unparalleled proficiency in recovering NeRFs in visually degraded scenes where RGB-based methods fall short. We conduct extensive experiments to demonstrate that Thermal-NeRF can achieve superior quality compared to existing methods. Furthermore, we contribute a dataset for IR-based NeRF applications, paving the way for future research in IR NeRF reconstruction, see https://github.com/Cerf-Volant425/Thermal-NeRF. Tianxiang Ye, Qi Wu 0007, Junyuan Deng, Liu Liu 0012, Songpengcheng Xia, Wenxian Yu, Ling Pei |
IROS | 2 |
| 2023 | NeRF-LOAM: Neural Implicit Representation for Large-Scale Incremental LiDAR Odometry and MappingabstractSimultaneously odometry and mapping using LiDAR data is an important task for mobile systems to achieve full autonomy in large-scale environments. However, most existing LiDAR-based methods prioritize tracking quality over reconstruction quality. Although the recently developed neural radiance fields (NeRF) have shown promising advances in implicit reconstruction for indoor environments, the problem of simultaneous odometry and mapping for large-scale scenarios using incremental LiDAR data remains unexplored. To bridge this gap, in this paper, we propose a novel NeRF-based LiDAR odometry and mapping approach, NeRF-LOAM, consisting of three modules neural odometry, neural mapping, and mesh reconstruction. All these modules utilize our proposed neural signed distance function, which separates LiDAR points into ground and non-ground points to reduce Z-axis drift, optimizes odometry and voxel embeddings concurrently, and in the end generates dense smooth mesh maps of the environment. Moreover, this joint optimization allows our NeRF-LOAM to be pre-trained free and exhibit strong generalization abilities when applied to different environments. Extensive evaluations on three publicly available datasets demonstrate that our approach achieves state-of-the-art odometry and mapping performance, as well as a strong generalization in large-scale environments utilizing LiDAR data. Furthermore, we perform multiple ablation studies to validate the effectiveness of our network design. The implementation of our approach will be made available at https://github.com/JunyuanDeng/NeRF-LOAM. Junyuan Deng, Qi Wu 0007, Xieyuanli Chen, Songpengcheng Xia, Wenxian Yu, Ling Pei |
ICCV | 2 |
| 2021 | Grayscale And Normal Guided Depth Completion With A Low-Cost LidarabstractIn this paper, we introduce DenseLivox, a dataset with dense and accurate depth as ground truth. To our best knowledge, it is the first dataset with dense ground truth designed for LiDAR depth completion using a low-cost LiDAR. Also, we develop a simple yet effective multi-task learning network to tackle the problem of depth completion. Compared to the works in the literature, our model’s uniqueness is that it completes a depth map, a normal map, and a grayscale image simultaneously. To address the area with heavy noises, we use modified Huber loss to smooth these outliers’ effect. We evaluate our method on DenseLivox and show that accuracy is greatly improved with the grayscale and normal guidance. Our method outperforms other depth-only methods and is comparable to the methods that take RGB and depth as input. Qingyang Yu, Qi Wu 0007, Ling Pei |
ICIP | 3 |
| 2021 | MARS: Mixed Virtual and Real Wearable Sensors for Human Activity Recognition With Multidomain Deep Learning ModelabstractTogether with the rapid development of the Internet of Things, human activity recognition (HAR) using wearable inertial measurement units (IMUs) becomes a promising technology for many research areas. Recently, deep-learning-based methods pave a new way of understanding and performing analysis of the complex data in the HAR system. However, the performance of these methods is mostly based on the quality and quantity of the collected data. In this article, we innovatively propose to build a large data set based on virtual IMUs and then address technical issues by introducing a multiple-domain deep learning framework consisting of three technical parts. In the first part, we propose to learn the single-frame human activity from the noisy IMU data with hybrid convolutional neural networks in the semisupervised form. For the second part, the extracted data features are fused according to the principle of uncertainty-aware consistency, which reduces the uncertainty by weighting the importance of the features. The transfer learning is performed in the last part based on the newly released archive of motion capture as surface shapes data set, containing abundant synthetic human poses, which enhances the variety and diversity of the training data set and is beneficial for the process of training and feature transfer in the proposed method. The efficiency and effectiveness of the proposed method have been demonstrated in the real deep inertial poser data set. The experimental results show that the proposed methods can surprisingly converge within a few iterations and outperform all competing methods. Ling Pei, Songpengcheng Xia, Fanyi Xiao, Qi Wu 0007, Wenxian Yu, Robert C. Qiu |
IEEE Internet Things J. | 5 |