VLDB 2026 Research / reviewers in the wild / expert
Bing Wang 0013
dblp:06/1909-13
· DBLP profile ↗
45ranked-venue papers
4as first author
42since 2021 · last 2026
0000-0003-0977-0426ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 3 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 17 since 2021Systems, architecture and hardware · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Computer networks · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Collaborative Feature Matching with Progressive Correspondence LearningabstractAccurate feature matching between image pairs is fundamental for various computer vision applications. In detector-base process, the feature matcher aims to find the optimal feature correspondences, and the match filter is used for further removing mismatches. However, their connection is rarely exploited since they are usually treated as two separate issues in previous method, which may lead to suboptimal results. In this paper, we propose an end-to-end collaborative feature matching (CFM) method, which contains a keypoint learning (KL) module and a correspondence learning (CL) module, to bridge the gap between two types of works. The former improves the discrimination of keypoints, and provides high-quality dynamic matches for CL module. The latter further captures the rich context of matches, and gives effective feedback to KL module. These two modules can reinforce each other in a progressive manner. Besides, we develop an efficient version of CFM, named ECFM, using an adaptive sampling strategy to avoid the negative influence of uninformative keypoints. Experimental results indicate that both methods outperform the state-of-the-art competitors in the tasks of relative pose estimation and visual localization. Xin Liu 0091, Yanbing Han, Rong Qin 0001, Bing Wang 0013, Jufeng Yang |
AAAI | 4 |
| 2026 | Joint 3D modeling and motion estimation of dynamic wind-turbine blades from monocular UAV videos
Zhizhang Zhou, Wenbo Hu 0009, Weihao Hong, Xuanqing Huang, Bing Wang 0013, You Dong |
Adv. Eng. Informatics | 5 |
| 2026 | Efficient Scene Modeling via Structure-Aware and Region-Prioritized 3D GaussiansabstractReconstructing 3D scenes with high fidelity and efficiency remains a central pursuit in computer vision and graphics. Recent advances in 3D Gaussian Splatting (3DGS) enable photorealistic rendering with Gaussian primitives, yet the modeling process remains governed predominantly by photometric supervision. This reliance often leads to irregular spatial distribution and indiscriminate primitive adjustments that largely ignore underlying geometric context. In this work, we rethink Gaussian modeling from a geometric standpoint and introduce Mini-Splatting2, an efficient scene modeling framework that couples structure-aware distribution and region-prioritized optimization, driving 3DGS into a geometry-regulated paradigm. The structure-aware distribution enforces spatial regularity through structured reorganization and representation sparsity, ensuring balanced structural coverage for compact organization. The region-prioritized optimization improves training discrimination through geometric saliency and computational selectivity, fostering appropriate structural emergence for fast convergence. These mechanisms alleviate the long-standing tension among representation compactness, convergence acceleration, and rendering fidelity. Extensive experiments demonstrate that Mini-Splatting2 achieves up to 4× fewer Gaussians and 3× faster optimization while maintaining state-of-the-art visual quality, paving the way towards structured and efficient 3D Gaussian modeling. Guangchi Fang, Bing Wang 0013 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | GrowSP++: Growing Superpoints and Primitives for Unsupervised 3D Semantic SegmentationabstractWe study the problem of 3D semantic segmentation from raw point clouds. Unlike existing methods which primarily rely on a large amount of human annotations for training neural networks, we proposes GrowSP++, an unsupervised method to successfully identify complex semantic classes for every point in 3D scenes, without needing any type of human labels. Our method is composed of three major components: 1) a feature extractor incorporating 2D-3D feature distillation, 2) a superpoint constructor featuring progressively growing superpoints, and 3) a semantic primitive constructor with an additional growing strategy. The key to our method is the superpoint constructor together with the progressive growing strategy on both superpoints and semantic primitives, driving the feature extractor to progressively learn similar features for 3D points belonging to the same semantic class. We extensively evaluate our method on five challenging indoor and outdoor datasets, demonstrating state-of-the-art performance over all unsupervised baselines. We hope our work could inspire more advanced methods for unsupervised 3D semantic learning. Weisheng Dai, Bing Wang 0013, Bo Li 0037, Bo Yang 0027 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion PriorsabstractSurface normal estimation serves as a cornerstone for a spectrum of computer vision applications. While numerous efforts have been devoted to static image scenarios, ensuring temporal coherence in video-based normal estimation remains a formidable challenge. Instead of merely augmenting existing methods with temporal components, we present NormalCrafter to leverage the inherent temporal priors of video diffusion models. To secure high-fidelity normal estimation across sequences, we propose Semantic Feature Regularization (SFR), which aligns diffusion features with semantic cues, encouraging the model to concentrate on the intrinsic semantics of the scene. Moreover, we introduce a two-stage training protocol that leverages both latent and pixel space learning to preserve spatial accuracy while maintaining long temporal context. Extensive evaluations demonstrate the efficacy of our method, showcasing a superior performance in generating temporally consistent normal sequences with intricate details from diverse videos. Yanrui Bin, Xinya Chen, Bing Wang 0013 |
ICCV | 5 |
| 2025 | MambaMl: Exploring State Space Models for Multi-Label Image Classification
Xuelin Zhu, Jiuxin Cao, Bing Wang 0013 |
ICCV | 4 |
| 2025 | OFF3D:Object-Centric Feature Field for 3D Scene Segmentationabstract3D scene segmentation is a fundamental but challenging task for 3D scene understanding in computer vision. However, the majority of existing methods rely heavily on labor-intensive and costly human-annotated 2D or 3D labels. To address this issue, we propose a novel Object-Centric Feature Field for 3d scene segmentation, called OFF3D. Given the machine-generated inconsistent 2D masks, OFF3D can build an implicit object feature field to achieve consistent segmentation for all objects in the 3D scene. Specifically, OFF3D utilizes the continuous object features to represent object properties of each 3D point and proposes a Pseudo-Segment Clustering module to roughly locate objects in the 3D scene at the segment level and design a Confidence-Weighted Contrastive loss to achieve precise object segmentation by pixel-level feature optimization. Extensive experiments demonstrate the effectiveness of our method compared to state-of-the-art methods both quantitatively and qualitatively. Qinwei Lin, Bing Wang 0013, Jun Xu 0019, Haoqian Wang |
ICME | 2 |
| 2025 | AVIP: Acoustic-Visual-Inertial-Pressure Fusion-based Underwater Localization System with Multi-Centric CalibrationabstractUnderwater localization is a crucial capability for ensuring robust and accurate vehicle navigation. Although various well-developed localization systems exist, their primary focus is on ground and aerial applications. The challenges posed by underwater environments, such as sparse textures and dynamic disturbances, enable the multi-modal fusion a promising solution for localization. This paper presents AVIP, a localization method that fuses Acoustic, Visual, Inertial, and Pressure modalities for underwater applications. To integrate the information from all modalities during initialization, visual and inertial modalities are alternately assigned as centric sensors to pairwise predict and update estimations of other modalities. The multi-centric calibration problem is addressed through factor graph optimization, which is fully integrated into the graph-based AVIP system as the calibration factor. To evaluate the performance and compare to state-of-the-art approaches, the proposed method is evaluated using semi-physical datasets recorded by a BlueROV2 robot and public real-world datasets. Extensive experiments demonstrate that AVIP achieves superior localization accuracy and exhibits adaptability across a range of sensor configurations. Yuanbo Xue, Dejin Zhang, Chih-Yung Wen, Bing Wang 0013 |
IROS | 5 |
| 2025 | U-DECN: End-to-End Underwater Object Detection ConvNet With Improved Denoising TrainingabstractUnderwater object detection has higher requirements of running speed and deployment efficiency for the detector due to its specific environmental challenges. NMS of two- or one-stage object detectors and transformer architecture of query-based end-to-end object detectors are not conducive to deployment on underwater embedded devices with limited processing power. As for the detrimental effect of underwater color cast noise, recent underwater object detectors make network architecture or training complex, which also hinders their application and deployment on unmanned underwater vehicles. In this paper, we propose the Underwater DECO with improved deNoising training (U-DECN), the query-based end-to-end object detector (with ConvNet encoder-decoder architecture) for underwater color cast noise that addresses the above problems. We integrate advanced technologies from DETR variants into DECO and design optimization methods specifically for the ConvNet architecture, including Deformable Convolution in SIM and Separate Contrastive DeNoising Forward methods. To address the underwater color cast noise issue, we propose an Underwater Color DeNoising Query method to improve the generalization of the model for the biased object feature information by different color cast noise. Our U-DECN, with ResNet-50 backbone, achieves the best 64.0 AP on DUO and the best 58.1 AP on RUOD, and 21 FPS (5 times faster than Deformable DETR and DINO 4 FPS) on NVIDIA AGX Orin by TensorRT FP16, outperforming the other state-of-the-art query-based end-to-end object detectors. The code is available at https://github.com/LEFTeyex/U-DECN. Zhuoyan Liu, Bo Wang 0015, Bing Wang 0013, Ye Li 0027 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | From Light to Position: An Underwater Visual-Inertial Positioning Method Using Visible LightabstractReliable localization is essential for the efficiency and safety of autonomous underwater operations. In this study, we propose a novel localization framework centered on a strapdown inertial navigation system, which receives high-precision pose corrections from image-based visible-light positioning. A structured LED array is deployed as a stable underwater positioning reference, and an adaptive optical signal processing method is developed to enhance the extraction of light spot features under challenging scenarios. To address the distortions introduced by cross-medium refraction, we introduce a compensation model that restores accurate camera pose estimation without the need for cumbersome recalibration. Extensive experimental evaluations demonstrate that the proposed system achieves millimeter-level accuracy under favorable scenarios and maintains centimeter-level robustness in unfavorable observation scenarios. Owing to its high-precision, robustness, and cost-effectiveness, the proposed approach holds significant promise for advancing autonomous underwater navigation and next-generation intelligent robotic systems. Fanyi Meng 0004, Zheng Cong, Bing Wang 0013, Dejin Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | NeuV-SLAM: Fast Neural Multiresolution Voxel Optimization for RGBD Dense SLAMabstractWe introduce NeuV-SLAM, a novel dense simultaneous localization and mapping pipeline based on neural multiresolution voxels, characterized by ultra-fast convergence and incremental expansion capabilities. This pipeline utilizes RGBD images as input to construct multiresolution neural voxels, achieving rapid convergence while maintaining robust incremental scene reconstruction and camera tracking. Central to our methodology is to propose a novel implicit representation, termedVDFthat combines the implementation of neural signed distance field (SDF) voxels with an SDF activation strategy. This approach entails the direct optimization of color features and SDF values anchored within the voxels, substantially enhancing the rate of scene convergence. To ensure the acquisition of clear edge delineation, SDF activation is designed, which maintains exemplary scene representation fidelity even under constraints of voxel resolution. Furthermore, in pursuit of advancing rapid incremental expansion with low computational overhead, we developedhashMV, a novel hash-based multiresolution voxel management structure. This architecture is complemented by a strategically designed voxel generation technique that synergizes with a two-dimensional scene prior. Our empirical evaluations, conducted on the Replica and ScanNet Datasets, substantiate NeuV-SLAM's exceptional efficacy in terms of convergence speed, tracking accuracy, scene reconstruction, and rendering quality. Wenzhi Guo, Bing Wang 0013, Lijun Chen 0006 |
IEEE Trans. Multim. | 2 |
| 2025 | Learning Selective Sensor Fusion for State EstimationabstractAutonomous vehicles and mobile robotic systems are typically equipped with multiple sensors to provide redundancy. By integrating the observations from different sensors, these mobile agents are able to perceive the environment and estimate system states, e.g., locations and orientations. Although deep learning (DL) approaches for multimodal odometry estimation and localization have gained traction, they rarely focus on the issue of robust sensor fusion-a necessary consideration to deal with noisy or incomplete sensor observations in the real world. Moreover, current deep odometry models suffer from a lack of interpretability. To this extent, we propose SelectFusion, an end-to-end selective sensor fusion module that can be applied to useful pairs of sensor modalities, such as monocular images and inertial measurements, depth images, and light detection and ranging (LIDAR) point clouds. Our model is a uniform framework that is not restricted to specific modality or task. During prediction, the network is able to assess the reliability of the latent features from different sensor modalities and to estimate trajectory at both scale and global pose. In particular, we propose two fusion modules-a deterministic soft fusion and a stochastic hard fusion-and offer a comprehensive study of the new strategies compared with trivial direct fusion. We extensively evaluate all fusion strategies both on public datasets and on progressively degraded datasets that present synthetic occlusions, noisy and missing data, and time misalignment between sensors, and we investigate the effectiveness of the different fusion strategies in attending the most reliable features, which in itself provides insights into the operation of the various models. Changhao Chen, Stefano Rosa, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | PointCG: Self-Supervised Point Cloud Learning via Joint Completion and GenerationabstractThe core of self-supervised point cloud learning lies in setting up appropriate pretext tasks, to construct a pre-training framework that enables the encoder to perceive 3D objects effectively. In this article, we integrate two prevalent methods, masked point modeling (MPM) and 3D-to-2D generation, as pretext tasks within a pre-training framework. We leverage the spatial awareness and precise supervision offered by these two methods to address their respective limitations: ambiguous supervision signals and insensitivity to geometric information. Specifically, the proposed framework, abbreviated as PointCG, consists of a Hidden Point Completion (HPC) module and an Arbitrary-view Image Generation (AIG) module. We first capture visible points from arbitrary views as inputs by removing hidden points. Then, HPC extracts representations of the inputs with an encoder and completes the entire shape with a decoder, while AIG is used to generate rendered images based on the visible points' representations. Extensive experiments demonstrate the superiority of the proposed method over the baselines in various downstream tasks. Our code will be made available upon acceptance. Yun Liu 0002, Peng Li 0064, Xuefeng Yan 0001, Liangliang Nan, Bing Wang 0013, Honghua Chen, Lina Gong, Wei Zhao 0039, Mingqiang Wei |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Mini-Splatting: Representing Scenes with a Constrained Number of Gaussians
Guangchi Fang, Bing Wang 0013 |
ECCV (77) | 2 |
| 2024 | Equi-GSPR: Equivariant SE(3) Graph Network Model for Sparse Point Cloud Registration
Xueyang Kang, Zhaoliang Luan, Kourosh Khoshelham, Bing Wang 0013 |
ECCV (4) | 4 |
| 2024 | FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth EstimatorsabstractMatching cross-modality features between images and point clouds is a fundamental problem for image-to-point cloud registration. However, due to the modality difference between images and points, it is difficult to learn robust and discriminative cross-modality features by existing metric learning methods for feature matching. Instead of applying metric learning on cross-modality data, we propose to unify the modality between images and point clouds by pretrained large-scale models first, and then establish robust correspondence within the same modality. We show that the intermediate features, called diffusion features, extracted by depth-to-image diffusion models are semantically consistent between images and point clouds, which enables the building of coarse but robust cross-modality correspondences. We further extract geometric features on depth maps produced by the monocular depth estimator. By matching such geometric features, we significantly improve the accuracy of the coarse correspondences produced by diffusion features. Extensive experiments demonstrate that without any task-specific training, direct utilization of both features produces accurate image-to-point cloud registration. On three public indoor and outdoor benchmarks, the proposed method averagely achieves a 20.6 percent improvement in Inlier Ratio, a $3.0\times$ higher Inlier Number, and a 48.6 percent improvement in Registration Recall than existing state-of-the-arts. The code and additional results are available at \url{https://whu-usi3dv.github.io/FreeReg/}. Haiping Wang 0004, Yuan Liu 0025, Bing Wang 0013, Yujing Sun 0001, Zhen Dong 0005, Wenping Wang 0001, Bisheng Yang |
ICLR | 3 |
| 2024 | Learning to Catch Reactive Objects with a Behavior PredictorabstractTracking and catching moving objects is an important ability for robots in a dynamic world. Whilst some objects have highly predictable state evolution e.g., the ballistic trajectory of a tennis ball, reactive targets alter their behavior in response to motion of the manipulator. Reactive applications range from gently capturing living animals such as snakes or fish for biological investigations, to smoothly interacting with and assisting a person. Existing works for dynamic catching usually perform target prediction followed by planning, but seldom account for highly non-linear reactive behaviors. Alternatively, Reinforcement Learning (RL) based methods simply treat the target and its motion as part of the observation of the world-state, but perform poorly due to the weak reward signal. In this work, we blend the approach of an explicit, yet learned, target state predictor with RL. We further show how a tightly coupled predictor which ‘observes’ the state of the robot leads to significantly improved anticipatory action, especially with targets that seek to evade the robot following a simple policy. Experiments show that our method achieves an 86.4% (open plane area) and a 73.8% (room) success rate on evasive objects, outperforming monolithic reinforcement learning and other techniques. We also demonstrate the efficacy of our approach across varied targets and trajectories. All code, data, and additional videos are at this GitHub link: https://kl-research.github.io/dyncatch. Kai Lu 0003, Jia-Xing Zhong, Bo Yang 0027, Bing Wang 0013, Andrew Markham |
ICRA | 4 |
| 2024 | RenderOcc: Vision-Centric 3D Occupancy Prediction with 2D Rendering Supervisionabstract3D occupancy prediction holds significant promise in the fields of robot perception and autonomous driving, which quantifies 3D scenes into grid cells with semantic labels. Recent works mainly utilize complete occupancy labels in 3D voxel space for supervision. However, the expensive annotation process and sometimes ambiguous labels have severely constrained the usability and scalability of 3D occupancy models. To address this, we present RenderOcc, a novel paradigm for training 3D occupancy models only using 2D labels. Specifically, we extract a NeRF-style 3D volume representation from multi-view images, and employ volume rendering techniques to establish 2D renderings, thus enabling direct 3D supervision from 2D semantics and depth labels. Additionally, we introduce an Auxiliary Ray method to tackle the issue of sparse viewpoints in autonomous driving scenarios, which leverages sequential frames to construct comprehensive 2D rendering for each object. To our best knowledge, RenderOcc is the first attempt to train multi-view 3D occupancy models only using 2D labels, reducing the dependence on costly 3D occupancy annotations. Extensive experiments demonstrate that RenderOcc achieves comparable performance to models fully supervised with 3D labels, underscoring the significance of this approach in real-world applications. Our code is available at https://github.com/pmj110119/RenderOcc. Mingjie Pan, Jiaming Liu 0003, Renrui Zhang, Peixiang Huang, Xiaoqi Li 0020, Hongwei Xie, Bing Wang 0013, Li Liu 0069, Shanghang Zhang |
ICRA | 7 |
| 2024 | An Adaptive Weighted GNSS/VINS/Wi-Fi RTT-based Seamless Positioning System for SmartphoneabstractExisting positioning methods have limitations in accuracy and reliability for seamless positioning. Global Navigation Satellite System (GNSS) faces multipath and signal blockage issues, especially indoors. Wi-Fi positioning solutions are mostly restricted to indoors due to large infrastructure requirements. Infrastructure-independent positioning systems, such as visual-inertial navigation systems (VINS), are limited to providing relative pose estimation and are affected by environmental luminance. This paper presents an approach for smartphone seamless positioning by adaptively fusing multi-sensor data from GNSS, Wi-Fi Round Trip Time (RTT) and VINS using factor graph optimization (FGO). The adaptive weighted FGO algorithm optimizes the estimation of the state variables by minimizing the loss function of selected factors with scaled covariance. Combining the strengths of GNSS, Wi-Fi RTT, and VINS, the proposed system achieves improved positioning accuracy and robustness, and experimental results demonstrate our method’s effectiveness in various scenarios. Meiling Su, Bing Wang 0013, Sugata Ahad, Guohao Zhang, Li-Ta Hsu |
IPIN | 2 |
| 2024 | A Degradation-Robust Keyframe Selection Method Based on Image Quality Evaluation for Visual LocalizationabstractLocalization information is increasingly crucial for incorporating location context into Internet of Things (IoT) data. As an important task in visual localization, keyframe selection helps effective augmentation of visual odometry. Although considerable progress has been made in the research field of keyframe selection, they have rarely focused on dealing with degraded input sensory data in the real world. To this extent, this work proposes a novel concept by incorporating image quality evaluation into the visual localization so that the keyframe selection module can identify images that may cause undesirable effects and take measures to avoid the catastrophic impact of degraded images. The quality for each image is estimated online using deep classifier trained with the image-itself, image-differential, and external information. Since no model-specific knowledge is needed, our method is applicable to any visual localization system. By creating a challenging dataset based on current public datasets under autonomous driving and unmanned aerial vehicles (UAV) scenarios and using it to evaluate our method, we obtain estimated trajectories that are closer to the original situation while validating its robustness to challenging degraded environments. Jianfan Chen, Qingquan Li 0001, Bing Wang 0013, Dejin Zhang |
IEEE Internet Things J. | 4 |
| 2024 | Drone-NeRF: Efficient NeRF based 3D scene reconstruction for large-scale drone survey
Bing Wang 0013, Changhao Chen |
Image Vis. Comput. | 2 |
| 2024 | Attention-enhanced joint learning network for micro-video venue classification
Bing Wang 0013, Xianglin Huang, Gang Cao 0001, Lifang Yang, Zhulin Tao |
Multim. Tools Appl. | 1 |
| 2024 | DarkLoc+: Thermal Image-Based Indoor Localization for Dark Environments With Relative Geometry ConstraintsabstractThermal images capture temperature information of the environments instead of texture, making it well suitable for obtaining position in dark environments. Many methods have been proposed to handle RGB images, while thermal image-based localization methods are not well studied. To address it, we propose DarkLoc+, a thermal image-based indoor localization method based on the attention model and relative constraints between images under a learning-based localization framework. To be specific, we utilize self-attention to extract reprehensive features from thermal images and exploit relative constraints to enforce the convolutional neural networks (CNNs) to predict global poses. Relative pose loss(RelLoss)and relative regression loss are designed to work with global poses to constrain the network in feature and pose space simultaneously. We evaluate the proposed method on the public thermal images indoor dataset and our own dataset. The experimental results demonstrate that our method can obtain accurate position information. Baoding Zhou, Yufeng Xiao, Qing Li 0029, Bing Wang 0013, Longmin Pan, Dejin Zhang, Jiasong Zhu, Qingquan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | WHU-Railway3D: A Diverse Dataset and Benchmark for Railway Point Cloud Semantic SegmentationabstractPoint cloud semantic segmentation (PCSS) shows great potential in generating accurate 3D semantic maps for digital twin railways. Deep learning-based methods have seen substantial advancements, driven by numerous PCSS datasets. Nevertheless, existing datasets tend to neglect railway scenes, with limitations in scale, categories, and scene diversity. This motivated us to establish WHU-Railway3D, a diverse PCSS dataset specifically designed for railway scenes. WHU-Railway3D is categorized into urban, rural, and plateau railways based on scene complexity and semantic class distribution. The dataset spans approximately 30 km with 4.6 billion points labeled into 11 classes, such as rails, masts, overhead lines, and fences. In addition to 3D coordinates, WHU-Railway3D provides rich attribute information such as reflected intensity, scanning angle, and number of returns. Cutting-edge methods are extensively evaluated on the dataset, followed by in-depth analysis. Lastly, key challenges and potential future work are identified to stimulate further innovative research. The dataset is accessible athttps://github.com/WHU-USI3DV/WHU-Railway3D. Yuzhou Zhou, Bing Wang 0013, Jianping Li 0004, Zhen Dong 0005, Chenglu Wen, Zhiliang Ma, Bisheng Yang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Deep Learning for Visual Localization and Mapping: A SurveyabstractDeep-learning-based localization and mapping approaches have recently emerged as a new research direction and receive significant attention from both industry and academia. Instead of creating hand-designed algorithms based on physical models or geometric theories, deep learning solutions provide an alternative to solve the problem in a data-driven way. Benefiting from the ever-increasing volumes of data and computational power on devices, these learning methods are fast evolving into a new area that shows potential to track self-motion and estimate environmental models accurately and robustly for mobile agents. In this work, we provide a comprehensive survey and propose a taxonomy for the localization and mapping methods using deep learning. This survey aims to discuss two basic questions: whether deep learning is promising for localization and mapping, and how deep learning should be applied to solve this problem. To this end, a series of localization and mapping topics are investigated, from the learning-based visual odometry and global relocalization to mapping, and simultaneous localization and mapping (SLAM). It is our hope that this survey organically weaves together the recent works in this vein from robotics, computer vision, and machine learning communities and serves as a guideline for future researchers to apply deep learning to tackle the problem of visual localization and mapping. Changhao Chen, Bing Wang 0013, Xiaoxuan Lu 0001, Agathoniki Trigoni, Andrew Markham |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | GrowSP: Unsupervised Semantic Segmentation of 3D Point CloudsabstractWe study the problem of 3D semantic segmentation from raw point clouds. Unlike existing methods which primarily rely on a large amount of human annotations for training neural networks, we propose the first purely unsupervised method, called GrowSP, to successfully identify complex semantic classes for every point in 3D scenes, without needing any type of human labels or pretrained models. The key to our approach is to discover 3D semantic elements via progressive growing of superpoints. Our method consists of three major components, 1) the feature extractor to learn per-point features from input point clouds, 2) the superpoint constructor to progressively grow the sizes of superpoints, and 3) the semantic primitive clustering module to group superpoints into semantic elements for the final semantic segmentation. We extensively evaluate our method on multiple datasets, demonstrating superior performance over all unsupervised baselines and approaching the classic fully-supervised PointNet. We hope our work could inspire more advanced methods for unsupervised 3D semantic learning. Bo Yang 0027, Bing Wang 0013, Bo Li 0037 |
CVPR | 3 |
| 2023 | DM-NeRF: 3D Scene Geometry Decomposition and Manipulation from 2D Images
Bing Wang 0013, Bo Yang 0027 |
ICLR | 1 |
| 2023 | Decoupling Skill Learning from Robotic Control for Generalizable Object ManipulationabstractRecent works in robotic manipulation through reinforcement learning (RL) or imitation learning (IL) have shown potential for tackling a range of tasks e.g., opening a drawer or a cupboard. However, these techniques generalize poorly to unseen objects. We conjecture that this is due to the high-dimensional action space for joint control. In this paper, we take an alternative approach and separate the task of learning ‘what to do’ from ‘how to do it’ i.e., whole-body control. We pose the RL problem as one of determining the skill dynamics for a disembodied virtual manipulator interacting with articulated objects. The whole-body robotic kinematic control is optimized to execute the high-dimensional joint motion to reach the goals in the workspace. It does so by solving a quadratic programming (QP) model with robotic singularity and kinematic constraints. Our experiments on manipulating complex articulated objects show that the proposed approach is more generalizable to unseen objects with large intra-class variations, outperforming previous approaches. The evaluation results indicate that our approach generates more compliant robotic motion and outperforms the pure RL and IL baselines in task success rates. Additional information and videos are available at https://kl-research.github.io/decoupskill. Kai Lu 0003, Bo Yang 0027, Bing Wang 0013, Andrew Markham |
ICRA | 3 |
| 2023 | CubeLearn: End-to-End Learning for Human Motion Recognition From Raw mmWave Radar SignalsabstractmmWave FMCW radar has attracted a huge amount of research interest for human-centered applications in recent years, such as human gesture and activity recognition. Most existing pipelines are built upon conventional discrete Fourier transform (DFT) preprocessing and deep neural network classifier hybrid methods, with a majority of previous works focusing on designing the downstream classifier to improve overall accuracy. In this work, we take a step back and look at the preprocessing module. To avoid the drawbacks of conventional DFT preprocessing, we propose a complex-weighted learnable preprocessing module, named CubeLearn, to directly extract features from raw radar signal and build an end-to-end deep neural network for mmWave FMCW radar motion recognition applications. Extensive experiments show that our CubeLearn module consistently improves the classification accuracies of different pipelines, especially, benefiting those simpler models, which are more likely to be used on edge devices due to their computational efficiency. We provide ablation studies on initialization methods and structure of the proposed module, as well as an evaluation of the running time on PC and edge devices. This work also serves as a comparison of different approaches toward data cube slicing. Through our task-agnostic design, we propose a first step toward a generic end-to-end solution for radar recognition problems. Peijun Zhao, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham |
IEEE Internet Things J. | 3 |
| 2023 | DS-Net++: Dynamic Weight Slicing for Efficient Inference in CNNs and Vision TransformersabstractDynamic networks have shown their promising capability in reducing theoretical computation complexity by adapting their architectures to the input during inference. However, their practical runtime usually lags behind the theoretical acceleration due to inefficient sparsity. In this paper, we explore a hardware-efficient dynamic inference regime, named dynamic weight slicing, that can generalized well on multiple dimensions in both CNNs and transformers (e.g. kernel size, embedding dimension, number of heads, etc.). Instead of adaptively selecting important weight elements in a sparse way, we pre-define dense weight slices with different importance level by nested residual learning. During inference, weights are progressively sliced beginning with the most important elements to less important ones to achieve different model capacity for inputs with diverse difficulty levels. Based on this conception, we present DS-CNN++ and DS-ViT++, by carefully designing the double headed dynamic gate and the overall network architecture. We further propose dynamic idle slicing to address the drastic reduction of embedding dimension in DS-ViT++. To ensure sub-network generality and routing fairness, we propose a disentangled two-stage optimization scheme. In Stage I, in-place bootstrapping (IB) and multi-view consistency (MvCo) are proposed to stablize and improve the training of DS-CNN++ and DS-ViT++ supernet, respectively. In Stage II, sandwich gate sparsification (SGS) is proposed to assist the gate training. Extensive experiments on 4 datasets and 3 different network architectures demonstrate our methods consistently outperform the state-of-the-art static and dynamic model compression methods by a large margin (up to 6.6%). Typically, we achieves 2-4× computation reduction and up to 61.5% real-world acceleration on MobileNet, ResNet-50 and Vision Transformer, with minimal accuracy drops on ImageNet. Code release: https://github.com/changlin31/DS-Net. Guangrun Wang, Bing Wang 0013, Xiaodan Liang, Zhihui Li 0001, Xiaojun Chang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | RoReg: Pairwise Point Cloud Registration With Oriented Descriptors and Local RotationsabstractWe present RoReg, a novel point cloud registration framework that fully exploits oriented descriptors and estimated local rotations in the whole registration pipeline. Previous methods mainly focus on extracting rotation-invariant descriptors for registration but unanimously neglect the orientations of descriptors. In this paper, we show that the oriented descriptors and the estimated local rotations are very useful in the whole registration pipeline, including feature description, feature detection, feature matching, and transformation estimation. Consequently, we design a novel oriented descriptor RoReg-Desc and apply RoReg-Desc to estimate the local rotations. Such estimated local rotations enable us to develop a rotation-guided detector, a rotation coherence matcher, and a one-shot-estimation RANSAC, all of which greatly improve the registration performance. Extensive experiments demonstrate that RoReg achieves state-of-the-art performance on the widely-used 3DMatch and 3DLoMatch datasets, and also generalizes well to the outdoor ETH dataset. In particular, we also provide in-depth analysis on each component of RoReg, validating the improvements brought by oriented descriptors and the estimated local rotations. Source code and supplementary material are available at https://github.com/HpWang-whu/RoReg. Haiping Wang 0004, Yuan Liu 0025, Qingyong Hu, Bing Wang 0013, Zhen Dong 0005, Yulan Guo, Wenping Wang 0001, Bisheng Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Knowledge Distillation via the Target-aware TransformerabstractKnowledge distillation becomes a de facto standard to improve the performance of small neural networks. Most of the previous works propose to regress the representational features from the teacher to the student in a one-to-one spatial matching fashion. However, people tend to overlook the fact that, due to the architecture differences, the semantic information on the same spatial location usually vary. This greatly undermines the underlying assumption of the one-to-one distillation approach. To this end, we propose a novel one-to-all spatial matching knowledge distillation approach. Specifically, we allow each pixel of the teacher feature to be distilled to all spatial locations of the student features given its similarity, which is generated from a target-aware transformer. Our approach surpasses the state-of-the-art methods by a significant margin on various computer vision benchmarks, such as ImageNet, Pascal VOC and COCOStuff10k. Code is available at https://github.com/sihaoevery/TaT. Sihao Lin, Hongwei Xie, Bing Wang 0013, Kaicheng Yu, Xiaojun Chang, Xiaodan Liang |
CVPR | 3 |
| 2022 | No Pain, Big Gain: Classify Dynamic Point Cloud Sequences with Static Models by Fitting Feature-level Space-time SurfacesabstractScene flow is a powerful tool for capturing the motion field of 3D point clouds. However, it is difficult to directly apply flow-based models to dynamic point cloud classification since the unstructured points make it hard or even impossible to efficiently and effectively trace point-wise correspondences. To capture 3D motions without explicitly tracking correspondences, we propose a kinematics-inspired neural network (Kinet) by generalizing the kinematic concept of ST-surfaces to the feature space. By unrolling the normal solver of ST-surfaces in the feature space, Kinet implicitly encodes feature-level dynamics and gains advantages from the use of mature back-bones for static point cloud processing. With only minor changes in network structures and low computing overhead, it is painless to jointly train and deploy our framework with a given static model. Experiments on NvGesture, SHREC'17, MSRAction-3D, and NTU-RGBD demonstrate its efficacy in performance, efficiency in both the number of parameters and computational complexity, as well as its versatility to various static backbones. Noticeably, Kinet achieves the accuracy of 93.27% on MSRAction-3D with only 3.20M parameters and 10.35G FLOPS. The code is available at https://github.com/jx-zhong-for-academic-purpose/Kinet. Jia-Xing Zhong, Kaichen Zhou, Qingyong Hu, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham |
CVPR | 4 |
| 2022 | AutoPlace: Robust Place Recognition with Single-chip Automotive RadarabstractThis paper presents a novel place recognition approach to autonomous vehicles by using low-cost, single-chip automotive radar. Aimed at improving recognition robustness and fully exploiting the rich information provided by this emerging automotive radar, our approach follows a principled pipeline that comprises (1) dynamic points removal from instant Doppler measurement, (2) spatial-temporal feature embedding on radar point clouds, and (3) retrieved candidates refinement from Radar Cross Section measurement. Extensive experimental results on the public nuScenes dataset demonstrate that existing visual/LiDAR/spinning radar place recognition approaches are less suitable for single-chip automotive radar. In contrast, our purpose-built approach for automotive radar consistently outperforms a variety of baseline methods via a comprehensive set of metrics, providing insights into the efficacy when used in a realistic system. Kaiwen Cai, Bing Wang 0013, Xiaoxuan Lu 0001 |
ICRA | 2 |
| 2022 | Graph-Based Thermal-Inertial SLAM With Probabilistic Neural NetworksabstractSimultaneous localization and mapping (SLAM) system typically employs vision-based sensors to observe the surrounding environment. However, the performance of such systems highly depends on the ambient illumination conditions. In scenarios with adverse visibility or in the presence of airborne particulates (e.g., smoke, dust, etc.), alternative modalities such as those based on thermal imaging and inertial sensors are more promising. In this article, we propose the first complete thermal–inertial SLAM system that combines neural abstraction in the SLAM front end with robust pose-graph optimization in the SLAM back end. We model the sensor abstraction in the front end by employing probabilistic deep learning parameterized by mixture density networks (MDNs). Our key strategies to successfully model this encoding from thermal imagery are the usage of normalized 14-b radiometric data, the incorporation of hallucinated visual (RGB) features, and the inclusion of feature selection to estimate the MDN parameters. To enable a full SLAM system, we also design an efficient global image descriptor that is able to detect loop closures from thermal embedding vectors. We performed extensive experiments and analysis using three datasets, namely self-collected ground robot and hand-held data taken in indoor environment, and one public dataset (SubT-tunnel) collected in underground tunnel. Finally, we demonstrate that an accurate thermal–inertial SLAM system can be realized in conditions of both benign and adverse visibility. Muhamad Risqi Utama Saputra, Xiaoxuan Lu 0001, Pedro Porto Buarque de Gusmão, Bing Wang 0013, Andrew Markham, Agathoniki Trigoni |
IEEE Trans. Robotics | 4 |
| 2021 | VMLoc: Variational Fusion For Learning-Based Multimodal Camera LocalizationabstractRecent learning-based approaches have achieved impressive results in the field of single-shot camera localization. However, how best to fuse multiple modalities (e.g., image and depth) and to deal with degraded or missing input are less well studied. In particular, we note that previous approaches towards deep fusion do not perform significantly better than models employing a single modality. We conjecture that this is because of the naive approaches to feature space fusion through summation or concatenation which do not take into account the different strengths of each modality. To address this, we propose an end-to-end framework, termed VMLoc, to fuse different sensor inputs into a common latent space through a variational Product-of-Experts (PoE) followed by attention-based fusion. Unlike previous multimodal variational works directly adapting the objective function of vanilla variational auto-encoder, we show how camera localization can be accurately estimated through an unbiased objective function based on importance weighting. Our model is extensively evaluated on RGB-D datasets and the results prove the efficacy of our model. The source code is available at https://github.com/Zalex97/VMLoc. Kaichen Zhou, Changhao Chen, Bing Wang 0013, Muhamad Risqi Utama Saputra, Agathoniki Trigoni, Andrew Markham |
AAAI | 3 |
| 2021 | Dynamic Slimmable NetworkabstractCurrent dynamic networks and dynamic pruning methods have shown their promising capability in reducing theoretical computation complexity. However, dynamic sparse patterns on convolutional filters fail to achieve actual acceleration in real-world implementation, due to the extra burden of indexing, weight-copying, or zero-masking. Here, we explore a dynamic network slimming regime, named Dynamic Slimmable Network (DS-Net), which aims to achieve good hardware-efficiency via dynamically adjusting filter numbers of networks at test time with respect to different inputs, while keeping filters stored statically and contiguously in hardware to prevent the extra burden. Our DS-Net is empowered with the ability of dynamic inference by the proposed double-headed dynamic gate that comprises an attention head and a slimming head to predictively adjust network width with negligible extra computation cost. To ensure generality of each candidate architecture and the fairness of gate, we propose a disentangled two-stage training scheme inspired by one-shot NAS. In the first stage, a novel training technique for weight-sharing networks named In-place Ensemble Bootstrapping is proposed to improve the supernet training efficacy. In the second stage, Sandwich Gate Sparsification is proposed to assist the gate training by identifying easy and hard samples in an online way. Extensive experiments demonstrate our DS-Net consistently outperforms its static counterparts as well as state-of-the-art static and dynamic model compression methods by a large margin (up to 5.9%). Typically, DS-Net achieves 2-4× computation reduction and 1.62× real-world acceleration over ResNet-50 and MobileNet with minimal accuracy drops on ImageNet.1 Guangrun Wang, Bing Wang 0013, Xiaodan Liang, Zhihui Li 0001, Xiaojun Chang |
CVPR | 3 |
| 2021 | P2-Net: Joint Description and Detection of Local Features for Pixel and Point MatchingabstractAccurately describing and detecting 2D and 3D key-points is crucial to establishing correspondences across images and point clouds. Despite a plethora of learning-based 2D or 3D local feature descriptors and detectors having been proposed, the derivation of a shared descriptor and joint keypoint detector that directly matches pixels and points remains under-explored by the community. This work takes the initiative to establish fine-grained correspondences between 2D images and 3D point clouds. In order to directly match pixels and points, a dual fully-convolutional framework is presented that maps 2D and 3D inputs into a shared latent representation space to simultaneously describe and detect keypoints. Furthermore, an ultra-wide reception mechanism and a novel loss function are designed to mitigate the intrinsic information variations between pixel and point local regions. Extensive experimental results demonstrate that our framework shows competitive performance in fine-grained matching between images and point clouds and achieves state-of-the-art results for the task of indoor visual localization. Our source code is available at https://github.com/BingCS/P2-Net. Bing Wang 0013, Changhao Chen, Zhaopeng Cui, Jie Qin 0004, Xiaoxuan Lu 0001, Zhengdi Yu, Peijun Zhao, Zhen Dong 0005, Fan Zhu 0001, Agathoniki Trigoni, Andrew Markham |
ICCV | 1 |
| 2021 | BossNAS: Exploring Hybrid CNN-transformers with Block-wisely Self-supervised Neural Architecture SearchabstractA myriad of recent breakthroughs in hand-crafted neural architectures for visual recognition have highlighted the urgent need to explore hybrid architectures consisting of diversified building blocks. Meanwhile, neural architecture search methods are surging with an expectation to reduce human efforts. However, whether NAS methods can efficiently and effectively handle diversified search spaces with disparate candidates (e.g. CNNs and transformers) is still an open question. In this work, we present Block-wisely Self-supervised Neural Architecture Search (BossNAS), an unsupervised NAS method that addresses the problem of in-accurate architecture rating caused by large weight-sharing space and biased supervision in previous methods. More specifically, we factorize the search space into blocks and utilize a novel self-supervised training scheme, named ensemble bootstrapping, to train each block separately before searching them as a whole towards the population center. Additionally, we present HyTra search space, a fabric-like hybrid CNN-transformer search space with searchable down-sampling positions. On this challenging search space, our searched model, BossNet-T, achieves up to 82.5% accuracy on ImageNet, surpassing EfficientNet by 2.4% with comparable compute time. Moreover, our method achieves superior architecture rating accuracy with 0.78 and 0.76 Spearman correlation on the canonical MBConv search space with ImageNet and on NATS-Bench size search space with CIFAR-100, respectively, surpassing state-of-the-art NAS methods.1 Guangrun Wang, Jiefeng Peng, Bing Wang 0013, Xiaodan Liang, Xiaojun Chang |
ICCV | 5 |
| 2021 | Exploring Inter-Channel Correlation for Diversity-preserved Knowledge DistillationabstractKnowledge Distillation has shown very promising ability in transferring learned representation from the larger model (teacher) to the smaller one (student). Despite many efforts, prior methods ignore the important role of retaining inter-channel correlation of features, leading to the lack of capturing intrinsic distribution of the feature space and sufficient diversity properties of features in the teacher network. To solve the issue, we propose the novel Inter-Channel Correlation for Knowledge Distillation (ICKD), with which the diversity and homology of the feature space of the student network can align with that of the teacher network. The correlation between these two channels is interpreted as diversity if they are irrelevant to each other, otherwise homology. Then the student is required to mimic the correlation within its own embedding space. In addition, we introduce the grid-level inter-channel correlation, making it capable of dense prediction tasks. Extensive experiments on two vision tasks, including ImageNet classification and Pascal VOC segmentation, demonstrate the superiority of our ICKD, which consistently outperforms many existing methods, advancing the state-of-the-art in the fields of Knowledge Distillation. To our knowledge, we are the first method based on knowledge distillation boosts ResNet18 beyond 72% Top-1 accuracy on ImageNet classification. Code is available at: https://github.com/ADLab-AutoDrive/ICKD. Li Liu 0069, Qingle Huang, Sihao Lin, Hongwei Xie, Bing Wang 0013, Xiaojun Chang, Xiaodan Liang |
ICCV | 5 |
| 2021 | 3D Motion Capture of an Unmodified Drone with Single-chip Millimeter Wave RadarabstractAccurate motion capture of aerial robots in 3D is a key enabler for autonomous operation in indoor environments such as warehouses or factories, as well as driving forward research in these areas. The most commonly used solutions at present are optical motion capture (e.g. VICON) and Ultrawide-band (UWB), but these are costly and cumbersome to deploy, due to their requirement of multiple cameras/anchors spaced around the tracking area. They also require the drone to be modified to carry an active or passive marker. In this work, we present an inexpensive system that can be rapidly installed, based on single-chip millimeter wave (mmWave) radar. Importantly, the drone does not need to be modified or equipped with any markers, as we exploit the Doppler signals from the rotating propellers. Furthermore, 3D tracking is possible from a single point, greatly simplifying deployment. We develop a novel deep neural network and demonstrate decimeter level 3D tracking at 10Hz, achieving better performance than classical baselines. Our hope is that this low-cost system will act to catalyse inexpensive drone research and increased autonomy. Peijun Zhao, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham |
ICRA | 3 |
| 2021 | DynaNet: Neural Kalman Dynamical Model for Motion Estimation and PredictionabstractDynamical models estimate and predict the temporal evolution of physical systems. State-space models (SSMs) in particular represent the system dynamics with many desirable properties, such as being able to model uncertainty in both the model and measurements, and optimal (in the Bayesian sense) recursive formulations, e.g., the Kalman filter. However, they require significant domain knowledge to derive the parametric form and considerable hand tuning to correctly set all the parameters. Data-driven techniques, e.g., recurrent neural networks, have emerged as compelling alternatives to SSMs with wide success across a number of challenging tasks, in part due to their impressive capability to extract relevant features from rich inputs. They, however, lack interpretability and robustness to unseen conditions. Thus, data-driven models are hard to be applied in safety-critical applications, such as self-driving vehicles. In this work, we present DynaNet, a hybrid deep learning and time-varying SSM, which can be trained end-to-end. Our neural Kalman dynamical model allows us to exploit the relative merits of both SSM and deep neural networks. We demonstrate its effectiveness in the estimation and prediction on a number of physically challenging tasks, including visual odometry, sensor fusion for visual-inertial navigation, and motion prediction. In addition, we show how DynaNet can indicate failures through investigation of properties, such as the rate of innovation (Kalman gain). Changhao Chen, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | AtLoc: Attention Guided Camera LocalizationabstractDeep learning has achieved impressive results in camera localization, but current single-image techniques typically suffer from a lack of robustness, leading to large outliers. To some extent, this has been tackled by sequential (multi-images) or geometry constraint approaches, which can learn to reject dynamic objects and illumination conditions to achieve better performance. In this work, we show that attention can be used to force the network to focus on more geometrically robust objects and features, achieving state-of-the-art performance in common benchmark, even if using only a single image as input. Extensive experimental evidence is provided through public indoor and outdoor datasets. Through visualization of the saliency maps, we demonstrate how the network learns to reject dynamic objects, yielding superior global camera pose regression performance. The source code is avaliable at https://github.com/BingCS/AtLoc. Bing Wang 0013, Changhao Chen, Xiaoxuan Lu 0001, Peijun Zhao, Agathoniki Trigoni, Andrew Markham |
AAAI | 1 |
| 2020 | Heart Rate Sensing with a Robot Mounted mmWave RadarabstractHeart rate monitoring at home is a useful metric for assessing health e.g. of the elderly or patients in post-operative recovery. Although non-contact heart rate monitoring has been widely explored, typically using a static, wall-mounted device, measurements are limited to a single room and sensitive to user orientation and position. In this work, we propose mBeats, a robot mounted millimeter wave (mmWave) radar system that provide periodic heart rate measurements under different user poses, without interfering in a users daily activities. mBeats contains a mmWave servoing module that adaptively adjusts the sensor angle to the best reflection pro le. Furthermore, mBeats features a deep neural network predictor, which can estimate heart rate from the lower leg and additionally provides estimation uncertainty. Through extensive experiments, we demonstrate accurate and robust operation of mBeats in a range of scenarios. We believe by integrating mobility and adaptability, mBeats can empower many down-stream healthcare applications at home, such as palliative care, post-operative rehabilitation and telemedicine. Peijun Zhao, Xiaoxuan Lu 0001, Bing Wang 0013, Changhao Chen, Linhai Xie, Agathoniki Trigoni, Andrew Markham |
ICRA | 3 |
| 2020 | See through smoke: robust indoor mapping with low-cost mmWave radarabstractThis paper presents the design, implementation and evaluation of milliMap, a single-chip millimetre wave (mmWave) radar based indoor mapping system targetted towards low-visibility environments to assist in emergency response. A unique feature of milliMap is that it only leverages a low-cost, off-the-shelf mmWave radar, but can reconstruct a dense grid map with accuracy comparable to lidar, as well as providing semantic annotations of objects on the map. milliMap makes two key technical contributions. First, it autonomously overcomes the sparsity and multi-path noise of mmWave signals by combining cross-modal supervision from a co-located lidar during training and the strong geometric priors of indoor spaces. Second, it takes the spectral response of mmWave reflections as features to robustly identify different types of objects e.g. doors, walls etc. Extensive experiments in different indoor environments show that milliMap can achieve a map reconstruction error less than 0.2m and classify key semantics with an accuracy of ~ 90%, whilst operating through dense smoke. Xiaoxuan Lu 0001, Stefano Rosa, Peijun Zhao, Bing Wang 0013, Changhao Chen, John A. Stankovic, Agathoniki Trigoni, Andrew Markham |
MobiSys | 4 |