Danyang Li 0005

dblp:172/7877-5 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
22since 2021 · last 2026
0000-0002-7527-4669ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 19 · 8 first-author · 17 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 E-Cube: Event Enhanced Efficient Video Streaming for Drones
abstract
This paper presents E-Cube, a framework designed to enhance the efficiency of mobile video streaming systems. Our core insight is that content correlations among frames, which are critical for video streaming but challenging to extract in dynamic scenes, are often already available as intermediate outputs from other mobile computing subsystems. By adopting a cross-subsystem design, E-Cube repurposes these intermediate results obtained from event cameras to reduce computation, energy consumption, and bandwidth usage of existing video streaming systems, while simultaneously improving video quality. We implement E-Cube on a drone-embedded chip through software-hardware co-design, and plugged E-Cube into three prevalent drone video streaming systems for industrial inspection. Evaluations in one of the world's largest oil fields and on public datasets demonstrate E-Cube can save over 35% network bandwidth overhead and reduce drone video streaming energy consumption by over 12% for H.265 and AV1, and 47% for advanced H.266, all while achieving improved video quality.
Jingao Xu, Longfei Shangguan, Danyang Li 0005, Yunhao Liu 0001, Zheng Yang 0002
EuroSys3
2026 SpecOffload: Unlocking Latent GPU Capacity for LLM Inference on Resource-Constrained Devices
abstract
Efficient LLM inference on resource-constrained devices presents significant challenges in compute and memory utilization. Due to limited GPU memory, existing systems offload model weights to CPU memory, incurring substantial I/O overhead between the CPU and GPU. This leads to two major inefficiencies: (1) GPU cores are underutilized, often remaining idle while waiting for data to be loaded; and (2) GPU memory has low impact on performance, as reducing its capacity has minimal effect on overall throughput.In this paper, we propose SpecOffload, a high-throughput inference engine that embeds speculative decoding into offloading. Our key idea is to unlock latent GPU resources for storing and executing a draft model used for speculative decoding, thus accelerating inference at near-zero additional cost. To support this, we carefully orchestrate the interleaved execution of target and draft models in speculative decoding within the offloading pipeline, and propose a planner to manage tensor placement and select optimal parameters. Compared to the best baseline, SpecOffload improves GPU core utilization by 4.49x and boosts inference throughput by 2.54x. Our code is available at https://github.com/MobiSense/SpecOffload-public .
Xiangwen Zhuge, Fan Dang 0001, Danyang Li 0005, Tianxiang Hao 0001, Qiang Ma 0007, Yahui Han, Zheng Yang 0002
IWQoS5
2026 Edge-Assisted Real-Time Motion Capture System
Chen Qian 0009, Danyang Li 0005, Jingao Xu, Zheng Yang 0002, Qiang Ma 0007
IEEE Trans. Mob. Comput.2
2025 OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
abstract
Grounding natural language instructions to visual observations is fundamental for embodied agents operating in open-world environments. Recent advances in visual-language mapping have enabled generalizable semantic representations by leveraging visionlanguage models (VLMs). However, these methods often fall short in aligning free-form language commands with specific scene instances, due to limitations in both instance-level semantic consistency and instruction interpretation. We present OpenMap, a zero-shot open-vocabulary visual-language map designed for accurate instruction grounding in navigation tasks. To address semantic inconsistencies across views, we introduce a Structural-Semantic Consensus constraint that jointly considers global geometric structure and vision-language similarity to guide robust 3D instancelevel aggregation. To improve instruction interpretation, we propose an LLM-assisted Instruction-to-Instance Grounding module that enables fine-grained instance selection by incorporating spatial context and expressive target descriptions. We evaluate OpenMap on ScanNet200 and Matterport3D, covering both semantic mapping and instruction-to-target retrieval tasks. Experimental results show that OpenMap outperforms state-of-the-art baselines in zero-shot settings, demonstrating the effectiveness of our method in bridging free-form language and 3D perception for embodied navigation.
Danyang Li 0005, Zenghui Yang, Guangpeng Qi, Songtao Pang, Guangyong Shang, Qiang Ma 0007, Zheng Yang 0002
ACM Multimedia1
2025 OpenMoCap: Rethinking Optical Motion Capture under Real-world Occlusion
abstract
Optical motion capture is a foundational technology driving advancements in cutting-edge fields such as virtual reality and film production. However, system performance suffers severely under large-scale marker occlusions common in real-world applications. An in-depth analysis identifies two primary limitations of current models: (i) the lack of training datasets accurately reflecting realistic marker occlusion patterns, and (ii) the absence of training strategies designed to capture long-range dependencies among markers. To tackle these challenges, we introduce the CMU-Occlu dataset, which incorporates ray tracing techniques to realistically simulate practical marker occlusion patterns. Furthermore, we propose OpenMoCap, a novel motion-solving model designed specifically for robust motion capture in environments with significant occlusions. Leveraging a marker-joint chain inference mechanism, OpenMoCap enables simultaneous optimization and construction of deep constraints between markers and joints. Extensive comparative experiments demonstrate that OpenMoCap consistently outperforms competing methods across diverse scenarios, while the CMU-Occlu dataset opens the door for future studies in robust motion solving. The proposed OpenMoCap is integrated into the MoSen MoCap system for practical deployment. The code is released at: https://github.com/qianchen214/OpenMoCap.
Chen Qian 0009, Danyang Li 0005, Xinran Yu, Zheng Yang 0002, Qiang Ma 0007
ACM Multimedia2
2025 Taming Event Cameras With Bio-Inspired Architecture and Algorithm: A Case for Drone Obstacle Avoidance
abstract
Fast and accurate obstacle avoidance is crucial to drone safety. Yet existing on-board sensor modules such as frame cameras and radars are ill-suited for doing so due to their low temporal resolution or limited field of view. This paper presentsBioDrone, a new design paradigm for drone obstacle avoidance using stereo event cameras. At the heart of BioDrone are three simple yet effective system designs inspired by the mammalian visual system, namely, a chiasm-inspired event filtering, a lateral geniculate nucleus (LGN)-inspired event matching, and a dorsal stream-inspired obstacle tracking. We implement BioDrone on FPGA through software-hardware co-design and deploy it on an industrial drone. In comparative experiments against two state-of-the-art event-based systems, BioDrone consistently achieves an obstacle detection rate of$> $90%, and an obstacle tracking error of$<$5.8 cm across all flight modes with an end-to-end latency of$<$6.4 ms, outperforming both baselines by over 44%.
Danyang Li 0005, Jingao Xu, Zheng Yang 0002, Yishujie Zhao, Yunhao Liu 0001, Longfei Shangguan
IEEE Trans. Mob. Comput.1
2024 EventBoost: Event-based Acceleration Platform for Real-time Drone Localization and Tracking
abstract
Drones have demonstrated their pivotal role in various applications such as search-and-rescue, smart logistics, and industrial inspection, with accurate localization playing an indispensable part. However, in high dynamic range and rapid motion scenarios, traditional visual sensors often face challenges in pose estimation. Event cameras, with their high temporal resolution, present a fresh opportunity for perception in such challenging environments. Current efforts resort to event-visual fusion to enhance the drone’s sensing capability. Yet, the lack of efficient event-visual fusion algorithms and corresponding acceleration hardware causes the potential of event cameras to remain underutilized. In this paper, we introduce EventBoost, an acceleration platform designed for drone-based applications with event-image fusion. We propose a suit of novel algorithms through software-hardware co-design on Zynq SoC, aimed at enhancing real-time localization precision and speed. EventBoost achieves enhanced visual fusion precision and markedly elevated processing efficiency. The performance comparison with two state-of-the-art systems shows EventBoost achieves 24.33% improvement in accuracy with 30 ms latency on resource-constrained platforms.
Jingao Xu, Danyang Li 0005, Zheng Yang 0002, Yunhao Liu 0001
INFOCOM3
2024 edgeSLAM2: Rethinking Edge-Assisted Visual SLAM with On-Chip Intelligence
abstract
Edge-assisted visual SLAM stands as a pivotal enabler for emerging mobile applications, such as search-and-rescue, smart logistics, and industrial inspection. Limited by the computing capability of lightweight mobile devices like MAVs, current innovations balance system accuracy and efficiency by allocating lightweight and time-sensitive tracking tasks to mobile devices, while offloading the more resource-intensive yet delay-tolerant map optimization tasks to the edge. However, our pilot study in a large-scale oil field reveals several limitations of such a tracking-optimization decoupled paradigm, arising due to the disruption of inter-dependencies between the two tasks concerning data, resources, and threads.In this paper, we design and implement edgeSLAM2, an innovative system that reshapes the edge-assisted visual SLAM paradigm by tightly integrating tracking and partial-yet-crucial optimization on mobile. edgeSLAM2 harnesses the hierarchical and heterogeneous computing units offered by the latest commercial systems-on-chip (SoCs) to enhance the computational capacity of mobile devices, which in turn, allows edgeSLAM2 to design a suit of novel algorithms for map sync, optimization, and tracking that accommodate such architectural upgrade. By fully embracing the on-chip intelligence, edgeSLAM2 simultaneously enhances system accuracy and efficiency through software-hardware co-design. We deploy edgeSLAM2 on an industrial drone and conduct comprehensive experiments in a large-scale oil field over three months. The results show that edgeSLAM2 surpasses comparative methods by achieving an 80% reduction in bandwidth consumption, a 32% improvement in accuracy, and a 26% reduction in tracking delay.
Danyang Li 0005, Yishujie Zhao, Jingao Xu, Shengkai Zhang, Longfei Shangguan, Zheng Yang 0002
INFOCOM1
2024 LeoVR: Motion-Inspired Visual-LiDAR Fusion for Environment Depth Estimation
abstract
Environment depth estimation by fusing camera and radar enables a broad spectrum of applications such as autonomous driving, environmental perception, context-aware localization and navigation. Various pioneering approaches have been proposed to achieve accurate and dense depth estimation by integrating vision and LiDAR through deep learning. However, due to the challenges of sparse sampling of in-vehicle LiDARs, high ground-truth annotation overhead, and severe dynamics in real environments, existing solutions have not yet achieved widespread deployment on commercial autonomous vehicles. In this paper, we propose LeoVR, a motion-inspired self-supervised visual-LiDAR fusion approach that enables accurate environment depth estimation. Leveraging the vehicle motion information, LeoVR employs two effective system frameworks to$(i)$optimize the depth estimation results, and$(ii)$provide supervision signals for DNN training. We fully implemented LeoVR on both a robotic testbed and a commercial vehicle and conducted extensive experiments over an 8-month period. The results demonstrate that LeoVR achieves remarkable performance with an average depth estimation error of 0.17$m$, outperforming existing state-of-the-art solutions by$\gt $45.9%. Besides, even cold-start in real environments by self-supervised training, LeoVR still achieves an average error of 0.2$m$, outperforming the related works by$\gt $47.8% and comparable to supervised training methods.
Danyang Li 0005, Jingao Xu, Zheng Yang 0002, Qiang Ma 0007, Li Zhang 0028
IEEE Trans. Mob. Comput.1
2024 Reshaping Edge-Assisted Visual SLAM by Embracing On-Chip Intelligence
abstract
Edge-assisted visual SLAM plays a crucial role in enabling innovative mobile applications, such as autonomous swarm inspection, search-and-rescue, and smart logistics. Constrained by the computational capacities of lightweight mobile devices, current approaches delegate lightweight, time-sensitive tracking tasks to the mobile end while offloading resource-intensive, latency-tolerant map optimization tasks to the edge. However, our pilot study reveals several limitations of the tracking-optimization decoupled paradigm, stemming from the disruption of inter-dependencies between the two tasks. In this paper, we design and implement edgeSLAM2, an innovative system that reshapes the edge-assisted visual SLAM paradigm by tightly integrating tracking and partial-yet-crucial optimization on mobile. edgeSLAM2 harnesses the heterogeneous computing units offered by the commercial systems-on-chip (SoCs) to enhance the computational capacity of mobile devices, which in turn, allows edgeSLAM2 to design a suit of novel algorithms for map sync, optimization, and tracking that accommodate such architectural upgrade. By capitalizing on the full potential of on-chip intelligence, edgeSLAM2 supports both solitary and collaborative SLAM with accuracy and immediacy, underpinned by a cohesive software-hardware co-design. We deploy edgeSLAM2 on drones for industrial inspection. Comprehensive experiments in one of the world’s largest oil fields over three months demonstrate its superior performance.
Danyang Li 0005, Yishujie Zhao, Jingao Xu, Shengkai Zhang, Longfei Shangguan, Qiang Ma 0007, Zheng Yang 0002
IEEE Trans. Mob. Comput.1
2024 Train Once, Locate Anytime for Anyone: Adversarial Learning-based Wireless Localization
abstract
Among numerous indoor localization systems, WiFi fingerprint-based localization has been one of the most attractive solutions, which is known to be free of extra infrastructure and specialized hardware. To push forward this approach for wide deployment, three crucial goals on high deployment ubiquity, high localization accuracy, and low maintenance cost are desirable. However, due to severe challenges about signal variation, device heterogeneity, and database degradation root in environmental dynamics, pioneer works usually make a trade-off among them. In this article, we propose iToLoc, a deep learning-based localization system that achieves all three goals simultaneously. Once trained, iToLoc will provide accurate localization service for everyone using different devices and under diverse network conditions, and automatically update itself to maintain reliable performance anytime. iToLoc is purely based on WiFi fingerprints without relying on specific infrastructures. The core components of iToLoc are a domain adversarial neural network and a co-training-based semi-supervised learning framework. Extensive experiments across 7 months with eight different devices demonstrate that iToLoc achieves remarkable performance with an accuracy of 1.92 m and >95% localization success rate. Even 7 months after the original fingerprint database was established, the rate still maintains >90%, which significantly outperforms previous works.
Danyang Li 0005, Jingao Xu, Zheng Yang 0002, Chengpei Tang
ACM Trans. Sens. Networks1
2024 HCCNet: Hybrid Coupled Cooperative Network for Robust Indoor Localization
abstract
Accurate localization of unmanned aerial vehicle (UAV) is critical for navigation in GPS-denied regions, which remains a highly challenging topic in recent research. This article describes a novel approach to multi-sensor hybrid coupled cooperative localization network (HCCNet) system that combines multiple types of sensors including camera, ultra-wideband (UWB), and inertial measurement unit (IMU) to address this challenge. The camera and IMU can automatically determine the position of UAV based on the perception of surrounding environments and their own measurement data. The UWB node and the UWB wireless sensor network (WSN) in indoor environments jointly determine the global position of UAV, and the proposed dynamic random sample consensus (D-RANSAC) algorithm can optimize UWB localization accuracy. To fully exploit UWB localization results, we provide an HCCNet system which combines the local pose estimator of visual inertial odometry (VIO) system with global constraints from UWB localization results. Experimental results show that the proposed D-RANSAC algorithm can achieve better accuracy than other UWB-based algorithms. The effectiveness of the proposed HCCNet method is verified by a mobile robot in real world and some simulation experiments in indoor environments.
Li Zhang 0028, Danyang Li 0005, Zheng Yang 0002
ACM Trans. Sens. Networks3
2023 FlyTracker: Motion Tracking and Obstacle Detection for Drones Using Event Cameras
abstract
Location awareness in environments is one of the key parts for drones’ applications and have been explored through various visual sensors. However, standard cameras easily suffer from motion blur under high moving speeds and low-quality image under poor illumination, which brings challenges for drones to perform motion tracking. Recently, a kind of bio-inspired sensors called event cameras emerge, offering advantages like high temporal resolution, high dynamic range and low latency, which motivate us to explore their potential to perform motion tracking in limited scenarios. In this paper, we propose FlyTracker, aiming at developing visual sensing ability for drones of both individual and circumambient location-relevant contextual, by using a monocular event camera. In FlyTracker, background-subtraction-based method is proposed to distinguish moving objects from background and fusion-based photometric features are carefully designed to obtain motion information. Through multilevel fusion of events and images, which are heterogeneous visual data, FlyTracker can effectively and reliably track the 6-DoF pose of the drone as well as monitor relative positions of moving obstacles. We evaluate performance of FlyTracker in different environments and the results show that FlyTracker is more accurate than the state-of-the-art baselines.
Yue Wu 0030, Jingao Xu, Danyang Li 0005, Yadong Xie, Fan Li 0001, Zheng Yang 0002
INFOCOM3
2023 Taming Event Cameras with Bio-Inspired Architecture and Algorithm: A Case for Drone Obstacle Avoidance
abstract
Fast and accurate obstacle avoidance is crucial to drone safety. Yet existing on-board sensor modules such as frame cameras and radars are ill-suited for doing so due to their low temporal resolution or limited field of view. This paper presents BioDrone, a new design paradigm for drone obstacle avoidance using stereo event cameras. At the heart of BioDrone is two simple yet effective system design inspired by the mammalian visual system, namely, a chiasm-inspired signal processing pipeline for fast event filtering and obstacle detection, and a lateral geniculate nucleus (LGN)-inspired event matching algorithm for accurate obstacle localization. To make BioDrone a practical solution, we further take significant engineering efforts to deploy the software stack on FPGA through software and hardware co-design. The performance comparison with two state-of-the-art event-based obstacle avoidance systems shows BioDrone achieves a consistently high obstacle detection rate of 96.1%. The average localization error of BioDrone is 6.8cm with a 4.7ms latency, outperforming both baselines by over 40%.
Jingao Xu, Danyang Li 0005, Zheng Yang 0002, Yishujie Zhao, Yunhao Liu 0001, Longfei Shangguan
MobiCom2
2023 Edge Assisted Mobile Semantic Visual SLAM
abstract
Localization and navigation play a key role in many location-based services and have attracted numerous research efforts. In recent years, visual SLAM has been prevailing for autonomous driving. However, the ever-growing computation resources demanded by SLAM impede its applications to resource-constrained mobile devices. In this paper, we present the design, implementation, and evaluation ofedgeSLAM, an edge-assisted real-time semantic visual SLAM service running on mobile devices.edgeSLAMleverages the state-of-the-art semantic segmentation algorithm to enhance localization and mapping accuracy, and speeds up the computation-intensive SLAM and semantic segmentation algorithms by computation offloading. The key innovations ofedgeSLAMinclude an efficient computation offloading strategy, an opportunistic data sharing method, an adaptive task scheduling algorithm, and a multi-user support mechanism. We fully implementedgeSLAMand plan to open-source it. Extensive experiments are conducted under 3 datasets. The results show thatedgeSLAMcan run on mobile devices at 35fps and achieve 5cm localization accuracy from real-world experiments, outperforming existing solutions by more than 15%. We also demonstrate the usability ofedgeSLAMthrough 2 case studies of pedestrian localization and robot navigation. To the best of our knowledge,edgeSLAMis the first edge-assisted real-time semantic visual SLAM for mobile devices.
Jingao Xu, Danyang Li 0005, Longfei Shangguan, Yunhao Liu 0001, Zheng Yang 0002
IEEE Trans. Mob. Comput.3
2022 UWB/IMU Fusion Localization Strategy Based on Continuity of Movement
Li Zhang 0028, Jinhui Bao, Jingao Xu, Danyang Li 0005
MobiQuitous4
2022 Motion inspires notion: self-supervised visual-LiDAR fusion for environment depth estimation
abstract
Environment depth estimation by fusing camera and radar enables a broad spectrum of applications such as autonomous driving, environmental perception, context-aware localization and navigation. Various pioneering approaches have been proposed to achieve accurate and dense depth estimation by integrating vision and LiDAR through deep learning. However, due to the challenges of sparse sampling of in-vehicle LiDARs, high ground-truth annotation overhead, and severe dynamics in real environments, existing solutions have not yet achieved widespread deployment on commercial autonomous vehicles. In this paper, we propose LeoVR, a visual-LiDAR fusion based self-supervised approach that enables accurate environment depth estimation. LeoVR digs into the vehicle's motion information and designs two effective system frameworks based on it to (i) optimize the depth estimation results, and (ii) provide supervision signals to train a DNN. We fully implement LeoVR on a robotic testbed and commercial vehicle to conduct extensive experiments across 6 months. The results demonstrate that LeoVR achieves remarkable performance with an average depth estimation error of 0.17m, outperforming existing state-of-the-art solutions by > 43%. Besides, even cold-start in real environments by self-supervised training, LeoVR still achieves an average error of 0.21m, outperforming the related works by > 45% and comparable to those supervised training methods.
Danyang Li 0005, Jingao Xu, Zheng Yang 0002, Qian Zhang 0017, Qiang Ma 0007, Li Zhang 0028
MobiSys1
2022 Wireless Localization with Spatial-Temporal Robust Fingerprints
abstract
Indoor localization has gained increasing attention in the era of the Internet of Things. Among various technologies, WiFi fingerprint-based localization has become a mainstream solution. However, RSS fingerprints suffer from critical drawbacks of spatial ambiguity and temporal instability that root in multipath effects and environmental dynamics, which degrade the performance of these systems and therefore impede their wide deployment in the real world. Pioneering works overcome these limitations at the costs of ubiquity as they mostly resort to additional information or extra user constraints. In this article, we present the design and implementation of ViViPlus, an indoor localization system purely based on WiFi fingerprints, which jointly mitigates spatial ambiguity and temporal instability and derives reliable performance without impairing the ubiquity. The key idea is to embrace the spatial awareness of RSS values in a novel form of RSS Spatial Gradient (RSG) matrix for enhanced WiFi fingerprints. We devise techniques for the representation, construction, and localization of the proposed fingerprint form and integrate them all in a practical system. Extensive experiments across 7 months in different environments demonstrate that ViViPlus significantly improves the accuracy in localization scenarios by about 30% to 50% compared with the state-of-the-art approaches.
Danyang Li 0005, Jingao Xu, Zheng Yang 0002, Chenshu Wu, Nicholas D. Lane
ACM Trans. Sens. Networks1
2021 Multi-Region Indoor Localization Based on WVP System
abstract
Indoor localization has attracted increasingly attention in the era of Internet of Things. Single indoor localization method based on WiFi fingerprint, surveillance camera or pedestrian dead reckoning suffers from low accuracy, limited tracking region or accumulative errors. Pioneering works over-come these limitations at the costs of ubiquity as they mostly resort to additional information or extra user constraints. In the large indoor region, it is important to quickly get pedestrian detection and tracking. In this paper, an indoor localization and tracking system has been presented which integrates WiFi fingerprint, Vision of surveillance camera and Pedestrian Dead Reckoning(WVP system for short). This WVP system achieves high accuracy in dynamic indoor environment. Importantly, WVP employs a motion sequence-based matching algorithm to confirm pedestrian identity. WVP outputs enhanced accuracy and overcomes the corresponding drawbacks of each subsystem simultaneously. Experimental results show that WVP can effectively track pedestrians in multi-region, and has great robustness, and the positioning accuracy is decimeter. It also performs well in complex environment.
Li Zhang 0028, Jinhui Bao, Qiuyu Wang, Jingao Xu, Danyang Li 0005, Yaodong Yang 0005
ICPADS6
2021 Train Once, Locate Anytime for Anyone: Adversarial Learning based Wireless Localization
abstract
Among numerous indoor localization systems, WiFi fingerprint-based localization has been one of the most attractive solutions, which is known to be free of extra infrastructure and specialized hardware. To push forward this approach for wide deployment, three crucial goals on delightful deployment ubiquity, high localization accuracy, and low maintenance cost are desirable. However, due to severe challenges about signal variation, device heterogeneity, and database degradation root in environmental dynamics, pioneer works usually make a trade-off among them. In this paper, we propose iToLoc, a deep learning based localization system that achieves all three goals simultaneously. Once trained, iToLoc will provide accurate localization service for everyone using different devices and under diverse network conditions, and automatically update itself to maintain reliable performance anytime. iToLoc is purely based on WiFi fingerprints without relying on specific infrastructures. The core components of iToLoc are a domain adversarial neural network and a co-training based semi-supervised learning framework. Extensive experiments across 7 months with 8 different devices demonstrate that iToLoc achieves remarkable performance with an accuracy of 1.92m and > 95% localization success rate. Even 7 months after the original fingerprint database was established, the rate still maintains > 90%, which significantly outperforms previous works.
Danyang Li 0005, Jingao Xu, Zheng Yang 0002, Yumeng Lu, Qian Zhang 0017, Xinglin Zhang 0001
INFOCOM1
2021 FollowUpAR: enabling follow-up effects in mobile AR applications
abstract
Existing smartphone-based Augmented Reality (AR) systems are able to render virtual effects on static anchors. However, today's solutions lack the ability to render follow-up effects attached to moving anchors since they fail to track the 6 degrees of freedom (6-DoF) poses of them. We find an opportunity to accomplish the task by leveraging sensors capable of generating sparse point clouds on smartphones and fusing them with vision-based technologies. However, realizing this vision is non-trivial due to challenges in modeling radar error distributions and fusing heterogeneous sensor data. This study proposes FollowUpAR, a framework that integrates vision and sparse measurements to track object 6-DoF pose on smartphones. We derive a physical-level theoretical radar error distribution model based on an in-depth understanding of its hardware-level working principles and design a novel factor graph competent in fusing heterogeneous data. By doing so, FollowUpAR enables mobile devices to track anchor's pose accurately. We implement FollowUpAR on commodity smartphones and validate its performance with 800,000 frames in a total duration of 15 hours. The results show that FollowUpAR achieves a remarkable rotation tracking accuracy of 2.3° with a translation accuracy of 2.9mm, outperforming most existing tracking systems and comparable to state-of-the-art learning-based solutions. FollowUpAR can be integrated into ARCore and enable smartphones to render follow-up AR effects to moving objects.
Jingao Xu, Guoxuan Chi, Zheng Yang 0002, Danyang Li 0005, Qian Zhang 0017, Qiang Ma 0007
MobiSys4
2021 Enabling Surveillance Cameras to Navigate
abstract
Smartphone localization is essential to a wide spectrum of applications in the era of mobile computing. The ubiquity of smartphone mobile cameras and surveillance ambient cameras holds promise for offering sub-meter accuracy localization services thanks to the maturity of computer vision techniques. In general, ambient-camera-based solutions are able to localize pedestrians in video frames at fine-grained, but the tracking performance under dynamic environments remains unreliable. On the contrary, mobile-camera-based solutions are capable of continuously tracking pedestrians; however, they usually involve constructing a large volume of image database, a labor-intensive overhead for practical deployment. We observe an opportunity of integrating these two most promising approaches to overcome above limitations and revisit the problem of smartphone localization with a fresh perspective. However, fusing mobile-camera-based and ambient-camera-based systems is non-trivial due to disparity of camera in terms of perspectives, parameters and incorrespondence of localization results. In this article, we propose iMAC, an integrated mobile cameras and ambient cameras based localization system that achieves sub-meter accuracy and enhanced robustness with zero-human start-up effort. The key innovation of iMAC is a well-designed fusing frame to eliminate disparity of cameras including a construction of projection map function to automatically calibrate ambient cameras, an instant crowd fingerprints model to describe user motion patterns, and a confidence-aware matching algorithm to associate results from two sub-systems. We fully implement iMAC on commodity smartphones and validate its performance in five different scenarios. The results show that iMAC achieves a remarkable localization accuracy of 0.68 m, outperforming the state-of-the-art systems by >75%.
Jingao Xu, Guoxuan Chi, Danyang Li 0005, Xinglin Zhang 0001, Qiang Ma 0007, Zheng Yang 0002
ACM Trans. Sens. Networks4
2020 Enabling Surveillance Cameras to Navigate
abstract
Smartphone localization is essential to a wide spectrum of applications in the era of mobile computing. The ubiquity of smartphone mobile cameras and surveillance ambient cameras holds promise for offering sub-meter accuracy localization services thanks to the maturity of computer vision techniques. In general, ambient-camera-based solutions are able to localize pedestrians in video frames at fine-grained, but the tracking performance under dynamic environments remains unreliable. On the contrary, mobile-camera-based solutions are capable of continuously tracking pedestrians, however, they usually involve constructing a large volume of image database, a labor-intensive overhead for practical deployment. We observe an opportunity of integrating these two most promising approaches to overcome above limitations and revisit the problem of smartphone localization with a fresh perspective. However, fusing mobile-camera-based and ambient-camera-based systems is non-trivial due to disparity of camera in terms of perspectives, parameters and incorrespondence of localization results. In this paper, we propose iMAC, an integrated mobile cameras and ambient cameras based localization system that achieves sub-meter accuracy and enhanced robustness with zero-human start-up effort. The key innovation of iMAC is a well-designed fusing frame to eliminate disparity of cameras including a construction of projection map function to automatically calibrate ambient cameras, an instant crowd fingerprints model to describe user motion patterns, and a confidence-aware matching algorithm to associate results from two sub-systems. We fully implement iMAC on commodity smart-phones and validate its performance in five different scenarios. The results show that iMAC achieves a remarkable localization accuracy of 0.68m, outperforming the state-of-the-art systems by > 75%.
Jingao Xu, Guoxuan Chi, Danyang Li 0005, Xinglin Zhang 0001, Qiang Ma 0007, Zheng Yang 0002
ICCCN4
2020 Edge Assisted Mobile Semantic Visual SLAM
abstract
Localization and navigation play a key role in many location-based services and have attracted numerous research efforts from both academic and industrial community. In recent years, visual SLAM has been prevailing for robots and autonomous driving cars. However, the ever-growing computation resource demanded by SLAM impedes its application to resource-constrained mobile devices. In this paper we present the design, implementation, and evaluation of edgeSLAM, an edge assisted real-time semantic visual SLAM service running on mobile devices. edgeSLAM leverages the state-of-the-art semantic segmentation algorithm to enhance localization and mapping accuracy, and speeds up the computation-intensive SLAM and semantic segmentation algorithms by computation offloading. The key innovations of edgeSLAM include an efficient computation offloading strategy, an opportunistic data sharing mechanism, and an adaptive task scheduling algorithm. We fully implement edgeSLAM on an edge server and different types of mobile devices (2 types of smartphones and a development board). Extensive experiments are conducted under 3 data sets, and the results show that edgeSLAM is able to run on mobile devices at 35fps frame rate and achieves a 5cm localization accuracy, outperforming existing solutions by more than 15%. We also demonstrate the usability of edgeSLAM through 2 case studies of pedestrian localization and robot navigation. To the best of our knowledge, edgeSLAM is the first real-time semantic visual SLAM for mobile devices.
Jingao Xu, Danyang Li 0005, Kehong Huang, Chen Qian 0009, Longfei Shangguan, Zheng Yang 0002
INFOCOM3