VLDB 2026 Research / reviewers in the wild / expert
Ruipeng Gao
dblp:137/0080
· DBLP profile ↗
55ranked-venue papers
14as first author
35since 2021 · last 2026
0000-0002-2490-6654ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 36 · 12 first-author · 17 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An enhanced deep learning framework with Large Separable Kernel Attention and Reparameterized Dual Convolution for real-time cold-crack detection in laser cladding
Jinyang Du, Ruipeng Gao, Jiabao Zhao, Yuechen Meng |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | An Alarm Prediction Framework Based on the Propagation Dependency and Reinforcement Learning Toward Cloud ServicesabstractIn cloud platforms, the occurrence of failures in cloud services is usually accompanied by a substantial number of alarms, posing challenges to the operation, management, and maintenance. To address this issue, this paper proposes an alarm prediction framework based on propagation dependency and reinforcement learning for cloud services. First, a three-layer cloud service alarm knowledge graph is constructed, encompassing the device layer, service layer, and alarm layer, with the objective of integrating alarm data. Then, to filter out irrelevant alarms, an alarm compression schema is designed, taking into account both structural and semantic information. It achieves the preservation of propagation information by mining frequent alarm subtrees. After that, a reinforcement learning-based alarm prediction model is devised, including an alarm-driven reinforcement learning network and a global alarm discriminator. The former formulates the alarm propagation problem as a reinforcement learning inferencing task, and designs an immediate reward function to simulate the local propagation patterns of alarms. The latter guides the agent to learn alarm evolution patterns from a global perspective through an adversarial approach. Experimental results on two datasets demonstrate the effectiveness of our alarm compression and prediction models. Peng Qi 0006, Dan Tao, Ruipeng Gao |
IEEE Trans. Cloud Comput. | 3 |
| 2026 | DeskPred: Two-Stage Video Stream Bandwidth Prediction for Cold-Start and Training Forgetting in Cloud DesktopsabstractAs a cloud-hosted virtual desktop service, cloud desktop supports various fields such as telecommuting, collaborative development, while enabling real-time user interaction through video stream. The stability of this process is determined by bandwidth, which significantly influences the user experience. Therefore, precise bandwidth prediction of video streams is essential in cloud desktops. This work proposes DeskPred for video stream transmission in cloud desktops, focusing on dynamic bandwidth prediction. In the startup stage, the limited data amount poses a challenge for achieving precise bandwidth predictions. We propose an Affinity-based Federated Learning algorithm, which leverages the historical records of high-affinity users for assisted training, all while protecting user privacy. During the long-term adjustment stage, we propose a Fluctuation-based Adaptive Incremental Prediction algorithm for independent training to address the issue of pattern forgetting. The algorithm considers both periodic features and instantaneous features, incorporating new patterns while revisiting previous knowledge through the memory module and Adversarial Elastic Weight Consolidation. We have verified DeskPred through an actual cloud desktop project supported by Lenovo Research. Through experiments conducted on a total of over 18 million data items (approximately 10 GB), DeskPred achieves the highest total score of 71.11%, making it highly suitable for cloud desktop environments. Zuodong Jin, Dan Tao, Peng Qi 0006, Ruipeng Gao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Harmonizing Time-of-Use Pricing and Rigid Driver Schedule in Battery Swapping ServiceabstractElectric scooters (ESs) serve as smart city terminals by uploading real-time charging data to improve mobility and energy efficiency. Battery swapping service has emerged as a transformative solution for energy replenishmen of ESs. Time-of-use (TOU) pricing is commonly adopted to encourage off-peak electricity usage by offering different rates. However, the rigid work schedules of ES drivers always hinder the optimization of battery swapping, leading to escalated charging costs. Existing literature fails to account for the practical constraints of battery charging in real-world scenarios. To tackle thisparadox, we studied battery swap stations in Chengdu. Our findings reveal that the cost of swapping stations could be remarkably reduced by optimizingintrinsic, energy-greedy charging power strategies rather than by altering theextrinsic, rigid swapping schedules. Further analysis indicates 63% of charged batteries remain idle, presenting an opportunity to optimize charging power for cost savings without disrupting swapping schedules. Motivated by these insights, we propose PowerRL, anewadaptive charging power strategy empowered by spatio-temporal reinforcement learning. Specifically, we employ a spatio-temporal gated recurrent unit to predict the demand for battery swapping, which is then used to optimize charging power through reinforcement learning. Applied to the large-scale dataset, PowerRL could reduce per-unit electricity costs by 25.7% while still accommodating driver swapping schedules, thereby harmonizing TOU pricing with the rigidity of ES swapping service. Dongshang Deng, Chaocan Xiang, Chengyi Gu, Bincan Yu, Xuangou Wu, Ruipeng Gao |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2026 | Enabling Service Monitoring in Industrial IoT: A Knowledge-Integrated Service Status Perception Framework
Peng Qi 0006, Dan Tao, Ruipeng Gao |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2025 | KFCalibNet: A KansFormer-Based Self-Calibration Network for Camera and LiDARabstractIn autonomous driving and robotic navigation, multi-sensor fusion technology has become increasingly mainstream, with precise sensor calibration as its foundation. Traditional calibration methods rely on manual effort or specific targets, limiting adaptability to complex environments. Learning-based calibration methods still face challenges, such as insufficient overlap between the fields of view (FoV) of multiple sensors and suboptimal cross-modal feature association, which hinder accurate parameter regression. Unlike traditional CNN-based networks, we propose a KansFormer-based self-Calibration Network for camera and LiDAR (KFCalibNet) that replaces fixed activation functions and linear transformations with learnable nonlinear activation functions. This enables the extraction of more fine-grained features from both image and point cloud, significantly enhancing the network's robustness in scenarios with limited FoV overlap. We also employ a multihead attention (MHA) module to compute correlations between image and point cloud features, significantly enhancing cross-modal feature association. To reduce learning complexity, we designed KansFormer with FastKAN as the feedforward network, enabling deep fusion and regression of fine-grained cross-modal features for accurate extrinsic calibration. KFCalibNet achieves an absolute average calibration error of 0.0965 cm in translation and 0.0234° in rotation on the KITTI Odometry dataset, outperforming existing state-of-the-art calibration methods. Moreover, its accuracy and generalization capability have been validated across multiple real-world railway lines. Zejing Xu, Ruipeng Gao, Dan Tao, Peng Qi 0006 |
ICRA | 3 |
| 2025 | Lightweight Self-Supervised Monocular Depth Estimation Based on Multi-Scale Hybrid AttentionabstractSelf-supervised monocular depth estimation has gained significant traction due to its capacity to train models without the need for depth annotations. However, CNNs are inherently constrained by limited receptive fields, restricting their ability to infer beyond local contexts. To address this, hybrid architectures that integrate Transformers have demonstrated exceptional performance by capturing long-range dependencies. Nevertheless, such large backbones impose substantial demands on both computation and storage resources. In this study, we propose a lightweight and computation-efficient depth estimation network that synergistically combines traditional convolutions with attention mechanisms. Our model is capable of conducting both local and global reasoning, thereby preserving high accuracy while significantly reducing computational overhead. We present our novel architectural components, including convolutional samplers, cross-feature extraction modules, and interactive fusion modules, which collectively enhance the extraction and fusion of multi-dimensional and multi-scale features. Our approach establishes new benchmarks on the KITTI dataset, achieving an inference accuracy of 98.4% while reducing the parameter count by 52.4% relative to Monodepth2. Xue Yi, Ruipeng Gao |
IJCNN | 2 |
| 2025 | LuminaLink: Enabling Low Cost Secure Visible Light Communication with Birefringence
Ruxin Lin, Yelin Cui, Ruipeng Gao, Xiaoqiang Zhu, Jiqiang Liu, Lingkun Li |
INFOCOM | 5 |
| 2025 | Rethinking Discrepancy Analysis: Anomaly Detection via Meta-Learning Powered Dual-Source Representation DifferentiationabstractIndustrial environments pose distinctive challenges for anomaly detection, primarily stemming from the complexities associated with high dimensionality and the dynamic nature of data patterns over time. These properties determine that the model’s proper convergence on unlabeled data is unpromising, consequently leading to less efficient discrimination of anomalies in previous anomaly detection (AD) works. To address this problem, we present AnoDual, a novel, meta-learning AD framework. From the perspective of data reconstruction, we introduce the multi-memory enhanced VAE reconstructor M2ER, which learns to extract the most salient patterns in unlabeled noisy data through a self-supervised manner. This design eases impacts from potential anomalous components during data reconstruction, and enhances the discernibility of anomalies. To address performance degradation caused by the numerical deviation based AD scheme in most existing works, we design a dual-source self-supervised discriminator DSD, which examines characteristics in the domain of representations. This model actively assesses discrepancies between data pairs and representation pairs in parallel, and conducts AD on a fine-grained scale. In this way, anomalies that used to be unnoticed due to a less prominent numerical deviation can be spotted. Besides, we propose a meta-learning powered training pipeline to enable model training even when no real label is available, which is common in the industry. Extensive experiments on five large-scale real-world industrial datasets suggest that AnoDual achieves an average F1-Score with a substantial increment of 3.39 %, outperforming the latest state-of-the-art baseline. Note to Practitioners—A generative model plus a numerical threshold based detection approach currently takes a significant share in both academia and the industry. However, the performance of this workflow is not promising in actual applications, with multiple factors contributing to this situation. The proper convergence of such generative models is difficult when the training material contains noisy samples - an over-expressed generative model would result in less significant reconstruction discrepancies for anomalies that are hard to notice. In addition, selecting a numerical threshold, which is used to spot anomalies, requires multiple laborious attempts, and can hardly adapt to an ever-changing pattern in industrial environments. These circumstances make it challenging to apply prior works in practical production, which, in turn, urges the need to develop an effective methodology to address the need for industrial anomaly detection. This manuscript includes a novel, meta-learning powered framework AnoDual, which is tailored for industrial scenarios. This framework discards the conventional design of comparing the reconstruction error numerically, but introduces a solution based on the differentiation of the representations. Besides, the multi-head attention enhanced variational autoencoder also leads to a much more pronounced discrepancy for anomalous samples, which benefits their successful detection. Providing a flexible and robust way to detect anomalies on deployed IoT assets, this work can be further transformed to serve applications in many other domains. Muyan Yao, Dan Tao, Peng Qi 0006, Ruipeng Gao |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Anomaly Detection for MEC Enabled Hierarchical Industrial IoT With Transformer Enhanced Variational Auto EncoderabstractMost existing works in Industrial Internet of Things (IIoT) anomaly detection either depend on computationally intensive models that exceed the capabilities of multiaccess edge computing (MEC) servers, or lightweight models that lack robustness, making them unadaptable in IIoT infrastructures. To address these challenges, we proposeTHREADS, a hierarchical anomaly detection framework tailored for IIoT applications. TheInstance threadutilizes an efficient variational auto encoder to produce instant feedback and offloads most of the workload to MECs. On the other hand, theShadow threademploys an attention-enhanced transformer discriminator to examine low-confidence results in the cloud. Experimental results on five large-scale datasets showTHREADSachieves an averageF1-Scoreof 0.8537 in the hierarchical mode where most of the workloads are handled by MECs, and the random access memory and CPU usage is reduced by up to 29% and 88%, respectively. Meanwhile,THREADSachieves anF1-Scoreof 0.8563 in a cloud-based mode, consistently outperforming state-of-the-art approaches. Muyan Yao, Dan Tao, Ruipeng Gao, Peng Qi 0006 |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | A Novel Adversarial Augmentation-based Domain LLM Framework for Perceiving IoT Service StatesabstractAccurately perceiving service states of the Internet of Things (IoT) is crucial for maintaining system stability and long-term sustainability. Recently, Large Language Models (LLMs) have emerged as a novel technological advancement, offering unprecedented possibilities in various domains due to their advanced comprehension capabilities. However, the limited data availability problem poses a challenge to the integration of domain knowledge into LLMs and the enhancement of their ability to perceive service states. In this article, a novel adversarial augmentation-based domain LLM framework for IoT services is proposed to solve this problem. We first construct an LLM fine-tuning dataset through an elaborate hierarchical prompt template, which integrates task-specific instructions, heterogeneous service data, and statistical features to the service contextual description. Then, inspired by using specialized compact small-models to enhance the capabilities of LLMs, we design two dedicated domain-specific small-models, including a service topology prediction model and a running status prediction model, to capture evolution patterns of services from different perspectives. On this basis, we design an LLM fine-tuning framework based on the domain adversarial distillation, which comprises an LLM parameter tuning module and a small-model adversarial distillation module. The former leverages the Low-Rank Adaptation (LoRA) algorithm to fine-tune the LLM. The latter distills the domain knowledge of small-models into the LLM through aligning the latent vector distributions of the LLM with those of small-models. Following the adversarial training, the LLM will be equipped with domain capabilities of small-models. Experimental results on two datasets demonstrate the effectiveness of our LLM framework. Peng Qi 0006, Zuodong Jin, Dan Tao, Ruipeng Gao |
ACM Trans. Internet Things | 4 |
| 2025 | DeskTransfer: Predicting Multi-Scenario Video Stream Throughput in Cloud Desktop Based on Transfer AutoencoderabstractThe popularity of cloud services has provided a new medium for video streams transmission. Cloud desktops, as a representative multimedia application, facilitate interaction between users and cloud via video streams, garnering widespread adoption in various fields. The network condition directly affects the transmission. Therefore, accurate throughput prediction helps guide the allocation of network resources, avoiding a decline in user experience due to insufficient resources and waste caused by excessive resources. Recent works focus more on the temporal characteristics of throughput. However, we believe that throughput of video streaming is significantly influenced by usage scenario. In this paper, we propose a transfer-based autoencoder framework DeskTransfer for throughput prediction in frequent switching cloud desktop scenarios. Specifically, we construct the Scenario Autoencoder and Throughput Autoencoder to respectively learn the scenario and throughput features from historical usage records. By adopting an adversarial mechanism, we design transfer algorithm using latent vectors, enabling the model suitable for multiple scenarios. We collect real-world data from a project cooperated with Lenovo Research for experiment and compare our solution with leading methods on public datasets to validate its effectiveness. Zuodong Jin, Peng Qi 0006, Ruipeng Gao, Yanzhe Jing, Dan Tao |
IEEE Trans. Multim. | 3 |
| 2025 | Scalable Large Model for Unlabeled Anomaly Detection With Trio-Attention U-Transformer and Manifold-Learning Siamese DiscriminatorabstractTo identify pattern deviations in large-scale industrial infrastructures, anomaly detection is crucial yet challenging. Previous research has not adequately addressed the characteristics and deployment considerations in these complex scenarios. In this paper, we presentInoU, a scalable anomaly detection framework to process unlabeled multivariate time-series data. We incorporate a VAE filter to ease impacts from noisy components in training materials. We propose a scalable trio-attention U-Transformer to construct the typical representation of high-dimensional streams and produce pseudo labels that enable the later training process. The ultra perception and intra-/ inter-flow attention mechanisms are delicately designed to aggregate information from different flows with variable granularities while keeping a global view of the data. Its nested structure helps to maintain high efficiency even when the model is scaled down. We introduce a Siamese discriminator that projects target data into manifolds, and collates discrepancies at the embedding level. This paradigm elevates detection performance far beyond segment-wise error comparison in prior works. We apply contrastive and adversarial learning techniques to optimize manifold projection and detection performance when processing unseen samples. Extensive experiments on five large-scale datasets demonstrate the effectiveness ofInoUwith an averageF1-Scoreimprovement of 5.58%, significantly outperforming the state-of-the-art. Muyan Yao, Dan Tao, Peng Qi 0006, Ruipeng Gao |
IEEE Trans. Serv. Comput. | 4 |
| 2024 | Smartphone-based Indoor Pedestrian Tracking via TransformerabstractIn a GPS-blocked environment, pedestrians often lose directions and struggle with finding the correct path. We propose an indoor pedestrian tracking method based on smartphone sensors, including the IMU (Inertial Measurement Unit) and the barometer. The system tracks pedestrians’ indoor walking trajectories by using a single smartphone. Without prerequisite infrastructures, the system utilizes Transformer for speed inference and integrates the gyroscope to estimate real-time bearing. During the trajectory estimation process, the system combines the speed and the bearing to track the trajectory. To enable multi-floor tracking, we propose a stair detection algorithm based on the barometer and the accelerometer. We also implement backtracking based on prior trajectories, which allows the pedestrian to get back to the starting position. We evaluate our system in an underground parking lot and multi-story indoor buildings. The results show our advantages. Xueqi Li 0006, Kejia Li, Jiayao Liu, Ruipeng Gao |
CSCWD | 4 |
| 2024 | DBLG: An Innovative Deep-Broad Learning and GAN Framework for CSI Fingerprint Database RefinementabstractWith the rapid development of Integrated Sensing and Communication in 6G, Channel State Information (CSI)-based fingerprint indoor localization technology is becoming crucial. However, during the offline phase, the fingerprint database update process using crowdsourcing techniques is prone to noise interference and incomplete coverage, and fitting Gaussian regression models requires extensive computational resources. In this paper, we propose a fine-grained CSI fingerprint database update method based on Deep-Broad Learning system (DeepBLS) and Generative Adversarial Networks (GAN), termed DBLG. Firstly, we employ the combined DeepBLS network for the initial construction of the global CSI fingerprint database. Subsequently, we utilize GAN to extract features from the raw data, predict and update the global CSI fingerprint database, and construct a high-precision fingerprint database using confidence coefficients. Finally, We implement the proposed algorithm in two real-world environments and conduct extensive experiments to verify its performance. Compared to several existing methods, our approach shows superior performance in updating the CSI finger-print database, achieving a 46.78 % improvement in localization accuracy. Mingbo Zhang, Lingyun Lu, Xiaoqiang Zhu, Lingkun Li, Ruipeng Gao |
MSN | 5 |
| 2024 | An Adaptive Cloud Resource Quota Scheme Based on Dynamic Portraits and Task-Resource MatchingabstractDue to the unrestricted location of cloud resources, an increasing number of users are opting to apply for them. However, determining the appropriate resource quota has always been a challenge for applicants. Excessive quotas can result in resource wastage, while insufficient quotas can pose stability risks. Therefore, it's necessary to propose an adaptive quota scheme for cloud resource. Most existing researches have designed fixed quota schemes for all users, without considering the differences among users. To solve this, we propose an adaptive cloud quota scheme through dynamic portraits and task-resource optimal matching. Specifically, we first aggregate information from text, statistical, and fractal three dimensions to establish dynamic portraits. On this basis, the bidirectional mixture of experts (Bi-MoE) model is designed to match the most suitable resource combinations for tasks. Moreover, we define the time-varying rewards and utilize portrait-based reinforcement learning (PRL) to obtain the optimal quotas, which ensures stability and reduces waste. Extensive simulation results demonstrate that the proposed scheme achieves a memory utilization rate of around 70%. Additionally, it shows improvements in task execution stability, throughput, and the percentage of effective execution time. Zuodong Jin, Dan Tao, Peng Qi 0006, Ruipeng Gao |
IEEE Trans. Cloud Comput. | 4 |
| 2024 | Real-World Large-Scale Cellular Localization for Pickup Position Recommendation at Black-HoleabstractIndoor localization availability is still sporadic in industry, especially at the black-hole, i.e., there only exist cellular signals, no GPS or WiFi signals. Based on our 2-year observations at the DiDi ride-hailing platform in China, there are$ 68\,\text{k}$orders everyday created at black-hole. In this paper, we presentTransparentLoc, a large-scale cellular localization system for pickup position recommendation of the DiDi platform. Specifically, we design a CNN model for real-time localization based on a crowdsourcing fingerprint set constructed by outdoor trajectories and abnormal cell tower detection. Then we leverage a DeepFM model to recommend an optimal pickup position for passengers. We share our 2-year experience with 50 million orders across 13 million devices in 4541 cities to address practical challenges including sparse cell towers, unbalanced user fingerprints, temporal variations, and abnormal cell towers in terms of four major service metrics, i.e., pickup position error, over-30-meters ratio, cancel ratio, and call ratio. The large-scale evaluations show that our system achieves a$ 0.54\,\text{m}$lower median pickup position error compared to the iOS built-in cellular localization system, regardless of environmental changes, smartphone brands/models, time, and cellular providers. Additionally, the over-30-meters ratio, cancel ratio, and call ratio have significant reductions of 0.88%, 0.88%, and 5.13%, respectively. Ruipeng Gao, Shuli Zhu, Lingkun Li, Xuyu Wang, Yuqin Jiang, Naiqiang Tan, Peng Qi 0006, Jiqiang Liu, Dan Tao |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Deep Compressed Sensing based Data Imputation for Urban Environmental MonitoringabstractData imputation is prevalent in crowdsensing, especially for Internet of Things (IoT) devices. On the one hand, data collected from sensors will inevitably be affected or damaged by unpredictability. On the other hand, extending the active time of sensor networks has urgently aspired environmental monitoring. Using neural networks to design a data imputation algorithm can take advantage of the prior information stored in the models. This paper proposes a preprocessing algorithm to extract a subset for training a neural network on an IoT dataset, including time window determination, sensor aggregation, sensor exclusion and data frame shape selection. Moreover, we propose a data imputation algorithm using deep compressed sensing with generative models. It explores novel representation matrices and can impute data in the case of a high missing ratio situation. Finally, we test our subset extraction algorithm and data imputation algorithm on the EPFL SensorScope dataset, respectively, and they effectively improve the accuracy and robustness even with extreme data loss. Qingyi Chang, Dan Tao, Jiangtao Wang 0001, Ruipeng Gao |
ACM Trans. Sens. Networks | 4 |
| 2023 | Energy-Efficient Transmission Scheduling with Guaranteed Data Imputation in MHealth SystemsabstractTransmission scheduling is a fundamental energy-saving problem in many wireless sensor networks (WSN), especially for mHealth systems with multiple distributed sensing modalities. Recent work has made significant progress to balance the transmission efficiency and timeliness, e.g., periodically entering sleep modes to reduce the power, but such corresponded missing samples impede data integrity and timeliness for realtime diagnosis. In this paper, we intuitively combine transmission scheduling with data imputation, and propose a novel software-hardware cooperated framework for energy-efficient mHealth systems. Specially, we devise a Wasserstein Generative Adversarial Imputation Network (WGAIN) model to impute missing samples, which exploits heterogeneous correlation, temporal dependency, and missing patterns by a divide-and-conquer strategy. We also propose a dropout-based uncertainty approximation method inside the imputation model, prove its equivalence to Gaussian process with variational inference, and explore a heuristic transmission scheduling algorithm for lifecycle management among heterogeneous sensory modules. Extensive experiments on the MIT-BIH dataset and our mHealth prototype have demonstrated the effectiveness compared with state-of-the-art. Ruipeng Gao, Haoyue Zhao, Zonglin Xie, Dan Tao |
IWQoS | 1 |
| 2023 | Experience: Large-scale Cellular Localization for Pickup Position Recommendation at Black-holeabstractLocation awareness is the basis for enabling pickup service at ride-hailing platforms. In contrast to the almost pervasive coverage outdoors, indoor localization availability is still sporadic in industry since it largely relies on RF signatures from certain IT infrastructure, e.g., WiFi access points. Based on our 2-year observations at DiDi ride-hailing platform in China, there are 68k orders everyday created at black-hole, i.e., where only cellular signals exist. In this paper, we present the design, development, and deployment of TransparentLoc, a large-scale cellular localization system for pickup position recommendation, and share our 2-year experience with 50 million orders across 13 million devices in 4541 cities to address practical challenges including sparse cell towers, unbalanced user fingerprints, and temporal variations. Our system outperforms the iOS built-in cellular localization system in terms of four major service metrics, regardless of environmental changes, smartphone brands/models, time, and cellular providers. Shuli Zhu, Lingkun Li, Xuyu Wang, Changcheng Liu, Yuqin Jiang, Zengwei Huo, Jiqiang Liu, Dan Tao, Ruipeng Gao |
MobiCom | 10 |
| 2023 | PresSafe: Barometer-Based On-Screen Pressure-Assisted Implicit Authentication for SmartphonesabstractGraphic-pattern-based implicit authentication has been successfully exploited to elevate the security of smartphones. On-screen pressure is one of the key features in such an approach since it can reveal users’ touch pattern. However, state-of-the-art approaches rely on a system API to obtain on-screen pressure, which is not adequately accurate and cannot meet the demands of robust implicit authentication. To bridge this gap, we propose PresSafe, a novel implicit authentication system that utilizes the smartphone’s built-in barometer sensor to measure pressure during the unlocking process, and to utilize the pressure data in authentication. A key technical challenge in utilizing barometer sensing, however, is to understand the user activity through measured pressure. To overcome this challenge, PresSafe leverages barometer data along with data from other conventional but heterogeneous ambient sensors to produce accurate and robust user activity descriptions. PresSafe utilizes a transfer-learning-based hybrid workflow to integrate user activity representation learning with a lightweight classical authentication algorithm to obtain a unified model. This approach offloads the computational cost from the terminal and addresses privacy concerns. To ensure applicability of our approach despite data heterogeneity and insufficient training data, we utilize a channel-adaptive data processing mechanism. Extensive experiments utilizing more than 70000 records from 23 volunteers in six different locations show that PresSafe achieves an FAR of 0.45%, an FRR of 0.49%, and an EER of 0.47%, which clearly demonstrate its superiority over several existing solutions. Muyan Yao, Dan Tao, Ruipeng Gao, Jiangtao Wang 0001, Abdelsalam Helal, Shiwen Mao |
IEEE Internet Things J. | 3 |
| 2023 | Dual-norm based dynamic graph diffusion network for temporal prediction
Fuyong Sun, Weiwei Xing, Xiaofei Tian, Ruipeng Gao, Wei Lu 0010 |
Inf. Process. Manag. | 4 |
| 2023 | Sensing-gain constrained participant selection mechanism for mobile crowdsensing
Dan Tao, Ruipeng Gao |
Pers. Ubiquitous Comput. | 2 |
| 2022 | PeTrack: Smartphone-based Pedestrian Tracking in Underground Parking LotabstractAlthough location awareness is prevalent outdoors due to GNSS systems and devices, pedestrians are back into darkness in indoor buildings such as underground parking lots. Frequently we forget where we park the car and get confused by such maze-like structure. In order to track pedestrians without any additional equipment and map support, we propose PeTrack which is a smartphone-only approach that collects the inertial measurement unit (IMU) data for long-term tracking. Our intuition is to train the tracking model with crowdsourced outdoor trajectories, and infer customized user's trace with only inertial readings at indoors. Specially, we propose an inertial sequence learning framework with outdoor geo-tags. We also exploit opportunistic landmark detection and structure cues to refine the trajectory. We have developed a prototype and conducted experiments in an underground parking lot, and results have shown our effectiveness. Xiaotong Ren, Shuli Zhu, Chuize Meng, Dan Tao, Ruipeng Gao |
MSN | 7 |
| 2022 | Crowdsourced Image Driven PM2.5 Estimation based on Hybrid 3-Channel Feature MapabstractIndustrialization has resulted in a relatively high airborne PM2.5 concentration in most developing areas, causing severe consequences due to its physico-chemical properties. In this paper, we propose a crowdsourced image driven PM2.5concentration estimation approach based on meteorological en-hanced 3-channel feature map. To fulfill this framework, we first use dark channel prior and Rayleigh's law of atmospheric scattering to correct the proportion of the sky area in the crowdsourced images and then construct hybrid 3-channel feature maps. We then use deep learning models to perform feature extraction on hand-engineered feature maps and provide fine-grained PM2.5 concentration estimation. We also conduct a series of experiments to evaluate the performance of our model. Evaluation on dataset collected at 8 sites over nearly 2 years demonstrates that, our system achieves an MAE of 23.27$\mu\mathrm{g}/\mathrm{m}^{3}$, outperforming baseline solutions. Muyan Yao, Ruipeng Gao, Dan Tao |
MSN | 3 |
| 2022 | BatMapper-Plus: Smartphone-Based Multi-level Indoor Floor Plan Construction via Acoustic Ranging and Inertial Sensing
Chuize Meng, Mengning Wu, Dan Tao, Ruipeng Gao |
WASA (2) | 6 |
| 2022 | TimeBird: Context-Aware Graph Convolution Network for Traffic Incident Duration Prediction
Fuyong Sun, Ruipeng Gao, Weiwei Xing, Yaoxue Zhang, Wei Lu 0010 |
WASA (1) | 2 |
| 2022 | Attentive Auto-encoder for Content-Aware Music Recommendation
Dan Tao, Chenwang Zheng, Ruipeng Gao |
CCF Trans. Pervasive Comput. Interact. | 4 |
| 2022 | Unsupervised Learning of Monocular Depth and Ego-Motion in Outdoor/Indoor EnvironmentsabstractVisual-based unsupervised learning[1]–[3]has emerged as a promising approach in estimating monocular depth and ego-motion, avoiding intensive efforts on collecting and labeling the ground truth. However, they are still restrained by the brightness constancy assumption among video sequences, especially susceptible with frequent illumination variations or nearby textureless surroundings in indoor environments. In this article, we selectively combine the complementary strength of visual and inertial measurements, i.e., videos extract static and distinct features while inertial readings depict scale-consistent and environment-agnostic movements, and propose a novel unsupervised learning framework to predict both monocular depth and ego-motion trajectory simultaneously. This challenging task is solved by learning both forward and backward inertial sequences to eliminate inevitable noises, and reweighting visual and inertial features via gated neural networks in various environments or with user-specific moving dynamics. In addition, we also employ structure cues to produce scene depths from a single image and explore structure consistency constraints to calibrate the depth estimates in indoor buildings. Experiments on the outdoor KITTI data set and our dedicated indoor prototype reveal that our approach consistently outperforms the state of the art on both depth and ego-motion estimates. To the best of our knowledge, this is the first work to fuse visual and inertial data without any supervision signals for monocular depth and ego-motion estimation, and our solution remain effective and robust even in textureless indoor scenarios. Ruipeng Gao, Weiwei Xing, Lei Liu 0059 |
IEEE Internet Things J. | 1 |
| 2022 | CTTE: Customized Travel Time Estimation via Mobile CrowdsensingabstractEstimating the origin-destination travel time is a fundamental problem in many location-based services for vehicles, e.g., ride-hailing, vehicle dispatching, and route planning. Recent work has made significant progress to accuracy, but they largely rely on GPS trajectories which are too coarse to model many personalized driving behaviors, e.g., differentiating novice and veteran drivers. In this paper, we propose Customized Travel Time Estimation (CTTE) that fuses GPS trajectories, smartphone inertial data, and road network within a deep recurrent neural network. It constructs a road link traffic database with topology representation, speed statistics, and query distribution. It also calibrates inertial readings, estimates the arbitrary phone’s pose in car, and detects multiple aggressive driving events (e.g., bump judders, sharp turns, sharp slopes, frequent lane shifts, overspeeds, and sudden brakes). Finally, we demonstrate our solution on two typical transportation problems, i.e., predicting traffic speed at holistic level and estimating customized travel time at personal level, within a multi-task learning structure. Experiments on two large-scale real-world traffic datasets from DiDi platform show our effectiveness compared with the state-of-the-art. Ruipeng Gao, Fuyong Sun, Weiwei Xing, Dan Tao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Robust Human Face Authentication Leveraging Acoustic Sensing on SmartphonesabstractUser authentication on smartphones is the key to many applications, which must satisfy both security and convenience. We propose a novel user authentication systemEchoPrint, which leverages acoustics and vision for secure and convenient user authentication, without requiring any special hardware.EchoPrintactively emits almost inaudible acoustic signals from the earpiece speaker to “illuminate” the user's face and authenticates the user by the unique features extracted from the echoes bouncing off the 3D facial contour. To combat changes in phone-holding poses thus echoes, a convolutional neural network (CNN) is trained to extract reliable acoustic features, which are further combined with visual facial features extracted from state-of-the-art face recognition deep models to feed a binary support vector machine (SVM) classifier for final authentication. Because the echo features depend on 3D facial geometries,EchoPrintis not easily spoofed by images or videos like 2D visual face recognition systems. It needs only commodity hardware, thus avoiding the extra costs of special sensors in solutions like FaceID. Experiments with 62 volunteers and non-human objects such as images, photos, and sculptures show thatEchoPrintachieves 93.75 percent balanced accuracy and 93.50 percent F-score, while the average precision is 98.05 percent using acoustic features and basic facial landmarks. The precision is further improved to 99.96 percent with sophisticated visual features. Bing Zhou 0001, Zongxing Xie, Jay Lohokare, Ruipeng Gao, Fan Ye 0003 |
IEEE Trans. Mob. Comput. | 5 |
| 2021 | DCCP: Deep Convolutional Neural Networks for Cellular Network PositioningabstractAlthough location awareness is prevalent outdoors due to the GPS, we get confused and disoriented in many blocked environments such as in urban canyons and under multi-level flyovers. A straightforward solution is to employ cellular signals for positioning, but the cellular signatures are always sparse and uneven in vast region, and vary among different devices and postures. In this paper, we propose DCCP, a novel cellular network positioning approach that transforms the localization problem into a corresponding object recognition task in geographic space. Specially, we elicit the receptive region of each cellular station via crowdsourced user queries, and exploit neighbour base stations to derive a multi-dimensional feature map. We also devise a CNN model to learn local correlations among nearby map grids, and employ it for cellular positioning. Extensive experiments on two real-world traffic datasets from the DiDi platform have demonstrated our effectiveness compared with the state-of-the-art. This is the first approach to use only user queries instead of RF signatures for cellular network positioning, and our system meets requirements of the E911. Qinkun Zhong, Ruipeng Gao, Buyi Yin, Lei Liu 0059 |
GLOBECOM | 4 |
| 2021 | Smartphone-based Vehicle Tracking without GPS: Experience and ImprovementsabstractNowadays, GPS and other global positioning systems have been widely developed, enabling accurate and convenient outdoor location-based services for vehicles. However, there are still two percents of areas in urban city that cannot be covered by satellites, e.g., underground parking lots, tunnels, and multi-level flyovers. Current positioning methods always rely on inertial dead-reckoning methods, but the performance is seriously affected by the low-quality inertial sensors embedded in crowdsourced smartphones. Based on our series of experiments with thousands of smartphones, we observe that the accuracy of existing inertial dead-reckoning methods is terribly affected by many factors, e.g., arbitrary and unknown placements of smartphones in car, inconstant inertial noises, and the diversity of smartphones and vehicles. In this paper, we explore a novel smartphone-based inertial sequence learning approach to infer vehicle's location in real time. We also propose a customized model refinement mechanism for individual drivers. Extensive experiments on DiDi ride-hailing platform have proved the effectiveness of our solution. Shuli Zhu, Qinkun Zhong, Ruipeng Gao, Lei Liu 0059 |
ICPADS | 4 |
| 2021 | Editorial for special issue on mobile intelligence: sensing, computing and networking
Chenren Xu, Ruipeng Gao, Shijia Pan, Pei Zhang 0001 |
CCF Trans. Pervasive Comput. Interact. | 2 |
| 2021 | Glow in the Dark: Smartphone Inertial Odometry for Vehicle Tracking in GPS Blocked EnvironmentsabstractAlthough vehicle location-based services are prevalent outdoors, we are back into darkness in many GPS blocked environments, such as tunnels, indoor parking garages, and multilevel flyovers. Existing smartphone-based solutions usually adopt inertial dead reckoning to infer the trajectory, but low-quality inertial sensors in phones are plagued by heavy noises, causing unbounded localization errors through double integrations for movements. In this article, we propose VeTorch, a smartphone inertial odometry that devises an inertial sequence learning framework to track vehicles in real time when GPS signal is not available. Specifically, we transform the inertial dynamics from the phone to the vehicle regardless of the arbitrary phone's placement in the car and explore a temporal convolutional network to learn the vehicle's moving dependencies directly from the inertial data. To tackle the heterogeneous smartphone properties and driving habits, we propose a federated learning-based active model training mechanism to produce customized models for individual smartphones, without incurring user privacy issues. We implement a highly efficient prototype and conduct extensive experiments on two large-scale real-world traffic data sets collected by a modern ride-hailing platform. Our results outperform the state-of-the-art vehicular inertial dead-reckoning solutions on both accuracy and efficiency. Ruipeng Gao, Shuli Zhu, Weiwei Xing, Lei Liu 0059 |
IEEE Internet Things J. | 1 |
| 2020 | Towards Scalable Indoor Map Construction and Refinement using Acoustics on SmartphonesabstractThe lack of digital floor plans is a huge obstacle to pervasive indoor location based services (LBS). Recent floor plan construction work crowdsources mobile sensing data from smartphone users for scalability. However, they incur long time (e.g., weeks or months) and tremendous efforts in data collection. In this paper, we propose BatMapper, which explores a previously untapped sensing modality-acoustics-forfast, fine grained, and low cost floor plan construction. We design sound signals suitable for heterogeneous microphones on commodity smartphones, and acoustic signal processing techniques to produce accurate distance measurements to nearby objects. We further develop robust probabilistic echo-object association, recursive outlier removal, and probabilistic resampling algorithms to identify the correspondence between distances and objects, thus the geometry of corridors and rooms. We compensate minute hand sway movements to identify small surface recessions, thus detecting doors automatically. Experiments in real buildings show BatMapperachieves 1 - 2 cm distance accuracy in ranges up around 4 m; a 2~3 minute walk generates fine grained corridor shapes, detects doors at 92 percent precision and 1~2 mlocation error at 90-percentile; and tens of seconds of measurement gestures produce room geometry with errors <; 0:3 m at 80-percentile, at 1 - 2 orders of magnitude less data amounts and user efforts. Bing Zhou 0001, Mohammed Elbadry, Ruipeng Gao, Fan Ye 0003 |
IEEE Trans. Mob. Comput. | 3 |
| 2019 | Aggressive Driving Saves More Time? Multi-task Learning for Customized Travel Time EstimationabstractEstimating the origin-destination travel time is a fundamental problem in many location-based services for vehicles, e.g., ride-hailing, vehicle dispatching, and route planning. Recent work has made significant progress to accuracy but they largely rely on GPS traces which are too coarse to model many personalized driving events. In this paper, we propose Customized Travel Time Estimation (CTTE) that fuses GPS traces, smartphone inertial data, and road network within a deep recurrent neural network. It constructs a link traffic database with topology representation, speed statistics, and query distribution. It also uses inertial data to estimate the arbitrary phone's pose in car, and detects fine-grained driving events. The multi-task learning structure predicts both traffic speed at public level and customized travel time at personal level. Extensive experiments on two real-world traffic datasets from Didi Chuxing have demonstrated our effectiveness. Ruipeng Gao, Xiaoyu Guo 0001, Fuyong Sun, Jiayan Zhu |
IJCAI | 1 |
| 2019 | Fast and Resilient Indoor Floor Plan Construction with a Single UserabstractA lack of floor plans is a fundamental obstacle to ubiquitous indoor location-based services. Recent work have made significant progress to accuracy, but they largely rely on slow crowdsensing that may take weeks or even months to collect enough data. In this paper, we propose Knitter that can generate accurate floor maps by a single random user’s one hour data collection efforts, and demonstrate how such maps can be used for indoor navigation. Knitter extracts high quality floor layout information from single images, calibrates user trajectories, and filters outliers. It uses a multi-hypothesis map fusion framework that updates landmark positions/orientations and accessible areas incrementally according to evidences from each measurement. Our experiments on three different large buildings (up to$140\times 50\;\mathrm{m}^2$) with 30+ users show that Knitter produces correct map topology, with landmark location errors of$3\sim 5\;\mathrm{m}$and orientation errors of$4\sim 6^\circ$, both at 90-percentile. Our results are comparable to the state-of-the-art at more than$20\times$speed up: data collection in each of the three buildings can finish in about one hour even by a novice user trained just a few minutes. Ruipeng Gao, Bing Zhou 0001, Fan Ye 0003, Yizhou Wang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2018 | EchoPrint: Two-factor Authentication using Acoustics and Vision on SmartphonesabstractUser authentication on smartphones must satisfy both security and convenience, an inherently difficult balancing art. Apple's FaceID is arguably the latest of such efforts, at the cost of additional hardware (e.g., dot projector, flood illuminator and infrared camera). We propose a novel user authentication system EchoPrint, which leverages acoustics and vision for secure and convenient user authentication, without requiring any special hardware. EchoPrint actively emits almost inaudible acoustic signals from the earpiece speaker to "illuminate" the user's face and authenticates the user by the unique features extracted from the echoes bouncing off the 3D facial contour. To combat changes in phone-holding poses thus echoes, a Convolutional Neural Network (CNN) is trained to extract reliable acoustic features, which are further combined with visual facial landmark locations to feed a binary Support Vector Machine (SVM) classifier for final authentication. Because the echo features depend on 3D facial geometries, EchoPrint is not easily spoofed by images or videos like 2D visual face recognition systems. It needs only commodity hardware, thus avoiding the extra costs of special sensors in solutions like FaceID. Experiments with 62 volunteers and non-human objects such as images, photos, and sculptures show that EchoPrint achieves 93.75% balanced accuracy and 93.50% F-score, while the average precision is 98.05%, and no image/video based attack is observed to succeed in spoofing. Bing Zhou 0001, Jay Lohokare, Ruipeng Gao, Fan Ye 0003 |
MobiCom | 3 |
| 2017 | Knitter: Fast, resilient single-user indoor floor plan constructionabstractLacking of floor plans is a fundamental obstacle to ubiquitous indoor location-based services. Recent work have made significant progress to accuracy, but they largely rely on slow crowdsensing that may take weeks or even months to collect enough data. In this paper, we propose Knitter that can generate accurate floor maps by a single random user's one hour data collection efforts. Knitter extracts high quality floor layout information from single images, calibrates user trajectories and filters outliers. It uses a multi-hypothesis map fusion framework that updates landmark positions/orientations and accessible areas incrementally according to evidences from each measurement. Our experiments on 3 different large buildings and 30+ users show that Knitter produces correct map topology, and 90-percentile landmark location and orientation errors of 3 ~ 5m and 4 ~ 6°, comparable to the state-of-the-art at more than 20× speed up: data collection can finish in about one hour even by a novice user trained just a few minutes. Ruipeng Gao, Bing Zhou 0001, Fan Ye 0003, Yizhou Wang 0001 |
INFOCOM | 1 |
| 2017 | Demo: Acoustic Sensing Based Indoor Floor Plan Construction Using SmartphonesabstractThis demo presents BatMapper, an acoustics sensing technology for fast, fine-grained and low cost floor plan construction. BatMapper operates by emitting sound signal and capturing its reflections by two microphones on smartphones. We develop robust probabilistic echo-object association and outlier removal algorithms to identify the correspondence between distances and objects, thus the geometry of corridors. We compensate minute hand sway movements to identify small surface recessions, thus detecting doors automatically. Additionally, we leverage structure cues in indoor environments for user trace calibration. The demo will enable any person to hold the smartphone and walk along a corridor to map the corridor shape and detect doors in real-time. Bing Zhou 0001, Mohammed Elbadry, Ruipeng Gao, Fan Ye 0003 |
MobiCom | 3 |
| 2017 | BatMapper: Acoustic Sensing Based Indoor Floor Plan Construction Using SmartphonesabstractThe lack of digital floor plans is a huge obstacle to pervasive indoor location based services (LBS). Recent floor plan construction work crowdsources mobile sensing data from smartphone users for scalability. However, they incur long time (e.g., weeks or months) and tremendous efforts in data collection, and many rely on images thus suffering technical and privacy limitations. In this paper, we propose BatMapper, which explores a previously untapped sensing modality -- acoustics -- for fast, fine grained and low cost floor plan construction. We design sound signals suitable for heterogeneous microphones on commodity smartphones, and acoustic signal processing techniques to produce accurate distance measurements to nearby objects. We further develop robust probabilistic echo-object association, recursive outlier removal and probabilistic resampling algorithms to identify the correspondence between distances and objects, thus the geometry of corridors and rooms. We compensate minute hand sway movements to identify small surface recessions, thus detecting doors automatically. Experiments in real buildings show BatMapper achieves 1-2cm distance accuracy in ranges up around 4m; a 2-3 minute walk generates fine grained corridor shapes, detects doors at 92% precision and 1~2m location error at 90-percentile; and tens of seconds of measurement gestures produce room geometry with errors <0.3m at 80-percentile, at 1-2 orders of magnitude less data amounts and user efforts. Bing Zhou 0001, Mohammed Elbadry, Ruipeng Gao, Fan Ye 0003 |
MobiSys | 3 |
| 2017 | Modeling and Evaluation of the Incentive Scheme in "E-photo"
Shijie Ni, Zhuorui Yong, Ruipeng Gao |
MSN | 3 |
| 2017 | BatTracker: High Precision Infrastructure-free Mobile Device Tracking in Indoor EnvironmentsabstractContinuous tracking of the device location in 3D space is a popular form of user input, especially for virtual/augmented reality (VR/AR), video games and health rehabilitation. Conventional inertial based approaches are well known for inaccuracy caused by large error drifts. Computer vision approaches can produce accuracy tracking but have privacy concerns and are subject to lighting conditions and computation complexity. Recent work exploits accurate acoustic distance measurements for high precision tracking. However, they require additional hardware (e.g., multiple external speakers), which adds to the costs and installation efforts, thus limiting the convenience and usability. In this paper, we propose BatTracker, which incorporates inertial and acoustic data for robust, high precision and infrastructure-free tracking in indoor environments. BatTracker leverages echoes from nearby objects and uses distance measurements from them to correct error accumulation in inertial based device position prediction. It incorporates Doppler shifts and echo amplitudes to reliably identify the association between echoes and objects despite noisy signals from multi-path reflection and cluttered environment. A probabilistic algorithm creates, prunes and evolves multiple hypotheses based on measurement evidences to accommodate uncertainty in device position. Experiments in real environments show that BatTracker can track a mobile device's movements in 3D space at sub-cm level accuracy, comparable to the state-of-the-art infrastructure based approaches, while eliminating the needs of any additional hardware. Bing Zhou 0001, Mohammed Elbadry, Ruipeng Gao, Fan Ye 0003 |
SenSys | 3 |
| 2017 | Smartphone-Based Real Time Vehicle Tracking in Indoor Parking StructuresabstractAlthough location awareness and turn-by-turn instructions are prevalent outdoors due to GPS, we are back into the darkness in uninstrumented indoor environments such as underground parking structures. We get confused, disoriented when driving in these mazes, and frequently forget where we parked, ending up circling back and forth upon return. In this paper, we propose VeTrack, asmartphone-only system that tracks the vehicle’s location in real time using the phone’s inertial sensors. It does not require any environment instrumentation or cloud backend. It uses a novel “shadow” trajectory tracing method to accurately estimate phone’s and vehicle’s orientations despite their arbitrary poses and frequent disturbances. We develop algorithms in a Sequential Monte Carlo framework to represent vehicle states probabilistically, and harness constraints by the garage map and detected landmarks to robustly infer the vehicle location. We also find landmark (e.g., speed bumps, turns) recognition methods reliable against noises, disturbances from bumpy rides, and even hand-held movements. We implement a highly efficient prototype and conduct extensive experiments in multiple parking structures of different sizes and structures, and collect data with multiple vehicles and drivers. We find that VeTrack can estimate the vehicle’s real time location with almost negligible latency, with error of$2\sim 4$parking spaces at the 80th percentile. Ruipeng Gao, Mingmin Zhao, Fan Ye 0003, Yizhou Wang 0001, Guojie Luo |
IEEE Trans. Mob. Comput. | 1 |
| 2016 | VeMap: Indoor Road Map Construction via Smartphone-Based Vehicle TrackingabstractSince GPS signal is not applicable indoors, vehicle tracking has proven a hassle in underground parking structures. Recent solutions highly rely on floor map to constraint inertial sensors noises. In this paper, we propose VeMap, a road map construction system using only smartphones inside vehicles. It saves effort-intensive and time-consuming business negotiations with building operators, and expensive personnel cost to gather such data. It fuses multiple sensors to calibrate inertial noises, and uses Dynamic Time Warping to align multiple trajectories. We represent the floor plan with occupancy grid mapping, and explore a vision-mobile joint algorithm to extract its skeleton and form the road map. VeMap is tested in a 250mx90m parking structure, and it can be directly used for driving navigation to free parking spaces. Ruipeng Gao, Guojie Luo, Fan Ye 0003 |
GLOBECOM | 1 |
| 2016 | Sextant: Towards Ubiquitous Indoor Localization Service by Photo-Taking of the EnvironmentabstractMainstream indoor localization technologies rely on RF signatures that require extensive human efforts to measure and periodically recalibrate signatures. The progress to ubiquitous localization remains slow. In this study, we explore Sextant, an alternative approach that leverages environmental reference objects such as store logos. A user uses a smartphone to obtain relative position measurements to such static reference objects for the system to triangulate the user location. Sextant leverages image matching algorithms to automatically identify the chosen reference objects by photo-taking, and we propose two methods to systematically address image matching mistakes that cause large localization errors. We formulate the benchmark image selection problem, prove its NP-completeness, and propose a heuristic algorithm to solve it. We also propose a couple of geographical constraints to further infer unknown reference objects. To enable fast deployment, we propose a lightweight site survey method for service providers to quickly estimate the coordinates of reference objects. Extensive experiments have shown that Sextant prototype achieves 2-5 m accuracy at 80-percentile, comparable to the industry state-of-the-art, while covering a 150 x 75 m mall and 300 x 200m train station requires a one time investment of only 2-3 man-hours from service providers. Ruipeng Gao, Fan Ye 0003, Guojie Luo, Kaigui Bian, Yizhou Wang 0001, Tao Wang 0004, Xiaoming Li 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2016 | Multi-Story Indoor Floor Plan Reconstruction via Mobile CrowdsensingabstractThe lack of floor plans is a critical reason behind the current sporadic availability of indoor localization service. Service providers have to go through effort-intensive and time-consuming business negotiations with building operators, or hire dedicated personnel to gather such data. In this paper, we propose Jigsaw, a floor plan reconstruction system that leverages crowdsensed data from mobile users. It extracts the position, size, and orientation information of individual landmark objects from images taken by users. It also obtains the spatial relation between adjacent landmark objects from inertial sensor data, then computes the coordinates and orientations of these objects on an initial floor plan. By combining user mobility traces and locations where images are taken, it produces complete floor plans with hallway connectivity, room sizes, and shapes. It also identifies different types of connection areas (e.g., escalators and stairs) between stories, and employs a refinement algorithm to correct detection errors. Our experiments on three stories of two large shopping malls show that the 90-percentile errors of positions and orientations of landmark objects are about 1~2m and 5~9°, while the hallway connectivity and connection areas between stories are 100 percent correct. Ruipeng Gao, Mingmin Zhao, Fan Ye 0003, Guojie Luo, Yizhou Wang 0001, Kaigui Bian, Tao Wang 0004, Xiaoming Li 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2015 | VeTrack: Real Time Vehicle Tracking in Uninstrumented Indoor EnvironmentsabstractAlthough location awareness and turn-by-turn instructions are prevalent outdoors due to GPS, we are back into the darkness in uninstrumented indoor environments such as underground parking structures. We get confused, disoriented when driving in these mazes, and frequently forget where we parked, ending up circling back and forth upon return.In this paper, we propose VeTrack, a smartphone-only system that tracks the vehicle's location in real time using the phone's inertial sensors. It does not require any environment instrumentation or cloud backend. It uses a novel "shadow" tracing method to accurately estimate the vehicle's trajectories despite arbitrary phone/vehicle poses and frequent disturbances. We develop algorithms in a Sequential Monte Carlo framework to represent vehicle states probabilistically, and harness constraints by the garage map and detected landmarks to robustly infer the vehicle location. We also find landmark (e.g., speed bumps, turns) recognition methods reliable against noises, disturbances from bumpy rides and even hand-held movements. We implement a highly efficient prototype and conduct extensive experiments in multiple parking structures of different sizes and structures, with multiple vehicles and drivers. We find that VeTrack can estimate the vehicle's real time location with almost negligible latency, with error of 2-4 parking spaces at 80-percentile. Mingmin Zhao, Ruipeng Gao, Fan Ye 0003, Yizhou Wang 0001, Guojie Luo |
SenSys | 3 |
| 2014 | Smartphone indoor localization by photo-taking of the environmentabstractExisting mainstream indoor localization technologies mainly rely on RF signatures and thus incur significant and recurring labor cost to measure the time-varying signature map. We have proposed a smartphone localization system using the embedded gyroscope for triangulation from nearby physical features (e.g., store logos) recognized from photo-taking. It requires a much reduced and one-time measurement, while incurs uncertain localization errors. In this paper, we propose two methods to systematically address image matching errors that cause unrecognized physical features and large errors in our system. We formulate the optimal benchmark image selection problem and propose a heuristic algorithm that finds the best benchmark images for high matching accuracy. We propose a couple of geographical constraints to further infer unknown physical features based on the observation that the features chosen by the user are close together. Experiments in a 150 × 75m shopping mall, 300 × 200m train station show that dramatically we cut down both maximum and general localization errors, and achieve 2-8m accuracy at 80-percentile even with only one benchmark image on the phone. Ruipeng Gao, Fan Ye 0003, Tao Wang 0004 |
ICC | 1 |
| 2014 | Towards ubiquitous indoor localization service leveraging environmental physical featuresabstractMainstream indoor localization technologies rely on RF signatures that require extensive human efforts to measure and periodically re-calibrate. Although recent crowdsourcing based work has started to address the issue, incentives are still lacking for wide user adoption. Thus the progress to ubiquitous localization remains slow. In this paper, we explore an alternative approach that leverages environmental physical features such as store logos or wall posters. A user uses a smartphone to obtain relative position measurements to such static reference points for the system to triangulate the user location. We study the principle of such localization, determine the suitable sensor, and devise guidelines for the user to choose reference points for better accuracy. To enable fast deployment, we propose a lightweight site survey method for service providers to quickly estimate the coordinates of reference points. We incorporate and enhance image matching algorithms with a heuristic technique to automatically identify chosen reference points at high accuracy. Extensive experiments have shown that the prototype achieves 4-5m accuracy at 80-percentile, comparable to the industry state-of-the-art, while covering a 150×75m mall and 300×200m train station requires a one time investment of only 2-3 man-hours from service providers. Ruipeng Gao, Kaigui Bian, Fan Ye 0003, Tao Wang 0004, Yizhou Wang 0001, Xiaoming Li 0001 |
INFOCOM | 2 |
| 2014 | Jigsaw: indoor floor plan reconstruction via mobile crowdsensingabstractThe lack of floor plans is a critical reason behind the current sporadic availability of indoor localization service. Service providers have to go through effort-intensive and time-consuming business negotiations with building operators, or hire dedicated personnel to gather such data. In this paper, we propose Jigsaw, a floor plan reconstruction system that leverages crowdsensed data from mobile users. It extracts the position, size and orientation information of individual landmark objects from images taken by users. It also obtains the spatial relation between adjacent landmark objects from inertial sensor data, then computes the coordinates and orientations of these objects on an initial floor plan. By combining user mobility traces and locations where images are taken, it produces complete floor plans with hallway connectivity, room sizes and shapes. Our experiments on 3 stories of 2 large shopping malls show that the 90-percentile errors of positions and orientations of landmark objects are about 1~2m and 5~9°, while the hallway connectivity is 100% correct. Ruipeng Gao, Mingmin Zhao, Fan Ye 0003, Yizhou Wang 0001, Kaigui Bian, Tao Wang 0004, Xiaoming Li 0001 |
MobiCom | 1 |
| 2014 | VeLoc: finding your car in the parking lotabstractWe present VeLoc, a smartphone-based vehicle localization approach that tracks the vehicle's parking location without GPS or WiFi signals. It uses only the embedded accelerometer and gyroscope sensors. VeLoc harnesses constraints imposed by the map and landmarks (e.g., speed bumps) recognized from inertial data, employs a Bayesian filtering framework to estimate the location of the vehicle. We have conducted experiments in three parking structures of different sizes and configurations, using three vehicles and three kinds of driving styles. We find that VeLoc can always localize the vehicle within 10m, which is sufficient for the driver to trigger a honk using the car key. Mingmin Zhao, Ruipeng Gao, Jiaxu Zhu, Fan Ye 0003, Yizhou Wang 0001, Kaigui Bian, Guojie Luo, Ming Zhang 0004 |
SenSys | 2 |
| 2014 | Interference Self-Coordination: A Proposal to Enhance Reliability of System-Level Information in OFDM-Based Mobile Networks via PCI PlanningabstractFor system-level information in OFDM-based mobile networks, the full-cell coverage requirement raises the issue of Inter-Cell Interference (ICI) in physical control channels. Unfortunately, conventional approaches (e.g., scheduling, beamforming) cannot effectively deal with this kind of ICI because of configuration limitations in physical control channels. To solve this challenge, this study defines a new concept of `Interference Self-Coordination (ISC)' for enhancing this system-level information delivery. Specifically, we apply ISC to relieve the ICI effect over the Physical Downlink Control Channel (PDCCH), the most important physical control channel in Long-Term Evolution (LTE/LTE-Advanced) networks, in which Physical Cell Identifier (PCI) planning naturally serves as a self-coordination mechanism. We prove that PCI planning is the most effective solution to relieve ICI over PDCCHs, under the constraint of non-orthogonal control regions between neighboring cells. Since PCI planning is an NP-Complete problem, heuristic-strategy algorithms are devised for implementation, which are self-adaptive to dynamic configurations of the network and system, and compatible with any applicable ICI solutions. The numerical analysis and system-level simulation experiment results demonstrate that this proposal can increase overall system throughput by up to 17%, and also improve cell-edge throughput by up to 42%. In the worst 5% of sectors, the proposal can still obtain 45% system throughput and 200% cell-edge throughput gain. Hemin Yang, Anpeng Huang, Ruipeng Gao, Tammy Chang, Linzhen Xie |
IEEE Trans. Wirel. Commun. | 3 |
| 2013 | A solution to relieve ICI effects on system control information in OFDM-based mobile networks: Conflict coordination on PDCCH via PCI planningabstractIn OFDM-based mobile networks, the full-cell coverage of SCI (System-level Control Information) must be guaranteed for a user that can communicate with its eNB (evolved Node Base station) system. The requirement of SCI coverage causes severe ICI (Inter-Cell Interference) effect. Furthermore, the negative effect is intensified owing to single-frequency networking and physical control region configuration in OFDM-based networks. To enhance the reliability of SCI under the coverage constraint, we investigated how SCI is carried in the Physical Downlink Control Channel (PDCCH) in the common search space, and proposed a solution of Conflict Coordination on PDCCH via PCI (Physical Cell Identifier) Planning (CCP3). In our proposal, PCI planning is used to relieve the ICI effect of the SCI. Thus, a heuristic algorithm is developed for the PCI planning since it is an NP-Complete optimization problem. The link-level simulation experiments and numerical analysis results demonstrate that the proposed CCP3can improve effective SINR (Signal to Interference plus Noise Ratio) of PDCCH significantly, and reduce RE (Resource Element)-occupation conflict probability of PDCCH around 50%. To the best of our knowledge, this study is the first effort to strengthen the SCI reliability from the view of networking optimization. Hemin Yang, Ruipeng Gao, Anpeng Huang, Linzhen Xie |
ICC | 2 |