VLDB 2026 Research / reviewers in the wild / expert
Lu Su 0001
dblp:63/4152-1
· DBLP profile ↗
145ranked-venue papers
7as first author
44since 2021 · last 2026
0000-0001-7223-543XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 77 · 5 first-author · 28 since 2021Databases, data management, data science and information retrieval · 31 · 5 since 2021Artificial intelligence and machine learning · 28 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 1 since 2021Systems, architecture and hardware · 12 · 2 since 2021Security and privacy · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ISAC Micro-Doppler Sensing of UAVs: Cramér-Rao Bound Analysis and Experiment Demonstration
Henglin Pu, Lu Su 0001, Husheng Li |
ICC | 2 |
| 2026 | Towards Privacy-Preserving and Heterogeneity-aware Split Federated Learning via Probabilistic MaskingabstractSplit Federated Learning (SFL) has emerged as an efficient alternative to traditional Federated Learning (FL) by reducing client-side computation through model partitioning. However, exchanging of intermediate activations and model updates introduces significant privacy risks, especially from data reconstruction attacks that recover original inputs from intermediate representations. Existing defenses using noise injection often degrade model performance. To overcome these challenges, we present PM-SFL, a scalable and privacy-preserving SFL framework that incorporates Probabilistic Mask training to add structured randomness without relying on explicit noise. This mitigates data reconstruction risks while maintaining model utility. To address data heterogeneity, PM-SFL employs personalized mask learning that tailors submodel structures to each client's local data. For system heterogeneity, we introduce a layer-wise knowledge compensation mechanism, enabling clients with varying resources to participate effectively under adaptive model splitting. Theoretical analysis confirms its privacy protection, and experiments on image and wireless sensing tasks demonstrate that PM-SFL consistently improves accuracy, communication efficiency, and robustness to privacy attacks, with particularly strong performance under data and system heterogeneity. Feijie Wu, Chenglin Miao, Tianchun Li, Qiming Cao, Jing Gao 0004, Lu Su 0001 |
KDD (1) | 8 |
| 2026 | Defending Autonomous Driving Perception against Adversarial Object-Based Attacks via Motion PlanningabstractAutonomous vehicles (AVs) rely on perception systems to detect surrounding objects using sensors such as cameras, LiDAR (Light Detection and Ranging), and millimeter-wave (mmWave) radar. However, recent studies have shown that attackers can deceive these systems by strategically placing adversarial objects (e.g., color patches, cardboard, or metal foil) in the driving environment. These attacks pose serious safety risks, yet existing defenses primarily focus on individual sensor modalities and lack generalizability across different sensing systems. To address this gap, we propose the first generalized defense mechanism capable of mitigating various attacks using adversarial objects. Our approach integrates real-time attack detection with trajectory adaptation, guiding the victim AV to positions where the attack is less effective. The defense mechanism combines a deep reinforcement learning (DRL)-based motion planning model, which dynamically adjusts the AV’s trajectory, with an uncertainty-aware filtering scheme that refines perception outputs to enhance detection robustness. Extensive experiments in both simulated and real-world environments demonstrate that our defense mechanism effectively mitigates adversarial object-based attacks across different sensing modalities and sensor fusion while maintaining safe and smooth driving behavior. Zihao Liu 0001, Yan Zhang 0133, Yi Zhu 0012, Lu Su 0001, Chunming Qiao, Chenglin Miao |
SenSys | 4 |
| 2026 | OTFS-ISAC System With Sub-Nyquist ADC Sampling RateabstractIntegrated sensing and communication (ISAC) has emerged as a pivotal technology for next-generation wireless communication and radar systems, enabling high-resolution sensing and high-throughput communication with shared spectrum and hardware. However, achieving a fine radar resolution often requires high-rate analog-to-digital converters (ADCs) and substantial storage, making it both expensive and impractical for many commercial applications. To address these challenges, this paper proposes an orthogonal time frequency space (OTFS)-based ISAC architecture that operates at reduced ADC sampling rates, yet preserves accurate radar estimation and supports simultaneous communication. The proposed architecture introduces pilot symbols directly in the delay-Doppler (DD) domain to leverage the transformation mapping between the DD and time-frequency (TF) domains to keep selected subcarriers active while others are inactive, allowing the radar receiver to exploit under-sampling aliasing and recover the original DD signal at much lower sampling rates. To further enhance the radar accuracy, we develop an iterative interference estimation and cancellation algorithm that mitigates data symbol interference. We propose a code-based spreading technique that distributes data across the DD domain to preserve the maximum unambiguous radar sensing range. For communication, we implement a complete transceiver pipeline optimized for reduced sampling rate system, including synchronization, channel estimation, and iterative data detection. Experimental results from a software-defined radio (SDR)-based testbed confirm that our method substantially lowers the required sampling rate without sacrificing radar sensing performance and ensures reliable communication. Henglin Pu, Ajay Kumar 0012, Lu Su 0001, Husheng Li |
IEEE J. Sel. Areas Commun. | 4 |
| 2026 | Space-Time-Frequency Synthetic Integrated Sensing and Communication NetworksabstractIntegrated sensing and communication (ISAC) promises high spectral and power efficiencies by sharing waveforms, spectrum, and hardware across sensing and data links. Yet commercial cellular networks struggle to deliver fine angular, range, and Doppler resolution due to limited aperture, bandwidth, and coherent observation time. In this paper, we propose a space-time-frequency synthetic ISAC architecture that fuses observations from distributed transmitters and receivers across time intervals and frequency bands. We develop a unified signal model for multistatic and monostatic configurations, derive Cramer-Rao lower bounds (CRLBs) for the estimations of position and velocity. The analysis shows how spatial diversity, multiband operation, and observation scheduling impact the Fisher information. We also compare the estimation performance between a concentrated maximum likelihood estimator (MLE) and a two stage information fusion (TSIF) method that first estimates per-path delay and radial speed and then fuses them by solving a weighted nonlinear least-squares problem via the Gauss-Newton algorithm. Numerical results show that MLE approaches the CRLB in the high signal-to-noise ratio (SNR) regime, while the two stage method remains competitive at moderate to high SNR but degrades at low SNR. A central finding is that fully synthesized network processing is essential, as estimations by individual base stations (BSs) followed by fusion are consistently inferior and unstable at low SNR. This framework offers a practical guidance for upgrading existing communication infrastructure into dense sensing networks. Henglin Pu, Lu Su 0001, Husheng Li |
IEEE Trans. Wirel. Commun. | 3 |
| 2025 | Enhancing Non-line-of-sight ISAC with Position-Aware BeamformingabstractMillimeter-wave (mmWave) technology represents a promising avenue in integrated sensing and communication (ISAC), leveraging wide bandwidth to accommodate growing demands for high data-rate communication and high-resolution radar sensing. However, in non-line-of-sight (NLOS) scenarios, mmWave signals suffer from severe attenuation, and integrating precise radar sensing into bandwidth-limited communication systems remains an open problem. To address this, we propose an NLOS-ISAC system that jointly provides accurate target positioning and robust communication. Our approach synthesizes many narrowband signals into a virtual wideband radar via stepped-frequency techniques, eliminating the need for additional hardware. By fusing time-of-flight (ToF) measurements of multipath reflections with environmental maps, the system achieves submeter positioning accuracy for NLOS targets. Building on these position estimates, we then employ position-aware beamforming to significantly enhance the NLOS communication throughput. Extensive experimental results demonstrate that the proposed NLOS-ISAC system reliably achieves high-accuracy localization while improving data rates in challenging NLOS environments. Henglin Pu, Karim A. Said, Lingjia Liu 0001, Lu Su 0001, Husheng Li |
GLOBECOM | 5 |
| 2025 | OTFS-Based ISAC with Reduced Sampling Rate: Algorithm and ExperimentabstractIntegrated Sensing and Communication (ISAC) is emerging as a promising technique for future communication and radar systems. A high-resolution estimation is essential for radar applications, necessitating large bandwidth. However, a large bandwidth requires dedicated Analog-to-Digital Converters (ADCs) with high sampling rates and substantial storage capacities to fully capture waveform signals, making it prohibitively expensive and impractical for many commercial systems. Therefore, it is crucial to develop methods for radar estimation with reduced sampling rates. In this paper, we propose a method to achieve accurate radar estimations for Orthogonal Time Frequency Space (OTFS)-based ISAC systems with reduced sampling rates. Specifically, we design pilot symbols in the delay-Doppler domain to ensure that the pilots in the time-frequency (TF) domain occupy specific subcarriers, leaving the remaining ones inactive. By leveraging the aliasing effect caused by undersampling, the receiver captures all aliased active subcarrier information with a reduced ADC sampling rate, enabling the restoration of the original delay-Doppler signal. Finally, we introduce an iterative estimation and interference cancellation algorithm to mitigate the interference from data symbols and provide accurate estimation results. Extensive experiments demonstrate that our method effectively reduces the sampling rate without compromising the ISAC performance. Henglin Pu, Lu Su 0001, Husheng Li |
ICC | 3 |
| 2025 | Resource Allocation for OTFS with High MobilityabstractOrthogonal Time Frequency Space (OTFS) modulation has emerged as a promising waveform for next-generation wireless communications. OTFS demonstrates robust non-fading properties even under doubly-dispersive conditions, making it a preferred modulation scheme in complex scenarios involving high mobility. Despite its advantages, developing a low-complexity multiaccess OTFS protocol for environments with a large number of user terminals (UTs) and high mobility remains challenging due to severe multiuser interference (MUI) and rapidly changing channel states. In this paper, we propose a low-complexity algorithm to find a suboptimal solution for the joint resource block and power allocation problem in multiple access OTFS systems. Specifically, we decompose the resource block allocation and power allocation into two sub-problems. First, we perform the resource block allocation algorithm by considering the channel conditions. Then, we employ a two-stage water-filling algorithm to achieve a suboptimal power allocation. Numerical results indicate that our proposed resource allocation scheme achieves higher sum-rate than existing schemes in both 8-user and 16-user systems while reducing the complexity from exponential to linear in terms of the number of users. This makes our approach practical even in high mobility and larger user scenarios. Henglin Pu, Lu Su 0001, Husheng Li |
ICC | 3 |
| 2025 | Towards Federated RLHF with Aggregated Client Preference for LLMsabstractReinforcement learning with human feedback (RLHF) fine-tunes a pretrained large language model (LLM) using user preference data, enabling it to generate content aligned with human preferences. However, due to privacy concerns, users may be reluctant to share sensitive preference data. To address this, we propose utilizing Federated Learning (FL) techniques, allowing large-scale preference collection from diverse real-world users without requiring them to transmit data to a central server. Our federated RLHF methods (i.e., FedBis and FedBiscuit) encode each client’s preferences into binary selectors and aggregate them to capture common preferences. In particular, FedBiscuit overcomes key challenges, such as preference heterogeneity and reward hacking, through innovative solutions like grouping clients with similar preferences to reduce heterogeneity and using multiple binary selectors to enhance LLM output quality. To evaluate the performance of the proposed methods, we establish the first federated RLHF benchmark with a heterogeneous human preference dataset. Experimental results show that by integrating the LLM with aggregated client preferences, FedBis and FedBiscuit significantly enhance the professionalism and readability of the generated content. Feijie Wu, Xiaoze Liu, Haoyu Wang 0004, Lu Su 0001, Jing Gao 0004 |
ICLR | 5 |
| 2025 | RAM-Hand: Robust Acoustic Multi-Hand Pose Reconstruction Using a Microphone ArrayabstractUsing 3D hand poses as the input of user interfaces can enable many novel human-computer interaction applications. However, conventional solutions for precisely reconstructing the hand poses are either vision-based, which are compute-intensive and may cause privacy issues, or wearable devices-based, which are intrusive to users. In this paper, we propose RAM-Hand, a Robust Acoustic 3D Multi-Hand pose reconstruction system built on a microphone array. Our RAM-Hand system can support multiple hands and is designed to be highly adaptable to new scenarios even when training data is limited. Specifically, it should robustly accommodate variations in environment, subject, and hand positions. To achieve this, on one hand, we propose a customized signal processing pipeline to segment multiple hands' reflections and extract the features corresponding to each hand, then feed those features into a transformer-based neural network for precise pose reconstruction. On the other hand, to tackle the challenge that the training data is limited, we propose a series of data augmentation methods to generate virtual training data, and utilize contrastive learning to ensure our model behaves well on new subjects. We conduct extensive experiments on a real-world microphone array testbed to evaluate the performance of the proposed system. The results show that our RAM-Hand system can localize each hand joint with an average error of 10.71 mm, handle multiple hands, and generalize well to the above mentioned new scenarios. Henglin Pu, Qiming Cao, Tianci Liu 0003, Zhengxin Jiang, Hongfei Xue, Lu Su 0001 |
SenSys | 9 |
| 2025 | SigCan: Toward Reliable ToF Estimation Leveraging Multipath Signal Cancellation on Commodity WiFi DevicesabstractThe widespread deployment of WiFi infrastructure has facilitated the development of Time-of-Flight (ToF) based sensing applications. ToF estimation, however, is a challenging task due to the complexity of multipath effect. In this paper, we propose a phase difference based method for ToF estimation and uncover the potential of signal cancellation to mitigate the impact of multipath and noise on phase differences among subcarriers. To separate the moving target path from the complex multipath for ToF estimation, we suggest employing specific elimination methods tailored to the characteristics of different signal components. For dynamic multipath, we observe that when a given subcarrier propagates along two paths to the receiver, with path lengths differing by half a wavelength, the phase difference introduced by these two paths cancels each other out. Therefore, we propose two metrics to identify signals that satisfy this condition, utilizing both frequency diversity and spatial diversity. Additionally, we propose leveraging time diversity to eliminate the static multipath component and reduce the impact of noise. We implemented the methods with off-the-shelf WiFi devices and achieved mean errors of 15.36 cm and 21.05 cm for distance estimation in outdoor and indoor scenarios, outperforming state-of-the-art ToF estimation method by 50% error reduction. Yang Li 0162, Dan Wu 0007, Leye Wang, Lu Su 0001, Daqing Zhang 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Multiple-Metric Frame Optimization for OTFS with Experimental DemonstrationabstractOrthogonal Time Frequency Space (OTFS) modulation emerges as a promising waveform for next generation wireless communications. The efficacy of OTFS, in both communication and sensing realms, is critically dependent on the design of its frame structure. This study delves into the design and optimization of a pilot-symbol-aided OTFS frame, with an emphasis on enhancing spectrum efficiency, minimizing the Peak-to-Average Power Ratio (PAPR), and reducing bit error rate (BER). Specifically, we first analyze the impact of specific channel characteristics on BER. Subsequently, we engage in a detailed exploration of the interplay between frame parameters and performance metrics, namely the spectrum efficiency, PAPR, and BER. We eventually propose an optimization framework for embedded-pilot OTFS frames, aimed at attaining optimal performance across these metrics. Through both simulations and experiments, we demonstrate that the optimized OTFS frame architecture offers significant improvements in BER and our optimization framework provides profound insights into determining the optimal frame parameters tailored to specific use cases, with an emphasis on varying priorities. Henglin Pu, Lu Su 0001, Husheng Li |
GLOBECOM | 3 |
| 2024 | Towards Poisoning Fair RepresentationsabstractFair machine learning seeks to mitigate model prediction bias against certain demographic subgroups such as elder and female.
Recently, fair representation learning (FRL) trained by deep neural networks has demonstrated superior performance, whereby representations containing no demographic information are inferred from the data and then used as the input to classification or other downstream tasks.
Despite the development of FRL methods, their vulnerability under data poisoning attack, a popular protocol to benchmark model robustness under adversarial scenarios, is under-explored. Data poisoning attacks have been developed for classical fair machine learning methods which incorporate fairness constraints into shallow-model classifiers.
Nonetheless, these attacks fall short in FRL due to notably different fairness goals and model architectures.
This work proposes the first data poisoning framework attacking FRL. We induce the model to output unfair representations that contain as much demographic information as possible by injecting carefully crafted poisoning samples into the training data.
This attack entails a prohibitive bilevel optimization, wherefore an effective approximated solution is proposed. A theoretical analysis on the needed number of poisoning samples is derived and sheds light on defending against the attack. Experiments on benchmark fairness datasets and state-of-the-art fair representation learning models demonstrate the superiority of our attack. Tianci Liu 0003, Haoyu Wang 0004, Feijie Wu, Hengtong Zhang, Pan Li 0005, Lu Su 0001, Jing Gao 0004 |
ICLR | 6 |
| 2024 | Malicious Attacks against Multi-Sensor Fusion in Autonomous DrivingabstractMulti-sensor fusion has been widely used by autonomous vehicles (AVs) to integrate the perception results from different sensing modalities including LiDAR, camera and radar. Despite the rapid development of multi-sensor fusion systems in autonomous driving, their vulnerability to malicious attacks have not been well studied. Although some prior works have studied the attacks against the perception systems of AVs, they only consider a single sensing modality or a camera-LiDAR fusion system, which can not attack the sensor fusion system based on LiDAR, camera, and radar. To fill this research gap, in this paper, we present the first study on the vulnerability of multi-sensor fusion systems that employ LiDAR, camera, and radar. Specifically, we propose a novel attack method that can simultaneously attack all three types of sensing modalities using a single type of adversarial object. The adversarial object can be easily fabricated at low cost, and the proposed attack can be easily performed with high stealthiness and flexibility in practice. Extensive experiments based on a real-world AV testbed show that the proposed attack can continuously hide a target vehicle from the perception system of a victim AV using only two small adversarial objects. Yi Zhu 0012, Chenglin Miao, Hongfei Xue, Yunnan Yu, Lu Su 0001, Chunming Qiao |
MobiCom | 5 |
| 2024 | FIARSE: Model-Heterogeneous Federated Learning via Importance-Aware Submodel ExtractionabstractIn federated learning (FL), accommodating clients' varied computational capacities poses a challenge, often limiting the participation of those with constrained resources in global model training. To address this issue, the concept of model heterogeneity through submodel extraction has emerged, offering a tailored solution that aligns the model's complexity with each client's computational capacity. In this work, we propose Federated Importance-Aware Submodel Extraction (FIARSE), a novel approach that dynamically adjusts submodels based on the importance of model parameters, thereby overcoming the limitations of previous static and dynamic submodel extraction methods. Compared to existing works, the proposed method offers a theoretical foundation for the submodel extraction and eliminates the need for additional information beyond the model parameters themselves to determine parameter importance, significantly reducing the overhead on clients. Extensive experiments are conducted on various datasets to showcase the superior performance of the proposed FIARSE. Feijie Wu, Yaqing Wang 0001, Tianci Liu 0003, Lu Su 0001, Jing Gao 0004 |
NeurIPS | 5 |
| 2024 | mmCLIP: Boosting mmWave-based Zero-shot HAR via Signal-Text AlignmentabstractMillimeter-wave (mmWave) based human activity recognition (HAR) systems have demonstrated promising performance in various applications, leveraging the power of deep neural networks. However, these systems are suffering from the scarcity of available mmWave data for model training. To address this challenge, we explore the possibility of transferring knowledge from large AI models built on massive text and visual data to enhance the generalizability of mmWave-based HAR models. Towards this end, we introduce mmCLIP, a novel system that aligns mmWave signal space and text space to facilitate zero-shot recognition for unseen activities. To enable this alignment, we employ cross-modality signal synthesis to augment mmWave signal data using large human mesh datasets and design an activity attribute decomposition and recomposition approach to characterize the semantic interconnections among activities. We conducted extensive experiments to demonstrate the effectiveness of our proposed framework. Qiming Cao, Hongfei Xue, Tianci Liu 0003, Haoyu Wang 0004, Xincheng Zhang, Lu Su 0001 |
SenSys | 7 |
| 2024 | Towards Efficient Heterogeneous Multi-Modal Federated Learning with Hierarchical Knowledge DisentanglementabstractMulti-modal sensing systems are becoming increasingly common in real-world applications like human activity recognition (HAR). To enable knowledge sharing among individuals, Federated Learning (FL) offers a solution as a distributed machine learning paradigm that retains user data locally, thereby safeguarding privacy. However, existing heterogeneous multi-modal Federated Learning (MMFL) solutions have yet to fully utilize all the potential knowledge-sharing opportunities, as they fail to capture fundamental common knowledge that is independent of both modality and client. In this paper, we propose Federated Hierarchical Knowledge Disentanglement (FedHKD), a new sensing system for heterogeneous multi-modal federated learning. FedHKD introduces a multi-stage training paradigm based on hierarchical knowledge disentanglement at both the modality and client levels. This design enhances collaboration among modality-heterogeneous clients while maintaining low storage overhead and high adaptation flexibility to new sensing modalities. Our evaluation of two public real-world multi-modal HAR datasets and a self-collected dataset demonstrates that FedHKD outperforms state-of-the-art baselines by up to 4.85% in accuracy while saving up to 2.29× in storage. Additionally, when adapting to new sensing modalities, it reduces communication overhead by up to 4.62×. Haoyu Wang 0004, Feijie Wu, Tianci Liu 0003, Qiming Cao, Lu Su 0001 |
SenSys | 6 |
| 2024 | Optimizing Long-Term Efficiency and Fairness in Ride-Hailing Under Budget Constraint via Joint Order Dispatching and Driver RepositioningabstractRide-hailing platforms (e.g., Uber and Didi Chuxing) have become increasingly popular in recent years.Efficiencyhas always been an important metric for such platforms. However, only focusing on efficiency inevitably ignores thefairnessof driver incomes, which could impair the sustainability of ride-hailing systems. To optimize such two essential objectives,order dispatchinganddriver repositioningplay an important role, as they impact not only the immediate, but also the future order-serving outcomes of drivers. In practice, the platform offers monetary incentives to drivers for completing the repositioning and has a budget for the repositioning cost. Therefore, in this paper, we aim to exploit joint order dispatching and driver repositioning to optimize both long-term efficiency and fairness in ride-hailing under the budget constraint. To this end, we propose JDRCL, a novel multi-agent reinforcement learning framework, which integrates a group-based action representation that copes with the variable action space, and a primal-dual iterative training algorithm to learn a constraint-satisfying policy that maximizes both the worst and the overall incomes of drivers. Furthermore, we prove the asymptotic convergence rate of our training algorithm. Extensive experiments based on three real-world ride-hailing order datasets show that JDRCL outperforms state-of-the-art baselines on both efficiency and fairness. Haiming Jin, Zhaoxing Yang, Lu Su 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | PhyFinAtt: An Undetectable Attack Framework Against PHY Layer Fingerprint-Based WiFi AuthenticationabstractWiFi connection has been suffering from MAC forgery attacks due to the loose authentication mechanism between access points (APs) and clients. To address this problem, the physical (PHY) layer information-based fingerprint has been adopted for safe WiFi authentication. Since such a fingerprint is constant and unique for each specific network interface card (NIC), it can effectively prevent MAC forgery attacks. However, the PHY layer information-based fingerprint is still vulnerable to malicious attacks as it is extracted from Channel State Information (CSI), and its stability can be affected by the wireless environment. In this paper, we propose a novel undetectable attack framework, called PhyFinAtt, base on which the attacker can undermine the stability of the PHY layer-based authentication fingerprints through human movement and further attack the WiFi authentication protocols. Specifically, we first demonstrate that human movement at a designated location can affect the PHY fingerprint. We then illustrate the impact of human movement on the PHY fingerprint and the relationship between the movement and the channel quality to ensure that the PHY fingerprint is destroyed by the movement in an undetected way without affecting normal communication. Extensive experiments in real-world scenarios show that our proposed attack can effectively disrupt the stability of the PHY fingerprints and significantly degrade the performance of the authentication protocols based on such fingerprints. To the best of our knowledge, this is the first study on effective attacks against the PHY information-based WiFi authentication protocols. Furthermore, we also present a practical defense mechanism without involving any additional equipment to mitigate attacks similar to PhyFinAtt. Jinyang Huang, Bin Liu 0016, Chenglin Miao, Xiang Zhang 0011, Jianchun Liu, Lu Su 0001, Zhi Liu 0002, Yu Gu 0003 |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Towards Smartphone-based 3D Hand Pose Reconstruction Using Acoustic SignalsabstractAccurately reconstructing 3D hand poses is a pivotal element for numerous Human-Computer Interaction applications. In this work, we propose SonicHand, the first smartphone-based 3D hand pose reconstruction system using purely inaudible acoustic signals. SonicHand incorporates signal processing techniques and a deep learning framework to address a series of challenges. First, it encodes the topological information of the hand skeleton as prior knowledge and utilizes a deep learning model to realistically and smoothly reconstruct the hand poses. Second, the system employs adversarial training to enhance the generalization ability of our system to be deployed in a new environment or for a new user. Third, we adopt a hand tracking method based on channel impulse response estimation. It enables our system to handle the scenario where the hand performs gestures while moving arbitrarily as a whole. We conduct extensive experiments on a smartphone testbed to demonstrate the effectiveness and robustness of our system from various dimensions. The experiments involve 10 subjects performing up to 12 different hand gestures in three distinctive environments. When the phone is held in one of the user’s hands, the proposed system can track joints with an average error of 18.64 mm. Chenglin Miao, Qiming Cao, Haoyu Wang 0004, Ke Sun 0012, Hongfei Xue, Lu Su 0001 |
ACM Trans. Sens. Networks | 9 |
| 2023 | SimFair: A Unified Framework for Fairness-Aware Multi-Label ClassificationabstractRecent years have witnessed increasing concerns towards unfair decisions made by machine learning algorithms. To improve fairness in model decisions, various fairness notions have been proposed and many fairness-aware methods are developed. However, most of existing definitions and methods focus only on single-label classification. Fairness for multi-label classification, where each instance is associated with more than one labels, is still yet to establish. To fill this gap, we study fairness-aware multi-label classification in this paper. We start by extending Demographic Parity (DP) and Equalized Opportunity (EOp), two popular fairness notions, to multi-label classification scenarios. Through a systematic study, we show that on multi-label data, because of unevenly distributed labels, EOp usually fails to construct a reliable estimate on labels with few instances. We then propose a new framework named Similarity s-induced Fairness (sγ -SimFair). This new framework utilizes data that have similar labels when estimating fairness on a particular label group for better stability, and can unify DP and EOp. Theoretical analysis and experimental results on real-world datasets together demonstrate the advantage of sγ -SimFair over existing methods on multi-label classification tasks. Tianci Liu 0003, Haoyu Wang 0004, Yaqing Wang 0001, Xiaoqian Wang 0001, Lu Su 0001, Jing Gao 0004 |
AAAI | 5 |
| 2023 | TileMask: A Passive-Reflection-based Attack against mmWave Radar Object Detection in Autonomous DrivingabstractIn autonomous driving, millimeter wave (mmWave) radar has been widely adopted for object detection because of its robustness and reliability under various weather and lighting conditions. For radar object detection, deep neural networks (DNNs) are becoming increasingly important because they are more robust and accurate, and can provide rich semantic information about the detected objects, which is critical for autonomous vehicles (AVs) to make decisions. However, recent studies have shown that DNNs are vulnerable to adversarial attacks. Despite the rapid development of DNN-based radar object detection models, there have been no studies on their vulnerability to adversarial attacks. Although some spoofing attack methods are proposed to attack the radar sensor by actively transmitting specific signals using some special devices, these attacks require sub-nanosecond-level synchronization between the devices and the radar and are very costly, which limits their practicability in real world. In addition, these attack methods can not effectively attack DNN-based radar object detection. To address the above problems, in this paper, we investigate the possibility of using a few adversarial objects to attack the DNN-based radar object detection models through passive reflection. These objects can be easily fabricated using 3D printing and metal foils at low cost. By placing these adversarial objects at some specific locations on a target vehicle, we can easily fool the victim AV's radar object detection model. The experimental results demonstrate that the attacker can achieve the attack goal by using only two adversarial objects and conceal them as car signs, which have good stealthiness and flexibility. To the best of our knowledge, this is the first study on the passive-reflection-based attacks against the DNN-based radar object detection models using low-cost, readily-available and easily concealable geometric shaped objects. Yi Zhu 0012, Chenglin Miao, Hongfei Xue, Zhengxiong Li, Yunnan Yu, Wenyao Xu, Lu Su 0001, Chunming Qiao |
CCS | 7 |
| 2023 | Multi-Objective Order Dispatch for Urban Crowd Sensing with For-Hire VehiclesabstractFor-hire vehicle-enabled crowd sensing (FVCS) has become a promising paradigm to conduct urban sensing tasks in recent years. FVCS platforms aim to jointly optimize both the order-serving revenue as well as sensing coverage and quality. However, such two objectives are often conflicting and need to be balanced according to the platforms’ preferences on both objectives. To address this problem, we propose a novel cooperative multi-objective multi-agent reinforcement learning framework, referred to as MOVDN, to serve as the first preference-configurable order dispatch mechanism for FVCS platforms. Specifically, MOVDN adopts a decomposed network structure, which enables agents to make distributed order selection decisions, and meanwhile aligns each agent’s local decision with the global objectives of the FVCS platform. Then, we propose a novel algorithm to train a single universal MOVDN that is optimized over the space of all preferences. This allows our trained model to produce the optimal policy for any preference. Furthermore, we provide the theoretical convergence guarantee and sample efficiency analysis of our algorithm. Extensive experiments on three real-world ride-hailing order datasets demonstrate that MOVDN outperforms strong baselines and can support the platform in decision-making effectively. Haiming Jin, Guiyun Fan, Yifei Wei, Lu Su 0001 |
INFOCOM | 6 |
| 2023 | Towards Generalized mmWave-based Human Pose Estimation through Signal AugmentationabstractThe unprecedented advance of wireless human sensing is enabled by the proliferation of the deep learning techniques, which, however, rely heavily on the completeness and representativeness of the data patterns contained in the training set. Thus, deep learning based wireless human perception models usually fail when the human subject is conducting activities that are unseen during the model training. To address this problem, we propose a novel wireless signal augmentation framework, named mmGPE, for Generalized mmWave-based Pose Estimation. In mmGPE, we adopt a physical simulator to generate mmWave FMCW signals. However, due to the imperfect simulation of the physical world, there is a big gap between the signals generated by the physical simulator and the real-world signals collected by the mmWave radar. To tackle this challenge, we propose to integrate the physical signal simulation with deep learning techniques. Specifically, we develop a deep learning-based signal refiner in mmGPE that is capable of bridging the gap and generating realistic signal data. Through extensive evaluations on a COTS mmWave testbed, our mmGPE system demonstrates high accuracy in generating human meshes for unseen activities. Hongfei Xue, Qiming Cao, Chenglin Miao, Yan Ju, Haochen Hu, Aidong Zhang 0001, Lu Su 0001 |
MobiCom | 7 |
| 2023 | Federated Transfer-Ordered-Personalized Learning for Driver Monitoring ApplicationabstractFederated learning (FL) shines through in the Internet of Things (IoT) with its ability to realize collaborative learning and improve learning efficiency by sharing client model parameters trained on local data. Although FL has been successfully applied to various domains, including driver monitoring applications (DMAs) on the Internet of Vehicles (IoV), its usages still face some open issues, such as data and system heterogeneity, large-scale parallelism communication resources, malicious attacks, and data poisoning. This article proposes a federated transfer–ordered–personalized learning (FedTOP) framework to address the above problems and test on two real-world data sets with and without system heterogeneity. The performance of the three extensions, transfer, ordered, and personalized, is compared by an ablation study and achieves 92.32% and 95.96% accuracy on the test clients of two data sets, respectively. Compared to the baseline, there is a 462% improvement in accuracy and a 37.46% reduction in communication resource consumption. The results demonstrate that the proposed FedTOP can be used as a highly accurate, streamlined, privacy-preserving, cybersecurity-oriented, and personalized framework for DMA. Liangqi Yuan, Lu Su 0001, Ziran Wang |
IEEE Internet Things J. | 2 |
| 2023 | PhaseAnti: An Anti-Interference WiFi-Based Activity Recognition System Using Interference-Independent Phase ComponentabstractDriven by a wide range of essential applications, significant achievements have recently been made to explore WiFi-based Human Activity Recognition (HAR) techniques that utilize the information collected by commercial off-the-shelf (COTS) WiFi infrastructures to infer human activities without the need for the subject to carry any devices. Although existing WiFi-based HAR systems achieve satisfactory performance in some instances, they are faced with a severe challenge that the impacts of ubiquitous Co-channel Interference (CCI) on WiFi signals are inevitable. This downgrades the performance of these HAR systems significantly. To address this challenge, we propose PhaseAnti in this paper, a novel WiFi-based HAR system to exploit the CCI-independent phase component, Nonlinear Phase Error Variation (NLPEV), of WiFi Channel State Information (CSI) to cope with the negative effects of CCI. The stability of NLPEV data and the sensibility of this component to motions are rigorously analyzed. Furthermore, validated by extensive properly designed experiments, this phase component across subcarriers is invariant under various CCI scenarios while sufficiently distinct for different motions. Therefore, the NLPEV data can be used and processed effectively to perform HAR in CCI scenarios. Extensive experiments with various daily activities in different indoor rooms demonstrate the superior effectiveness and generalizability of the proposed PhaseAnti system under various CCI scenarios. Specifically, PhaseAnti achieves a$ 96.5\%$recognition accuracy rate (RAR) on average in different CCI scenarios, which can improve up to a$ 16.7\%$RAR compared with the amplitude component in the presence of CCI. Furthermore, the recognition speed is 10.3 × faster than the state-of-the-art solution. Jinyang Huang, Bin Liu 0016, Chenglin Miao, Yan Lu 0001, Qijia Zheng, Yu Wu 0020, Jiancun Liu, Lu Su 0001, Chang Wen Chen |
IEEE Trans. Mob. Comput. | 8 |
| 2023 | VocalPrint: A mmWave-Based Unmediated Vocal Sensing System for Secure AuthenticationabstractWith the continuing growth of voice-controlled devices, voice metrics have been widely used for user identification. However, voice biometrics is vulnerable to replay attacks and ambient noise. We identify that the fundamental vulnerability in voice biometrics is rooted in its indirect sensing modality (e.g., microphone). In this paper, we presentVocalPrint, a resilient mmWave interrogation system which directly captures and analyzes the vocal vibrations for user authentication. Specifically,VocalPrintexploits the unique disturbance of the skin-reflect radio frequency (RF) signals around the near-throat region of the user, caused by the vocal vibrations. The complex ambient noise is isolated from the RF signal using a novel resilience-aware clutter suppression approach for preserving fine-grained vocal biometric properties. Afterward, we extract the vocal tract and vocal source features and input them into an ensemble classifier for authentication.VocalPrintis practical as it allows the effortless transition to a smartphone while having sufficient usability due to its non-contact nature. Our experimental results from 41 participants with different interrogation distances, orientations, and body motions show thatVocalPrintachieves over 96 percent authentication accuracy even under unfavorable conditions. We demonstrate the resilience of our system against complex noise interference and spoof attacks of various threat levels. Huining Li, Chenhan Xu, Aditya Singh Rathore, Zhengxiong Li, Hanbin Zhang, Chen Song 0001, Kun Wang 0005, Lu Su 0001, Feng Lin 0004, Kui Ren 0001, Wenyao Xu |
IEEE Trans. Mob. Comput. | 8 |
| 2023 | Optimizing Cross-Line Dispatching for Minimum Electric Bus FleetabstractRecent years have witnessed the increasing popularity of electric buses (e-buses) around the globe due to their environment friendly nature. However, various factors, such as the prohibitive purchasing costs and the scarcity of large-scale charging facilities, hinder the wider adoption of e-buses. Thus, to effectively cut the cost of building and maintaining urban e-bus systems, we optimize the dispatching strategy for urban e-bus systems to satisfy public transportation demands with the minimum e-bus fleet. Specifically, we propose to systematically exploit at city-scale cross-line dispatching, a smart dispatching strategy allowing one bus to serve multiple bus lines when necessary. Technically, we construct a novel and generalizable graph-theoretic model for urban e-bus systems integrating e-buses non-negligible charging time, the spatio-temporal constraints of bus trips, and various other real-world factors. We prove that it is NP-hard, and has no$(2-\epsilon)$-approximation algorithm. Next, we propose a polynomial-time algorithm solving the problem with a guaranteed approximation ratio. Furthermore, we conduct extensive experiments on a large-scale real-world bus dataset from Shenzhen, China, which validate the effectiveness of our algorithms. As shown by our experimental results, to serve 300 bus lines, our dispatching strategy needs 38.2% less e-buses than the one currently used in practice. Chonghuan Wang, Yiwen Song, Guiyun Fan, Haiming Jin, Lu Su 0001, Fan Zhang 0019, Xinbing Wang |
IEEE Trans. Mob. Comput. | 5 |
| 2022 | Joint Order Dispatch and Charging for Electric Self-Driving Taxi SystemsabstractNowadays, the rapid development of self-driving technology and its fusion with the current vehicle electrification process has given rise to electric self-driving taxis (es-taxis). Foreseeably, es-taxis will become a major force that serves the massive urban mobility demands not far into the future. Though promising, it is still a fundamental unsolved problem of effectively deciding when and where a city-scale fleet of es-taxis should be charged, so that enough es-taxis will be available whenever and wherever ride requests are submitted. Furthermore, charging decisions are far from isolated, but tightly coupled with the order dispatch process that matches orders with es-taxis. Therefore, in this paper, we investigate the problem of joint order dispatch and charging in es-taxi systems, with the objective of maximizing the ride-hailing platform’s long-term cumulative profit. Technically, such problem is challenging in a myriad of aspects, such as long-term profit maximization, partial statistical information on future orders, etc. We address the various arising challenges by meticulously integrating a series of methods, including distributionally robust optimization, primal-dual transformation, and second order conic programming to yield far-sighted decisions. Finally, we validate the effectiveness of our proposed methods though extensive experiments based on two large-scale real-world online ride-hailing order datasets. Guiyun Fan, Haiming Jin, Yiran Zhao 0001, Yiwen Song, Xiaoying Gan, Jiaxin Ding 0001, Lu Su 0001, Xinbing Wang |
INFOCOM | 7 |
| 2022 | Optimizing Long-Term Efficiency and Fairness in Ride-Hailing via Joint Order Dispatching and Driver RepositioningabstractThe ride-hailing service offered by mobility-on-demand platforms, such as Uber and Didi Chuxing, has greatly facilitated people's traveling and commuting, and become increasingly popular in recent years. Efficiency (e.g., gross merchandise volume) has always been an important metric for such platforms. However, only focusing on the efficiency inevitably ignores the fairness of driver incomes, which could impair the sustainability of the overall ride-hailing system in the long run. To optimize the aforementioned two essential metrics, order dispatching and driver repositioning play an important role, as they impact not only the immediate, but also the future order-serving outcomes of drivers. Thus, in this paper, we aim to exploit joint order dispatching and driver repositioning to optimize both the long-term efficiency and fairness for ride-hailing platforms. To address this problem, we propose a novel multi-agent reinforcement learning framework, referred to as JDRL, to help drivers make distributed order selection and repositioning decisions. Specifically, to cope with the variable action space, JDRL segments the action space into a fixed number of action groups, and fixes the policy output dimension for order selection as the number of action groups. In terms of the fairness criterion, JDRL adopts the max-min fairness, and augments the vanilla policy gradient to an iterative training algorithm that alternates between a minimization step and a policy improvement step to maximize both the worst and the overall performance of agents. In addition, we provide the theoretical convergence guarantee of our JDRL training algorithm even under non-convex policy networks and stochastic gradient updating. Extensive experiments are conducted with three public real-world ride-hailing order datasets, including over 2 million orders in Haikou, China, over 5 million orders in Chengdu, China, and over 6 million orders in New York City, USA. Experimental results show that JDRL demonstrates a consistent advantage compared to state-of-the-art baselines in terms of both efficiency and fairness. To the best of our knowledge, this is the first work that exploits joint order dispatching and driver repositioning to optimize both the long-term efficiency and fairness in a ride-hailing system. Haiming Jin, Zhaoxing Yang, Lu Su 0001, Xinbing Wang |
KDD | 4 |
| 2022 | M4esh: mmWave-Based 3D Human Mesh Construction for Multiple SubjectsabstractThe recent proliferation of various wireless sensing systems and applications demonstrates the advantages of radio frequency (RF) signals over traditional camera-based solutions that are faced with various challenges, such as occlusions and poor lighting conditions. Towards the ultimate goal of imaging human body using RF signals, researchers have been exploring the possibility of constructing the human mesh, a structure capturing not only the pose but also the shape of the human body, from RF signals. In this paper, we introduce M4esh, a novel system that utilizes commercial millimeter wave (mmWave) radar for multi-subject 3D human mesh construction. Our M4esh system can detect and track the subjects on a 2D energy map by predicting the subject bounding boxes on the map, and tackle the subjects' mutual occlusion through utilizing the location, velocity and size information of the subjects' bounding boxes from the previous frames as a clue to estimate the bounding box in the current frame. Through extensive experiments on a real-world COTS millimeter-wave testbed, we show that our proposed M4esh system can accurately localize the subjects and generate their human meshes, which demonstrate the superior effectiveness of the proposed M4esh system. Hongfei Xue, Qiming Cao, Yan Ju, Haochen Hu, Haoyu Wang 0004, Aidong Zhang 0001, Lu Su 0001 |
SenSys | 7 |
| 2022 | Towards Backdoor Attacks against LiDAR Object Detection in Autonomous DrivingabstractDue to the great advantage of LiDAR sensors in perceiving complex driving environments, LiDAR-based 3D object detection has recently drawn significant attention in autonomous driving. Although many advanced LiDAR object detection models have been developed, their designs are mainly based on deep learning approaches, which are usually data-hungry and expensive to train. Thus, it is common for some LiDAR perception system developers or self-driving car companies to collect training data from different sources (e.g., self-driving car users) or outsource the training work to a third party. However, these practices provide opportunities for backdoor attacks, where the attacker aims to inject a hidden trigger pattern into the victim detection model by poisoning its training set and let the model fail to detect objects when the trigger presents in the inference phase. Although backdoor attacks have posed serious security concerns, the vulnerability of LiDAR object detection to such attacks has not yet been studied. To fill the research gap, in this paper, we present the first study on backdoor attacks against LiDAR object detection in autonomous driving. Specifically, we propose a novel backdoor attack strategy based on which the attacker can achieve the attack goal by poisoning a small number of point cloud samples. In addition, the proposed attack strategy is physically realizable, and it allows the attacker to easily perform the attack using some common objects as the triggers. To make the poisoned samples difficult to be detected, we also design a stealthy attack strategy by creating some fake vehicle point clusters to hide the injected points in the point cloud. The desirable performance of our attacks is demonstrated through both simulation and real-world case study. Yan Zhang 0133, Yi Zhu 0012, Zihao Liu 0001, Chenglin Miao, Foad Hajiaghajani, Lu Su 0001, Chunming Qiao |
SenSys | 6 |
| 2022 | Towards the Inference of Travel Purpose with Heterogeneous Urban DataabstractIn people’s daily lives, travel takes up an important part, and many trips are generated everyday, such as going to school or shopping. With the widely adoption of GPS-integrated devices, a large amount of trips can be recorded with GPS trajectories. These trajectories are represented by sequences of geo-coordinates and can help us answer simple questions such as “where did you go”. However, there is another important question awaiting to be answered, that is “what did/will you do”, i.e., the trip purpose inference. In practice, people’s trip purposes are very important in understanding travel behaviors and estimating travel demands. Obviously, it is very challenging to infer trip purposes solely based on the trajectories, because the GPS devices are not accurate enough to pinpoint the venues visited. In this paper, we infer individual’s trip purposes by combining the knowledge from heterogeneous data sources including trajectories, POIs and social media data. The proposed Dynamic Bayesian Network model (DBN) captures three important factors: the sequential properties of trip activities, the functionality and POI popularity of trip end areas. In addition, we propose an efficient method with local candidate pools to identify POIs from geo-tagged social media messages, and learn the POI popularities from nearby social media data. Moreover, trip data is usually imbalanced across different activities. This data imbalance problem can cause serious challenges because theDBNmodel could be biased by those “popular” class labels. Considering this challenge, we propose an ensemble DBN method with sampling technique (eDBN) which results in more accurate inference. Furthermore, real-world trip data are continuously collected on a daily basis. The batch model would result in unnecessary computation because historical data need to be revisited. We handle this problem by proposing an incremental DBN method (iDBN) which is both effective and efficient. Extensive experiments are conducted on real-world data sets with trajectories of 8,361 residents and the 6.9 million geo-tagged tweets in the Bay area. Experimental results demonstrate the advantages of the proposed method on correctly inferring the trip purposes. Chuishi Meng, Qing He 0011, Lu Su 0001, Jing Gao 0004 |
IEEE Trans. Big Data | 4 |
| 2022 | Joint Charging and Relocation Recommendation for E-Taxi Drivers via Multi-Agent Mean Field Hierarchical Reinforcement LearningabstractNowadays, most of the taxi drivers have become users of the relocation recommendation service offered by online ride-hailing platforms (e.g., Uber and Didi Chuxing), which could oftentimes lead drivers to places with profitable orders. At the same time, electric taxis (e-taxis) are increasingly adopted and gradually replacing gasoline taxis in today’s public transportation systems due to their environmental-friendly nature. Though effective for traditional gasoline taxis, existing relocation recommendation schemes are rather suboptimal for e-taxi drivers’ user experience. On one hand, the existing schemes take no account of taxis’ refueling decisions, as the refueling durations of gasoline taxis are usually short enough to be ignored. However, the charging duration of the e-taxis spent at charging stations can be as long as hours. Obviously, an e-taxi’s battery could be easily depleted by the continuous relocations suggested by existing schemes, and thus will have to be charged for a long time afterwards, making the e-taxi driver miss numerous order-serving opportunities. On the other hand, charging posts are typically sparsely and unevenly distributed across a city. With no consideration of charging opportunities, existing schemes could probably send an e-taxi to an area with no charging post around, even though its battery is running low. To optimize e-taxi drivers’ user experience, in this paper, we design a jointcharging and relocation recommendation system for e-taxi drivers (CARE). We take the perspective of e-taxi drivers and formulate their decision making as a multi-agent reinforcement learning problem where each e-taxi driver aims to maximize his own cumulative rewards. More specifically, we propose a novelmulti-agent mean field hierarchical reinforcement learning (MFHRL)framework. The hierarchical architecture of MFHRL helps the proposed CARE provide far-sighted charging and relocation recommendations for e-taxi drivers. Besides, we integrate each hierarchical level of MFHRL separately with the mean field approximation to incorporate e-taxis’ mutual influences in decision making. We set up a simulator with one of the largest real-world e-taxi datasets in Shenzhen, China, which contains the GPS trajectory data and transaction data of 3848 e-taxis from June 1st to June 30th, 2017, coupled with 165 charging stations including 317 fast charging posts and 1421 slow charging posts. We adopt this simulator to generate 6 dynamic urban environments, which reflect the different real-world scenarios faced by e-taxi drivers. In all of these environments, we conduct extensive experiments to validate that the proposed MFHRL framework greatly outperforms all baselines by significantly increasing the rewards obtained by e-taxi drivers. Besides, we also show that the charging policy learned by MFHRL can effectively reduce the range anxiety of e-taxi drivers, which significantly boosts e-taxi drivers’ quality of experience. Enshu Wang, Zhaoxing Yang, Haiming Jin, Chenglin Miao, Lu Su 0001, Fan Zhang 0019, Chunming Qiao, Xinbing Wang |
IEEE Trans. Mob. Comput. | 6 |
| 2021 | Can We Use Arbitrary Objects to Attack LiDAR Perception in Autonomous Driving?abstractAs an effective way to acquire accurate information about the driving environment, LiDAR perception has been widely adopted in autonomous driving. The state-of-the-art LiDAR perception systems mainly rely on deep neural networks (DNNs) to achieve good performance. However, DNNs have been demonstrated vulnerable to adversarial attacks. Although there are a few works that study adversarial attacks against LiDAR perception systems, these attacks have some limitations in feasibility, flexibility, and stealthiness when being performed in real-world scenarios. In this paper, we investigate an easier way to perform effective adversarial attacks with high flexibility and good stealthiness against LiDAR perception in autonomous driving. Specifically, we propose a novel attack framework based on which the attacker can identify a few adversarial locations in the physical space. By placing arbitrary objects with reflective surface around these locations, the attacker can easily fool the LiDAR perception systems. Extensive experiments are conducted to evaluate the performance of the proposed attack, and the results show that our proposed attack can achieve more than 90% success rate. In addition, our real-world study demonstrates that the proposed attack can be easily performed using only two commercial drones. To the best of our knowledge, this paper presents the first study on the effect of adversarial locations on LiDAR perception models' behaviors, the first investigation on how to attack LiDAR perception systems using arbitrary objects with reflective surface, and the first attack against LiDAR perception systems using commercial drones in physical world. Potential defense strategies are also discussed to mitigate the proposed attacks. Yi Zhu 0012, Chenglin Miao, Tianhang Zheng, Foad Hajiaghajani, Lu Su 0001, Chunming Qiao |
CCS | 5 |
| 2021 | Profanity-Avoiding Training Framework for Seq2seq Models with Certified RobustnessabstractSeq2seq models have demonstrated their incredible effectiveness in a large variety of applications.However, recent research has shown that inappropriate language in training samples and well-designed testing cases can induce seq2seq models to output profanity.These outputs may potentially hurt the usability of seq2seq models and make the end-users feel offended.To address this problem, we propose a training framework with certified robustness to eliminate the causes that trigger the generation of profanity.The proposed training framework leverages merely a short list of profanity examples to prevent seq2seq models from generating a broader spectrum of profanity.The framework is composed of a patterneliminating training component to suppress the impact of language patterns with profanity in the training set, and a trigger-resisting training component to provide certified robustness for seq2seq models against intentionally injected profanity-triggering expressions in test samples.In the experiments, we consider two representative NLP tasks that seq2seq can be applied to, i.e., style transfer and dialogue generation.Extensive experimental results show that the proposed training framework can successfully prevent the NLP models from generating profanity. Hengtong Zhang, Tianhang Zheng, Yaliang Li, Jing Gao 0004, Lu Su 0001, Bo Li 0126 |
EMNLP (1) | 5 |
| 2021 | Heterogeneous Spatio-Temporal Graph Convolution Network for Traffic Forecasting with Missing ValuesabstractAccurate traffic prediction is indispensable for intelligent traffic management. The availability of large-scale road sensing data collected by connected wireless sensors and mobile devices have provided unrealized potential for traffic prediction. However, sensory data is often incomplete due to various factors in the process of data acquisition and transmission. The missingness of traffic data brings a key challenge to the traffic prediction task since the state-of-the-art ML-based traffic prediction models (e.g., Graph Convolutional Networks (GCN)) often rely on spatial and temporal completion of the data. Moreover, existing GCN-based methods usually build a static graph based on geographical distances and are limited in their ability to capture the time-evolving relationships amongst road segments. In this paper, we develop a heterogeneous spatio-temporal prediction framework for traffic prediction using incomplete historical data. In the framework, we build multiple graphs to explicitly model the dynamic correlations among road segments from both geographical and historical aspects, and employ recurrent neural networks to capture temporal correlations for each road segment. We impute missing values in a recurrent process, which is seamlessly embedded in the prediction framework so they can be jointly trained. The proposed framework is evaluated on a public dataset of static sensors and a private dataset collected by our roving sensor system. Experimental results show the effectiveness of the proposed framework compared to state-of-the-art methods, and indicate the potential to be deployed into real-world traffic prediction systems. Weida Zhong, Qiuling Suo, Xiaowei Jia, Aidong Zhang 0001, Lu Su 0001 |
ICDCS | 5 |
| 2021 | Data Poisoning Attacks Against Outcome Interpretations of Predictive ModelsabstractThe past decades have witnessed significant progress towards improving the accuracy of predictions powered by complex machine learning models. Despite much success, the lack of model interpretability prevents the usage of these techniques in life-critical systems such as medical diagnosis and self-driving systems. Recently, the interpretability issue has received much attention, and one critical task is to explain why a predictive model makes a specific decision. We refer to this task as outcome interpretation. Many outcome interpretation methods have been developed to produce human-understandable interpretations by utilizing intermediate results of the machine learning models, such as gradients and model parameters. Hengtong Zhang, Jing Gao 0004, Lu Su 0001 |
KDD | 3 |
| 2021 | Data Poisoning Attack against Recommender System Using Incomplete and Perturbed DataabstractRecent studies reveal that recommender systems are vulnerable to data poisoning attack due to their openness nature. In data poisoning attack, the attacker typically recruits a group of controlled users to inject well-crafted user-item interaction data into the recommendation model's training set to modify the model parameters as desired. Thus, existing attack approaches usually require full access to the training data to infer items' characteristics and craft the fake interactions for controlled users. However, such attack approaches may not be feasible in practice due to the attacker's limited data collection capability and the restricted access to the training data, which sometimes are even perturbed by the privacy preserving mechanism of the service providers. Such design-reality gap may cause failure of attacks. In this paper, we fill the gap by proposing two novel adversarial attack approaches to handle the incompleteness and perturbations in user-item interaction data. First, we propose a bi-level optimization framework that incorporates a probabilistic generative model to find the users and items whose interaction data is sufficient and has not been significantly perturbed, and leverage these users and items' data to craft fake user-item interactions. Moreover, we reverse the learning process of recommendation models and develop a simple yet effective approach that can incorporate context-specific heuristic rules to handle data incompleteness and perturbations. Extensive experiments on two datasets against three representative recommendation models show that the proposed approaches can achieve better attack performance than existing approaches. Hengtong Zhang, Changxin Tian, Yaliang Li, Lu Su 0001, Wayne Xin Zhao, Jing Gao 0004 |
KDD | 4 |
| 2021 | mmMesh: towards 3D real-time dynamic human mesh construction using millimeter-waveabstractIn this paper, we present mmMesh, the first real-time 3D human mesh estimation system using commercial portable millimeter-wave devices. mmMesh is built upon a novel deep learning framework that can dynamically locate the moving subject and capture his/her body shape and pose by analyzing the 3D point cloud generated from the mmWave signals that bounce off the human body. The proposed deep learning framework addresses a series of challenges. First, it encodes a 3D human body model, which enables mmMesh to estimate complex and realistic-looking 3D human meshes from sparse point clouds. Second, it can accurately align the 3D points with their corresponding body segments despite the influence of ambient points as well as the error-prone nature and the multi-path effect of the RF signals. Third, the proposed model can infer missing body parts from the information of the previous frames. Our evaluation results on a commercial mmWave sensing testbed show that our mmMesh system can accurately localize the vertices on the human mesh with an average error of 2.47 cm. The superior experimental results demonstrate the effectiveness of our proposed human mesh construction system. Hongfei Xue, Yan Ju, Chenglin Miao, Yijiang Wang, Aidong Zhang 0001, Lu Su 0001 |
MobiSys | 7 |
| 2021 | Adversarial Attacks against LiDAR Semantic Segmentation in Autonomous DrivingabstractToday, most autonomous vehicles (AVs) rely on LiDAR (Light Detection and Ranging) perception to acquire accurate information about their immediate surroundings. In LiDAR-based perception systems, semantic segmentation plays a critical role as it can divide LiDAR point clouds into meaningful regions according to human perception and provide AVs with semantic understanding of the driving environments. However, an implicit assumption for existing semantic segmentation models is that they are performed in a reliable and secure environment, which may not be true in practice. In this paper, we investigate adversarial attacks against LiDAR semantic segmentation in autonomous driving. Specifically, we propose a novel adversarial attack framework based on which the attacker can easily fool LiDAR semantic segmentation by placing some simple objects (e.g., cardboard and road signs) at some locations in the physical space. We conduct extensive real-world experiments to evaluate the performance of our proposed attack framework. The experimental results show that our attack can achieve more than 90% success rate in real-world driving environments. To the best of our knowledge, this is the first study on physically realizable adversarial attacks against LiDAR point cloud semantic segmentation with real-world evaluations. Yi Zhu 0012, Chenglin Miao, Foad Hajiaghajani, Mengdi Huai, Lu Su 0001, Chunming Qiao |
SenSys | 5 |
| 2021 | Truth Discovery With Multi-Modal Data in Social SensingabstractThis article proposes unsupervised truth-finding algorithms that combine consideration of multi-modal content features with analysis of propagation patterns to evaluate the veracity of observations in social sensing applications. A key social sensing challenge is to develop effective algorithms for estimating both the reliability of sources and the veracity of their observations without prior knowledge. In contrast to prior solutions that use labeled examples to learn content features that are correlated with veracity, our approach is entirely unsupervised. Hence, given no prior training data, we jointly learn the importance of different content features together with the veracity of observations using propagation patterns as an indicator of perceived content reliability. A novel penalized expectation maximization (PEM) algorithm is proposed to improve the quality of estimation results for observations bolstered by multiple features. In addition, we develop a constrained expectation maximum likelihood with multiple features (CEM-MultiF) that introduces a novel constraint to boost the probability of correctness of some claims. Finally, we evaluate the performance of the proposed algorithms, called EM-Multi, CEM-Multi and PEM-MultiF, respectively, on real-world data sets collected from Twitter. The evaluation results demonstrate that the proposed algorithms outperform the existing fact-finding approaches, and offer tunable knobs for controlling robustness/performance trade-offs in the presence of malicious sources. Huajie Shao, Dachun Sun, Shuochao Yao, Lu Su 0001, Zhibo Wang 0001, Dongxin Liu, Shengzhong Liu, Lance M. Kaplan, Tarek F. Abdelzaher |
IEEE Trans. Computers | 4 |
| 2021 | Who Is in Control? Practical Physical Layer Attack and Defense for mmWave-Based Sensing in Autonomous VehiclesabstractWith the wide bandwidths in millimeter wave (mmWave) frequency band that results in unprecedented accuracy, mmWave sensing has become vital for many applications, especially in autonomous vehicles (AVs). In addition, mmWave sensing has superior reliability compared to other sensing counterparts such as camera and LiDAR, which is essential for safety-critical driving. Therefore, it is critical to understand the security vulnerabilities and improve the security and reliability of mmWave sensing in AVs. To this end, we perform the end-to-end security analysis of a mmWave-based sensing system in AVs, by designing and implementing practical physical layer attack and defense strategies in a state-of-the-art mmWave testbed and an AV testbed in real-world settings. Various strategies are developed to take control of the victim AV by spoofing its mmWave sensing module, including adding fake obstacles at arbitrary locations and faking the locations of existing obstacles. Five real-world attack scenarios are constructed to spoof the victim AV and force it to make dangerous driving decisions leading to a fatal crash. Field experiments are conducted to study the impact of the various attack scenarios using a Lincoln MKZ-based AV testbed, which validate that the attacker can indeed assume control of the victim AV to compromise its security and safety. To defend the attacks, we design and implement a challenge-response authentication scheme and a RF fingerprinting scheme to reliably detect aforementioned spoofing attacks. Sarankumar Balakrishnan, Lu Su 0001, Arupjyoti Bhuyan, Pu Wang 0001, Chunming Qiao |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Driver Behavior-aware Parking Availability Crowdsensing System Using Truth DiscoveryabstractSpot-level parking availability information (the availability of each spot in a parking lot) is in great demand, as it can help reduce time and energy waste while searching for a parking spot. In this article, we propose a crowdsensing system called SpotE that can provide spot-level availability in a parking lot using drivers’ smartphone sensors. SpotE only requires the sensor data from drivers’ smartphones, which avoids the high cost of installing additional sensors and enables large-scale outdoor deployment. We propose a new model that can use the parking search trajectory and final destination (e.g., an exit of the parking lot) of a single driver in a parking lot to generate the probability profile that contains the probability of each spot being occupied in a parking lot. To deal with conflicting estimation results generated from different drivers, due to the variance in different drivers’ parking behaviors, a novel aggregation approach SpotE-TD is proposed. The proposed aggregation method is based on truth discovery techniques and can handle the variety in Quality of Information of different vehicles. We evaluate our proposed method through a real-life deployment study. Results show that SpotE-TD can efficiently provide spot-level parking availability information with a 20% higher accuracy than the state-of-the-art. Yi Zhu 0012, Shaohan Hu, Weida Zhong, Lu Su 0001, Chunming Qiao |
ACM Trans. Sens. Networks | 5 |
| 2020 | Estimation of Road Transverse Slope Using Crowd-Sourced Data from SmartphonesabstractIntegration of information on road transverse geometric features such as cross slope and superelevation in digital maps can widen the scope of its applications, which is primarily navigation, by enabling driving safety and efficiency applications such as Advanced Driver Assistance Systems (ADAS). The huge scale and dynamic nature of road networks make sensing such road geometric features a challenging task. Traditional methods oftentimes suffer from high cost, limited scalability and update frequency, as well as poor sensing accuracy. To overcome these problems, we propose a cost-effective and scalable road transverse slope estimation framework using sensor data from smartphones. Based on error characteristics of smartphone sensors, we intelligently combine data from accelerometer, gyroscope and GPS to estimate road transverse slope profile of a road segment. To improve accuracy and robustness of the system, the estimations of road transverse slope from multiple sources/vehicles are crowd-sourced to compensate for the effects of varying quality of sensor data from different sources. Extensive experimental evaluation on a test route of 9km demonstrates the superior performance of our proposed method, achieving 350% improvement on road transverse slope estimation accuracy over existing methods, with 90% of errors below 0.5°. Abhinav Khare, Haiming Jin, Adel W. Sadek, Lu Su 0001, Chunming Qiao |
SIGSPATIAL/GIS | 5 |
| 2020 | Towards Differentially Private Truth Discovery for Crowd Sensing SystemsabstractNowadays, crowd sensing becomes increasingly more popular due to the ubiquitous usage of mobile devices. However, the quality of such human-generated sensory data varies significantly among different users. To better utilize sensory data, the problem of truth discovery, whose goal is to estimate user quality and infer reliable aggregated results through quality-aware data aggregation, has emerged as a hot topic. Although the existing truth discovery approaches can provide reliable aggregated results, they fail to protect the private information of individual users. Moreover, crowd sensing systems typically involve a large number of participants, making encryption or secure multi-party computation based solutions difficult to deploy. To address these challenges, in this paper, we propose an efficient privacy-preserving truth discovery mechanism with theoretical guarantees of both utility and privacy. The key idea of the proposed mechanism is to perturb data from each user independently and then conduct weighted aggregation among users’ perturbed data. The proposed approach is able to assign user weights based on information quality, and thus the aggregated results will not deviate much from the true results even when large noise is added. We adapt local differential privacy definition to this privacy-preserving task and demonstrate the proposed mechanism can satisfy local differential privacy while preserving high aggregation accuracy. We formally quantify utility and privacy trade-off and further verify the claim by experiments on both synthetic data and a real-world crowd sensing system. Yaliang Li, Houping Xiao, Zhan Qin, Chenglin Miao, Lu Su 0001, Jing Gao 0004, Kui Ren 0001, Bolin Ding |
ICDCS | 5 |
| 2020 | Road Grade Estimation Using Crowd-Sourced Smartphone DataabstractEstimates of road grade/slope can add another dimension of information to existing 2D digital road maps. Integration of road grade information will widen the scope of digital map’s applications, which is primarily used for navigation, by enabling driving safety and efficiency applications such as Advanced Driver Assistance Systems (ADAS), eco-driving, etc. The huge scale and dynamic nature of road networks make sensing road grade a challenging task. Traditional methods oftentimes suffer from limited scalability and update frequency, as well as poor sensing accuracy. To overcome these problems, we propose a cost-effective and scalable road grade estimation framework using sensor data from smartphones. Based on our understanding of the error characteristics of smartphone sensors, we intelligently combine data from accelerometer, gyroscope and vehicle speed data from OBD-II/smartphone’s GPS to estimate road grade. To improve accuracy and robustness of the system, the estimations of road grade from multiple sources/vehicles are crowd-sourced to compensate for the effects of varying quality of sensor data from different sources. Extensive experimental evaluation on a test route of 9km demonstrates the superior performance of our proposed method, achieving 5× improvement on road grade estimation accuracy over baselines, with 90% of errors below 0.3°. Shaohan Hu, Weida Zhong, Adel W. Sadek, Lu Su 0001, Chunming Qiao |
IPSN | 5 |
| 2020 | Towards 3D human pose construction using wifiabstractThis paper presents WiPose, the first 3D human pose construction framework using commercial WiFi devices. From the pervasive WiFi signals, WiPose can reconstruct 3D skeletons composed of the joints on both limbs and torso of the human body. By overcoming the technical challenges faced by traditional camera-based human perception solutions, such as lighting and occlusion, the proposed WiFi human sensing technique demonstrates the potential to enable a new generation of applications such as health care, assisted living, gaming, and virtual reality. WiPose is based on a novel deep learning model that addresses a series of technical challenges. First, WiPose can encode the prior knowledge of human skeleton into the posture construction process to ensure the estimated joints satisfy the skeletal structure of the human body. Second, to achieve cross environment generalization, WiPose takes as input a 3D velocity profile which can capture the movements of the whole 3D space, and thus separate posture-specific features from the static objects in the ambient environment. Finally, WiPose employs a recurrent neural network (RNN) and a smooth loss to enforce smooth movements of the generated skeletons. Our evaluation results on a real-world WiFi sensing testbed with distributed antennas show that WiPose can localize each joint on the human skeleton with an average error of 2.83cm, achieving a 35% improvement in accuracy over the state-of-the-art posture construction model designed for dedicated radar sensors. Hongfei Xue, Chenglin Miao, Sen Lin 0009, Chong Tian, Srinivasan Murali, Haochen Hu, Lu Su 0001 |
MobiCom | 10 |
| 2020 | VocalPrint: exploring a resilient and secure voice authentication via mmWave biometric interrogationabstractWith the continuing growth of voice-controlled devices, voice metrics have been widely used for user identification. However, voice biometrics is vulnerable to replay attacks and ambient noise. We identify that the fundamental vulnerability in voice biometrics is rooted in its indirect sensing modality (e.g., microphone). In this paper, we present VocalPrint, a resilient mmWave interrogation system which directly captures and analyzes the vocal vibrations for user authentication. Specifically, VocalPrint exploits the unique disturbance of the skin-reflect radio frequency (RF) signals around the near-throat region of the user, caused by the vocal vibrations during communication. The complex ambient noise is isolated from the RF signal using a novel resilience-aware clutter suppression approach for preserving fine-grained vocal biometric properties. Afterward, we extract the text-independent vocal tract and vocal source features and input them to an ensemble classifier for user authentication. VocalPrint is practical as it leverages a low-cost, portable, and energy-efficient hardware allowing effortless transition to a smartphone while having sufficient usability as typical voice authentication systems due to its non-contact nature. Our experimental results from 41 participants with different interrogation distances, orientations, and body motions show that VocalPrint can achieve over 96% authentication accuracy even under unfavorable conditions. We demonstrate the resilience of our system against complex noise interference and spoof attacks of various threat levels. Huining Li, Chenhan Xu, Aditya Singh Rathore, Zhengxiong Li, Hanbin Zhang, Chen Song 0001, Kun Wang 0005, Lu Su 0001, Feng Lin 0004, Kui Ren 0001, Wenyao Xu |
SenSys | 8 |
| 2020 | WaveSpy: Remote and Through-wall Screen Attack via mmWave SensingabstractDigital screens, such as liquid crystal displays (LCDs), are vulnerable to attacks (e.g., "shoulder surfing") that can bypass security protection services (e.g., firewall) to steal confidential information from intended victims. The conventional practice to mitigate these threats is isolation. An isolated zone, without accessibility, proximity, and line-of-sight, seems to bring personal devices to a truly secure place.In this paper, we revisit this historical topic and re-examine the security risk of screen attacks in an isolation scenario mentioned above. Specifically, we identify and validate a new and practical side-channel attack for screen content via liquid crystal nematic state estimation using a low-cost radio-frequency sensor. By leveraging the relationship between the screen content and the states of liquid crystal arrays in displays, we develop WaveSpy, an end-to-end portable through-wall screen attack system. WaveSpy comprises a low-cost, energy-efficient and light-weight millimeter-wave (mmWave) probe which can remotely collect the liquid crystal state response to a set of mmWave stimuli and facilitate screen content inference, even when the victim’s screen is placed in an isolated zone. We intensively evaluate the performance and practicality of WaveSpy in screen attacks, including over 100 different types of content on 30 digital screens of modern electronic devices. WaveSpy achieves an accuracy of 99% in screen content type recognition and a success rate of 87.77% in Top-3 sensitive information retrieval under real-world scenarios, respectively. Furthermore, we discuss several potential defense mechanisms to mitigate screen eavesdropping similar to WaveSpy. Zhengxiong Li, Fenglong Ma, Aditya Singh Rathore, Zhuolin Yang 0001, Baicheng Chen, Lu Su 0001, Wenyao Xu |
SP | 6 |
| 2020 | Learning Distance Metrics from Probabilistic InformationabstractThe goal of metric learning is to learn a good distance metric that can capture the relationships among instances, and its importance has long been recognized in many fields. An implicit assumption in the traditional settings of metric learning is that the associated labels of the instances are deterministic. However, in many real-world applications, the associated labels come naturally with probabilities instead of deterministic values, which makes it difficult for the existing metric-learning methods to work well in these applications. To address this challenge, in this article, we study how to effectively learn the distance metric from datasets that contain probabilistic information, and then propose several novel metric-learning mechanisms for two types of probabilistic labels, i.e., the instance-wise probabilistic label and the group-wise probabilistic label. Compared with the existing metric-learning methods, our proposed mechanisms are capable of learning distance metrics directly from the probabilistic labels with high accuracy. We also theoretically analyze the proposed mechanisms and conduct extensive experiments on real-world datasets to verify the desirable properties of these mechanisms. Mengdi Huai, Chenglin Miao, Yaliang Li, Qiuling Suo, Lu Su 0001, Aidong Zhang 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2019 | Unsupervised Fact-finding with Multi-modal Data in Social Sensing
Huajie Shao, Shuochao Yao, Yiran Zhao 0001, Lu Su 0001, Zhibo Wang 0001, Dongxin Liu, Shengzhong Liu, Lance M. Kaplan, Tarek F. Abdelzaher |
FUSION | 4 |
| 2019 | Deep Metric Learning: The Generalization Analysis and an Adaptive AlgorithmabstractAs an effective way to learn a distance metric between pairs of samples, deep metric learning (DML) has drawn significant attention in recent years. The key idea of DML is to learn a set of hierarchical nonlinear mappings using deep neural networks, and then project the data samples into a new feature space for comparing or matching. Although DML has achieved practical success in many applications, there is no existing work that theoretically analyzes the generalization error bound for DML, which can measure how good a learned DML model is able to perform on unseen data. In this paper, we try to fill up this research gap and derive the generalization error bound for DML. Additionally, based on the derived generalization bound, we propose a novel DML method (called ADroDML), which can adaptively learn the retention rates for the DML models with dropout in a theoretically justified way. Compared with existing DML works that require predefined retention rates, ADroDML can learn the retention rates in an optimal way and achieve better performance. We also conduct experiments on real-world datasets to verify the findings derived from the generalization error bound and demonstrate the effectiveness of the proposed adaptive DML method. Mengdi Huai, Hongfei Xue, Chenglin Miao, Liuyi Yao, Lu Su 0001, Changyou Chen, Aidong Zhang 0001 |
IJCAI | 5 |
| 2019 | Data Poisoning Attack against Knowledge Graph EmbeddingabstractKnowledge graph embedding (KGE) is a technique for learning continuous embeddings for entities and relations in the knowledge graph. Due to its benefit to a variety of downstream tasks such as knowledge graph completion, question answering and recommendation, KGE has gained significant attention recently. Despite its effectiveness in a benign environment, KGE's robustness to adversarial attacks is not well-studied. Existing attack methods on graph data cannot be directly applied to attack the embeddings of knowledge graph due to its heterogeneity. To fill this gap, we propose a collection of data poisoning attack strategies, which can effectively manipulate the plausibility of arbitrary targeted facts in a knowledge graph by adding or deleting facts on the graph. The effectiveness and efficiency of the proposed attack strategies are verified by extensive evaluations on two widely-used benchmarks. Hengtong Zhang, Tianhang Zheng, Jing Gao 0004, Chenglin Miao, Lu Su 0001, Yaliang Li, Kui Ren 0001 |
IJCAI | 5 |
| 2019 | Dynamic Task Pricing in Multi-Requester Mobile Crowd Sensing with Markov Correlated EquilibriumabstractThe recent proliferation of human-carried mobile devices has given rise to mobile crowd sensing (MCS) systems, where a myriad of data requesters outsource their sensing tasks to a crowd of workers via a cloud-based platform. In order to incentivize participation, requesters typically compensate workers with specific amount of payments. Clearly, setting an appropriate task price is critical for a requester to attract enough worker participation without unnecessary expenses. Therefore, we investigate the problem of task pricing in MCS systems with multi-requester price competition, and also dynamically arriving workers. Task pricing in such scenario is challenging, because of each requester's incomplete information about the others, uncertainty of future information, etc. So as to address these challenges, we use Markov game to model requesters' competitive task pricing, and Markov correlated equilibrium (MCE) as the solution concept. We propose that the platform uses the social cost-minimizing MCE to coordinate requesters' prices, which is self-enforcing, and optimizes the system-wide objective of social cost. Technically, we propose a computationally efficient algorithm to compute an approximately optimal MCE. Furthermore, through extensive performance evaluation, we show numerically that our algorithm yields close-to-minimum social cost in very short running time. Haiming Jin, Hongpeng Guo, Lu Su 0001, Klara Nahrstedt, Xinbing Wang |
INFOCOM | 3 |
| 2019 | Automating CSI Measurement with UAVs: from Problem Formulation to Energy-Optimal SolutionabstractIndoor localization has been an active research area given the popularity of Location-Based Services. The CSI fingerprinting based approach is one of the most practical and effective approaches since it can provide adequate accuracy with low overhead for users. The key drawback that limits its wide application is the huge amount of human effort required to build the fingerprint map. This paper is the first to explore addressing this limitation by automating CSI map construction using an Unmanned Aerial Vehicle (UAV). Given the limited battery capacity of commodity UAVs, it is extremely important yet challenging to optimize energy efficiency for the UAV during the CSI measurement task. To address this challenge, we formulate an energy optimization problem based on a novel graph model that includes the cost of possible actions for UAVs. We then transform the formulated problem to the classic Generalized Traveling Salesman Problem (GTSP), which can be solved efficiently. We implement the system on an off-the-shelf programmable drone equipped with a CSI measurement module. We achieve great energy efficiency improvement over the conventional coverage path planning algorithm. Meanwhile, accurate indoor localization can be achieved using the CSI data collected by our UAV system. Sixu Piao, Zhongjie Ba, Lu Su 0001, Dimitrios Koutsonikolas, Shi Li 0001, Kui Ren 0001 |
INFOCOM | 3 |
| 2019 | DeepFusion: A Deep Learning Framework for the Fusion of Heterogeneous Sensory DataabstractIn recent years, significant research efforts have been spent towards building intelligent and user-friendly IoT systems to enable a new generation of applications capable of performing complex sensing and recognition tasks. In many of such applications, there are usually multiple different sensors monitoring the same object. Each of these sensors can be regarded as an information source and provides us a unique "view" of the observed object. Intuitively, if we can combine the complementary information carried by multiple sensors, we will be able to improve the sensing performance. Towards this end, we propose DeepFusion, a unified multi-sensor deep learning framework, to learn informative representations of heterogeneous sensory data. DeepFusion can combine different sensors' information weighted by the quality of their data and incorporate cross-sensor correlations, and thus can benefit a wide spectrum of IoT applications. To evaluate the proposed DeepFusion model, we set up two real-world human activity recognition testbeds using commercialized wearable and wireless sensing devices. Experiment results show that DeepFusion can outperform the state-of-the-art human activity recognition methods. Hongfei Xue, Chenglin Miao, Ye Yuan 0006, Fenglong Ma, Xin Ma 0006, Yijiang Wang, Shuochao Yao, Wenyao Xu, Aidong Zhang 0001, Lu Su 0001 |
MobiHoc | 11 |
| 2019 | STFNets: Learning Sensing Signals from the Time-Frequency Perspective with Short-Time Fourier Neural NetworksabstractRecent advances in deep learning motivate the use of deep neural networks in Internet-of-Things (IoT) applications. These networks are modelled after signal processing in the human brain, thereby leading to significant advantages at perceptual tasks such as vision and speech recognition. IoT applications, however, often measure physical phenomena, where the underlying physics (such as inertia, wireless signal propagation, or the natural frequency of oscillation) are fundamentally a function of signal frequencies, offering better features in the frequency domain. This observation leads to a fundamental question: For IoT applications, can one develop a new brand of neural network structures that synthesize features inspired not only by the biology of human perception but also by the fundamental nature of physics? Hence, in this paper, instead of using conventional building blocks (e.g., convolutional and recurrent layers), we propose a new foundational neural network building block, the Short-Time Fourier Neural Network (STFNet). It integrates a widely-used time-frequency analysis method, the Short-Time Fourier Transform, into data processing to learn features directly in the frequency domain, where the physics of underlying phenomena leave better footprints. STFNets bring additional flexibility to time-frequency analysis by offering novel nonlinear learnable operations that are spectral-compatible. Moreover, STFNets show that transforming signals to a domain that is more connected to the underlying physics greatly simplifies the learning process. We demonstrate the effectiveness of STFNets with extensive experiments on a wide range of sensing inputs, including motion sensors, WiFi, ultrasound, and visible light. STFNets significantly outperform the state-of-the-art deep learning models in all experiments. A STFNet, therefore, demonstrates superior capability as the fundamental building block of deep neural networks for IoT applications for various sensor inputs 1. Shuochao Yao, Ailing Piao, Yiran Zhao 0001, Huajie Shao, Shengzhong Liu, Dongxin Liu, Jinyang Li 0004, Tianshi Wang 0002, Shaohan Hu, Lu Su 0001, Jiawei Han 0001, Tarek F. Abdelzaher |
WWW | 11 |
| 2019 | A hybrid self-attention deep learning framework for multivariate sleep stage classificationabstractBACKGROUND: Sleep is a complex and dynamic biological process characterized by different sleep patterns. Comprehensive sleep monitoring and analysis using multivariate polysomnography (PSG) records has achieved significant efforts to prevent sleep-related disorders. To alleviate the time consumption caused by manual visual inspection of PSG, automatic multivariate sleep stage classification has become an important research topic in medical and bioinformatics. RESULTS: We present a unified hybrid self-attention deep learning framework, namely HybridAtt, to automatically classify sleep stages by capturing channel and temporal correlations from multivariate PSG records. We construct a new multi-view convolutional representation module to learn channel-specific and global view features from the heterogeneous PSG inputs. The hybrid attention mechanism is designed to further fuse the multi-view features by inferring their dependencies without any additional supervision. The learned attentional representation is subsequently fed through a softmax layer to train an end-to-end deep learning model. CONCLUSIONS: We empirically evaluate our proposed HybridAtt model on a benchmark PSG dataset in two feature domains, referred to as the time and frequency domains. Experimental results show that HybridAtt consistently outperforms ten baseline methods in both feature spaces, demonstrating the effectiveness of HybridAtt in the task of sleep stage classification. Ye Yuan 0006, Kebin Jia, Fenglong Ma, Guangxu Xun, Yaqing Wang 0001, Lu Su 0001, Aidong Zhang 0001 |
BMC Bioinform. | 6 |
| 2019 | Towards Confidence Interval Estimation in Truth DiscoveryabstractThe demand for automatic extraction of true information (i.e., truths) from conflicting multi-source data has soared recently. A variety of truth discovery methods have witnessed great successes via jointly estimating source reliability and truths. All existing truth discovery methods focus on providing a point estimator for each object's truth, but in many real-world applications, confidence interval estimation of truths is more desirable, since confidence interval contains richer information. To address this challenge, in this paper, we propose a novel truth discovery method (ETCIBoot) to construct confidence interval estimates as well as identify truths, where the bootstrapping techniques are nicely integrated into the truth discovery procedure. Due to the properties of bootstrapping, the estimators obtained by ETCIBoot are more accurate and robust compared with the state-of-the-art truth discovery approaches. The proposed framework is further adapted to deal with large-scale truth discovery task in distributed paradigm. Theoretically, we prove the asymptotical consistency of the confidence interval obtained by ETCIBoot. Experimentally, we demonstrate that ETCIBoot is not only effective in constructing confidence intervals but also able to obtain better truth estimates. Houping Xiao, Jing Gao 0004, Qi Li 0012, Fenglong Ma, Lu Su 0001, Yunlong Feng, Aidong Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2019 | Thanos: Incentive Mechanism with Quality Awareness for Mobile Crowd SensingabstractRecent years have witnessed the emergence of mobile crowd sensing (MCS) systems, which leverage the public crowd equipped with various mobile devices for large scale sensing tasks. In this paper, we study a critical problem in MCS systems, namely, incentivizing worker participation. Different from existing work, we propose an incentive framework for MCS systems, named Thanos, that incorporates a crucial metric, called workers' quality of information (QoI). Due to various factors (e.g., sensor quality and environment noise), the quality of the sensory data contributed by individual workers varies significantly. Obtaining high quality data with little expense is always the ideal of MCS platforms. Technically, our design of Thanos is based on reverse combinatorial auctions. We investigate both the single- and multi-minded combinatorial auction models. For the former, we design a truthful, individual rational, and computationally efficient mechanism that ensures a close-to-optimal social welfare. For the latter, we design an iterative descending mechanism that satisfies individual rationality and computational efficiency, and approximately maximizes the social welfare with a guaranteed approximation ratio. Through extensive simulations, we validate our theoretical analysis on the various desirable properties guaranteed by Thanos. Haiming Jin, Lu Su 0001, Hongpeng Guo, Klara Nahrstedt, Jinhui Xu 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2019 | Dolphin: Real-Time Hidden Acoustic Signal Capture with SmartphonesabstractDual-channel screen-camera communication has been proposed to enable simultaneous screen viewing and hidden screen-camera communication. However, it strictly requires a well-controlled camera-screen alignment and an obstacle-free access. In this paper, we propose Dolphin, a novel real-time acoustics-based dual-channel communication system. Leveraging masking effects of human auditory system and readily available audio signals, Dolphin enables real-time unobtrusive speaker-microphone data communication without affecting the primary audio-hearing experience of human users. Compared with screen-camera communication, Dolphin supports non-line-of-sight transmissions and more flexible speaker-microphone alignments. Dolphin can also automatically adapt the data rate to various channel conditions. We further develop a secure data broadcasting scheme on Dolphin, where only designated privileged users can recover the embedded information in the acoustic signals. Our Dolphin prototype, built using COTS (Commercial Off-The-Shelf) smartphones, realizes (potentially secure) real-time hidden information communication, supports up to 8-meter signal capture distance and +900 listening angle, and achieves an average goodput of 240 bps at 2 m. Man Zhou 0004, Qian Wang 0002, Kui Ren 0001, Dimitrios Koutsonikolas, Lu Su 0001, Yanjiao Chen |
IEEE Trans. Mob. Comput. | 5 |
| 2019 | Data-Driven Pricing for Sensing Effort Elicitation in Mobile Crowd Sensing SystemsabstractThe recent proliferation of human-carried mobile devices has given rise to mobile crowd sensing (MCS) systems that outsource sensory data collection to the public crowd. In order to identify truthful values from (crowd) workers' noisy or even conflicting sensory data, truth discovery algorithms, which jointly estimate workers' data quality and the underlying truths through quality-aware data aggregation, have drawn significant attention. However, the power of these algorithms could not be fully unleashed in MCS systems, unless workers' strategic reduction of their sensing effort is properly tackled. To address this issue, in this paper, we propose a payment mechanism, named Theseus, that deals with workers' such strategic behavior, and incentivizes high-effort sensing from workers. We ensure that, at the Bayesian Nash Equilibrium of the non-cooperative game induced by Theseus, all participating workers will spend their maximum possible effort on sensing, which improves their data quality. As a result, the aggregated results calculated subsequently by truth discovery algorithms based on workers' data will be highly accurate. Additionally, Theseus bears other desirable properties, including individual rationality and budget feasibility. We validate the desirable properties of Theseus through theoretical analysis, as well as extensive simulations. Haiming Jin, Baoxiang He, Lu Su 0001, Klara Nahrstedt, Xinbing Wang |
IEEE/ACM Trans. Netw. | 3 |
| 2019 | Privacy-Preserving Truth Discovery in Crowd Sensing SystemsabstractThe recent proliferation of human-carried mobile devices has given rise to the crowd sensing systems. However, the sensory data provided by individual participants are usually not reliable. To better utilize such sensory data, the topic of truth discovery, whose goal is to estimate user quality and infer reliable aggregated results through quality-aware data aggregation, has drawn significant attention. Though able to improve aggregation accuracy, existing truth discovery approaches fail to address the privacy concerns of individual users. In this article, we propose a novel privacy-preserving truth discovery (PPTD) framework, which can protect not only users’ sensory data but also their reliability scores derived by the truth discovery approaches. The key idea of the proposed framework is to perform weighted aggregation on users’ encrypted data using a homomorphic cryptosystem, which can guarantee both high accuracy and strong privacy protection. In order to deal with large-scale data, we also propose to parallelize PPTD with MapReduce framework. Additionally, we design an incremental PPTD scheme for the scenarios where the sensory data are collected in a streaming manner. Extensive experiments based on two real-world crowd sensing systems demonstrate that the proposed framework can generate accurate aggregated results while protecting users’ private information. Chenglin Miao, Lu Su 0001, Yaliang Li, Suxin Guo, Zhan Qin, Houping Xiao, Jing Gao 0004, Kui Ren 0001 |
ACM Trans. Sens. Networks | 3 |
| 2018 | Leveraging the Power of Informative Users for Local Event DetectionabstractDetecting local events (e.g., protests, accidents) in real-time is an important task needed by a wide spectrum of real-world applications. In recent years, with the proliferation of social media platforms, we can access massive geo- tagged social messages, which can serve as a precious resource for timely local event detection. However, existing local event detection methods either suffer from unsatisfactory performances or need intensive annotations. These limitations make existing methods impractical for large-scale applications. Through the analysis of real-world datasets, we found that the informativeness level of social media users, which is neglected by existing work, plays a highly critical role in distilling event-related information from noisy social media contexts. Motivated by this finding, we propose an unsupervised framework, named LEDetect, to estimate the informativeness level of social media users and leverage the power of highly informative users for local event detection. Experiments on a large-scale real-world dataset show that the proposed LEDetect model can improve the performance of event detection compared with the state-of-the-art unsupervised approach. Also, we use case studies to show that the events discovered by the proposed model are of high quality and the extracted highly informative users are reasonable. Hengtong Zhang, Fenglong Ma, Yaliang Li, Chao Zhang 0014, Yaqing Wang 0001, Jing Gao 0004, Lu Su 0001 |
ASONAM | 8 |
| 2018 | Multivariate Sleep Stage Classification using Hybrid Self-Attentive Deep Learning Networks
Ye Yuan 0006, Kebin Jia, Fenglong Ma, Guangxu Xun, Yaqing Wang 0001, Lu Su 0001, Aidong Zhang 0001 |
BIBM | 6 |
| 2018 | Towards Personalized Learning in Mobile Sensing SystemsabstractNowadays, mobile devices have become an important part of our daily life. Numerous mobile sensing applications are enabled by various mobile platforms, which leverage machine learning techniques to detect or classify the events of interest such as human activities and health conditions. To achieve this, each user is required to provide a considerable amount of training samples. However, in practice, a large portion of the users may provide only a few or even zero labels, due to various reasons such as privacy concern or simply laziness. A straightforward solution to this problem is to gather the data of all the users in a central database, and train a global classifier from the combined data. Such global classifier, however, may not work well since it ignores the variety in different users' data. To address this challenge, we propose PLOS, a Personalized Learning framework for mObile Sensing applications. PLOS can jointly model the commonness shared among the users as well as the differences between them, which are inferred from both the label information and the underlying structures of individual data. We further develop the distributed PLOS where the raw data of the users are locally processed so that the users only need to send model parameters to the server. Through extensive experiments on both synthetic data and real mobile sensing systems, we show that the proposed PLOS framework is scalable and efficient in energy, computation, and communication costs, and can achieve more accurate classification results compared with the baseline methods. Qi Li 0012, Lu Su 0001, Chenglin Miao, Quanquan Gu, Wenyao Xu |
ICDCS | 3 |
| 2018 | ApDeepSense: Deep Learning Uncertainty Estimation without the Pain for IoT ApplicationsabstractRecent advances in deep-learning-based applications have attracted a growing attention from the IoT community. These highly capable learning models have shown significant improvements in expected accuracy of various sensory inference tasks. One important and yet overlooked direction remains to provide uncertainty estimates in deep learning outputs. Since robustness and reliability of sensory inference results are critical to IoT systems, uncertainty estimates are indispensable for IoT applications. To address this challenge, we develop ApDeepSense, an effective and efficient deep learning uncertainty estimation method for resource-constrained IoT devices. ApDeepSense leverages an implicit Bayesian approximation that links neural networks to deep Gaussian processes, allowing output uncertainty to be quantified. Our approach is shown to significantly reduce the execution time and energy consumption of uncertainty estimation thanks to a novel layer-wise approximation that replaces the traditional computationally intensive sampling-based uncertainty estimation methods. ApDeepSense is designed for neural net-works trained using dropout; one of the most widely used regularization methods in deep learning. No additional training is needed for uncertainty estimation purposes. We evaluate ApDeepSense using four IoT applications on Intel Edison devices. Results show that ApDeepSense can reduce around 88.9% of the execution time and 90.0% of the energy consumption, while producing more accurate uncertainty estimates compared with state-of-the-art methods. Shuochao Yao, Yiran Zhao 0001, Huajie Shao, Chao Zhang 0014, Aston Zhang, Dongxin Liu, Shengzhong Liu, Lu Su 0001, Tarek F. Abdelzaher |
ICDCS | 8 |
| 2018 | MuVAN: A Multi-view Attention Network for Multivariate Temporal DataabstractRecent advances in attention networks have gained enormous interest in time series data mining. Various attention mechanisms are proposed to soft-select relevant timestamps from temporal data by assigning learnable attention scores. However, many real-world tasks involve complex multivariate time series that continuously measure target from multiple views. Different views may provide information of different levels of quality varied over time, and thus should be assigned with different attention scores as well. Unfortunately, the existing attention-based architectures cannot be directly used to jointly learn the attention scores in both time and view domains, due to the data structure complexity. Towards this end, we propose a novel multi-view attention network, namely MuVAN, to learn fine-grained attentional representations from multivariate temporal data. MuVAN is a unified deep learning model that can jointly calculate the two-dimensional attention scores to estimate the quality of information contributed by each view within different timestamps. By constructing a hybrid focus procedure, we are able to bring more diversity to attention, in order to fully utilize the multi-view information. To evaluate the performance of our model, we carry out experiments on three real-world benchmark datasets. Experimental results show that the proposed MuVAN model outperforms the state-of-the-art deep representation approaches in different real-world tasks. Analytical results through a case study demonstrate that MuVAN can discover discriminative and meaningful attention scores across views over time, which improves the feature representation of multivariate temporal data. Ye Yuan 0006, Guangxu Xun, Fenglong Ma, Yaqing Wang 0001, Nan Du 0001, Kebin Jia, Lu Su 0001, Aidong Zhang 0001 |
ICDM | 7 |
| 2018 | A Constrained Maximum Likelihood Estimator for Unguided Social SensingabstractThis paper develops a constrained expectation maximization algorithm (CEM) that improves the accuracy of truth estimation in unguided social sensing applications. Unguided social sensing refers to the act of leveraging naturally occurring observations on social media as “sensor measurements”, when the sources post at will and not in response to specific sensing campaigns or surveys. A key challenge in social sensing, in general, lies in estimating the veracity of reported observations, when the sources reporting these observations are of unknown reliability and their observations themselves cannot be readily verified. This problem is known as fact-finding. Unsupervised solutions have been proposed to the fact-finding problem that explore notions of internal data consistency in order to estimate observation veracity. This paper observes that unguided social sensing gives rise to a new (and very simple) constraint that dramatically reduces the space of feasible fact-finding solutions, hence significantly improving the quality of fact-finding results. The constraint relies on a simple approximate test of source independence, applicable to unguided sensing, and incorporates information about the number of independent sources of an observation to constrain the posterior estimate of its probability of correctness. Two different approaches are developed to test the independence of sources for purposes of applying this constraint, leading to two flavors of the CEM algorithm, we call CEM and CEM-Jaccard. We show using both simulation and real data sets collected from Twitter that by forcing the algorithm to converge to a solution in which the constraint is satisfied, the quality of solutions is significantly improved. Huajie Shao, Shuochao Yao, Yiran Zhao 0001, Chao Zhang 0014, Jinda Han, Lance M. Kaplan, Lu Su 0001, Tarek F. Abdelzaher |
INFOCOM | 7 |
| 2018 | Metric Learning from Probabilistic LabelsabstractMetric learning aims to learn a good distance metric that can capture the relationships among instances, and its importance has long been recognized in many fields. In the traditional settings of metric learning, an implicit assumption is that the associated labels of the instances are deterministic. However, in many real-world applications, the associated labels come naturally with probabilities instead of deterministic values. Thus, the existing metric learning methods cannot work well in these applications. To tackle this challenge, in this paper, we study how to effectively learn the distance metric from datasets that contain probabilistic information, and then propose two novel metric learning mechanisms for two types of probabilistic labels, i.e., the instance-wise probabilistic label and the group-wise probabilistic label. Compared with the existing metric learning methods, our proposed mechanisms are capable of learning distance metrics directly from the probabilistic labels with high accuracy. We also theoretically analyze the two proposed mechanisms and provide theoretical bounds on the sample complexity for both of them. Additionally, extensive experiments based on real-world datasets are conducted to verify the desirable properties of the proposed mechanisms. Mengdi Huai, Chenglin Miao, Yaliang Li, Qiuling Suo, Lu Su 0001, Aidong Zhang 0001 |
KDD | 5 |
| 2018 | An Efficient Two-Layer Mechanism for Privacy-Preserving Truth DiscoveryabstractSoliciting answers from online users is an efficient and effective solution to many challenging tasks. Due to the variety in the quality of users, it is important to infer their ability to provide correct answers during aggregation. Therefore, truth discovery methods can be used to automatically capture the user quality and aggregate user-contributed answers via a weighted combination. Despite the fact that truth discovery is an effective tool for answer aggregation, existing work falls short of the protection towards the privacy of participating users. To fill this gap, we propose perturbation-based mechanisms that provide users with privacy guarantees and maintain the accuracy of aggregated answers. We first present a one-layer mechanism, in which all the users adopt the same probability to perturb their answers. Aggregation is then conducted on perturbed answers but the aggregation accuracy could drop accordingly. To improve the utility, a two-layer mechanism is proposed where users are allowed to sample their own probabilities from a hyper distribution. We theoretically compare the one-layer and two-layer mechanisms, and prove that they provide the same privacy guarantee while the two-layer mechanism delivers better utility. This advantage is brought by the fact that the two-layer mechanism can utilize the estimated user quality information from truth discovery to reduce the accuracy loss caused by perturbation, which is confirmed by experimental results on real-world datasets. Experimental results also demonstrate the effectiveness of the proposed two-layer mechanism in privacy protection with tolerable accuracy loss in aggregation. Yaliang Li, Chenglin Miao, Lu Su 0001, Jing Gao 0004, Qi Li 0012, Bolin Ding, Zhan Qin, Kui Ren 0001 |
KDD | 3 |
| 2018 | EANN: Event Adversarial Neural Networks for Multi-Modal Fake News DetectionabstractAs news reading on social media becomes more and more popular, fake news becomes a major issue concerning the public and government. The fake news can take advantage of multimedia content to mislead readers and get dissemination, which can cause negative effects or even manipulate the public events. One of the unique challenges for fake news detection on social media is how to identify fake news on newly emerged events. Unfortunately, most of the existing approaches can hardly handle this challenge, since they tend to learn event-specific features that can not be transferred to unseen events. In order to address this issue, we propose an end-to-end framework named Event Adversarial Neural Network (EANN), which can derive event-invariant features and thus benefit the detection of fake news on newly arrived events. It consists of three main components: the multi-modal feature extractor, the fake news detector, and the event discriminator. The multi-modal feature extractor is responsible for extracting the textual and visual features from posts. It cooperates with the fake news detector to learn the discriminable representation for the detection of fake news. The role of event discriminator is to remove the event-specific features and keep shared features among events. Extensive experiments are conducted on multimedia datasets collected from Weibo and Twitter. The experimental results show our proposed EANN model can outperform the state-of-the-art methods, and learn transferable feature representations. Yaqing Wang 0001, Fenglong Ma, Zhiwei Jin, Ye Yuan 0006, Guangxu Xun, Kishlay Jha, Lu Su 0001, Jing Gao 0004 |
KDD | 7 |
| 2018 | TextTruth: An Unsupervised Approach to Discover Trustworthy Information from Multi-Sourced Text DataabstractTruth discovery has attracted increasingly more attention due to its ability to distill trustworthy information from noisy multi-sourced data without any supervision. However, most existing truth discovery methods are designed for structured data, and cannot meet the strong need to extract trustworthy information from raw text data as text data has its unique characteristics. The major challenges of inferring true information on text data stem from the multifactorial property of text answers (i.e., an answer may contain multiple key factors) and the diversity of word usages (i.e., different words may have the same semantic meaning). To tackle these challenges, in this paper, we propose a novel truth discovery method, named "TextTruth", which jointly groups the keywords extracted from the answers of a specific question into multiple interpretable factors, and infers the trustworthiness of both answer factors and answer providers. After that, the answers to each question can be ranked based on the estimated trustworthiness of factors. The proposed method works in an unsupervised manner, and thus can be applied to various application scenarios that involve text data. Experiments on three real-world datasets show that the proposed TextTruth model can accurately select trustworthy answers, even when these answers are formed by multiple factors. Hengtong Zhang, Yaliang Li, Fenglong Ma, Jing Gao 0004, Lu Su 0001 |
KDD | 5 |
| 2018 | Towards Environment Independent Device Free Human Activity RecognitionabstractDriven by a wide range of real-world applications, significant efforts have recently been made to explore device-free human activity recognition techniques that utilize the information collected by various wireless infrastructures to infer human activities without the need for the monitored subject to carry a dedicated device. Existing device free human activity recognition approaches and systems, though yielding reasonably good performance in certain cases, are faced with a major challenge. The wireless signals arriving at the receiving devices usually carry substantial information that is specific to the environment where the activities are recorded and the human subject who conducts the activities. Due to this reason, an activity recognition model that is trained on a specific subject in a specific environment typically does not work well when being applied to predict another subject's activities that are recorded in a different environment. To address this challenge, in this paper, we propose EI, a deep-learning based device free activity recognition framework that can remove the environment and subject specific information contained in the activity data and extract environment/subject-independent features shared by the data collected on different subjects under different environments. We conduct extensive experiments on four different device free activity recognition testbeds: WiFi, ultrasound, 60 GHz mmWave, and visible light. The experimental results demonstrate the superior effectiveness and generalizability of the proposed EI framework. Chenglin Miao, Fenglong Ma, Shuochao Yao, Yaqing Wang 0001, Ye Yuan 0006, Hongfei Xue, Chen Song 0001, Xin Ma 0006, Dimitrios Koutsonikolas, Wenyao Xu, Lu Su 0001 |
MobiCom | 12 |
| 2018 | Towards Data Poisoning Attacks in Crowd Sensing SystemsabstractWith the proliferation of sensor-rich mobile devices, crowd sensing has emerged as a new paradigm of collecting information from the physical world. However, the sensory data provided by the participating workers are usually not reliable. In order to identify truthful values from the crowd sensing data, the topic of truth discovery, whose goal is to estimate each worker's reliability and infer the underlying truths through weighted data aggregation, is widely studied. Since truth discovery incorporates workers' reliability into the aggregation procedure, it shows robustness to the data poisoning attacks, which are usually conducted by the malicious workers who aim to degrade the effectiveness of the crowd sensing systems through providing malicious sensory data. However, truth discovery is not perfect in all cases. In this paper, we study how to effectively conduct two types of data poisoning attacks, i.e., the availability attack and the target attack, against a crowd sensing system empowered with the truth discovery mechanism. We develop an optimal attack framework in which the attacker can not only maximize his attack utility but also disguise the introduced malicious workers as normal ones such that they cannot be detected easily. The desirable performance of the proposed framework is verified through extensive experiments conducted on a real-world crowd sensing system. Chenglin Miao, Qi Li 0012, Houping Xiao, Mengdi Huai, Lu Su 0001 |
MobiHoc | 6 |
| 2018 | Online Truth Discovery on Time Series DataabstractTruth discovery, with the goal of inferring true information from massive data through aggregating the information from multiple data sources, has attracted significant attention in recent years. It has demonstrated great advantages in real applications since it can automatically learn the reliability degrees of the data sources without supervision and in turn helps to find more reliable information. In many applications, however, the data may arrive in a stream and present various temporal patterns. Unfortunately, there is no existing truth discovery work that can handle such time series data. To tackle this challenge, we propose a novel online truth discovery framework that incorporates the predictions on the time series data into the truth estimation process. By jointly considering the multi-source information and the temporal patterns of the time series data, the proposed framework can improve the accuracy of the truth discovery results as well as the time series prediction. The effectiveness of the proposed framework is validated on both synthetic and real-world datasets. Liuyi Yao, Lu Su 0001, Qi Li 0012, Yaliang Li, Fenglong Ma, Jing Gao 0004, Aidong Zhang 0001 |
SDM | 2 |
| 2018 | FastDeepIoT: Towards Understanding and Optimizing Neural Network Execution Time on Mobile and Embedded DevicesabstractDeep neural networks show great potential as solutions to many sensing application problems, but their excessive resource demand slows down execution time, pausing a serious impediment to deployment on low-end devices. To address this challenge, recent literature focused on compressing neural network size to improve performance. We show that changing neural network size does not proportionally affect performance attributes of interest, such as execution time. Rather, extreme run-time nonlinearities exist over the network configuration space. Hence, we propose a novel framework, called FastDeepIoT, that uncovers the non-linear relation between neural network structure and execution time, then exploits that understanding to find network configurations that significantly improve the trade-off between execution time and accuracy on mobile and embedded devices. FastDeepIoT makes two key contributions. First, FastDeepIoT automatically learns an accurate and highly interpretable execution time model for deep neural networks on the target device. This is done without prior knowledge of either the hardware specifications or the detailed implementation of the used deep learning library. Second, FastDeepIoT informs a compression algorithm how to minimize execution time on the profiled device without impacting accuracy. We evaluate FastDeepIoT using three different sensing-related tasks on two mobile devices: Nexus 5 and Galaxy Nexus. FastDeepIoT further reduces the neural network execution time by 48% to 78% and energy consumption by 37% to 69% compared with the state-of-the-art compression algorithms. Shuochao Yao, Yiran Zhao 0001, Huajie Shao, Shengzhong Liu, Dongxin Liu, Lu Su 0001, Tarek F. Abdelzaher |
SenSys | 6 |
| 2018 | Attack under Disguise: An Intelligent Data Poisoning Attack Mechanism in CrowdsourcingabstractAs an effective way to solicit useful information from the crowd, crowdsourcing has emerged as a popular paradigm to solve challenging tasks. However, the data provided by the participating workers are not always trustworthy. In real world, there may exist malicious workers in crowdsourcing systems who conduct the data poisoning attacks for the purpose of sabotage or financial rewards. Although data aggregation methods such as majority voting are conducted on workers» labels in order to improve data quality, they are vulnerable to such attacks as they treat all the workers equally. In order to capture the variety in the reliability of workers, the Dawid-Skene model, a sophisticated data aggregation method, has been widely adopted in practice. By conducting maximum likelihood estimation (MLE) using the expectation maximization (EM) algorithm, the Dawid-Skene model can jointly estimate each worker»s reliability and conduct weighted aggregation, and thus can tolerate the data poisoning attacks to some degree. However, the Dawid-Skene model still has weakness. In this paper, we study the data poisoning attacks against such crowdsourcing systems with the Dawid-Skene model empowered. We design an intelligent attack mechanism, based on which the attacker can not only achieve maximum attack utility but also disguise the attacking behaviors. Extensive experiments based on real-world crowdsourcing datasets are conducted to verify the desirable properties of the proposed mechanism. Chenglin Miao, Qi Li 0012, Lu Su 0001, Mengdi Huai, Jing Gao 0004 |
WWW | 3 |
| 2018 | Harvest Energy from the Water: A Self-Sustained Wireless Water Quality Sensing SystemabstractWater quality data is incredibly important and valuable, but its acquisition is not always trivial. A promising solution is to distribute a wireless sensor network in water to measure and collect the data; however, a drawback exists in that the batteries of the system must be replaced or recharged after being exhausted. To mitigate this issue, we designed a self-sustained water quality sensing system that is powered by renewable bioenergy generated from microbial fuel cells (MFCs). MFCs collect the energy released from native magnesium oxidizing microorganisms (MOMs) that are abundant in natural waters. The proposed energy-harvesting technology is environmentally friendly and can provide maintenance-free power to sensors for several years. Despite these benefits, an MFC can only provide microwatt-level power that is not sufficient to continuously power a sensor. To address this issue, we designed a power management module to accumulate energy when the input voltage is as low as 0.33V. We also proposed a radio-frequency (RF) activation technique to remotely activate sensors that otherwise are switched off in default. With this innovative technique, a sensor’s energy consumption in sleep mode can be completely avoided. Additionally, this design can enable on-demand data acquisitions from sensors. We implement the proposed system and evaluate its performance in a stream. In 3-month field experiments, we find the system is able to reliably collect water quality data and is robust to environment changes. Qi Chen 0018, Ye Liu 0004, Guangchi Liu, Qing Yang 0003, Xianming Shi, Lu Su 0001, Quanlong Li |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2018 | Cooperative and Integrated Vehicle and Intersection Control for Energy Efficiency (CIVIC-E2)abstractRecent advances in connected vehicle technologies enable vehicles and signal controllers to cooperate and improve the traffic management at intersections. This paper explores the opportunity for cooperative and integrated vehicle and intersection control for energy efficiency (CIVIC-E2) to contribute to a more sustainable transportation system. We propose a two-level approach that jointly optimizes the traffic signal timing and vehicles' approach speed, with the objective being to minimize total energy consumption for all vehicles passing through an isolated intersection. More specifically, at the intersection level, a dynamic programming algorithm is designed to find the optimal signal timing by explicitly considering the arrival time and energy profile of each vehicle. At the vehicle level, a model predictive control strategy is adopted to ensure that vehicles pass through the intersection in a timely fashion. Our simulation study has shown that the proposed CIVIC-E2system can significantly improve intersection performance under various traffic conditions. Compared with conventional fixed-time and actuated signal control strategies, the proposed algorithm can reduce energy consumption and queue length by up to 31% and 95%, respectively. Yunfei Hou, Salaheldeen M. S. Seliman, Enshu Wang, Jeffrey D. Gonder, Eric Wood, Qing He 0011, Adel W. Sadek, Lu Su 0001, Chunming Qiao |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2018 | Incentive Mechanism for Privacy-Aware Data Aggregation in Mobile Crowd Sensing Systems
Haiming Jin, Lu Su 0001, Houping Xiao, Klara Nahrstedt |
IEEE/ACM Trans. Netw. | 2 |
| 2018 | Towards Quality Aware Information Integration in Distributed Sensing SystemsabstractIn this paper, we present GDA, a generalized decision aggregation framework that integrates information from distributed sensor nodes for decision making in a resource efficient manner. Different from traditional approaches, our proposed GDA framework is able to not only estimate the reliability of each sensor, but also take advantage of its confidence information, and thus achieves higher decision accuracy. Targeting generalized problem domains, our framework can naturally handle the scenarios where different sensor nodes observe different sets of events whose numbers of possible classes may also be different. GDA also makes no assumption about the availability level of ground truth label information, while being able to take advantage of any if present. For these reasons, our approach can be applied to a much broader spectrum of sensing scenarios. In this paper, we also propose two extensions of the GDA framework, i.e., incremental GDA (I-GDA) and parallel GDA (P-GDA) to deal with streaming and large-scale data. The advantages of our proposed methods are demonstrated through both theoretic analysis and extensive experiments. Chenglin Miao, Lu Su 0001, Qi Li 0012, Shaohan Hu, Shiguang Wang, Jing Gao 0004, Hengchang Liu, Tarek F. Abdelzaher, Jiawei Han 0001, Xue (Steve) Liu, Yan Gao 0010, Lance M. Kaplan |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Travel purpose inference with GPS trajectories, POIs, and geo-tagged social media dataabstractIn our daily lives, travel takes up an important part, and many trips are generated everyday, such as going to school or shopping. With the widely adoption of GPS-integrated devices, a large amount of trips can be recorded with GPS trajectories. These trajectories are represented by sequences of geo-coordinates and can help us answer simple questions such as “where did you go”. However, there is another important question awaiting to be answered, that is “what did/will you do”, i.e., the trip purpose inference. In practice, people's trip purposes are very important in understanding travel behaviors and estimating travel demands. Obviously, it is very challenging to infer trip purposes solely based on the trajectories, because the GPS devices are not accurate enough to pinpoint the venues visited. In this paper, we infer individual's trip purposes by combining the knowledge from heterogeneous data sources including trajectories, POIs and social media data. The proposed dynamic Bayesian network model captures three important factors: the sequential properties of trip activities, the functionality and POI popularity of trip end areas. Extensive experiments are conducted on real-world data sets with trajectories of 8,361 residents and the 6.9 million geo-tagged tweets in the Bay area. Experimental results demonstrate the advantages of the proposed method on correctly inferring the trip purposes. Chuishi Meng, Qing He 0011, Lu Su 0001, Jing Gao 0004 |
IEEE BigData | 4 |
| 2017 | City-wide Traffic Volume Inference with Loop Detector Data and Taxi TrajectoriesabstractThe traffic volume on road segments is a vital property of the transportation efficiency. City-wide traffic volume information can benefit people with their everyday life, and help the government on better city planning. However, there are no existing methods that can monitor the traffic volume of every road, because they are either too expensive or inaccurate. Fortunately, nowadays we can collect a large amount of urban data which provides us the opportunity to tackle this problem. In this paper, we propose a novel framework to infer the city-wide traffic volume information with data collected by loop detectors and taxi trajectories. Although these two data sets are incomplete, sparse and from quite different domains, the proposed spatio-temporal semi-supervised learning model can take the full advantages of both data and accurately infer the volume of each road. In order to provide a better interpretation on the inference results, we also derive the confidence of the inference based on spatio-temporal properties of traffic volume. Real-world data was collected from 155 loop detectors and 6,918 taxis over a period of 17 days in Guiyang China. The experiments performed on this large urban data set demonstrate the advantages of the proposed framework on correctly inferring the traffic volume in a city-wide scale. Chuishi Meng, Xiuwen Yi, Lu Su 0001, Jing Gao 0004, Yu Zheng 0004 |
SIGSPATIAL/GIS | 3 |
| 2017 | You Can Hear But You Cannot Steal: Defending Against Voice Impersonation Attacks on SmartphonesabstractVoice, as a convenient and efficient way of information delivery, has a significant advantage over the conventional keyboard-based input methods, especially on small mobile devices such as smartphones and smartwatches. However, the human voice could often be exposed to the public, which allows an attacker to quickly collect sound samples of targeted victims and further launch voice impersonation attacks to spoof those voice-based applications. In this paper, we propose the design and implementation of a robust software-only voice impersonation defense system, which is tailored for mobile platforms and can be easily integrated with existing off-the-shelf smart devices. In our system, we explore magnetic field emitted from loudspeakers as the essential characteristic for detecting machine-based voice impersonation attacks. Furthermore, we use a state-of-the-art automatic speaker verification system to defend against human imitation attacks. Finally, our evaluation results show that our system achieves simultaneously high accuracy (100%) and low equal error rates (EERs) (0%) in detecting the machine-based voice impersonation attack on smartphones. Si Chen 0009, Kui Ren 0001, Sixu Piao, Cong Wang 0001, Qian Wang 0002, Jian Weng 0001, Lu Su 0001, David Mohaisen |
ICDCS | 7 |
| 2017 | When Smart TV Meets CRN: Privacy-Preserving Fine-Grained Spectrum AccessabstractDynamic spectrum sharing techniques applied in the UHF TV band have been developed to allow secondary WiFi transmission in areas with active TV users. This technique of dynamically controlling the exclusion zone enables vastly increasing secondary spectrum re-use, compared to the "TV white space" model where TV transmitters determine the exclusion zone and only "idle" channels can be re-purposed. However, in current such dynamic spectrum sharing systems, the sensitive operation parameters of both primary TV users (PUs) and secondary users (SUs) need to be shared with the spectrum database controller (SDC) for the purpose of realizing efficient spectrum allocation. Since such SDC server is not necessarily operated by a trusted third party, those current systems might cause essential threatens to the privacy requirement from both PUs and SUs. To address this privacy issue, this paper proposes a privacy-preserving spectrum sharing system between PUs and SUs, which realizes the spectrum allocation decision process using efficient multi-party computation (MPC) technique. In this design, the SDC only performs secure computation over encrypted input from PUs and SUs such that none of the PU or SU operation parameters will be revealed to SDC. The evaluation of its performance illustrates that our proposed system based on efficient MPC techniques can perform dynamic spectrum allocation process between PUs and SUs efficiently while preserving users' privacy. Chaowen Guan, David Mohaisen, Lu Su 0001, Kui Ren 0001, Yaling Yang |
ICDCS | 4 |
| 2017 | Discovering Truths from Distributed DataabstractIn the big data era, the information about the same object collected from multiple sources is inevitably conflicting. The task of identifying true information (i.e., the truths) among conflicting data is referred to as truth discovery, which incorporates the estimation of source reliability degrees into the aggregation of multi-source data. However, in many real-world applications, large-scale data are distributed across multiple servers. Traditional truth discovery approaches cannot handle this scenario due to the constraints of communication overhead and privacy concern. Another limitation of most existing work is that they ignore the differences among objects, i.e., they treat all the objects equally. This limitation would be exacerbated in distributed environments where significant differences exist among the objects. To tackle the aforementioned issues, in this paper, we propose a novel distributed truth discovery framework (DTD), which can effectively and efficiently aggregate conflicting data stored across distributed servers, with the differences among the objects as well as the importance level of each server being considered. The proposed framework consists of two steps: the local truth computation step conducted by each local server and the central truth estimation step taking place in the central server. Specifically, we introduce the uncertainty values to model the differences among objects, and propose a new uncertainty-based truth discovery method (UbTD) for calculating the true information of objects in each local server. The outputs of the local truth computation step include the estimated local truths and the variances of objects, which are the input information of the central truth estimation step. To infer the final true information in the central server, we propose a new algorithm to aggregate the outputs of all the local servers with the quality of different local servers taken into account. The proposed distributed truth discovery framework can infer object truths without delivering any raw data to the central server, and thus can reduce communication overhead as well as preserve data privacy. Experimental results on three real world datasets show that the proposed DTD framework can efficiently estimate object truths with accuracy guarantee, and the proposed UbTD algorithm significantly outperforms the state-of-the-art batch truth discovery approaches. Yaqing Wang 0001, Fenglong Ma, Lu Su 0001, Jing Gao 0004 |
ICDM | 3 |
| 2017 | CENTURION: Incentivizing multi-requester mobile crowd sensingabstractThe recent proliferation of increasingly capable mobile devices has given rise to mobile crowd sensing (MCS) systems that outsource the collection of sensory data to a crowd of participating workers that carry various mobile devices. Aware of the paramount importance of effectively incentivizing participation in such systems, the research community has proposed a wide variety of incentive mechanisms. However, different from most of these existing mechanisms which assume the existence of only one data requester, we consider MCS systems with multiple data requesters, which are actually more common in practice. Specifically, our incentive mechanism is based on double auction, and is able to stimulate the participation of both data requesters and workers. In real practice, the incentive mechanism is typically not an isolated module, but interacts with the data aggregation mechanism that aggregates workers' data. For this reason, we propose CENTURION, a novel integrated framework for multi-requester MCS systems, consisting of the aforementioned incentive and data aggregation mechanism. CENTURION's incentive mechanism satisfies truthfulness, individual rationality, computational efficiency, as well as guaranteeing non-negative social welfare, and its data aggregation mechanism generates highly accurate aggregated results. The desirable properties of CENTURION are validated through both theoretical analysis and extensive simulations. Haiming Jin, Lu Su 0001, Klara Nahrstedt |
INFOCOM | 2 |
| 2017 | A lightweight privacy-preserving truth discovery framework for mobile crowd sensing systemsabstractThe recent proliferation of human-carried mobile devices has given rise to the mobile crowd sensing (MCS) systems. However, the sensory data provided by the participating workers are usually not reliable. As an efficient technique to extract truthful information from unreliable data, truth discovery has drawn significant attention. Currently, the privacy concern of the participating workers poses a major challenge on the design of truth discovery mechanisms. Although the existing mechanism can conduct truth discovery with high accuracy and strong privacy guarantee, tremendous overhead is incurred on the worker side. In this paper, we propose a novel lightweight privacy preserving truth discovery framework, L-PPTD, which is implemented by involving two non-colluding cloud platforms and adopting additively homomorphic cryptosystem. This framework not only achieves the protection of each worker's sensory data and reliability information but also introduces little overhead to the workers. In order to further reduce each worker's overhead in the scenarios where only the sensory data need to be protected, we propose another more lightweight framework named L2-PPTD. The desirable performance of the proposed frameworks is verified through extensive experiments conducted on real world MCS systems. Chenglin Miao, Lu Su 0001, Yaliang Li, Miaomiao Tian 0001 |
INFOCOM | 2 |
| 2017 | Unsupervised Discovery of Drug Side-Effects from Heterogeneous Data SourcesabstractDrug side-effects become a worldwide public health concern, which are the fourth leading cause of death in the United States. Pharmaceutical industry has paid tremendous effort to identify drug side-effects during the drug development. However, it is impossible and impractical to identify all of them. Fortunately, drug side-effects can also be reported on heterogeneous platforms (i.e., data sources), such as FDA Adverse Event Reporting System and various online communities. However, existing supervised and semi-supervised approaches are not practical as annotating labels are expensive in the medical field. In this paper, we propose a novel and effective unsupervised model Sifter to automatically discover drug side-effects. Sifter enhances the estimation on drug side-effects by learning from various online platforms and measuring platform-level and user-level quality simultaneously. In this way, Sifter demonstrates better performance compared with existing approaches in terms of correctly identifying drug side-effects. Experimental results on five real-world datasets show that Sifter can significantly improve the performance of identifying side-effects compared with the state-of-the-art approaches. Fenglong Ma, Chuishi Meng, Houping Xiao, Qi Li 0012, Jing Gao 0004, Lu Su 0001, Aidong Zhang 0001 |
KDD | 6 |
| 2017 | Theseus: Incentivizing Truth Discovery in Mobile Crowd Sensing SystemsabstractThe recent proliferation of human-carried mobile devices has given rise to mobile crowd sensing (MCS) systems that outsource sensory data collection to the public crowd. In order to identify truthful values from (crowd) workers' noisy or even conflicting sensory data, truth discovery algorithms, which jointly estimate workers' data quality and the underlying truths through quality-aware data aggregation, have drawn significant attention. However, the power of these algorithms could not be fully unleashed in MCS systems, unless workers' strategic reduction of their sensing effort is properly tackled. To address this issue, in this paper, we propose a payment mechanism, named Theseus, that deals with workers' such strategic behavior, and incentivizes high-effort sensing from workers. We ensure that, at the Bayesian Nash Equilibrium of the non-cooperative game induced by Theseus, all participating workers will spend their maximum possible effort on sensing, which improves their data quality. As a result, the aggregated results calculated subsequently by truth discovery algorithms based on workers' data will be highly accurate. Additionally, Theseus bears other desirable properties, including individual rationality and budget feasibility. We validate the desirable properties of Theseus through theoretical analysis, as well as extensive simulations. Haiming Jin, Lu Su 0001, Klara Nahrstedt |
MobiHoc | 2 |
| 2017 | DeepIoT: Compressing Deep Neural Network Structures for Sensing Systems with a Compressor-Critic FrameworkabstractRecent advances in deep learning motivate the use of deep neutral networks in sensing applications, but their excessive resource needs on constrained embedded devices remain an important impediment. A recently explored solution space lies in compressing (approximating or simplifying) deep neural networks in some manner before use on the device. We propose a new compression solution, called DeepIoT, that makes two key contributions in that space. First, unlike current solutions geared for compressing specific types of neural networks, DeepIoT presents a unified approach that compresses all commonly used deep learning structures for sensing applications, including fully-connected, convolutional, and recurrent neural networks, as well as their combinations. Second, unlike solutions that either sparsify weight matrices or assume linear structure within weight matrices, DeepIoT compresses neural network structures into smaller dense matrices by finding the minimum number of non-redundant hidden elements, such as filters and dimensions required by each layer, while keeping the performance of sensing applications the same. Importantly, it does so using an approach that obtains a global view of parameter redundancies, which is shown to produce superior compression. The compressed model generated by DeepIoT can directly use existing deep learning libraries that run on embedded and mobile systems without further modifications. We conduct experiments with five different sensing-related tasks on Intel Edison devices. DeepIoT outperforms all compared baseline algorithms with respect to execution time and energy consumption by a significant margin. It reduces the size of deep neural networks by 90% to 98.9%. It is thus able to shorten execution time by 71.4% to 94.5%, and decrease energy consumption by 72.2% to 95.7%. These improvements are achieved without loss of accuracy. The results underscore the potential of DeepIoT for advancing the exploitation of deep neural networks on resource-constrained embedded devices. Shuochao Yao, Yiran Zhao 0001, Aston Zhang, Lu Su 0001, Tarek F. Abdelzaher |
SenSys | 4 |
| 2017 | VehSense: Slippery Road Detection Using SmartphonesabstractThis paper investigates a new application of vehicular sensing: detecting and reporting the slippery road conditions. We describe a system and associated algorithm to monitor vehicle skidding events using smartphones and OBD-II (On board Diagnostics) adapters. This system, which we call the VehSense, gathers data from smartphone inertial sensors and vehicle wheel speed sensors, and processes the data to monitor slippery road conditions in real-time. Specifically, two speed readings are collected: 1) ground speed, which is estimated by vehicle acceleration and rotation, and 2) wheel speed, which is retrieved from the OBD-II interface. The mismatch between these two speeds is used to infer a skidding event. Without tapping into vehicle manufactures' proprietary data (e.g., antilock braking system), VehSense is compatible with most of the passenger vehicles, and thus can be easily deployed. We evaluate our system on snow-covered roads at Buffalo, and show that it can detect vehicle skidding effectively. Yunfei Hou, Tong Guan, Shaohan Hu, Lu Su 0001, Chunming Qiao |
VTC Spring | 5 |
| 2017 | On Exploiting Structured Human Interactions to Enhance Sensing Accuracy in Cyber-physical SystemsabstractIn this article, we describe a general methodology for enhancing sensing accuracy in cyber-physical systems that involve structured human interactions in noisy physical environment. We define structured human interactions as domain-specific workflow. A novel workflow-aware sensing model is proposed to jointly correct unreliable sensor data and keep track of states in a workflow. We also propose a new inference algorithm to handle cases with partially known states and objects as supervision. Our model is evaluated with extensive simulations. As a concrete application, we develop a novel log service called Emergency Transcriber , which can automatically document operational procedures followed by teams of first responders in emergency response scenarios. Evaluation shows that our system has significant improvement over commercial off-the-shelf (COTS) sensors and keeps track of workflow states with high accuracy in noisy physical environment. Shaohan Hu, Shiguang Wang, Renato Mancuso 0001, Minje Kim 0001, Po-Liang Wu, Lu Su 0001, Lui Sha, Tarek F. Abdelzaher |
ACM Trans. Cyber Phys. Syst. | 8 |
| 2017 | A Weighted Crowdsourcing Approach for Network Quality Measurement in Cellular Data NetworksabstractWith ubiquitous smartphone usages, it is important for network providers to provide high-quality service to every user in the network. To make more effective planning and scheduling, network providers need an accurate estimate of network quality for base stations and cells from the perspective of user experience. Traditional drive testing approach provides a quality measurement for each area and the quality measurement is obtained from the equipment in a moving vehicle. This approach suffers from the limitations of high costs, low coverage, and out-of-date values. In this paper, we propose a novel crowdsourcing approach for the task of network quality estimation, which incurs little costs and provides timely and accurate quality estimation. The proposed approach collects quality measurements from individual end users within a certain network or cell coverage area, and then aggregates these measurements to obtain a global measurement of network quality. We propose an effective aggregation scheme which infers the information weights of end users and incorporates such weights into the estimation of network quality. Experiments are conducted on two datasets collected from citywide 3G networks, which involve 616,796 users and 22,715 cells. We validate the effectiveness of the proposed approach compared with baseline method. From the aggregated measurement results, we observe some interesting patterns about network quality, which can be explained by network usage and traffic behavior. We also show that proposed approach runs in linear time. Yaliang Li, Jing Gao 0004, Patrick P. C. Lee, Lu Su 0001, Caifeng He, Wei Fan 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2016 | Influence-Aware Truth DiscoveryabstractIn the age of big data, information for the same entity can be obtained from different sources, which is inevitably conflicting. Therefore, aggregation methods are needed to identify the trustworthy information from such conflicting data. Truth discovery, which improves the aggregation results by estimating source trustworthiness and discovering truths simultaneously, has become an emerging field. Most truth discovery methods assume that sources make their claims independently, which may not be true in practice. As a matter of fact, influences among sources are ubiquitous and the claims made by one source may be influenced by others. Although there is some work that considers source correlation, those methods are designed to handle categorical claims, which is not general enough to represent the complicated real world applications. To tackle these challenges in truth discovery, we propose an unsupervised probabilistic model named IATD. The model takes source correlations as prior for influence derivation. To model influences among sources, we introduce "claim trustworthiness", which fuses the trustworthiness of the source which provides the claim and the trustworthiness of its influencers. Besides, the proposed model can handle different data types using different distributions in the probabilistic model. Experiments on real-world datasets show that IATD model can improve the aggregation performance compared with the state-of-the-art truth discovery approaches. The properties of IATD model are further illustrated using simulated datasets. Hengtong Zhang, Qi Li 0012, Fenglong Ma, Houping Xiao, Yaliang Li, Jing Gao 0004, Lu Su 0001 |
CIKM | 7 |
| 2016 | Enabling Privacy-Preserving Incentives for Mobile Crowd Sensing SystemsabstractRecent years have witnessed the proliferation of mobile crowd sensing (MCS) systems that leverage the public crowd equipped with various mobile devices (e.g., smartphones, smartglasses, smartwatches) for large scale sensing tasks. Because of the importance of incentivizing worker participation in such MCS systems, several auction-based incentive mechanisms have been proposed in past literature. However, these mechanisms fail to consider the preservation of workers' bid privacy. Therefore, different from prior work, we propose a differentially private incentive mechanism that preserves the privacy of each worker's bid against the other honest-but-curious workers. The motivation of this design comes from the concern that a worker's bid usually contains her private information that should not be disclosed. We design our incentive mechanism based on the single-minded reverse combinatorial auction. Specifically, we design a differentially private, approximately truthful, individual rational, and computationally efficient mechanism that approximately minimizes the platform's total payment with a guaranteed approximation ratio. The advantageous properties of the proposed mechanism are justified through not only rigorous theoretical analysis but also extensive simulations. Haiming Jin, Lu Su 0001, Bolin Ding, Klara Nahrstedt, Nikita Borisov |
ICDCS | 2 |
| 2016 | On Source Dependency Models for Reliable Social Sensing: Algorithms and Fundamental Error BoundsabstractThis paper develops a simplified dependency model for sources on social networks that is shown to improve the quality of fact-finding -- assessing veracity of observations shared on social media. Recent literature developed a mathematical approach for exploiting social networks, such as Twitter, as noisy sensor networks that report observations on the state of the physical world. It was shown that the quality of state estimation from such noisy data, known as fact-finding, was a function of assumptions made regarding the independence of sources or lack thereof. When sources propagate information they hear from others (without verification), correlated errors may arise that degrade fact-finding performance. This work advances the state of the art by developing a simplified model of dependencies between sources and designing an improved dependency-aware estimator to assess veracity of observations, taking into account the observed dependency structure. A fundamental error bound is derived for this estimator to understand the gap in its performance from optimal. It is shown that the new estimator outperforms state of the art fact-finders and, in some cases, yields an accuracy close to the fundamental error bound. Shuochao Yao, Shaohan Hu, Shen Li 0002, Yiran Zhao 0001, Lu Su 0001, Lance M. Kaplan, Aylin Yener, Tarek F. Abdelzaher |
ICDCS | 5 |
| 2016 | Demonstration Abstract: A Novel Human Tracking and Localization System Based on Pyroelectric Infrared SensorsabstractIn this Demo, we present the design and implementation of a novel human localization and tracking system based on pyroelectric infrared(PIR) sensors. It is a new approach to setup several low-cost sensors to improve the simplicity of deployment. We first present a novel design in mechanical architecture, circuit board and sensors which enables our PIR sensor based system to detect the moving objects and track human beings in real time. Different from the previous work which only adopts binary output, we further explore the usage of analog signals from the sensors. This improvement is beneficial to accurate estimation of the distance between the device and human body. We have implemented several prototypes for evaluation and demonstration. According to the experimental results, our system can achieve an accuracy of 0.113m in an area of 12m × 6m. Guo Liu, Jianwei Niu 0002, Lu Su 0001 |
IPSN | 6 |
| 2016 | Recursive Ground Truth Estimator for Social Data StreamsabstractThe paper develops a recursive state estimator for social network data streams that allows exploitation of social networks, such as Twitter, as sensor networks to reliably observe physical events. Recent literature suggested using social networks as sensor networks leveraging the fact that much of the information upload on the former constitutes acts of sensing. A significant challenge identified in that context was that source reliability is often unknown, leading to uncertainty regarding the veracity of reported observations. Multiple truth finding systems were developed to solve this problem, generally geared towards batch analysis of offline datasets. This work complements the present batch approaches by developing an online recursive state estimator that recovers ground truth from streaming data. In this paper, we model physical world state by a set of binary signals (propositions, called assertions, about world state) and the social network as a noisy medium, where distortion, fabrication, omissions, and duplication are introduced. Our recursive state estimator is designed to recover the original binary signal (the true propositions) from the received noisy signal, essentially decoding the unreliable social network output to obtain the best estimate of ground truth in the physical world. Results show that the estimator is both effective and efficient at recovering the original signal with a high degree of accuracy. The estimator gives rise to a novel situation awareness tool that can be used for reliably following unfolding events in real time, using dynamically arriving social network data. Shuochao Yao, Md. Tanvir Al Amin, Lu Su 0001, Shaohan Hu, Shen Li 0002, Shiguang Wang, Yiran Zhao 0001, Tarek F. Abdelzaher, Lance M. Kaplan, Charu C. Aggarwal, Aylin Yener |
IPSN | 3 |
| 2016 | Sparse approximations of directed information graphsabstractGiven a network of agents interacting over time, which few interactions best characterize the dynamics of the whole network? We propose an algorithm that finds the optimal sparse approximation of a network. The user controls the level of sparsity by specifying the total number of edges. The networks are represented using directed information graphs, a graphical model that depicts causal influences between agents in a network. Goodness of approximation is measured with Kullback-Leibler divergence. The algorithm finds the best approximation with no assumptions on the topology or the class of the joint distribution. Christopher J. Quinn, Ali Pinar, Jing Gao 0004, Lu Su 0001 |
ISIT | 4 |
| 2016 | Towards Confidence in the Truth: A Bootstrapping based Truth Discovery ApproachabstractThe demand for automatic extraction of true information (i.e., truths) from conflicting multi-source data has soared recently. A variety of truth discovery methods have witnessed great successes via jointly estimating source reliability and truths. All existing truth discovery methods focus on providing a point estimator for each object's truth, but in many real-world applications, confidence interval estimation of truths is more desirable, since confidence interval contains richer information. To address this challenge, in this paper, we propose a novel truth discovery method (ETCIBoot) to construct confidence interval estimates as well as identify truths, where the bootstrapping techniques are nicely integrated into the truth discovery procedure. Due to the properties of bootstrapping, the estimators obtained by ETCIBoot are more accurate and robust compared with the state-of-the-art truth discovery approaches. Theoretically, we prove the asymptotical consistency of the confidence interval obtained by ETCIBoot. Experimentally, we demonstrate that ETCIBoot is not only effective in constructing confidence intervals but also able to obtain better truth estimates. Houping Xiao, Jing Gao 0004, Qi Li 0012, Fenglong Ma, Lu Su 0001, Yunlong Feng, Aidong Zhang 0001 |
KDD | 5 |
| 2016 | A Truth Discovery Approach with Theoretical GuaranteeabstractIn the information age, people can easily collect information about the same set of entities from multiple sources, among which conflicts are inevitable. This leads to an important task, truth discovery, i.e., to identify true facts (truths) via iteratively updating truths and source reliability. However, the convergence to the truths is never discussed in existing work, and thus there is no theoretical guarantee in the results of these truth discovery approaches. In contrast, in this paper we propose a truth discovery approach with theoretical guarantee. We propose a randomized gaussian mixture model (RGMM) to represent multi-source data, where truths are model parameters. We incorporate source bias which captures its reliability degree into RGMM formulation. The truth discovery task is then modeled as seeking the maximum likelihood estimate (MLE) of the truths. Based on expectation-maximization (EM) techniques, we propose population-based (i.e., on the limit of infinite data) and sample-based (i.e., on a finite set of samples) solutions for the MLE. Theoretically, we prove that both solutions are contractive to an ε-ball around the MLE, under certain conditions. Experimentally, we evaluate our method on both simulated and real-world datasets. Experimental results show that our method achieves high accuracy in identifying truths with convergence guarantee. Houping Xiao, Jing Gao 0004, Zhaoran Wang 0001, Lu Su 0001, Han Liu 0001 |
KDD | 5 |
| 2016 | Messages behind the sound: real-time hidden acoustic signal capture with smartphonesabstractWith the ever-increasing use of smart devices, recent research endeavors have led to unobtrusive screen-camera communication channel designs, which allow simultaneous screen viewing and hidden screen-camera communication. Such practices, albeit innovative and effective, require well-controlled alignment of camera and screen and obstacle-free access. Qian Wang 0002, Kui Ren 0001, Man Zhou 0004, Tao Lei 0005, Dimitrios Koutsonikolas, Lu Su 0001 |
MobiCom | 6 |
| 2016 | Real-time hidden acoustic signal capture with smartphones: demoabstractWith the ever-increasing use of smart devices, recent research endeavors have led to unobtrusive screen-camera communication channel designs, which allow simultaneous screen viewing and hidden screen-camera communication. Such practices, albeit innovative and effective, require well-controlled alignment of camera and screen and obstacle-free access. In this demo, we present Dolphin, a novel form of real-time acoustics-based dual-channel communication, which uses a speaker and the microphones on off-the-shelf smartphones to achieve concurrent audible and hidden communication. By leveraging masking effects of the human auditory system and readily available audio signals in our daily lives, Dolphin ensures real-time unobtrusive speaker-microphone data communication, while, at the same time, it overcomes the main limitations of existing screen-camera links. Qian Wang 0002, Kui Ren 0001, Man Zhou 0004, Tao Lei 0005, Dimitrios Koutsonikolas, Lu Su 0001 |
MobiCom | 6 |
| 2016 | Towards distributed ensemble clustering for networked sensing systems: a novel geometric approachabstractGiven a set of different clustering solutions to a unified dataset, ensemble clustering is to aggregate them to yield a more accurate and robust solution. In recent years, ensemble clustering has been extensively studied and successfully applied to many areas. In this paper, we study a new variant of ensemble clustering, distributed ensemble clustering, motivated by the proliferation of networked sensing systems where communication is enabled between only connected nodes. Our goal is to aggregate the clustering solutions produced by the sensor nodes that observe the same set of objects. Different from traditional ensemble clustering problems, distributed ensemble clustering aims to achieve not only accurate clustering results, but also low communication cost among the nodes. To this end, we build a novel geometric optimization model that can be efficiently solved with theoretical quality guarantee. The proposed approach, bearing nice geometric properties, can be easily adapted to distributed settings without any sacrifice of clustering quality, and facilitates a dimension reduction procedure which can significantly reduce the communication complexity. We validate our approach on two benchmark datasets. Experimental results suggest that our approach can efficiently solve the distributed ensemble clustering problem, and outperform the baselines on both clustering accuracy and communication cost. Hu Ding 0003, Lu Su 0001, Jinhui Xu 0001 |
MobiHoc | 2 |
| 2016 | INCEPTION: incentivizing privacy-preserving data aggregation for mobile crowd sensing systemsabstractThe recent proliferation of human-carried mobile devices has given rise to mobile crowd sensing (MCS) systems that outsource the collection of sensory data to the public crowd equipped with various mobile devices. A fundamental issue in such systems is to effectively incentivize worker participation. However, instead of being an isolated module, the incentive mechanism usually interacts with other components which may affect its performance, such as data aggregation component that aggregates workers' data and data perturbation component that protects workers' privacy. Therefore, different from past literature, we capture such interactive effect, and propose INCEPTION, a novel MCS system framework that integrates an incentive, a data aggregation, and a data perturbation mechanism. Specifically, its incentive mechanism selects workers who are more likely to provide reliable data, and compensates their costs for both sensing and privacy leakage. Its data aggregation mechanism also incorporates workers' reliability to generate highly accurate aggregated results, and its data perturbation mechanism ensures satisfactory protection for workers' privacy and desirable accuracy for the final perturbed results. We validate the desirable properties of INCEPTION through theoretical analysis, as well as extensive simulations. Haiming Jin, Lu Su 0001, Houping Xiao, Klara Nahrstedt |
MobiHoc | 2 |
| 2016 | Tackling the Redundancy and Sparsity in Crowd Sensing ApplicationsabstractDriven by the proliferation of sensor-rich mobile devices, crowd sensing has emerged as a new paradigm of gathering information about the physical world. In crowd sensing applications, user observations are usually unevenly distributed across the monitored entities, and this gives rise to two major challenges -- redundancy and sparsity. On one hand, multiple users may observe the same entity, and their observations are sometimes conflicting with each other due to the unreliable nature of human-carried sensors. On the other hand, crowd sensing data are usually very sparse, and there may exist considerable number of entities that never receive any observations from users. Some existing work studies these two challenges separately. However, we can gain great benefits by dealing with them jointly. In this paper, we develop an integrated framework to estimate the true values of entities from redundant and sparse data in crowd sensing applications. In this framework, we propose an effective algorithm to infer the "missing" observations for each entity, and aggregate both user-contributed and inferred observations to discover the true values of entities. We conduct extensive experiments on real-world crowd sensing systems to demonstrate the advantages of the proposed framework on correctly inferring entity truths from redundant and sparse data. Chuishi Meng, Houping Xiao, Lu Su 0001 |
SenSys | 3 |
| 2016 | Crowdsourcing High Quality Labels with a Tight BudgetabstractIn the past decade, commercial crowdsourcing platforms have revolutionized the ways of classifying and annotating data, especially for large datasets. Obtaining labels for a single instance can be inexpensive, but for large datasets, it is important to allocate budgets wisely. With limited budgets, requesters must trade-off between the quantity of labeled instances and the quality of the final results. Existing budget allocation methods can achieve good quantity but cannot guarantee high quality of individual instances under a tight budget. However, in some scenarios, requesters may be willing to label fewer instances but of higher quality. Moreover, they may have different requirements on quality for different tasks. To address these challenges, we propose a flexible budget allocation framework called Requallo. Requallo allows requesters to set their specific requirements on the labeling quality and maximizes the number of labeled instances that achieve the quality requirement under a tight budget. The budget allocation problem is modeled as a Markov decision process and a sequential labeling policy is produced. The proposed policy greedily searches for the instance to query next as the one that can provide the maximum reward for the goal. The Requallo framework is further extended to consider worker reliability so that the budget can be better allocated. Experiments on two real-world crowdsourcing tasks as well as a simulated task demonstrate that when the budget is tight, the proposed Requallo framework outperforms existing state-of-the-art budget allocation methods from both quantity and quality aspects. Qi Li 0012, Fenglong Ma, Jing Gao 0004, Lu Su 0001, Christopher J. Quinn |
WSDM | 4 |
| 2016 | Conflicts to Harmony: A Framework for Resolving Conflicts in Heterogeneous Data by Truth DiscoveryabstractIn many applications, one can obtain descriptions about the same objects or events from a variety of sources. As a result, this will inevitably lead to data or information conflicts. One important problem is to identify the true information (i.e., thetruths) among conflicting sources of data. It is intuitive to trust reliable sources more when deriving the truths, but it is usually unknown which one is more reliablea priori. Moreover, each source possesses a variety of properties with different data types. An accurate estimation of source reliability has to be made by modeling multiple properties in a unified model. Existing conflict resolution work either does not conduct source reliability estimation, or models multiple properties separately. In this paper, we propose to resolve conflicts among multiple sources of heterogeneous data types. We model the problem using an optimization framework where truths and source reliability are defined as two sets of unknown variables. The objective is to minimize the overall weighted deviation between the truths and the multi-source observations where each source is weighted by its reliability. Different loss functions can be incorporated into this framework to recognize the characteristics of various data types, and efficient computation approaches are developed. The proposed framework is further adapted to deal with streaming data in an incremental fashion and large-scale data in MapReduce model. Experiments on real-world weather, stock, and flight data as well as simulated multi-source data demonstrate the advantage of jointly modeling different data types in the proposed framework. Yaliang Li, Qi Li 0012, Jing Gao 0004, Lu Su 0001, Bo Zhao 0001, Wei Fan 0001, Jiawei Han 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2016 | Experiences with GreenGPS - Fuel-Efficient Navigation Using Participatory SensingabstractParticipatory sensing services based on mobile phones constitute an important growing area of mobile computing. Most services start small and hence are initially sparsely deployed. Unless a mobile service adds value while sparsely deployed, it may not survive conditions of sparse deployment. The paper offers a generic solution to this problem and illustrates this solution in the context ofGreenGPS; a navigation service that allows drivers to find the most fuel-efficient routes customized for their vehicles between arbitrary end-points. Specifically, when the participatory sensing service is sparsely deployed, we demonstrate a general framework for generalization from sparse collected data to produce models extending beyond the current data coverage. This generalization allows the mobile service to offer value under broader conditions. GreenGPS uses our developed participatory sensing infrastructure and generalization algorithms to perform inexpensive data collection, aggregation, and modeling in an end-to-end automated fashion. The models are subsequently used by our backend engine to predict customized fuel-efficient routes for both members and non-members of the service. GreenGPS is offered as a mobile phone application and can be easily deployed and used by individuals. A preliminary study of our green navigation idea was performed in[1], however, the effort was focused on a proof-of-concept implementation that involved substantial offline and manual processing. In contrast, the results and conclusions in the current paper are based on a more advanced and accurate model and extensive data from a real-world phone-based implementation and deployment, which enables reliable and automatic end-to-end data collection and route recommendation. The system further benefits from lower cost and easier deployment. To evaluate the green navigation service efficiency, we conducted a user subject study consisting of 22 users driving different vehicles over the course of several months in Urbana-Champaign, IL. The experimental results using the collected data suggest that fuel savings of 21.5 over the fastest, 11.2 percent over the shortest, and 8.4 percent over the Garmin eco routes can be achieved by following GreenGPS green routes. The study confirms that our navigation service can survive conditions of sparse deployment and at the same time achieve accurate fuel predictions and lead to significant fuel savings. Fatemeh Saremi, Omid Fatemieh, Hossein Ahmadi 0001, Tarek F. Abdelzaher, Raghu K. Ganti, Hengchang Liu, Shaohan Hu, Shen Li 0002, Lu Su 0001 |
IEEE Trans. Mob. Comput. | 10 |
| 2016 | Participatory Sensing Meets Opportunistic Sharing: Automatic Phone-to-Phone Communication in VehiclesabstractThis paper explores direct phone-to-phone communication (via WiFi interface) among vehicles to support participatory sensing applications. Sensing data usually contains location, speed, and fuel consumption of the car, and has a long time delay between collected and transferred to the server. Direct communication among phones aboard is important in reducing data transfer delay time and sharing participatory sensing information in an inexpensive manner. We design a practical and optimized communication mechanism for direct phone-to-phone data transfer among phones aboard that strategically enables phone-to-phone and/or phone-to-WiFiAP communications by optimally toggling the phones between the normal client and the hotspot modes. We take advantage of the WiFi hotspot functionality on smartphones, and hence require neither involvement of participants nor changes to existing wireless infrastructure and protocols. An analytical model is established to optimize toggling between client and hotspot modes for optimal system efficiency. We fully implement this system on off-the-shelf Google Galaxy Nexus and Nexus S phones. Through a 35-vehicle two-month deployment study, as well as simulation experiments using the real-world T-drive 9,211-taxicab dataset, we show that our solution significantly reduces data transfer delay time and maintains over 80 percent system efficiency under varying system parameters. Xiaoshan Sun, Shaohan Hu, Lu Su 0001, Tarek F. Abdelzaher, Pan Hui 0001, Wei Zheng 0011, Hengchang Liu, John A. Stankovic |
IEEE Trans. Mob. Comput. | 3 |
| 2015 | On Exploiting Logical Dependencies for Minimizing Additive Cost Metrics in Resource-Limited CrowdsensingabstractWe develop data retrieval algorithms for crowd-sensing applications that reduce the underlying network bandwidth consumption or any additive cost metric by exploiting logical dependencies among data items, while maintaining the level of service to the client applications. Crowd sensing applications refer to those where local measurements are performed by humans or devices in their possession for subsequent aggregation and sharing purposes. In this paper, we focus on resource-limited crowd sensing, such as disaster response and recovery scenarios. The key challenge in those scenarios is to cope with resource constraints. Unlike the traditional application design, where measurements are sent to a central aggregator, in resource limited scenarios, data will typically reside at the source until requested to prevent needless transmission. Many applications exhibit dependencies among data items. For example, parts of a city might tend to get flooded together because of a correlated low elevation, and some roads might become useless for evacuation if a bridge they lead to fails. Such dependencies can be encoded as logic expressions that obviate retrieval of some data items based on values of others. Our algorithm takes logical data dependencies into consideration such that application queries are answered at the central aggregation node, while network bandwidth usage is minimized. The algorithms consider multiple concurrent queries and accommodate retrieval latency constraints. Simulation results show that our algorithm outperforms several baselines by significant margins, maintaining the level of service perceived by applications in the presence of resource-constraints. Shaohan Hu, Shen Li 0002, Shuochao Yao, Lu Su 0001, Ramesh Govindan, Reginald L. Hobbs, Tarek F. Abdelzaher |
DCOSS | 4 |
| 2015 | Experiences with eNav: a low-power vehicular navigation systemabstractThis paper presents experiences with eNav, a smartphone-based vehicular GPS navigation system that has an energy-saving location sensing mode capable of drastically reducing navigation energy needs. Traditional navigation systems sample the phone's GPS at a fixed rate (usually around 1Hz), regardless of factors such as current vehicle speed and distance from the next navigation waypoint. This practice results in a large energy consumption and unnecessarily reduces the attainable length of a navigation session, if the phone is left unplugged. The paper investigates two questions. First, would drivers be willing to sacrifice some of the affordances of modern navigation systems in order to prolong battery life? Second, how much energy could be saved using straightforward alternative localization mechanisms, applied to complement GPS for vehicular navigation? According to a survey we conducted of 500 drivers, as much as 91% of drivers said they would like to have a vehicular navigation application with an energy saving mode. To meet this need, eNav exploits on-board accelerometers for approximate location sensing when the vehicle is sufficiently far from the next navigation waypoint (or is stopped). A user test-study of eNav shows that it results in roughly the same user experience as standard GPS navigation systems, while reducing navigation energy consumption by almost 80%. We conclude that drivers find an energy-saving mode on phone-based vehicular navigation applications desirable, even at the expense of some loss of functionality, and that significant savings can be achieved using straightforward location sensing mechanisms that avoid frequent GPS sampling. Shaohan Hu, Lu Su 0001, Shen Li 0002, Shiguang Wang, Chenji Pan, Siyu Gu, Md. Tanvir Al Amin, Hengchang Liu, Suman Nath, Romit Roy Choudhury, Tarek F. Abdelzaher |
UbiComp | 2 |
| 2015 | Scalable social sensing of interdependent phenomenaabstractThe proliferation of mobile sensing and communication devices in the possession of the average individual generated much recent interest in social sensing applications. Significant advances were made on the problem of uncovering ground truth from observations made by participants of unknown reliability. The problem, also called fact-finding commonly arises in applications where unvetted individuals may opt in to report phenomena of interest. For example, reliability of individuals might be unknown when they can join a participatory sensing campaign simply by downloading a smartphone app. This paper extends past social sensing literature by offering a scalable approach for exploiting dependencies between observed variables to increase fact-finding accuracy. Prior work assumed that reported facts are independent, or incurred exponential complexity when dependencies were present. In contrast, this paper presents the first scalable approach for accommodating dependency graphs between observed states. The approach is tested using real-life data collected in the aftermath of hurricane Sandy on availability of gas, food, and medical supplies, as well as extensive simulations. Evaluation shows that combining expected correlation graphs (of outages) with reported observations of unknown reliability, results in a much more reliable reconstruction of ground truth from the noisy social sensing data. We also show that correlation graphs can help test hypotheses regarding underlying causes, when different hypotheses are associated with different correlation patterns. For example, an observed outage profile can be attributed to a supplier outage or to excessive local demand. The two differ in expected correlations in observed outages, enabling joint identification of both the actual outages and their underlying causes. Shiguang Wang, Lu Su 0001, Shen Li 0002, Shaohan Hu, Md. Tanvir Al Amin, Shuochao Yao, Lance M. Kaplan, Tarek F. Abdelzaher |
IPSN | 2 |
| 2015 | On the Discovery of Evolving TruthabstractIn the era of big data, information regarding the same objects can be collected from increasingly more sources. Unfortunately, there usually exist conflicts among the information coming from different sources. To tackle this challenge, truth discovery, i.e., to integrate multi-source noisy information by estimating the reliability of each source, has emerged as a hot topic. In many real world applications, however, the information may come sequentially, and as a consequence, the truth of objects as well as the reliability of sources may be dynamically evolving. Existing truth discovery methods, unfortunately, cannot handle such scenarios. To address this problem, we investigate the temporal relations among both object truths and source reliability, and propose an incremental truth discovery framework that can dynamically update object truths and source weights upon the arrival of new data. Theoretical analysis is provided to show that the proposed method is guaranteed to converge at a fast rate. The experiments on three real world applications and a set of synthetic data demonstrate the advantages of the proposed method over state-of-the-art truth discovery methods. Yaliang Li, Qi Li 0012, Jing Gao 0004, Lu Su 0001, Bo Zhao 0001, Wei Fan 0001, Jiawei Han 0001 |
KDD | 4 |
| 2015 | FaitCrowd: Fine Grained Truth Discovery for Crowdsourced Data AggregationabstractIn crowdsourced data aggregation task, there exist conflicts in the answers provided by large numbers of sources on the same set of questions. The most important challenge for this task is to estimate source reliability and select answers that are provided by high-quality sources. Existing work solves this problem by simultaneously estimating sources' reliability and inferring questions' true answers (i.e., the truths). However, these methods assume that a source has the same reliability degree on all the questions, but ignore the fact that sources' reliability may vary significantly among different topics. To capture various expertise levels on different topics, we propose FaitCrowd, a fine grained truth discovery model for the task of aggregating conflicting data collected from multiple users/sources. FaitCrowd jointly models the process of generating question content and sources' provided answers in a probabilistic model to estimate both topical expertise and true answers simultaneously. This leads to a more precise estimation of source reliability. Therefore, FaitCrowd demonstrates better ability to obtain true answers for the questions compared with existing approaches. Experimental results on two real-world datasets show that FaitCrowd can significantly reduce the error rate of aggregation compared with the state-of-the-art multi-source aggregation approaches due to its ability of learning topical expertise from question content and collected answers. Fenglong Ma, Yaliang Li, Qi Li 0012, Minghui Qiu, Jing Gao 0004, Shi Zhi, Lu Su 0001, Bo Zhao 0001, Heng Ji 0001, Jiawei Han 0001 |
KDD | 7 |
| 2015 | Quality of Information Aware Incentive Mechanisms for Mobile Crowd Sensing SystemsabstractRecent years have witnessed the emergence of mobile crowd sensing (MCS) systems, which leverage the public crowd equipped with various mobile devices for large scale sensing tasks. In this paper, we study a critical problem in MCS systems, namely, incentivizing user participation. Different from existing work, we incorporate a crucial metric, called users' quality of information (QoI), into our incentive mechanisms for MCS systems. Due to various factors (e.g., sensor quality, noise, etc.) the quality of the sensory data contributed by individual users varies significantly. Obtaining high quality data with little expense is always the ideal of MCS platforms. Technically, we design incentive mechanisms based on reverse combinatorial auctions. We investigate both the single-minded and multi-minded combinatorial auction models. For the former, we design a truthful, individual rational and computationally efficient mechanism that approximately maximizes the social welfare with a guaranteed approximation ratio. For the latter, we design an iterative descending mechanism that achieves close-to-optimal social welfare while satisfying individual rationality and computational efficiency. Through extensive simulations, we validate our theoretical analysis about the close-to-optimal social welfare and fast running time of our mechanisms. Haiming Jin, Lu Su 0001, Klara Nahrstedt, Jinhui Xu 0001 |
MobiHoc | 2 |
| 2015 | Data Acquisition for Real-Time Decision-Making under Freshness ConstraintsabstractThe paper describes a novel algorithm for timely sensor data retrieval in resource-poor environments under freshness constraints. Consider a civil unrest, national security, or disaster management scenario, where a dynamic situation evolves and a decision-maker must decide on a course of action in view of latest data. Since the situation changes, so is the best course of action. The scenario offers two interesting constraints. First, one should be able to successfully compute the course of action within some appropriate time window, which we call the decision deadline. Second, at the time the course of action is computed, the data it is based on must be fresh (i.e., within some corresponding validity interval). We call it the freshness constraint. These constraints create an interesting novel problem of timely data retrieval. We address this problem in resource-scarce environments, where network resource limitations require that data objects (e.g., pictures and other sensor measurements pertinent to the decision) generally remain at the sources. Hence, one must decide on (i) which objects to retrieve and (ii) in what order, such that the cost of deciding on a valid course of action is minimized while meeting data freshness and decision deadline constraints. Such an algorithm is reported in this paper. The algorithm is shown in simulation to reduce the cost of data retrieval compared to a host of baselines that consider time or resource constraints. It is applied in the context of minimizing cost of finding unobstructed routes between specified locations in a disaster zone by retrieving data on the health of individual route segments. Shaohan Hu, Shuochao Yao, Haiming Jin, Yiran Zhao 0001, Yitao Hu, Nooreddin Naghibolhosseini, Shen Li 0002, Akash Kapoor, William Dron, Lu Su 0001, Amotz Bar-Noy, Pedro A. Szekely, Ramesh Govindan, Reginald L. Hobbs, Tarek F. Abdelzaher |
RTSS | 11 |
| 2015 | Truth Discovery on Crowd Sensing of Correlated EntitiesabstractWith the popular usage of mobile devices and smartphones, crowd sensing becomes pervasive in real life when human acts as sensors to report their observations about entities. For the same entity, users may report conflicting information, and thus it is important to identify the true information and the reliable users. This task, referred to as truth discovery, has recently attracted much attention. Existing work typically assumes independence among entities. However, correlations among entities are commonly observed in many applications. Such correlation information is crucial in the truth discovery task. When entities are not observed by enough reliable users, it is impossible to obtain true information. In such cases, it is important to propagate trustworthy information from correlated entities that have been observed by reliable users. We formulate the task of truth discovery on correlated entities as an optimization problem in which both truths and user reliability are modeled as variables. The correlation among entities adds to the difficulty of solving this problem. In light of the challenge, we propose both sequential and parallel solutions. In the sequential solution, we partition entities into disjoint independent sets and derive iterative approaches based on block coordinate descent. In the parallel solution, we adapt the solution to MapReduce programming model, which can be executed on Hadoop clusters. Experiments on real-world crowd sensing applications show the advantages of the proposed method on discovering truths from conflicting information reported on correlated entities. Chuishi Meng, Yaliang Li, Jing Gao 0004, Lu Su 0001, Hu Ding 0003 |
SenSys | 5 |
| 2015 | Cloud-Enabled Privacy-Preserving Truth Discovery in Crowd Sensing SystemsabstractThe recent proliferation of human-carried mobile devices has given rise to the crowd sensing systems. However, the sensory data provided by individual participants are usually not reliable. To identify truthful values from the crowd sensing data, the topic of truth discovery, whose goal is to estimate user quality and infer truths through quality-aware data aggregation, has drawn significant attention. Though able to improve aggregation accuracy, existing truth discovery approaches fail to take into consideration an important issue in their design, i.e., the protection of individual users' private information. In this paper, we propose a novel cloud-enabled privacy-preserving truth discovery (PPTD) framework for crowd sensing systems, which can achieve the protection of not only users' sensory data but also their reliability scores derived by the truth discovery approaches. The key idea of the proposed framework is to perform weighted aggregation on users' encrypted data using homomorphic cryptosystem. In order to deal with large-scale data, we also propose to parallelize PPTD with MapReduce framework. Through extensive experiments on not only synthetic data but also real world crowd sensing systems, we justify the guarantee of strong privacy and high accuracy of our proposed framework. Chenglin Miao, Lu Su 0001, Yaliang Li, Suxin Guo, Zhan Qin, Houping Xiao, Jing Gao 0004, Kui Ren 0001 |
SenSys | 3 |
| 2015 | SmartRoad: Smartphone-Based Crowd Sensing for Traffic Regulator Detection and IdentificationabstractIn this article we present SmartRoad, a crowd-sourced road sensing system that detects and identifies traffic regulators, traffic lights, and stop signs, in particular. As an alternative to expensive road surveys, SmartRoad works on participatory sensing data collected from GPS sensors from in-vehicle smartphones. The resulting traffic regulator information can be used for many assisted-driving or navigation systems. In order to achieve accurate detection and identification under realistic and practical settings, SmartRoad automatically adapts to different application requirements by (i) intelligently choosing the most appropriate information representation and transmission schemes, and (ii) dynamically evolving its core detection and identification engines to effectively take advantage of any external ground truth information or manual label opportunity. We implemented SmartRoad on a vehicular smartphone test bed, and deployed it on 35 external volunteer users’ vehicles for two months. Experiment results show that SmartRoad can robustly, effectively, and efficiently carry out the detection and identification tasks. Shaohan Hu, Lu Su 0001, Hengchang Liu, Tarek F. Abdelzaher |
ACM Trans. Sens. Networks | 2 |
| 2015 | Power-Based Diagnosis of Node Silence in Remote High-End Sensing SystemsabstractTroubleshooting unresponsive sensor nodes is a significant challenge in remote sensor network deployments. While prior work often targets low-end sensor networks, this article introduces a novel diagnostic tool, called the telediagnostic powertracer, geared for remote high-end sensing systems. Leveraging special properties of high-end systems, this in situ troubleshooting tool uses external power measurements to determine the internal health condition of an unresponsive node and the most likely cause of its failure. We develop our own low-cost power meter with low-bandwidth radio, propose both passive and active sampling schemes to measure the power consumption of the host node, and then report the measurements to a base station, hence allowing remote (i.e., tele-) diagnosis. The tool was deployed and tested in a remote solar-powered sensing system for acoustic and visual environmental monitoring. It was shown to successfully distinguish between several categories of failures that cause unresponsive behavior including energy depletion, antenna damage, radio disconnection, system crashes, and anomalous reboots. It was also able to determine the internal health conditions of an unresponsive node, such as the presence or absence of sensing and data storage activities (for each of multiple applications). The article explores the feasibility of building such a remote diagnostic tool from the standpoint of economy, scale, and diagnostic accuracy. The main novelty lies in its use of power consumption as a side channel, which has more availability than other I/O ports, to diagnose sensing system failures. Yong Yang 0009, Lu Su 0001, Mohammad Maifi Hasan Khan, Michael LeMay, Tarek F. Abdelzaher, Jiawei Han 0001 |
ACM Trans. Sens. Networks | 2 |
| 2014 | Data Extrapolation in Social Sensing for Disaster ResponseabstractThis paper complements the large body of social sensing literature by developing means for augmenting sensing data with inference results that "fill-in" missing pieces. Unlike trend-extrapolation methods, we focus on prediction in disaster scenarios where disruptive trend changes occur. A set of prediction heuristics (and a standard trend extrapolation algorithm) are compared that use either predominantly-spatial or predominantly-temporal correlations for data extrapolation purposes. The evaluation shows that none of them do well consistently. This is because monitored system state, in the aftermath of disasters, alternates between periods of relative calm and periods of disruptive change (e.g., aftershocks). A good prediction algorithm, therefore, needs to intelligently combine time-based data extrapolation during periods of calm, and spatial data extrapolation during periods of change. The paper develops such an algorithm. The algorithm is tested using data collected during the New York City crisis in the aftermath of Hurricane Sandy in November 2012. Results show that consistently good predictions are achieved. The work is unique in addressing the bi-modal nature of damage propagation in complex systems subjected to stress, and offers a simple solution to the problem. Siyu Gu, Chenji Pan, Hengchang Liu, Shen Li 0002, Shaohan Hu, Lu Su 0001, Shiguang Wang, Dong Wang 0002, Md. Tanvir Al Amin, Ramesh Govindan, Charu C. Aggarwal, Raghu K. Ganti, Mudhakar Srivatsa, Amotz Bar-Noy, Peter Terlecky, Tarek F. Abdelzaher |
DCOSS | 6 |
| 2014 | Centaur: Dynamic message dissemination over online social networksabstractWe present the design, implementation, and evaluation of Centaur, an application-level user-assisted message dissemination solution for Online Social Networks (OSN). Characteristics of OSNs make their message dissemination distinct from scenarios like multicast streaming and P2P file sharing. First, updates issued by each user are sporadic and the “online” follower set is highly dynamic. Hence, it is unnecessarily expensive to maintain always-alive multicast topologies. Second, the key advantage of OSNs over traditional media is realtime update, which would be greatly shadowed if it takes long to construct well-shaped dissemination structures. Therefore, in contrast to the multitude of prior multicast solutions, Centaur constructs location-aware dissemination trees locally for each incoming message. We implement a prototype with Cirrus and evaluate it with Twitter data. Experiment results show that Centaur achieves 98% delivery ratio and few seconds of delay with only around one tenth server traffic compared to centralized solutions used in many current OSNs. Shen Li 0002, Lu Su 0001, Yerzhan Suleimenov, Hengchang Liu, Tarek F. Abdelzaher, Guihai Chen |
ICCCN | 2 |
| 2014 | WOHA: Deadline-Aware Map-Reduce Workflow Scheduling Framework over Hadoop ClustersabstractIn this paper, we present WOHA, an efficient scheduling framework for deadline-aware Map-Reduce workflows. In data centers, complex backend data analysis often utilizes a workflow that contains tens or even hundreds of interdependent Map-Reduce jobs. Meeting deadlines of these workflows is usually of crucial importance to businesses (for example, workflows tightly linked to time-sensitive advertisement placement optimizations can directly affect revenue). Popular Map-Reduce implementations, such as Hadoop, deal with independent Map-Reduce jobs rather than workflows of jobs. In order to simplify the process of submitting workflows, solutions like Oozie emerge, which take a workflow configuration file as input and automatically submit its Hadoop jobs at the right time. The information separation that Hadoop only handles resource allocation and Oozie workflow topology, although preventing the Hadoop master node from getting involved with complex workflow analysis, may unnecessarily lengthen the workflow spans and thus cause more deadline misses. To address this problem and at the same time honor the efficiency of Hadoop master node, WOHA allows client nodes to locally generate scheduling plans which are later used as resource allocation hints by the master node. Under this framework design, we propose a novel scheduling algorithm that improves deadline satisfaction ratio by dynamically assigning priorities among workflows based on their progresses. We implement WOHA by extending Hadoop-1.2.1. Our experiments over an 80-server cluster show that WOHA manages to increase the deadline satisfaction ratio by 10% compared to state-of-the-art solutions, and scales up to tens of thousands of concurrently running workflows. Shen Li 0002, Shaohan Hu, Shiguang Wang, Lu Su 0001, Tarek F. Abdelzaher, Indranil Gupta, Richard Pace |
ICDCS | 4 |
| 2014 | Cost-Minimizing Mobile Access Point Deployment in Workflow-Based Mobile Sensor NetworksabstractIn mission-based mobile environments such as airplane maintenance, workflow-based mobile sensor networks emerge, where mobile users (MUs) with sensing devices visit sequences of mission-driven locations defined by workflows, and demand the gathering of sensory data within mission durations. To satisfy this demand in a cost-efficient manner, mobile access point (AP) deployment needs to be part of the overall solution. Therefore, we study the mobile AP deployment in workflow-based mobile sensor networks. We categorize MUs' workflows according to a priori knowledge of MUs' staying durations at mission locations into complete and incomplete information workflows. In both categories, we formulate the cost-minimizing mobile AP deployment problem into multiple (mixed) integer optimization problems, satisfying MUs' QoS constraints. We prove that the formulated optimization problems are NP-hard and design approximation algorithms with guaranteed approximation ratios. We demonstrate using simulations that the AP deployment cost calculated using our algorithms is 50-60% less than the stationary baseline approach and fairly close to the optimal AP deployment cost. In addition, the run times of our approximation algorithms are only 10-25% of those of the branch-and-bound algorithm used to derive the optimal AP deployment cost. Haiming Jin, Lu Su 0001, Klara Nahrstedt |
ICNP | 3 |
| 2014 | Towards automatic phone-to-phone communication for vehicular networking applicationsabstractThis paper explores direct phone-to-phone communication (via WiFi interface) among vehicles to support mobile sensing applications. Direct communication among drivers' phones is important in improving data collection efficiency and sharing participatory sensing information in an inexpensive manner. We design a practical and optimized communication mechanism for direct phone-to-phone data transfer among drivers' phones that strategically enables phone-to-phone and/or phone-to-WiFiAP communications by optimally toggles the phone between the normal client and the hotspot modes. We take advantage of the WiFi hotspot functionality on smartphones, and hence require neither involvement of participants nor changes to existing wireless infrastructure and protocols. An analytical model is established to optimize toggling between client and hotspot modes for optimal system efficiency. We fully implement this system on off-the-shelf Google Galaxy Nexus and Nexus S phones. Through a 35-vehicle 2-month deployment study, as well as simulation experiments using the real-world T-drive 9,211-taxicab dataset, we show that our solution significantly reduces data transfer delay time and maintains over 80% efficiency under varying system parameters. We even achieve 90% for parameter settings of the latest smartphones. Shaohan Hu, Hengchang Liu, Lu Su 0001, Tarek F. Abdelzaher, Pan Hui 0001, Wei Zheng 0011, Zhiheng Xie, John A. Stankovic |
INFOCOM | 3 |
| 2014 | Poster abstract: eNav: a smartphone-based energy efficient vehicular navigation system
Shaohan Hu, Lu Su 0001, Shen Li 0002, Shiguang Wang, Chenji Pan, Siyu Gu, Md. Tanvir Al Amin, Hengchang Liu, Suman Nath, Romit Roy Choudhury, Tarek F. Abdelzaher |
IPSN | 2 |
| 2014 | Robust confidentiality preserving data delivery in federated coalition networksabstractFederated coalition networks are formed by interconnected nodes belonging to different friendly-but-curious parties cooperating for common objectives. Each party has its policy regarding what information may be accessed by which other parties. Data delivery in coalition networks must provide both confidentiality and robustness. First, data should remain confidential when passing through intermediate nodes belonging to parties not authorized to see its content. Second, data delivery has to be robust against dynamic topology changes caused by frequent node churn and failures. We utilize the technique of linear network coding to transform the original data into multiple coded packets and send them along different paths in a way such that no other party can reconstruct the data. This lightweight approach provides confidentiality and robustness for friendly-but-curious coalitions with much less complexity than cryptography methods. In addition, we formulate an optimization problem to find minimum-cost paths, and use column generation framework to address the huge number of variables. Based on the proposed algorithms, we develop a Robust Confidentiality Preserving (R-CP) data delivery protocol. Our evaluation demonstrates that the proposed method can find the optimum solution in several seconds for networks of a few thousands nodes, and deliver data at a high success rate. Lu Su 0001, Fan Ye 0003, Peng Liu 0005, Oktay Günlük, Tom Bcrman, Seraphin B. Calo, Tarek F. Abdelzaher |
Networking | 1 |
| 2014 | Generalized Decision Aggregation in Distributed Sensing SystemsabstractIn this paper, we present GDA, a generalized decision aggregation framework that integrates information from distributed sensor nodes for decision making in a resource efficient manner. Traditional approaches that target similar problems only take as input the discrete label information from individual sensors that observe the same events. Different from them, our proposed GDA framework is able to take advantage of the confidence information of each sensor about its decision, and thus achieves higher decision accuracy. Targeting generalized problem domains, our framework can naturally handle the scenarios where different sensor nodes observe different sets of events whose numbers of possible classes may also be different. GDA also makes no assumption about the availability level of ground truth label information, while being able to take advantage of any if present. For these reasons, our approach can be applied to a much broader spectrum of sensing scenarios. The advantages of our proposed framework are demonstrated through both theoretic analysis and extensive experiments. Lu Su 0001, Qi Li 0012, Shaohan Hu, Shiguang Wang, Jing Gao 0004, Hengchang Liu, Tarek F. Abdelzaher, Jiawei Han 0001, Xue (Steve) Liu, Yan Gao 0010, Lance M. Kaplan |
RTSS | 1 |
| 2014 | Towards Cyber-Physical Systems in Social Spaces: The Data Reliability ChallengeabstractToday's cyber-physical systems (CPS) increasingly operate in social spaces. Examples include transportation systems, disaster response systems, and the smart grid, where humans are the drivers, survivors, or users. Much information about the evolving system can be collected from humans in the loop, a practice that is often called crowd-sensing. Crowd-sensing has not traditionally been considered a CPS topic, largely due to the difficulty in rigorously assessing its reliability. This paper aims to change that status quo by developing a mathematical approach for quantitatively assessing the probability of correctness of collected observations (about an evolving physical system), when the observations are reported by sources whose reliability is unknown. The paper extends prior literature on state estimation from noisy inputs, that often assumed unreliable sources that fall into one or a small number of categories, each with the same (possibly unknown) background noise distribution. In contrast, in the case of crowd-sensing, not only do we assume that the error distribution is unknown but also that each (human) sensor has its own possibly different error distribution. Given the above assumptions, we rigorously estimate data reliability in crowd-sensing systems, hence enabling their exploitation as state estimators in CPS feedback loops. We first consider applications where state is described by a number of binary variables, then extend the approach trivially to multivalued variables. The approach also extends prior work that addressed the problem in the special case of systems whose state does not change over time. Evaluation results, using both simulation and a real-life case-study, demonstrate the accuracy of the approach. Shiguang Wang, Dong Wang 0002, Lu Su 0001, Lance M. Kaplan, Tarek F. Abdelzaher |
RTSS | 3 |
| 2014 | A Confidence-Aware Approach for Truth Discovery on Long-Tail DataabstractIn many real world applications, the same item may be described by multiple sources. As a consequence, conflicts among these sources are inevitable, which leads to an important task: how to identify which piece of information is trustworthy, i.e., the truth discovery task. Intuitively, if the piece of information is from a reliable source, then it is more trustworthy, and the source that provides trustworthy information is more reliable. Based on this principle, truth discovery approaches have been proposed to infer source reliability degrees and the most trustworthy information (i.e., the truth) simultaneously. However, existing approaches overlook the ubiquitous long-tail phenomenon in the tasks, i.e., most sources only provide a few claims and only a few sources make plenty of claims, which causes the source reliability estimation for small sources to be unreasonable. To tackle this challenge, we propose a confidence-aware truth discovery (CATD) method to automatically detect truths from conflicting data with long-tail phenomenon. The proposed method not only estimates source reliability, but also considers the confidence interval of the estimation, so that it can effectively reflect real source reliability for sources with various levels of participation. Experiments on four real world tasks as well as simulated multi-source long-tail datasets demonstrate that the proposed method outperforms existing state-of-the-art truth discovery approaches by successful discounting the effect of small sources. Qi Li 0012, Yaliang Li, Jing Gao 0004, Lu Su 0001, Bo Zhao 0001, Murat Demirbas, Wei Fan 0001, Jiawei Han 0001 |
Proc. VLDB Endow. | 4 |
| 2013 | Poster abstract: SmartRoad: a crowd-sourced traffic regulator detection and identification systemabstractIn this paper we present SmartRoad, a crowd-sourced sensing system that detects and identifies traffic regulators, traffic lights and stop signs in particular. As an alternative to expensive road surveys, SmartRoad works on participatory sensing data collected from GPS sensors from in-vehicle smartphones. The resulting traffic regulator information can be used for many assisted-driving or navigation systems. We implement SmartRoad on a vehicular smartphone testbed, and deploy on 35 external volunteer users' vehicles for two months. Experiment results show that SmartRoad can robustly, effectively and efficiently carry out its detection and identification tasks without consuming excessive communication energy/bandwidth or requiring too much ground truth information. Shaohan Hu, Lu Su 0001, Hengchang Liu, Tarek F. Abdelzaher |
IPSN | 2 |
| 2013 | On the Detectability of Node Grouping in NetworksabstractIn typical studies of node grouping detection, the grouping is presumed to have a certain type of correlation with the network structure (e.g., densely connected groups of nodes that are loosely connected in between). People have defined different fitness measures (modularity, conductance, etc.) to quantify such correlation, and group the nodes by optimizing a certain fitness measure. However, a particular grouping with desired semantics, as the target of the detection, is not promised to be detectable by each measure. We study a fundamental problem in the process of node grouping discovery: Given a particular grouping in a network, whether and to what extent it can be discovered with a given fitness measure. We propose two approaches of testing the detectability, namely ranking-based and correlation-based randomization tests. Our methods are evaluated on both synthetic and real datasets, which shows the proposed methods can effectively predict the detectability of groupings of various types, and support explorative process of node grouping discovery. Jiawei Han 0001, Ming Ji, Lu Su 0001, Chi Wang 0001, Hongning Wang |
SDM | 5 |
| 2013 | Extrapolation from participatory sensing dataabstractIn this demo, a learning system, called Metis, is presented that extrapolates missing pieces in participatory sensing data. The work addresses the challenge of incomplete coverage in participatory sensing applications, where lack of complete control over participant mobility and sensing patterns may create coverage gaps in space and in time. Metis learns the underlying spatiotemporal patterns of the measured phenomenon from available incomplete observations, and uses these patterns to infer missing data. We describe the overall system design and demonstrate the system using data collected during the New York City gas crisis in the aftermath of Hurricane Sandy. Hengchang Liu, Siyu Gu, Chenji Pan, Wei Zheng 0011, Shen Li 0002, Shaohan Hu, Shiguang Wang, Dong Wang 0002, Md. Tanvir Al Amin, Lu Su 0001, Zhiheng Xie, Ramesh Govindan, Amotz Bar-Noy, Tarek F. Abdelzaher |
SenSys | 10 |
| 2012 | Quality of Information Based Data Selection and Transmission in Wireless Sensor NetworksabstractIn this paper, we provide a quality of information (QoI) based data selection and transmission service for classification missions in sensor networks. We first identify the two aspects of QoI, data reliability and data redundancy, and then propose metrics to estimate them. In particular, reliability implies the degree to which a sensor node contributes to the classification mission, and can be estimated through exploring the agreement between this node and the majority of others. On the other hand, redundancy represents the information overlap among different sensor nodes, and can be measured via investigating the similarity of their clustering results. Based on the proposed QoI metrics, we formulate an optimization problem that aims at maximizing the reliability of sensory data while eliminating their redundancies under the constraint of network resources. We decompose this problem into a data selection sub problem and a data transmission sub problem, and develop a distributed algorithm to solve them separately. The advantages of our schemes are demonstrated through the simulations on not only synthetic data but also a set of real audio records. Lu Su 0001, Shaohan Hu, Shen Li 0002, Jing Gao 0004, Tarek F. Abdelzaher, Jiawei Han 0001 |
RTSS | 1 |
| 2011 | Consensus extraction from heterogeneous detectors to improve performance over network traffic anomaly detectionabstractNetwork operators are continuously confronted with malicious events, such as port scans, denial-of-service attacks, and spreading of worms. Due to the detrimental effects caused by these anomalies, it is critical to detect them promptly and effectively. There have been numerous softwares, algorithms, or rules developed to conduct anomaly detection over traffic data. However, each of them only has limited descriptions of the anomalies, and thus suffers from high false positive/false negative rates. In contrast, the combination of multiple atomic detectors can provide a more powerful anomaly capturing capability when the base detectors complement each other. In this paper, we propose to infer a discriminative model by reaching consensus among multiple atomic anomaly detectors in an unsupervised manner when there are very few or even no known anomalous events for training. The proposed algorithm produces a perevent based non-trivial weighted combination of the atomic detectors by iteratively maximizing the probabilistic consensus among the output of the base detectors applied to different traffic records. The resulting model is different and not obtainable using Bayesian model averaging or weighted voting. Through experimental results on three network anomaly detection datasets, we show that the combined detector improves over the base detectors by 10% to 20% in accuracy. Jing Gao 0004, Wei Fan 0001, Deepak S. Turaga, Olivier Verscheure, Xiaoqiao Meng, Lu Su 0001, Jiawei Han 0001 |
INFOCOM | 6 |
| 2011 | Power watermarking: Facilitating power-based diagnosis of node silence in remote high-end sensing systems
Yong Yang 0009, Lu Su 0001, Mohammad Maifi Hasan Khan, Michael LeMay, Tarek F. Abdelzaher, Jiawei Han 0001 |
IPSN | 2 |
| 2011 | Towards optimal rate allocation for data aggregation in wireless sensor networksabstractThis paper aims at achieving optimal rate allocation for data aggregation in wireless sensor networks. We first formulate this rate allocation problem as a network utility maximization problem. Due to its non-convexity, we take a couple of variable substitutions on the original problem and transform it into an approximate problem, which is convex. We then apply duality theory to decompose this approximate problem into a rate control subproblem and a scheduling subproblem. Based on this decomposition, a distributed algorithm for joint rate control and scheduling is designed, and proved to approach arbitrarily close to the optimum of the approximate problem. Finally, we show that our approximate solution can achieve near-optimal performance through both theoretical analysis and simulations. Lu Su 0001, Yan Gao 0010, Yong Yang 0009, Guohong Cao |
MobiHoc | 1 |
| 2011 | Hierarchical aggregate classification with limited supervision for data reduction in wireless sensor networksabstractThe main challenge of designing classification algorithms for sensor networks is the lack of labeled sensory data, due to the high cost of manual labeling in the harsh locales where a sensor network is normally deployed. Moreover, delivering all the sensory data to the sink would cost enormous energy. Therefore, although some classification techniques can deal with limited label information, they cannot be directly applied to sensor networks since they are designed for centralized databases. To address these challenges, we propose a hierarchical aggregate classification (HAC) protocol which can reduce the amount of data sent by each node while achieving accurate classification in the face of insufficient label information. In this protocol, each sensor node locally makes cluster analysis and forwards only its decision to the parent node. The decisions are aggregated along the tree, and eventually the global agreement is achieved at the sink node. In addition, to control the tradeoff between the communication energy and the classification accuracy, we design an extended version of HAC, called the constrained hierarchical aggregate classification (cHAC) protocol. cHAC can achieve more accurate classification results compared with HAC, at the cost of more energy consumption. The advantages of our schemes are demonstrated through the experiments on not only synthetic data but also a real testbed. Lu Su 0001, Jing Gao 0004, Yong Yang 0009, Tarek F. Abdelzaher, Bolin Ding, Jiawei Han 0001 |
SenSys | 1 |
| 2010 | SolarCode: Utilizing Erasure Codes for Reliable Data Delivery in Solar-powered Wireless Sensor NetworksabstractSolar-powered sensor nodes have incentive to spend extra energy, especially when the battery is fully charged, because this energy surplus would be wasted otherwise. In this paper, we consider the problem of utilizing such energy surplus to adaptively adjust the redundancy level of erasure codes used in communication, so that the delivery reliability is improved while the network lifetime is still conserved. We formulate the problem as maximizing the end-to-end packet delivery probability under energy constraints. This formulated problem is hard to solve because of the combinatorics involved and the special curvature of its objective function. By exploiting its inherent properties, we propose an effective solution called SolarCode, which has a constant approximation ratio. We evaluate SolarCode in the context of our solar-powered sensor network testbed. Experiments show that SolarCode is successful in utilizing energy surplus and leads to higher data delivery reliability. Yong Yang 0009, Lu Su 0001, Yan Gao 0010, Tarek F. Abdelzaher |
INFOCOM | 2 |
| 2009 | oCast: Optimal Multicast Routing Protocol for Wireless Sensor NetworksabstractIn this paper, we describe oCast, an energy-optimal multicast routing protocol for wireless sensor networks. The general minimum-energy multicast problem is NP-hard. Intermittent connectivity that results from duty-cycling further complicates the problem. Nevertheless, we present both a centralized and distributed algorithm that are provably optimal when the number of destinations is small. This model is motivated by scenarios where sensors report to a small number of base stations or where data needs to be replicated on a small number of other nodes. We further propose an extended version of oCast, called Delay Bounded oCast (DB-oCast), which can discover optimal multicast trees under a predefined delay bound. Finally, we demonstrate the advantages of our schemes through both theoretical analysis and simulations. Lu Su 0001, Bolin Ding, Yong Yang 0009, Tarek F. Abdelzaher, Guohong Cao, Jennifer C. Hou |
ICNP | 1 |
| 2008 | Routing in intermittently connected sensor networksabstractTo prolong the lifetime of sensor networks, various scheduling schemes have been designed to reduce the number of active sensors. However, some scheduling strategies, such as partial coverage scheduling and target coverage scheduling, may result in disconnected network topologies, due to the low density of the active nodes. In such cases, traditional routing algorithms cannot be applied, and the shortest path discovered by these algorithms may not have the minimum packet delivery latency. In this paper, we address the problem of finding minimum latency routes in intermittently connected sensor networks by proposing an on-demand minimum latency (ODML) routing algorithm. Since on-demand routing algorithm does not work well when the source and destination frequently communicate with each other, we propose two proactive minimum latency routing algorithms: optimal-PML and quick-PML. Theoretical analysis and simulation results show that (1) ODML can effectively identify minimum latency routes which have much smaller latency than the shortest path, and (2) optimal-PML can minimize the routing message overhead and quick-PML can significantly reduce the route acquisition delay. Lu Su 0001, Changlei Liu, Guohong Cao |
ICNP | 1 |