VLDB 2026 Research / reviewers in the wild / expert
Shaohan Hu
dblp:33/8173
· DBLP profile ↗
56ranked-venue papers
9as first author
10since 2021 · last 2026
0000-0002-2877-2665ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 18 · 4 first-author · 4 since 2021Systems, architecture and hardware · 11 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Physical Self-Supervised Learning: IMU Sensing without Manual LabelsabstractDeep neural networks have become a promising approach for IMU-based sensing, but their scalability is fundamentally limited by costly labeled data and poor robustness to heterogeneous devices, placements, and users. Existing unsupervised and self-supervised methods reduce but do not remove this dependence, still requiring labeled data for domain adaptation and largely ignoring known physical structure. We propose physical self-supervised learning, an autoencoder-style paradigm for label-free IMU sensing. We replace the conventional neural decoder with an auto-adaptive physics decoder—a learnable family of kinematic equations that enforces explicit physical structure while adapting across environments—and adopt a hybrid two-stage IMU encoder with reconstruction in a structured latent space to mitigate sensor noise. Our framework further introduces probabilistic frequency-spatial constraints to disentangle sensor and object motion, a multi-view kinematic tree to exploit sparse physical self-supervised signals, and an uncertainty-aware formulation to handle the inherent ambiguity of IMU inference. Evaluated on inertial tracking and full-body motion capture over public datasets and realistic deployments, physical self-supervised learning reduces errors by up to 5× for tracking and 4× for motion capture in challenging generalization scenarios, consistently outperforming state-of-the-art supervised and self-supervised baselines without any labels. Yuyang Leng, Renyuan Liu, Shaohan Hu, Peijun Zhao, Chun-Fu Chen 0001, Songqing Chen, Shuochao Yao |
MobiSys | 3 |
| 2025 | PASS: Private Attributes Protection with Stochastic Data SubstitutionabstractThe growing Machine Learning (ML) services require extensive collections of user data, which may inadvertently include people’s private information irrelevant to the services. Various studies have been proposed to protect private attributes by removing them from the data while maintaining the utilities of the data for downstream tasks. Nevertheless, as we theoretically and empirically show in the paper, these methods reveal severe vulnerability because of a common weakness rooted in their adversarial training based strategies. To overcome this limitation, we propose a novel approach, PASS, designed to stochastically substitute the original sample with another one according to certain probabilities, which is trained with a novel loss function soundly derived from information-theoretic objective defined for utility-preserving private attributes protection. The comprehensive evaluation of PASS on various datasets of different modalities, including facial images, human activity sensory signals, and voice recording datasets, substantiates PASS’s effectiveness and generalizability. Yizhuo Chen, Chun-Fu Chen 0001, Hsiang Hsu, Shaohan Hu, Tarek F. Abdelzaher |
ICML | 4 |
| 2025 | DAF: An Efficient End-to-End Dynamic Activation Framework for on-Device DNN TrainingabstractRecent advancements in on-device training for deep neural networks have underscored the critical need for efficient activation compression to overcome the memory constraints of mobile and edge devices. As activations dominate memory usage during training and are essential for gradient computation, compressing them without compromising accuracy remains a key research challenge. While existing methods for dynamic activation quantization promise theoretical memory savings, their practical deployment is impeded by system-level challenges such as computational overhead and memory fragmentation. Renyuan Liu, Yuyang Leng, Kaiyan Liu, Shaohan Hu, Chun-Fu Chen 0001, Peijun Zhao, Heechul Yun, Shuochao Yao |
MobiSys | 4 |
| 2024 | Dropout-Based Rashomon Set Exploration for Efficient Predictive Multiplicity EstimationabstractPredictive multiplicity refers to the phenomenon in which classification tasks may admit multiple competing models that achieve almost-equally-optimal performance, yet generate conflicting outputs for individual samples.
This presents significant concerns, as it can potentially result in systemic exclusion, inexplicable discrimination, and unfairness in practical applications.
Measuring and mitigating predictive multiplicity, however, is computationally challenging due to the need to explore all such almost-equally-optimal models, known as the Rashomon set, in potentially huge hypothesis spaces.
To address this challenge, we propose a novel framework that utilizes dropout techniques for exploring models in the Rashomon set.
We provide rigorous theoretical derivations to connect the dropout parameters to properties of the Rashomon set, and empirically evaluate our framework through extensive experimentation.
Numerical results show that our technique consistently outperforms baselines in terms of the effectiveness of predictive multiplicity metric estimation, with runtime speedup up to $20\times \sim 5000\times$.
With efficient Rashomon set exploration and metric estimation, mitigation of predictive multiplicity is then achieved through dropout ensemble and model selection. Hsiang Hsu, Guihong Li, Shaohan Hu, Chun-Fu Chen 0001 |
ICLR | 3 |
| 2024 | MaSS: Multi-attribute Selective Suppression for Utility-preserving Data Transformation from an Information-theoretic PerspectiveabstractThe growing richness of large-scale datasets has been crucial in driving the rapid advancement and wide adoption of machine learning technologies. The massive collection and usage of data, however, pose an increasing risk for people’s private and sensitive information due to either inadvertent mishandling or malicious exploitation. Besides legislative solutions, many technical approaches have been proposed towards data privacy protection. However, they bear various limitations such as leading to degraded data availability and utility, or relying on heuristics and lacking solid theoretical bases. To overcome these limitations, we propose a formal information-theoretic definition for this utility-preserving privacy protection problem, and design a data-driven learnable data transformation framework that is capable of selectively suppressing sensitive attributes from target datasets while preserving the other useful attributes, regardless of whether or not they are known in advance or explicitly annotated for preservation. We provide rigorous theoretical analyses on the operational bounds for our framework, and carry out comprehensive experimental evaluations using datasets of a variety of modalities, including facial images, voice audio clips, and human activity motion sensor signals. Results demonstrate the effectiveness and generalizability of our method under various configurations on a multitude of tasks. Our source code is available at this URL. Yizhuo Chen, Chun-Fu Chen 0001, Hsiang Hsu, Shaohan Hu, Marco Pistoia, Tarek F. Abdelzaher |
ICML | 4 |
| 2024 | DynaSpa: Exploiting Spatial Sparsity for Efficient Dynamic DNN Inference on DevicesabstractRecent advancements in exploring machine learning models' dynamic spatial sparsity have demonstrated great potential for superior efficiency and adaptability without compromising accuracy when compared to conventional static-and-dense DNNs. However, realizing theoretical inference acceleration under practical deployment environments is still faced with significant system challenges. Current vendor libraries and tensor compilers fall short due to their extra data copy operations or insufficient computation schemes, especially for DNN operators with dynamic spatial sparsity. Renyuan Liu, Yuyang Leng, Shilei Tian, Shaohan Hu, Chun-Fu Chen 0001, Shuochao Yao |
SenSys | 4 |
| 2022 | Utility-Preserving Biometric Information Anonymization
Bill Moriarty, Chun-Fu Chen 0001, Shaohan Hu, Sean J. Moran, Marco Pistoia, Vincenzo Piuri, Pierangela Samarati |
ESORICS (2) | 3 |
| 2021 | Quantum Machine Learning for Finance ICCAD Special Session PaperabstractQuantum computers are expected to surpass the computational capabilities of classical computers during this decade, and achieve disruptive impact on numerous industry sectors, particularly finance. In fact, finance is estimated to be the first industry sector to benefit from Quantum Computing not only in the medium and long terms, but even in the short term. This review paper presents the state of the art of quantum algorithms for financial applications, with particular focus to those use cases that can be solved via Machine Learning. Marco Pistoia, Syed Farhan Ahmad, Akshay Ajagekar, Alexander Buts, Shouvanik Chakrabarti, Dylan Herman, Shaohan Hu, Andrew Jena, Pierre Minssen, Pradeep Niroula, Arthur G. Rattew, Romina Yalovetzky |
ICCAD | 7 |
| 2021 | Multi-labelled proteins recognition for high-throughput microscopy images using deep convolutional neural networksabstractBACKGROUND: Proteins are of extremely vital importance in the human body, and no movement or activity can be performed without proteins. Currently, microscopy imaging technologies developed rapidly are employed to observe proteins in various cells and tissues. In addition, due to the complex and crowded cellular environments as well as various types and sizes of proteins, a considerable number of protein images are generated every day and cannot be classified manually. Therefore, an automatic and accurate method should be designed to properly solve and analyse protein images with mixed patterns. RESULTS: In this paper, we first propose a novel customized architecture with adaptive concatenate pooling and "buffering" layers in the classifier part, which could make the networks more adaptive to training and testing datasets, and develop a novel hard sampler at the end of our network to effectively mine the samples from small classes. Furthermore, a new loss is presented to handle the label imbalance based on the effectiveness of samples. In addition, in our method, several novel and effective optimization strategies are adopted to solve the difficult training-time optimization problem and further increase the accuracy by post-processing. CONCLUSION: Our methods outperformed the SOTA method of multi-labelled protein classification on the HPA dataset, GapNet-PL, by above 2% in the F1 score. Therefore, experimental results based on the test set split from the Human Protein Atlas dataset show that our methods have good performance in automatically classifying multi-class and multi-labelled high-throughput microscopy protein images. Boheng Zhang, Shaohan Hu, Fa Zhang 0001, Zhiyong Liu 0002 |
BMC Bioinform. | 3 |
| 2021 | Driver Behavior-aware Parking Availability Crowdsensing System Using Truth DiscoveryabstractSpot-level parking availability information (the availability of each spot in a parking lot) is in great demand, as it can help reduce time and energy waste while searching for a parking spot. In this article, we propose a crowdsensing system called SpotE that can provide spot-level availability in a parking lot using drivers’ smartphone sensors. SpotE only requires the sensor data from drivers’ smartphones, which avoids the high cost of installing additional sensors and enables large-scale outdoor deployment. We propose a new model that can use the parking search trajectory and final destination (e.g., an exit of the parking lot) of a single driver in a parking lot to generate the probability profile that contains the probability of each spot being occupied in a parking lot. To deal with conflicting estimation results generated from different drivers, due to the variance in different drivers’ parking behaviors, a novel aggregation approach SpotE-TD is proposed. The proposed aggregation method is based on truth discovery techniques and can handle the variety in Quality of Information of different vehicles. We evaluate our proposed method through a real-life deployment study. Results show that SpotE-TD can efficiently provide spot-level parking availability information with a 20% higher accuracy than the state-of-the-art. Yi Zhu 0012, Shaohan Hu, Weida Zhong, Lu Su 0001, Chunming Qiao |
ACM Trans. Sens. Networks | 3 |
| 2020 | Parallelization of Classical Numerical optimization in Quantum Variational AlgorithmsabstractNumerical optimization has been extensively used in many real-world applications related to Scientific Computing, Artificial Intelligence and, more recently, Quantum Computing. However, existing optimizers conduct their internal computations sequentially, which affects their performance. We observed a general pattern that enabled us to parallelize such internal computations and achieve significant speedup. We designed a novel parallelization algorithm for optimizers, which consists of pattern detection, prediction, precomputation, and caching. Importantly, our design does not require any change to the optimizers. Instead, it simply modifies the function to be optimized, thereby leading to several engineering advantages, including simplicity, modularity and portability. We implemented this solution and included it in the Qiskit Aqua open-source project. In this paper, we present an evaluation on both standard benchmarks and real-world quantum-computing applications. The evaluation results confirm that our approach (1) incurs negligible overhead, (2) effectively speeds up optimization, and (3) does not affect the accuracy of the results or the convergence of the optimizers. Marco Pistoia, Peng Liu 0010, Chun-Fu Chen 0001, Shaohan Hu, Stephen P. Wood |
ICST | 4 |
| 2020 | Road Grade Estimation Using Crowd-Sourced Smartphone DataabstractEstimates of road grade/slope can add another dimension of information to existing 2D digital road maps. Integration of road grade information will widen the scope of digital map’s applications, which is primarily used for navigation, by enabling driving safety and efficiency applications such as Advanced Driver Assistance Systems (ADAS), eco-driving, etc. The huge scale and dynamic nature of road networks make sensing road grade a challenging task. Traditional methods oftentimes suffer from limited scalability and update frequency, as well as poor sensing accuracy. To overcome these problems, we propose a cost-effective and scalable road grade estimation framework using sensor data from smartphones. Based on our understanding of the error characteristics of smartphone sensors, we intelligently combine data from accelerometer, gyroscope and vehicle speed data from OBD-II/smartphone’s GPS to estimate road grade. To improve accuracy and robustness of the system, the estimations of road grade from multiple sources/vehicles are crowd-sourced to compensate for the effects of varying quality of sensor data from different sources. Extensive experimental evaluation on a test route of 9km demonstrates the superior performance of our proposed method, achieving 5× improvement on road grade estimation accuracy over baselines, with 90% of errors below 0.3°. Shaohan Hu, Weida Zhong, Adel W. Sadek, Lu Su 0001, Chunming Qiao |
IPSN | 2 |
| 2019 | Efficient Circuits for Quantum Search over 2D Square Lattice ArchitectureabstractQuantum computing has increasingly drawn interest and investments from the academic, industrial, and governmental research communities worldwide. Among quantum algorithms, Quantum Search is important for its quadratic speedup over its classical-computing counterpart. A key ingredient in its implementation is the Multi-Control Toffoli (MCT) gate, which creates a Boolean product of control variables and XORs it into the target. On an idealized quantum computer, all-to-all connectivity would eliminate the need to use SWAP gates to communicate information. This is, however, not affordable in the current Noisy Intermediate-Scale Quantum (NISQ) computing era. In this work, we discuss how to efficiently implement MCT gates on 2D Square Lattices (2DSL), suitable for superconducting circuits, by taking advantage of relative-phase Toffoli gates and H-tree layouts to drastically reduce resulting circuits' depths and the amount of SWAPping required. Shaohan Hu, Dmitri Maslov, Marco Pistoia, Jay M. Gambetta |
DAC | 1 |
| 2019 | Eugene: Towards Deep Intelligence as a ServiceabstractThe paper discusses an emerging suite of machine intelligence services that are of increasing importance in the highly instrumented world of the Internet of Things (IoT). The suite, called Eugene, would offer a form of intelligent behavior (based on deep neural networks) to otherwise simple embedded devices; the clients of the service. These devices would benefit from service resources to learn from data and to perform intelligent inference, classification, prediction, and estimation tasks that they are too limited to carry out on their own. The paper discusses the taxonomy of such services and the state of implementation, as well as the various challenges entailed, including scheduling, caching (of intelligent functions), and cooperative learning. Shuochao Yao, Kasthuri Jayarajah, Archan Misra, Tarek F. Abdelzaher, Yiran Zhao 0001, Ailing Piao, Huajie Shao, Dongxin Liu, Shengzhong Liu, Shaohan Hu, Dulanga Weerakoon |
ICDCS | 11 |
| 2019 | Classifying Mixed Patterns of Proteins in High-Throughput Microscopy Images Using Deep Neural Networks
Boheng Zhang, Shaohan Hu, Fa Zhang 0001 |
ICIC (1) | 3 |
| 2019 | Delving into Precise Attention in Image Captioning
Shaohan Hu, Shenglei Huang, Guolong Wang 0001, Zheng Qin 0003 |
ICONIP (5) | 1 |
| 2019 | Watch and Ask: Video Question Generation
Shenglei Huang, Shaohan Hu, Bencheng Yan |
ICONIP (3) | 2 |
| 2019 | An Expert Validation Framework for Improving the Quality of Crowdsourced Clustering
Liu Jiang, Zheng Qin 0003, Pengbo Shen, Shaohan Hu |
ICONIP (5) | 5 |
| 2019 | SADeepSense: Self-Attention Deep Learning Framework for Heterogeneous On-Device Sensors in Internet of Things ApplicationsabstractDeep neural networks are becoming increasingly popular in Internet of Things (IoT) applications. Their capabilities of fusing multiple sensor inputs and extracting temporal relationships can enhance intelligence in a wide range of applications. However, one key problem is the missing of adaptation to heterogeneous on-device sensors. These low-end sensors on IoT devices possess different accuracies, granularities, and amounts of information, whose sensing qualities are heterogeneous and vary over time. The existing deep learning frameworks for IoT applications usually treat every sensor input equally over time or increase model capacity in an ad-hoc manner, lacking the ability to identify and exploit the sensor heterogeneities. In this work, we propose SADeepSense, a deep learning framework that can automatically balance the contributions of multiple sensor inputs over time by exploiting their sensing qualities. SADeepSense makes two key contributions. First, SADeepSense employs the self-attention mechanism to learn the correlations among different sensors over time with no additional supervision. The correlations are then applied to infer the sensing qualities and to reassign model concentrations in multiple sensors over time. Second, instead of directly learning the sensing qualities and contributions, SADeepSense generates the residual concentrations that are deviated from the equal contributions, which helps to stabilize the training process. We demonstrate the effectiveness of SADeepSense with two representative IoT sensing tasks: heterogeneous human activity recognition with motion sensors and gesture recognition with the wireless signal. SADeepSense consistently outperforms the state-of-the-art methods by a clear margin. In addition, we show that SADeepSense only imposes little additional resource-consumption burden on embedded devices compared to the corresponding state-of-the-art framework. Shuochao Yao, Yiran Zhao 0001, Huajie Shao, Dongxin Liu, Shengzhong Liu, Ailing Piao, Shaohan Hu, Su Lu, Tarek F. Abdelzaher |
INFOCOM | 8 |
| 2019 | STFNets: Learning Sensing Signals from the Time-Frequency Perspective with Short-Time Fourier Neural NetworksabstractRecent advances in deep learning motivate the use of deep neural networks in Internet-of-Things (IoT) applications. These networks are modelled after signal processing in the human brain, thereby leading to significant advantages at perceptual tasks such as vision and speech recognition. IoT applications, however, often measure physical phenomena, where the underlying physics (such as inertia, wireless signal propagation, or the natural frequency of oscillation) are fundamentally a function of signal frequencies, offering better features in the frequency domain. This observation leads to a fundamental question: For IoT applications, can one develop a new brand of neural network structures that synthesize features inspired not only by the biology of human perception but also by the fundamental nature of physics? Hence, in this paper, instead of using conventional building blocks (e.g., convolutional and recurrent layers), we propose a new foundational neural network building block, the Short-Time Fourier Neural Network (STFNet). It integrates a widely-used time-frequency analysis method, the Short-Time Fourier Transform, into data processing to learn features directly in the frequency domain, where the physics of underlying phenomena leave better footprints. STFNets bring additional flexibility to time-frequency analysis by offering novel nonlinear learnable operations that are spectral-compatible. Moreover, STFNets show that transforming signals to a domain that is more connected to the underlying physics greatly simplifies the learning process. We demonstrate the effectiveness of STFNets with extensive experiments on a wide range of sensing inputs, including motion sensors, WiFi, ultrasound, and visible light. STFNets significantly outperform the state-of-the-art deep learning models in all experiments. A STFNet, therefore, demonstrates superior capability as the fundamental building block of deep neural networks for IoT applications for various sensor inputs 1. Shuochao Yao, Ailing Piao, Yiran Zhao 0001, Huajie Shao, Shengzhong Liu, Dongxin Liu, Jinyang Li 0004, Tianshi Wang 0002, Shaohan Hu, Lu Su 0001, Jiawei Han 0001, Tarek F. Abdelzaher |
WWW | 10 |
| 2018 | Towards Quality Aware Information Integration in Distributed Sensing SystemsabstractIn this paper, we present GDA, a generalized decision aggregation framework that integrates information from distributed sensor nodes for decision making in a resource efficient manner. Different from traditional approaches, our proposed GDA framework is able to not only estimate the reliability of each sensor, but also take advantage of its confidence information, and thus achieves higher decision accuracy. Targeting generalized problem domains, our framework can naturally handle the scenarios where different sensor nodes observe different sets of events whose numbers of possible classes may also be different. GDA also makes no assumption about the availability level of ground truth label information, while being able to take advantage of any if present. For these reasons, our approach can be applied to a much broader spectrum of sensing scenarios. In this paper, we also propose two extensions of the GDA framework, i.e., incremental GDA (I-GDA) and parallel GDA (P-GDA) to deal with streaming and large-scale data. The advantages of our proposed methods are demonstrated through both theoretic analysis and extensive experiments. Chenglin Miao, Lu Su 0001, Qi Li 0012, Shaohan Hu, Shiguang Wang, Jing Gao 0004, Hengchang Liu, Tarek F. Abdelzaher, Jiawei Han 0001, Xue (Steve) Liu, Yan Gao 0010, Lance M. Kaplan |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2017 | On the improvement of classifying EEG recordings using neural networksabstractThis paper presents improved results on classifying electroencephalography (EEG) recordings using deep learning. The task is to classify movements that the subject is thinking about (motor imagery), using only the recorded electrical activities on the scalp. The challenges are: poor signal-to-noise ratio; interference from numerous sources such as electrical line noise, muscle activity, and eye movements; considerable variability between individuals and even recording sessions. Traditional signal processing techniques such as frequency band analysis, common spatial pattern (CSP) algorithm or independent component analysis (ICA) fall short due to their limited capacity. Thanks to the rise of big data in healthcare, medical recordings now come in abundance. Therefore deep learning which relies on large amounts of training data is becoming the new cutting edge tool. We present a significant improvement of classification accuracy on the Brain-Computer Interfaces Competition IV dataset (2a), and compare the results of various state of the art neural network structures. Yiran Zhao 0001, Shuochao Yao, Shaohan Hu, Shiyu Chang, Raghu K. Ganti, Mudhakar Srivatsa, Shen Li 0002, Tarek F. Abdelzaher |
IEEE BigData | 3 |
| 2017 | Demo: Unsupervised Fill-level Estimation for Smart Trash Removal Systems
Yiran Zhao 0001, Shuochao Yao, Shen Li 0002, Shaohan Hu, Huajie Shao, Tarek F. Abdelzaher |
EWSN | 4 |
| 2017 | Decision-Driven Execution: A Distributed Resource Management Paradigm for the Age of IoTabstractThis paper introduces a novel paradigm for resource management in distributed systems, called decision-driven execution. The paradigm is appropriate for mission-driven systems, where the goal is to enable faster, leaner, and more effective decision making. All resource consumption, in this paradigm, is tied to the needs of making decisions on alternative courses of action. A point of departure from traditional architectures lies in interfaces that allow applications to specify their underlying decision logic. This specification, in turn, allows the system to reason about most effective means to meet information needs of decisions, resulting in simultaneous optimization of decision accuracy, cost, and speed. The paper discusses the overall vision of decision-driven execution, outlining preliminary work and novel challenges. Tarek F. Abdelzaher, Md. Tanvir Al Amin, Amotz Bar-Noy, William Dron, Ramesh Govindan, Reginald L. Hobbs, Shaohan Hu, Jung-Eun Kim, Jongdeog Lee, Kelvin Marcus, Shuochao Yao, Yiran Zhao 0001 |
ICDCS | 7 |
| 2017 | VehSense: Slippery Road Detection Using SmartphonesabstractThis paper investigates a new application of vehicular sensing: detecting and reporting the slippery road conditions. We describe a system and associated algorithm to monitor vehicle skidding events using smartphones and OBD-II (On board Diagnostics) adapters. This system, which we call the VehSense, gathers data from smartphone inertial sensors and vehicle wheel speed sensors, and processes the data to monitor slippery road conditions in real-time. Specifically, two speed readings are collected: 1) ground speed, which is estimated by vehicle acceleration and rotation, and 2) wheel speed, which is retrieved from the OBD-II interface. The mismatch between these two speeds is used to infer a skidding event. Without tapping into vehicle manufactures' proprietary data (e.g., antilock braking system), VehSense is compatible with most of the passenger vehicles, and thus can be easily deployed. We evaluate our system on snow-covered roads at Buffalo, and show that it can detect vehicle skidding effectively. Yunfei Hou, Tong Guan, Shaohan Hu, Lu Su 0001, Chunming Qiao |
VTC Spring | 4 |
| 2017 | DeepSense: A Unified Deep Learning Framework for Time-Series Mobile Sensing Data ProcessingabstractMobile sensing and computing applications usually require time-series inputs from sensors, such as accelerometers, gyroscopes, and magnetometers. Some applications, such as tracking, can use sensed acceleration and rate of rotation to calculate displacement based on physical system models. Other applications, such as activity recognition, extract manually designed features from sensor inputs for classification. Such applications face two challenges. On one hand, on-device sensor measurements are noisy. For many mobile applications, it is hard to find a distribution that exactly describes the noise in practice. Unfortunately, calculating target quantities based on physical system and noise models is only as accurate as the noise assumptions. Similarly, in classification applications, although manually designed features have proven to be effective, it is not always straightforward to find the most robust features to accommodate diverse sensor noise patterns and heterogeneous user behaviors. To this end, we propose DeepSense, a deep learning framework that directly addresses the aforementioned noise and feature customization challenges in a unified manner. DeepSense integrates convolutional and recurrent neural networks to exploit local interactions among similar mobile sensors, merge local interactions of different sensory modalities into global interactions, and extract temporal relationships to model signal dynamics. DeepSense thus provides a general signal estimation and classification framework that accommodates a wide range of applications. We demonstrate the effectiveness of DeepSense using three representative and challenging tasks: car tracking with motion sensors, heterogeneous human activity recognition, and user identification with biometric motion analysis. DeepSense significantly outperforms the state-of-the-art methods for all three tasks. In addition, we show that DeepSense is feasible to implement on smartphones and embedded devices thanks to its moderate energy consumption and low latency. Shuochao Yao, Shaohan Hu, Yiran Zhao 0001, Aston Zhang, Tarek F. Abdelzaher |
WWW | 2 |
| 2017 | On Exploiting Structured Human Interactions to Enhance Sensing Accuracy in Cyber-physical SystemsabstractIn this article, we describe a general methodology for enhancing sensing accuracy in cyber-physical systems that involve structured human interactions in noisy physical environment. We define structured human interactions as domain-specific workflow. A novel workflow-aware sensing model is proposed to jointly correct unreliable sensor data and keep track of states in a workflow. We also propose a new inference algorithm to handle cases with partially known states and objects as supervision. Our model is evaluated with extensive simulations. As a concrete application, we develop a novel log service called Emergency Transcriber , which can automatically document operational procedures followed by teams of first responders in emergency response scenarios. Evaluation shows that our system has significant improvement over commercial off-the-shelf (COTS) sensors and keeps track of workflow states with high accuracy in noisy physical environment. Shaohan Hu, Shiguang Wang, Renato Mancuso 0001, Minje Kim 0001, Po-Liang Wu, Lu Su 0001, Lui Sha, Tarek F. Abdelzaher |
ACM Trans. Cyber Phys. Syst. | 3 |
| 2017 | Efficient 3G/4G Budget Utilization in Mobile Sensing ApplicationsabstractThis paper explores efficient 3G/4G budget utilization in mobile sensing applications. Distinct from previous research work that either relied on limited WiFi access points or assumed the availability of unlimited 3G/4G communication capability, we offer a more practical mobile sensing system that leverages potential 3G/4G budgets that participants contribute at will, and uses it efficiently customized for the needs of multiple mobile sensing applications with heterogeneous sensitivity to environmental changes. We address the challenge that the information of data generation and WiFi encounters is not a priori knowledge, and propose an online decision making algorithm that takes advantage of participants' historical data. Three typical mobile sensing applications, vehicular application, mobile health and video sharing application are explored. Experimental results demonstrate that our proposed algorithms lead to significantly better system performance compared to alternative solutions for both applications. Shaohan Hu, Wei Zheng 0011, Tarek F. Abdelzaher, Pan Hui 0001, Zhiheng Xie, Hengchang Liu, John A. Stankovic |
IEEE Trans. Mob. Comput. | 2 |
| 2016 | On Source Dependency Models for Reliable Social Sensing: Algorithms and Fundamental Error BoundsabstractThis paper develops a simplified dependency model for sources on social networks that is shown to improve the quality of fact-finding -- assessing veracity of observations shared on social media. Recent literature developed a mathematical approach for exploiting social networks, such as Twitter, as noisy sensor networks that report observations on the state of the physical world. It was shown that the quality of state estimation from such noisy data, known as fact-finding, was a function of assumptions made regarding the independence of sources or lack thereof. When sources propagate information they hear from others (without verification), correlated errors may arise that degrade fact-finding performance. This work advances the state of the art by developing a simplified model of dependencies between sources and designing an improved dependency-aware estimator to assess veracity of observations, taking into account the observed dependency structure. A fundamental error bound is derived for this estimator to understand the gap in its performance from optimal. It is shown that the new estimator outperforms state of the art fact-finders and, in some cases, yields an accuracy close to the fundamental error bound. Shuochao Yao, Shaohan Hu, Shen Li 0002, Yiran Zhao 0001, Lu Su 0001, Lance M. Kaplan, Aylin Yener, Tarek F. Abdelzaher |
ICDCS | 2 |
| 2016 | Recursive Ground Truth Estimator for Social Data StreamsabstractThe paper develops a recursive state estimator for social network data streams that allows exploitation of social networks, such as Twitter, as sensor networks to reliably observe physical events. Recent literature suggested using social networks as sensor networks leveraging the fact that much of the information upload on the former constitutes acts of sensing. A significant challenge identified in that context was that source reliability is often unknown, leading to uncertainty regarding the veracity of reported observations. Multiple truth finding systems were developed to solve this problem, generally geared towards batch analysis of offline datasets. This work complements the present batch approaches by developing an online recursive state estimator that recovers ground truth from streaming data. In this paper, we model physical world state by a set of binary signals (propositions, called assertions, about world state) and the social network as a noisy medium, where distortion, fabrication, omissions, and duplication are introduced. Our recursive state estimator is designed to recover the original binary signal (the true propositions) from the received noisy signal, essentially decoding the unreliable social network output to obtain the best estimate of ground truth in the physical world. Results show that the estimator is both effective and efficient at recovering the original signal with a high degree of accuracy. The estimator gives rise to a novel situation awareness tool that can be used for reliably following unfolding events in real time, using dynamically arriving social network data. Shuochao Yao, Md. Tanvir Al Amin, Lu Su 0001, Shaohan Hu, Shen Li 0002, Shiguang Wang, Yiran Zhao 0001, Tarek F. Abdelzaher, Lance M. Kaplan, Charu C. Aggarwal, Aylin Yener |
IPSN | 4 |
| 2016 | An Experimental Evaluation of Datacenter Workloads On Low-Power Embedded Micro ServersabstractThis paper presents a comprehensive evaluation of an ultra-low power cluster, built upon the Intel Edison based micro servers. The improved performance and high energy efficiency of micro servers have driven both academia and industry to explore the possibility of replacing conventional brawny servers with a larger swarm of embedded micro servers. Existing attempts mostly focus on mobile-class micro servers, whose capacities are similar to mobile phones. We, on the other hand, target on sensor-class micro servers, which are originally intended for uses in wearable technologies, sensor networks, and Internet-of-Things. Although sensor-class micro servers have much less capacity, they are touted for minimal power consumption (< 1 Watt), which opens new possibilities of achieving higher energy efficiency in datacenter workloads. Our systematic evaluation of the Edison cluster and comparisons to conventional brawny clusters involve careful workload choosing and laborious parameter tuning, which ensures maximum server utilization and thus fair comparisons. Results show that the Edison cluster achieves up to 3.5x improvement on work-done-per-joule for web service applications and data-intensive MapReduce jobs. In terms of scalability, the Edison cluster scales linearly on the throughput of web service workloads, and also shows satisfactory scalability for MapReduce workloads despite coordination overhead. Yiran Zhao 0001, Shen Li 0002, Shaohan Hu, Shuochao Yao, Huajie Shao, Tarek F. Abdelzaher |
Proc. VLDB Endow. | 3 |
| 2016 | Experiences with GreenGPS - Fuel-Efficient Navigation Using Participatory SensingabstractParticipatory sensing services based on mobile phones constitute an important growing area of mobile computing. Most services start small and hence are initially sparsely deployed. Unless a mobile service adds value while sparsely deployed, it may not survive conditions of sparse deployment. The paper offers a generic solution to this problem and illustrates this solution in the context ofGreenGPS; a navigation service that allows drivers to find the most fuel-efficient routes customized for their vehicles between arbitrary end-points. Specifically, when the participatory sensing service is sparsely deployed, we demonstrate a general framework for generalization from sparse collected data to produce models extending beyond the current data coverage. This generalization allows the mobile service to offer value under broader conditions. GreenGPS uses our developed participatory sensing infrastructure and generalization algorithms to perform inexpensive data collection, aggregation, and modeling in an end-to-end automated fashion. The models are subsequently used by our backend engine to predict customized fuel-efficient routes for both members and non-members of the service. GreenGPS is offered as a mobile phone application and can be easily deployed and used by individuals. A preliminary study of our green navigation idea was performed in[1], however, the effort was focused on a proof-of-concept implementation that involved substantial offline and manual processing. In contrast, the results and conclusions in the current paper are based on a more advanced and accurate model and extensive data from a real-world phone-based implementation and deployment, which enables reliable and automatic end-to-end data collection and route recommendation. The system further benefits from lower cost and easier deployment. To evaluate the green navigation service efficiency, we conducted a user subject study consisting of 22 users driving different vehicles over the course of several months in Urbana-Champaign, IL. The experimental results using the collected data suggest that fuel savings of 21.5 over the fastest, 11.2 percent over the shortest, and 8.4 percent over the Garmin eco routes can be achieved by following GreenGPS green routes. The study confirms that our navigation service can survive conditions of sparse deployment and at the same time achieve accurate fuel predictions and lead to significant fuel savings. Fatemeh Saremi, Omid Fatemieh, Hossein Ahmadi 0001, Tarek F. Abdelzaher, Raghu K. Ganti, Hengchang Liu, Shaohan Hu, Shen Li 0002, Lu Su 0001 |
IEEE Trans. Mob. Comput. | 8 |
| 2016 | Participatory Sensing Meets Opportunistic Sharing: Automatic Phone-to-Phone Communication in VehiclesabstractThis paper explores direct phone-to-phone communication (via WiFi interface) among vehicles to support participatory sensing applications. Sensing data usually contains location, speed, and fuel consumption of the car, and has a long time delay between collected and transferred to the server. Direct communication among phones aboard is important in reducing data transfer delay time and sharing participatory sensing information in an inexpensive manner. We design a practical and optimized communication mechanism for direct phone-to-phone data transfer among phones aboard that strategically enables phone-to-phone and/or phone-to-WiFiAP communications by optimally toggling the phones between the normal client and the hotspot modes. We take advantage of the WiFi hotspot functionality on smartphones, and hence require neither involvement of participants nor changes to existing wireless infrastructure and protocols. An analytical model is established to optimize toggling between client and hotspot modes for optimal system efficiency. We fully implement this system on off-the-shelf Google Galaxy Nexus and Nexus S phones. Through a 35-vehicle two-month deployment study, as well as simulation experiments using the real-world T-drive 9,211-taxicab dataset, we show that our solution significantly reduces data transfer delay time and maintains over 80 percent system efficiency under varying system parameters. Xiaoshan Sun, Shaohan Hu, Lu Su 0001, Tarek F. Abdelzaher, Pan Hui 0001, Wei Zheng 0011, Hengchang Liu, John A. Stankovic |
IEEE Trans. Mob. Comput. | 2 |
| 2015 | On Exploiting Logical Dependencies for Minimizing Additive Cost Metrics in Resource-Limited CrowdsensingabstractWe develop data retrieval algorithms for crowd-sensing applications that reduce the underlying network bandwidth consumption or any additive cost metric by exploiting logical dependencies among data items, while maintaining the level of service to the client applications. Crowd sensing applications refer to those where local measurements are performed by humans or devices in their possession for subsequent aggregation and sharing purposes. In this paper, we focus on resource-limited crowd sensing, such as disaster response and recovery scenarios. The key challenge in those scenarios is to cope with resource constraints. Unlike the traditional application design, where measurements are sent to a central aggregator, in resource limited scenarios, data will typically reside at the source until requested to prevent needless transmission. Many applications exhibit dependencies among data items. For example, parts of a city might tend to get flooded together because of a correlated low elevation, and some roads might become useless for evacuation if a bridge they lead to fails. Such dependencies can be encoded as logic expressions that obviate retrieval of some data items based on values of others. Our algorithm takes logical data dependencies into consideration such that application queries are answered at the central aggregation node, while network bandwidth usage is minimized. The algorithms consider multiple concurrent queries and accommodate retrieval latency constraints. Simulation results show that our algorithm outperforms several baselines by significant margins, maintaining the level of service perceived by applications in the presence of resource-constraints. Shaohan Hu, Shen Li 0002, Shuochao Yao, Lu Su 0001, Ramesh Govindan, Reginald L. Hobbs, Tarek F. Abdelzaher |
DCOSS | 1 |
| 2015 | Experiences with eNav: a low-power vehicular navigation systemabstractThis paper presents experiences with eNav, a smartphone-based vehicular GPS navigation system that has an energy-saving location sensing mode capable of drastically reducing navigation energy needs. Traditional navigation systems sample the phone's GPS at a fixed rate (usually around 1Hz), regardless of factors such as current vehicle speed and distance from the next navigation waypoint. This practice results in a large energy consumption and unnecessarily reduces the attainable length of a navigation session, if the phone is left unplugged. The paper investigates two questions. First, would drivers be willing to sacrifice some of the affordances of modern navigation systems in order to prolong battery life? Second, how much energy could be saved using straightforward alternative localization mechanisms, applied to complement GPS for vehicular navigation? According to a survey we conducted of 500 drivers, as much as 91% of drivers said they would like to have a vehicular navigation application with an energy saving mode. To meet this need, eNav exploits on-board accelerometers for approximate location sensing when the vehicle is sufficiently far from the next navigation waypoint (or is stopped). A user test-study of eNav shows that it results in roughly the same user experience as standard GPS navigation systems, while reducing navigation energy consumption by almost 80%. We conclude that drivers find an energy-saving mode on phone-based vehicular navigation applications desirable, even at the expense of some loss of functionality, and that significant savings can be achieved using straightforward location sensing mechanisms that avoid frequent GPS sampling. Shaohan Hu, Lu Su 0001, Shen Li 0002, Shiguang Wang, Chenji Pan, Siyu Gu, Md. Tanvir Al Amin, Hengchang Liu, Suman Nath, Romit Roy Choudhury, Tarek F. Abdelzaher |
UbiComp | 1 |
| 2015 | Scalable social sensing of interdependent phenomenaabstractThe proliferation of mobile sensing and communication devices in the possession of the average individual generated much recent interest in social sensing applications. Significant advances were made on the problem of uncovering ground truth from observations made by participants of unknown reliability. The problem, also called fact-finding commonly arises in applications where unvetted individuals may opt in to report phenomena of interest. For example, reliability of individuals might be unknown when they can join a participatory sensing campaign simply by downloading a smartphone app. This paper extends past social sensing literature by offering a scalable approach for exploiting dependencies between observed variables to increase fact-finding accuracy. Prior work assumed that reported facts are independent, or incurred exponential complexity when dependencies were present. In contrast, this paper presents the first scalable approach for accommodating dependency graphs between observed states. The approach is tested using real-life data collected in the aftermath of hurricane Sandy on availability of gas, food, and medical supplies, as well as extensive simulations. Evaluation shows that combining expected correlation graphs (of outages) with reported observations of unknown reliability, results in a much more reliable reconstruction of ground truth from the noisy social sensing data. We also show that correlation graphs can help test hypotheses regarding underlying causes, when different hypotheses are associated with different correlation patterns. For example, an observed outage profile can be attributed to a supplier outage or to excessive local demand. The two differ in expected correlations in observed outages, enabling joint identification of both the actual outages and their underlying causes. Shiguang Wang, Lu Su 0001, Shen Li 0002, Shaohan Hu, Md. Tanvir Al Amin, Shuochao Yao, Lance M. Kaplan, Tarek F. Abdelzaher |
IPSN | 4 |
| 2015 | The Packing Server for real-time scheduling of MapReduce workflowsabstractThis paper develops new schedulability bounds for a simplified MapReduce workflow model. MapReduce is a distributed computing paradigm, deployed in industry for over a decade. Different from conventional multiprocessor platforms, MapReduce deployments usually span thousands of machines, and a MapReduce job may contain as many as tens of thousands of parallel segments. State-of-the-art MapReduce workflow schedulers operate in a best-effort fashion, but the need for real-time operation has grown with the emergence of real-time analytic applications. MapReduce workflow details can be captured by the generalized parallel task model from recent real-time literature. Under this model, the best-known result guarantees schedulability if the task set utilization stays below 50% of total capacity, and the deadline to critical path length ratio, which we call the stretch φ, surpasses 2. This paper improves this bound further by introducing a hierarchical scheduling scheme based on the novel notion of a Packing Server, inspired by servers for aperiodic tasks. The Packing Server consists of multiple periodically replenished budgets that can execute in parallel and that appear as independent tasks to the underlying scheduler. Hence, the original problem of scheduling MapReduce workflows reduces to that of scheduling independent tasks. We prove that the utilization bound for schedulability of MapReduce workflows is UB· φ-β/φ , where UBis the utilization bound of the underlying independent task scheduling policy, and β is a tunable parameter that controls the maximum individual budget utilization. By leveraging past schedulability results for independent tasks on multiprocessors, we improve schedulable utilization of DAG workflows above 50% of total capacity, when the number of processors is large and the largest server budget is (sufficiently) smaller than its deadline. This surpasses the best known bounds for the generalized parallel task model. Our evaluation using a Yahoo! MapReduce trace as well as a physical cluster of 46 machines confirms the validity of the new utilization bound for MapReduce workflows. Shen Li 0002, Shaohan Hu, Tarek F. Abdelzaher |
RTAS | 2 |
| 2015 | Data Acquisition for Real-Time Decision-Making under Freshness ConstraintsabstractThe paper describes a novel algorithm for timely sensor data retrieval in resource-poor environments under freshness constraints. Consider a civil unrest, national security, or disaster management scenario, where a dynamic situation evolves and a decision-maker must decide on a course of action in view of latest data. Since the situation changes, so is the best course of action. The scenario offers two interesting constraints. First, one should be able to successfully compute the course of action within some appropriate time window, which we call the decision deadline. Second, at the time the course of action is computed, the data it is based on must be fresh (i.e., within some corresponding validity interval). We call it the freshness constraint. These constraints create an interesting novel problem of timely data retrieval. We address this problem in resource-scarce environments, where network resource limitations require that data objects (e.g., pictures and other sensor measurements pertinent to the decision) generally remain at the sources. Hence, one must decide on (i) which objects to retrieve and (ii) in what order, such that the cost of deciding on a valid course of action is minimized while meeting data freshness and decision deadline constraints. Such an algorithm is reported in this paper. The algorithm is shown in simulation to reduce the cost of data retrieval compared to a host of baselines that consider time or resource constraints. It is applied in the context of minimizing cost of finding unobstructed routes between specified locations in a disaster zone by retrieving data on the health of individual route segments. Shaohan Hu, Shuochao Yao, Haiming Jin, Yiran Zhao 0001, Yitao Hu, Nooreddin Naghibolhosseini, Shen Li 0002, Akash Kapoor, William Dron, Lu Su 0001, Amotz Bar-Noy, Pedro A. Szekely, Ramesh Govindan, Reginald L. Hobbs, Tarek F. Abdelzaher |
RTSS | 1 |
| 2015 | Pyro: A Spatial-Temporal Big-Data Storage System
Shen Li 0002, Shaohan Hu, Raghu K. Ganti, Mudhakar Srivatsa, Tarek F. Abdelzaher |
USENIX ATC | 2 |
| 2015 | Reliable social sensing with physical constraints: analytic bounds and performance evaluation
Dong Wang 0002, Tarek F. Abdelzaher, Lance M. Kaplan, Raghu K. Ganti, Shaohan Hu, Hengchang Liu |
Real Time Syst. | 5 |
| 2015 | SmartRoad: Smartphone-Based Crowd Sensing for Traffic Regulator Detection and IdentificationabstractIn this article we present SmartRoad, a crowd-sourced road sensing system that detects and identifies traffic regulators, traffic lights, and stop signs, in particular. As an alternative to expensive road surveys, SmartRoad works on participatory sensing data collected from GPS sensors from in-vehicle smartphones. The resulting traffic regulator information can be used for many assisted-driving or navigation systems. In order to achieve accurate detection and identification under realistic and practical settings, SmartRoad automatically adapts to different application requirements by (i) intelligently choosing the most appropriate information representation and transmission schemes, and (ii) dynamically evolving its core detection and identification engines to effectively take advantage of any external ground truth information or manual label opportunity. We implemented SmartRoad on a vehicular smartphone test bed, and deployed it on 35 external volunteer users’ vehicles for two months. Experiment results show that SmartRoad can robustly, effectively, and efficiently carry out the detection and identification tasks. Shaohan Hu, Lu Su 0001, Hengchang Liu, Tarek F. Abdelzaher |
ACM Trans. Sens. Networks | 1 |
| 2014 | Data Extrapolation in Social Sensing for Disaster ResponseabstractThis paper complements the large body of social sensing literature by developing means for augmenting sensing data with inference results that "fill-in" missing pieces. Unlike trend-extrapolation methods, we focus on prediction in disaster scenarios where disruptive trend changes occur. A set of prediction heuristics (and a standard trend extrapolation algorithm) are compared that use either predominantly-spatial or predominantly-temporal correlations for data extrapolation purposes. The evaluation shows that none of them do well consistently. This is because monitored system state, in the aftermath of disasters, alternates between periods of relative calm and periods of disruptive change (e.g., aftershocks). A good prediction algorithm, therefore, needs to intelligently combine time-based data extrapolation during periods of calm, and spatial data extrapolation during periods of change. The paper develops such an algorithm. The algorithm is tested using data collected during the New York City crisis in the aftermath of Hurricane Sandy in November 2012. Results show that consistently good predictions are achieved. The work is unique in addressing the bi-modal nature of damage propagation in complex systems subjected to stress, and offers a simple solution to the problem. Siyu Gu, Chenji Pan, Hengchang Liu, Shen Li 0002, Shaohan Hu, Lu Su 0001, Shiguang Wang, Dong Wang 0002, Md. Tanvir Al Amin, Ramesh Govindan, Charu C. Aggarwal, Raghu K. Ganti, Mudhakar Srivatsa, Amotz Bar-Noy, Peter Terlecky, Tarek F. Abdelzaher |
DCOSS | 5 |
| 2014 | WOHA: Deadline-Aware Map-Reduce Workflow Scheduling Framework over Hadoop ClustersabstractIn this paper, we present WOHA, an efficient scheduling framework for deadline-aware Map-Reduce workflows. In data centers, complex backend data analysis often utilizes a workflow that contains tens or even hundreds of interdependent Map-Reduce jobs. Meeting deadlines of these workflows is usually of crucial importance to businesses (for example, workflows tightly linked to time-sensitive advertisement placement optimizations can directly affect revenue). Popular Map-Reduce implementations, such as Hadoop, deal with independent Map-Reduce jobs rather than workflows of jobs. In order to simplify the process of submitting workflows, solutions like Oozie emerge, which take a workflow configuration file as input and automatically submit its Hadoop jobs at the right time. The information separation that Hadoop only handles resource allocation and Oozie workflow topology, although preventing the Hadoop master node from getting involved with complex workflow analysis, may unnecessarily lengthen the workflow spans and thus cause more deadline misses. To address this problem and at the same time honor the efficiency of Hadoop master node, WOHA allows client nodes to locally generate scheduling plans which are later used as resource allocation hints by the master node. Under this framework design, we propose a novel scheduling algorithm that improves deadline satisfaction ratio by dynamically assigning priorities among workflows based on their progresses. We implement WOHA by extending Hadoop-1.2.1. Our experiments over an 80-server cluster show that WOHA manages to increase the deadline satisfaction ratio by 10% compared to state-of-the-art solutions, and scales up to tens of thousands of concurrently running workflows. Shen Li 0002, Shaohan Hu, Shiguang Wang, Lu Su 0001, Tarek F. Abdelzaher, Indranil Gupta, Richard Pace |
ICDCS | 2 |
| 2014 | Towards automatic phone-to-phone communication for vehicular networking applicationsabstractThis paper explores direct phone-to-phone communication (via WiFi interface) among vehicles to support mobile sensing applications. Direct communication among drivers' phones is important in improving data collection efficiency and sharing participatory sensing information in an inexpensive manner. We design a practical and optimized communication mechanism for direct phone-to-phone data transfer among drivers' phones that strategically enables phone-to-phone and/or phone-to-WiFiAP communications by optimally toggles the phone between the normal client and the hotspot modes. We take advantage of the WiFi hotspot functionality on smartphones, and hence require neither involvement of participants nor changes to existing wireless infrastructure and protocols. An analytical model is established to optimize toggling between client and hotspot modes for optimal system efficiency. We fully implement this system on off-the-shelf Google Galaxy Nexus and Nexus S phones. Through a 35-vehicle 2-month deployment study, as well as simulation experiments using the real-world T-drive 9,211-taxicab dataset, we show that our solution significantly reduces data transfer delay time and maintains over 80% efficiency under varying system parameters. We even achieve 90% for parameter settings of the latest smartphones. Shaohan Hu, Hengchang Liu, Lu Su 0001, Tarek F. Abdelzaher, Pan Hui 0001, Wei Zheng 0011, Zhiheng Xie, John A. Stankovic |
INFOCOM | 1 |
| 2014 | Poster abstract: eNav: a smartphone-based energy efficient vehicular navigation system
Shaohan Hu, Lu Su 0001, Shen Li 0002, Shiguang Wang, Chenji Pan, Siyu Gu, Md. Tanvir Al Amin, Hengchang Liu, Suman Nath, Romit Roy Choudhury, Tarek F. Abdelzaher |
IPSN | 1 |
| 2014 | Generalized Decision Aggregation in Distributed Sensing SystemsabstractIn this paper, we present GDA, a generalized decision aggregation framework that integrates information from distributed sensor nodes for decision making in a resource efficient manner. Traditional approaches that target similar problems only take as input the discrete label information from individual sensors that observe the same events. Different from them, our proposed GDA framework is able to take advantage of the confidence information of each sensor about its decision, and thus achieves higher decision accuracy. Targeting generalized problem domains, our framework can naturally handle the scenarios where different sensor nodes observe different sets of events whose numbers of possible classes may also be different. GDA also makes no assumption about the availability level of ground truth label information, while being able to take advantage of any if present. For these reasons, our approach can be applied to a much broader spectrum of sensing scenarios. The advantages of our proposed framework are demonstrated through both theoretic analysis and extensive experiments. Lu Su 0001, Qi Li 0012, Shaohan Hu, Shiguang Wang, Jing Gao 0004, Hengchang Liu, Tarek F. Abdelzaher, Jiawei Han 0001, Xue (Steve) Liu, Yan Gao 0010, Lance M. Kaplan |
RTSS | 3 |
| 2014 | Community Similarity Networks
Nicholas D. Lane, Hong Lu 0006, Shaohan Hu, Tanzeem Choudhury, Andrew T. Campbell, Feng Zhao 0001 |
Pers. Ubiquitous Comput. | 4 |
| 2013 | MINERVA: Information-Centric Programming for Social SensingabstractIn this paper, we introduce Minerva; an information-centric programming paradigm and toolkit for social sensing. The toolkit is geared for smartphone applications whose main objective is to collect and share information about the physical world. Information-centric programming refers to a publish-subscribe paradigm that maximizes the amount of information delivered. Unlike a traditional publish-subscribe system where publishers are assumed to have independent content, Minerva is geared for social sensing applications where different sources (participants sharing sensor data) often overlap in information they share. For example, through lack of coordination, they might collect redundant pictures of the same scene or redundant speed measurements of the same street. The main contribution of Minerva, therefore, lies in a data prioritization scheme that maximizes information delivery from publishers to subscribers by reducing redundancy, taking into account the non-independent nature of content. The algorithm is implemented on Android phones on top of the recently introduced named data networking framework. Evaluation results from both two smartphone-based experiments and a large-scale real data driven simulation demonstrate that the prioritization algorithm outperforms other candidates in terms of information coverage. Shiguang Wang, Shaohan Hu, Shen Li 0002, Hengchang Liu, Md. Yusuf Sarwar Uddin, Tarek F. Abdelzaher |
ICCCN | 2 |
| 2013 | Proteus: Power Proportional Memory Cache Cluster in Data CentersabstractIn this paper, we describe the design, implementation and evaluation of Proteus, a power-proportional cache cluster which eliminates the delay penalty during server provisioning dynamics. To speed up data center services, a cache cluster is used in front of the database tier, providing fast in-cache data access. Since the number of cache servers is large, building power-proportional cache clusters can lead to considerable monetary savings. Dynamic server provisioning, one common methodology for realizing power proportionality in data centers, calls for agile load balancing schemes and smart in-cache data migration algorithms when applied to cache clusters. Otherwise, it induces unacceptable delay spikes due to data re-allocation among cache servers. Proteus addresses both challenges by using a specifically designed virtual nodes placement algorithm and an amortized data migration policy. We implement Proteus, and evaluate it on a 40-server cluster using real Wikipedia data and workload traces. The results show that, with Proteus, the load distribution is much more evenly balanced compared to the case of applying unmodified consistent hashing. At the same time, Proteus induces almost no extra delay during provisioning transitions, which is a significant advantage over other state-of-the-art solutions. Shen Li 0002, Shiguang Wang, Shaohan Hu, Fatemeh Saremi, Tarek F. Abdelzaher |
ICDCS | 4 |
| 2013 | Efficient 3G budget utilization in mobile participatory sensing applicationsabstractThis paper explores efficient 3G budget utilization in mobile participatory sensing applications. 1 Distinct from previous research work that either rely on limited WiFi access points or assume the availability of unlimited 3G communication capability, we offer a more practical participatory sensing system that leverages potential 3G budgets that participants contribute at will, and uses it efficiently customized for the needs of multiple participatory sensing applications with heterogeneous sensitivity to environmental changes. We address the challenge that the information of data generation and WiFi encounters is not a priori knowledge, and propose an online decision making algorithm that takes advantage of participants' historical data. We also develop a heuristic algorithm to consume less energy and reduce the storage overhead while maintaining efficient 3G budget utilization. Experimental results from a 30-participant deployment demonstrate that, even when the budget is as small as 2.5% of a popular data plan, these two algorithms achieve higher utility of uploaded data compared to the baseline solution, especially, they increase the utility of received data by 151.4% and 137.8% for those sensitive applications. Hengchang Liu, Shaohan Hu, Wei Zheng 0011, Zhiheng Xie, Shiguang Wang, Pan Hui 0001, Tarek F. Abdelzaher |
INFOCOM | 2 |
| 2013 | Poster abstract: SmartRoad: a crowd-sourced traffic regulator detection and identification systemabstractIn this paper we present SmartRoad, a crowd-sourced sensing system that detects and identifies traffic regulators, traffic lights and stop signs in particular. As an alternative to expensive road surveys, SmartRoad works on participatory sensing data collected from GPS sensors from in-vehicle smartphones. The resulting traffic regulator information can be used for many assisted-driving or navigation systems. We implement SmartRoad on a vehicular smartphone testbed, and deploy on 35 external volunteer users' vehicles for two months. Experiment results show that SmartRoad can robustly, effectively and efficiently carry out its detection and identification tasks without consuming excessive communication energy/bandwidth or requiring too much ground truth information. Shaohan Hu, Lu Su 0001, Hengchang Liu, Tarek F. Abdelzaher |
IPSN | 1 |
| 2013 | Exploitation of Physical Constraints for Reliable Social SensingabstractThis paper develops and evaluates algorithms for exploiting physical constraints to improve the reliability of social sensing. Social sensing refers to applications where a group of sources (e.g., individuals and their mobile devices) volunteer to collect observations about the physical world. A key challenge in social sensing is that the reliability of sources and their devices is generally unknown, which makes it non-trivial to assess the correctness of collected observations. To solve this problem, the paper adopts a cyber-physical approach, where assessment of correctness of individual observations is aided by knowledge of physical constraints on both sources and observed variables to compensate for the lack of information on source reliability. We cast the problem as one of maximum likelihood estimation. The goal is to jointly estimate both (i) the latent physical state of the observed environment, and (ii) the inferred reliability of individual sources such that they are maximally consistent with both provenance information (who claimed what) and physical constraints. We evaluate the new framework through a real-world social sensing application. The results demonstrate significant performance gains in estimation accuracy of both source reliability and observation correctness. Dong Wang 0002, Tarek F. Abdelzaher, Lance M. Kaplan, Raghu K. Ganti, Shaohan Hu, Hengchang Liu |
RTSS | 5 |
| 2013 | Extrapolation from participatory sensing dataabstractIn this demo, a learning system, called Metis, is presented that extrapolates missing pieces in participatory sensing data. The work addresses the challenge of incomplete coverage in participatory sensing applications, where lack of complete control over participant mobility and sensing patterns may create coverage gaps in space and in time. Metis learns the underlying spatiotemporal patterns of the measured phenomenon from available incomplete observations, and uses these patterns to infer missing data. We describe the overall system design and demonstrate the system using data collected during the New York City gas crisis in the aftermath of Hurricane Sandy. Hengchang Liu, Siyu Gu, Chenji Pan, Wei Zheng 0011, Shen Li 0002, Shaohan Hu, Shiguang Wang, Dong Wang 0002, Md. Tanvir Al Amin, Lu Su 0001, Zhiheng Xie, Ramesh Govindan, Amotz Bar-Noy, Tarek F. Abdelzaher |
SenSys | 6 |
| 2012 | Quality of Information Based Data Selection and Transmission in Wireless Sensor NetworksabstractIn this paper, we provide a quality of information (QoI) based data selection and transmission service for classification missions in sensor networks. We first identify the two aspects of QoI, data reliability and data redundancy, and then propose metrics to estimate them. In particular, reliability implies the degree to which a sensor node contributes to the classification mission, and can be estimated through exploring the agreement between this node and the majority of others. On the other hand, redundancy represents the information overlap among different sensor nodes, and can be measured via investigating the similarity of their clustering results. Based on the proposed QoI metrics, we formulate an optimization problem that aims at maximizing the reliability of sensory data while eliminating their redundancies under the constraint of network resources. We decompose this problem into a data selection sub problem and a data transmission sub problem, and develop a distributed algorithm to solve them separately. The advantages of our schemes are demonstrated through the simulations on not only synthetic data but also a set of real audio records. Lu Su 0001, Shaohan Hu, Shen Li 0002, Jing Gao 0004, Tarek F. Abdelzaher, Jiawei Han 0001 |
RTSS | 2 |
| 2011 | Enabling large-scale human activity inference on smartphones using community similarity networks (csn)abstractSensor-enabled smartphones are opening a new frontier in the development of mobile sensing applications. The recognition of human activities and context from sensor-data using classification models underpins these emerging applications. However, conventional approaches to training classifiers struggle to cope with the diverse user populations routinely found in large-scale popular mobile applications. Differences between users (e.g., age, sex, behavioral patterns, lifestyle) confuse classifiers, which assume everyone is the same. To address this, we propose Community Similarity Networks (CSN), which incorporates inter-person similarity measurements into the classifier training process. Under CSN every user has a unique classifier that is tuned to their own characteristics. CSN exploits crowd-sourced sensor-data to personalize classifiers with data contributed from other similar users. This process is guided by similarity networks that measure different dimensions of inter-person similarity. Our experiments show CSN outperforms existing approaches to classifier training under the presence of population diversity. Nicholas D. Lane, Hong Lu 0006, Shaohan Hu, Tanzeem Choudhury, Andrew T. Campbell, Feng Zhao 0001 |
UbiComp | 4 |
| 2010 | Resource-Bounded Information Extraction: Acquiring Missing Feature Values on Demand
Pallika H. Kanani, Andrew McCallum, Shaohan Hu |
PAKDD (1) | 3 |