Shen Li 0002

dblp:22/1835-2 · DBLP profile ↗
← Back
31ranked-venue papers
8as first author
0since 2021 · last 2017
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 7 first-authorComputer networks · 11 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
8 papers
Internet of things and sensor networks · 69% Internet architecture and protocols · 19% Vehicular, aerial and satellite networks · 9%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Performance modeling and evaluation · 34% Cloud and datacenter computing · 22% Storage systems · 15%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational social science and digital humanities · 50% Smart cities and intelligent transportation · 50%
Human-computer interaction and pervasive computing
2 papers
Ubiquitous computing and smart environments · 55% Interaction techniques and input · 22% Wearable and physiological sensing · 22%
Databases, data mining, and information retrieval
4 papers
Data stream processing · 45% Spatial and temporal data management · 39% Data integration and cleaning · 9%
Artificial intelligence
1 paper
Trustworthy machine learning · 100%

Topics — the 25 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Internet of things and sensor networks › mobile crowdsensing
participatory sensing
0.422015
Scalable social sensing of interdependent phenomena · IPSN 2015
Extrapolation from participatory sensing data · SenSys 2013
Smart cities and intelligent transportation › route planning
eco-routing
0.212016
Experiences with GreenGPS - Fuel-Efficient Navigation Using Participatory Sensing · IEEE Trans. Mob. Comput. 2016
Smart cities and intelligent transportation › route planning
route recommendation
0.212016
Experiences with GreenGPS - Fuel-Efficient Navigation Using Participatory Sensing · IEEE Trans. Mob. Comput. 2016
Computational social science and digital humanities
social media analysis
0.212016
Recursive Ground Truth Estimator for Social Data Streams · IPSN 2016
Computational social science and digital humanities › social computing
social sensing
0.212016
Recursive Ground Truth Estimator for Social Data Streams · IPSN 2016
Ubiquitous computing and smart environments › mobile crowdsourcing
participatory sensing
0.212016
Experiences with GreenGPS - Fuel-Efficient Navigation Using Participatory Sensing · IEEE Trans. Mob. Comput. 2016
Cloud and datacenter computing › datacenter operations
datacenter workload characterization
0.212016
An Experimental Evaluation of Datacenter Workloads On Low-Power Embedded Micro Servers · Proc. VLDB Endow. 2016
Performance modeling and evaluation › performance evaluation methodology
energy efficiency evaluation
0.212016
An Experimental Evaluation of Datacenter Workloads On Low-Power Embedded Micro Servers · Proc. VLDB Endow. 2016
Performance modeling and evaluation
workload characterization
0.212016
An Experimental Evaluation of Datacenter Workloads On Low-Power Embedded Micro Servers · Proc. VLDB Endow. 2016
Wearable and physiological sensing
energy-efficient sensing
0.212015
Experiences with eNav: a low-power vehicular navigation system · UbiComp 2015
Interaction techniques and input
mobile interaction
0.212015
Experiences with eNav: a low-power vehicular navigation system · UbiComp 2015
Ubiquitous computing and smart environments › automotive user interfaces
vehicle navigation system
0.212015
Experiences with eNav: a low-power vehicular navigation system · UbiComp 2015
Internet of things and sensor networks › mobile crowdsensing
truth discovery
0.212015
Scalable social sensing of interdependent phenomena · IPSN 2015
Distributed systems
distributed resource management
0.212015
Data Acquisition for Real-Time Decision-Making under Freshness Constraints · RTSS 2015
Storage systems
distributed storage
0.212015
Pyro: A Spatial-Temporal Big-Data Storage System · USENIX ATC 2015
Internet of things and sensor networks › wireless sensor network
data collection
0.212014
Poster abstract: information-maximizing data collection in social sensing using named-data · IPSN 2014
Internet architecture and protocols
information-centric networking
0.212014
Poster abstract: information-maximizing data collection in social sensing using named-data · IPSN 2014
Internet architecture and protocols › information-centric networking
named data networking
0.212014
Poster abstract: information-maximizing data collection in social sensing using named-data · IPSN 2014
Vehicular, aerial and satellite networks › intelligent transportation systems
vehicular navigation
0.212014
Poster abstract: eNav: a smartphone-based energy efficient vehicular navigation system · IPSN 2014
Internet of things and sensor networks
quality of information
0.112012
Quality of Information Based Data Selection and Transmission in Wireless Sensor Networks · RTSS 2012
Internet of things and sensor networks
wireless sensor network
0.112012
Quality of Information Based Data Selection and Transmission in Wireless Sensor Networks · RTSS 2012
Cloud and datacenter computing › datacenter architecture
datacenter server
0.112016
An Experimental Evaluation of Datacenter Workloads On Low-Power Embedded Micro Servers · Proc. VLDB Endow. 2016
Ubiquitous computing and smart environments
location sensing
0.112015
Experiences with eNav: a low-power vehicular navigation system · UbiComp 2015
Wireless sensing and localization › location-based services
smartphone-based navigation
0.112014
Poster abstract: eNav: a smartphone-based energy efficient vehicular navigation system · IPSN 2014
Data integration and cleaning › missing data
missing value imputation
0.012013
Extrapolation from participatory sensing data · SenSys 2013

Methods — techniques the papers use, named apart from their topics

recursive state estimation · 0.8binary signal decoding · 0.8participatory sensing · 0.5generalization algorithms · 0.5simulation · 0.4scheduling algorithm · 0.4estimation theory · 0.4parameter tuning · 0.2experimental benchmarking · 0.2dependency graph · 0.2correlation graph · 0.2accelerometer-based localization · 0.2smartphone sensing · 0.2information maximization · 0.2spatiotemporal learning · 0.2spatio-temporal learning · 0.2optimization · 0.1distributed algorithm · 0.1
YearPublicationVenuePosition
2017 On the improvement of classifying EEG recordings using neural networks
abstract
This paper presents improved results on classifying electroencephalography (EEG) recordings using deep learning. The task is to classify movements that the subject is thinking about (motor imagery), using only the recorded electrical activities on the scalp. The challenges are: poor signal-to-noise ratio; interference from numerous sources such as electrical line noise, muscle activity, and eye movements; considerable variability between individuals and even recording sessions. Traditional signal processing techniques such as frequency band analysis, common spatial pattern (CSP) algorithm or independent component analysis (ICA) fall short due to their limited capacity. Thanks to the rise of big data in healthcare, medical recordings now come in abundance. Therefore deep learning which relies on large amounts of training data is becoming the new cutting edge tool. We present a significant improvement of classification accuracy on the Brain-Computer Interfaces Competition IV dataset (2a), and compare the results of various state of the art neural network structures.
Yiran Zhao 0001, Shuochao Yao, Shaohan Hu, Shiyu Chang, Raghu K. Ganti, Mudhakar Srivatsa, Shen Li 0002, Tarek F. Abdelzaher
IEEE BigData7
2017 Demo: Unsupervised Fill-level Estimation for Smart Trash Removal Systems
Yiran Zhao 0001, Shuochao Yao, Shen Li 0002, Shaohan Hu, Huajie Shao, Tarek F. Abdelzaher
EWSN3
2017 Stark: Optimizing In-Memory Computing for Dynamic Dataset Collections
abstract
Emerging distributed in-memory computing frameworks, such as Apache Spark, can process a huge amount of cached data within seconds. This remarkably high efficiency requires the system to well balance data across tasks and ensure data locality. However, it is challenging to satisfy these requirements for applications that operate on a collection of dynamically loaded and evicted datasets. The dynamics may lead to time-varying data volume and distribution, which would frequently invoke expensive data re-partition and transfer operations, resulting in high overhead and large delay. To address this problem, we present Stark, a system specifically designed for optimizing in-memory computing on dynamic dataset collections. Stark enforces data locality for transformations spanning multiple datasets (e.g., join and cogroup) to avoid unnecessary data replications and shuffles. Moreover, to accommodate fluctuating data volume and skeweddata distribution, Stark delivers elasticity into partitions to balance task execution time andreduce job makespan. Finally, Stark achieves bounded failure recovery latency byoptimizing the data checkpointing strategy. Evaluations on a 50-server cluster show that Stark reduces the job makespan by 4X and improves system throughput by 6X compared to Spark.
Shen Li 0002, Md. Tanvir Al Amin, Raghu K. Ganti, Mudhakar Srivatsa, Shanhao Hu, Yiran Zhao 0001, Tarek F. Abdelzaher
ICDCS1
2017 Optimizing Source Selection in Social Sensing in the Presence of Influence Graphs
abstract
This paper addresses the problem of choosing the right sources to solicit data from in sensing applications involving broadcast channels, such as those crowdsensing applications where sources share their observations on social media. The goal is to select sources such that expected fusion error is minimized. We assume that soliciting data from a source incurs a cost and that the cost budget is limited. Contrary to other formulations of this problem, we focus on the case where some sources influence others. Hence, asking a source to make a claim affects the behavior of other sources as well, according to an influence model. The paper makes two contributions. First, we develop an analytic model for estimating expected fusion error, given a particular influence graph and solution to the source selection problem. Second, we use that model to search for a solution that minimizes expected fusion error, formulating it as a zero-one integer non-linear programming (INLP) problem. To scale the approach, the paper further proposes a novel reliability-based pruning heuristic (RPH) and a similarity-based lossy estimation (SLE) algorithm that significantly reduce the complexity of the INLP algorithm at the cost of a modest approximation. The analytically computed expected fusion error is validated using both simulations and real-world data from Twitter, demonstrating a good match between analytic predictions and empirical measurements. It is also shown that our method outperforms baselines in terms of resulting fusion error.
Huajie Shao, Shiguang Wang, Shen Li 0002, Shuochao Yao, Yiran Zhao 0001, Md. Tanvir Al Amin, Tarek F. Abdelzaher, Lance M. Kaplan
ICDCS3
2016 On Source Dependency Models for Reliable Social Sensing: Algorithms and Fundamental Error Bounds
abstract
This paper develops a simplified dependency model for sources on social networks that is shown to improve the quality of fact-finding -- assessing veracity of observations shared on social media. Recent literature developed a mathematical approach for exploiting social networks, such as Twitter, as noisy sensor networks that report observations on the state of the physical world. It was shown that the quality of state estimation from such noisy data, known as fact-finding, was a function of assumptions made regarding the independence of sources or lack thereof. When sources propagate information they hear from others (without verification), correlated errors may arise that degrade fact-finding performance. This work advances the state of the art by developing a simplified model of dependencies between sources and designing an improved dependency-aware estimator to assess veracity of observations, taking into account the observed dependency structure. A fundamental error bound is derived for this estimator to understand the gap in its performance from optimal. It is shown that the new estimator outperforms state of the art fact-finders and, in some cases, yields an accuracy close to the fundamental error bound.
Shuochao Yao, Shaohan Hu, Shen Li 0002, Yiran Zhao 0001, Lu Su 0001, Lance M. Kaplan, Aylin Yener, Tarek F. Abdelzaher
ICDCS3
2016 Recursive Ground Truth Estimator for Social Data Streams
abstract
The paper develops a recursive state estimator for social network data streams that allows exploitation of social networks, such as Twitter, as sensor networks to reliably observe physical events. Recent literature suggested using social networks as sensor networks leveraging the fact that much of the information upload on the former constitutes acts of sensing. A significant challenge identified in that context was that source reliability is often unknown, leading to uncertainty regarding the veracity of reported observations. Multiple truth finding systems were developed to solve this problem, generally geared towards batch analysis of offline datasets. This work complements the present batch approaches by developing an online recursive state estimator that recovers ground truth from streaming data. In this paper, we model physical world state by a set of binary signals (propositions, called assertions, about world state) and the social network as a noisy medium, where distortion, fabrication, omissions, and duplication are introduced. Our recursive state estimator is designed to recover the original binary signal (the true propositions) from the received noisy signal, essentially decoding the unreliable social network output to obtain the best estimate of ground truth in the physical world. Results show that the estimator is both effective and efficient at recovering the original signal with a high degree of accuracy. The estimator gives rise to a novel situation awareness tool that can be used for reliably following unfolding events in real time, using dynamically arriving social network data.
Shuochao Yao, Md. Tanvir Al Amin, Lu Su 0001, Shaohan Hu, Shen Li 0002, Shiguang Wang, Yiran Zhao 0001, Tarek F. Abdelzaher, Lance M. Kaplan, Charu C. Aggarwal, Aylin Yener
IPSN5
2016 An Experimental Evaluation of Datacenter Workloads On Low-Power Embedded Micro Servers
abstract
This paper presents a comprehensive evaluation of an ultra-low power cluster, built upon the Intel Edison based micro servers. The improved performance and high energy efficiency of micro servers have driven both academia and industry to explore the possibility of replacing conventional brawny servers with a larger swarm of embedded micro servers. Existing attempts mostly focus on mobile-class micro servers, whose capacities are similar to mobile phones. We, on the other hand, target on sensor-class micro servers, which are originally intended for uses in wearable technologies, sensor networks, and Internet-of-Things. Although sensor-class micro servers have much less capacity, they are touted for minimal power consumption (< 1 Watt), which opens new possibilities of achieving higher energy efficiency in datacenter workloads. Our systematic evaluation of the Edison cluster and comparisons to conventional brawny clusters involve careful workload choosing and laborious parameter tuning, which ensures maximum server utilization and thus fair comparisons. Results show that the Edison cluster achieves up to 3.5x improvement on work-done-per-joule for web service applications and data-intensive MapReduce jobs. In terms of scalability, the Edison cluster scales linearly on the throughput of web service workloads, and also shows satisfactory scalability for MapReduce workloads despite coordination overhead.
Yiran Zhao 0001, Shen Li 0002, Shaohan Hu, Shuochao Yao, Huajie Shao, Tarek F. Abdelzaher
Proc. VLDB Endow.2
2016 Experiences with GreenGPS - Fuel-Efficient Navigation Using Participatory Sensing
abstract
Participatory sensing services based on mobile phones constitute an important growing area of mobile computing. Most services start small and hence are initially sparsely deployed. Unless a mobile service adds value while sparsely deployed, it may not survive conditions of sparse deployment. The paper offers a generic solution to this problem and illustrates this solution in the context ofGreenGPS; a navigation service that allows drivers to find the most fuel-efficient routes customized for their vehicles between arbitrary end-points. Specifically, when the participatory sensing service is sparsely deployed, we demonstrate a general framework for generalization from sparse collected data to produce models extending beyond the current data coverage. This generalization allows the mobile service to offer value under broader conditions. GreenGPS uses our developed participatory sensing infrastructure and generalization algorithms to perform inexpensive data collection, aggregation, and modeling in an end-to-end automated fashion. The models are subsequently used by our backend engine to predict customized fuel-efficient routes for both members and non-members of the service. GreenGPS is offered as a mobile phone application and can be easily deployed and used by individuals. A preliminary study of our green navigation idea was performed in[1], however, the effort was focused on a proof-of-concept implementation that involved substantial offline and manual processing. In contrast, the results and conclusions in the current paper are based on a more advanced and accurate model and extensive data from a real-world phone-based implementation and deployment, which enables reliable and automatic end-to-end data collection and route recommendation. The system further benefits from lower cost and easier deployment. To evaluate the green navigation service efficiency, we conducted a user subject study consisting of 22 users driving different vehicles over the course of several months in Urbana-Champaign, IL. The experimental results using the collected data suggest that fuel savings of 21.5 over the fastest, 11.2 percent over the shortest, and 8.4 percent over the Garmin eco routes can be achieved by following GreenGPS green routes. The study confirms that our navigation service can survive conditions of sparse deployment and at the same time achieve accurate fuel predictions and lead to significant fuel savings.
Fatemeh Saremi, Omid Fatemieh, Hossein Ahmadi 0001, Tarek F. Abdelzaher, Raghu K. Ganti, Hengchang Liu, Shaohan Hu, Shen Li 0002, Lu Su 0001
IEEE Trans. Mob. Comput.9
2015 On Exploiting Logical Dependencies for Minimizing Additive Cost Metrics in Resource-Limited Crowdsensing
abstract
We develop data retrieval algorithms for crowd-sensing applications that reduce the underlying network bandwidth consumption or any additive cost metric by exploiting logical dependencies among data items, while maintaining the level of service to the client applications. Crowd sensing applications refer to those where local measurements are performed by humans or devices in their possession for subsequent aggregation and sharing purposes. In this paper, we focus on resource-limited crowd sensing, such as disaster response and recovery scenarios. The key challenge in those scenarios is to cope with resource constraints. Unlike the traditional application design, where measurements are sent to a central aggregator, in resource limited scenarios, data will typically reside at the source until requested to prevent needless transmission. Many applications exhibit dependencies among data items. For example, parts of a city might tend to get flooded together because of a correlated low elevation, and some roads might become useless for evacuation if a bridge they lead to fails. Such dependencies can be encoded as logic expressions that obviate retrieval of some data items based on values of others. Our algorithm takes logical data dependencies into consideration such that application queries are answered at the central aggregation node, while network bandwidth usage is minimized. The algorithms consider multiple concurrent queries and accommodate retrieval latency constraints. Simulation results show that our algorithm outperforms several baselines by significant margins, maintaining the level of service perceived by applications in the presence of resource-constraints.
Shaohan Hu, Shen Li 0002, Shuochao Yao, Lu Su 0001, Ramesh Govindan, Reginald L. Hobbs, Tarek F. Abdelzaher
DCOSS2
2015 Experiences with eNav: a low-power vehicular navigation system
abstract
This paper presents experiences with eNav, a smartphone-based vehicular GPS navigation system that has an energy-saving location sensing mode capable of drastically reducing navigation energy needs. Traditional navigation systems sample the phone's GPS at a fixed rate (usually around 1Hz), regardless of factors such as current vehicle speed and distance from the next navigation waypoint. This practice results in a large energy consumption and unnecessarily reduces the attainable length of a navigation session, if the phone is left unplugged. The paper investigates two questions. First, would drivers be willing to sacrifice some of the affordances of modern navigation systems in order to prolong battery life? Second, how much energy could be saved using straightforward alternative localization mechanisms, applied to complement GPS for vehicular navigation? According to a survey we conducted of 500 drivers, as much as 91% of drivers said they would like to have a vehicular navigation application with an energy saving mode. To meet this need, eNav exploits on-board accelerometers for approximate location sensing when the vehicle is sufficiently far from the next navigation waypoint (or is stopped). A user test-study of eNav shows that it results in roughly the same user experience as standard GPS navigation systems, while reducing navigation energy consumption by almost 80%. We conclude that drivers find an energy-saving mode on phone-based vehicular navigation applications desirable, even at the expense of some loss of functionality, and that significant savings can be achieved using straightforward location sensing mechanisms that avoid frequent GPS sampling.
Shaohan Hu, Lu Su 0001, Shen Li 0002, Shiguang Wang, Chenji Pan, Siyu Gu, Md. Tanvir Al Amin, Hengchang Liu, Suman Nath, Romit Roy Choudhury, Tarek F. Abdelzaher
UbiComp3
2015 Scalable social sensing of interdependent phenomena
abstract
The proliferation of mobile sensing and communication devices in the possession of the average individual generated much recent interest in social sensing applications. Significant advances were made on the problem of uncovering ground truth from observations made by participants of unknown reliability. The problem, also called fact-finding commonly arises in applications where unvetted individuals may opt in to report phenomena of interest. For example, reliability of individuals might be unknown when they can join a participatory sensing campaign simply by downloading a smartphone app. This paper extends past social sensing literature by offering a scalable approach for exploiting dependencies between observed variables to increase fact-finding accuracy. Prior work assumed that reported facts are independent, or incurred exponential complexity when dependencies were present. In contrast, this paper presents the first scalable approach for accommodating dependency graphs between observed states. The approach is tested using real-life data collected in the aftermath of hurricane Sandy on availability of gas, food, and medical supplies, as well as extensive simulations. Evaluation shows that combining expected correlation graphs (of outages) with reported observations of unknown reliability, results in a much more reliable reconstruction of ground truth from the noisy social sensing data. We also show that correlation graphs can help test hypotheses regarding underlying causes, when different hypotheses are associated with different correlation patterns. For example, an observed outage profile can be attributed to a supplier outage or to excessive local demand. The two differ in expected correlations in observed outages, enabling joint identification of both the actual outages and their underlying causes.
Shiguang Wang, Lu Su 0001, Shen Li 0002, Shaohan Hu, Md. Tanvir Al Amin, Shuochao Yao, Lance M. Kaplan, Tarek F. Abdelzaher
IPSN3
2015 The Packing Server for real-time scheduling of MapReduce workflows
abstract
This paper develops new schedulability bounds for a simplified MapReduce workflow model. MapReduce is a distributed computing paradigm, deployed in industry for over a decade. Different from conventional multiprocessor platforms, MapReduce deployments usually span thousands of machines, and a MapReduce job may contain as many as tens of thousands of parallel segments. State-of-the-art MapReduce workflow schedulers operate in a best-effort fashion, but the need for real-time operation has grown with the emergence of real-time analytic applications. MapReduce workflow details can be captured by the generalized parallel task model from recent real-time literature. Under this model, the best-known result guarantees schedulability if the task set utilization stays below 50% of total capacity, and the deadline to critical path length ratio, which we call the stretch φ, surpasses 2. This paper improves this bound further by introducing a hierarchical scheduling scheme based on the novel notion of a Packing Server, inspired by servers for aperiodic tasks. The Packing Server consists of multiple periodically replenished budgets that can execute in parallel and that appear as independent tasks to the underlying scheduler. Hence, the original problem of scheduling MapReduce workflows reduces to that of scheduling independent tasks. We prove that the utilization bound for schedulability of MapReduce workflows is UB· φ-β/φ , where UBis the utilization bound of the underlying independent task scheduling policy, and β is a tunable parameter that controls the maximum individual budget utilization. By leveraging past schedulability results for independent tasks on multiprocessors, we improve schedulable utilization of DAG workflows above 50% of total capacity, when the number of processors is large and the largest server budget is (sufficiently) smaller than its deadline. This surpasses the best known bounds for the generalized parallel task model. Our evaluation using a Yahoo! MapReduce trace as well as a physical cluster of 46 machines confirms the validity of the new utilization bound for MapReduce workflows.
Shen Li 0002, Shaohan Hu, Tarek F. Abdelzaher
RTAS1
2015 Data Acquisition for Real-Time Decision-Making under Freshness Constraints
abstract
The paper describes a novel algorithm for timely sensor data retrieval in resource-poor environments under freshness constraints. Consider a civil unrest, national security, or disaster management scenario, where a dynamic situation evolves and a decision-maker must decide on a course of action in view of latest data. Since the situation changes, so is the best course of action. The scenario offers two interesting constraints. First, one should be able to successfully compute the course of action within some appropriate time window, which we call the decision deadline. Second, at the time the course of action is computed, the data it is based on must be fresh (i.e., within some corresponding validity interval). We call it the freshness constraint. These constraints create an interesting novel problem of timely data retrieval. We address this problem in resource-scarce environments, where network resource limitations require that data objects (e.g., pictures and other sensor measurements pertinent to the decision) generally remain at the sources. Hence, one must decide on (i) which objects to retrieve and (ii) in what order, such that the cost of deciding on a valid course of action is minimized while meeting data freshness and decision deadline constraints. Such an algorithm is reported in this paper. The algorithm is shown in simulation to reduce the cost of data retrieval compared to a host of baselines that consider time or resource constraints. It is applied in the context of minimizing cost of finding unobstructed routes between specified locations in a disaster zone by retrieving data on the health of individual route segments.
Shaohan Hu, Shuochao Yao, Haiming Jin, Yiran Zhao 0001, Yitao Hu, Nooreddin Naghibolhosseini, Shen Li 0002, Akash Kapoor, William Dron, Lu Su 0001, Amotz Bar-Noy, Pedro A. Szekely, Ramesh Govindan, Reginald L. Hobbs, Tarek F. Abdelzaher
RTSS8
2015 Pyro: A Spatial-Temporal Big-Data Storage System
Shen Li 0002, Shaohan Hu, Raghu K. Ganti, Mudhakar Srivatsa, Tarek F. Abdelzaher
USENIX ATC1
2014 Data Extrapolation in Social Sensing for Disaster Response
abstract
This paper complements the large body of social sensing literature by developing means for augmenting sensing data with inference results that "fill-in" missing pieces. Unlike trend-extrapolation methods, we focus on prediction in disaster scenarios where disruptive trend changes occur. A set of prediction heuristics (and a standard trend extrapolation algorithm) are compared that use either predominantly-spatial or predominantly-temporal correlations for data extrapolation purposes. The evaluation shows that none of them do well consistently. This is because monitored system state, in the aftermath of disasters, alternates between periods of relative calm and periods of disruptive change (e.g., aftershocks). A good prediction algorithm, therefore, needs to intelligently combine time-based data extrapolation during periods of calm, and spatial data extrapolation during periods of change. The paper develops such an algorithm. The algorithm is tested using data collected during the New York City crisis in the aftermath of Hurricane Sandy in November 2012. Results show that consistently good predictions are achieved. The work is unique in addressing the bi-modal nature of damage propagation in complex systems subjected to stress, and offers a simple solution to the problem.
Siyu Gu, Chenji Pan, Hengchang Liu, Shen Li 0002, Shaohan Hu, Lu Su 0001, Shiguang Wang, Dong Wang 0002, Md. Tanvir Al Amin, Ramesh Govindan, Charu C. Aggarwal, Raghu K. Ganti, Mudhakar Srivatsa, Amotz Bar-Noy, Peter Terlecky, Tarek F. Abdelzaher
DCOSS4
2014 The Information Funnel: Exploiting Named Data for Information-Maximizing Data Collection
abstract
This paper describes the exploitation of hierarchical data names to achieve information-utility maximizing data collection in social sensing applications. We describe a novel transport abstraction, called the information funnel. It encapsulates a data collection protocol for social sensing that maximizes a measure of delivered information utility, that is the minimized data redundancy, by diversifying the data objects to be collected. The abstraction leverages named-data networking, a communication paradigm where data objects are named instead of hosts. We argue that this paradigm is especially suited for utility-maximizing transport in resource constrained environments, because hierarchical data names give rise to a notion of distance between named objects that is a function of only the topology of the name tree. This distance, in turn, can expose similarities between named objects that can be leveraged for minimizing redundancy among objects transmitted over bottlenecks, thereby maximizing their aggregate utility. With a proper hierarchical name space design, our protocol prioritizes transmission of data objects over bottlenecks to maximize information utility, with very weak assumptions on the utility function. This prioritization is achieved merely by comparing data name prefixes, without knowing application-level name semantics, which makes it generalizable across a wide range of applications. Evaluation results show that the information funnel improves the utility of the collected data objects compared to other lossy protocols.
Shiguang Wang, Tarek F. Abdelzaher, Santhosh Gajendran, Ajith Herga, Sachin Kulkarni, Shen Li 0002, Hengchang Liu, Chethan Suresh, Abhishek Sreenath, William Dron, Alice Leung, Ramesh Govindan, John P. Hancock
DCOSS6
2014 Centaur: Dynamic message dissemination over online social networks
abstract
We present the design, implementation, and evaluation of Centaur, an application-level user-assisted message dissemination solution for Online Social Networks (OSN). Characteristics of OSNs make their message dissemination distinct from scenarios like multicast streaming and P2P file sharing. First, updates issued by each user are sporadic and the “online” follower set is highly dynamic. Hence, it is unnecessarily expensive to maintain always-alive multicast topologies. Second, the key advantage of OSNs over traditional media is realtime update, which would be greatly shadowed if it takes long to construct well-shaped dissemination structures. Therefore, in contrast to the multitude of prior multicast solutions, Centaur constructs location-aware dissemination trees locally for each incoming message. We implement a prototype with Cirrus and evaluate it with Twitter data. Experiment results show that Centaur achieves 98% delivery ratio and few seconds of delay with only around one tenth server traffic compared to centralized solutions used in many current OSNs.
Shen Li 0002, Lu Su 0001, Yerzhan Suleimenov, Hengchang Liu, Tarek F. Abdelzaher, Guihai Chen
ICCCN1
2014 WOHA: Deadline-Aware Map-Reduce Workflow Scheduling Framework over Hadoop Clusters
abstract
In this paper, we present WOHA, an efficient scheduling framework for deadline-aware Map-Reduce workflows. In data centers, complex backend data analysis often utilizes a workflow that contains tens or even hundreds of interdependent Map-Reduce jobs. Meeting deadlines of these workflows is usually of crucial importance to businesses (for example, workflows tightly linked to time-sensitive advertisement placement optimizations can directly affect revenue). Popular Map-Reduce implementations, such as Hadoop, deal with independent Map-Reduce jobs rather than workflows of jobs. In order to simplify the process of submitting workflows, solutions like Oozie emerge, which take a workflow configuration file as input and automatically submit its Hadoop jobs at the right time. The information separation that Hadoop only handles resource allocation and Oozie workflow topology, although preventing the Hadoop master node from getting involved with complex workflow analysis, may unnecessarily lengthen the workflow spans and thus cause more deadline misses. To address this problem and at the same time honor the efficiency of Hadoop master node, WOHA allows client nodes to locally generate scheduling plans which are later used as resource allocation hints by the master node. Under this framework design, we propose a novel scheduling algorithm that improves deadline satisfaction ratio by dynamically assigning priorities among workflows based on their progresses. We implement WOHA by extending Hadoop-1.2.1. Our experiments over an 80-server cluster show that WOHA manages to increase the deadline satisfaction ratio by 10% compared to state-of-the-art solutions, and scales up to tens of thousands of concurrently running workflows.
Shen Li 0002, Shaohan Hu, Shiguang Wang, Lu Su 0001, Tarek F. Abdelzaher, Indranil Gupta, Richard Pace
ICDCS1
2014 Poster abstract: eNav: a smartphone-based energy efficient vehicular navigation system
Shaohan Hu, Lu Su 0001, Shen Li 0002, Shiguang Wang, Chenji Pan, Siyu Gu, Md. Tanvir Al Amin, Hengchang Liu, Suman Nath, Romit Roy Choudhury, Tarek F. Abdelzaher
IPSN3
2014 Poster abstract: information-maximizing data collection in social sensing using named-data
Shiguang Wang, Tarek F. Abdelzaher, Santhosh Gajendran, Ajith Herga, Sachin Kulkarni, Shen Li 0002, Hengchang Liu, Chethan Suresh, Abhishek Sreenath, William Dron, Alice Leung, Ramesh Govindan, John P. Hancock
IPSN6
2014 Using humans as sensors: an estimation-theoretic perspective
Dong Wang 0002, Md. Tanvir Al Amin, Shen Li 0002, Tarek F. Abdelzaher, Lance M. Kaplan, Siyu Gu, Chenji Pan, Hengchang Liu, Charu C. Aggarwal, Raghu K. Ganti, Xinlei (Oscar) Wang, Prasant Mohapatra, Boleslaw K. Szymanski, Hieu Khac Le
IPSN3
2014 Message Passing Algorithm for the Generalized Assignment Problem
Mindi Yuan, Shen Li 0002, Yannis Pavlidis
NPC3
2013 MINERVA: Information-Centric Programming for Social Sensing
abstract
In this paper, we introduce Minerva; an information-centric programming paradigm and toolkit for social sensing. The toolkit is geared for smartphone applications whose main objective is to collect and share information about the physical world. Information-centric programming refers to a publish-subscribe paradigm that maximizes the amount of information delivered. Unlike a traditional publish-subscribe system where publishers are assumed to have independent content, Minerva is geared for social sensing applications where different sources (participants sharing sensor data) often overlap in information they share. For example, through lack of coordination, they might collect redundant pictures of the same scene or redundant speed measurements of the same street. The main contribution of Minerva, therefore, lies in a data prioritization scheme that maximizes information delivery from publishers to subscribers by reducing redundancy, taking into account the non-independent nature of content. The algorithm is implemented on Android phones on top of the recently introduced named data networking framework. Evaluation results from both two smartphone-based experiments and a large-scale real data driven simulation demonstrate that the prioritization algorithm outperforms other candidates in terms of information coverage.
Shiguang Wang, Shaohan Hu, Shen Li 0002, Hengchang Liu, Md. Yusuf Sarwar Uddin, Tarek F. Abdelzaher
ICCCN3
2013 Proteus: Power Proportional Memory Cache Cluster in Data Centers
abstract
In this paper, we describe the design, implementation and evaluation of Proteus, a power-proportional cache cluster which eliminates the delay penalty during server provisioning dynamics. To speed up data center services, a cache cluster is used in front of the database tier, providing fast in-cache data access. Since the number of cache servers is large, building power-proportional cache clusters can lead to considerable monetary savings. Dynamic server provisioning, one common methodology for realizing power proportionality in data centers, calls for agile load balancing schemes and smart in-cache data migration algorithms when applied to cache clusters. Otherwise, it induces unacceptable delay spikes due to data re-allocation among cache servers. Proteus addresses both challenges by using a specifically designed virtual nodes placement algorithm and an amortized data migration policy. We implement Proteus, and evaluate it on a 40-server cluster using real Wikipedia data and workload traces. The results show that, with Proteus, the load distribution is much more evenly balanced compared to the case of applying unmodified consistent hashing. At the same time, Proteus induces almost no extra delay during provisioning transitions, which is a significant advantage over other state-of-the-art solutions.
Shen Li 0002, Shiguang Wang, Shaohan Hu, Fatemeh Saremi, Tarek F. Abdelzaher
ICDCS1
2013 Extrapolation from participatory sensing data
abstract
In this demo, a learning system, called Metis, is presented that extrapolates missing pieces in participatory sensing data. The work addresses the challenge of incomplete coverage in participatory sensing applications, where lack of complete control over participant mobility and sensing patterns may create coverage gaps in space and in time. Metis learns the underlying spatiotemporal patterns of the measured phenomenon from available incomplete observations, and uses these patterns to infer missing data. We describe the overall system design and demonstrate the system using data collected during the New York City gas crisis in the aftermath of Hurricane Sandy.
Hengchang Liu, Siyu Gu, Chenji Pan, Wei Zheng 0011, Shen Li 0002, Shaohan Hu, Shiguang Wang, Dong Wang 0002, Md. Tanvir Al Amin, Lu Su 0001, Zhiheng Xie, Ramesh Govindan, Amotz Bar-Noy, Tarek F. Abdelzaher
SenSys5
2012 Joint Optimization of Computing and Cooling Energy: Analytic Model and a Machine Room Case Study
abstract
Total energy minimization in data centers (including both computing and cooling energy) requires modeling the interactions between computing decisions (such as load distribution) and heat transfer in the room, since load acts as heat sources whose distribution in space affects cooling energy. This paper presents the first closed-form analytic optimal solution for load distribution in a machine rack that minimizes the sum of computing and cooling energy. We show that by considering actuation knobs on both computing and cooling sides, it is possible to reduce energy cost comparing to state of the art solutions that do not offer holistic energy optimization. The above can be achieved while meeting both throughput requirements and maximum CPU temperature constraints. Using a thorough evaluation on a real test bed of 20 machines, we demonstrate that our simple model adequately captures the thermal behavior and energy consumption of the system. We further show that our approach saves more energy compared to the state of the art in the field.
Shen Li 0002, Hieu Khac Le, Nam Pham, Jin Heo, Tarek F. Abdelzaher
ICDCS1
2012 Quality of Information Based Data Selection and Transmission in Wireless Sensor Networks
abstract
In this paper, we provide a quality of information (QoI) based data selection and transmission service for classification missions in sensor networks. We first identify the two aspects of QoI, data reliability and data redundancy, and then propose metrics to estimate them. In particular, reliability implies the degree to which a sensor node contributes to the classification mission, and can be estimated through exploring the agreement between this node and the majority of others. On the other hand, redundancy represents the information overlap among different sensor nodes, and can be measured via investigating the similarity of their clustering results. Based on the proposed QoI metrics, we formulate an optimization problem that aims at maximizing the reliability of sensory data while eliminating their redundancies under the constraint of network resources. We decompose this problem into a data selection sub problem and a data transmission sub problem, and develop a distributed algorithm to solve them separately. The advantages of our schemes are demonstrated through the simulations on not only synthetic data but also a set of real audio records.
Lu Su 0001, Shaohan Hu, Shen Li 0002, Jing Gao 0004, Tarek F. Abdelzaher, Jiawei Han 0001
RTSS3
2011 Aggregation-Friendly Data Collection Protocol of Wireless Sensor Networks for PoI Monitoring
abstract
Point of Interest (PoI) coverage problem has been discussed lately. Both practical strategies and theoretical analysis have been proposed. We however argue that coverage solution alone is not enough. In real applications, data from the covered area need to be collected and sent back to the sink. We notice that existing protocols are for general purpose use and not specially designed for PoI monitoring. Accordingly, we propose Smaller tree Collection Protocol (SCP) for data collection in this kind of applications. SCP is a distributed approximation for building a minimum Steiner tree connecting all PoI nodes and the sink. SCP aims to activate minimum relay nodes and supports data aggregation better than other collection protocol, e.g., Collection Tree Protocol (CTP), Backpressure Collection Protocol (BCP) and opportunistic routing protocols. We also design low cost distributed mechanisms to handle PoI dynamics. Extensive simulations have been performed to evaluate the proposed protocol.
Xiaobing Wu, Shen Li 0002, Guihai Chen
ICCCN2
2011 Understanding Vicious Cycles in Server Clusters
abstract
In this paper, we present an automated on-line service for troubleshooting performance problems in server clusters caused by unintended vicious cycles. The tool complements a large volume of prior performance troubleshooting and diagnostic literature for server farms that identifies problems arising due to resource bottlenecks or failed components. We show that unintended interactions between components in large-scale systems can cause performance problems even in the absence of bottlenecks or failures. Our tool leverages discriminative sequence mining to identify anomalous sequences of events that are candidates for blame for the performance problem. The tool looks for patterns consistent with "vicious cycles" or unstable behavior, as such patterns, when present, are most likely to be problematic. It highlights candidates that are semantically conflicting, such as those arising when different performance management mechanisms make adjustments in conflicting directions. Our approach offers two key advantages in performance troubleshooting. First, it does not require detailed prior knowledge of the underlying system to diagnose the problem. Second, contrary to simple statistical techniques, such as correlation analysis, that work well for continuous variables, our scheme can also identify chains of events (labels) that may explain the root cause of a problem. Our service is deployed on a web server testbed of 17 machines. To make the comparison of our scheme to prior work more concrete, we first reproduce two real-life problem scenarios reported in earlier literature, then explore a third, new case study. In all cases, our tool reports the patterns that explain the cause of the problem without requiring detailed a priori knowledge.
Mohammad Maifi Hasan Khan, Jin Heo, Shen Li 0002, Tarek F. Abdelzaher
ICDCS3
2009 ERN: Emergence Rescue Navigation with Wireless Sensor Networks
abstract
Navigation with wireless sensor networks (WSNs) can help people escape safely from an emergency. Previous navigation algorithms attempt to find safe and efficient escape paths for individuals under various environmental dynamics but ignore possible congestion caused by the individuals rushing for the exits. Moreover, all the previous works have overlooked the fact that the emergency rescue force can take actions strategically in order to save people out of danger. We propose ERN, Emergence Rescue Navigation algorithm by treating WSNs as navigation infrastructure. ERN takes both pedestrian congestion and rescue force flexibility into account. A directed graph is used to model the emergency regions. Human's movements are regarded as network flows on the graph. By calculating the maximum flow and minimum cut on the graph, the system can provide firemen rescue commands to eliminate key dangerous areas, which may significantly reduce congestion and save trapped people. We have performed extensive simulations under dynamic environments to evaluate the effectiveness and response time of ERN. Simulation results show that with ERN people in emergency are evacuated much faster and less congestion is observed.
Shen Li 0002, Andong Zhan, Xiaobing Wu, Guihai Chen
ICPADS1
2009 iTracking: Accurate Light-based Location-tracking in Wireless Sensor Networks
abstract
Most previous localization and tracking systems in wireless sensor networks are based on RF signals, ultrasounds, and UWB. However, systems using RF signals suffer accuracy fluctuations and systems using ultrasound or UWB need extra hardware. In our paper, we propose to utilize light, an easily accessible and pervasive resource in our daily life, and off-the-shelf TelosB Motes to achieve the goal of stability and high accuracy in localization and tracking. The iTracking system is a mobile light source location-tracking system based on light intensity. Our main contribution is the first demonstration to track mobile light sources based on light intensity with high accuracy, which may provide an alternative method for localization and tracking or inspire game designers. We examine point light based tracking and flashlight based tracking in our demonstrations, which proves these two main sources of light in our daily life can achieve centimeter-level location-tracking requirements. Our future work will focus on enabling the iTracking system to writing Chinese characters in flashlight.
Andong Zhan, Shen Li 0002, Lubin Guan, Panlong Yang, Xiaobing Wu, Guihai Chen
MASS2