EDBT 2026 Demo / reviewers in the wild / expert
Lakshminarayanan Subramanian
dblp:85/5401 · also Lakshmi Subramanian
· DBLP profile ↗
72ranked-venue papers
10as first author
17since 2021 · last 2026
0000-0001-8101-1243ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 26 · 7 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 6 since 2021Artificial intelligence and machine learning · 16 · 8 since 2021Databases, data management, data science and information retrieval · 11 · 4 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 since 2021Systems, architecture and hardware · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Security and privacy · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-Regulating Cars: Automating Traffic Control in Free Flow Road NetworksabstractFree-flow road networks, such as suburban highways, are increasingly experiencing traffic congestion due to growing commuter inflow and limited infrastructure. Traditional control mechanisms—traffic signals or local heuristics—are ineffective or infeasible in these high-speed, signal-free environments. We introduce self-regulating cars, a reinforcement learning-based traffic control protocol that dynamically modulates vehicle speeds to optimize throughput and prevent congestion, without requiring new physical infrastructure. Our approach integrates classical traffic flow theory, gap acceptance models, and microscopic simulation into a physics-informed RL framework. By abstracting roads into super-segments, the agent captures emergent flow dynamics and learns robust speed modulation policies from instantaneous traffic observations. Evaluated in the high-fidelity PTV Vissim simulator on a real-world highway network, our method improves total throughput by 5%, reduces average delay by 13%, and decreases total stops by 3% compared to the no-control setting. It also achieves smoother, congestion-resistant flow while generalizing across varied traffic patterns—demonstrating its potential for scalable, ML-driven traffic management. Ankit Bhardwaj 0001, Rohail Asim, Sachin Chauhan, Yasir Zaki, Lakshminarayanan Subramanian |
AAAI | 5 |
| 2025 | The Privacy Quagmire: Where Computer Scientists and Lawyers May DisagreeabstractPrivacy policies dictate how systems handle user data, yet engineers struggle to verify compliance because policies use intentionally vague legal language. Current automated analyzers extract data practices using NLP but fail when policies say things like "share data for legitimate purposes" - terms that have no computational definition. This mismatch between legal flexibility and formal verification creates a fundamental barrier to automated compliance checking. We identify four systematic challenges: vague terms, evolving terminology, exception patterns that appear contradictory, and external legal dependencies. We propose an approach that preserves this ambiguity, where we use LLMs to extract structured parameters and convert them to first-order logic while keeping vague conditions as explicit placeholders for human interpretation. Our system can extract hundreds of data practices and reveals hidden complexities in TikTok and Meta policies, though the resulting formulas remain too complex for SMT solvers. This demonstrates the promise and fundamental limits of formalizing the legal text. Yunwei Zhao, Varun Chandrasekaran, Thomas Wies, Lakshminarayanan Subramanian |
HotNets | 4 |
| 2025 | MAML: Towards a Faster Web in Developing RegionsabstractThe web experience in developing regions remains subpar, primarily due to the growing complexity of modern webpages and insufficient optimization by content providers. Users in these regions typically rely on low-end devices and limited bandwidth, which results in a poor user experience as they download and parse webpages bloated with excessive third-party CSS and JavaScript (JS). To address these challenges, we introduce the Mobile Application Markup Language (MAML), a flat layout-based web specification language that reduces computational and data transmission demands, while replacing the excessive bloat from JS with a new scripting language centered on essential (and popular) web functionalities. Last but not least, MAML is backward compatible as it can be transpiled to minimal HTML/JavaScript/CSS and thus work with legacy browsers. We benchmark MAML in terms of page load times and sizes, using a translator which can automatically port any webpage to MAML. When compared to the popular Google AMP, across 100 testing webpages, MAML offers webpage speedups by tens of seconds under challenging network conditions thanks to its significant size reductions. Next, we run a competition involving 25 university students porting 50 of the above webpages to MAML using a web-based editor we developed. This experiment verifies that, with little developer effort, MAML is quite effective in maintaining the visual and functional correctness of the originating webpages. Ayush Pandey 0002, Matteo Varvello, Syed Ishtiaque Ahmed, Shurui Zhou, Lakshminarayanan Subramanian, Yasir Zaki |
WWW | 5 |
| 2024 | NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative DataabstractManuel Tonneau, Pedro Quinta De Castro, Karim Lasri, Ibrahim Farouq, Lakshmi Subramanian, Victor Orozco-Olvera, Samuel Fraiberger. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Manuel Tonneau, Pedro Vitor Quinta de Castro, Karim Lasri, Ibrahim Farouq, Lakshminarayanan Subramanian, Víctor Orozco-Olvera, Samuel P. Fraiberger |
ACL (1) | 5 |
| 2024 | The GAIUS Experience: Powering a Hyperlocal Mobile Web for Communities in Emerging Regions
Rohail Asim, Arjuna Sathiaseelan, Arko Chatterjee, Mukund Lal, Yasir Zaki, Lakshminarayanan Subramanian |
ICTD | 6 |
| 2022 | Targeted Policy Recommendations using Outcome-aware ClusteringabstractPolicy recommendations using observational data typically rely on estimating an econometric model on a sample of observations drawn from an entire population. However, different policy actions could potentially be optimal for different subgroups of a population. In this paper, we propose outcome-aware clustering, a new methodology to segment a population into different clusters and derive cluster-level policy recommendations. Outcome-aware clustering differs from conventional clustering algorithms across two basic dimensions. First, given a specific outcome of interest, outcome-aware clustering segments the population based on selecting a small set of features that closely relate with the outcome variable. Second, the clustering algorithm aims to generate near-homogeneous clusters based on a combination of cluster size-balancing constraints, inter and intra-cluster distances in the reduced feature space. We generate targeted policy recommendations for each outcome-aware cluster based on a standard multivariate regression of a condensed set of actionable policy features (which may partially overlap or differ from the features used for segmentation) from the observational data. We implement our outcome-aware clustering method on the Living Standards Measurement Study - Integrated Surveys on Agriculture (LSMS-ISA) dataset to generate targeted policy recommendations for improving farmers outcomes in sub-Saharan Africa. Based on a detailed analysis of the LSMS-ISA, we derive outcome-aware clusters of farmer populations across three sub-Saharan African countries and show that the targeted policy recommendations at the cluster level significantly differ from policies that are generated at the population level. Ananth Balashankar, Samuel P. Fraiberger, Eric Deregt, Marelize Gorgens, Lakshminarayanan Subramanian |
COMPASS | 5 |
| 2022 | To Block or Not to Block: Accelerating Mobile Web Pages On-The-Fly Through JavaScript ClassificationabstractThe increasing complexity of JavaScript (JS) in modern mobile web pages has become a performance bottleneck for low-end mobile phone users, especially in developing regions. In this paper we propose SlimWeb, a novel approach that automatically derives lightweight versions of mobile web pages on-the-fly by eliminating non-essential JavaScript that does not impact the core page content and interactive functionality. SlimWeb consists of a JavaScript classification service powered by a supervised Machine Learning (ML) model that provides insights into each JavaScript element embedded in a web page. SlimWeb aims to improve the web browsing experience by predicting the class of each element, such that essential elements are preserved and non-essential elements are blocked by the browsers using the service. We motivate SlimWeb’s core design via a preference survey where 306 users overwhelmingly preferred having faster page load times over fetching various categories of non-essential JavaScript. We evaluate SlimWeb across 500 popular web pages in a developing region on real cellular networks, along with a user experience study with 20 real-world users and a usage willingness survey of 588 users. Evaluation results show that SlimWeb achieves 50% reduction in page load time compared to the original pages, and more than 30% reduction compared to competing solutions, while achieving high similarity scores to the original pages measured via a qualitative evaluation study with 62 users. SlimWeb improves the overall user experience metric (defined by Google Lighthouse combining first contentful paint, time to interactive, speed index) by more than 60% compared to the original pages, while maintaining 90-100% of the visual and functional components of most pages. Moumena Chaqfeh, Waleed Hashmi, Patrick Inshuti, Manesha Ramesh, Matteo Varvello, Lakshminarayanan Subramanian, Fareed Zaffar, Yasir Zaki |
ICTD | 7 |
| 2022 | Learning Pollution Maps from Mobile Phone ImagesabstractAir pollution monitoring and management is one of the key challenges for urban sectors, especially in developing countries. Measuring pollution levels requires significant investment in reliable and durable instrumentation and subsequent maintenance. On the other hand, there have been many attempts by researchers to develop image-based pollution measurement models which have shown significant results and established the feasibility of the idea. But, taking image-level models to a city-level system presents new challenges, which include scarcity of high-quality annotated data and a high amount of label noise. In this paper, we present a low-cost, end-to-end system for learning pollution maps using images captured through a mobile phone. We demonstrate our system for parts of New Delhi and Ghaziabad. We use transfer learning to overcome the problem of data scarcity. We investigate the effects of label noise in detail and introduce the metric of in-interval accuracy to evaluate our models in presence of noise. We use distributed averaging to learn pollution maps and mitigate the effects of noise to some extent. We also develop haze-based interpretable models which have comparable performance to mainstream models. With only 382 images from Delhi and Ghaziabad and single-scene dataset from Beijing and Shanghai, we are able to achieve a mean absolute error of 44 ug/m^3 in PM2.5 concentration on a test set of 267 images and an in-interval accuracy of 67% on predictions. Going further, we learn pollution maps with a mean absolute error as low as 35 ug/m^3 and in-interval accuracy as high as 74% significantly mitigating the image models' error. We also show that the noise in pollution labels emerging from unreliable sensing instrumentation forms a significant barrier to the realization of an ideal air pollution monitoring system. Our codebase can be found at https://github.com/ankitbha/pollution_with_images. Ankit Bhardwaj 0001, Shiva R. Iyer, Yash Jalan, Lakshminarayanan Subramanian |
IJCAI | 4 |
| 2022 | Muzeel: assessing the impact of JavaScript dead code elimination on mobile web performanceabstractTo quickly create interactive web pages, developers heavily rely on (large) general-purpose JavaScript libraries. This practice bloats web pages with complex unused functions dead code which are unnecessarily downloaded and processed by the browser. The identification and the elimination of these functions is an open problem, which this paper tackles with Muzeel, a black-box approach requiring neither knowledge of the code nor execution traces. While the state-of-the-art solutions stop analyzing JavaScript when the page loads, the core design principle of Muzeel is to address the challenge of dynamically analyzing JavaScript after the page is loaded, by emulating all possible user interactions with the page, such that the used functions (executed when interactivity events fire) are accurately identified, whereas unused functions are filtered out and eliminated. We run Muzeel against 15,000 popular web pages and show that half of the 300,000 JavaScript files used in these pages have at least 70% of unused functions, accounting for 55% of the files' sizes. To assess the impact of dead code elimination on Mobile Web performance, we serve 200 Muzeel-ed pages to several Android phones and browsers, under variable network conditions. Our evaluation shows that Muzeel can speed up page loads by 25-30% thanks to a combination of lower CPU and bandwidth usage. Most importantly, we show that such savings are achieved while maintaining the pages' visual appearance and interactive functionality. Jesutofunmi Kupoluyi, Moumena Chaqfeh, Matteo Varvello, Russell Coke, Waleed Hashmi, Lakshminarayanan Subramanian, Yasir Zaki |
IMC | 6 |
| 2022 | QLUE: A Computer Vision Tool for Uniform Qualitative Evaluation of Web PagesabstractThe increasing complexity of the web has attracted a number of solutions to offer optimized versions of web pages that are lighter to process and faster to load. These solutions have been quantitatively evaluated to show significant speed-ups in load times and/or considerable savings in bandwidth/memory consumption. However, while these solutions often produce optimized versions from existing pages, they rarely evaluate the impact of their optimizations on the original content and functionality. Additionally, due to the lack of a unified metric to evaluate the similarity of the pages generated by these solutions in comparison to the original pages, it is not yet possible to fairly compare the results obtained from different user studies campaigns, unless recruiting the exact same users, which is extremely challenging. In this paper, we demonstrate the lack of qualitative evaluation metrics, and propose QLUE (QuaLitative Uniform Evaluation), a tool that automates the qualitative evaluation of web pages generated by web complexity solutions with respect to their original versions using computer vision. QLUE evaluates the content and the functionality of these pages separately using two metrics: QLUE’s Structural Similarity, to assess the former, and QLUE’s Functional Similarity to assess the latter—a task that is proven to be challenging for humans given the complex functional dependencies in modern pages. Our results show that QLUE computes comparable content and functional scores to those provided by humans. Specifically, 90% of a set of 100 pages were given a similarity score between 90% and 100% by human evaluators, while QLUE shows similar scores for more than 75% of the same pages. In terms of time complexity, QLUE shows that it is capable of evaluating an optimized web page in a few minutes. Waleed Hashmi, Moumena Chaqfeh, Lakshminarayanan Subramanian, Yasir Zaki |
WWW | 3 |
| 2022 | JSAnalyzer: A Web Developer Tool for Simplifying Mobile Web Pages through Non-critical JavaScript EliminationabstractThe amount of JavaScript used in web pages has substantially grown in the past decade, leading to large and complex pages that are computationally intensive for handheld mobile devices. Due to the increasing usage of these devices to access today’s web, and to accommodate the needs of a large number of mobile web users who solely rely on low-end devices, we propose “JSAnalyzer,” an easy-to-use tool that enables web developers to quickly optimize JavaScript usage in their pages and to generate simpler versions of these pages for mobile web users. JSAnalyzer is motivated by the widespread use of non-critical JavaScript elements, i.e., those that have negligible (if any) impact on the page’s visual content and interactive functionality. JSAnalyzer allows the developer to selectively enable or disable JavaScript elements in any given page while visually observing their impact on the page to (1) accurately identify any non-critical JavaScript elements and (2) create a simplified page with these elements removed. Our quantitative evaluation shows that, given a low-end mobile phone, JSAnalyzer achieves an increase of nearly 90% in Google’s lighthouse performance score while reducing the page load time by 30%. A qualitative study of 22 users shows that the lighter pages produced by JSAnalyzer maintain more than 90% visual similarity compared to the original pages. Moreover, JSAnalyzer was evaluated by 69 developers, showing that it scores nearly 90% in terms of usefulness and usability while retaining the page’s content and functionality. Finally, we show that JSAnalyzer outperforms state-of-the-art solutions in terms of timing speedups and resource savings. Moumena Chaqfeh, Russell Coke, Jacinta Hu, Waleed Hashmi, Lakshminarayanan Subramanian, Talal Rahwan, Yasir Zaki |
ACM Trans. Web | 5 |
| 2021 | Learning Faithful Representations of Causal GraphsabstractAnanth Balashankar, Lakshminarayanan Subramanian. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ananth Balashankar, Lakshminarayanan Subramanian |
ACL/IJCNLP (1) | 2 |
| 2021 | CONTRA: Contrarian statistics for controlled variable selectionabstractThe holdout randomization test (HRT) discovers a set of covariates most predictive of a response. Given the covariate distribution, HRTs can explicitly control the false discovery rate (FDR). However, if this distribution is unknown and must be estimated from data, HRTs can inflate the FDR. To alleviate the inflation of FDR, we propose the contrarian randomization test (CONTRA), which is designed explicitly for scenarios where the covariate distribution must be estimated from data and may even be misspecified. Our key insight is to use an equal mixture of two “contrarian” probabilistic models in determining the importance of a covariate. One model is fit with the real data, while the other is fit using the same data, but with the covariate being tested replaced with samples from an estimate of the covariate distribution. CONTRA is flexible enough to achieve a power of 1 asymptotically, can reduce the FDR compared to state-of-the-art CVS methods when the covariate distribution is misspecified, and is computationally efficient in high dimensions and large sample sizes. We further demonstrate the effectiveness of CONTRA on numerous synthetic benchmarks, and highlight its capabilities on a genetic dataset. Mukund Sudarshan, Aahlad Manas Puli, Lakshminarayanan Subramanian, Sriram Sankararaman, Rajesh Ranganath |
AISTATS | 3 |
| 2021 | Deep Significance Clustering (DICE) a Heterogenous Population
Yufang Huang, Peter A. D. Steel, Kelly M. Axsom, Sri Lekha Tummalapalli, Alison Hermann, Rochelle Joly, Fei Wang 0001, Jyotishman Pathak, Lakshminarayanan Subramanian, Yiye Zhang |
AMIA | 11 |
| 2021 | CollectiveTeach: A System To Generate And Sequence Web-Annotated Lesson PlansabstractDespite an abundance of educational resources on the Web, there exists a gap between teachers and the efficient utilization of these resources. A fundamental component of teaching is the preparation of a lesson plan—an organized sequence of educational content—and for the most part, the task of generating lesson plans today is manual and laborious. To address this gap, we present CollectiveTeach, a platform that enables educators to generate lesson plans. CollectiveTeach has two main facets: (i) an information retrieval engine that gathers relevant documents pertaining to a topic, and (ii) a framework to sequence the retrieved documents into coherent lesson plans. We present a novel architecture that leverages information retrieval algorithms, data mining techniques, and user feedback to generate automated lesson plans. We built and deployed CollectiveTeach for 3 popular undergraduate Computer Science subjects: Algorithms, Operating Systems, and Machine Learning, on a corpus of ∼ 100,000 web pages. Further, we evaluated the platform in 3 phases: (1) computing the precision of the documents retrieved, (2) a user study with 10 participants who assessed lesson plans returned by CollectiveTeach based on appropriateness, quality, and coverage and (3) benchmarking our sequencing approach against the Beam-Search approach. Our results show that CollectiveTeach achieves high precision in retrieving content relevant to a user’s query, users are satisfied with the appropriateness, coverage, and reliability of the generated lesson plans and that our sequencing approach is effective. These results indicate that CollectiveTeach is a promising platform that could enrich the lesson plan generation process and encourage collaboration amongst the community of educators and learners. Rishabh Ranawat, Ashwin Venkataraman, Lakshminarayanan Subramanian |
COMPASS | 3 |
| 2021 | Enhancing Neural Recommender Models through Domain-Specific ConcordanceabstractRecommender models trained on historical observational data alone can be brittle when domain experts subject them to counterfactual evaluation. In many domains, experts can articulate common, high-level mappings or rules between categories of inputs (user's history) and categories of outputs (preferred recommendations). One challenge is to determine how to train recommender models to adhere to these rules. In this work, we introduce the goal of domain-specific concordance: the expectation that a recommender model follow a set of expert-defined categorical rules. We propose a regularization-based approach that optimizes for robustness on rule-based input perturbations. To test the effectiveness of this method, we apply it in a medication recommender model over diagnosis-medicine categories, and in movie and music recommender models, on rules over categories based on movie tags and song genres. We demonstrate that we can increase the category-based robustness distance by up to 126% without degrading accuracy, but rather increasing it by up to 12% compared to baseline models in the popular MIMIC-III, MovieLens-20M and Last.fm Million Song datasets. Ananth Balashankar, Alex Beutel, Lakshminarayanan Subramanian |
WSDM | 3 |
| 2021 | Deep significance clustering: a novel approach for identifying risk-stratified and predictive patient subgroupsabstractOBJECTIVE: Deep significance clustering (DICE) is a self-supervised learning framework. DICE identifies clinically similar and risk-stratified subgroups that neither unsupervised clustering algorithms nor supervised risk prediction algorithms alone are guaranteed to generate. MATERIALS AND METHODS: Enabled by an optimization process that enforces statistical significance between the outcome and subgroup membership, DICE jointly trains 3 components, representation learning, clustering, and outcome prediction while providing interpretability to the deep representations. DICE also allows unseen patients to be predicted into trained subgroups for population-level risk stratification. We evaluated DICE using electronic health record datasets derived from 2 urban hospitals. Outcomes and patient cohorts used include discharge disposition to home among heart failure (HF) patients and acute kidney injury among COVID-19 (Cov-AKI) patients, respectively. RESULTS: Compared to baseline approaches including principal component analysis, DICE demonstrated superior performance in the cluster purity metrics: Silhouette score (0.48 for HF, 0.51 for Cov-AKI), Calinski-Harabasz index (212 for HF, 254 for Cov-AKI), and Davies-Bouldin index (0.86 for HF, 0.66 for Cov-AKI), and prediction metric: area under the Receiver operating characteristic (ROC) curve (0.83 for HF, 0.78 for Cov-AKI). Clinical evaluation of DICE-generated subgroups revealed more meaningful distributions of member characteristics across subgroups, and higher risk ratios between subgroups. Furthermore, DICE-generated subgroup membership alone was moderately predictive of outcomes. DISCUSSION: DICE addresses a gap in current machine learning approaches where predicted risk may not lead directly to actionable clinical steps. CONCLUSION: DICE demonstrated the potential to apply in heterogeneous populations, where having the same quantitative risk does not equate with having a similar clinical profile. Yufang Huang, Peter A. D. Steel, Kelly M. Axsom, Sri Lekha Tummalapalli, Fei Wang 0001, Jyotishman Pathak, Lakshminarayanan Subramanian, Yiye Zhang |
J. Am. Medical Informatics Assoc. | 9 |
| 2020 | Forecasting Sparse Traffic Congestion Patterns Using Message-Passing RNNSabstractThe ability to forecast traffic congestion ahead of time given road conditions has remained a prominent problem in road traffic analysis. In this work, we leverage mobility traces of public transport vehicles tracked by the New York City MTA and formulate Message-Passing Recurrent Neural Nets (MPRNN) to produce long-term traffic forecasting on data that is sparse but wide in coverage. We model the interactions among road segments spread over the entirety of Manhattan, New York over a period of 3 months, such that traffic conditions can be propagated to > 90% of examined segments from just a few observations. In comparison to other competing algorithms, MPRNN achieves the lowest mean error of <; 0.3 mph when predicting ahead in 10 minute intervals, for up to 3 road segments ahead (message passing across 3 hops). The MPRNN model further offers compelling results when forecasting traffic speeds several hours ahead given distant observations up to approximately 1 kilometer away (three consecutive bus stops) with a mean error of about 2 mph. Shiva R. Iyer, Ulzee An, Lakshminarayanan Subramanian |
ICASSP | 3 |
| 2020 | A hyperlocal mobile web for the next 3 billion usersabstractDespite increasing mobile Internet penetration in developing regions, growing web page complexity and the lack of optimization from remote content providers make the web experience poor in these areas. The high relative bandwidth cost, poor network performance, and lack of relevant local content combine to dampen the demand for the Internet and services it enables. In this paper we propose GAIUS, a content ecosystem enabling efficient creation and dissemination of locally relevant web content. At its core, GAIUS consists of the following innovations: a locally sustainable content ecosystem, and MAML, a web specification language that simplifies web pages to reduce costs and lower barriers for content creation. Arjuna Sathiaseelan, Arko Chatterjee, Mukund Lal, Yasir Zaki, Lakshminarayanan Subramanian |
MobiCom | 5 |
| 2020 | JSCleaner: De-Cluttering Mobile Webpages Through JavaScript CleanupabstractA significant fraction of the World Wide Web suffers from the excessive usage of JavaScript (JS). Based on an analysis of popular webpages, we observed that a considerable number of JS elements utilized by these pages are not essential for their visual and functional features. In this paper, we propose JSCleaner, a JavaScript de-cluttering engine that aims at simplifying webpages without compromising their content or functionality. JSCleaner relies on a rule-based classification algorithm that classifies JS into three main categories: non-critical, replaceable, and critical. JSCleaner removes non-critical JS from a webpage, translates replaceable JS elements with their HTML outcomes, and preserves critical JS. Our quantitative evaluation of 500 popular webpages shows that JSCleaner achieves around 30% reduction in page load times coupled with a 50% reduction in the number of requests and the page size. In addition, our qualitative user study of 103 evaluators shows that JSCleaner preserves 95% of the page content similarity, while maintaining nearly 88% of the page functionality (the remaining 12% did not have a major impact on the user browsing experience). Moumena Chaqfeh, Yasir Zaki, Jacinta Hu, Lakshminarayanan Subramanian |
WWW | 4 |
| 2020 | Sepsis in the era of data-driven medicine: personalizing risks, diagnoses, treatments and prognosesabstractSepsis is a series of clinical syndromes caused by the immunological response to infection. The clinical evidence for sepsis could typically attribute to bacterial infection or bacterial endotoxins, but infections due to viruses, fungi or parasites could also lead to sepsis. Regardless of the etiology, rapid clinical deterioration, prolonged stay in intensive care units and high risk for mortality correlate with the incidence of sepsis. Despite its prevalence and morbidity, improvement in sepsis outcomes has remained limited. In this comprehensive review, we summarize the current landscape of risk estimation, diagnosis, treatment and prognosis strategies in the setting of sepsis and discuss future challenges. We argue that the advent of modern technologies such as in-depth molecular profiling, biomedical big data and machine intelligence methods will augment the treatment and prevention of sepsis. The volume, variety, veracity and velocity of heterogeneous data generated as part of healthcare delivery and recent advances in biotechnology-driven therapeutics and companion diagnostics may provide a new wave of approaches to identify the most at-risk sepsis patients and reduce the symptom burden in patients within shorter turnaround times. Developing novel therapies by leveraging modern drug discovery strategies including computational drug repositioning, cell and gene-therapy, clustered regularly interspaced short palindromic repeats -based genetic editing systems, immunotherapy, microbiome restoration, nanomaterial-based therapy and phage therapy may help to develop treatments to target sepsis. We also provide empirical evidence for potential new sepsis targets including FER and STARD3NL. Implementing data-driven methods that use real-time collection and analysis of clinical variables to trace, track and treat sepsis-related adverse outcomes will be key. Understanding the root and route of sepsis and its comorbid conditions that complicate treatment outcomes and lead to organ dysfunction may help to facilitate identification of most at-risk patients and prevent further deterioration. To conclude, leveraging the advances in precision medicine, biomedical data science and translational bioinformatics approaches may help to develop better strategies to diagnose and treat sepsis in the next decade. Andrew C. Liu, Krishna Patel, Ramya Dhatri Vunikili, Kipp W. Johnson, Fahad Jibrin Abdu, Shivani Kamath Belman, Benjamin S. Glicksberg, Pratyush Tandale, Roberto Fontanez, Oommen K. Mathew, Andrew Kasarskis, Priyabrata Mukherjee, Lakshminarayanan Subramanian, Joel Dudley, Khader Shameer |
Briefings Bioinform. | 13 |
| 2019 | Reconstructing the MERS disease outbreak from newsabstractDisease surveillance is critical for mobilizing health care resources and deciding on isolation measures to contain the spread of infectious diseases. Because ground truth signals of rare and deadly diseases are sparse, it can be useful to enrich surveillance systems using measures of social and environmental factors which are known to influence the spread of a disease. One approach to measure such factors is by using real time news streams. In this study, we model the epidemiological transmission of the Middle Eastern Respiratory Syndrome (MERS) disease during the outbreak that occurred from 2013 to 2018 in the Arabian peninsula. Using the GDELT news event database, we show that conflict related signals allow us to reconstruct the time series of newly infected cases per week. This reduces the residual sum of squared errors by a factor of 3.36 as compared to a standard epidemiological model. We also capture interpretable time-sensitive factors which illustrate the importance of using real time news stream to model the evolution of a disease such as MERS and facilitate early and effective policy interventions. Ananth Balashankar, Aashish Dugar, Lakshminarayanan Subramanian, Samuel P. Fraiberger |
COMPASS | 3 |
| 2019 | Identifying Predictive Causal Factors from News StreamsabstractAnanth Balashankar, Sunandan Chakraborty, Samuel Fraiberger, Lakshminarayanan Subramanian. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ananth Balashankar, Sunandan Chakraborty, Samuel P. Fraiberger, Lakshminarayanan Subramanian |
EMNLP/IJCNLP (1) | 4 |
| 2019 | GAIUS: a new mobile content creation and diffusion ecosystem for emerging regionsabstractDespite increasing mobile Internet penetration in emerging regions, the growth of web page complexity combined with inadequate content delivery mechanisms and lack of relevant local content make the web experience poor for users in these regions. In this paper, we propose GAIUS, a content ecosystem that enables the efficient creation and dissemination of locally relevant web content. GAIUS aims at creating a community content exchange that enables users to easily create and consume relevant web content with the help of a mobile application. Based on an early deployment of GAIUS in various regions, we show that our ecosystem enables localized content creation and diffusion efficiently. Talal Ahmad, Yasir Zaki, Thomas Pötsch, Jay Chen, Arjuna Sathiaseelan, Lakshminarayanan Subramanian |
ICTD | 6 |
| 2019 | VACCINE: Using Contextual Integrity For Data Leakage DetectionabstractModern enterprises rely on Data Leakage Prevention (DLP) systems to enforce privacy policies that prevent unintentional flow of sensitive information to unauthorized entities. However, these systems operate based on rule sets that are limited to syntactic analysis and therefore completely ignore the semantic relationships between participants involved in the information exchanges. For similar reasons, these systems cannot enforce complex privacy policies that require temporal reasoning about events that have previously occurred. Yan Shvartzshnaider, Zvonimir Pavlinovic, Ananth Balashankar, Thomas Wies, Lakshminarayanan Subramanian, Helen Nissenbaum, Prateek Mittal |
WWW | 5 |
| 2018 | GreenApps: A Platform For Cellular Edge ApplicationsabstractThis paper presents the design, implementation and deployment of GreenApps, a ground-up platform that enables off-grid, near off-line and highly available cellular applications in rural contexts. The GreenApps platform enables rapid development and deployment of almost-always available applications that can be executed locally on top of open-sourced cellular base stations. The GreenApps platform has been deployed in two rural regions in Kumawu, Ghana and Pearl Lagoon, Nicaragua and supports different community-centric applications. Talal Ahmad, Edwin Reed-Sanchez, Fatima Zarinni, Alfred Afutu, Kessir Adjaho, Yaw Nyarko, Lakshminarayanan Subramanian |
COMPASS | 7 |
| 2017 | xCache: Rethinking Edge Caching for Developing RegionsabstractEnd-users in emerging markets experience poor web performance due to a combination of three factors: high server response time, limited edge bandwidth and the complexity of web pages. The absence of cloud infrastructure in developing regions and the limited bandwidth experienced by edge nodes constrain the effectiveness of conventional caching solutions for these contexts. This paper describes the design, implementation and deployment of xCache, a cloud-managed Internet caching architecture that aims to proactively profile popular web pages and maintain the liveness of popular content at software defined edge caches to enhance the cache hit rate with minimal bandwidth overhead. xCache uses a Cloud Controller that continuously analyzes active cloud-managed web pages and derives an object-group representation of web pages based on the objects of a page. Using this object-group representation, xCache computes a bandwidth-aware utility measure to derive the most valuable configuration for each edge cache. Our preliminary real-world deployment across university campuses in three developing regions demonstrates its potential compared to conventional caching by improving cache hit rates by about 15%. Our evaluations of xCache have also shown that it can be applied in conjunction with other web optimizations solutions like Shandian, and can improve page load times by more than 50%. Ali Raza 0003, Yasir Zaki, Thomas Pötsch, Jay Chen, Lakshminarayanan Subramanian |
ICTD | 5 |
| 2017 | The Fake vs Real Goods Problem: Microscopy and Machine Learning to the RescueabstractCounterfeiting of physical goods is a global problem amounting to nearly 7% of world trade. While there have been a variety of overt technologies like holograms and specialized barcodes and covert technologies like taggants and PUFs, these solutions have had a limited impact on the counterfeit market due to a variety of factors - clonability, cost or adoption barriers. In this paper, we introduce a new mechanism that uses machine learning algorithms on microscopic images of physical objects to distinguish between genuine and counterfeit versions of the same product. The underlying principle of our system stems from the idea that microscopic characteristics in a genuine product or a class of products (corresponding to the same larger product line), exhibit inherent similarities that can be used to distinguish these products from their corresponding counterfeit versions. A key building block for our system is a wide-angle microscopy device compatible with a mobile device that enables a user to easily capture the microscopic image of a large area of a physical object. Based on the captured microscopic images, we show that using machine learning algorithms (ConvNets and bag of words), one can generate a highly accurate classification engine for separating the genuine versions of a product from the counterfeit ones; this property also holds for "super-fake" counterfeits observed in the marketplace that are not easily discernible from the human eye. We describe the design of an end-to-end physical authentication system leveraging mobile devices, portable hardware and a cloud-based object verification ecosystem. We evaluate our system using a large dataset of 3 million images across various objects and materials such as fabrics, leather, pills, electronics, toys and shoes. The classification accuracy is more than 98% and we show how our system works with a cellphone to verify the authenticity of everyday objects. Ashlesh Sharma, Vidyuth Srinivasan, Vishal Kanchan, Lakshminarayanan Subramanian |
KDD | 4 |
| 2017 | More than a Feeling: The MiFace Framework for Defining Facial Communication MappingsabstractFacial expressions transmit a variety of social, grammatical, and affective signals. For technology to leverage this rich source of communication, tools that better model the breadth of information they convey are required. MiFace is a novel framework for creating expression lexicons that map signal values to parameterized facial muscle movements. In traditional mapping paradigms using posed photographs, naïve judges select from predetermined label sets and movements are inferred by trained experts. The set of generally accepted expressions established in this way is limited to six basic displays of affect. In contrast, our approach generatively simulates muscle movements on a 3D avatar. By applying natural language processing techniques to crowdsourced free-response labels for the resulting images, we efficiently converge on an expression's value across signal categories. Two studies returned 218 discriminable facial expressions with 51 unique labels. The six basic emotions are included, but we additionally define such nuanced expressions as embarrassed, curious, and hopeful. Crystal Butler, Stephanie Michalowicz, Lakshminarayanan Subramanian, Winslow Burleson |
UIST | 3 |
| 2017 | Identifying Unreliable and Adversarial Workers in Crowdsourced Labeling TasksabstractWe study the problem of identifying unreliable and adversarial workers in crowdsourcing systems where workers (or users) provide labels for tasks (or items). Most existing studies assume that worker responses follow specific probabilistic models; however, recent evidence shows the presence of workers adopting non-random or even malicious strategies. To account for such workers, we suppose that workers comprise a mixture of honest and adversarial workers. Honest workers may be reliable or unreliable, and they provide labels according to an unknown but explicit probabilistic model. Adversaries adopt labeling strategies different from those of honest workers, whether probabilistic or not. We propose two reputation algorithms to identify unreliable honest workers and adversarial workers from only their responses. Our algorithms assume that honest workers are in the majority, and they classify workers with outlier label patterns as adversaries. Theoretically, we show that our algorithms successfully identify unreliable honest workers, workers adopting deterministic strategies, and worst- case sophisticated adversaries who can adopt arbitrary labeling strategies to degrade the accuracy of the inferred task labels. Empirically, we show that filtering out outliers using our algorithms can significantly improve the accuracy of several state-of-the-art label aggregation algorithms in real-world crowdsourcing datasets. Srikanth Jagabathula, Lakshminarayanan Subramanian, Ashwin Venkataraman |
J. Mach. Learn. Res. | 2 |
| 2016 | The Effects of the Content of FOMC Communications on US Treasury RatesabstractThis study measures the effects of Federal Open Market Committee text content on the direction of short-and medium-term interest rate movements.Because the words relevant to short-and medium-term interest rates differ, we apply a supervised approach to learn distinct sets of topics for each dependent variable being examined.We generate predictions with and without controlling for factors relevant to interest rate movements, and our prediction results average across multiple training-test splits.Using data from 1999-2016, we achieve 93% and 64% accuracy in predicting Target and Effective Federal Funds Rate movements and 38%-40% accuracy in predicting longer term Treasury Rate movements.We obtain lower but comparable accuracies after controlling for other macroeconomic and market factors. Chris Rohlfs, Sunandan Chakraborty, Lakshminarayanan Subramanian |
EMNLP | 3 |
| 2016 | Learning Privacy Expectations by Crowdsourcing Contextual Informational NormsabstractDesigning programmable privacy logic frameworks that correspond to social, ethical, and legal norms has been a fundamentally hard problem. Contextual integrity (CI) (Nissenbaum, 2010) offers a model for conceptualizing privacy that is able to bridge technical design with ethical, legal, and policy approaches. While CI is capable of capturing the various components of contextual privacy in theory, it is challenging to discover and formally express these norms in operational terms. In the following, we propose a crowdsourcing method for the automated discovery of contextual norms. To evaluate the effectiveness and scalability of our approach, we conducted an extensive survey on Amazon's Mechanical Turk (AMT) with more than 450 participants and 1400 questions. The paper has three main takeaways: First, we demonstrate the ability to generate survey questions corresponding to privacy norms within any context. Second, we show that crowdsourcing enables the discovery of norms from these questions with strong majoritarian consensus among users. Finally, we demonstrate how the norms thus discovered can be encoded into a formal logic to automatically verify their consistency. Yan Shvartzshnaider, Schrasing Tong, Thomas Wies, Paula Kift, Helen Nissenbaum, Lakshminarayanan Subramanian, Prateek Mittal |
HCOMP | 6 |
| 2016 | Predicting Socio-Economic Indicators using News EventsabstractMany socio-economic indicators are sensitive to real-world events. Proper characterization of the events can help to identify the relevant events that drive fluctuations in these indicators. In this paper, we propose a novel generative model of real-world events and employ it to extract events from a large corpus of news articles. We introduce the notion of an event class, which is an abstract grouping of similarly themed events. These event classes are manifested in news articles in the form of event triggers which are specific words that describe the actions or incidents reported in any article. We use the extracted events to predict fluctuations in different socio-economic indicators. Specifically, we focus on food prices and predict the price of 12 different crops based on real-world events that potentially influence food price volatility, such as transport strikes, festivals etc. Our experiments demonstrate that incorporating event information in the prediction tasks reduces the root mean square error (RMSE) of prediction by 22% compared to the standard ARIMA model. We also predict sudden increases in the food prices (i.e. spikes) using events as features, and achieve an average 5-10% increase in accuracy compared to baseline models, including an LDA topic-model based predictive model. Sunandan Chakraborty, Ashwin Venkataraman, Srikanth Jagabathula, Lakshminarayanan Subramanian |
KDD | 4 |
| 2015 | Solar vs diesel: where to draw the line for cell towers?abstractCellular networks in developing regions continue to rely heavily on diesel for energy to provide network coverage due to the paucity of reliable grid power which directly impacts the network's economic viability and long-term sustainability. At the other extreme, solar powered cellular installations have gained prominence but have faced their own adoption challenges including inability to provide adequate and reliable 24×7 power supply, need for large land footprints and lack of efficient power storage. In this paper, we perform a detailed economic cost analysis comparing diesel powered cellular networks with solar powered cellular networks. The key goal of this paper is to establish the cross-over boundary beyond which solar powered installations are better than diesel powered alternatives. We perform a detailed analysis based on actual diesel consumption data from a large telecom operator in a developing region. Using our model, we can also easily perform an extended analysis based on future projections on solar efficiencies and future cellular network designs. Talal Ahmad, Shankar Kalyanaraman, Fareeha Amjad, Lakshminarayanan Subramanian |
ICTD | 4 |
| 2015 | Extreme Web Caching for Faster Web BrowsingabstractModern web pages are very complex; each web page consists of hundreds of objects that are linked from various servers all over the world. While mechanisms such as caching reduce the overall number of end-to-end requests saving bandwidth and loading time, there is still a large portion of content that is re-fetched -- despite not having changed. In this demo, we present Extreme Cache, a web caching architecture that enhances the web browsing experience through a smart pre-fetching engine. Our extreme cache tries to predict the rate of change of web page objects to bring cacheable content closer to the user. Ali Raza 0003, Yasir Zaki, Thomas Pötsch, Jay Chen, Lakshminarayanan Subramanian |
SIGCOMM | 5 |
| 2015 | Adaptive Congestion Control for Unpredictable Cellular NetworksabstractLegacy congestion controls including TCP and its variants are known to perform poorly over cellular networks due to highly variable capacities over short time scales, self-inflicted packet delays, and packet losses unrelated to congestion. To cope with these challenges, we present Verus, an end-to-end congestion control protocol that uses delay measurements to react quickly to the capacity changes in cellular networks without explicitly attempting to predict the cellular channel dynamics. The key idea of Verus is to continuously learn a delay profile that captures the relationship between end-to-end packet delay and outstanding window size over short epochs and uses this relationship to increment or decrement the window size based on the observed short-term packet delay variations. While the delay-based control is primarily for congestion avoidance, Verus uses standard TCP features including multiplicative decrease upon packet loss and slow start. Through a combination of simulations, empirical evaluations using cellular network traces, and real-world evaluations against standard TCP flavors and state of the art protocols like Sprout, we show that Verus outperforms these protocols in cellular channels. In comparison to TCP Cubic, Verus achieves an order of magnitude (> 10x) reduction in delay over 3G and LTE networks while achieving comparable throughput (sometimes marginally higher). In comparison to Sprout, Verus achieves up to 30% higher throughput in rapidly changing cellular networks. Yasir Zaki, Thomas Pötsch, Jay Chen, Lakshminarayanan Subramanian, Carmelita Görg |
SIGCOMM | 4 |
| 2014 | TAQ: enhancing fairness and performance predictability in small packet regimesabstractTCP congestion control algorithms implicitly assume that the per-flow throughput is at least a few packets per round trip time. Environments where this assumption does not hold, which we refer to as small packet regimes, are common in the contexts of wired and cellular networks in developing regions. In this paper we show that in small packet regimes TCP flows experience severe unfairness, high packet loss rates, and flow silences due to repetitive timeouts. We propose an approximate Markov model to describe TCP behavior in small packet regimes to characterize the TCP breakdown region that leads to repetitive timeout behavior. To enhance TCP performance in such regimes, we propose Timeout Aware Queuing (TAQ), a readily deployable in-network middlebox approach that uses a multi-level adaptive priority queuing algorithm to reduce the probability of timeouts, improve fairness and performance predictability. We demonstrate the effectiveness of TAQ across a spectrum of small packet regime network conditions using simulations, a prototype implementation, and testbed experiments. Jay Chen, Lakshminarayanan Subramanian, Janardhan R. Iyengar, Bryan Ford |
EuroSys | 2 |
| 2014 | Dissecting Web Latency in GhanaabstractWeb access is prohibitively slow in many developing regions despite substantial effort to increase bandwidth and network penetration. In this paper, we explore the fundamental bottlenecks that cause poor web performance from a client's perspective by carefully dissecting webpage load latency contributors in Ghana. Based on our measurements from 2012 to 2014, we find several interesting issues that arise due to the increasing complexity of web pages and number of server redirections required to completely render the assets of a page. We observe that, rather than bandwidth, the primary bottleneck of web performance in Ghana is the lack of good DNS servers and caching infrastructure. The main bottlenecks are: (a) Recursive DNS query resolutions; (b) HTTP redirections; (c) TLS/SSL handshakes. We experiment with a range of well-known end-to-end latency optimizations and find that simple DNS caching, redirection caching, and the use of SPDY can all yield substantial improvements to user-perceived latency. Yasir Zaki, Jay Chen, Thomas Pötsch, Talal Ahmad, Lakshminarayanan Subramanian |
Internet Measurement Conference | 5 |
| 2014 | Tackling societal grand challenges using mobile computingabstractMobile computing has deeply influenced the lives of almost every human in the world on a daily basis. In celebration of the 20th anniversary of Mobicom, Mobicom 2014 features an exciting panel on the topic of tackling societal grand challenges using mobile computing. The panel features five excellent researchers who have had a tremendous impact on the field of mobile computing over the years and whose research work has directly addressed pressing societal problems. The panel discussion will feature an in-depth discussion of the grand challenges that we face in society today and the role of mobile computing as a frontier platform for addressing these challenges. The panel will cover both a retrospective and futuristic perspective where the retrospective aspects highlight the diverse contributions of the panelists and the futuristic aspects will feature of a discussion of their individual views of what are the next big societal problems that Mobicom as a community should tackle. Lakshminarayanan Subramanian, Sanjit Biswas, Gaetano Borriello, Prabal Dutta, Dina Katabi, Randy H. Katz |
MobiCom | 1 |
| 2014 | Reputation-based Worker Filtering in Crowdsourcing
Srikanth Jagabathula, Lakshminarayanan Subramanian, Ashwin Venkataraman |
NIPS | 2 |
| 2011 | PaperSpeckle: microscopic fingerprinting of paperabstractPaper forgery is among the leading causes of corruption in many developing regions. In this paper, we introduce PaperSpeckle, a robust system that leverages the natural randomness property present in paper to generate a fingerprint for any piece of paper. Our goal in developing PaperSpeckle is to build a low-cost paper based authentication mechanism for applications in rural regions such as microfinance, healthcare, land ownership records, supply chain services and education which heavily rely on paper based records. Unlike prior paper fingerprinting techniques that have extracted fingerprints based on the fiber structure of paper, PaperSpeckle uses the texture speckle pattern, a random bright/dark region formation at the microscopic level when light falls on to the paper, to extract a unique fingerprint to identify paper. In PaperSpeckle, we show how to extract a "repeatable" texture speckle pattern of a microscopic region of a paper using low-cost machinery involving paper, pen and a cheap microscope. Using extensive testing on different types of paper, we show that PaperSpeckle can produce a robust repeatable fingerprint even if paper is damaged due to crumpling, printing or scribbling, soaking in water or aging with time. Ashlesh Sharma, Lakshminarayanan Subramanian, Eric A. Brewer |
CCS | 2 |
| 2011 | Optimal Sybil-resilient node admission controlabstractMost existing large-scale networked systems on the Internet such as peer-to-peer systems are vulnerable to Sybil attacks where a single adversary can introduce many bogus identities. One promising defense of Sybil attacks is to perform social-network based admission control to bound the number of Sybil identities admitted. SybilLimit, the best known Sybil admission control mechanism, can restrict the number of Sybil identities admitted per attack edge to O(log n) with high probability assuming O(n/ log n) attack edges. In this paper, we propose Gatekeeper, a decentralized Sybil-resilient admission control protocol that significantly improves over SybilLimit. Gatekeeper is optimal for the case of O(1) attack edges and admits only O(1) Sybil identities (with high probability) in a random expander social networks (real-world social networks exhibit expander properties). In the face of O(k) attack edges (for any k ∈ O(n/ log n)), Gatekeeper admits O(log k) Sybils per attack edge. This result provides a graceful continuum across the spectrum of attack edges. We demonstrate the effectiveness of Gatekeeper experimentally on real-world social networks and synthetic topologies. Dinh Nguyen Tran, Jinyang Li 0001, Lakshminarayanan Subramanian, Sherman S. M. Chow |
INFOCOM | 3 |
| 2011 | WiRE: a new rural connectivity paradigmabstractMany rural areas in developing regions remain largely disconnected from the rest of the world due to low purchasing power and the exorbitant cost of existing connectivity solutions. Wireless Rural Extensions (WiRE) is a low-power rural wireless network architecture that provides inexpensive, self-sustainable, and high-bandwidth connectivity. WiRE relies on a high-bandwidth directional wireless backbone with local distribution networks to provide focused IP coverage. WiRE also provides cellular connectivity using OpenBTS-based GSM microcells. It supports a naming and addressing framework that inter-operates with traditional telecom networks and enables a wide range of mobile services on a common IP framework. The entire name network can be built by integrating a range of off-the-shelf components and existing open source tools. Aditya Dhananjay, Matt Tierney, Jinyang Li 0001, Lakshminarayanan Subramanian |
SIGCOMM | 4 |
| 2011 | TCP behavior in sub packet regimesabstractMany network links in developing regions operate in the sub-packet regime, an environment where the typical per-flow throughput is less than 1 packet per round-trip time. TCP and other common congestion control protocols break down in the sub-packet regime, resulting in severe unfairness, high packet loss rates, and flow silences due to repetitive timeouts. To understand TCP's behavior in this regime, we propose a model particularly tailored to high packet loss-rates and relatively small congestion window sizes. We validate the model under a variety of network conditions. Jay Chen, Janardhan R. Iyengar, Lakshminarayanan Subramanian, Bryan Ford |
SIGMETRICS | 3 |
| 2010 | SIMbaLink: towards a sustainable and feasible solar rural electrification systemabstractRural areas lack sustainable electrification solutions. Although solar solutions hold promise, they are fundamentally constrained by high maintenance costs (due to low user densities, equipment failure, poor handling) and a complete lack of accountability. In this paper, we describe our experiences deploying more than 5,000 Solar Home Systems in Ethiopia and the sustain-ability problems we faced. Towards developing a decentralized and sustainable solar solution, we have designed SIMbaLink, an extremely low-cost real-time solar monitoring system that significantly reduces both the maintenance costs and the time to repair. By explicitly exposing the real-time status of a solar system to all parties concerned, SIMbaLink addresses the lack of accountability and trust concerns. SIMbaLink can be easily integrated with existing solar systems and can reduce equipment failure rates through early detection of system malfunctions. Nahana Schelling, Meredith J. Hasson, Sara Leeun Huong, Ariel Nevarez, Wei-Chih Lu, Matt Tierney, Lakshminarayanan Subramanian, Harald Schützeichel |
ICTD | 7 |
| 2010 | SMS-based web search for low-end mobile devicesabstractShort Messaging Service (SMS) based mobile information services have become increasingly common around the world, especially in emerging regions among users with low-end mobile devices. This paper presents the design and implementation of SMSFind, an SMS-based search system that enables users to obtain extremely concise (one SMS) message of 140 bytes) and appropriate search responses for queries across arbitrary topics in one round of interaction. SMSFind is designed to complement existing SMS-based search services that are either limited in the topics they recognize or involve a human in the loop. Jay Chen, Lakshminarayanan Subramanian, Eric A. Brewer |
MobiCom | 2 |
| 2010 | Hermes: data transmission over unknown voice channelsabstractWhile the cellular revolution has made voice connectivity ubiquitous in the developing world, data services are largely absent or are prohibitively expensive. In this paper, we present Hermes1, a point-to-point data connectivity solution that works by modulating data onto acoustic signals that are sent over a cellular voice call. The main challenge is that most voice codecs greatly distort signals that are not voice-like; furthermore, the backhaul can be highly heterogeneous and of low quality, thereby introducing unpredictable distortions. Hermes modulates data over the extremely narrow-band approximately 3kHz bandwidth) acoustic carrier, while being severely constrained by the requirement that the resulting sound signals are voice-like, as far as the voice codecs are concerned. Hermes uses a robust data transcoding and modulation scheme to detect and correct errors in the face of bit flips, insertions and deletions; it also adapts the modulation parameters to the observed bit error rate on the actual voice channel. Through real-world experiments, we show that Hermes achieves approximately 1.2 kbps goodput which when compared to SMS, improves throughput by a factor of 5× and reduces the cost-per-byte by over a factor of 50x Aditya Dhananjay, Ashlesh Sharma, Michael Paik, Jay Chen, Trishank Karthik Kuppusamy, Jinyang Li 0001, Lakshminarayanan Subramanian |
MobiCom | 7 |
| 2010 | Brief announcement: improving social-network-based sybil-resilient node admission controlabstractWe present Gatekeeper, a decentralized protocol that performs Sybil-resilient node admission control based on a social network. Gatekeeper can admit most honest nodes while limiting the number of Sybils admitted per attack edge to O(log k), where k is the number of attack edges. Our result improves over SybilLimit [3] by a factor of log n in the face of O(1) attack edges. Even when the number of attack edges reaches O(n/ log n), Gatekeeper only admits O(log n) Sybils per attack edge, similar to that achieved by SybilLimit. Dinh Nguyen Tran, Jinyang Li 0001, Lakshminarayanan Subramanian, Sherman S. M. Chow |
PODC | 3 |
| 2009 | Web search over low bandwidthabstractWeb search and browsing have been streamlined for a comfortable experience when the network connection is fast. Existing systems, however, are not optimized for scenarios where connectivity is poor, as is the case for many users in developing regions where fast connections are expensive, rare, or unavailable. We examined the challenges faced by users behind one of these connections in a previous study, and found that the existing web interface is incapable of providing a good experience when the connection is extremely slow. In this demonstration we present a prototype implementation of a web browsing system that incorporates what we learned. Our system helps the users search offline as much as possible, and in a single search query when information from the Internet is required. Jay Chen, Lakshminarayanan Subramanian, Jinyang Li 0001 |
ICTD | 2 |
| 2009 | ATMosphere: A system for ATM microdeposit services in rural contextsabstractThis paper describes strategies to lower the cost of providing Automated Teller Machine microdeposit services in rural contexts. Microdeposits represent a growing market in the developing world, but the cost of running a conventional ATM network is prohibitive due to the capital investment required to deploy networks and terminals. Michael Paik, Lakshminarayanan Subramanian |
ICTD | 2 |
| 2009 | The case for SmartTrackabstractNearly 40 million people in Africa suffer from HIV/AIDS. African governments and international aid agencies have been working to combat this epidemic by vigorously promoting Highly Active Anti-Retroviral Therapy (HAART) programs. Despite the enormous subsidies offered by governments along with free Anti-RetroViral (ARV) drugs supplied by agencies, the introduction and implementation of HAART programs on a large scale has been limited by two fundamental problems: (a) lack of adherence to the ARV therapy regimen; (b) lack of accountability in drug distribution due to theft, corruption and counterfeit medication. In this paper, we motivate the case for SmartTrack, a telehealth project which aims to address these two problems facing HAART programs. The goal of SmartTrack is to create a highly reliable, secure and ultra low-cost cellphone-based distributed drug information system that can be used for tracking the flow and consumption of ARV drugs in HAART programs. In this paper, we assess the potential benefit of SmartTrack using a detailed needs-assessment study performed in Ghana, using interviews with 516 HIV-positive rural patients in a number of locations across the country. We find that a system like SmartTrack would immensely benefit both patients and healthcare providers, and can ultimately lead to improved patient outcomes and better accountability. Michael Paik, Ashlesh Sharma, Arthur Meacham, Giulio Quarta, John Trahanas, Brian Levine, Mary Ann Hopkins, Barbara Rapchak, Lakshminarayanan Subramanian |
ICTD | 10 |
| 2009 | Two-Party Computation Model for Privacy-Preserving Queries over Distributed Databases
Sherman S. M. Chow, Jie-Han Lee, Lakshminarayanan Subramanian |
NDSS | 3 |
| 2009 | Sybil-Resilient Online Content Voting
Dinh Nguyen Tran, Bonan Min, Jinyang Li 0001, Lakshminarayanan Subramanian |
NSDI | 4 |
| 2009 | Practical, distributed channel assignment and routing in dual-radio mesh networksabstractRealizing the full potential of a multi-radio mesh network involves two main challenges: how to assign channels to radios at each node to minimize interference and how to choose high throughput routing paths in the face of lossy links, variable channel conditions and external load. This paper presents ROMA, a practical, distributed channel assignment and routing protocol that achieves good multi-hop path performance between every node and one or more designated gateway nodes in a dual-radio network. ROMA assigns non-overlapping channels to links along each gateway path to eliminate intra-path interference. ROMA reduces inter-path interference by assigning different channels to paths destined for different gateways whenever possible. Evaluations on a 24-node dual-radio testbed show that ROMA achieves high throughput in a variety of scenarios. Aditya Dhananjay, Hui Zhang 0013, Jinyang Li 0001, Lakshminarayanan Subramanian |
SIGCOMM | 4 |
| 2009 | ELMR: lightweight mobile health recordsabstractCell phones are increasingly being used as common clients for a wide suite of distributed, database-centric healthcare applications in developing regions. This is particularly true for rural developing regions where the bulk of the healthcare is handled by health workers due to lack of doctors; the widespread availability of cellular services have made mobile devices as an important computing platform for enabling healthcare applications for these health workers. Unfortunately, the current SQL model for distributed client/server systems is far too heavy-weight for these applications, particularly in light of the high communications cost and extremely limited data transmission capacity available in these environments. Amey Purandare, Jay Chen, Arthur Meacham, Lakshminarayanan Subramanian |
SIGMOD Conference | 5 |
| 2009 | Economics of Rural Wireless Networks
Lakshminarayanan Subramanian |
VTC Fall | 1 |
| 2009 | RuralCafe: web search in the rural developing worldabstractThe majority of people in rural developing regions do not have access to the World Wide Web. Traditional network connectivity technologies have proven to be prohibitively expensive in these areas. The emergence of new long-range wireless technologies provide hope for connecting these rural regions to the Internet. However, the network connectivity provided by these new solutions are by nature intermittent due to high network usage rates, frequent power-cuts and the use of delay tolerant links. Typical applications, especially interactive applications like web search, do not tolerate intermittent connectivity. In this paper, we present the design and implementation of RuralCafe, a system intended to support efficient web search over intermittent networks. RuralCafe enables users to perform web search asynchronously and find what they are looking for in one round of intermittency as opposed to multiple rounds of search/downloads. RuralCafe does this by providing an expanded search query interface which allows a user to specify additional query terms to maximize the utility of the results returned by a search query. Given knowledge of the limited available network resources, RuralCafe performs optimizations to prefetch pages to best satisfy a search query based on a user's search preferences. In addition, RuralCafe does not require modifications to the web browser, and can provide single round search results tailored to various types of networks and economic constraints. We have implemented and evaluated the effectiveness of RuralCafe using queries from logs made to a large search engine, queries made by users in an intermittent setting, and live queries from a small testbed deployment. We have also deployed a prototype of RuralCafe in Kerala, India. Jay Chen, Lakshminarayanan Subramanian, Jinyang Li 0001 |
WWW | 2 |
| 2008 | An adaptive, high performance mac for long-distance multihop wireless networksabstractWe consider the problem of efficientMAC design for long-distance WiFi-based mesh networks. In such networks it is common to find long propagation delays, the use of directional antennas, and the presence of inter-link interference. Prior work has shown that these characteristics make traditional CSMA-based MACs a poor choice for long-distance mesh networks and this finding has led to several recent research efforts exploring the use of TDMA-based approaches to media access. In this paper we first identify, and then address, several shortcomings of current TDMA-based proposals. First, because they use fixed-length transmission slots, current TDMA-based solutions do not adapt to dynamic variations in traffic load leading to inefficiencies in both throughput and delay. As we show in this paper, the throughput achieved by existing solutions falls far short of the optimal achievable network throughput. Finally, due to the scheduling constraints imposed by inter-link interference, current TDMA-based solutions only apply to bipartite network topologies. Sergiu Nedevschi, Rabin K. Patra, Sonesh Surana, Sylvia Ratnasamy, Lakshminarayanan Subramanian, Eric A. Brewer |
MobiCom | 5 |
| 2008 | Beyond Pilots: Keeping Rural Wireless Networks Alive
Sonesh Surana, Rabin K. Patra, Sergiu Nedevschi, Manuel Ramos, Lakshminarayanan Subramanian, Yahel Ben-David, Eric A. Brewer |
NSDI | 5 |
| 2008 | One more bit is enough
Yong Xia 0007, Lakshminarayanan Subramanian, Ion Stoica, Shivkumar Kalyanaraman |
IEEE/ACM Trans. Netw. | 2 |
| 2007 | Packet Loss Characterization in WiFi-Based Long Distance NetworksabstractDespite the increasing number of WiFi-based Long Distance (WiLD) network deployments, there is a lack of understanding of how WiLD networks perform in practice. In this paper, we perform a systematic study to investigate the commonly cited sources of packet loss induced by the wireless channel and by the 802.11 MAC protocol. The channel induced losses that we study are external WiFi, non-WiFi and multipath interference. The protocol induced losses that we study are protocol timeouts and the breakdown of CSMA over WiLD links. Our results are based on measurements performed on two real-world WiLD deployments and a wireless channel emulator. The two deployments allow us to compare measurements across rural and urban settings. The channel emulator allows us to study each source of packet loss in isolation in a controlled environment. Based on our experiments we observe that the presence of external WiFi interference leads to significant amount of packet loss in WiLD links. In addition to identifying the sources of packet loss, we analyze the loss variability across time. We also explore the solution space and propose a range of MAC and network layer adaptation algorithms to mitigate the channel and protocol induced losses. The key lessons from this study were also used in the design of a TDMA based MAC protocol for high performance long distance multihop wireless networks [12]. Anmol Sheth, Sergiu Nedevschi, Rabin K. Patra, Sonesh Surana, Eric A. Brewer, Lakshminarayanan Subramanian |
INFOCOM | 6 |
| 2007 | WiLDNet: Design and Implementation of High Performance WiFi Based Long Distance Networks
Rabin K. Patra, Sergiu Nedevschi, Sonesh Surana, Anmol Sheth, Lakshminarayanan Subramanian, Eric A. Brewer |
NSDI | 5 |
| 2006 | Rethinking Wireless in the Developing World
Lakshminarayanan Subramanian, Sonesh Surana, Rabin K. Patra, Sergiu Nedevschi, Melissa Densmore, Eric A. Brewer, Anmol Sheth |
HotNets | 1 |
| 2005 | Reliable broadcast in unknown fixed-identity networksabstractIn this paper, we formulate a new theoretical problem, namely the reliable broadcast problem in unknown fixed-identity networks. This problem arises in the context of developing decentralized security mechanisms in a specific-class of distributed systems: Consider an undirected graph G connecting n nodes where each node is aware of only its neighbors but not of the entire graph. Additionally, each node has a unique identity and cannot fake its identity to its neighbors. Assume that k among the n nodes act in an adversarial manner and the remaining n-k are good nodes. Under what constraints does there exist a distributed algorithm Γ that enables every good node v to reliably broadcast a message m(v) to all other good nodes in G? While good nodes follow the algorithm Γ, an adversary can additionally discard messages, generate spurious messages or collude with other adversaries.In this paper, we prove two results on this problem. First, we provide a distributed algorithm Γ that can achieve reliable broadcast in an unknown fixed-identity network in the presence of k adversaries if G is 2k+1 vertex connected. Additionally, a minimum vertex connectivity of 2k+1 is a necessary condition for achieving reliable broadcast. Next, we study the problem of reliable broadcast in sparse networks (1-connected and 2-connected) in the presence of a single adversary i.e. k=1. In sparse networks, we show that a single adversary can partition the good nodes into groups such that nodes within a group can reliably broadcast to each other but nodes across groups cannot. For 1-connected and 2-connected graphs, we prove lower bounds on the number of such groups and provide a distributed algorithm to achieve these lower bounds. We also show that in a power-law random graph G(n,α), a single adversary can partition at most O(n1/α x (log n)(5- α)/(3-α)) good nodes from the remaining set of good nodes.Addressing this problem has practical implications to two real-world problems of paramount importance: (a) developing decentralized security measures to protect Internet routing against adversaries; (b) achieving decentralized public key distribution in static networks. Prior works on Byzantine agreement [17, 11, 23, 13, 3, 4, 24] are not applicable for this problem since they assume that either G is known, or that every pair of nodes can directly communicate, or that nodes use a key distribution infrastructure to sign messages. A solution to our problem can be extended to solve the byzantine agreement problem in unknown fixed-identity networks. Lakshminarayanan Subramanian, Randy H. Katz, Volker Roth 0002, Scott Shenker, Ion Stoica |
PODC | 1 |
| 2005 | HLP: a next generation inter-domain routing protocolabstractIt is well-known that BGP, the current inter-domain routing protocol, has many deficiencies. This paper describes a hybrid link-state and path-vector protocol called HLP as an alternative to BGP that has vastly better scalability, isolation and convergence properties. Using current BGP routing information, we show that HLP, in comparison to BGP, can reduce the churn-rate of route updates by a factor 400 as well as isolate the effect of routing events to a region 100 times smaller than that of BGP. For a majority of Internet routes, HLP guarantees worst-case linear-time convergence. We also describe a prototype implementation of HLP on top of the XORP router platform. HLP is not intended to be a finished and final proposal for a replacement for BGP, but is instead offered as a starting point for debates about the nature of the next-generation inter-domain routing protocol. Lakshminarayanan Subramanian, Matthew Caesar 0001, Cheng Tien Ee, Mark Handley, Z. Morley Mao, Scott Shenker, Ion Stoica |
SIGCOMM | 1 |
| 2005 | One more bit is enoughabstractAchieving efficient and fair bandwidth allocation while minimizing packet loss in high bandwidth-delay product networks has long been a daunting challenge. Existing end-to-end congestion control (eg TCP) and traditional congestion notification schemes (eg TCP+AQM/ECN) have significant limitations in achieving this goal. While the recently proposed XCP protocol addresses this challenge, XCP requires multiple bits to encode the congestion-related information exchanged between routers and end-hosts. Unfortunately, there is no space in the IP header for these bits, and solving this problem involves a non-trivial and time-consuming standardization process.In this paper, we design and implement a simple, low-complexity protocol, called Variable-structure congestion Control Protocol (VCP), that leverages only the existing two ECN bits for network congestion feedback, and yet achieves comparable performance to XCP, ie high utilization, low persistent queue length, negligible packet loss rate, and reasonable fairness. On the downside, VCP converges significantly slower to a fair allocation than XCP. We evaluate the performance of VCP using extensive ns2 simulations over a wide range of network scenarios. To gain insight into the behavior of VCP, we analyze a simple fluid model, and prove a global stability result for the case of a single bottleneck link shared by flows with identical round-trip times. Yong Xia 0007, Lakshminarayanan Subramanian, Ion Stoica, Shivkumar Kalyanaraman |
SIGCOMM | 2 |
| 2004 | Listen and Whisper: Security Mechanisms for BGP (Awarded Best Student Paper!)
Lakshminarayanan Subramanian, Volker Roth 0002, Ion Stoica, Scott Shenker, Randy H. Katz |
NSDI | 1 |
| 2004 | OverQoS: An Overlay Based Architecture for Enhancing Internet QoS
Lakshminarayanan Subramanian, Ion Stoica, Hari Balakrishnan, Randy H. Katz |
NSDI | 1 |
| 2002 | Characterizing the Internet Hierarchy from Multiple Vantage PointsabstractThe delivery of IP traffic through the Internet depends on the complex interactions between thousands of autonomous systems (AS) that exchange routing information using the border gateway protocol (BGP). This paper investigates the topological structure of the Internet in terms of customer-provider and peer-peer relationships between autonomous systems, as manifested in BGP routing policies. We describe a technique for inferring AS relationships by exploiting partial views of the AS graph available from different vantage points. Next we apply the technique to a collection of ten BGP routing tables to infer the relationships between neighboring autonomous systems. Based on these results, we analyze the hierarchical structure of the Internet and propose a five-level classification of AS. Our characterization differs from previous studies by focusing on the commercial relationships between autonomous systems rather than simply the connectivity between the nodes. Lakshminarayanan Subramanian, Sharad Agarwal, Jennifer Rexford, Randy H. Katz |
INFOCOM | 1 |
| 2002 | Geographic Properties of Internet Routing
Lakshminarayanan Subramanian, Venkat N. Padmanabhan, Randy H. Katz |
USENIX ATC, General Track | 1 |
| 2001 | An investigation of geographic mapping techniques for internet hostsabstractIn this paper, we ask whether it is possible to build an IP address to geographic location mapping service for Internet hosts. Such a service would enable a large and interesting class of location-aware applications. This is a challenging problem because an IP address does not inherently contain an indication of location.We present and evaluate three distinct techniques, collectively referred to as IP2Geo, for determining the geographic location of Internet hosts. The first technique, Geo Track, infers location based on the DNS names of the target host or other nearby network nodes. The second technique, GeoPing, uses network delay measurements from geographically distributed locations to deduce the coordinates of the target host. The third technique, GeoCluster, combines partial (and possibly inaccurate) host-to-location mapping information and BGP prefix information to infer the location of the target host. Using extensive and varied data sets, we evaluate the performance of these techniques and identify fundamental challenges in deducing geographic location from the IP address of an Internet host. Venkat N. Padmanabhan, Lakshminarayanan Subramanian |
SIGCOMM | 2 |
| 2000 | An architecture for building self-configurable systemsabstractDeveloping wireless sensor networks can enable information gathering, information processing and reliable monitoring of a variety of environments for both civil and military applications. It is however necessary to agree upon a basic architecture for building sensor network applications. This paper presents a general classification of sensor network applications based on their network configurations and discusses some of their architectural requirements. We propose a generic architecture for a specific subclass of sensor applications which we define as self-configurable systems where a large number of sensors coordinate amongst themselves to achieve a large sensing task. Throughout this paper we assume a certain subset of the sensors to be immobile. This paper lists the general architectural and infra-structural components necessary for building this class of sensor applications. Given the various architectural components, we present an algorithm that self-organizes the sensors into a network in a transparent manner. Some of the basic goals of our algorithm include minimizing power utilization, localizing operations and tolerating node and link failures. Lakshminarayanan Subramanian, Randy H. Katz |
MobiHoc | 1 |