VLDB 2026 Research / reviewers in the wild / expert
Lakshmish Ramaswamy
dblp:r/LakshmishRamaswamy · also Lakshmish Macheeri Ramaswamy
· DBLP profile ↗
70ranked-venue papers
18as first author
17since 2021 · last 2026
0000-0002-4567-4186ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 15 · 7 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 13 · 3 first-authorSystems, architecture and hardware · 8 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 2 since 2021Computer networks · 6 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 5Security and privacy · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | $\varDelta $-NeRF: Incremental Refinement of Neural Radiance Fields Through Residual Control and Knowledge Transfer
Kriti Ghosh, Devjyoti Chakraborty, Lakshmish Ramaswamy, Suchendra M. Bhandarkar, In Kee Kim, Nancy O'Hare, Deepak Mishra 0005 |
ICPR (9) | 3 |
| 2025 | CDAC: Content-Driven Access Control Architecture for Smart Farms
Ghadeer I. Yassin, Lakshmish Ramaswamy |
ICISSP (1) | 2 |
| 2024 | F4D: Factorized 4D Convolutional Neural Network for Efficient Video-Level Representation Learning
Mohammad Al-Saad, Lakshmish Ramaswamy, Suchendra M. Bhandarkar |
ICAART (3) | 2 |
| 2024 | Cost-Aware TrE-ND: Tri-embed Noise Detection for Enhancing Data Quality of Knowledge Graph
Jumana Alsubhi, Abdulrahman Ahmed Gharawi, Lakshmish Ramaswamy |
ICAART (3) | 3 |
| 2024 | JS-Siamese: Generalized Zero Shot Learning for IMU-based Human Activity Recognition
Mohammad Al-Saad, Lakshmish Ramaswamy, Suchendra M. Bhandarkar |
ICPR (15) | 2 |
| 2024 | An Empirical Evaluation of the Impact of Solar Correction in NeRFs for Satellite Imagery
Devjyoti Chakraborty, Kriti Ghosh, Zaki Sukma, In Kee Kim, Lakshmish Ramaswamy, Suchendra M. Bhandarkar, Deepak Mishra 0005 |
ICPR (18) | 5 |
| 2024 | CroMA: Enhancing Fault-Resilience of Machine Learning-Coupled IoT ApplicationsabstractMachine learning (ML)-coupled Internet of Things (IoT) applications are becoming increasingly popular in many domains. However, faulty sensor readings pose a significant challenge to IoT-ML applications. Noise is inherent in IoT environments due to issues such as harsh IoT environments and hardware limitations. Unfortunately, the performance of the ML applications quickly degrades when faced with faulty sensor readings. Most current techniques to handle faulty sensor readings require many models, and they attempt to correct faulty data using the entire sample containing both readings of non-faulty and faulty sensors, making them unsuitable for edge ML-IoT applications, where real-time processing and low latency are critical. Unfortunately, these techniques fail, especially when multiple sensors' readings are concurrently faulty (due to simultaneous sensor failures). With the goal of building robust edge IoT-coupled ML applications, this paper proposes CroMA - a proactive approach for overcoming simultaneous failure of sensors. CroMA is unique in that it is based upon a masked autoencoder trained on several randomly masked portions of training samples, corresponding to sensors' readings. This enables the autoencoder to learn to predict sensors' readings through the unmasked portions. CroMA incorporates a novel technique to optimize its prediction accuracy by leveraging sensor correlations-based masking and identically treats all fault types as one type via masking. CroMA's network design includes gated recurrent units (GRUs) to capture long-short-term dependencies and correlations between sensors' readings. We confirm the effectiveness of the CroMA through a series of experiments involving three distinct datasets. The experimental results affirm that CroMA effectively corrects sensor faults, thereby preserving the classification accuracy of the IoT ML-based application in the face of sensor failures. Yousef AlShehri, Lakshmish Ramaswamy |
SEC | 2 |
| 2023 | Modular Deep Learning for Big Data: Motivation, Challenges, and ApproachesabstractDeep learning (DL) technologies have the potential to fundamentally transform many Big Data domains ranging from agriculture and healthcare to transportation and entertainment. However, a large percentage of “non-tech” sectors have been unable to fully harness the power of DL technologies because of the lack of expertise, resources and training data for developing, tuning and managing DL models. The monolithic model paradigm that dominates the current DL landscape has a number of drawbacks that have impeded wider adoption of DL in various Big Data domains. This paper presents an alternate vision, namely modular DL paradigm. Semantically meaningful DL modules form the basic building blocks of this paradigm, and specific DL tasks are accomplished by creating pipelines of these modules. This paper motivates the need for modular DL through a case study that includes an experimental analysis examine the trade-offs between modular and monolithic DL paradigms. We discuss the challenges in designing a modular DL framework for robust and cost-effective Big Data analytics, and outline approaches to address the challenges. Samiyuru Menik, Lakshmish Ramaswamy |
IEEE Big Data | 2 |
| 2023 | Cost-Aware Ensemble Learning Approach for Overcoming Noise in Labeled Data
Abdulrahman Ahmed Gharawi, Jumana Alsubhi, Lakshmish Ramaswamy |
ICAART (3) | 3 |
| 2023 | Continual Optimization of In-Production Machine Learning Systems Through Semantic Analysis of User Feedback
Hemadri Jayalath, Ghadeer I. Yassin, Lakshmish Ramaswamy, Sheng Li 0001 |
ICAART (3) | 3 |
| 2023 | DynaES: Dynamic Energy Scheduling for Energy Harvesting Environmental SensorsabstractLow-cost sensors and IoT technologies have facilitated the deployment of environmental sensors to collect and analyze various factors, such as soil properties. Due to the lack of electric power networks in many deployment locations, these sensors rely on energy harvesting (EH) systems that use natural energy sources such as solar power. Specifically, solar-powered EH systems benefit from timely obtaining weather information for efficient future energy scheduling. However, obtaining weather information for EH environmental sensors is always challenging, as they are commonly deployed in harsh environments without network access. To address this problem, we present DynaES, a novel energy scheduling method for EH sensors without relying on online weather forecasts. DynaES comprises two components: a DC power gain estimator that predicts future power gain by individually estimating changes in environmental parameters and ensembling them, and a dynamic energy scheduler that distributes energy to sensors based on priority and adjusts sensing intervals and frequency. We evaluate DynaES via simulation-based studies on real-world datasets and compare its performance against state-of-the-art baselines. Evaluation results show that DynaES accurately predicts future energy gain with low estimation errors and enables $1.8 \times$ – $4 \times$ more frequent sensing operations with shorter sensing intervals while achieving longer operation hours without complete battery drains. Jianwei Hao, Emmanuel Oni, In Kee Kim, Lakshmish Ramaswamy |
IPCCC | 4 |
| 2023 | Reaching for the Sky: Maximizing Deep Learning Inference Throughput on Edge Devices with AI Multi-TenancyabstractThe wide adoption of smart devices and Internet-of-Things (IoT) sensors has led to massive growth in data generation at the edge of the Internet over the past decade. Intelligent real-time analysis of such a high volume of data, particularly leveraging highly accurate deep learning (DL) models, often requires the data to be processed as close to the data sources (or at the edge of the Internet) to minimize the network and processing latency. The advent of specialized, low-cost, and power-efficient edge devices has greatly facilitated DL inference tasks at the edge. However, limited research has been done to improve the inference throughput (e.g., number of inferences per second) by exploiting various system techniques. This study investigates system techniques, such as batched inferencing, AI multi-tenancy, and cluster of AI accelerators, which can significantly enhance the overall inference throughput on edge devices with DL models for image classification tasks. In particular, AI multi-tenancy enables collective utilization of edge devices’ system resources (CPU, GPU) and AI accelerators (e.g., Edge Tensor Processing Units; EdgeTPUs). The evaluation results show that batched inferencing results in more than 2.4× throughput improvement on devices equipped with high-performance GPUs like Jetson Xavier NX. Moreover, with multi-tenancy approaches, e.g., concurrent model executions (CME) and dynamic model placements (DMP), the DL inference throughput on edge devices (with GPUs) and EdgeTPU can be further improved by up to 3× and 10×, respectively. Furthermore, we present a detailed analysis of hardware and software factors that change the DL inference throughput on edge devices and EdgeTPUs, thereby shedding light on areas that could be further improved to achieve high-performance DL inference at the edge. Jianwei Hao, Piyush Subedi, Lakshmish Ramaswamy, In Kee Kim |
ACM Trans. Internet Techn. | 3 |
| 2022 | Effective & Efficient Access Control in Smart Farms: Opportunities, Challenges & Potential Approaches
Ghadeer I. Yassin, Lakshmish Ramaswamy |
ICISSP | 2 |
| 2022 | SECOE: Alleviating Sensors Failure in Machine Learning-Coupled IoT SystemsabstractInternet of Things (IoT) domains are characterized by continuous streams of data originating from diverse, geographically distributed sensors. Sensor/network failures that result in data stream interruptions is a major challenge in applying ML techniques to IoT domains. Unfortunately, the performance of many ML applications quickly degrades when faced with data incompleteness. With the aim of building robust IoT-coupled ML applications, this paper proposes SECOE – a unique, proactive approach for alleviating potentially simultaneous sensor failures. The fundamental idea behind SECOE is to create a carefully chosen ensemble of ML models in which each model is trained assuming a set of failed sensors. SECOE includes a novel technique to minimize the number of models in the ensemble by harnessing the correlations among sensors. We demonstrate the efficacy of the SECOE approach through a series of experiments involving two distinct datasets. Yousef AlShehri, Lakshmish Ramaswamy |
ICMLA | 2 |
| 2022 | A Layer Decomposition Approach to Inference Time Prediction of Deep Learning ArchitecturesabstractIn recent years, deep learning models have been widely adopted in lots of fields. such as computer vision, pattern recognition, and classification problems like plant disease classification. Due to the large diversity among the computing devices that these models may run on, we need to choose between the appropriate device based on cost and performance. Furthermore, finding the suitable optimal device for a given project is a complex process that needs significant time and resources. Prediction of inference latency DNN models is necessary for many tasks where measuring the latency on real devices is either infeasible or too costly. This is a very challenging problem, and most existing approaches fail to achieve high accuracy of prediction. While some research has been carried out to predict the inference time of DNN models – most existing techniques assume that training time is linearly related to the number of floating-point operations. This paper designs and develops a framework to predict the inference time for deep learning models and is generic to be easily extended for a large set of devices. Our key idea is decomposing a given model inference into layers and conducting layer-level prediction. Our experiments demonstrate that this strategy provides significant benefits in terms of prediction accuracy. Ola Mustafa Alqahtani, Lakshmish Ramaswamy |
ICMLA | 2 |
| 2022 | Impact of Labeling Noise on Machine Learning: A Cost-aware Empirical StudyabstractSince the emergence of large datasets, machine learning models have demonstrated excellent performance in a wide range of applications. This accomplishment was made possible by the availability of large amounts of labeled datasets. Finding high-quality labeled datasets, on the other hand, is difficult to obtain. Acquiring high-quality datasets with limited class label noise becomes an important task since noisy datasets can affect the performance and structure of machine learning models. However, it is extremely difficult to reduce label noise significantly in real-world datasets unless using expensive expert annotators. This work studies the influence of varying degrees of label noise on the complexity and accuracy of machine learning models, based on considerable testing and research. It also explores how to reduce labeling costs while maintaining the desired accuracy. Abdulrahman Ahmed Gharawi, Jumana Alsubhi, Lakshmish Ramaswamy |
ICMLA | 3 |
| 2021 | AI Multi-Tenancy on Edge: Concurrent Deep Learning Model Executions and Dynamic Model Placements on Edge DevicesabstractMany real-world applications are widely adopting the edge computing paradigm due to its low latency and better privacy protection. With notable success in AI and deep learning (DL), edge devices and AI accelerators play a crucial role in deploying DL inference services at the edge of the Internet. While prior works quantified various edge devices’ efficiency, most studies focused on the performance of edge devices with single DL tasks. Therefore, there is an urgent need to investigate AI multi-tenancy on edge devices, required by many advanced DL applications for edge computing. This work investigates two techniques – concurrent model executions and dynamic model placements – for AI multi-tenancy on edge devices. With image classification as an example scenario, we empirically evaluate AI multi-tenancy on various edge devices, AI accelerators, and DL frameworks to identify its benefits and limitations. Our results show that multi-tenancy significantly improves DL inference throughput by up to 3.3 × − 3.8 × on Jetson TX2. These AI multi-tenancy techniques also open up new opportunities for flexible deployment of multiple DL services on edge devices and AI accelerators. Piyush Subedi, Jianwei Hao, In Kee Kim, Lakshmish Ramaswamy |
CLOUD | 4 |
| 2019 | Extracting Topics from Semi-structured Data for Enhancing Enterprise Knowledge Graphs
Neda Abolhassani, Lakshmish Ramaswamy |
CollaborateCom | 2 |
| 2019 | A Mobile and Web-Based Approach for Targeted and Proactive Participatory Sensing
Navid Hashemi Tonekaboni, Lakshmish Ramaswamy, Sakshi Sachdev |
CollaborateCom | 2 |
| 2018 | A Multi-Cloud Cyber Infrastructure for Monitoring Global Proliferation of Cyanobacterial Harmful Algal BloomsabstractCyanobacterial Harmful Algal Blooms (CyanoHABs) are a major water quality and public health issue in inland waters as they hamper recreational activities, degrade aquatic habitats, and potentially affect human health via toxic contamination. Despite the risks posed to environment, human and animal health, currently, there is lack of rapid monitoring program to periodically evaluate the spatial distribution of cyanobacteria in inland waters. This study integrated multiple clouds including community cloud (via social media data), sensor cloud (wireless hyperspectral sensor and satellite sensor) and computational cloud to design and implement techniques for early detection of CyanoHABs in inland waters. Social cloud data helped to identify the geographical locations frequently affected by CyanoHABs and sensor clouds helped in verifying those locations. This integrated monitoring system would be very useful for lake resource managements and state agencies by reducing their budget cost for rapid detection and frequent monitoring of CyanoHABs across inland waters. Deepak Mishra 0005, Lakshmish Ramaswamy, Abhishek Kumar 0020, Suchendra M. Bhandarkar, Sunil Narumalani |
IGARSS | 2 |
| 2017 | Caching for Pattern Matching Queries in Time Evolving Graphs: Challenges and ApproachesabstractPattern matching is an important class of problems related to graphs. It is a fundamental problem for many applications and has been extensively studied in literature. With the advent of huge graphs, the challenges in this domain have increased manifold. Consequently a lot of recent research has led to new architectures and approaches for optimized solutions to the pattern matching problem. A vast majority of these graphs hardly remain static and constantly evolve over time (like social networks, web graphs, etc). Recently, caching has been studied in the context of static graphs to optimize the throughput of query processing systems. In this paper, we list the challenges in caching in the context of Time Evolving Graphs (TEGs). Amongst others, one major challenge is consistency which entails to making sure the cache is consistent with the streaming changes. We propose an approach to successfully implement caching that addresses those issues and based on the initial results, we see significant gains in the overall performance of system. M. Usman Nisar, Sahar Voghoei, Lakshmish Ramaswamy |
ICDCS | 3 |
| 2016 | TAU-FIVE: A Multi-tiered Architecture for Data Quality and Energy-Sustainability in Sensor NetworksabstractCurrent research on wireless sensor networks "WSNs" in the Internet of Things "IoT" has focused on performance, scalability and energy efficiency. Innovations in these areas have many challenges due to the increasing volume of smart device data streams in the internet of Everything "IoE". Data feeds from future IoE systems such as the internet of vehicles, smart homes and smart-cities will need real time consolidation. This merger of technologies will require innovative big data algorithms and architectures that authenticate the data streams. A primary concern is in dynamically quantifying the data quality "DQ" of the streams while constructing real-time metrics to assess the energy efficiency "EE" of these IoE devices. In order to define the relationship between sensor stream DQ and EE, we propose our multi-tiered cloud-service architecture TAU-FIVE. The technical contributions of our framework includes data quality and energy efficiency models based on 7 DQ attributes and multiple reprogrammable smart sensors that dynamically modify and regulate the DQ and EE of a WSN. Our research maintains that WSN's can balance sustainability with quality of service by creating real-time metrics that merge energy usage with data stream integrity. This equilibrium will impact energy awareness in the IoT as the multitude of batch device data streams are integrated with the variety of social and professional networks and evolve into the IoE. Victor Lawson, Lakshmish Ramaswamy |
DCOSS | 2 |
| 2016 | A Hierarchical Meta-Classifier for Human Activity RecognitionabstractThis paper proposes a multi-level meta-classifier for identifying human activities based on accelerometer data. The training data consists of 77 subjects performing a combination of 23 different activities and monitored using a single hip-worn triaxial accelerometer. Time and frequency based features were extracted from two-second windows of raw accelerometer data and a subset of the features, together with demographic information, was selected for classification. The activities were divided into five activity groups: non-ambulatory activities, walking, running, climbing upstairs, and climbing downstairs. Multiple classification techniques were tested for each classifier level and groups. Random forests were found to perform comparatively better at each level. Based upon those tests, a 3-level hierarchical classifier, consisting of 5 random forest classifiers, was built. At the first level, the non-ambulatory activities are separated from the rest. At the second, the ambulatory activities are divided into four activity groups. At the final level, the activities are classified individually. Accuracy on test sets was found to be approximately 87% overall for individual activities and 94% at the activity group level. These results compare favorably to contemporary results in classifying human activity. Anzah H. Niazi, Delaram Yazdansepas, Jennifer L. Gay, Frederick W. Maier, Lakshmish Ramaswamy, Khaled Rasheed, Matthew P. Buman |
ICMLA | 5 |
| 2015 | A rule-engine-based approach to a personalized OTC drug safety frameworkabstractIn this paper we present a rule-engine-based framework to support personalized and context-aware OTC medication safety advice. We describe how high-level architectural choices impact the system. We report on our prototype implementation and initial experiments. We also describe a means of representing drug safety rules in an appropriate language for automated decision making and some of the challenges associated with rule authorship. Michael D. Scott 0001, Lakshmish Ramaswamy, Sarabpreet Kaur Dhillon |
HealthCom | 2 |
| 2014 | Effective caching techniques for accelerating pattern matching queriesabstractUsing caching techniques to improve response time of queries is a proven approach in many contexts. However, it is not well explored for subgraph pattern matching queries, mainly because of subtleties enforced by traditional pattern matching models. Indeed, efficient caching can greatly impact the query answering performance for massive graphs in any query engine whether it is centralized or distributed. This paper investigates the capabilities of the newly introduced pattern matching models in graph simulation family for this purpose. We propose a novel caching technique, and show how the results of a query can be used to answer the new similar queries according to the similarity measure that is introduced. Using large real-world graphs, we experimentally verify the efficiency of the proposed technique in answering subgraph pattern matching queries. Arash Fard, Satya Manda, Lakshmish Ramaswamy, John A. Miller 0001 |
IEEE BigData | 3 |
| 2014 | Performance modeling of computation and communication tradeoffs in vertex-centric graph processing clustersabstractDistributed vertex-centric graph processing systems have been recently proposed to perform different types of analytics on large graphs. These systems utilize the parallelism of shared nothing clusters. In this work we propose a novel model for the performance cost of such clusters. We also define n Amir Abdolrashidi, Lakshmish Ramaswamy, David S. Narron |
CollaborateCom | 2 |
| 2014 | DQS-Cloud: A Data Quality-Aware autonomic cloud for sensor servicesabstractWith the advent of Internet of Things, the field of domain sensing is increasingly being servitized. In order to effectively support this servitization, there is a growing need for a powerful and easy-to-use infrastructure that enables seamless sharing of sensor data in real-time. In this paper, we Abhishek Kothari, Vinay Boddula, Lakshmish Ramaswamy, Neda Abolhassani |
CollaborateCom | 3 |
| 2014 | Editorial
Lakshmish Ramaswamy, Barbara Carminati, Lujo Bauer, Dongwan Shin, James B. D. Joshi, Calton Pu, Dimitris Gritzalis |
Comput. Secur. | 1 |
| 2014 | Preface
Barbara Carminati, Lakshmish Ramaswamy, Anna Cinzia Squicciarini, James B. D. Joshi, Calton Pu |
Int. J. Cooperative Inf. Syst. | 2 |
| 2014 | Collaborative caching for efficient dissemination of personalized video streams in resource constrained environments
Suchendra M. Bhandarkar, Lakshmish Ramaswamy, Hari Devulapally |
Multim. Syst. | 2 |
| 2014 | Editorial: Collaborative Computing: Networking, Applications and Worksharing (CollaborateCom 2012)
Lakshmish Ramaswamy, Barbara Carminati, James B. D. Joshi, Calton Pu |
Mob. Networks Appl. | 1 |
| 2013 | A distributed vertex-centric approach for pattern matching in massive graphsabstractGraph pattern matching is fundamentally important to many applications such as analyzing hyper-links in the World Wide Web, mining associations in online social networks, and substructure search in biochemistry. Most existing graph pattern matching algorithms are highly computation intensive, and do not scale to extremely large graphs that characterize many emerging applications. In recent years, graph processing frameworks such as Pregel have sought to harness shared nothing clusters for processing massive graphs through a vertex-centric, Bulk Synchronous Parallel (BSP) programming model. However, developing scalable and efficient BSP-based algorithms for pattern matching is very challenging because this problem does not naturally align with a vertex-centric programming paradigm. This paper presents novel distributed algorithms based on the vertex-centric programming paradigm for a set of pattern matching models, namely, graph simulation, dual simulation and strong simulation. Our algorithms are fine-tuned to consider the challenges of pattern matching on massive data graphs. Furthermore, we introduce a new pattern matching model, called strict simulation, which outperforms strong simulation in terms of scalability while preserving its important properties. We investigate potential performance bottlenecks and propose several techniques to mitigate them. This paper also presents an extensive set of experiments involving massive graphs (millions of vertices and billions of edges) to study the effects of various parameters on the scalability and performance of the proposed algorithms. The results demonstrate that our techniques are highly effective in alleviating performance bottlenecks and yield significant scalability benefits. Arash Fard, M. Usman Nisar, Lakshmish Ramaswamy, John A. Miller 0001, Matthew Saltz |
IEEE BigData | 3 |
| 2013 | SCISSOR: scalable and efficient reachability query processing in time-evolving hierarchiesabstractA time-evolving hierarchy (TEH) consists of multiple snapshots of the hierarchy (collection of one or more trees) as it evolves over time. It is often important to test reachability between a given pair of vertices in an arbitrary (possibly past) snapshot of the hierarchy. While interval-based indexing has been a popular strategy for reachability testing in static hierarchies, a straightforward extension of this strategy to TEHs is impractical because of the exorbitant indexing overheads. In this paper, we propose SCISSOR (selective snapshot indexing with progressive solution refinement), which, to the best of our knowledge is the first time and space efficient framework for answering reachability queries in TEHs. The main idea here is to maintain indexes only for a selective interspersed subset of TEH snapshots. A query on a non-indexed snapshot will be answered by utilizing the index of a temporally-nearby indexed snapshot and analyzing the structural changes that have occurred between the two snapshots. We also present a experimental study demonstrating the scalability and efficiency of the SCISSOR framework in terms of both indexing costs and query latencies. Phani Rohit Mullangi, Lakshmish Ramaswamy |
CIKM | 2 |
| 2013 | A content-context-centric approach for detecting vandalism in WikipediaabstractCollaborative online social media (CSM) applications such as Wikipedia have not only revolutionized the World Wide Web, but they also have had a hugely positive effect on modern free societies. Unfortunately, Wikipedia has also become target to a wide-variety of vandalism attacks. Most existing vand Lakshmish Ramaswamy, Raga Sowmya Tummalapenta, Kang Li 0001, Calton Pu |
CollaborateCom | 1 |
| 2012 | Preface
Barbara Carminati, Lakshmish Ramaswamy, Calton Pu, James B. D. Joshi |
CollaborateCom | 2 |
| 2012 | Towards efficient query processing on massive time-evolving graphsabstractTime evolving graph (TEG) is increasingly being used as a paradigm for modeling and analyzing dynamic relationships in many emerging domains such as online social networks, World Wide Web and evolutionary genomics. A time-evolving graph consists of a sequence of snapshots of the graph as it evolves Arash Fard, Amir Abdolrashidi, Lakshmish Ramaswamy, John A. Miller 0001 |
CollaborateCom | 3 |
| 2012 | Comet: Decentralized Complex Event Detection in Mobile Delay Tolerant NetworksabstractIncreased commodity use of mobile devices has the potential to enable mission-critical monitoring applications. However, these mobile-enabled monitoring applications have to often work in environments where a delay-tolerant network (DTN) is the only feasible communication paradigm. Detection of complex (composite) events is fundamental to monitoring applications. However, the existing plan-based CED techniques are mostly centralized, and hence are inherently unscalable for DTNs. In this paper, we create Comet â" a decentralized plan-based, efficient and scalable CED for DTNs. Comet shares the task of detecting complex events (CEs) among multiple nodes, with each node detecting a part of the CE by aggregating two or more primitive events or sub-CEs. Comet uses a unique h-function to construct cost and delay efficient CED trees. As finding an optimal CED plan requires exponential-time, Comet finds near-optimal detection plans for individual CEs through a novel multi-level push-pull conversion algorithm. Performance results show that Comet reduces cost by up to 89% compared to pushing all primitive events and over 60% compared to a two-level exhaustive search algorithm. Jianxia Chen, Lakshmish Ramaswamy, David K. Lowenthal, Shivkumar Kalyanaraman |
MDM | 2 |
| 2012 | DATEM: Towards QoS Aware Event Notification Framework for Delay Tolerant NetworksabstractDelay tolerant networks (DTNs), which are wireless networks prone to long delays and frequent disruptions are increasingly becoming common. We are exploring the challenges involved in designing effective event notification framework on DTNs. This paper presents the design and evaluation of a delay-tolerant event notification framework called DATEM based on the philosophy that failures and delays are too common to be effectively masked from applications. DATEM represents a comprehensive redesign of traditional event routing mechanisms and system management techniques. First, the framework exposes the delay- and failure-prone nature of the underlying network to the applications via an enhanced set of API and let them handle these parameters in an application specific manner. Second, it incorporates a novel event routing protocol that enables the nodes of the delay tolerant network to collaboratively achieve QoS requested by the applications. Third, a utility-aware event prioritization scheme is designed to prioritize event-packets that have the largest impact on the performance of DATEM framework. Our experimental study shows that DATEM provides significant performance benefits when compared with traditional event notification schemes. Phani Rohit Mullangi, Lakshmish Ramaswamy, Raga Sowmya Tummalapenta |
MDM | 2 |
| 2012 | Collaborative caching for efficient dissemination of personalized video streams in resource constrained environmentsabstractThe ever increasing deployment of broadband networks and simultaneous proliferation of low cost video capturing and multimedia enabled mobile devices have triggered a wave of novel mobile multimedia applications, resulting in the development of large scale systems for delivery of video streams to heterogeneous resource constrained mobile clients. Invariably, the video streams need to be personalized to provide a resource constrained mobile device with video content that is most relevant to the client's request while simultaneously satisfying the client-side and system-wide resource constraints. In this paper we present the design and implementation of a distributed system, consisting of several geographically distributed video personalization servers and proxy caches, for efficient dissemination of personalized video in a resource constrained mobile environment. With the objective of optimizing cache performance, a novel cache replacement policy and multi-stage client request aggregation strategy, both of which are specifically tailored for personalized video content, are proposed. A novel latency-biased collaborative caching protocol based on counting Bloom filters is designed for further enhancing the scalability and efficiency of disseminating personalized video content. The benefits and costs associated with collaborative caching for disseminating personalized video content to resource constrained and geographically distributed clients are analyzed and experimentally verified. The impact of different levels of collaboration amongst the caches and, the advantages of using multiple video personalization servers with varying degrees of mirrored content on the efficiency of personalized video delivery, are also studied. Experimental results demonstrate that the proposed collaborative caching scheme, coupled with the proposed personalization-aware cache replacement and client request aggregation strategies, provides a means for efficient dissemination of personalized video streams in resource constrained environments. Suchendra M. Bhandarkar, Lakshmish Ramaswamy, Hari Devulapally |
MMSys | 2 |
| 2011 | Video personalization in heterogeneous and resource-constrained environments
Suchendra M. Bhandarkar, Kang Li 0001, Lakshmish Ramaswamy |
Multim. Syst. | 4 |
| 2011 | The CoQUOS Approach to Continuous Queries in Unstructured OverlaysabstractThe current peer-to-peer (P2P) content distribution systems are constricted by their simple on-demand content discovery mechanism. The utility of these systems can be greatly enhanced by incorporating two capabilities, namely a mechanism through which peers can register their long term interests with the network so that they can be continuously notified of new data items, and a means for the peers to advertise their contents. Although researchers have proposed a few unstructured overlay-based publish-subscribe systems that provide the above capabilities, most of these systems require intricate indexing and routing schemes, which not only make them highly complex but also render the overlay network less flexible toward transient peers. This paper argues that for many P2P applications, implementing full-fledged publish-subscribe systems is an overkill. For these applications, we study the alternate continuous query paradigm, which is a best-effort service providing the above two capabilities. We present a scalable and effective middleware, called CoQUOS, for supporting continuous queries in unstructured overlay networks. Besides being independent of the overlay topology, CoQUOS preserves the simplicity and flexibility of the unstructured P2P network. Our design of the CoQUOS system is characterized by two novel techniques, namely cluster-resilient random walk algorithm for propagating the queries to various regions of the network and dynamic probability-based query registration scheme to ensure that the registrations are well distributed in the overlay. Further, we also develop effective and efficient schemes for providing resilience to the churn of the P2P network and for ensuring a fair distribution of the notification load among the peers. This paper studies the properties of our algorithms through theoretical analysis. We also report series of experiments evaluating the effectiveness and the costs of the proposed schemes. Lakshmish Ramaswamy, Jianxia Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2010 | Elusive vandalism detection in wikipedia: a text stability-based approachabstractThe open collaborative nature of wikis encourages participation of all users, but at the same time exposes their content to vandalism. The current vandalism-detection techniques, while effective against relatively obvious vandalism edits, prove to be inadequate in detecting increasingly prevalent sophisticated (or elusive) vandal edits. We identify a number of vandal edits that can take hours, even days, to correct and propose a text stability-based approach for detecting them. Our approach is focused on the likelihood of a certain part of an article being modified by a regular edit. In addition to text-stability, our machine learning-based technique also takes into account edit patterns. We evaluate the performance of our approach on a corpus comprising of 15000 manually labeled edits from the Wikipedia Vandalism PAN corpus. The experimental results show that text-stability is able to improve the performance of the selected machine-learning algorithms significantly. Qinyi Wu, Danesh Irani, Calton Pu, Lakshmish Ramaswamy |
CIKM | 4 |
| 2010 | CAEVA: A customizable and adaptive event aggregation framework for collaborative broker overlaysabstractThe publish-subscribe (pub-sub) paradigm is maturing and integrating into community-oriented collaborative applications. Because of this, pub-sub systems are faced with an event stream that may potentially contain large numbers of redundant and partial messages. Most pub-sub systems view partial and Jianxia Chen, Lakshmish Ramaswamy, David K. Lowenthal, Shivkumar Kalyanaraman |
CollaborateCom | 2 |
| 2010 | Enhancing Scalability and Performance of Mashups Through Merging and Operator ReorderingabstractRecently, mashups are gaining tremendous popularity as an important Web 2.0 application. Mashups provide end-users with an opportunity to create personalized Web services which aggregate and manipulate data from multiple diverse sources distributed across the Web. However, this increase in personalization also results in new scalability and performance challenges. Surprisingly, there are very few studies on the performance aspect of mashups. In this paper, we propose two novel techniques to enhance the scalability and performance of mashup platforms. The first is an efficient mashup merging scheme that avoids duplicate computations and unnecessary data retrievals by detecting common operator sequences in different mashups and executing them together. Second, we propose a canonical form-based mashup reordering scheme that not only transforms individual mashups to their most efficient forms but also increases the effectiveness of mashup merging. This paper also reports a number of experiments studying the benefits and costs of the proposed techniques. Osama Al-Haj Hassan, Lakshmish Ramaswamy, John A. Miller 0001 |
ICWS | 2 |
| 2009 | Efficient dissemination of personalized video content in resource-constrained environmentsabstractVideo streaming on mobile devices such as PDA's, laptop PCs, pocket PCs and cell phones is becoming increasingly popular. These mobile devices are typically constrained by their battery capacity, bandwidth, screen resolution and video decoding and rendering capabilities. Consequently, video personal Piyush Parate, Lakshmish Ramaswamy, Suchendra M. Bhandarkar, Siddhartha Chattopadhyay, Hari Devulapally |
CollaborateCom | 2 |
| 2009 | Rapid Identification Approach for Reusable SOA Assets Using Component Business MapsabstractSubstantial savings from asset reuse result when the right assets are identified in the very early stages of a client engagement. Unfortunately, advanced identification approaches (known by having high precision and recall, such as behavior-based approaches) cannot be adopted in these early stages, because at these early stages, there is no many details nor much understanding about the client functional requirements. On the other hand, unstructured keyword-based identification approaches are known of having low precision and recall. To overcome this problem, we argue that assets descriptions should have explicit information about the business activities realized by the assets. To be able to capture this information in a machine understandable format, this paper proposes a model for describing reusable assets functional scopes using component business maps (CBMs), in which the asset scope is represented as a hierarchy of CBM elements. Adopting this scope model, the paper proposes a rapid identification approach for reusable assets that retrieves assets based on their CBM projections. We believe the proposed approach provides better precision and recall when compared to unstructured keyword-based approaches. Islam Elgedawy, Lakshmish Ramaswamy |
ICWS | 2 |
| 2009 | MACE: A Dynamic Caching Framework for MashupsabstractThe recent surge of popularity has established mashups as an important category of Web 2.0 applications. Mashups are essentially Web services that are often created by end-users. They aggregate and manipulate data from sources around the World Wide Web. Surprisingly, there are very few studies on the scalability and performance of mashups. In this paper, we study caching as a vehicle for enhancing the scalability and the efficiency of mashups. Although caching has long been used to improve the performance of Web services, mashups pose some unique challenges that necessitate a more dynamic approach to caching. Towards this end, we present MACE - a cache specifically designed for mashups. In designing the MACE framework this paper makes three technical contributions. First, we present a model for representing mashups and analyzing their performance. Second, we propose an indexing scheme that enables efficient reuse of cached data for newly created mashups. Finally, this paper also describes a novel caching policy that analyzes the costs and benefits of caching data at various stages of different mashups and selectively stores data that is most effective in improving system scalability. We report experiments studying the performance of the MACE system. Osama Al-Haj Hassan, Lakshmish Ramaswamy, John A. Miller 0001 |
ICWS | 2 |
| 2009 | CAESAR: A Context-Aware, Social Recommender System for Low-End Mobile DevicesabstractMobile-enabled social networks applications are becoming increasingly popular. Most of the current social network applications have been designed for high-end mobile devices, and they rely upon features such as GPS, capabilities of the world wide web, and rich media support. However, a significant fraction of mobile user base, especially in the developing world, own low-end devices that are only capable of voice and short text messages (SMS). In this context, a natural question is whether one can design meaningful social network-based applications that can work well with these simple devices, and if so, what the real challenges are. Towards answering these questions, this paper presents a social network-based recommender system that has been explicitly designed to work even with devices that just support phone calls and SMS. Our design of the social network based recommender system incorporates three features that complement each other to derive highly targeted ads. First, we analyze information such as customer's address books to estimate the level of social affinity among various users. This social affinity information is used to identify the recommendations to be sent to an individual user. Second, we combine the social affinity information with the spatio-temporal context of users and historical responses of the user to further refine the set of recommendations and to decide when a recommendation would be sent. Third, social affinity computation and spatio-temporal contextual association are continuously tuned through user feedback. We outline the challenges in building such a system, and outline approaches to deal with such challenges. Lakshmish Ramaswamy, Deepak P 0001, Ramana Polavarapu, Kutila Gunasekera, Dinesh Garg, Karthik Visweswariah, Shivkumar Kalyanaraman |
Mobile Data Management | 1 |
| 2009 | Privacy-Aware Collaborative Spam FilteringabstractWhile the concept of collaboration provides a natural defense against massive spam e-mails directed at large numbers of recipients, designing effective collaborative anti-spam systems raises several important research challenges. First and foremost, since e-mails may contain confidential information, any collaborative anti-spam approach has to guarantee strong privacy protection to the participating entities. Second, the continuously evolving nature of spam demands the collaborative techniques to be resilient to various kinds of camouflage attacks. Third, the collaboration has to be lightweight, efficient, and scalable. Toward addressing these challenges, this paper presents ALPACAS-a privacy-aware framework for collaborative spam filtering. In designing the ALPACAS framework, we make two unique contributions. The first is a feature-preserving message transformation technique that is highly resilient against the latest kinds of spam attacks. The second is a privacy-preserving protocol that provides enhanced privacy guarantees to the participating entities. Our experimental results conducted on a real e-mail data set shows that the proposed framework provides a 10 fold improvement in the false negative rate over the Bayesian-based Bogofilter when faced with one of the recent kinds of spam attacks. Further, the privacy breaches are extremely rare. This demonstrates the strong privacy protection provided by the ALPACAS system. Kang Li 0001, Zhenyu Zhong, Lakshmish Ramaswamy |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2008 | Replication in Overlay Networks: A Multi-objective Optimization Approach
Osama Al-Haj Hassan, Lakshmish Ramaswamy, John A. Miller 0001, Khaled Rasheed, E. Rodney Canfield |
CollaborateCom | 2 |
| 2008 | ALPACAS: A Large-Scale Privacy-Aware Collaborative Anti-Spam SystemabstractWhile the concept of collaboration provides a natural defense against massive spam emails directed at large numbers of recipients, designing effective collaborative anti-spam systems raises several important research challenges. First and foremost, since emails may contain confidential information, any collaborative anti-spam approach has to guarantee strong privacy protection to the participating entities. Second, the continuously evolving nature of spam demands the collaborative techniques to be resilient to various kinds of camouflage attacks. Third, the collaboration has to be lightweight, efficient, and scalable. Towards addressing these challenges, this paper presents ALPACAS - a privacy-aware framework for collaborative spam filtering. In designing the ALPACAS framework, we make two unique contributions. The first is a feature-preserving message transformation technique that is highly resilient against the latest kinds of spam attacks. The second is a privacy-preserving protocol that provides enhanced privacy guarantees to the participating entities. Our experimental results conducted on a real email dataset shows that the proposed framework provides a 10 fold improvement in the false negative rate over the Bayesian-based Bogofilter when faced with one of the recent kinds of spam attacks. Further, the privacy breaches are extremely rare. This demonstrates the strong privacy protection provided by the ALPACAS system. Zhenyu Zhong, Lakshmish Ramaswamy, Kang Li 0001 |
INFOCOM | 2 |
| 2008 | PeerCast: Churn-resilient end system multicast on heterogeneous overlay networks
Jianjun Zhang 0001, Ling Liu 0001, Lakshmish Ramaswamy, Calton Pu |
J. Netw. Comput. Appl. | 3 |
| 2007 | Message replication in unstructured peer-to-peer networkabstractRecently, unstructured peer-to-peer (P2P) applications have become extremely popular. Searching in these networks has been a hot research topic. Flooding-based searching, which has been the basis of real-world P2P networks is inherently inefficient and unscalable. Replication has proven to be an effective strategy to improve efficiency and scalability of unstructured P2P networks. Previous research has largely focused on replicating resources or their references. This paper considers a replication solution from a different perspective; we investigate replicating messages and its effect on overloading problem. We propose two message replication strategies. The distance-based message replication technique replicates the query messages at different topological regions of the network. The landmarks-based technique further optimizes the performance by considering both the topology as well as the physical proximities of the peers of the overlay. Our experiments show that the proposed techniques substantially reduce the message traffic in the overlay while maintaining query performance. Osama Al-Haj Hassan, Lakshmish Ramaswamy |
CollaborateCom | 2 |
| 2007 | A Semantic Framework for Identifying Events in a Service Oriented ArchitectureabstractWe propose a semantic framework for automatically identifying events as a step towards developing an adaptive middleware for Service Oriented Architecture (SOA). Current related research focuses on adapting to events that violate certain non-functional objectives of the service requestor. Given the large of number of events that can happen during the execution of a service, identifying events that can impact the non-functional objectives of a service request is a key challenge. To address this problem we propose an approach that allows service requestors to create semantically rich service requirement descriptions, called semantic templates. We propose a formal model for expressing semantic templates and for measuring the relevance of an event to both the action being performed and the nonfunctional objectives. This model is extended to adjust the relevance of the events based on feedback from the underlying adaptation framework. We present an algorithm that utilizes multiple ontologies for identifying relevant events and present our evaluations that measure the efficiency of both the event identification and the subsequent adaptation scheme. Karthik Gomadam, Ajith Ranabahu, Lakshmish Ramaswamy, Amit P. Sheth, Kunal Verma |
ICWS | 3 |
| 2007 | CoQUOS: Lightweight Support for Continuous Queries in Unstructured OverlaysabstractThe utility and the effectiveness of peer-to-peer (P2P) content distribution systems can be greatly enhanced by augmenting their ad-hoc content discovery mechanisms with two capabilities, namely a mechanism to enable the peers to register their queries and receive notifications when corresponding data-items are added to the network and a means for the peers to advertise their new content. While P2P-based publish-sub scribe systems can infuse these capabilities, developing full-fledged publish-subscribe systems on top of unstructured P2P networks requires complex techniques, and it is often an overkill for many P2P applications. For these applications, we study the alternate continuous query paradigm, which is functionally similar to publish-subscribe systems, but provides best-effort notification guarantees. This paper presents CoQUOS - a scalable and lightweight middleware to support continuous queries in unstructured P2P networks. A key strength of the CoQUOS system is that it can be implemented on any unstructured overlay network. Moreover, CoQUOS preserves the simplicity and flexibility of the overlay network. Central to our design of the CoQUOS middleware is a completely decentralized scheme to register a query at different regions of the P2P network. This mechanism includes two novel components, namely cluster resilient random walk algorithm for propagating query to various regions of the network and dynamic probability-based query registration technique for ensuring that the registrations are well distributed. Our experiments show that the proposed techniques are highly effective and their overheads are low. Lakshmish Ramaswamy, Jianxia Chen, Piyush Parate |
IPDPS | 1 |
| 2007 | A Utility-Aware Middleware Architecture for Decentralized Group Communication Applications
Jianjun Zhang 0001, Ling Liu 0001, Lakshmish Ramaswamy, Gong Zhang 0008, Calton Pu |
Middleware | 3 |
| 2007 | A framework for encoding and caching of video for quality adaptive progressive downloadabstractProgressive download of multimedia objects over the Internet (e.g. www.youtube.com), where the video is downloaded and viewed during the download process, has become an increasingly popular alternative to multimedia streaming. Due to the fluctuating bandwidth and latency of the Internet, progressive download is often not fast enough, often resulting in intermittent stalling of the video. In this paper, we first propose a variation of the existing MPEG Fine Grained Scalability (FGS) profile to create a layered video representation that is suitable for progressive download in an environment characterized by varying bitrate. We also propose an efficient caching scheme that is specifically tailored for the proposed layered video representation. The proposed layered version of the Greedy-Dual-Size cache replacement policy is shown to reduce the latency observed by the client during progressive download of video in a varying bitrate environment. Experimental results demonstrate that the proposed caching scheme improves the latency of progressive video downloads as well as the server efficiency. Siddhartha Chattopadhyay, Lakshmish Ramaswamy, Suchendra M. Bhandarkar |
ACM Multimedia | 2 |
| 2007 | Message Diffusion in Unstructured Overlay NetworksabstractMany unstructured overlay-based peer-to-peer (P2P) applications require techniques that can effectively send messages to various topological regions of the overlay. While searching in unstructured P2P networks has been widely studied in literature, the problem of diffusing messages to various parts of an arbitrary overlay network has received surprisingly little research attention. In this paper we analyze the message diffusion problem and make two technical contributions towards addressing it. First, we propose a novel message propagation technique called the cluster resilient random walk (CRW). While the CRW technique preserves the overall framework of random walks, at each step of message forwarding, it favors the neighbors that are more likely to send the message deeper into the network. Second, in order to ensure effective message diffusion in networks with small cuts, we introduce a unique message fission technique in which messages are split when they reach peers connecting two or more topological regions of the network. Our experiments show that the proposed technique are very effective in diffusing messages across overlay networks of various topologies. Jianxia Chen, Lakshmish Ramaswamy, Archana Meka |
NCA | 2 |
| 2007 | Scalable Delivery of Dynamic Content Using a Cooperative Edge Cache GridabstractIn recent years, edge computing has emerged as a popular mechanism to deliver dynamic Web content to clients. However, many existing edge cache networks have not been able to harness the full potential of edge computing technology. In this paper, we argue and experimentally demonstrate that cooperation among the individual edge caches coupled with scalable server-driven document consistency mechanisms can significantly enhance the capabilities and performance of edge cache networks in delivering fresh dynamic content. However, designing large-scale cooperative edge cache networks presents many research challenges. Toward addressing these challenges, this paper presents cooperative edge cache grid (cooperative EC grid, for short)-a large-scale cooperative edge cache network for efficiently delivering highly dynamic Web content with varying server update frequencies. The design of the cooperative EC grid focuses on the scalability and reliability of dynamic content delivery in addition to cache hit rates, and it incorporates several novel features. We introduce the concept of cache clouds as a generic framework of cooperation in large-scale edge cache networks. The architectural design of the cache clouds includes dynamic hashing-based document lookup and update protocols, which dynamically balance lookup and update loads among the caches in the cloud. We also present cooperative techniques for making the document lookup and update protocols resilient to the failures of individual caches. This paper reports a series of simulation-based experiments which show that the overheads of cooperation in the cooperative EC grid are very low, and our architecture and techniques enhance the performance of the cooperative edge networks. Lakshmish Ramaswamy, Ling Liu 0001, Arun Iyengar |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2006 | Cooperative Data Placement and Replication in Edge Cache NetworksabstractCooperation among individual caches has proven to be an effective strategy to improve the scalability and performance of edge cache networks delivering dynamic Web content. To date, research in the area of cooperative edge caching has mainly focused on serving client requests and maintaining freshness of cached documents. However, designing mechanisms to effectively manage the available resources is an important challenge that can have significant impact on the performance of an edge cache network. In this paper we propose a novel data placement scheme, called the utility-based placement scheme, which is not only sensitive to the ongoing cooperation in the edge cache but also takes into account the various costs and benefits of storing a data-item at an individual edge cache. At the heart of proposed scheme is a utility function that quantifies the usefulness of storing a data-item at a particular edge cache. Experiments show that the proposed scheme provides significant performance benefits Lakshmish Ramaswamy, Arun Iyengar, Jianxia Chen |
CollaborateCom | 1 |
| 2006 | Efficient Formation of Edge Cache Groups for Dynamic Content DeliveryabstractCost-effective cooperation among a network of edge caches is widely accepted as an effective mechanism for enhancing the scalability, performance, and reliability of edge cache networks. However, the problem of how to form cache groups for achieving effective and efficient cooperation in edge cache networks has largely been unexplored. In this paper, we identify two important factors that need to be considered while forming cooperative groups, namely, network proximities of edge caches and network distances of the caches to the origin server. We propose two novel cache clustering schemes for accurately partitioning the caches of a given edge cache network into specified number of cache groups. The first scheme, called the Selective Landmarks scheme (SL scheme), accurately partitions the edge cache network into cooperative groups based on the network proximities of the caches. The second cache group formation scheme, called Server Distance sensitive Selective Landmarks scheme (SDSL scheme), provides a careful combination network proximities and server distances. Our experiments indicate that the proposed techniques can yield significant performance benefits. Lakshmish Ramaswamy, Ling Liu 0001, Jianjun Zhang 0001 |
ICDCS | 1 |
| 2005 | Efficient delivery of dynamic content: the cooperative EC grid projectabstractThe exponential growth of dynamic Web content has posed serious challenges to the scalability of the World Wide Web. While caching on the edge of the Internet has emerged as a popular technique to address these challenges, many of the present-day edge caching systems do not harness the complete benefits of edge computing. Our research efforts in the cooperative edge cache grid project are aimed at utilizing collaboration among edge caches as a means to further enhance the capabilities and the performance of edge cache network. This paper outlines the cooperative EC grid project including its architecture, fundamental concepts, and various techniques that have been designed for supporting low-cost cooperation among the edge caches Lakshmish Ramaswamy, Jianxia Chen |
CollaborateCom | 1 |
| 2005 | Cache Clouds: Cooperative Caching of Dynamic Documents in Edge NetworksabstractCaching on the edge of the Internet is becoming a popular technique to improve the scalability and efficiency of delivering dynamic web content. In this paper we study the challenges in designing a large scale cooperative edge cache network, focusing on mechanisms and methodologies for efficient cooperation among caches to improve the overall performance of the edge cache network. This paper makes three original contributions. First, we introduce the concept of cache clouds, which forms the fundamental framework for cooperation among caches in the edge network. Second, we present dynamic hashing-based protocols for document lookups and updates within each cache cloud, which are not only efficient, but also effective in dynamically balancing lookup and update loads among the caches in the cloud. Third, we outline a utility-based mechanism for placing dynamic documents within a cache cloud. Our experiments indicate that these techniques can significantly improve the performance of the edge cache networks. Lakshmish Ramaswamy, Ling Liu 0001, Arun Iyengar |
ICDCS | 1 |
| 2005 | Automatic Fragment Detection in Dynamic Web Pages and Its Impact on CachingabstractConstructing Web pages from fragments has been shown to provide significant benefits for both content generation and caching. In order for a Web site to use fragment-based content generation, however, good methods are needed for fragmenting the Web pages. Manual fragmentation of Web pages is expensive, error prone, and unscalable. This paper proposes a novel scheme to automatically detect and flag fragments that are cost-effective cache units in Web sites serving dynamic content. Our approach analyzes Web pages with respect to their information sharing behavior, personalization characteristics, and change patterns. We identify fragments which are shared among multiple documents or have different lifetime or personalization characteristics. Our approach has three unique features. First, we propose a framework for fragment detection, which includes a hierarchical and fragment-aware model for dynamic Web pages and a compact and effective data structure for fragment detection. Second, we present an efficient algorithm to detect maximal fragments that are shared among multiple documents. Third, we develop a practical algorithm that effectively detects fragments based on their lifetime and personalization characteristics. This paper shows the results when the algorithms are applied to real Web sites. We evaluate the proposed scheme through a series of experiments, showing the benefits and costs of the algorithms. We also study the impact of using the fragments detected by our system on key parameters such as disk space utilization, network bandwidth consumption, and load on the origin servers. Lakshmish Ramaswamy, Arun Iyengar, Ling Liu 0001, Fred Douglis |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2005 | A Distributed Approach to Node Clustering in Decentralized Peer-to-Peer NetworksabstractConnectivity-based node clustering has wide-ranging applications in decentralized peer-to-peer (P2P) networks such as P2P file sharing systems, mobile ad-hoc networks, P2P sensor networks, and so forth. This paper describes a connectivity-based distributed node clustering scheme (CDC). This scheme presents a scalable and efficient solution for discovering connectivity-based clusters in peer networks. In contrast to centralized graph clustering algorithms, the CDC scheme is completely decentralized and it only assumes the knowledge of neighbor nodes instead of requiring a global knowledge of the network (graph) to be available. An important feature of the CDC scheme is its ability to cluster the entire network automatically or to discover clusters around a given set of nodes. To cope with the typical dynamics of P2P networks, we provide mechanisms to allow new nodes to be incorporated into appropriate existing clusters and to gracefully handle the departure of nodes in the clusters. These mechanisms enable the CDC scheme to be extensible and adaptable in the sense that the clustering structure of the network adjusts automatically as nodes join or leave the system. We provide detailed experimental evaluations of the CDC scheme, addressing its effectiveness in discovering good quality clusters and handling the node dynamics. We further study the types of topologies that can benefit best from the connectivity-based distributed clustering algorithms like CDC. Our experiments show that utilizing message-based connectivity structure can considerably reduce the messaging cost and provide better utilization of resources, which in turn improves the quality of service of the applications executing over decentralized peer-to-peer networks. Lakshmish Ramaswamy, Bugra Gedik, Ling Liu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2004 | Automatic detection of fragments in dynamically generated web pagesabstractDividing web pages into fragments has been shown to provide significant benefits for both content generation and caching. In order for a web site to use fragment-based content generation, however, good methods are needed for dividing web pages into fragments. Manual fragmentation of web pages is expensive, error prone, and unscalable. This paper proposes a novel scheme to automatically detect and flag fragments that are cost-effective cache units in web sites serving dynamic content. We consider the fragments to be interesting if they are shared among multiple documents or they have different lifetime or personalization characteristics. Our approach has three unique features. First, we propose a hierarchical and fragment-aware model of the dynamic web pages and a data structure that is compact and effective for fragment detection. Second, we present an efficient algorithm to detect maximal fragments that are shared among multiple documents. Third, we develop a practical algorithm that effectively detects fragments based on their lifetime and personalization characteristics. We evaluate the proposed scheme through a series of experiments, showing the benefits and costs of the algorithms. We also study the impact of adopting the fragments detected by our system on disk space utilization and network bandwidth consumption. Lakshmish Ramaswamy, Arun Iyengar, Ling Liu 0001, Fred Douglis |
WWW | 1 |
| 2004 | An Expiration Age-Based Document Placement Scheme for Cooperative Web CachingabstractThe sharing of caches among proxies is an important technique to reduce Web traffic, alleviate network bottlenecks, and improve response time of document requests. Most existing work on cooperative caching has been focused on serving misses collaboratively. Very few have studied the effect of cooperation on document placement schemes and its potential enhancements on cache hit ratio and latency reduction. We propose a new document placement scheme which takes into account the contentions at individual caches in order to limit the replication of documents within a cache group and increase document hit ratio. The main idea of this new scheme is to view the aggregate disk space of the cache group as a global resource of the group and uses the concept of cache expiration age to measure the contention of individual caches. The decision of whether to cache a document at a proxy is made collectively among the caches that already have a copy of this document. We refer to this new document placement scheme as the Expiration Age-based scheme (EA scheme). The EA scheme effectively reduces the replication of documents across the cache group, while ensuring that a copy of the document always resides in a cache where it is likely to stay for the longest time. We report our study on the potentials and limits of the EA scheme using both analytic modeling and trace-based simulation. The analytical model compares and contrasts the existing (ad hoc) placement scheme of cooperative proxy caches with our new EA scheme and indicates that the EA scheme improves the effectiveness of aggregate disk usage, thereby increasing the average time duration for which documents stay in the cache. The trace-based simulations show that the EA scheme yields higher hit rates and better response times compared to the existing document placement schemes used in most of the caching proxies. Lakshmish Ramaswamy, Ling Liu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2003 | Techniques for efficient fragment detection in web pagesabstractThe existing approaches to fragment-based publishing, delivery and caching of web pages assume that the web pages are manually fragmented at their respective web sites. However manual fragmentation of web pages is expensive, error prone, and not scalable. This paper proposes a novel scheme to automatically detect and flag possible fragments in a web site. Our approach is basedonananalysisofthewebpagesdynamicallygeneratedat given web sites with respect to their information sharing behavior, personalization characteristics and change patterns. Categories and Subject Descriptors: H.3.3 [Information Systems- Information storage and retrieval]: Information search and retrieval Lakshmish Ramaswamy, Arun Iyengar, Ling Liu 0001, Fred Douglis |
CIKM | 1 |
| 2003 | Connectivity Based Node Clustering in Decentralized Peer-to-Peer NetworksabstractConnectivity based node clustering has wide ranging applications in decentralized peer-to-peer (P2P) networks such as P2P file sharing systems, mobile ad-hoc networks, P2P sensor networks and so forth. We describe a connectivity-based distributed node clustering scheme (CDC). This scheme presents a scalable and an efficient solution for discovering connectivity based clusters in peer networks. In contrast to centralized graph clustering algorithms, the CDC scheme is completely decentralized and it only assumes the knowledge of neighbor nodes, instead of requiring a global knowledge of the network (graph) to be available. An important feature of the CDC scheme is its ability to cluster the entire network automatically or to discover clusters around a given set of nodes. We provide experimental evaluations of the CDC scheme, addressing its effectiveness in discovering good quality clusters. Our experiments show that utilizing message-based connectivity structure can considerably reduce the messaging cost, and provide better utilization of resources, which in turn improves the quality of service of the applications executing over decentralized peer-to-peer networks. Lakshmish Ramaswamy, Bugra Gedik, Ling Liu 0001 |
Peer-to-Peer Computing | 1 |
| 2002 | A New Document Placement Scheme for Cooperative Caching on the InternetabstractMost existing work on cooperative caching has been-focused on serving misses collaboratively. Very few have studied the effect of cooperation on document placement schemes and its potential enhancements on cache hit ratio and latency reduction. In this paper we propose a new document placement scheme, called the Expiration Age based scheme (EA scheme), which takes into account the contentions at individual caches in order to limit the replication of documents within a cache group and increase document hit ratio. The main idea of this new scheme is to view the aggregate disk space of the cache group as a global resource of the group, and uses the concept of cache expiration age to measure the contention of individual caches. The decision of whether to cache a document at a proxy is made collectively, among the caches that already have a copy of this document. The experiments show that the EA scheme yields higher hit rates and better response times compared to the existing document placement schemes used in most of the caching proxies. Lakshmish Ramaswamy, Ling Liu 0001 |
ICDCS | 1 |