Calton Pu

dblp:p/CaltonPu · DBLP profile ↗
← Back
219ranked-venue papers
20as first author
16since 2021 · last 2026
0000-0002-6616-8987ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 51 · 7 first-authorSoftware engineering, systems software and programming languages · 44 · 3 first-author · 5 since 2021Systems, architecture and hardware · 40 · 6 first-author · 1 since 2021Computer networks · 23 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 1 since 2021Human-computer interaction and ubiquitous computing · 20 · 1 first-author · 2 since 2021Security and privacy · 18 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 9 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Adaptive Skill-Mastery Feedback Loops in AI-Generated Courses
abstract
This demo presents the next-generation version of KimBilet.com, an educational platform that leverages generative AI to create adaptive, mastery-driven learning experiences. Building on last year's personalized course generation, the new system introduces a fine-grained skill taxonomy and a feedback loop that evaluates and responds to learner performance in real time.
Aibek Musaev, Kerimbek Musaev, Mirbek Dzhumaliev, Calton Pu
SIGCSE (2)4
2025 Leveraging Generative AI for Personalized Learning Experiences
abstract
This demo presents KimBilet.com, an educational platform that utilizes generative AI to create personalized educational content on demand. Catering to high-school and college students, instructors, job seekers, and lifelong learners, the system generates customized courses based on user prompts, covering any topic of interest. Each course may include a sequence of AI-created lessons and quizzes, providing detailed feedback for every quiz option to enhance understanding. The platform supports intuitive navigation through keyboard shortcuts and allows users to jump between course items seamlessly. It also maintains a history of completed quizzes to help users track their learning progress. Future enhancements include topic suggestions based on past interests, support for coding exercises, and multilingual support. This demo will showcase how KimBilet.com leverages AI to offer adaptive learning experiences, engage attendees through interactive exploration, and discuss its potential applications in educational settings. Participants will gain insights into integrating AI-driven tools into teaching and learning processes to address diverse educational needs.
Mirbek Dzhumaliev, Aibek Musaev, Calton Pu
SIGCSE (2)3
2025 Challenges and Experiences in Data Integration to Support Research on Nonprofit Organizations
abstract
Nonprofit organizations are important contributors to the US economy and social well-being. The Nonprofit Organization Research Panel Project (NORPP) Manager has been developing and sharing datasets and software tools to facilitate data-driven research on nonprofit organizations. The project has two major thrusts: (1) large-scale survey panels, e.g., the Annual National Survey of Nonprofit Trends and Impacts, from 2021 to the current (2024); and (2) NORPP Analytics Integration Platform (NAIP) to collect, process, and query a wide range of socioeconomic indicator datasets, such as IRS 990 form and census data. Technical challenges that include data heterogeneity, data quality, and protection of sensitive data have made the expansion and maintenance of NAIP datasets both labor-intensive and time-consuming. We are currently exploring new technologies, including Large Language Models such as GPT series, to automate database query generation, schema adaptation, and data quality assurance processes.
Calton Pu, Teresa Derrick-Mills, Lewis Faulk
ACM Trans. Internet Techn.1
2024 Sync-Millibottleneck Attack on Microservices Cloud Architecture
abstract
The modern web services landscape is characterized by numerous fine-grained, loosely coupled microservices with increasingly stringent low-latency requirements. However, this architecture also brings new performance vulnerabilities. In this paper, we introduce a novel low-volume application layer DDoS attack called the Sync-Millibottleneck (SyncM) attack, specifically targeting microservices. The goal of this attack is to cause a long-tail latency problem that violates the service-level agreement (SLA) while evading state-of-the-art DDoS detection/defense mechanisms. The SyncM attack exploits two unique features of microservices architecture: (1) the shared frontend gateway that directs user requests to mid-tier/backend microservices, and (2) the co-existence of multiple logically independent execution paths, each with its own bottleneck resource. By creating synchronized millibottlenecks (i.e., sub-second duration bottlenecks) on multiple independent execution paths, SyncM attack can cause the queuing effect in each execution path to be propagated and superimposed in the shared frontend gateway. As a result, SyncM triggers surprisingly high latency spikes in the system, even when all system resources are far from saturation, making it challenging to trace the cause of performance instability.
Xuhang Gu, Qingyang Wang 0001, Qiben Yan 0001, Jianshu Liu, Calton Pu
AsiaCCS5
2023 An Integrated Cloud-Edge-Device Adaptive Deep Learning Service for Cross-Platform Web
abstract
Deep learning shows great promise in providing more intelligence to the cross-platform web. However, insufficient infrastructure, heavy models, and intensive computation limit the use of deep learning with low-performing web browsers. We propose DeepAdapter, an integrated cloud-edge-device framework that ties the edge, the remote cloud, with the device by cross-platform web technology for adaptive deep learning services towards lower latency, lower mobile energy, and higher system throughput. DeepAdapter consists of context-aware pruning, service updating, and online scheduling. First, the offline pruning module provides a context-aware pruning algorithm that incorporates the latency, the network condition, and the device's computing capability to fit various contexts. Second, the service updating module optimizes branch model cache on the edge for massive mobile users and updates the new model pruning requirements. Third, the online scheduling module matches optimal branch models for mobile users. Also, a two-stage DRL-based online scheduling method named DeepScheduler can handle high concurrent requests between edge centers and remote cloud by designing the reward prediction model. Extensive experiments show that DeepAdapter can decrease average latency by 1.33x, reduce average mobile energy by 1.4x, and improve system throughput by 2.1x with considerable accuracy.
Yakun Huang, Xiuquan Qiao, Jian Tang 0008, Pei Ren, Ling Liu 0001, Calton Pu, Junliang Chen 0001
IEEE Trans. Mob. Comput.6
2023 Special Section on "Advances in Cyber-Manufacturing: Architectures, Challenges, & Future Research Directions"
abstract
introduction Share on Special Section on “Advances in Cyber-Manufacturing: Architectures, Challenges, & Future Research Directions” Editors: Gautam Srivastava Brandon University, Canada Brandon University, CanadaSearch about this author , Jerry Chun-Wei Lin Silesian Uiversity of Technology, Poland Silesian Uiversity of Technology, PolandSearch about this author , Calton Pu Georgia Tech, USA Georgia Tech, USASearch about this author , Yudong Zhang University of Leicester, UK University of Leicester, UKSearch about this author Authors Info & Claims ACM Transactions on Internet TechnologyVolume 23Issue 4Article No.: 49pp 1–4https://doi.org/10.1145/3627990Published:17 November 2023Publication History 0citation0DownloadsMetricsTotal Citations0Total Downloads0Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Gautam Srivastava 0001, Jerry Chun-Wei Lin, Calton Pu, Yudong Zhang 0001
ACM Trans. Internet Techn.3
2022 ShadowSync: latency long tail caused by hidden synchronization in real-time LSM-tree based stream processing systems
abstract
Mission-critical, real-time, continuous stream processing applications that interact with the real world have stringent latency requirements. For example, e-commerce websites like Amazon improve their marketing strategy by performing real-time advertising based on customers' behavior, and latency long tail can cause significant revenue loss. Recent work [39] showed a positive correlation between latency long tail and variance in the execution time of synchronous invocation chains (critical paths) in microservices benchmarks. This paper shows that asynchronous, very short but intense resource demands (called millibottlenecks) outside of critical paths can also cause significant latency long tail.
Shungeng Zhang, Qingyang Wang 0001, Yasuhiko Kanemasa, Julius Michaelis, Jianshu Liu, Calton Pu
Middleware6
2022 Edge AR X5: An Edge-Assisted Multi-User Collaborative Framework for Mobile Web Augmented Reality in 5G and Beyond
abstract
Multi-user mobile Augmented Reality (AR) has been successfully used in various fields as a novel visual interaction technology. But current mainstream wearable device-based and app-based solutions are still facing cross-platform, real-time communication, and intensive computing requirements. Mobile Web technology is envisioned to be a promising supporting technology for cross-platform application of mobile AR especially in 5G networks, which provide pervasive communication and computing resources thereby forming a formidable framework for the practical application of multi-user mobile Web AR. However, the problem of how to use these new techniques properly to achieve efficient communication and computing collaboration is obviously paramount in order for multi-user mobile Web AR to be realized in 5G networks. In this article, we propose the first edge-assisted multi-user collaborative framework for mobile Web AR in the 5G era. First, we propose a heuristic mechanism BA-CPP for efficient communication planning, which allows multi-user interaction synchronization to be achieved. Second, we introduce a motion-aware key frame selection mechanism called Mo-KFP to optimize the computational efficiency of the edge system, and simultaneously alleviate the initialization problem by collaborating with nearby mobile devices using the Device-to-Device (D2D) communication technique. Experiments are conducted in a real-world 5G network, and the results demonstrate the superiority of our proposed collaborative framework.
Pei Ren, Xiuquan Qiao, Yakun Huang, Ling Liu 0001, Calton Pu, Schahram Dustdar, Junliang Chen 0001
IEEE Trans. Cloud Comput.5
2022 A Lightweight Collaborative Deep Neural Network for the Mobile Web in Edge Cloud
abstract
Enabling deep learning technology on the mobile web can improve the user’s experience for achieving web artificial intelligence in various fields. However, heavy DNN models and limited computing resources of the mobile web are now unable to support executing computationally intensive DNNs when deploying in a cloud computing platform. With the help of promising edge computing, we propose a lightweight collaborative deep neural network for the mobile web, named LcDNN, which contributes to three aspects: (1) We design a composite collaborative DNN that reduces the model size, accelerates inference, and reduces mobile energy cost by executing a lightweight binary neural network (BNN) branch on the mobile web. (2) We provide a jointly training method for LcDNN and implement an energy-efficient inference library for executing the BNN branch on the mobile web. (3) To further promote the resource utilization of the edge cloud, we develop a DRL-based online scheduling scheme to obtain an optimal allocation for LcDNN. The experimental results show that LcDNN outperforms existing approaches for reducing the model size by about 16x to 29x. It also reduces the end-to-end latency and mobile energy cost with acceptable accuracy and improves the throughput and resource utilization of the edge cloud.
Yakun Huang, Xiuquan Qiao, Pei Ren, Ling Liu 0001, Calton Pu, Schahram Dustdar, Junliang Chen 0001
IEEE Trans. Mob. Comput.5
2022 Introduction to the Special Issue on Multiagent Systems and Services in the Internet of Things
abstract
research-article Share on Introduction to the Special Issue on Multiagent Systems and Services in the Internet of Things Authors: Andrei Ciortea University of St. Gallen, Switzerland University of St. Gallen, Switzerland 0000-0003-0721-4135Search about this author , Xiaomin Zhu National University of Defense Technology, China National University of Defense Technology, China 0000-0003-1301-7840Search about this author , Calton Pu Georgia Institute of Technology, USA Georgia Institute of Technology, USA 0000-0002-6616-8987Search about this author , Munindar P. Singh North Carolina State University, USA North Carolina State University, USA 0000-0003-3599-3893Search about this author Authors Info & Claims ACM Transactions on Internet TechnologyVolume 22Issue 4November 2022 Article No.: 99pp 1–3https://doi.org/10.1145/3584744Published:03 March 2023Publication History 0citation0DownloadsMetricsTotal Citations0Total Downloads0Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Andrei Ciortea, Xiaomin Zhu 0001, Calton Pu, Munindar P. Singh
ACM Trans. Internet Techn.3
2022 Fine-Grained Elastic Partitioning for Distributed DNN Towards Mobile Web AR Services in the 5G Era
abstract
Web-based Deep Neural Networks (DNNs) enhance the ability of object recognition and has attracted considerable attention in mobile Web AR and other services. However, neither performing the DNN inference on mobile Web browsers locally nor offloading computations to the cloud can strike a balance between accuracy and efficiency; generally, rude methods are often accompanied by unsatisfactory accuracy. Collaborative approaches seem to fill this gap by coordinating the distributed hierarchical computing resources, especially in the 5G era, but it still faces challenges in the current solutions, such as the lack of (1) full use of 5G resources for the one point DNN computation partitioning schemes; (2) fine-grained branching mechanism; (3) efficient partitioning method; and (4) multi-objective optimization. To this end, we present the fine-grained elastic computation partitioning mechanism for distributed DNN in 5G networks. First, we elaborate two collaborative scenarios. Second, we study the DNN branching mechanism at layer granularity. Next, we propose a DNN computation partitioning algorithm based on deep reinforcement learning. Finally, we develop a mobile Web AR application as a proof of concept. The experiments were conducted in an actually deployed 5G trial network, and the results show the superiority of this collaborative approach. The common theme is, under the premise that Quality of Service (QoS) is satisfied, to balance multiple interests by orchestrating computations across heterogeneous computing platforms.
Pei Ren, Xiuquan Qiao, Yakun Huang, Ling Liu 0001, Calton Pu, Schahram Dustdar
IEEE Trans. Serv. Comput.5
2022 A Comparative Measurement Study of Deep Learning as a Service Framework
abstract
Big data powered Deep Learning (DL) and its applications have blossomed in recent years, fueled by three technological trends: a large amount of digitized data openly accessible, a growing number of DL software frameworks in open source and commercial markets, and a selection of affordable parallel computing hardware devices. However, no single DL framework, to date, dominates in terms of performance and accuracy even for baseline classification tasks on standard datasets, making the selection of a DL framework an overwhelming task. This paper takes a holistic approach to conduct empirical comparison and analysis of four representative DL frameworks with three unique contributions.First, given a selection of CPU-GPU configurations, we show that for a specific DL framework, different configurations of its hyper-parameters may have a significant impact on both performance and accuracy of DL applications.Second, to the best of our knowledge, this study is the first to identify the opportunities for improving the training time performance and the accuracy of DL frameworks by configuring parallel computing libraries and tuning individual and multiple hyper-parameters.Third, we also conduct a comparative measurement study on the resource consumption patterns of four DL frameworks and their performance and accuracy implications, including CPU and memory usage, and their correlations to varying settings of hyper-parameters under different configuration combinations of hardware, parallel computing libraries. We argue that this measurement study provides in-depth empirical comparison and analysis of four representative DL frameworks, and offers practical guidance for service providers to deploying and delivering DL as a Service (DLaaS) and for application developers and DLaaS consumers to select the right DL frameworks for the right DL workloads.
Yanzhao Wu 0001, Ling Liu 0001, Calton Pu, Wenqi Cao, Semih Sahin, Wenqi Wei 0001, Qi Zhang 0009
IEEE Trans. Serv. Comput.3
2021 SINETStream: Enabling Research IoT Applications with Portability, Security and Performance Requirements
abstract
Demands for Internet of Things (IoT) platforms are increasing along with the expansion of data-driven sciences, and academic network infrastructure such as national research and education networks (NRENs) are now being encouraged to aggressively support diverse IoT research projects. In Japan, the National Institute of Informatics (NII) operates an NREN called "SINET5," which is an academic backbone network linking more than 900 universities and research institutions. In addition to SINET5's high-speed 100 Gbps backbone network, SINET5 also provides a mobile network called "Mobile SINET" as IoT infrastructure. To provide crucial support to IoT research projects and to facilitate the development and deployment of IoT applications, we present herein our experience with our software library, "SINETStream." In this paper, we provide an overview of SINETStream and discuss lessons learned from examples of application deployment over Mobile SINET. SINETStream provides a common and simple application program interface (API) for various message brokers, security functions, and performance tuning support features. These functions improve the portability of applications and help application developers to remove hindrances to the development of secure and efficient IoT applications. The experimental results provided herein show that application developers can use SINETStream functions within a reasonable overhead. We also show how the combination of SINETStream and Mobile SINET enables users to develop and deploy highly confidential and efficient IoT applications.
Atsuko Takefusa, Ikki Fujiwara, Hiroshi Yoshida, Kento Aida, Calton Pu
COMPSAC6
2021 Introduction to the Special Section on Data Science for Cyber-Physical Systems
abstract
Introduction to the Special Section on Data Science for Cyber-Physical Systems
Francesco Piccialli, Nik Bessis, Gwanggil Jeon, Calton Pu
ACM Trans. Internet Techn.4
2021 Introduction to the Special Section on Security and Privacy of Medical Data for Smart Healthcare
abstract
introduction Share on Introduction to the Special Section on Security and Privacy of Medical Data for Smart Healthcare Editors: Amit Kumar Singh National Institute of Technology Patna, India National Institute of Technology Patna, IndiaSearch about this author , Jonathan Wu University of Windsor, Canada University of Windsor, CanadaSearch about this author , Ali Al-Haj Princess Sumaya University for Technology, Jordan Princess Sumaya University for Technology, JordanSearch about this author , Calton Pu Georgia Institute of Technology, USA Georgia Institute of Technology, USASearch about this author Authors Info & Claims ACM Transactions on Internet TechnologyVolume 21Issue 3August 2021 Article No.: 53pp 1–4https://doi.org/10.1145/3460870Online:09 June 2021Publication History 3citation88DownloadsMetricsTotal Citations3Total Downloads88Last 12 Months88Last 6 weeks5 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Amit Kumar Singh 0001, Q. M. Jonathan Wu, Ali Al-Haj 0001, Calton Pu
ACM Trans. Internet Techn.4
2021 Robust and Reliable Process-Aware Information Systems
abstract
Over recent years, several sophisticated Process-Aware Information Systems (PAIS) have been proposed for managing business processes and automating large-scale scientific (e-Science) processes. Much of this success is due to their ability to provide generic functionality for modeling, execution and monitoring processes. These functionalities work well when process execution follows a well-behaved path towards achieving the models objectives. However, exceptions and anomalous situations that fall outside of the well-behaved execution path still pose a significant challenge to PAIS. The treatment for such exceptions usually involves interventions in systems by human operators, which result in significant additional cost for businesses. In this paper, we introduce a cost-aware recovery composition method that is able to find and follow recovery paths that reduce the cost of exception handling. From a practical point of view, our proposal reduces complexity and the need for manual interventions to handle exceptions. Finally, the feasibility of recovery mechanism is discussed from its implementation into WED-flow framework.
André Luís Schwerz, Rafael Liberato, Calton Pu, João Eduardo Ferreira
IEEE Trans. Serv. Comput.3
2020 DeepAdapter: A Collaborative Deep Learning Framework for the Mobile Web Using Context-Aware Network Pruning
abstract
Deep learning shows great promise in providing more intelligence to the mobile web, but insufficient infrastructure, heavy models, and intensive computation limit the use of deep learning in mobile web applications. In this paper, we present DeepAdapter, a collaborative framework that ties the mobile web with an edge server and a remote cloud server to allow executing deep learning on the mobile web with lower processing latency, lower mobile energy, and higher system throughput. DeepAdapter provides a context-aware pruning algorithm that incorporates the latency, the network condition and the computing capability of the mobile device to fit the resource constraints of the mobile web better. It also provides a model cache update mechanism improving the model request hit rate for mobile web users. At runtime, it matches an appropriate model with the mobile web user and provides a collaborative mechanism to ensure accuracy. Our results show that DeepAdapter decreases average latency by 1.33x, reduces average mobile energy consumption by 1.4x, and improves system throughput by 2.1x with a considerable accuracy. Its contextaware pruning algorithm also improves inference accuracy by up to 0.3% with a smaller and faster model.
Yakun Huang, Xiuquan Qiao, Jian Tang 0008, Pei Ren, Ling Liu 0001, Calton Pu, Junliang Chen 0001
INFOCOM6
2020 DoubleFaceAD: A New Datastore Driver Architecture to Optimize Fanout Query Performance
abstract
The broad adoption of fanout queries on distributed datastores has made asynchronous event-driven datastore drivers a natural choice due to reduced multithreading overhead. However, through extensive experiments using the latest datastore drivers (e.g., MongoDB, HBase, DynamoDB) and YCSB benchmark, we show that an asynchronous datastore driver can cause unexpected performance degradation especially in fanout-query scenarios. For example, the default MongoDB asynchronous driver adopts the latest Java asynchronous I/O library, which uses a hidden on-demand JVM level thread pool to process fanout query responses, causing a surprising multithreading overhead when the query response size is large. A second instance is the traditional wisdom of modular design of an application server and the embedded asynchronous datastore driver can cause an im-balanced workload between the two components due to lack of coordination, incurring frequent unnecessary system calls. To address the revealed problems, we introduce DoubleFaceAD--a new asynchronous datastore driver architecture that integrates the management of both upstream and downstream workload traffic through a few shared reactor threads, with fanout-query-aware priority-based scheduling to reduce the overall query waiting time. Our experimental results on two representative application scenarios (YCSB and DBLP) show DoubleFaceAD outperforms all other types of datastore drivers up to 34% on throughput and 1.9× faster on 99th percentile response time.
Shungeng Zhang, Qingyang Wang 0001, Yasuhiko Kanemasa, Jianshu Liu, Calton Pu
Middleware5
2020 ODIN: Automated Drift Detection and Recovery in Video Analytics
Abhijit Suprem, Joy Arulraj, Calton Pu, João Eduardo Ferreira
Proc. VLDB Endow.3
2020 Beyond Artificial Reality: Finding and Monitoring Live Events from Social Sensors
abstract
With billions of active social media accounts and millions of live video cameras, live new big data offer many opportunities for smart applications. However, the main consumers of the new big data have been humans. We envision the research on live knowledge , to automatically acquire real-time, validated, and actionable information. Live knowledge presents two significant and diverging technical challenges: big noise and concept drift. We describe the EBKA (evidence-based knowledge acquisition) approach, illustrated by the LITMUS landslide information system. LITMUS achieves both high accuracy and wide coverage, demonstrating the feasibility and promise of EBKA approach to achieve live knowledge.
Calton Pu, Abhijit Suprem, Rodrigo Alves Lima, Aibek Musaev, De Wang, Danesh Irani, Steve Webb, João Eduardo Ferreira
ACM Trans. Internet Techn.1
2019 Demystifying Learning Rate Policies for High Accuracy Training of Deep Neural Networks
abstract
Learning Rate (LR) is an important hyper-parameter to tune for effective training of deep neural networks (DNNs). Even for the baseline of a constant learning rate, it is non-trivial to choose a good constant value for training a DNN. Dynamic learning rates involve multi-step tuning of LR values at various stages of the training process and offer high accuracy and fast convergence. However, they are much harder to tune. In this paper, we present a comprehensive study of 13 learning rate functions and their associated LR policies by examining their range parameters, step parameters, and value update parameters. We propose a set of metrics for evaluating and selecting LR policies, including the classification confidence, variance, cost, and robustness, and implement them in LRBench, an LR benchmarking system. LRBench can assist end-users and DNN developers to select good LR policies and avoid bad LR policies for training their DNNs. We tested LRBench on Caffe, an open source deep learning framework, to showcase the tuning optimization of LR policies. Evaluated through extensive experiments, we attempt to demystify the tuning of LR policies by identifying good LR policies with effective LR value ranges and step sizes for LR update schedules.
Yanzhao Wu 0001, Ling Liu 0001, Juhyun Bae, Ka-Ho Chow 0001, Arun Iyengar, Calton Pu, Wenqi Wei 0001, Lei Yu 0002, Qi Zhang 0009
IEEE BigData6
2019 Integration of Machine Learning Techniques as Auxiliary Diagnosis of Inherited Metabolic Disorders: Promising Experience with Newborn Screening Data
Bo Lin 0008, Jianwei Yin, Qiang Shu, Shuiguang Deng, Ying Li 0001, Pingping Jiang, Rulai Yang, Calton Pu
CollaborateCom8
2019 A Lightweight Collaborative Recognition System with Binary Convolutional Neural Network for Mobile Web Augmented Reality
abstract
Lightweight and precise recognition is a key component of web-based augmented reality (Web AR) applications. Although edge-based distributed deep learning approach is now possible to achieve satisfactory recognition for Web AR applications, it puts significant pressure on the computation and energy consumption of the mobile web browser, especially the app-based embedded browser. Thus, reducing the model size and accelerating the inference are regarded as the two fundamental challenges to enable this edge-based collaborative recognition system efficiently. In this paper, we propose a lightweight collaborative recognition system (LCRS) for Web AR applications. LCRS contributes to three aspects: (1) we design a composite deep neural network for reducing the model size and inference latency by introducing binary convolutional neural network; (2) we provide a joint training method to co-train the general branch and the binary branch; (3) we develop a JavaScript library for the mobile web browser to execute and accelerate inference of the binary branch, which also provides a collaborative mechanism between the mobile web browser and the edge server. We have conducted extensive experiments using several well-known networks and datasets. The experimental results have shown that the proposed system outperforms the existing approaches in terms of reducing the model size by about 16x to 29x, and it also reduces end-to-end latency and outpaces the existing state-of-the-art approaches by over 3x to 60x when applying it in practical Web AR cases.
Yakun Huang, Xiuquan Qiao, Pei Ren, Ling Liu 0001, Calton Pu, Junliang Chen 0001
ICDCS5
2019 Differentially Private Model Publishing for Deep Learning
abstract
Deep learning techniques based on neural networks have shown significant success in a wide range of AI tasks. Large-scale training datasets are one of the critical factors for their success. However, when the training datasets are crowdsourced from individuals and contain sensitive information, the model parameters may encode private information and bear the risks of privacy leakage. The recent growing trend of the sharing and publishing of pre-trained models further aggravates such privacy risks. To tackle this problem, we propose a differentially private approach for training neural networks. Our approach includes several new techniques for optimizing both privacy loss and model accuracy. We employ a generalization of differential privacy called concentrated differential privacy(CDP), with both a formal and refined privacy loss analysis on two different data batching methods. We implement a dynamic privacy budget allocator over the course of training to improve model accuracy. Extensive experiments demonstrate that our approach effectively improves privacy loss accounting, training efficiency and model quality under a given privacy budget.
Lei Yu 0002, Ling Liu 0001, Calton Pu, Mehmet Emre Gursoy, Stacey Truex
IEEE Symposium on Security and Privacy3
2019 Mitigating Tail Response Time of n-Tier Applications: The Impact of Asynchronous Invocations
abstract
Consistent low response time is essential for e-commerce due to intense competitive pressure. However, practitioners of web applications have often encountered the long-tail response time problem in cloud data centers as the system utilization reaches moderate levels (e.g., 50%). Our fine-grained measurements of an open source n-tier benchmark application (RUBBoS) show such long response times are often caused by Cross-tier Queue Overflow (CTQO). Our experiments reveal the CTQO is primarily created by the synchronous nature of RPC-style call/response inter-tier communications, which create strong inter-tier dependencies due to the request processing chain of classic n-tier applications composed of synchronous RPC/thread-based servers. We remove gradually the dependencies in n-tier applications by replacing the classic synchronous servers (e.g., Apache, Tomcat, and MySQL) with their corresponding event-driven asynchronous version (e.g., Nginx, XTomcat, and XMySQL) one-by-one. Our measurements with two application scenarios (virtual machine co-location and background monitoring interference) show that replacing a subset of asynchronous servers will shift the CTQO, without significant improvements in long-tail response time. Only when all the servers become asynchronous the CTQO is resolved. In synchronous n-tier applications, long-tail response times resulting from CTQO arise at utilization as low as 43%. On the other hand, the completely asynchronous n-tier system can disrupt CTQO and remove the long tail latency at utilization as high as 83%.
Qingyang Wang 0001, Shungeng Zhang, Yasuhiko Kanemasa, Calton Pu
ACM Trans. Internet Techn.4
2018 A Comparative Study of Containers and Virtual Machines in Big Data Environment
abstract
Container technique is gaining increasing attention in recent years and has become an alternative to traditional virtual machines. Some of the primary motivations for the enterprise to adopt the container technology include its conveniency to encapsulate and deploy applications, lightweight operations, as well as efficiency and flexibility in resources sharing. However, there still lacks an in-depth and systematic comparison study on how big data applications, such as Spark jobs, perform between a container environment and a virtual machine environment. In this paper, by running various Spark applications with different configurations, we evaluate the two environments from many interesting aspects, such as how convenient the execution environment can be set up, what are makespans of different workloads running in each setup, how efficient the hardware resources, such as CPU and memory, are utilized, and how well each environment can scale. The results show that compared with virtual machines, containers provide a more easy-to-deploy and scalable environment for big data workloads. The research work in this paper can help practitioners and researchers to make more informed decisions on tuning their cloud environment and configuring the big data applications, so as to achieve better performance and higher resources utilization.
Qi Zhang 0009, Ling Liu 0001, Calton Pu, Qiwei Dou, Liren Wu, Wei Zhou 0011
IEEE CLOUD3
2018 Evaluating User Satisfaction with Typography Designs via Mining Touch Interaction Data in Mobile Reading
abstract
Previous work has demonstrated that typography design has a great influence on users' reading experience. However, current typography design guidelines are mainly for general purpose, while the individual needs are nearly ignored. To achieve personalized typography designs, an important and necessary step is accurately evaluating user satisfaction with the typography designs. Current evaluation approaches, e.g., asking for users' opinions directly, however, interrupt the reading and affect users' judgments. In this paper, we propose a novel method to address this challenge by mining users' implicit feedbacks, e.g., touch interaction data. We conduct two mobile reading studies in Chinese to collect the touch interaction data from 91 participants. We propose various features based on our three hypotheses to capture meaningful patterns in the touch behaviors. The experiment results show the effectiveness of our evaluation models with higher accuracy on comparing with the baseline under three text difficulty levels, respectively.
Jianwei Yin, Shuiguang Deng, Ying Li 0001, Calton Pu, Zhiling Luo
CHI5
2018 Efficient Shared Memory Orchestration towards Demand Driven Memory Slicing
abstract
Memory is increasingly becoming a bottleneck for big data and latency-sensitive applications in virtualized systems. Memory efficiency is critical for high-performance execution of virtual machines (VMs). Mechanisms proposed for improving memory utilization often rely on an accurate estimation of VM working set size at runtime, which is difficult under changing workloads. This paper explores opportunities for improving memory efficiency and their impacts on the performance of VM executions. First, we show that if each VM is initialized with an application-specified lower bound memory, then by maintaining a shared memory region across VMs in the presence of temporal memory usage variations on the host, those VMs under high memory pressure can minimize their performance loss by opportunistically and transparently harvesting idle memory on other VMs. Second, we show that by enabling on-demand VM memory allocation and deallocation in the presence of changing workloads, VM performance degradation due to memory swapping can be reduced effectively, compared to the conventional VM configuration scenario, in which all VMs are allocated with the upper-bound of memory requested by their applications. Third, we show that by providing shared memory pipes between co-located VMs, the inter-VM communication can speed up by avoiding unnecessary overhead of communication via the network. We develop MemLego, a lightweight shared memory based system, to achieve all these benefits without requiring any modification to user applications and the OSes. We demonstrate the effectiveness of these opportunities through extensive experiments on unmodified Redis and MemCached. Using MemLego, the throughput of Redis and Memcached improves by up to 4x over the native system without MemLego, up to 2 orders of magnitude when the applications working set size does not fit in memory.
Qi Zhang 0009, Ling Liu 0001, Calton Pu, Wenqi Cao, Semih Sahin
ICDCS3
2018 Towards Bandwidth Guarantee for Virtual Clusters Under Demand Uncertainty in Multi-Tenant Clouds
abstract
In the cloud, multiple tenants share the resource of datacenters and their applications compete with each other for scarce network bandwidth. Current studies have shown that the lack of bandwidth guarantee causes unpredictable network performance, leading to poor application performance. To address this issue, several virtual network abstractions have been proposed which allow the tenants to reserve virtual clusters with specified bandwidth between the Virtual Machines (VMs) in the datacenters. However, all these existing proposals require the tenants to deterministically characterize the bandwidth demands in the abstractions, which can be difficult and result in inefficient bandwidth reservation due to the demand uncertainty. In this paper, we explore a virtual cluster abstraction with stochastic bandwidth characterization to address the bandwidth demand uncertainty. We propose Stochastic Virtual Cluster (SVC), which models the bandwidth demand between VMs in a probabilistic way. Based on SVC, we develop a stochastic framework for virtual cluster allocation, in which the admitted virtual cluster's bandwidth demands are satisfied with a high probability. Efficient VM allocation algorithms are proposed to implement the framework while reducing the possibility of link congestion through minimizing the maximum bandwidth occupancy of a virtual cluster on physical links. Using simulations, we show that SVC achieves the trade-off between the job concurrency and the average job running time, and demonstrate its effectiveness for accommodating cloud application workloads with highly volatile bandwidth demands and its improvement to work-conserving bandwidth enforcement.
Lei Yu 0002, Haiying Shen, Zhipeng Cai 0001, Ling Liu 0001, Calton Pu
IEEE Trans. Parallel Distributed Syst.5
2018 Editorial Preface: Special Issue on Mobile & Cloud Computing Services
abstract
The four papers in this special section provide deep research results to report the advance in mobile and cloud computing services. In recent years, cloud computing has become a scalable services consumption and delivery platform in the field of Services Computing. The technical foundations of cloud computing include Service-Oriented Architecture (SOA) and virtualizations of hardware and software. The goal of cloud computing is to share resources among the cloud service consumers, the cloud service providers, and the cloud vendors in the cloud value chain.
Jia Zhang 0001, Stephen S. Yau, Calton Pu, Onur Altintas
IEEE Trans. Serv. Comput.3
2017 Tail Attacks on Web Applications
abstract
As the extension of Distributed Denial-of-Service (DDoS) attacks to application layer in recent years, researchers pay much interest in these new variants due to a low-volume and intermittent pattern with a higher level of stealthiness, invaliding the state-of-the-art DDoS detection/defense mechanisms. We describe a new type of low-volume application layer DDoS attack--Tail Attacks on Web Applications. Such attack exploits a newly identified system vulnerability of n-tier web applications (millibottlenecks with sub-second duration and resource contention with strong dependencies among distributed nodes) with the goal of causing the long-tail latency problem of the target web application (e.g., 95th percentile response time > 1 second) and damaging the long-term business of the service provider, while all the system resources are far from saturation, making it difficult to trace the cause of performance degradation.
Huasong Shan, Qingyang Wang 0001, Calton Pu
CCS3
2017 milliScope: A Fine-Grained Monitoring Framework for Performance Debugging of n-Tier Web Services
abstract
Modern distributed systems are often considered to be black boxes that greatly limit the potential to understand behaviors at the level of detail necessary to diagnose some of the most important types of performance problems. Recently researchers have found abnormal response time delays, one to two orders of magnitude longer than the average response time, that exist in short periods and cause economic loss for service providers. These very short bottlenecks are hard to detect due to their short life spans and their variety of possible reasons. In this paper, we propose milliScope (mScope), the first millisecond-granularity software-based resource and event monitoring for distributed systems that achieves both performance, low overhead at high frequency, and high accuracy matched with other firmware monitoring tool. More specifically, milliScope is a fine-grained monitoring framework to collaborate multiple mScopeMonitors for event and resource monitoring to reconstruct the flow of each client request and profile execution performance in a distributed system. We utilize the resource mScopeMonitors for system resource monitoring, and we develop our own event mScopeMonitors to identify the execution boundary in a lightweight, precise and systematic methodology. The semantic and syntactic of these monitoring logs with arbitrary formats are enriched by our multistage data transformation tool, mScopeDataTransformer, which unifies the diverse monitoring logs into a dynamic data warehouse, mScopeDB, for advanced analysis. We conduct several illustrative scenarios in which milliScope successfully diagnoses the response time anomalies caused by very short bottlenecks using a representative web application benchmark (RUBBoS).
Chien-An Lai, Josh Kimball, Qingyang Wang 0001, Calton Pu
ICDCS5
2017 Performance Analysis of Cloud Computing Centers Serving Parallelizable Rendering Jobs Using M/M/c/r Queuing Systems
abstract
Performance analysis is crucial to the successful development of cloud computing paradigm. And it is especially important for a cloud computing center serving parallelizable application jobs, for determining a proper degree of parallelism could reduce the mean service response time and thus improve the performance of cloud computing obviously. In this paper, taking the cloud based rendering service platform as an example application, we propose an approximate analytical model for cloud computing centers serving parallelizable jobs using M/M/c/r queuing systems, by modeling the rendering service platform as a multi-station multi-server system. We solve the proposed analytical model to obtain a complete probability distribution of response time, blocking probability and other important performance metrics for given cloud system settings. Thus this model can guide cloud operators to determine a proper setting, such as the number of servers, the buffer size and the degree of parallelism, for achieving specific performance levels. Through extensive simulations based on both synthetic data and real-world workload traces, we show that our proposed analytical model can provide approximate performance prediction results for cloud computing centers serving parallelizable jobs, even those job arrivals follow different distributions.
Xiulin Li, Li Pan 0001, Jiwei Huang, Shijun Liu, Yuliang Shi, Calton Pu
ICDCS7
2017 LITMUS: Towards Multilingual Reporting of Landslides
abstract
LITMUS is a real-time online and openly accessible service that collects high quality information on landslide events from social media. This service uses disaster related keywords, such as "landslide" and "mudslide", to analyze messages posted by English speaking users. However, comprehensive coverage of disasters must include multilingual support as there are events that are reported in languages other than English. We discuss and evaluate possible implementations of such support using "native" and "translated" approaches. "Native" approach involves a complete reimplementation of the existing infrastructure in another language whereas in the "translated" approach the existing infrastructure can be used without modification. As an illustration, we present a demo that extends LITMUS to implement a "native" approach for multilingual reporting of landslide events.
Aibek Musaev, Qixuan Hou, Calton Pu
ICDCS4
2017 Towards Multilingual Automated Classification Systems
abstract
In this paper we propose and evaluate three approaches for automated classification of texts in over 60 languages without the need for a manually annotated dataset in those languages. All approaches are based on the randomized Explicit Semantic Analysis method using multilingual Wikipedia articles as their knowledge repository. We evaluate the proposed approaches by classifying a Twitter dataset in English and Portuguese into relevant and irrelevant items with respect to landslide as a natural disaster, where the highest achieved F1-score is 0.93. These approaches can be used in various applications where multilingual classification is needed, including multilingual disaster reporting using Social Media to improve coverage and increase confidence. As illustration, we present a demonstration that combines data from physical sensors and social networks to detect landslide events reported in English and Portuguese.
Aibek Musaev, Calton Pu
ICDCS2
2017 REX: Rapid Ensemble Classification System for Landslide Detection Using Social Media
abstract
We study the problem of using Social Media to detect natural disasters, of which we are interested in a special kind, namely landslides. Employing information from Social Media presents unique research challenges, as there exists a considerable amount of noise due to multiple meanings of the search keywords, such as "landslide" and "mudslide". To tackle these challenges, we propose REX, a rapid ensemble classification system which can filter out noisy information by implementing two key ideas: (I) a new method for constructing independent classifiers that can be used for rapid ensemble classification of Social Media texts, where each classifier is built using randomized Explicit Semantic Analysis; and (II) a self-correction approach which takes advantage of the observation that the majority label assigned to Social Media texts belonging to a large event is highly accurate. We perform experiments using real data from Twitter over 1.5 years to show that REX classification achieves 0.98 in F-measure, which outperforms the standard Bag-of-Words algorithm by an average of 0.14 and the state-of-the-art Word2Vec algorithm by 0.04. We also release the annotated datasets used in the experiments as a contribution to the research community containing 282k labeled items.
Aibek Musaev, De Wang, Jiateng Xie, Calton Pu
ICDCS4
2017 The Millibottleneck Theory of Performance Bugs, and Its Experimental Verification
abstract
The performance of n-tier web-facing applications often suffer from response time long-tail problem. With relatively low resource utilization (less than 50%) and the majority of requests returning within a few milliseconds, a non-negligible num-ber of normally short requests may take seconds to return. We propose the millibottleneck theory of performance bugs (that lead to long-tail problems). Several case studies have confirmed the millibottlenecks (that last a few tens to hundreds of milliseconds) as causal agents of long requests. A concrete example (garbage collection) illustrates the experimental verification of millibottlenecks. An open source fine-grain monitoring toolkit is being devel-oped to facilitate the experimental research on millibottlenecks.
Calton Pu, Josh Kimball, Chien-An Lai, Jack Li 0001, Junhee Park, Qingyang Wang 0001, Deepal Jayasinghe, PengCheng Xiong, Simon Malkowski, Qinyi Wu, Gueyoung Jung, Younggyun Koh, Galen S. Swint
ICDCS1
2017 A Study of Long-Tail Latency in n-Tier Systems: RPC vs. Asynchronous Invocations
abstract
Long-tail latency of web-facing applications continues to be a serious problem. Most of the previously published research addresses two classes of long latency problems: uneven workloads such as web search, and resource saturation in single nodes. We describe an experimental study of a third class of long tail latency problems that are specific to distributed systems: Cross-Tier Queue Overflow (CTQO) due to a combination of millibottlenecks (with sub-second duration) and tightly-coupled servers in n-tier systems (e.g., Apache, Tomcat, and MySQL) using RPC-style request-response communications. Our experiments show that the appearance of millibottlenecks (e.g., created by short workload bursts) in one server often causes another server (which has no saturated resources) in the synchronous invocation chain to fill up its queues (CTQO) and drop packets, creating very long response time queries. CTQO can be reduced or avoided by replacing the server dropping packets with an asynchronous server. In synchronous n-tier system experiments, long tail latency due to CTQO can be reproduced consistently at utilization as low as 43%. In contrast, when all n-tier servers are replaced by asynchronous versions, CTQO and consequent dropped packets remain absent at utilization levels as high as 83%, despite the same millibottlenecks.
Qingyang Wang 0001, Chien-An Lai, Yasuhiko Kanemasa, Shungeng Zhang, Calton Pu
ICDCS5
2017 Automated Performance Evaluation for Multi-tier Cloud Service Systems Subject to Mixed Workloads
abstract
In multi-tier cloud service systems, performance evaluation relies on numerous experiments in order to collect key metrics such as resources usage. The approach may result in highly time-consuming in practice. In this paper, we propose an automated framework for performance tracking, data management and analysis to minimize human intervention in multi-tier cloud service systems. The framework support fine-grained analysis of the mixed workloads through the Discrete-time Markov-modulated Poisson process (DMMPP). A general multi-tier application is theoretically formulated as a queueing network to evaluate the performance. The effectiveness of the model has been validated through extensive experiments conducted in the RUBiS benchmark system.
Xudong Zhao 0004, Jiwei Huang, Lei Liu 0003, Shijun Liu, Calton Pu, Li-Zhen Cui 0001
ICDCS5
2017 Limitations of Load Balancing Mechanisms for N-Tier Systems in the Presence of Millibottlenecks
abstract
The scalability of n-tier systems relies on effective load balancing to distribute load among the servers of the same tier. We found that load balancing mechanisms (and some policies) in servers used in typical n-tier systems (e.g., Apache and Tomcat) have issues of instability when very long response time (VLRT) requests appear due to millibottlenecks, very short bottlenecks that last only tens to hundreds of milliseconds. Experiments with standard n-tier benchmarks show that during millibottlenecks, some load balancing policy/mechanism combinations make the mistake of sending new requests to the node(s) suffering from millibottlenecks, instead of the idle nodes as load balancers are supposed to do. Several of these mistakes are due to the implicit assumptions made by load balancing policies and mechanisms on the stability of system state. Our study shows that appropriate remedies at policy and mechanism levels can avoid these mistakes during millibottlenecks and remove the VLRT requests, thus improving the average response time by a factor of 12.
Jack Li 0001, Josh Kimball, Junhee Park, Chien-An Lai, Calton Pu, Qingyang Wang 0001
ICDCS6
2017 Landslide Information Service Based on Composition of Physical and Social Sensors
abstract
Modern world data come from an increasing number of sources, including data from physical sensors like weather satellites and seismographs as well as social networks and web logs. While progress has been made in the filtering of individual social networks, there are significant advantages in the integration of big data from multiple sources. For physical events, the integration of physical sensors and social network data can improve filtering efficiency and quality of results beyond what is feasible in each individual data stream. Disasters are representative physical events with real world impact. As illustration and demonstration, we have built the LITMUS landslide information service that combines data from both physical sensors and social networks in real-time. LITMUS filters and combines reliable but indirect physical data with direct report social media data on landslides to achieve high quality and wide coverage of landslide information.
Aibek Musaev, Calton Pu
ICDE2
2017 Real-Time Soft Resource Allocation in Multi-Tier Web Service Systems
abstract
Soft resource allocation is an important factor of system configuration which plays a critical role in guaranteeing the performance of multi-tier web service systems. There is a tradeoff between real-time performance and resource consumption, and thus the real-time adjustment of soft resource allocation in response to dynamic workload is quite challenging. In this paper, we propose a real-time soft resource allocation method that integrates both model-based analysis and real-time optimization. Specifically, a multi-tier web service system is firstly formulated by a queueing network model, and theoretical analyses are provided. Then, an optimization approach for real-time soft resource allocation is designed by applying sliding window techniques, in order to cope with dynamic workloads and performance demands. Based on the RUBiS benchmark system, model parameters are obtained by measurements and the efficacy of our approach is finally validated.
Xudong Zhao 0004, Jiwei Huang, Lei Liu 0003, Yuliang Shi, Shijun Liu, Calton Pu, Li-Zhen Cui 0001
ICWS6
2017 An Experimental Study of a Biosequence Big Data Analysis Service
abstract
With the development of next-generation sequencing (NGS), DNA/RNA sequencing has become cheaper and more efficient. Today, a whole human genome can be sequenced under $1,000, providing opportunities for large-scale bioinformatic analysis on big datasets. However, most of existing bioinformatic analysis tools are programmed for single server based computing platform and not suitable to process such big datasets. As Hadoop MapReduce and Spark are gaining popularity as cluster computing based big data processing platform, more and more bioinformatic applications start to explore cluster computing platform for large scale data analysis. In this paper we present an in-depth experimental study on deploying Spark clusters for high performance bioinformatic short sequence reconstruction. Our experimental results enable us to answer a number of challenging and yet most frequently asked questions regarding efficient management of bioinformatic data analysis services on Spark systems. Example questions include how to best split big dataset into multiple partitions, and how to distribute data partitions and bioinformatic analysis tasks on a Spark cluster for carrying out a high performance distributed analysis job? What types of memory models are effective for bioinformatic data analysis services on a Spark cluster? Why do different bioinformatic data analysis operations exhibit different throughput performance on the same Spark cluster? We conjecture that this experimental study not only demonstrates the feasibility of high performance bioinformatic data analysis on Spark platform, but also will help bioinformatic application developers to make more informed decisions on both design and configuration of Spark Cluster, managing and tuning parameters of Spark runtime system for enhancing the performance of large scale big data analytics.
Wei Zhou 0011, Ling Liu 0001, Calton Pu, Qingyang Wang 0001, Wenkun Xiang, Shaowen Yao 0001
ICWS3
2017 Dynamic Differential Location Privacy with Personalized Error Bounds
Lei Yu 0002, Ling Liu 0001, Calton Pu
NDSS3
2017 ASSER: An Efficient, Reliable, and Cost-Effective Storage Scheme for Object-Based Cloud Storage Systems
abstract
High reliability, efficient I/O performance and flexible consistency provided with low storage cost are all desirable properties of cloud storage systems. Due to the inherent conflicts, however, simultaneously achieving optimum on all these properties is impractical. N-way Replication and Erasure Coding, two extensively-applied storage schemes with high reliability, adopt opposite and unbalanced strategies on the tradeoff among these properties, thus considerably restraining their effectiveness on wide range of workloads. To address the aforementioned obstacle, we propose a novel storage scheme called ASSER, an ASSembling chain of Erasure coding and Replication. ASSER stores each object in two parts: a full copy and a certain amount of erasure-coded segments. We establish dedicated read/write protocols for ASSER leveraging the unique structural advantages. On the basis of elementary protocols, we implement sequential and PRAM (Pipeline-RAM) consistency to make ASSER feasible for various services with different performance/consistency requirements. Evaluation results demonstrate that under the same fault tolerance and consistency level, ASSER outperforms N-way replication and pure erasure coding in I/O throughput under diverse system and workload configurations with superior performance stability. More importantly, ASSER delivers stably efficient I/O performance at much lower storage cost than the other comparatives.
Jianwei Yin, Shuiguang Deng, Ying Li 0001, Wei Lo, Kexiong Dong, Albert Y. Zomaya, Calton Pu
IEEE Trans. Computers8
2016 Coarse-Grained Information Flow Control on Hybrid Clouds
abstract
Recently, more and more enterprises have adopted hybrid cloud strategies to simultaneously enjoy the security of on-premise clouds and the low cost of public clouds. The key challenge of hybrid clouds, though, stems from the difficulty of specifying where the data should be stored and where the information could flow efficiently. In order to meet security concerns and performance requirements, we introduce a coarse-grained information flow control (CIFC) model to limit storing, accessing, and disclosing of confidential data in public clouds. The CIFC model aims at providing information control implicitly, without the large overhead of periodically checking access privileges. Moreover, since the CIFC model may request redistributing data whenever the secrecy level of a dataset changes, we formulate the data redistribution problem as an optimization problem and propose the Partition Biased Sampling Algorithm (PBSA) for its solution. We implemented the CIFC model on top of Spark, and our results show that Spark applications can achieve 1.4 to 2.1 times better performance by utilizing the additional computational capacity of public cloud to process non-sensitive data. Furthermore, we integrate the PBSA algorithm into Spark and demonstrate a saving of more than 35% in execution time, compared to the Spark default data distribution strategy.
Chien-An Lai, Asser N. Tantawi, Calton Pu
CLOUD3
2016 Enabling Elastic Stream Processing in Shared Clusters
abstract
Distributed data stream processing has become an increasingly popular computational framework due to many emerging applications which require real-time processing of data such as dynamic content delivery and security event analysis. These distributed data stream processing applications are often run on shared, multi-tenant clusters as companies try to consolidate from dedicated clusters for each application (batch and streaming) to a single cluster using a global cluster manager such as Hadoop YARN. In shared cluster environments, guaranteeing the quality of service constraints for throughput and response time for both stream processing applications and batch applications is a significant challenge. Stream processing applications often face an elastic demand where the input rate can vary drastically. The typical solution to solve workload elasticity is to guarantee enough resources to the application, but this solution is not possible when resources are being shared among multiple applications. In this paper, we present an approach for supporting elastic scaling of distributed data stream processing applications and efficiently scheduling and coordinating stream processing with batch processing in shared clusters. Our solution consists of a congestion detection monitor which detects bottlenecks in the streaming system and a global state manager that performs non-disruptive, stateful scaling of streaming applications. We implemented our solution using Storm, a popular stream processing framework, and tested our implementation on a Hadoop YARN cluster using a real-time security event processing workload. Our experimental results show that our solution improves stream processing application throughput by 49% over default Storm while decreasing average request response times by 58%.
Jack Li 0001, Calton Pu, Yuan Chen 0001, Daniel Gmach, Dejan S. Milojicic
CLOUD2
2016 Performance Interference of Memory Thrashing in Virtualized Cloud Environments: A Study of Consolidated n-Tier Applications
abstract
Modern datacenters employ server virtualization and consolidation to reduce the cost of operation and to maximize profit. However, interference among consolidated virtual machines (VMs) has barred mission-critical applications due to unpredictable performance. Through extensive measurements of RUBBoS n-tier benchmark, we found a major source of performance unpredictability: the memory thrashing caused by VM consolidation can reduce the system throughput by 46% although memory was not over-committed. On a physical host with 4 consolidated VMs, we observed two distinct operational modes during a typical RUBBoS benchmark experiment. Over the first half of run-time session we found frequent CPU IOwait causing very long response time requests even though the system is under read-only CPU intensive workload, however, the latter half showed no such CPU abnormalities (IOwait). Using ElbaLens - a lightweight tracing tool, we conducted fine-grain analyses at time granularities as short as 50ms and found that the abnormal IOwait is caused by transient memory thrashing among consolidated VMs. The abnormal IOwait induces queue overflows that propagate through the entire n-tier system, resulting in very long response time requests due to frequent TCP retransmissions. We provide three practical techniques such as VM migration, memory reallocation, soft resource reallocation and show that they can mitigate the effects of performance interference among consolidated VMs.
Junhee Park, Qingyang Wang 0001, Jack Li 0001, Chien-An Lai, Calton Pu
CLOUD6
2015 Fast Iterative Graph Computation with Resource Aware Graph Parallel Abstractions
abstract
Iterative computation on large graphs has challenged system research from two aspects: (1) how to conduct high performance parallel processing for both in-memory and out-of-core graphs; and (2) how to handle large graphs that exceed the resource boundary of traditional systems by resource aware graph partitioning such that it is feasible to run large-scale graph analysis on a single PC. This paper presents GraphLego, a resource adaptive graph processing system with multi-level programmable graph parallel abstractions. GraphLego is novel in three aspects: (1) we argue that vertex-centric or edge-centric graph partitioning are ineffective for parallel processing of large graphs and we introduce three alternative graph parallel abstractions to enable a large graph to be partitioned at the granularity of subgraphs by slice, strip and dice based partitioning; (2) we use dice-based data placement algorithm to store a large graph on disk by minimizing non-sequential disk access and enabling more structured in-memory access; and (3) we dynamically determine the right level of graph parallel abstraction to maximize sequential access and minimize random access. GraphLego can run efficiently on different computers with diverse resource capacities and respond to different memory requirements by real-world graphs of different complexity. Extensive experiments show the competitiveness of GraphLego against existing representative graph processing systems, such as GraphChi, GraphLab and X-Stream.
Yang Zhou 0001, Ling Liu 0001, Kisung Lee, Calton Pu, Qi Zhang 0009
HPDC4
2015 Toward a Real-Time Service for Landslide Detection: Augmented Explicit Semantic Analysis and Clustering Composition Approaches
abstract
The use of Social Media for event detection, such as detection of natural disasters, has gained a booming interest from research community as Social Media has become an immensely important source of real-time information. However, it poses a number of challenges with respect to high volume, noisy information and lack of geo-tagged data. Extraction of high quality information (e.g., Accurate locations of events) while maintaining good performance (e.g., Low latency) are the major problems. In this paper, we propose two approaches for tackling these issues: an augmented Explicit Semantic Analysis approach for rapid classification and a composition of clustering algorithms for location estimation. Our experiments demonstrate over 98% in precision, recall and F-measure when classifying Social Media data while producing a 20% improvement in location estimation due to clustering composition approach. We implement these approaches as part of the landslide detection service LITMUS, which is live and openly accessible for continued evaluation and use.
Aibek Musaev, De Wang, Saajan Shridhar, Chien-An Lai, Calton Pu
ICWS5
2015 Clustering Service Networks with Entity, Attribute, and Link Heterogeneity
abstract
Many popular web service networks are content-rich in terms of heterogeneous types of entities and links, associated with incomplete attributes. Clustering such heterogeneous service networks demands new clustering techniques that can handle two heterogeneity challenges: (1) multiple types of entities co-exist in the same service network with multiple attributes, and (2) links between entities have diverse types and carry different semantics. Existing heterogeneous graph clustering techniques tend to pick initial centroids uniformly at random, specify the number k of clusters in advance, and fix k during the clustering process. In this paper, we propose Service Cluster, a novel heterogeneous service network clustering algorithm with four unique features. First, we incorporate various types of entity, attribute and link information into a unified distance measure. Second, we design a Discrete Steepest Descent method to naturally produce initial k and initial centroids simultaneously. Third, we propose a dynamic learning method to automatically adjust the link weights towards clustering convergence. Fourth, we develop an effective optimization strategy to identify new suitable k and k well-chosen centroids at each clustering iteration. Extensive evaluation on real datasets demonstrates that Service Cluster outperforms existing representative methods in terms of both effectiveness and efficiency.
Yang Zhou 0001, Ling Liu 0001, Calton Pu, Kisung Lee, Balaji Palanisamy, Emre Yigitoglu, Qi Zhang 0009
ICWS3
2015 Improving Preemptive Scheduling with Application-Transparent Checkpointing in Shared Clusters
abstract
Modern data center clusters are shifting from dedicated single framework clusters to shared clusters. In such shared environments, cluster schedulers typically utilize preemption by simply killing jobs in order to achieve resource priority and fairness during peak utilization. This can cause significant resource waste and delay job response time.
Jack Li 0001, Calton Pu, Yuan Chen 0001, Vanish Talwar, Dejan S. Milojicic
Middleware2
2015 Scaling iterative graph computations with GraphMap
abstract
In recent years, systems researchers have devoted considerable effort to the study of large-scale graph processing. Existing distributed graph processing systems such as Pregel, based solely on distributed memory for their computations, fail to provide seamless scalability when the graph data and their intermediate computational results no longer fit into the memory; and most distributed approaches for iterative graph computations do not consider utilizing secondary storage a viable solution. This paper presents GraphMap, a distributed iterative graph computation framework that maximizes access locality and speeds up distributed iterative graph computations by effectively utilizing secondary storage. GraphMap has three salient features: (1) It distinguishes data states that are mutable during iterative computations from those that are read-only in all iterations to maximize sequential access and minimize random access. (2) It entails a two-level graph partitioning algorithm that enables balanced workloads and locality-optimized data placement. (3) It contains a proposed suite of locality-based optimizations that improve computational efficiency. Extensive experiments on several real-world graphs show that GraphMap outperforms existing distributed memory-based systems for various iterative graph algorithms.
Kisung Lee, Ling Liu 0001, Karsten Schwan, Calton Pu, Qi Zhang 0009, Yang Zhou 0001, Emre Yigitoglu, Pingpeng Yuan
SC4
2015 MICS: Mingling Chained Storage Combining Replication and Erasure Coding
abstract
High reliability, low space cost, and efficient read/write performance are all desirable properties for cloud storage systems. Due to the inherent conflicts, however, simultaneously achieving optimality on these properties is unrealistic. Since reliable storage is indispensable prerequisite for services with high availability, tradeoff should therefore be made between space and read/write efficiency when storage scheme is designed. N-way Replication and Erasure Coding, two extensively-used storage schemes with high reliability, adopt opposite strategies on this tradeoff issue. However, unbalanced tradeoff designs of both schemes confine their effectiveness to limited types of workloads and system requirements. To mitigate such applicability penalty, we propose MICS, a MIngling Chained Storage scheme that combines structural and functional advantages from both N-way replication and erasure coding. Qualitatively, MICS provides efficient read/write performance and high reliability at reasonably low space cost. MICS stores each object in two forms: a full copy and certain amount of erasure-coded segments. We establish dedicated read/write protocols for MICS leveraging the unique structural advantages. Moreover, MICS provides high read/write efficiency with Pipeline Random-Access Memory consistency to guarantee reasonable semantics for services users. Evaluation results demonstrate that under same fault tolerance and consistency level, MICS outperforms N-way replication and pure erasure coding in I/O throughput by up to 34.1% and 51.3% respectively. Furthermore, MICS shows superior performance stability over diverse workload conditions, in which case the standard deviation of MICS is 70.1% and 29.3% smaller than those of other two schemes.
Jianwei Yin, Wei Lo, Ying Li 0001, Shuiguang Deng, Kexiong Dong, Calton Pu
SRDS7
2015 SmartSLA: Cost-Sensitive Management of Virtualized Resources for CPU-Bound Database Services
abstract
Virtualization-based multi-tenant database consolidation is an important technique for database-as-a-service (DBaaS) providers to minimize their total cost which is composed of SLA penalty cost, infrastructure cost and action cost. Due to the bursty and diverse tenant workloads, over-provisioning for the peak or under-provisioning for the off-peak often results in either infrastructure cost or service level agreement (SLA) penalty cost. Moreover, although the process of scaling out database systems will help DBaaS providers satisfy tenants' service level agreement, its indiscriminate use has performance implications or incurs action cost. In this paper, we propose SmartSLA, a cost-sensitive virtualized resource management system for CPU-bound database services which is composed of two modules. The system modeling module uses machine learning techniques to learn a model for predicting the SLA penalty cost for each tenant under different resource allocations. Based on the learned model, the resource allocating module dynamically adjusts the resource allocation by weighing the potential reduction of SLA penalty cost against increase of infrastructure cost and action cost. SmartSLA is evaluated by using the TPC-W and modified YCSB benchmarks with dynamic workload trace and multiple database tenants. The experimental results show that SmartSLA is able to minimize the total cost under time-varying workloads compared to the other cost-insensitive approaches.
PengCheng Xiong, Yun Chi, Shenghuo Zhu, Hyun Jin Moon, Calton Pu, Hakan Hacigümüs
IEEE Trans. Parallel Distributed Syst.5
2015 LITMUS: A Multi-Service Composition System for Landslide Detection
abstract
Landslides are an illustrative example of multi-hazards, which can be caused by earthquakes, rainfalls and human activity among other reasons. Detection of landslides presents a significant challenge, since there are no physical sensors that would detect landslides directly. A more recent approach in detection of natural hazards, such as earthquakes, involves the use of social media. We propose a multi-service composition approach and describe LITMUS, which is a landslide detection service that combines data from both physical and social information services by filtering and then joining the information flow from those services based on their spatiotemporal features. Our results show that with such approach LITMUS detects 25 out of 27 landslides reported by USGS in December 2013 and 40 more landslide locations unreported by USGS during this period. LITMUS is a prototype tool that is used to investigate and implement research ideas in the area of disaster detection. We list some of the current work being done on refining the system that allows us to identify 137 landslide locations unreported by USGS during a more recent period of September 2014. Finally, we describe a live demonstration that displays landslide detection results on a web map in real-time.
Aibek Musaev, De Wang, Calton Pu
IEEE Trans. Serv. Comput.3
2014 IO Performance Interference among Consolidated n-Tier Applications: Sharing Is Better Than Isolation for Disks
abstract
The performance unpredictability associated with migrating applications into cloud computing infrastructures has impeded this migration. For example, CPU contention between co-located applications has been shown to exhibit counter-intuitive behavior. In this paper, we investigate IO performance interference through the experimental study of consolidated n-tier applications leveraging the same disk. Surprisingly, we found that specifying a specific disk allocation, e.g., limiting the number of Input/Output Operations Per Second (IOPs) per VM, results in significantly lower performance than fully sharing disk across VMs. Moreover, we observe severe performance interference among VMs can not be totally eliminated even with a sharing strategy (e.g., response times for constant workloads still increase over 1,100%). By using a micro-benchmark (Filebench) and an n-tier application benchmark systems (RUBBoS), we demonstrate the existence of disk contention in consolidated environments, and how performance loss occurs when co-located database systems in order to maintain database consistency flush their logs from memory to disk. Potential solutions to these isolation issues are (1) to increase the log buffer size to amortize the disk IO cost (2) to decrease the number of write threads to alleviate disk contention. We validate these methods experimentally and find a 64% and 57% reduction in response time (or more generally, a reduction in performance interference) for constant and increasing workloads respectively.
Chien-An Lai, Qingyang Wang 0001, Josh Kimball, Jack Li 0001, Junhee Park, Calton Pu
IEEE CLOUD6
2014 The Impact of Software Resource Allocation on Consolidated n-Tier Applications
abstract
Consolidating several under-utilized user applications together to achieve higher utilization of hardware resources is important for cloud vendors to reduce cost and maximize profit. In this paper, we study the impact of tuning software resources (e.g., server thread pool size or connection pool size) on n-tier web application performance in a consolidated cloud environment. By measuring CPU utilizations and performance of two consolidated n-tier web application benchmark systems running RUBBoS, we found significant differences depending on the amount of soft resources allocated. When the two systems have different soft resource allocations and are fully utilized, the application with more software resources may steal up to 8% CPU from the co-resident application. Further analysis shows that the CPU stealing is due to more threads being scheduled for the system with higher software resources. By limiting the number of runnable active threads for the consolidated VMs, we were able to mitigate the performance interference. More generally, our results show that careful software resource allocation is a significant factor when deploying and tuning n-tier application performance in clouds.
Jack Li 0001, Qingyang Wang 0001, Chien-An Lai, Junhee Park, Daisaku Yokoyama, Calton Pu
IEEE CLOUD6
2014 Finding Optimized Deployment Strategy for Multitenant Services by Iterative Staging
abstract
A serious challenge that confronts multi-tenant service systems is finding an optimized deployment strategy according to their business scale and operating characteristics. The tenants want to rent high performance services and services providers demand minimizing cost at the same time of meeting the requirements of tenants. But there are often contradictions between high performance and low cost. Therefore, in order to balance the contradiction, this paper proposes a staging-based optimized deployment method. This method performs iterative optimization based on customized workload generation, continuously emulation and evaluation in a benchmark suite. We demonstrate our method by a case study on a multi-tenant Supplier Business Management (SBM) service, as well as evaluate the capability of our benchmark suite through two sets of experiments. Results from these experiments characterized the relationship between workloads and performance, which can help find optimized deployment strategies for multi-tenant applications. In the case study on multi-tenant SBM service system, we gain an optimized strategy that satisfies the requirement of tenants and makes the maximum use of the resources, which can give useful recommendations in real service instances deployment stage.
Jizun Liu, Ze-yu Di, Shijun Liu, Calton Pu, Lei Wu 0002, Li Pan 0001
APSCC4
2014 Landslide Detection Service Based on Composition of Physical and Social Information Services
abstract
Social media have been used in the detection and management of natural hazards such as earthquakes. However, disasters often lead to other kinds of disasters, forming multi-hazards. Landslide is an illustrative example of a multi-hazard, which may be caused by earthquakes, rainfalls, water erosion, among other reasons. Detecting such multi-hazards is a significant challenge, since physical sensors designed for specific disasters are insufficient for multi-hazards. We describe LITMUS -- a landslide detection service based on a multi-service composition approach that combines data from both physical and social information services by filtering and then joining the information flow from those services based on their spatiotemporal features. Our results show that with such approach LITMUS detects 25 out of 27 landslides reported by USGS in December and 40 more landslides unreported by USGS. Also, LITMUS provides a live demonstration that displays results on a web map.
Aibek Musaev, De Wang, Chien-An Cho, Calton Pu
ICWS4
2014 Editorial
Lakshmish Ramaswamy, Barbara Carminati, Lujo Bauer, Dongwan Shin, James B. D. Joshi, Calton Pu, Dimitris Gritzalis
Comput. Secur.6
2014 Preface
Barbara Carminati, Lakshmish Ramaswamy, Anna Cinzia Squicciarini, James B. D. Joshi, Calton Pu
Int. J. Cooperative Inf. Syst.5
2014 A Perspective of Evolution After Five Years: A Large-Scale Study of Web Spam Evolution
abstract
Identifying and detecting web spam is an ongoing battle between spam-researchers and spammers which has been going on since search engines allowed searching of web pages to the modern sharing of web links via social networks. A common challenge faced by spam-researchers is the fact that new techniques depend on requiring a corpus of legitimate and spam web pages. Although large corpora of legitimate web pages are available to researchers, the same cannot be said about web spam or spam web pages. In this paper, we introduce the Webb Spam Corpus 2011 — a corpus of approximately 330,000 spam web pages — which we make available to researchers in the fight against spam. By having a standard corpus available, researchers can collaborate better on developing and reporting results of spam filtering techniques. The corpus contains web pages crawled from links found in over 6.3 million spam emails. We analyze multiple aspects of this corpus including redirection, HTTP headers, web page content, and classification evaluation. We also provide insights into changes in web spam since the last Webb Spam Corpus was released in 2006. These insights include: (1) spammers manipulate social media in spreading spam; (2) HTTP headers and content also change over time; (3) spammers have evolved and adopted new techniques to avoid the detection based on HTTP header information.
De Wang, Danesh Irani, Calton Pu
Int. J. Cooperative Inf. Syst.3
2014 Editorial: Collaborative Computing: Networking, Applications and Worksharing (CollaborateCom 2012)
Lakshmish Ramaswamy, Barbara Carminati, James B. D. Joshi, Calton Pu
Mob. Networks Appl.4
2014 Variations in Performance and Scalability: An Experimental Study in IaaS Clouds Using Multi-Tier Workloads
abstract
The increasing popularity of clouds drives researchers to find answers to a large variety of new and challenging questions. Through extensive experimental measurements, we show variance in performance and scalability of clouds for two non-trivial scenarios. In the first scenario, we target the public Infrastructure as a Service (IaaS) clouds, and study the case when a multi-tier application is migrated from a traditional datacenter to one of the three IaaS clouds. To validate our findings in the first scenario, we conduct similar study with three private clouds built using three mainstream hypervisors. We used the RUBBoS benchmark application and compared its performance and scalability when hosted in Amazon EC2, Open Cirrus, and Emulab. Our results show that a best-performing configuration in one cloud can become the worst-performing configuration in another cloud. Subsequently, we identified several system level bottlenecks such as high context switching and network driver processing overheads that degraded the performance. We experimentally evaluate concrete alternative approaches as practical solutions to address these problems. We then built the three private clouds using a commercial hypervisor (CVM), Xen, and KVM respectively and evaluated performance characteristics using both RUBBoS and Cloudstone benchmark applications. The three clouds show significant performance variations; for instance, Xen outperforms CVM by 75 percent on the read-write RUBBoS workload and CVM outperforms Xen by over 10 percent on the Cloudstone workload. These observed problems were confirmed at a finer granularity through micro-benchmark experiments that measure component performance directly.
Deepal Jayasinghe, Simon Malkowski, Jack Li 0001, Qingyang Wang 0001, Zhikui Wang, Calton Pu
IEEE Trans. Serv. Comput.6
2013 An Experimental Study of Rapidly Alternating Bottlenecks in n-Tier Applications
abstract
Identifying the location of performance bottlenecks is a non-trivial challenge when scaling n-tier applications in computing clouds. Specifically, we observed that an n-tier application may experience significant performance loss when bottlenecks alternate rapidly between component servers. Such rapidly alternating bottlenecks arise naturally and often from resource dependencies in an n-tier system and bursty workloads. These rapidly alternating bottlenecks are difficult to detect because the saturation in each participating server may have a very short lifespan (e.g., milliseconds) compared to current system monitoring tools and practices with sampling at intervals of seconds or minutes. Using passive network tracing at fine-granularity (e.g., aggregate at every 50ms), we are able to correlate throughput (i.e., request service rate) and load (i.e., number of concurrent requests) in each server of an n-tier system. Our experimental results show conclusive evidence of rapidly alternating bottlenecks caused by system software (JVM garbage collection) and middleware (VM collocation).
Qingyang Wang 0001, Yasuhiko Kanemasa, Jack Li 0001, Deepal Jayasinghe, Toshihiro Shimizu, Masazumi Matsubara, Motoyuki Kawaba, Calton Pu
IEEE CLOUD8
2013 An infrastructure for automating large-scale performance studies and data processing
abstract
The Cloud has enabled the computing model to shift from traditional data centers to publicly shared computing infrastructure; yet, applications leveraging this new computing model can experience performance and scalability issues, which arise from the hidden complexities of the cloud. The most reliable path for better understanding these complexities is an empirically based approach that relies on collecting data from a large number of performance studies. Armed with this performance data, we can understand what has happened, why it happened, and more importantly, predict what will happen in the future. However, this approach presents challenges itself, namely in the form of data management. We attempt to mitigate these data challenges by fully automating the performance measurement process. Concretely, we have developed an automated infrastructure, which reduces the complexity of the large-scale performance measurement process by generating all the necessary resources to conduct experiments, to collect and process data and to store and analyze data. In this paper, we focus on the performance data management aspect of our infrastructure.
Deepal Jayasinghe, Josh Kimball, Siddharth Choudhary, Calton Pu
IEEE BigData5
2013 Real-time collaborative planning with big data: Technical challenges and in-place computing (invited paper)
abstract
There is increasing collaboration in new generation supply chain planning applications, where participants across a supply chain analyze and plan on a big volume of sales data over the internet together. To achieve real-time collaborative planning over big data, we have developed an unconventional t
Wenwey Hseush, Yi-Cheng Huang, Shih-Chang Hsu, Calton Pu
CollaborateCom4
2013 A content-context-centric approach for detecting vandalism in Wikipedia
abstract
Collaborative online social media (CSM) applications such as Wikipedia have not only revolutionized the World Wide Web, but they also have had a hugely positive effect on modern free societies. Unfortunately, Wikipedia has also become target to a wide-variety of vandalism attacks. Most existing vand
Lakshmish Ramaswamy, Raga Sowmya Tummalapenta, Kang Li 0001, Calton Pu
CollaborateCom4
2013 A study on evolution of email spam over fifteen years
abstract
Email spam is a persistent problem, especially today, with the increasing dedication and sophistication of spammers. Even popular social media sites such as Facebook, Twitter, and Google Plus are not exempt from email spam as they all interface with email systems. With an "arms-race'' between spamme
De Wang, Danesh Irani, Calton Pu
CollaborateCom3
2013 Click traffic analysis of short URL spam on Twitter
abstract
With an average of 80% length reduction, the URL shorteners have become the norm for sharing URLs on Twitter, mainly due to the 140-character limit per message. Unfortunately, spammers have also adopted the URL shorteners to camouflage and improve the user click-through of their spam URLs. In this p
De Wang, Shamkant B. Navathe, Ling Liu 0001, Danesh Irani, Acar Tamersoy, Calton Pu
CollaborateCom6
2013 KQguard: Binary-Centric Defense against Kernel Queue Injection Attacks
Jinpeng Wei, Feng Zhu 0015, Calton Pu
ESORICS3
2013 Detecting Transient Bottlenecks in n-Tier Applications through Fine-Grained Analysis
abstract
Identifying the location of performance bottlenecks is a non-trivial challenge when scaling n-tier applications in computing clouds. Specifically, we observed that an n-tier application may experience significant performance loss when there are transient bottlenecks in component servers. Such transient bottlenecks arise frequently at high resource utilization and often result from transient events (e.g., JVM garbage collection) in an n-tier system and bursty workloads. Because of their short lifespan (e.g., milliseconds), these transient bottlenecks are difficult to detect using current system monitoring tools with sampling at intervals of seconds or minutes. We describe a novel transient bottleneck detection method that correlates throughput (i.e., request service rate) and load (i.e., number of concurrent requests) of each server in an n-tier system at fine time granularity. Both throughput and load can be measured through passive network tracing at millisecond-level time granularity. Using correlation analysis, we can identify the transient bottlenecks at time granularities as short as 50ms. We validate our method experimentally through two case studies on transient bottlenecks caused by factors at the system software layer (e.g., JVM garbage collection) and architecture layer (e.g., Intel SpeedStep).
Qingyang Wang 0001, Yasuhiko Kanemasa, Jack Li 0001, Deepal Jayasinghe, Toshihiro Shimizu, Masazumi Matsubara, Motoyuki Kawaba, Calton Pu
ICDCS8
2013 RoadAlarm: A spatial alarm system on road networks
abstract
Spatial alarms are one of the fundamental functionalities for many LBSs. We argue that spatial alarms should be road network aware as mobile objects travel on spatially constrained road networks or walk paths. In this software system demonstration, we will present the first prototype system of ROADALARM - a spatial alarm processing system for moving objects on road networks. The demonstration system of ROAD-ALARM focuses on the three unique features of ROADALARM system design. First, we will show that the road network distance-based spatial alarm is best modeled using road network distance such as segment length-based and travel time-based distance. Thus, a road network spatial alarm is a star-like subgraph centered at the alarm target. Second, we will show the suite of ROADALARM optimization techniques to scale spatial alarm processing by taking into account spatial constraints on road networks and mobility patterns of mobile subscribers. Third, we will show that, by equipping the ROADALARM system with an activity monitoring-based control panel, we are able to enable the system administrator and the end users to visualize road network-based spatial alarms, mobility traces of moving objects and dynamically make selection or customization of the ROADALARM techniques for spatial alarm processing through graphical user interface. We show that the ROADALARM system provides both the general system architecture and the essential building blocks for location-based advertisements and location-based reminders.
Kisung Lee, Emre Yigitoglu, Ling Liu 0001, Binh Han, Balaji Palanisamy, Calton Pu
ICDE6
2013 Road network mix-zones for anonymous location based services
abstract
We present MobiMix, a road network based mix-zone framework to protect location privacy of mobile users traveling on road networks. An alternative and complementary approach to spatial cloaking based location privacy protection is to break the continuity of location exposure by introducing techniques, such as mix-zones, where no applications can trace user movements. However, existing mixzone proposals fail to provide effective mix-zone construction and placement algorithms that are resilient to timing and transition attacks. In MobiMix, mix-zones are constructed and placed by carefully taking into consideration of multiple factors, such as the geometry of the zones, the statistical behavior of the user population, the spatial constraints on movement patterns of the users, and the temporal and spatial resolution of the location exposure. In this demonstration, we first introduce a visualization of the location privacy risks of mobile users traveling on road networks and show how mixzone based anonymization breaks the continuity of location exposure to protect user location privacy. We demonstrate a suite of road network mix-zone construction and placement methods that provide higher level of resilience to timing and transition attacks on road networks. We show the effectiveness of the MobiMix approach through detailed visualization using traces produced by GTMobiSim on different scales of geographic maps.
Balaji Palanisamy, Sindhuja Ravichandran, Ling Liu 0001, Binh Han, Kisung Lee, Calton Pu
ICDE6
2013 vPerfGuard: an automated model-driven framework for application performance diagnosis in consolidated cloud environments
abstract
Many business customers hesitate to move all their applications to the cloud due to performance concerns. White-box diagnosis relies on human expert experience or performance troubleshooting "cookbooks" to find potential performance bottlenecks. Despite wide adoption, the scalability and adaptivity of such approaches remain severely constrained, especially in a highly-dynamic, consolidated cloud environment. Leveraging the rich telemetry collected from applications and systems in the cloud, and the power of statistical learning, vPerfGuard complements the existing approaches with a model-driven framework by: (1) automatically identifying system metrics that are most predictive of application performance, and (2) adaptively detecting changes in the performance and potential shifts in the predictive metrics that may accompany such a change. Although correlation does not imply causation, the predictive system metrics point to potential causes that can guide a cloud service provider to zero in on the root cause.
PengCheng Xiong, Calton Pu, Xiaoyun Zhu, Rean Griffith
ICPE2
2013 Editorial for CollaborateCom 2011 Special Issue
James Caverlee, Calton Pu, Dimitrios Georgakopoulos 0001, James B. D. Joshi
Mob. Networks Appl.2
2013 Who Is Your Neighbor: Net I/O Performance Interference in Virtualized Clouds
abstract
User-perceived performance continues to be the most important QoS indicator in cloud-based data centers today. Effective allocation of virtual machines (VMs) to handle both CPU intensive and I/O intensive workloads is a crucial performance management capability in virtualized clouds. Although a fair amount of researches have dedicated to measuring and scheduling jobs among VMs, there still lacks of in-depth understanding of performance factors that impact the efficiency and effectiveness of resource multiplexing and scheduling among VMs. In this paper, we present the experimental research on performance interference in parallel processing of CPU-intensive and network-intensive workloads on Xen virtual machine monitor (VMM). Based on our study, we conclude with five key findings which are critical for effective performance management and tuning in virtualized clouds. First, colocating network-intensive workloads in isolated VMs incurs high overheads of switches and events in Dom0 and VMM. Second, colocating CPU-intensive workloads in isolated VMs incurs high CPU contention due to fast I/O processing in I/O channel. Third, running CPU-intensive and network-intensive workloads in conjunction incurs the least resource contention, delivering higher aggregate performance. Fourth, performance of network-intensive workload is insensitive to CPU assignment among VMs, whereas adaptive CPU assignment among VMs is critical to CPU-intensive workload. The more CPUs pinned on Dom0 the worse performance is achieved by CPU-intensive workload. Last, due to fast I/O processing in I/O channel, limitation on grant table is a potential bottleneck in Xen. We argue that identifying the factors that impact the total demand of exchanged memory pages is important to the in-depth understanding of interference costs in Dom0 and VMM.
Xing Pu, Ling Liu 0001, Yiduo Mei, Sankaran Sivathanu, Younggyun Koh, Calton Pu, Yuanda Cao
IEEE Trans. Serv. Comput.6
2012 Expertus: A Generator Approach to Automate Performance Testing in IaaS Clouds
abstract
Cloud computing is an emerging technology paradigm that revolutionizes the computing landscape by providing on-demand delivery of software, platform, and infrastructure over the Internet. Yet, architecting, deploying, and configuring enterprise applications to run well on modern clouds remains a challenge due to associated complexities and non-trivial implications. The natural and presumably unbiased approach to these questions is thorough testing before moving applications to production settings. However, thorough testing of enterprise applications on modern clouds is cumbersome and error-prone due to a large number of relevant scenarios and difficulties in testing process. We address some of these challenges through Expertus---a flexible code generation framework for automated performance testing of distributed applications in Infrastructure as a Service (IaaS) clouds. Expertus uses a multi-pass compiler approach and leverages template-driven code generation to modularly incorporate different software applications on IaaS clouds. Expertus automatically handles complex configuration dependencies of software applications and significantly reduces human errors associated with manual approaches for software configuration and testing. To date, Expertus has been used to study three distributed applications on five IaaS clouds with over 10,000 different hardware, software, and virtualization configurations. The flexibility and extensibility of Expertus and our own experience on using it shows that new clouds, applications, and software packages can easily be incorporated.
Deepal Jayasinghe, Galen S. Swint, Simon Malkowski, Jack Li 0001, Qingyang Wang 0001, Junhee Park, Calton Pu
IEEE CLOUD7
2012 Challenges and Opportunities in Consolidation at High Resource Utilization: Non-monotonic Response Time Variations in n-Tier Applications
abstract
A central goal of cloud computing is high resource utilization through hardware sharing; however, utilization often remains modest in practice due to the challenges in predicting consolidated application performance accurately. We present a thorough experimental study of consolidated n-tier application performance at high utilization to address this issue through reproducible measurements. Our experimental method illustrates opportunities for increasing operational efficiency by making consolidated application performance more predictable in high utilization scenarios. The main focus of this paper are non-trivial dependencies between SLA-critical response time degradation effects and software configurations (i.e., readily available tuning knobs). Methodologically, we directly measure and analyze the resource utilizations, request rates, and performance of two consolidated n-tier application benchmark systems (RUBBoS) in an enterprise-level computer virtualization environment. We find that monotonically increasing the workload of an n-tier application system may unexpectedly spike the overall response time of another co-located system by 300 percent despite stable throughput. Based on these findings, we derive a software configuration best-practice to mitigate such non-monotonic response time variations by enabling higher request-processing concurrency (e.g., more threads) in all tiers. More generally, this experimental study increases our quantitative understanding of the challenges and opportunities in the widely used (but seldom supported, quantified, or even mentioned) hypothesis that applications consolidate with linear performance in cloud environments.
Simon Malkowski, Yasuhiko Kanemasa, Hanwei Chen, Masao Yamamoto, Qingyang Wang 0001, Deepal Jayasinghe, Calton Pu, Motoyuki Kawaba
IEEE CLOUD7
2012 Performance Analysis of Parallel Processing Systems with Horizontal Decomposition
abstract
Parallel processing is an important pattern in cluster systems. To analyze the performance of parallel processing systems, we leveraged the fork-join queueing network (FJQN) models. However, there are no easy solutions to these models, especially for the multi-class closed ones. In this paper, a novel and efficient method named horizontal decomposition has been proposed. The main idea of our method is to approximate a non-product-form FJQN with some closed and open product-form networks. So the computational complexity can be dramatically reduced compared with the traditional hierarchical decomposition approach. And the algorithms for solving single-class and multi-class closed FJQNs have been developed respectively based on the horizontal decomposition. With these algorithms, the response time and throughput of each service center in a FJQN can be approximately calculated. The evaluation results show that 90 percentile of relative errors of most service centers are less than 15% except for the shared ones. The evaluation results also showed that the number of iterations in the algorithm for the multi-class FJQNs almost grows linearly with the population of networks.
Hanwei Chen, Jianwei Yin, Calton Pu
CLUSTER3
2012 Preface
Barbara Carminati, Lakshmish Ramaswamy, Calton Pu, James B. D. Joshi
CollaborateCom3
2012 Evolutionary study of web spam: Webb Spam Corpus 2011 versus Webb Spam Corpus 2006
abstract
With over 2.5 hours a day spent browsing websites online[1] and with over a billion pages[2], identifying and detecting web spam is an important problem. Although large corpora of legitimate web pages are available to researchers, the same cannot be said about web spam or spam web pages. We int
De Wang, Danesh Irani, Calton Pu
CollaborateCom3
2012 Transactional Recovery Support for Robust Exception Handling in Business Process Services
abstract
Building mission critical applications and services (e.g. e-commerce) using process-oriented approaches has had successes and difficulties. These applications automated successfully the important frequent cases such as purchases, but the code needed for handling exceptions such as cancellations and failures tend to grow to disproportionate size and complexity. These difficulties lead to non-automated and expensive solutions such as call centers, which resolve data inconsistency problems manually. In this paper we describe the WED-flow (work, event, and data-flow) approach, which provides transactional recovery through incremental evolution of exception handling, by combining the concepts of advanced transaction models, events, and data states. By carefully recording the detailed data states of each execution step, WED-flow composes backward and forward recovery mechanisms as reusable exception handling services to preserve the consistency of all databases involved in the application with well-defined correctness properties. A practical application of the automated recovery in WED-flow is the real-time recovery of failed cases for mission-critical applications and services.
João Eduardo Ferreira, Kelly Rosa Braghetto, Osvaldo Kotaro Takai, Calton Pu
ICWS4
2012 Response Time Reliability in Cloud Environments: An Empirical Study of n-Tier Applications at High Resource Utilization
abstract
When running mission-critical web-facing applications (e.g., electronic commerce) in cloud environments, predictable response time, e.g., specified as service level agreements (SLA), is a major performance reliability requirement. Through extensive measurements of n-tier application benchmarks in a cloud environment, we study three factors that significantly impact the application response time predictability: bursty workloads (typical of web-facing applications), soft resource management strategies (e.g., global thread pool or local thread pool), and bursts in system software consumption of hardware resources (e.g., Java Virtual Machine garbage collection). Using a set of profit-based performance criteria derived from typical SLAs, we show that response time reliability is brittle, with large response time variations (order of several seconds) depending on each one of those factors. For example, for the same workload and hardware platform, modest increases in workload burstiness may result in profit drops of more than 50%. Our results show that profitbased performance criteria may contribute significantly to the successful delimitation of performance unreliability boundaries and thus support effective management of clouds.
Qingyang Wang 0001, Yasuhiko Kanemasa, Jack Li 0001, Deepal Jayasinghe, Motoyuki Kawaba, Calton Pu
SRDS6
2012 Toward a general defense against kernel queue hooking attacks
Jinpeng Wei, Calton Pu
Comput. Secur.2
2012 ACM/Springer Mobile Networks and Applications (MONET) Special Issue on "Collaborative Computing: Networking, Applications and Worksharing"
James B. D. Joshi, Elisa Bertino, Calton Pu, Heri Ramampiaro
Mob. Networks Appl.3
2012 Scaling Group Communication Services with Self-adaptive and Utility-driven Message Routing
Yuehua Wang, Ling Liu 0001, Calton Pu
Mob. Networks Appl.3
2011 Variations in Performance and Scalability When Migrating n-Tier Applications to Different Clouds
abstract
The increasing popularity of computing clouds continues to drive both industry and research to provide answers to a large variety of new and challenging questions. We aim to answer some of these questions by evaluating performance and scalability when an n-tier application is migrated from a traditional datacenter environment to an IaaS cloud. We used a representative n-tier macro-benchmark (RUBBoS) and compared its performance and scalability in three different test beds: Amazon EC2, Open Cirrus (an open scientific research cloud), and Emulab (academic research test bed). Interestingly, we found that the best-performing configuration in Emulab can become the worst-performing configuration in EC2. Subsequently, we identified the bottleneck components, high context switch overhead and network driver processing overhead, to be at the system level. These overhead problems were confirmed at a finer granularity through micro-benchmark experiments that measure component performance directly. We describe concrete alternative approaches as practical solutions for resolving these problems.
Deepal Jayasinghe, Simon Malkowski, Qingyang Wang 0001, Jack Li 0001, PengCheng Xiong, Calton Pu
IEEE CLOUD6
2011 ActiveSLA: a profit-oriented admission control framework for database-as-a-service providers
abstract
The system overload is a common problem in a Database-as-a-Serice (DaaS) environment because of unpredictable and bursty workloads from various clients. Due to the service delivery nature of DaaS, such system overload usually has direct economic impact on the service provider, who has to pay penalties if the system performance does not meet clients' service level agreements (SLAs). In this paper, we investigate techniques that prevent system overload by using admission control. We propose a profit-oriented admission control framework, called ActiveSLA, for DaaS providers. ActiveSLA is an end-to-end framework that consists of two components. First, a prediction module estimates the probability for a new query to finish the execution before its deadline. Second, based on the predicted probability, a decision module determines whether or not to admit the given query into the database system. The decision is made with the profit optimization objective, where the expected profit is derived from the service level agreements between a service provider and its clients. We present extensive real system experiments with standard database benchmarks, under different traffic patterns, DBMS settings, and SLAs. The results demonstrate that ActiveSLA is able to make admission control decisions that are both more accurate and more profit-effective than several state-of-the-art methods.
PengCheng Xiong, Yun Chi, Shenghuo Zhu, Jun'ichi Tatemura, Calton Pu, Hakan Hacigümüs
SoCC5
2011 Reverse Social Engineering Attacks in Online Social Networks
Danesh Irani, Marco Balduzzi, Davide Balzarotti, Engin Kirda, Calton Pu
DIMVA5
2011 Economical and Robust Provisioning of N-Tier Cloud Workloads: A Multi-level Control Approach
abstract
Resource provisioning for N-tier web applications in Clouds is non-trivial due to at least two reasons. First, there is an inherent optimization conflict between cost of resources and Service Level Agreement (SLA) compliance. Second, the resource demands of the multiple tiers can be different from each other, and varying along with the time. Resources have to be allocated to multiple (virtual) containers to minimize the total amount of resources while meeting the end-to-end performance requirements for the application. In this paper we address these two challenges through the combination of the resource controllers on both application and container levels. On the application level, a decision maker (i.e., an adaptive feedback controller) determines the total budget of the resources that are required for the application to meet SLA requirements as the workload varies. On the container level, a second controller partitions the total resource budget among the components of the applications to optimize the application performance (i.e., to minimize the round trip time). We evaluated our method with three different workload models -- open, closed, and semi-open - that were implemented in the RUBiS web application benchmark. Our evaluation indicates two major advantages of our method in comparison to previous approaches. First, fewer resources are provisioned to the applications to achieve the same performance. Second, our approach is robust enough to address various types of workloads with time-varying resource demand without reconfiguration.
PengCheng Xiong, Zhikui Wang, Simon Malkowski, Qingyang Wang 0001, Deepal Jayasinghe, Calton Pu
ICDCS6
2011 Intelligent management of virtualized resources for database systems in cloud environment
abstract
In a cloud computing environment, resources are shared among different clients. Intelligently managing and allocating resources among various clients is important for system providers, whose business model relies on managing the infrastructure resources in a cost-effective manner while satisfying the client service level agreements (SLAs). In this paper, we address the issue of how to intelligently manage the resources in a shared cloud database system and present SmartSLA, a cost-aware resource management system. SmartSLA consists of two main components: the system modeling module and the resource allocation decision module. The system modeling module uses machine learning techniques to learn a model that describes the potential profit margins for each client under different resource allocations. Based on the learned model, the resource allocation decision module dynamically adjusts the resource allocations in order to achieve the optimum profits. We evaluate SmartSLA by using the TPC-W benchmark with workload characteristics derived from real-life systems. The performance results indicate that SmartSLA can successfully compute predictive models under different hardware resource allocations, such as CPU and memory, as well as database specific resources, such as the number of replicas in the database systems. The experimental results also show that SmartSLA can provide intelligent service differentiation according to factors such as variable workloads, SLA levels, resource costs, and deliver improved profit margins.
PengCheng Xiong, Yun Chi, Shenghuo Zhu, Hyun Jin Moon, Calton Pu, Hakan Hacigümüs
ICDE5
2011 The Impact of Soft Resource Allocation on n-Tier Application Scalability
abstract
Good performance and efficiency, in terms of high quality of service and resource utilization for example, are important goals in a cloud environment. Through extensive measurements of an n-tier application benchmark (RUBBoS), we show that overall system performance is surprisingly sensitive to appropriate allocation of soft resources (e.g., server thread pool size). Inappropriate soft resource allocation can quickly degrade overall application performance significantly. Concretely, both under-allocation and over-allocation of thread pool can lead to bottlenecks in other resources because of non-trivial dependencies. We have observed some non-obvious phenomena due to these correlated bottlenecks. For instance, the number of threads in the Apache web server can limit the total useful throughput, causing the CPU utilization of the C-JDBC clustering middleware to decrease as the workload increases. We provide a practical iterative solution approach to this challenge through an algorithmic combination of operational queuing laws and measurement data. Our results show that soft resource allocation plays a central role in the performance scalability of complex systems such as n-tier applications in cloud environments.
Qingyang Wang 0001, Simon Malkowski, Deepal Jayasinghe, PengCheng Xiong, Calton Pu, Yasuhiko Kanemasa, Motoyuki Kawaba, Lilian Harada
IPDPS5
2010 Understanding Performance Interference of I/O Workload in Virtualized Cloud Environments
abstract
Server virtualization offers the ability to slice large, underutilized physical servers into smaller, parallel virtual machines (VMs), enabling diverse applications to run in isolated environments on a shared hardware platform. Effective management of virtualized cloud environments introduces new and unique challenges, such as efficient CPU scheduling for virtual machines, effective allocation of virtual machines to handle both CPU intensive and I/O intensive workloads. Although a fair number of research projects have dedicated to measuring, scheduling, and resource management of virtual machines, there still lacks of in-depth understanding of the performance factors that can impact the efficiency and effectiveness of resource multiplexing and resource scheduling among virtual machines. In this paper, we present our experimental study on the performance interference in parallel processing of CPU and network intensive workloads in the Xen Virtual Machine Monitors (VMMs). We conduct extensive experiments to measure the performance interference among VMs running network I/O workloads that are either CPU bound or network bound. Based on our experiments and observations, we conclude with four key findings that are critical to effective management of virtualized cloud environments for both cloud service providers and cloud consumers. First, running network-intensive workloads in isolated environments on a shared hardware platform can lead to high overheads due to extensive context switches and events in driver domain and VMM. Second, co-locating CPU-intensive workloads in isolated environments on a shared hardware platform can incur high CPU contention due to the demand for fast memory pages exchanges in I/O channel. Third, running CPU-intensive workloads and network-intensive workloads in conjunction incurs the least resource contention, delivering higher aggregate performance. Last but not the least, identifying factors that impact the total demand of the exchanged memory pages is critical to the in-depth understanding of the interference overheads in I/O channel in the driver domain and VMM.
Xing Pu, Ling Liu 0001, Yiduo Mei, Sankaran Sivathanu, Younggyun Koh, Calton Pu
IEEE CLOUD6
2010 Elusive vandalism detection in wikipedia: a text stability-based approach
abstract
The open collaborative nature of wikis encourages participation of all users, but at the same time exposes their content to vandalism. The current vandalism-detection techniques, while effective against relatively obvious vandalism edits, prove to be inadequate in detecting increasingly prevalent sophisticated (or elusive) vandal edits. We identify a number of vandal edits that can take hours, even days, to correct and propose a text stability-based approach for detecting them. Our approach is focused on the likelihood of a certain part of an article being modified by a regular edit. In addition to text-stability, our machine learning-based technique also takes into account edit patterns. We evaluate the performance of our approach on a corpus comprising of 15000 manually labeled edits from the Wikipedia Vandalism PAN corpus. The experimental results show that text-stability is able to improve the performance of the selected machine-learning algorithms significantly.
Qinyi Wu, Danesh Irani, Calton Pu, Lakshmish Ramaswamy
CIKM3
2010 Modeling the Runtime Integrity of Cloud Servers: A Scoped Invariant Perspective
abstract
One of the underpinnings of Cloud Computing security is the runtime integrity of individual Cloud servers. Due to the on-going discovery of runtime software vulnerabilities like buffer overflows, it is critical to be able to gauge the integrity of a Cloud server as it operates. In this paper, we propose scoped invariants as a primitive for analyzing the software system for its integrity properties. We report our experience with the modeling and detection of scoped invariants. The Xen Virtual Machine Manager is used for a case study. Our research detects a set of essential scoped invariants that are critical to the runtime integrity of Xen. One such property, that the addressable memory limit of a guest OS must not include Xen's code and data, is indispensable for Xen's guest isolation mechanism. The violation of this property demonstrates that the attacker only needs to modify a single byte in the Global Descriptor Table to achieve his goal.
Jinpeng Wei, Calton Pu, Carlos V. Rozas, Anand Rajan, Feng Zhu 0015
CloudCom2
2010 Using LOTOS for rigorous specifications of workflow patterns
abstract
Collaborative applications require understanding of the theoretical foundations. In case of workflow systems, one possibility to achieve this is an accurate description of workflow functionalities. Despite its growing popularity and success, it has not yet been evaluated whether Language of Temporal
Pedro Losco Takecian, João Eduardo Ferreira, Simon Malkowski, Calton Pu
CollaborateCom4
2010 An utility-driven routing scheme for scaling multicast applications
abstract
Multicast is a common platform for supporting group communication applications, such as IPTV, multimedia content delivery, and location-based advertisements. Distributed hash table (DHT) based overlay networks such as Chord and CAN presents a popular distributed computing architecture for multicast
Yuehua Wang, Ling Liu 0001, Calton Pu, Gong Zhang 0008
CollaborateCom3
2010 Modeling and implementing collaborative editing systems with transactional techniques
abstract
Many collaborative editing systems have been developed for coauthoring documents. These systems generally have different infrastructures and support a subset of interactions found in collaborative environments. In this paper, we propose a transactional framework with two advantages. First, the frame
Qinyi Wu, Calton Pu
CollaborateCom2
2010 Performance and availability aware regeneration for cloud based multitier applications
abstract
Virtual machine technology enables agile system deployments in which software components can be cheaply moved, replicated, and allocated hardware resources in a controlled fashion. This paper examines how these facilities can be used to provide enhanced solutions to the classic problem of ensuring high availability while maintaining performance. By regenerating software components to restore the redundancy of a system whenever failures occur, we achieve improved availability compared to a system with a fixed redundancy level. Moreover, by smartly controlling component placement and resource allocation using information about application control flow and performance predictions from queuing models, we ensure that the resulting performance degradation is minimized. We consider an environment in which a collection of multitier enterprise applications operates across multiple hosts, racks, clusters, and data centers to maximize failure independence. Simulation results show that our proposed approach provides better availability and significantly lower degradation of system response times compared to traditional approaches.
Gueyoung Jung, Kaustubh R. Joshi, Matti A. Hiltunen, Richard D. Schlichting, Calton Pu
DSN5
2010 Mistral: Dynamically Managing Power, Performance, and Adaptation Cost in Cloud Infrastructures
abstract
Server consolidation based on virtualization is an important technique for improving power efficiency and resource utilization in cloud infrastructures. However, to ensure satisfactory performance on shared resources under changing application workloads, dynamic management of the resource pool via online adaptation is critical. The inherent tradeoffs between power and performance as well as between the cost of an adaptation and its benefits make such management challenging. In this paper, we present Mistral, a holistic controller framework that optimizes power consumption, performance benefits, and the transient costs incurred by various adaptations and the controller itself to maximize overall utility. Mistral can handle multiple distributed applications and large-scale infrastructures through a multi-level adaptation hierarchy and scalable optimization algorithm. We show that our approach outstrips other strategies that address the tradeoff between only two of the objectives (power, performance, and transient costs).
Gueyoung Jung, Matti A. Hiltunen, Kaustubh R. Joshi, Richard D. Schlichting, Calton Pu
ICDCS5
2010 A partial persistent data structure to support consistency in real-time collaborative editing
abstract
Co-authored documents are becoming increasingly important for knowledge representation and sharing. Tools for supporting document co-authoring are expected to satisfy two requirements: 1) querying changes over editing histories; 2) maintaining data consistency among users. Current tools support either limited queries or are not suitable for loosely controlled collaborative editing scenarios. We address both problems by proposing a new persistent data structure-partial persistent sequence. The new data structure enables us to create unique character identifiers that can be used for associating meta-information and tracking their changes, and also design simple view synchronization algorithms to guarantee data consistency under the presence of concurrent updates. Experiments based on real-world collaborative editing traces show that our data structure uses disk space economically and provides efficient performance for document update and retrieval.
Qinyi Wu, Calton Pu, João Eduardo Ferreira
ICDE2
2010 Study of Static Classification of Social Spam Profiles in MySpace
Danesh Irani, Steve Webb, Calton Pu
ICWSM3
2010 Study on performance management and application behavior in virtualized environment
abstract
Control theory has been utilized in recent years to manage the resources in virtualized environment for applications with time-varying resource demand. The systems under control, including the servers and the applications, are taken as black-boxes, and the controllers are generally expected to be adaptive to the underline systems. However, little attention has been paid to the behaviors of the applications themselves, and most of time, single performance target such as the mean response time threshold has been tracked. In this paper, we experimentally show that more than one performance metrics have to be considered to characterize the quality of service that the end users receive when the performance is managed through dynamic resource allocation. Moreover, the behavior of the applications, especially that of the workload generators has significant effect on the quality of the service. Our study provides insights and guidance for end-to-end performance management problem in virtualized environment.
PengCheng Xiong, Zhikui Wang, Gueyoung Jung, Calton Pu
NOMS4
2010 Towards Flexible Event-Handling in Workflows through Data States
abstract
Despite recent advances in many real-time and workflow management systems (WFMS), event-handling is still a manual or semi-automated task. The integration of automated event processing with workflows remains an open research challenge to both academic and industrial communities. In this work, we propose a concrete approach that logs interactions between workflow component activities in the form of data states that accurately and efficiently store necessary information for event-handling. Our approach (called WED-flow) explicitly represents various dependencies and constraints of a WFMS in sophisticated data states. Due to the availability of this large amount of historic information, our approach is able to support a flexible event-handling in WFMS. In this paper we present definitions for workflow management systems that incorporate events, and characterize such systems using the WED-flow approach. We also present a scientific workflow example in genetic testing to illustrate the advantages of integrating events with workflow through the WED-flow approach.
João Eduardo Ferreira, Qinyi Wu, Simon Malkowski, Calton Pu
SERVICES4
2010 Modeling and preventing TOCTTOU vulnerabilities in Unix-style file systems
Jinpeng Wei, Calton Pu
Comput. Secur.2
2009 Towards Algorithmic Generation of Business Processes: From Business Step Dependencies to Process Algebra Expressions
Marcio K. Oikawa, João Eduardo Ferreira, Simon Malkowski, Calton Pu
BPM4
2009 JTangSynergy 3.0: A framework and software tool for integrating cross-organizational applications
abstract
Enterprise service bus is one of the most promising infrastructures for simplifying enterprise application integration (EAI). However, in order to exploit its full potential for integrating large scale cross-organizational applications, a flexible and reliable distributed environment management infr
Hanwei Chen, Jianwei Yin, Calton Pu
CollaborateCom4
2009 A new perspective on experimental analysis of N-tier systems: Evaluating database scalability, multi-bottlenecks, and economical operation
abstract
Economical configuration planning, component performance evaluation, and analysis of bottleneck phenomena in N-tier applications are serious challenges due to design requirements such as non-stationary workloads, complex non-modular relationships, and global consistency management when replicating d
Simon Malkowski, Markus Hedwig, Deepal Jayasinghe, Junhee Park, Yasuhiko Kanemasa, Calton Pu
CollaborateCom6
2009 A General Proximity Privacy Principle
abstract
This work presents a systematic study of the problem of protecting general proximity privacy, with findings applicable to most existing data models. Our contributions are multi-folded: we highlighted and formulated proximity privacy breaches in a data-model-neutral manner; we proposed a new privacy principle (epsiv,delta)k-dissimilarity, with theoretically guaranteed protection against linking attacks in terms of both exact and proximate QI-SA associations; we provided a theoretical analysis regarding the satisfiability of (epsiv,delta)k-dissimilarity, and pointed to promising solutions to fulfilling this principle.
Ting Wang 0006, Shicong Meng, Bhuvan Bamba, Ling Liu 0001, Calton Pu
ICDE5
2009 A Petri Net Siphon Based Solution to Protocol-Level Service Composition Mismatches
abstract
Protocol-level mismatch is one of the most important problems in service composition. The commonly used reachability exploration method focuses on verifying deadlock-freeness. When this property is violated, the states and traces in the reachability graph only give clues to re-design the composition. The process must then repeat itself until no deadlock is found. In this paper, multiple Web service interaction is modeled with a Petri net called composition net (C-net). The protocol-level mismatch problem is transformed into the deadlock structure problem of a C-net. If mismatches are found, a solution based on Petri net siphons is proposed. The proposed method is shown to achieve higher efficiency for resolving protocol-level mismatching issues than traditional ones do.
PengCheng Xiong, MengChu Zhou, Calton Pu
ICWS3
2009 A Cost-Sensitive Adaptation Engine for Server Consolidation of Multitier Applications
Gueyoung Jung, Kaustubh R. Joshi, Matti A. Hiltunen, Richard D. Schlichting, Calton Pu
Middleware5
2009 Improving Virtualized Windows Network Performance by Delegating Network Processing
abstract
Virtualized environments are important building blocks in consolidated data centers and cloud computing. Full virtualization (FV) allows unmodified guest OSes to run on virtualization-aware microprocessors. However, the significant overhead of device emulation in FV has caused high I/O overhead. Current implementations based on paravirtualization can only reduce such overhead partially. This paper describes the Linsock approach that applies the outsourcing method to speed up I/O in FV environments by combining different guest OS and host OS. Concretely, Linsock replaces the guest Windowspsila network processing with the host Linux kernel on the same machine. Linsock has been implemented on Linux Kernel-based Virtual Machine (KVM) as the host virtual machine (VM) environment. Our measurement results with Linsock show significant performance increase of more than 300% compared with device paravirtualization in a 10 Gbps Ethernet networking environment. In addition, Linsock also yields a fourfold increase in inter-VM communication performance.
Younggyun Koh, Calton Pu, Yasushi Shinjo, Hideki Eiraku, Go Saito, Daiyuu Nobori
NCA2
2008 Soft-Timer Driven Transient Kernel Control Flow Attacks and Defense
abstract
A new class of stealthy kernel-level malware, called transient kernel control flow attacks, uses dynamic soft timers to achieve significant work while avoiding any persistent changes to kernel code or data. We demonstrate that soft timers can be used to implement attacks such as a stealthy key logger and a CPU cycle stealer. To defend against these attacks, we propose an approach based on static analysis of the entire kernel, which identifies and catalogs all legitimate soft timer interrupt requests (STIR) in a database. At run-time, a reference monitor in a trusted virtual machine compares each STIR with the database, only allowing the execution of known good STIRs. Our defensive technique has no false negatives because it mediates every STIR execution and prevents execution of all unknown, illegitimate STIRs, and no false positives because the relevant kernel code analyzed was unambiguous. The overhead for this additional security is less than 7% for each of our benchmarks.
Jinpeng Wei, Bryan D. Payne, Jonathon Giffin, Calton Pu
ACSAC4
2008 Predicting web spam with HTTP session information
abstract
Web spam is a widely-recognized threat to the quality and security of the Web. Web spam pages pollute search engine indexes, burden Web crawlers and Web mining services, and expose users to dangerous Web-borne malware. To defend against Web spam, most previous research analyzes the contents of Web pages and the link structure of the Web graph. Unfortunately, these heavyweight approaches require full downloads of both legitimate and spam pages to be effective, making real-time deployment of these techniques infeasible for Web browsers, high-performance Web crawlers, and real-time Web applications. In this paper, we present a lightweight, predictive approach to Web spam classification that relies exclusively on HTTP session information (i.e., hosting IP addresses and HTTP session headers).
Steve Webb, James Caverlee, Calton Pu
CIKM3
2008 PeerCast: Churn-resilient end system multicast on heterogeneous overlay networks
Jianjun Zhang 0001, Ling Liu 0001, Lakshmish Ramaswamy, Calton Pu
J. Netw. Comput. Appl.4
2008 Remote specialization for efficient embedded operating systems
abstract
Prior to their deployment on an embedded system, operating systems are commonly tailored to reduce code size and improve runtime performance. Program specialization is a promising match for this process: it is predictable and modules, and it allows the reuse of previously implemented specializations. A specialization engine for embedded systems must overcome three main obstacles: (i) Reusing existing compilers for embedded systems, (ii) supporting specialization on a resource-limited system and (iii) coping with dynamic applications by supporting specialization on demand. In this article, we describe a runtime specialization infrastructure that addresses these problems. Our solution proposes: (i) Specialization in two phases of which the former generates specialized C templates and the latter uses a dedicated compiler to generate efficient native code. (ii) A virtualization mechanism that facilitates specialization of code at a remote location. (iii) An API and supporting OS extensions that allow applications to produce, manage and dispose of specialized code. We evaluate our work through two case studies: (i) The TCP/IP implementation of Linux and (ii) The TUX embedded web server. We report appreciable improvements in code size and performance. We also quantify the overhead of specialization and argue that a specialization server can scale to support a sizable workload.
Sapan Bhatia, Charles Consel, Calton Pu
ACM Trans. Program. Lang. Syst.3
2008 A Secure Information Flow Architecture for Web Service Platforms
abstract
Current Web service platforms (WSPs) often perform all Web services-related processing, including security-sensitive information handling, in the same protection domain. Consequently, the entire WSP may have access to security-sensitive information, forcing us to trust a large and complex piece of software. To address this problem, we propose ISO-WSP, a new information flow architecture that decomposes current WSPs into a small trusted T-WSP to handle security-sensitive data and a large, legacy untrusted U-WSP that provides the normal WSP functionality. To achieve end-to-end security, the application code is also decomposed into a small trusted part and the remaining untrusted code. The trusted part encapsulates all accesses to security-sensitive data through a secure functional interface (SFI). To ease the migration of legacy applications to ISO-WSP, we developed tools to translate direct manipulations of security-sensitive data by the untrusted part into SFI invocations. Using a prototype implementation based on the Apache Axis2 WSP, we show that ISO-WSP reduces software complexity of trusted components by a factor of five, while incurring a modest performance overhead of few milliseconds per request. We also show that existing applications can be migrated to run on ISO-WSP with a few tens of lines of new and modified code.
Jinpeng Wei, Lenin Singaravelu, Calton Pu
IEEE Trans. Serv. Comput.3
2007 Multiprocessors May Reduce System Dependability under File-Based Race Condition Attacks
abstract
Attacks exploiting race conditions have been considered rare and "low risk". However, the increasing popularity of multiprocessors has changed this situation: instead of waiting for the victim process to be suspended to carry out an attack, the attacker can now run on a dedicated processor and actively seek attack opportunities. This change from fortuitous encountering to active exploiting may greatly increase the success probability of race condition attacks. This point is exemplified by studying the TOCTTOU (Time-of- Check-to-Time-of-Use) race condition attacks in this paper. We first propose a probabilistic model for predicting TOCTTOU attack success rate on both uniprocessors and multiprocessors. Then we confirm the applicability of this model by carrying out TOCTTOU attacks against two widely used utility programs: vi and gedit. The success probability of attacking vi increases from low single digit percentage on a uniprocessor to almost 100% on a multiprocessor. Similarly, the success rate of attacking gedit jumps from almost zero to 83%. These case studies suggest that our model captures the sharply increased risks, and hence the decreased dependability of our systems, represented by race condition attacks such as TOCTTOU on the next generation multiprocessors.
Jinpeng Wei, Calton Pu
DSN2
2007 Categorization and Optimization of Synchronization Dependencies in Business Processes
abstract
The current approach for modeling synchronization in business processes relies on sequencing constructs, such as sequence, parallel etc. However, sequencing constructs obfuscate the true source of dependencies in a business process. Moreover, because of the nested structure and scattered code that results from using sequencing constructs, it is hard to add or delete additional constraints without over-specifying necessary constraints or invalidating existing ones. We propose a dataflow programming approach in which dependencies are explicitly modeled to guide activity scheduling. We first give a systematic categorization of dependencies: data, control, service and cooperation. Each dimension models dependency from its own point of view. Then we show that dependencies of various kinds can be first merged and then optimized to generate a minimal dependency set, which guarantees high concurrency and minimal maintenance cost for process execution.
Qinyi Wu, Calton Pu, Akhil Sahai, Roger S. Barga
ICDE2
2007 Services Computing in Daily Work: Service Engineering vs. Software Engineering
abstract
Today, more and more software are augmented with service-oriented packaging. At the same time, more and more business and government services are provided and offered in the form of software. However, there are debates on whether and how much service engineering has in common with software engineering.
Hemant K. Jain 0001, Calton Pu, Sridhar Iyengar, M. Brian Blake, Carl K. Chang
ICWS2
2007 Guarding Sensitive Information Streams through the Jungle of Composite Web Services
abstract
Complex and dynamic web service compositions may introduce unpredictable and unintentional sharing of security-sensitive data (e.g., credit card numbers) as well as unexpected vulnerabilities that cause information leak. This paper describes a fine-grain access policy specification of security-sensitive data items for each component web service. We propose the SF-Guard architecture to enforce these access policies at component web services. A prototype implementation of SF-Guard (on Apache Axis2) and its evaluation show that effective protection of security-sensitive information can be achieved at low overhead (a few percent addition to response time) while preserving the functionality of flexible web service composition.
Jinpeng Wei, Lenin Singaravelu, Calton Pu
ICWS3
2007 An Analysis of Performance Interference Effects in Virtual Environments
abstract
Virtualization is an essential technology in modern datacenters. Despite advantages such as security isolation, fault isolation, and environment isolation, current virtualization techniques do not provide effective performance isolation between virtual machines (VMs). Specifically, hidden contention for physical resources impacts performance differently in different workload configurations, causing significant variance in observed system throughput. To this end, characterizing workloads that generate performance interference is important in order to maximize overall utility. In this paper, we study the effects of performance interference by looking at system-level workload characteristics. In a physical host, we allocate two VMs, each of which runs a sample application chosen from a wide range of benchmark and real-world workloads. For each combination, we collect performance metrics and runtime characteristics using an instrumented Ken hypervisor. Through subsequent analysis of collected data, we identify clusters of applications that generate certain types of performance interference. Furthermore, we develop mathematical models to predict the performance of a new application from its workload characteristics. Our evaluation shows our techniques were able to predict performance with average error of approximately 5%
Younggyun Koh, Rob C. Knauerhase, Paul Brett, Mic Bowman, Zhihua Wen, Calton Pu
ISPASS6
2007 A Utility-Aware Middleware Architecture for Decentralized Group Communication Applications
Jianjun Zhang 0001, Ling Liu 0001, Lakshmish Ramaswamy, Gong Zhang 0008, Calton Pu
Middleware5
2007 QueST: querying music databases by acoustic and textual features
abstract
With continued growth of music content available on the Internet, music information retrieval has attracted increasing attention. An important challenge for music searching is its ability to support both keyword and content based queries efficiently and with high precision. In this paper, we present a music query system - QueST (Query by acouStic and Textual features) to support both keyword and content based retrieval in large music databases. QueST has two distinct features. First, it provides new index schemes that can efficiently handle various queries within a uniform architecture. Concretely, we propose a hybrid structure consisting of Inverted file and Signature file to support keyword search. For content based query, we introduce the notion of similarity to capture various music semantics like melody and genre. We extract acoustic features from a music object, and map it to multiple high-dimension spaces with respect to the similarity notion using PCA and RBF neural network. Second, we design a result fusion scheme, called the Quick Threshold Algorithm, to speed up the processing of complex queries involving both textual and multiple acoustic features. Our experimental results show that QueST offers higher accuracy and efficiency compared to existing algorithms.
Bin Cui 0001, Ling Liu 0001, Calton Pu, Jialie Shen 0001, Kian-Lee Tan
ACM Multimedia3
2006 Reducing TCB complexity for security-sensitive applications: three case studies
abstract
The large size and high complexity of security-sensitive applications and systems software is a primary cause for their poor testability and high vulnerability. One approach to alleviate this problem is to extract the security-sensitive parts of application and systems software, thereby reducing the size and complexity of software that needs to be trusted. At the system software level, we use the Nizza architecture which relies on a kernelized trusted computing base (TCB) and on the reuse of legacy code using trusted wrappers to minimize the size of the TCB. At the application level, we extract the security-sensitive portions of an already existing application into an AppCore. The AppCore is executed as a trusted process in the Nizza architecture while the rest of the application executes on a virtualized, untrusted legacy operating system. In three case studies of real-world applications (e-commerce transaction client, VPN gateway and digital signatures in an e-mail client), we achieved a considerable reduction in code size and complexity. In contrast to the few hundred thousand lines of current application software code running on millions of lines of systems software code, we have AppCores with tens of thousands of lines of code running on a hundred thousand lines of systems software code. We also show the performance penalty of AppCores to be modest (a few percent) compared to current software.
Lenin Singaravelu, Calton Pu, Hermann Härtig, Christian Helmuth
EuroSys2
2006 DSCWeaver: Synchronization-Constraint Aspect Extension to Procedural Process Specification Languages
abstract
BPEL is emerging as an open-standards language for Web service composition. However, its procedural style can lead to inflexible and tangled code for managing a crosscutting aspect - synchronization constraints that define permissible sequences of execution for activities in a process. In this paper, we present DSCWeaver, a tool that enables a synchronization-aspect extension to BPEL. It uses DSCL, a synchronization expression language, to specify constraints. DSCL has the desirable features of declarative syntax, fine granularity, and validation support. A designer can use DSCL to describe and validate the synchronization behavior and rely on DSCWeaver to generate BPEL code. We demonstrate the advantages of our approach in a service deployment process and evaluate its performance using two metrics: lines of code (LoC) and places to visit (PtV). Evaluation results show that our approach can effectively reduce development effort of process designers while providing performance competitive to un-woven BPEL code
Qinyi Wu, Calton Pu, Akhil Sahai, Roger S. Barga, Gueyoung Jung
ICWS2
2006 Issues in Bottleneck Detection in Multi-Tier Enterprise Applications
abstract
In this work, the performance of various machine learning classifiers with regard to bottleneck detection in enterprise, multi-tier applications governed by service level objectives is described. Specifically, in this paper, it demonstrates the effectiveness of three classifiers, a tree-augmented Naive Bayesian network, a J48 decision tree, and LogitBoost, using our bottleneck detection process, which delves into a new area of performance analysis based on the trends of metrics (first order derivative) rather than the metric value itself. Furthermore, the efficiency of each classifier by measuring the convergence speed, or the number of staging trials required in order to provide positive results is illustrated. Finally, the effectiveness of the classifiers used in the bottleneck detection process as each classifier strongly identifies the enterprise system bottleneck
Jason Parekh, Gueyoung Jung, Galen S. Swint, Calton Pu, Akhil Sahai
IWQoS4
2006 Efficient Packet Processing in User-Level OSes: A Study of UML
abstract
Network server consolidation has become popular through virtualization technology that builds secure, isolated network systems on shared hardware. One of the virtualization techniques used is that of user-level operating systems. (ULOSes) However, the isolation and security they bring comes at the price of performance, as virtualization introduces a number of overheads into the system. Such overheads can be surprisingly large, especially for complex OS modules like network protocol stacks. Our studies of the TCP/IP stack in user-mode Linux (UML), an implementation of a ULOS, attribute the resulting slow-downs to three main sources: the execution of privileged code, memory management across layers, and additional instructions to execute. To mitigate these bottlenecks, we present five optimization techniques, improving the network performance significantly, reducing packet processing latency by 60% and increasing network throughput by three folds. Furthermore, the network throughput of the improved ULOS is comparable to that of native Linux up to gigabit speeds
Younggyun Koh, Calton Pu, Sapan Bhatia, Charles Consel
LCN2
2006 Automated Staging for Built-to-Order Application Systems
abstract
The increasing complexity of enterprise and distributed systems demands automated design, testing, deployment, and monitoring of applications. Testing, or staging, in particular poses unique challenges. In this paper, we present the Elba project and Mulini generator. The goal of Elba is creating automated staging and testing of complex enterprise systems before deployment to production. Automating the staging process lowers the cost of testing applications. Feedback from staging, especially when coupled with appropriate resource costs, can be used to ensure correct functionality and provisioning for the application. The Elba project extracts test parameters from production specifications (such as SLAs) and deployment specifications, and via the Mulini generator, creates staging plans for the application. We then demonstrate Mulini on an example application, TPC-W, and show how information from automated staging and monitoring allows us to refine application deployments easily based on performance and cost.
Galen S. Swint, Gueyoung Jung, Calton Pu, Akhil Sahai
NOMS3
2005 DSL Weaving for Distributed Information Flow Systems
Calton Pu, Galen S. Swint
APWeb1
2005 Integration of collaborative information system in Internet applications using RiverFish architecture
abstract
Business process integration is a serious challenge in collaborative information systems due to the potential interference among them. This paper describes RiverFish architecture to solving integration problems in collaborative information systems that belong to e-commerce environment. DECA application has been used to show a good example of a non-trivial problem of this integration. DECA application controls the processing application involving several government agencies to illustrate a new application called DECA. Each step of DECA processes various levels of check points and stores the results into associated collaborative information systems. This application has served more than 2 million users since 2000, demonstrating the reliability and support for evolution of RiverFish approach
João Eduardo Ferreira, Osvaldo Kotaro Takai, Calton Pu
CollaborateCom3
2005 Collaborative Enterprise Applications (Panel)
Calton Pu
CollaborateCom1
2005 An experimental evaluation of spam filter performance and robustness against attack
abstract
In this paper, we show experimentally that learning filters are able to classify large corpora of spam and legitimate email messages with a high degree of accuracy. The corpora in our experiments contain about half a million spam messages and a similar number of legitimate messages, making them two orders of magnitude larger than the corpora used in current research. The use of such large corpora represents a collaborative approach to spam filtering because the corpora combine spam and legitimate messages from many different sources. First, we show that this collaborative approach creates very accurate spam filters. Then, we introduce an effective attack against these filters which successfully degrades their ability to classify spam. Finally, we present an effective solution to the above attack which involves retraining the filters to accurately identify the attack messages
Steve Webb, Subramanyam Chitti, Calton Pu
CollaborateCom3
2005 TOCTTOU Vulnerabilities in UNIX-Style File Systems: An Anatomical Study
Jinpeng Wei, Calton Pu
FAST2
2005 Constructing a proximity-aware power law overlay network
abstract
Peer-to-peer (P2P) networks offer a message exchanging overlay for distributed applications such as file sharing, application layer multicast, and publisher/subscriber system. The communication efficiency of the underlying overlay network is thus one of the primary factors that determine the performance of those applications. In this paper, we propose a P2P overlay network aiming at offering the low maintenance overhead of unstructured P2P networks and the scalability and communication efficiency of structured P2P networks. We design a distributed algorithm to construct low-diameter overlay networks with power law topologies. Peers consider both network proximity information and capacity of existing peers when choosing their P2P network neighbors. Using an application layer multicast system as our example, we demonstrate that our system can provide generic, scalable, and low diameter overlay networks for distributed applications that demand efficient P2P communication supports
Jianjun Zhang 0001, Ling Liu 0001, Calton Pu
GLOBECOM3
2005 Fine-Grain Adaptive Compression in Dynamically Variable Networks
abstract
Despite voluminous previous research on adaptive compression, we found significant challenges when attempting to fully utilize both network bandwidth and CPU. We describe the fine-grain (FG) mixing strategy that compresses and sends as much data as possible, and then uses any remaining bandwidth to send uncompressed packets. Experimental measurements show that FG mixing achieves significant gains in effective throughput, particularly at higher network bandwidths. However, non-trivial interactions between system components and layers (e.g., compression algorithms and middleware settings such as block size and buffer size) have significant impact on the overall system performance. Finally, the trade-offs and performance profiles of FG mixing are measured, observed, and found to be consistent over a wide range of combinations of compression algorithms (GZIP, LZO, BZ1P2), workload compression ratios (from 1 to 4), and network bandwidth (from 0 to 400 Mbps)
Calton Pu, Lenin Singaravelu
ICDCS1
2005 Comparison of Approaches to Service Deployment
abstract
IT today is driven by the trend of increasing scale and complexity. Utility and Grid computing models, PlanetLab, and traditional data centers, are reaching the scale of thousands of computers. Installed software consists of dozens of interdependent applications and services. As the complexity and scale of these systems continues to grow, it becomes increasingly difficult to administer and manage them. At the same time, the service deployment technologies are still based on scripts and configuration files with minimal ability to express dependencies, to document and to verify configurations. This results in hard-to-use and erroneous system configurations. Language- and model-based tools, such as SmartFrog and Radia, are proposed for addressing these deployment challenges, but it is unclear whether they are beneficial over traditional solutions. In this paper, we quantitatively compare manual, script-, language-, and model-based deployment solutions as a function of scale, complexity, and susceptibility to change. We also qualitatively compare them in terms of expressiveness and barrier to first use. We demonstrate that script-based solutions are well matched for large scale deployments, language-based for services of large complexity, and model-based for dynamic changes to the design. Finally, we offer a table summarizing rules of thumb regarding which solution to use in which case, subject to deployment needs.
Vanish Talwar, Qinyi Wu, Calton Pu, Wenchang Yan, Gueyoung Jung, Dejan S. Milojicic
ICDCS3
2005 Resilient Trust Management for Web Service Integration
abstract
In a distributed Web service integration environment, the selection of Web services should be based on their reputation and quality-of-service (QoS). Various trust models for web services have been proposed to evaluate the reputation of Web services/service providers. Current mechanisms are based on tracing the feedbacks to the past behaviors of Web services. However, very few of them consider the robustness and attack-resiliency of the trust models. In this paper, we present an attack resilient distributed trust management system in a Web service management environment. The proposed attack resilient trust model uses two vectors to capture the behavior and the trustworthiness of a Web service/service provider based on our analysis on the possible attacks against the trust models. We also present a set of experiments that show the effectiveness of our trust model in detecting malicious behavior of service providers.
Sungkeun Park, Ling Liu 0001, Calton Pu, Mudhakar Srivatsa, Jianjun Zhang 0001
ICWS3
2005 Clearwater: extensible, flexible, modular code generation
abstract
Distributed applications typically interact with a number of heterogeneous and autonomous components that evolve independently. Methodical development of such applications can benefit from approaches based on domain-specific languages (DSLs). However, the evolution and customization of heterogeneous components introduces significant challenges to accommodating the syntax and semantics of a DSL in addition to the heterogeneous platforms on which they must run. In this paper, we address the challenge of implementing code generators for two such DSLs that are flexible (resilient to changes in generators or input formats), extensible (able to support multiple output targets and multiple input variants), and modular (generated code can be re-written). Our approach, Clearwater, leverages XML and XSLT standards: XML supports extensibility and mutability for in-progress specification formats, and XSLT provides flexibility and extensibility for multiple target languages. Modularity arises from using XML meta-tags in the code generator itself, which supports controlled addition, subtraction, or replacement to the generated code via XML-weaving. We discuss the use of our approach and show its advantages in two non-trivial code generators: the Infopipe Stub Generator (ISG) to support distributed flow applications, and the Automated Composable Code Translator to support automated distributed application deployment. As an example, the ISG accepts as input an XML description and generates output for C, C++, or Java using a number of communications platforms such as sockets and publish-subscribe.
Galen S. Swint, Calton Pu, Gueyoung Jung, Wenchang Yan, Younggyun Koh, Qinyi Wu, Charles Consel, Akhil Sahai, Koichi Moriyama
ASE2
2005 Building a Semantic Web System for Scientific Applications: An Engineering Approach
Renato Fileto, Claudia Bauzer Medeiros, Calton Pu, Ling Liu 0001, Eduardo Delgado Assad
WISE3
2005 Achieving Efficiency and Portability in Systems Software: A Case Study on POSIX-Compliant Multithreaded Programs
abstract
Portable (standards-compliant) systems software is usually associated with unavoidable overhead from the standards-prescribed interface. For example, consider the POSIX Threads standard facility for using thread-specific data (TSD) to implement multithreaded code. The first TSD reference must be preceded by pthread/spl I.bar/getspecific( ), typically implemented as a function or macro with 40-50 instructions. This paper proposes a method that uses the runtime specialization'facility of the Tempo program specializer to convert such unavoidable source code into simple memory references of one or two instructions for execution. Consequently, the source code remains standard compliant and the executed code's performance is similar to direct global variable access. Measurements show significant performance gains over a range of code sizes. A random number generator (10 lines of C) shows a speedup of 4.8 times on a SPARC and 2.2 times on a Pentium. A time converter (2,800 lines) was sped up by 14 and 22 percent, respectively, and a parallel genetic algorithm system (14,000 lines) was sped up by 13 and 5 percent.
Yasushi Shinjo, Calton Pu
IEEE Trans. Software Eng.2
2004 Remote customization of systems code for embedded devices
abstract
Dedicated operating systems for embedded systems are fast being phased out due to their use of manual optimization, which provides high performance and small footprint, but also requires high maintenance and portability costs every time hardware evolves.In this paper, we describe an approach based on customization of generic operating system modules. Our approach uses a remote customization server to automatically generate highly optimized code that is then loaded and executed in the kernel of the embedded device. This process combines the advantages of generic systems software code (leveraging portability and evolution costs) with the advantages of customization (small footprint and low overhead).We have validated our customization infrastructure with a case study: the TCP/IP stack of the Linux kernel. We analyzed the performance and size of the customized code generated on three platforms: a Pentium III (600MHz), an ARM SA1100 (200Mhz) on a COMPAQ iPAQ, and a 486 (40MHz). The customized code runs about 25% faster and its size reduces by up to a factor of 20. The throughput of the protocol stack improves by up to 21%.
Sapan Bhatia, Charles Consel, Calton Pu
EMSOFT3
2004 Code Generation for WSLAs using AXpect
abstract
WSLAs can be viewed as describing the service aspect of Web services. By their nature, Web services are distributed. Therefore, integrating support code into a Web service application is potentially costly and error prone. Viewed from this AOP perspective, then, we present a method for integrating WSLAs into code generation using the AXpect weaver, the AOP technology for Infopipes. This helps to localize the code physically and therefore increase the eventual maintainability and enhance the reuse of the WSLA code. We then illustrate the weavers capability by using a WSLA document to codify constraints and metrics for a streaming image application that requires CPU resource monitoring.
Galen S. Swint, Calton Pu
ICWS2
2004 Automatic Specialization of Protocol Stacks
abstract
Abstract — Fast and optimized protocol stacks play a major role in the performance of network services. This role is especially important in embedded class systems, where performance metrics such as data throughput tend to be limited by the CPU. It is common on such systems, to have protocol stacks that are optimized by hand for better performance and smaller code footprint. In this paper, we propose a strategy to automate this process. Our approach uses program specialization, and enables appli-cations using the network to request specialized code based on the current usage scenario. The specialized code is generated dy-namically and loaded in the kernel to be used by the application. We have successfully applied our approach to the TCP/IP implementation in the Linux kernel and used the optimized protocol stack in existing applications. These applications were minimally modified to request the specialization of code based on the current usage context, and to use the specialized code generated instead of its generic version. Specialization can be performed locally, or deferred to a remote specialization server using a novel mechanism [1]. Experiments conducted on three platforms show that the specialized code runs about 25 % faster and its size reduces by up to 20 times. The throughput of the protocol stack improves by up to 21%. I.
Sapan Bhatia, Charles Consel, Anne-Françoise Le Meur, Calton Pu
LCN4
2004 Infopipes: The ISL/ISG Implementation Evaluation
abstract
We provide a performance comparison of generated Infopipes that have been translated and the Spi/XlP variant of Infopipe specification into executable code. Infopipes are abstractions to support information flow applications. These tools are evaluated through a realistic application: a continuous image streaming program. We implement the application in C and compare its performance to both a hand-written application and one that uses SunRPC.
Galen S. Swint, Calton Pu, Younggyun Koh, Ling Liu 0001, Wenchang Yan, Charles Consel, Koichi Moriyama, Jonathan Walpole
NCA2
2004 Reliable Peer-to-Peer End System Multicasting through Replication
abstract
A key challenge in peer-to-peer computing system is to provide decentralized and yet reliable services on top of a network of loosely coupled, weakly connected and possibly unreliable peers. This work presents an effective dynamic passive replication scheme designed to provide reliable multicast service in peer-cast, an efficient and self-configurable peer-to-peer end system multicast (ESM) system. We first describe the design of a distributed replication scheme, which enables reliable subscription and multicast dissemination of information in an environment of inherently unreliable peers. Then we present an analytical model to discuss its fault tolerance properties, and report a set of initial experiments, showing the feasibility and the effectiveness of the proposed approach.
Jianjun Zhang 0001, Ling Liu 0001, Calton Pu, Mostafa H. Ammar
Peer-to-Peer Computing3
2004 Efficient mediators with closures for handling dynamic interfaces in an imperative language
Yasushi Shinjo, Toshiyuki Kubo, Calton Pu
Inf. Softw. Technol.3
2004 A Systematic Approach to Flexible Specification, Composition, and Restructuring of Workflow Activities
abstract
We introduce the ActivityFlow specification language for flexible specification, composition, and coordination of workflow activities. The most interesting features of the ActivityFlow specification language include: (1) a collection of specification mechanisms, allowing workflow designers to use a uniform workflow specification interface to describe different types (i.e., ad-hoc, administrative, or production) of workflows involved in their organizational processes– this feature helps to increase the flexibility of workflow processes in accommodating various types of changes; (2) a set of activity modeling facilities, enabling workflow designers to describe the flow of work declaratively and incrementally, allowing to reason about correctness and security of complex workflow activities independently from their underlying implementation mechanisms; (3) an open architecture that supports user interaction as well as collaboration of workflow systems of different organizations, and a set of workflow activity restructuring operators to respond to dynamic changes of workflow activities. We end the paper with a series of simulation-based experiments that demonstrate the effectiveness of these restructuring operators and the implementation architecture of the ActivityFlow system.
Ling Liu 0001, Calton Pu, Duncan Dubugras Alcoba Ruiz
J. Database Manag.2
2003 BioSeek: Exploiting Source-Capability Information for Integrated Access to Multiple Bioinformatics Data Sources
abstract
Modern Bioinformatics data sources are widely used by molecular biologists for homology searching and new drug discovery. User-friendly and yet responsive access is one of the most desirable properties for integrated access to the rapidly growing, heterogeneous, and distributed collection of data sources. The increasing volume and diversity of digital information related to bioinformatics (such as genomes, protein sequences, protein structures, etc.) have led to a growing problem that conventional data management systems do not have, namely finding which information sources out of many candidate choices are the most relevant and most accessible to answer a given user query. We refer to this problem as the query routing problem. In this paper we introduce the notation and issues of query routing, and present a practical solution for designing a scalable query routing system based on multi-level progressive pruning strategies. The key idea is to create and maintain source capability profiles independently, and to provide algorithms that can dynamically discover relevant information sources for a given query through the smart use of source profiles. Compared to the keyword-based indexing techniques adopted in most of the search engines and software, our approach offers fine-granularity of interest matching, thus it is more powerful and effective for handling queries with complex conditions.
Ling Liu 0001, David Buttler, Terence Critchlow, Henrique Paques, Calton Pu, Daniel Rocco
BIBE6
2003 Spidle: A DSL Approach to Specifying Streaming Applications
Charles Consel, Hédi Hamdi, Laurent Réveillère, Lenin Singaravelu, Calton Pu
GPCE6
2003 Trigger Grouping: A Scalable Approach to Large Scale Information Monitoring
abstract
Information change monitoring services are becoming increasingly useful as more and more information is published on the Web. A major research challenge is how to make the service scalable to serve millions of monitoring requests. Such services usually use soft triggers to model users' monitoring requests. We have developed an effective trigger grouping scheme to optimize the trigger processing. The main idea behind this scheme is to reduce repeated computation by grouping monitoring requests of similar structures together. In this paper, we evaluate our approach using both measurements on real systems and simulations. The study shows significant performance gains using the trigger grouping approach. Moreover, the gains are critically dependent on group size and group size distribution (e.g., Zipf). We also discuss the benefit, trade-off, and runtime characteristics of the proposed approach.
Wei Tang 0005, Ling Liu 0001, Calton Pu
NCA3
2003 A Modeling and Execution Environment for Distributed Scientific Workflows
abstract
We illustrate how a domain scientist can perform a complex scientific task by interleaving data access, querying, and manipulation, as well as analytical steps and computations in complex, problem specific ways. We show how our system is used by a geneticist for solving the problem of discovering the so-called "co-regulated" genes by interlinking data and computation from several Web sites, local computations, as well as local and remote databases. The main distinctive features of our system (compared, e.g., to the ZOO environment (Ioannidis et al., 1996)) include: (i) executable workflows run as Web services; (ii) abstract workflows employ concept names and semantic types that are higher-level (and thus more "scientist friendly") than executable workflows; and (iii) our system supports automatic translation of the latter into the former.
Ilkay Altintas, Sangeeta Bhagwanani, David Buttler, Sandeep Chandra, Zhengang Cheng, Matthew Coleman, Terence Critchlow, Amarnath Gupta, Ling Liu 0001, Bertram Ludäscher, Calton Pu, Reagan W. Moore, Arie Shoshani, Mladen A. Vouk
SSDBM12
2003 Thread transparency in information flow middleware
abstract
Abstract Applications that process continuous information flows are challenging to write because the application programmer must deal with flow‐specific concurrency and timing requirements, necessitating the explicit management of threads, synchronization, scheduling and timing. We believe that middleware can ease this burden, but many middleware platforms do not match the structure of these applications, because they focus on control‐flow centric interaction models such as remote method invocation. Indeed, they abstract away from the very things that the information‐flow centric programmer must control. This paper describes Infopipes—a new high‐level abstraction for information flow applications—and a middleware framework that supports them. Infopipes handle the complexities associated with control flow and multi‐threading, relieving the programmer of these tasks. Starting from a high‐level description of an information flow pipeline, the framework determines which parts of a pipeline require separate threads or coroutines, and handles synchronization transparently to the application programmer. The framework also gives the programmer the freedom to write or reuse components in a passive style, even though the configuration will actually require the use of a thread or coroutine. Conversely, it is possible to write a component using a thread and know that the thread will be eliminated if it is not needed in a pipeline. This allows the most appropriate programming model to be chosen for a given task, and existing code to be reused irrespective of its activity model. Copyright © 2003 John Wiley & Sons, Ltd.
Rainer Koster, Andrew P. Black, Jie Huang 0040, Jonathan Walpole, Calton Pu
Softw. Pract. Exp.5
2003 POESIA: An ontological workflow approach for composing Web services in agriculture
Renato Fileto, Ling Liu 0001, Calton Pu, Eduardo Delgado Assad, Claudia Bauzer Medeiros
VLDB J.3
2002 Ginga: a self-adaptive query processing system
abstract
Article Share on Ginga: a self-adaptive query processing system Authors: Henrique Paques Georgia Institute of Technology Georgia Institute of TechnologyView Profile , Ling Liu Georgia Institute of Technology Georgia Institute of TechnologyView Profile , Calton Pu Georgia Institute of Technology Georgia Institute of TechnologyView Profile Authors Info & Claims CIKM '02: Proceedings of the eleventh international conference on Information and knowledge managementNovember 2002 Pages 655–658https://doi.org/10.1145/584792.584910Online:04 November 2002Publication History 3citation398DownloadsMetricsTotal Citations3Total Downloads398Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Henrique Paques, Ling Liu 0001, Calton Pu
CIKM3
2002 Guarding the next Internet frontier: countering denial of information attacks
abstract
As applications enabled by the Internet become information rich, ensuring access to quality information in the presence of potentially malicious entities will be a major challenge. Denial of information (DoI) attacks attempt to degrade the quality of information by deliberately introducing noise that appears to be useful information. The mere availability of information is insufficient if the user must find a needle in a haystack of noise that is created by an adversary to hide critical information. We focus on the characterization of information quality metrics that are relevant in the presence of DoI attacks. In particular, two complementary metrics are explored. Information regularity captures predictability in the patterns of information creation and access. The second metric, information quality trust, captures the known ability of an information source to meet the needs of its clients.
Mustaque Ahamad, Leo Mark, Wenke Lee, Edward Omicienski, Andre dos Santos, Ling Liu 0001, Calton Pu
NSPW7
2002 Enhancing Access Control with SysGuard, Reference Monitor Supporting Portable and Composable Kernel Module
abstract
To install security modules or reference monitors into operating system kernels is a common and effective way for enhancing access control for networks. However, security modules in conventional kernel-level reference monitors are usually not portable to other kernels and require detailed knowledge about kernel internals. Furthermore, different security modules are often not composable and conflict with each other. This paper describes a reference monitor called SysGuard that addresses these problems. SysGuard uses modules called guards that are invoked before or after the execution of system calls. Unlike kernel-specific security modules, guards are attached to standard system calls that enhance their portability. The guard scoping on a per-process basis improves composability of individual guards, and it is implemented efficiently by using a per-process jump table of system calls. This paper describes the implementation of restricted execution environments for networks by composing simple and portable guards, and shows the advantages of the SysGuard security framework.
Yasushi Shinjo, Kotaro Eiraku, Kozo Itano, Calton Pu
PRDC5
2002 Infopipes: An abstraction for multimedia streaming
Andrew P. Black, Jie Huang 0040, Rainer Koster, Jonathan Walpole, Calton Pu
Multim. Syst.5
2002 Information Modelling on the Web: A Scalable Solution
Ling Liu 0001, Wei Tang 0005, David Buttler, Calton Pu
World Wide Web4
2001 Active Streams-An Approach to Adaptive Distributed Systems
abstract
Summary form only given. An increasing number of distributed applications aim to provide services to users by interacting with a correspondingly growing set of data-intensive network services. To support such requirements, we believe that new services need to be customizable, applications need to be dynamically extensible, and both applications and services need to be able to adapt to variations in resource availability and demand. A comprehensive approach to building new distributed applications can facilitate this by considering the contents of the information flowing across the application and its services and by adopting a component-based model to application/service programming. It should provide for dynamic adaptation at multiple levels and points in the underlying platform; and, since the mapping of components to resources in dynamic environment is too complicated, it should relieve programmers of this task. We propose Active Streams, a middleware approach and its associated framework for building distributed applications and services that exhibit these characteristics.
Fabián E. Bustamante, Greg Eisenhauer, Patrick M. Widener, Karsten Schwan, Calton Pu
HotOS5
2001 A Fully Automated Object Extraction System for the World Wide Web
abstract
This paper presents a fully automated object extraction system Omini. A distinct feature of Omini is the suite of algorithms and the automatically learned information extraction rules for discovering and extracting objects from dynamic Web pages or static Web pages that contain multiple object instances. We evaluated the system using more than 2,000 Web pages over 40 sites. It achieves 100% precision (returns only correct objects) and excellent recall (between 99% and 98%, with very few significant objects left out). The object boundary identification algorithms are fast, about 0.1 second per page with a simple optimization.
David Buttler, Ling Liu 0001, Calton Pu
ICDCS3
2001 Thread Transparency in Information Flow Middleware
Rainer Koster, Andrew P. Black, Jie Huang 0040, Jonathan Walpole, Calton Pu
Middleware5
2001 OminiSearch: A Method for Searching Dynamic Content on the Web
abstract
No abstract available.
David Buttler, Ling Liu 0001, Calton Pu, Henrique Paques
SIGMOD Conference3
2001 An XML-enabled data extraction toolkit for web sources
Ling Liu 0001, Calton Pu
Inf. Syst.2
2001 Specialization tools and techniques for systematic optimization of system software
abstract
Specialization has been recognized as a powerful technique for optimizing operating systems. However, specialization has not been broadly applied beyond the research community because current techniques based on manual specialization, are time-consuming and error-prone. The goal of the work described in this paper is to help operating system tuners perform specialization more easily. We have built a specialization toolkit that assists the major tasks of specializing operating systems. We demonstrate the effectiveness of the toolkit by applying it to three diverse operating system components. We show that using tools to assist specialization enables significant performance optimizations without error-prone manual modifications. Our experience with the toolkit suggests new ways of designing systems that combine high performance and clean structure.
Dylan McNamee, Jonathan Walpole, Calton Pu, Crispin Cowan, Charles Krasic, Ashvin Goel, Perry Wagle, Charles Consel, Gilles Muller, Renaud Marlet
ACM Trans. Comput. Syst.3
2000 WebCQ: Detecting and Delivering Information Changes on the Web
abstract
WebCQ is a prototype system for large-scale Web information monitoring and delivery.It makes heavy use of the structure present i n h ypertext and the concept of continual queries.In this paper we discuss both mechanisms that We-bCQ uses to discover and detect changes to the World Wide Web (the Web) pages eciently, and the methods to notify users of interesting changes with a personalized customization.The WebCQ system consists of four main components: a c hange detection robot that discovers and detects changes, a proxy cache service that reduces communication tracs to the original information servers, a personalized presentation tool that highlights changes detected by W ebCQ sentinels, and a change noti cation service that delivers fresh information to the right users at the right time.A salient feature of our change detection robot is its ability to support various types of web page sentinels for detecting, presenting, and delivering interesting changes to web pages.This paper describes the WebCQ system with an emphasis on general issues in designing and engineering a large-scale information change monitoring system on the Web.
Ling Liu 0001, Calton Pu, Wei Tang 0005
CIKM2
2000 XWRAP: An XML-Enabled Wrapper Construction System for Web Information Sources
abstract
The paper describes the methodology and the software development of XWRAP, an XML-enabled wrapper construction system for semi-automatic generation of wrapper programs. By XML-enabled we mean that the metadata about information content that are implicit in the original Web pages will be extracted and encoded explicitly as XML tags in the wrapped documents. In addition, the query based content filtering process is performed against the XML documents. The XWRAP wrapper generation framework has three distinct features. First, it explicitly separates tasks of building wrappers that are specific to a Web source from the tasks that are repetitive for any source, and uses a component library to provide basic building blocks for wrapper programs. Second, it provides a user friendly interface program to allow wrapper developers to generate their wrapper code with a few mouse clicks. Third and most importantly, we introduce and develop a two-phase code generation framework. The first phase utilizes an interactive interface facility to encode the source-specific metadata knowledge identified by individual wrapper developers as declarative information extraction rules. The second phase combines the information extraction rules generated at the first phase with the XWRAP component library to construct an executable wrapper program for the given Web source. We report the initial experiments on performance of the XWRAP code generation system and the wrapper programs generated by XWRAP.
Ling Liu 0001, Calton Pu
ICDE2
2000 SubDomain: Parsimonious Server Security
Crispin Cowan, Steve Beattie, Greg Kroah-Hartman, Calton Pu, Perry Wagle, Virgil D. Gligor
LISA4
2000 Research challenges in environmental observation and forecasting systems
abstract
The availability of tremendous computation power coupled with widespread connectivity have fueled the development of real-time environmental observation and forecasting systems (EOFS).
David C. Steere, António M. Baptista, Dylan McNamee, Calton Pu, Jonathan Walpole
MobiCom4
2000 AQR-Toolkit: An Adaptive Query Routing Middleware for Distributed Data Intensive Systems
abstract
Query routing is an intelligent service that can direct query requests to appropriate servers that are capable of answering the queries. The goal of a query routing system is to provide efficient associative access to a large, heterogeneous, distributed collection of information providers by routing a user query to the most relevant information sources that can provide the best answer. Effective query routing not only minimizes the query response time and the overall processing cost, but also eliminates a lot of unnecessary communication overhead over the global networks and over the individual information sources.
Ling Liu 0001, Calton Pu, David Buttler, Henrique Paques, Wei Tang 0005
SIGMOD Conference2
2000 Guest Editors' Introduction - Papers from ICDE 1999
abstract
In the paper An Approach to Active Spatial Min- ing Based on Statistical Information, authors Wei Wang, Jiong Yang, and Richard Muntz propose data mining algo- rithm to efficiently support user-defined triggers on dy- namically evolving spatial data. It is shown that a new hi- erarchical triggering strategy which introduces subtriggers can improve the performance by three orders of magnitude compared with naive approach. The authors explored the new research area of active spatial mining. The paper A Database Approach for Modeling and Querying Video Data by Mohand-Said Hacid, Cyril De- cleir, and Jacques Kouloumdjian proposes the integrated
Masaru Kitsuregawa, Mike P. Papazoglou, Calton Pu
IEEE Trans. Knowl. Data Eng.3
2000 Correction to "Continual Queries for Internet Scale Event-Driven Information Delivery"
Ling Liu 0001, Calton Pu, Wei Tang 0005
IEEE Trans. Knowl. Data Eng.2
1999 A Feedback-driven Proportion Allocator for Real-Rate Scheduling
David C. Steere, Ashvin Goel, Joshua Gruenberg, Dylan McNamee, Calton Pu, Jonathan Walpole
OSDI5
1999 An XML-based Wrapper Generator for Web Information Extraction
Ling Liu 0001, David Buttler, Calton Pu, Wei Tang 0005
SIGMOD Conference4
1999 TAM: A System for Dynamic Transactional Activity Management
abstract
Article Free Access Share on TAM: a system for dynamic transactional activity management Authors: Tong Zhou Department of Computer Science & Engineering, Oregon Graduate Institute, Portland, OR Department of Computer Science & Engineering, Oregon Graduate Institute, Portland, ORView Profile , Ling Liu Department of Computer Science & Engineering, Oregon Graduate Institute, Portland, OR Department of Computer Science & Engineering, Oregon Graduate Institute, Portland, ORView Profile , Calton Pu Department of Computer Science & Engineering, Oregon Graduate Institute, Portland, OR Department of Computer Science & Engineering, Oregon Graduate Institute, Portland, ORView Profile Authors Info & Claims SIGMOD '99: Proceedings of the 1999 ACM SIGMOD international conference on Management of dataJune 1999Pages 571–573https://doi.org/10.1145/304182.304580Published:01 June 1999Publication History 2citation336DownloadsMetricsTotal Citations2Total Downloads336Last 12 Months16Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Ling Liu 0001, Calton Pu
SIGMOD Conference3
1999 Continual Queries for Internet Scale Event-Driven Information Delivery
abstract
We introduce the concept of continual queries, describe the design of a distributed event-driven continual query system-OpenCQ, and outline the initial implementation of OpenCQ on top of the distributed interoperable information mediation system DIOM. Continual queries are standing queries that monitor update of interest and return results whenever the update reaches specified thresholds. In OpenCQ, users may specify to the system the information they would like to monitor (such as the events or the update thresholds they are interested in). Whenever the information of interest becomes available, the system immediately delivers it to the relevant users; otherwise, the system continually monitors the arrival of the desired information and pushes it to the relevant users as it meets the specified update thresholds. In contrast to conventional pull-based data management systems such as DBMSs and Web search engines, OpenCQ exhibits two important features: it provides push-enabled, event-driven, content-sensitive information delivery capabilities; and it combines pull and push services in a unified framework. By event-driven we mean that the update events of interest to be monitored are specified by users or applications. By content-sensitive, we mean the evaluation of the trigger condition happens only when a potentially interesting change occurs. By push-enabled, we mean the active delivery of query results or triggering of actions without user intervention.
Ling Liu 0001, Calton Pu, Wei Tang 0005
IEEE Trans. Knowl. Data Eng.2
1998 Dynamic Restructuring of Transactional Workflow Activities: A Practical Implementation Method
abstract
Article Free Access Share on Dynamic restructuring of transactional workflow activities: a practical implementation method Authors: Tong Zhou Department of Computer Science & Engineering, Oregon Graduate Institute, P.O.Box 91000, Portland, OR Department of Computer Science & Engineering, Oregon Graduate Institute, P.O.Box 91000, Portland, ORView Profile , Calton Pu Department of Computer Science & Engineering, Oregon Graduate Institute, P.O.Box 91000, Portland, OR Department of Computer Science & Engineering, Oregon Graduate Institute, P.O.Box 91000, Portland, ORView Profile , Ling Liu Department of Computer Science & Engineering, Oregon Graduate Institute, P.O.Box 91000, Portland, OR Department of Computer Science & Engineering, Oregon Graduate Institute, P.O.Box 91000, Portland, ORView Profile Authors Info & Claims CIKM '98: Proceedings of the seventh international conference on Information and knowledge managementNovember 1998 Pages 378–385https://doi.org/10.1145/288627.288683Online:01 November 1998Publication History 6citation301DownloadsMetricsTotal Citations6Total Downloads301Last 12 Months4Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Calton Pu, Ling Liu 0001
CIKM2
1998 Fast, Optimized Sun RPC Using Automatic Program Specialization
abstract
Fast remote procedure call (RPC) is a major concern for distributed systems. Many studies aimed at efficient RPC consist of either new implementations of the RPC paradigm or manual optimization of critical sections of the code. This paper presents an experiment that achieves automatic optimization of an existing, commercial RPC implementation, namely the Sun RPC. The optimized Sun RPC is obtained by using an automatic program specializer. It runs up to 1.5 times faster than the original Sun RPC. Close examination of the specialized code does not reveal further optimization opportunities which would lead to significant improvements without major manual restructuring. The contributions of this work are: the optimized code is safely produced by an automatic tool and thus does not entail any additional maintenance; to the best of our knowledge this is the first successful specialization of mature, commercial, representative system code; and the optimized Sun RPC runs significantly faster than the original code.
Gilles Muller, Renaud Marlet, Nic Volanschi, Charles Consel, Calton Pu, Ashvin Goel
ICDCS5
1998 Methodical Restructuring of Complex Workflow Activities
abstract
We describe a family of activity-split and activity-join operations with a notion of validity. The key idea of introducing the set of activity-split and activity-join operations is to allow users to restructure ongoing activities in anticipation of uncertainty so that any significant performance loss due to unexpected unavailablity or delay of shared resources can be avoided or reduced through release of early committed resources or transferring ownership of uncommitted resources. To guarantee the correctness of new activities generated by activity-split or activity-join operations, we define the notion of validity of activity restructuring operations and identify the cases where the correctness is ensured and the cases where activity-split or activity-join are illegal due to the inconsistency incurred.
Ling Liu 0001, Calton Pu
ICDE2
1998 CQ: A Personalized Update Monitoring Toolkit
abstract
The CQ project at OGI, funded by DARPA, aims at developing a scalable toolkit and techniques for update monitoring and event-driven information delivery on the net. The main feature of the CQ project is a “personalized update monitoring” toolkit based on continual queries [3]. Comparing with the pure pull (such as DBMSs, various web search engines) and pure push (such as Pointcast, Marimba, Broadcast disks) technology, the CQ project can be seen as a hybrid approach that combines the pull and push technology by supporting personalized update monitoring through a combined client-pull and server-push paradigm.
Ling Liu 0001, Calton Pu, Wei Tang 0005, David Buttler, John Biggs, Paul Benninghoff, Fenghua Yu
SIGMOD Conference2
1998 View Consistency for Optimistic Replication
abstract
Optimistically replicated systems provide highly available data even when communication between data replicas is unreliable or unavailable. The high availability comes at the cost of allowing inconsistent accesses, since users can read and write old copies of data. Session guarantees have been used to reduce such inconsistencies. They preserve most of the availability benefits of optimistic systems. We generalize session guarantees to apply to persistent as well as distributed entities. We implement these guarantees, called view consistency, on Ficus an optimistically replicated file system. Our implementation enforces consistency on a per-file basis and does not require changes to individual applications. View consistency is enforced by clients accessing the data and thus requires minimal changes to the replicated data servers. We show that view consistency allows access to available and high performing data replicas and can be implemented efficiently. Experimental results show that the consistency overhead for clients ranges from 1% to 8% of application runtime for the benchmarks studied in the prototype system. The benefits of the system are an improvement in access times due to better replica selection and improved consistency guarantees over a purely optimistic system.
Ashvin Goel, Calton Pu, Gerald J. Popek
SRDS2
1998 Distributed Query Scheduling Service: An Architecture and Its Implementation
abstract
We present the systematic design and development of a distributed query scheduling service (DQS) in the context of DIOM, a distributed and interoperable query mediation system.26 DQS consists of an extensible architecture for distributed query processing, a three-phase optimization algorithm for generating efficient query execution schedules, and a prototype implementation. Functionally, two important execution models of distributed queries, namely moving query to data or moving data to query, are supported and combined into a unified framework, allowing the data sources with limited search and filtering capabilities to be incorporated through wrappers into the distributed query scheduling process. Algorithmically, conventional optimization factors (such as join order) are considered separately from and refined by distributed system factors (such as data distribution, execution location, heterogeneous host capabilities), allowing for stepwise refinement through three optimization phases: Compilation, parallelization, site selection and execution. A subset of DQS algorithms has been implemented in Java to demonstrate the practicality of the architecture and the usefulness of the distributed query scheduling algorithm in optimizing execution schedules for inter-site queries.
Ling Liu 0001, Calton Pu, Kirill Richine
Int. J. Cooperative Inf. Syst.2
1997 ActivityFlow: Towards Incremental Specification and Flexible Coordination of Workflow Activities
Ling Liu 0001, Calton Pu
ER2
1997 A Dynamic Query Scheduling Framework for Distributed and Evolving Information Systems
abstract
The rapid growth of the wide area network technology has led to an increasing number of information sources available online. To ensure the query services to scale up with such dynamic open environments, an advanced distributed information system must provide adequate support for dynamic interconnection between information consumers and information producers, instead of just functioning as a static data delivery system. We develop a distributed query scheduling framework to demonstrate the feasibility and the benefit for supporting interoperability and dynamic information gathering across heterogeneous information sources, without relying on an integrated view predefined over the participating information sources. We outline the mechanisms developed for the main components of our distributed query scheduling framework, such as query routing and query execution planning services. We also provide a concrete example to illustrate the issues on how the information consumers' query requests are dynamically processed and linked to the heterogeneous information sources and how the query scheduling framework scales up as the number of information sources increases.
Ling Liu 0001, Calton Pu
ICDCS2
1997 Code Generation through Annotation of Macromolecular Structure Data
John Biggs, Calton Pu, Philip E. Bourne
ISMB2
1997 An Adaptive Object-Oriented Approach to Integration and Access of Heterogeneous Information Sources
Ling Liu 0001, Calton Pu
Distributed Parallel Databases2
1997 Divergence Control Algorithms for Epsilon Serializability
abstract
The paper presents divergence control methods for epsilon serializability (ESR) in centralized databases. ESR alleviates the strictness of serializability (SR) in transaction processing by allowing for limited inconsistency. The bounded inconsistency is automatically maintained by divergence control (DC) methods in a way similar to SR is maintained by concurrency control (CC) mechanisms. However, DC for ESR allows more concurrency than CC for SR. The authors first demonstrate the feasibility of ESR by showing the design of three representative DC methods: two-phase locking, timestamp ordering and optimistic approaches. DC methods are designed by systematically enhancing CC algorithms in two stages: extension and relaxation. In the extension stage, a CC algorithm is analyzed to locate the places where it identifies non-SR conflicts of database operations. In the relaxation stage, the non-SR conflicts are relaxed to allow for controlled inconsistency. They then demonstrate the applicability Of ESR by presenting the design of DC methods using other most known inconsistency specifications, such as absolute value, age and total number of nonserializably read data items. In addition, they present a performance study using an optimistic divergence control algorithm as an example to show that a substantial improvement in concurrency can be achieved in ESR by allowing for a small amount of inconsistency.
Kun-Lung Wu, Philip S. Yu, Calton Pu
IEEE Trans. Knowl. Data Eng.3
1996 An Adaptive Approach to Query Mediation Across Heterogeneous Information Sources
abstract
The authors propose a query mediation framework to support customizable information gathering across heterogeneous and autonomous information sources. Instead of an integrated (and static) global schema, they propose an adaptive approach to interoperability which allows information consumers to represent their queries based on the customized personal view rather than at system-defined integrated view. The query mediation framework consists of five steps: query routing, query decomposition, parallel access plan generation, subquery translation and execution, and query result assembly. Concrete examples illustrate the challenges arising from heterogeneity in these five steps and how the framework scales up as the number of information sources grow and evolve.
Ling Liu 0001, Calton Pu, Yooshin Lee
CoopIS2
1996 Differential Evaluation of Continual Queries
abstract
We define continual queries as a useful tool for monitoring of updated information. Continual queries are standing queries that monitor the source data and notify the users whenever new data matches the query. In addition to periodic refresh, continual queries include Epsilon Transaction concepts to allow users to specify query refresh based on the magnitude of updates. To support efficient processing of continual queries, we propose a differential re-evaluation algorithm (DRA), which exploits the structure and information contained in both the query expressions and the database update operations. The DRA design can be seen as a synthesis of previous research on differential files, incremental view maintenance, and active databases.
Ling Liu 0001, Calton Pu, Roger S. Barga
ICDCS2
1996 What's in a WWW Link? - Panel
Amit P. Sheth, Robert Meersman, Erich J. Neuhold, Calton Pu, V. S. Subrahmanian
ICDE4
1995 The Distributed Interoperable Object Model and Its Application to Large-scale Interoperable Database Systems
abstract
A large-scale interoperable database system operating in a dynamic environment should provide uniform access user interface to its components, scalability to larger networks, evolution of database schema and applications, flexible composability of client and server components, and preserve component autonomy. To address the research issues presented by such systems, we introduce the Distributed Interoperable Object Model (DIOM). DIOM's main features include the explicit representation of and access to semantics in data sources through the DIOM base interfaces, the use of interface abstraction mechanisms (such as specialization, generalization, aggregation and import) to support incremental design and construction of compound interoperation interfaces, the deferment of conflict resolution to the query submission time instead of at the time of schema integration, and a clean interface between distributed interoperable objects that supports the independent evolution and management of such...
Ling Liu 0001, Calton Pu
CIKM2
1995 A Practical Technique for Asynchronous Transaction Processing
abstract
Asynchronous transaction processing extends traditional on-line transaction processing (TP) to improve performance of distributed systems by alleviating the serializability (SR) bottleneck. For example, epsilon serializability (ESR) uses divergence control algorithms to allow more concurrency by permitting limited non-SR interleavings. In a distributed environment, ESR relaxes commit and abort dependencies among transactions, allowing transactions to commit asynchronously. A second example, chopping up transactions allows more concurrency by dividing transactions into smaller pieces and thus reduces resource holding time. Chopping transactions enforces no commit protocols among pieces from one original transaction, allowing each piece to commit asynchronously. We combine the benefits of ESR and chopping transactions by designing three new methods that chop transactions and run them under ESR. The practical applicability of our technique is enhanced by two factors: (1) chopping transactions does not require changes in existing TP systems, and (2) ESR support has already been prototyped on a commercial TP system.
Wenwey Hseush, Calton Pu
ICDCS2
1995 Demonstrating the Effect of Software Feedback on a Distributed Real-Time MPEG Video Audio Player
abstract
No abstract available.
Shanwei Cen, Calton Pu, Richard Staehli, Crispin Cowan, Jonathan Walpole
ACM Multimedia2
1995 A Distributed Real-Time MPEG Video Audio Player
Shanwei Cen, Calton Pu, Richard Staehli, Crispin Cowan, Jonathan Walpole
NOSSDAV2
1995 Optimistic Incremental Specialization: Streamlining a Commercial Operating System
abstract
Conventionaloperating system code is written to deal with all possible system stat es, and performs considerable interpretation to determine the current system state before taking action.A consequence of this approach is that kernel calls which perform little actual work take a long time to execute.To address this problem, we use specialized operating system code that reduces interpretation for common cases, but still behaves correctly in the fully general case.We describe how specialized operating system code can be generated and bound wtcrementally as the information on which it depends becomes available.We extend our specialization techniques to include the notion of optimistic in c remental specialization n: a technique for generating specialized kernel code optimistically for system states that are likely to occur, but not certain.The ideas outlined in this paper allow the conventional kernel design tenet of "optimizing for the common case" to be extended to the domain of adaptive operating systems.We also show that aggressive use of specialization can produce in-kernel implementations of operating system functionality with performance comparable to user-level implementations.We demonstrate that these ideas are applicable in realworld operating systems by describing a re-implementation of the HP-UX file system.Our specialized read system call reduces the cost of a single byte read by a factor of 3, and an 8 KB read by 26~o, while preserving the semantics of the HP-UXread call.By relaxing the semantics of HP-UX read we were able to cut the cost of a single byte read system call by more than an order of magnitude.1
Calton Pu, Tito Autrey, Andrew P. Black, Charles Consel, Crispin Cowan, Jon Inouye, Lakshmi Kethana, Jonathan Walpole
SOSP1
1995 A Practical and Modular Implementation of Extended Transaction Models
Roger S. Barga, Calton Pu
VLDB2
1995 Divergence Control for Distributed Database Systems
Wenwey Hseush, Gail E. Kaiser, Calton Pu, Kun-Lung Wu, Philip S. Yu
Distributed Parallel Databases3
1995 A Formal Characterization of Epsilon Serializability
abstract
Epsilon serializability (ESR) is a generalization of classic serializability (SR). In this paper, we provide a precise characterization of ESR when queries that may view inconsistent data run concurrently with consistent update transactions. Our first goal is to understand the behavior of queries in the presence of conflicts and to show how ESR in fact is a generalization of SR. So, using the ACTA framework, we formally express the intertransaction conflicts that are recognized by ESR and through that define ESR, analogous to the manner in which conflict-based serializability is defined. Secondly, expressions are derived for the amount of inconsistency (in a data item) viewed by a query and its effects on the results of a query. These inconsistencies arise from concurrent updates allowed by ESR. Thirdly, in order to maintain the inconsistencies within bounds associated with each query, the expressions are used to determine the preconditions that operations have to satisfy. The results of a query, and the errors in it, depend on what a query does with the (possibly inconsistent) data viewed by it. One of the important byproducts of this work is the identification of different types of queries which lend themselves to an analysis of the effects of data inconsistency on the results of the query.
Krithi Ramamritham, Calton Pu
IEEE Trans. Knowl. Data Eng.2
1994 Multiversion Divergence Control of Time Fuzziness
abstract
Epsilon Serializability (ESR) has been proposed to manage and control inconsistency in extending the classic transaction processing. ESR increases system concurrency by tolerating a bounded amount of inconsistency. In this paper, we present multiversion divergence control (mvDC) algorithms that support ESR with not only value but also time fuzziness in multiversion databases. Unlike value fuzziness, accumulating time fuzziness is semantically different. A simple summation of the length of two time intervals may either underestimate the total time fuzziness, resulting in incorrect execution, or overestimate the total time fuzziness, unnecessarily degrading the effectiveness of mvESR. We present a new operation, called TimeUnion, to accurately accumulate the total time fuzziness. Because of the accurate control of time and value fuzziness by the mvDC algorithm, mvESR is very suitable for the use of multiversion databases for real-time applications that may tolerate a limited degree of data inconsistency but prefer more data recency.
Calton Pu, Miu K. Tsang, Kun-Lung Wu, Philip S. Yu
CIKM1
1994 Design and Application of a C++ Macromolecular Class Library
Weider Chang, Ilya N. Shindyalov, Calton Pu, Philip E. Bourne
ISMB3
1994 Design and application of PDBlib, a C++ macromolecular class library
abstract
PDBlib is an extensible object-oriented class library written in C++ for representing the three-dimensional structure of biological macromolecules. The software design strategy, features of many of the 129 classes currently distributed with the library, and two sample applications which use the library are described. Version 1.0 of the library represents the structural features of proteins, DNA, RNA and complexes thereof, at a level of detail on a par with that which can be parsed from a Protein Data Bank (PDB) entry. However, the memory-resident representation of the macromolecule is independent of the PDB entry and can be obtained from other sources, e.g. relational and object-oriented databases. PDBlib classes are organized into four categories: (i) classes that model the macromolecule; (ii) classes that enhance the extensibility of the library; (iii) classes that provide navigation facilities of the object-oriented macromolecular structure representation; and (iv) a class that loads a PDB file into the memory-resident object-oriented representation. A number of general-purpose procedures that return features of this representation and that are relevant to all biological disciplines are included in (i). The library has been used to develop PDBtool, a prototype structure verification tool, and PDBview, a structure rendering tool that requires no specialized graphics hardware and software. Current work centers on making the macromolecular structures represented by PDBlib persistent using a commercial object-oriented database and providing an additional class library, MMQLlib, to query those structures.
Weider Chang, Ilya N. Shindyalov, Calton Pu, Philip E. Bourne
Comput. Appl. Biosci.3
1994 Applying an information gathering architecture to Netfind: a white pages tool for a changing and growing Internet
abstract
The Internet is quickly becoming an indispensable means of communication and collaboration, based on applications such as electronic mail, remote information retrieval, and multimedia conferencing. A fundamental problem for such applications is supporting resource discovery in a fashion that keeps pace with the Internet's exponential growth in size and diversity. Netfind is a scalable tool that locates current electronic mail addresses and other information about Internet users. Since the time we first deployed Netfind in 1990, it has evolved considerably, making use of more types of information sources. As well as more sophisticated mechanisms to gather and cross-correlate information. In this paper, we describe these techniques, and present a general framework for gathering and harnessing widely distributed information in a diverse and growing Internet environment. At present, Netfind gathers information from 17 different types of sources, providing a particularly thorough demonstration of an information gathering architecture.>
Michael F. Schwartz, Calton Pu
IEEE/ACM Trans. Netw.2
1993 Distributed Divergence Control for Epsilon Serializability
abstract
Epsilon serializability (ESR) allows for more concurrency by permitting nonserializable interleavings of database operations among epsilon transactions (ETs). The authors present the design of distributed divergence control (DDC) algorithms for ESR in homogeneous and heterogeneous distributed databases. They first present a strict two-phase locking DDC algorithm (S2PLDDC) and an optimistic DDC algorithm (ODDC) for homogeneous distributed databases, where the local orderings of all the sub-ETs of a distributed ET are the same, and the total inconsistency of a distributed ET is simply the sum of that of all its sub-ETs. A superdatabase DDC algorithm is described for heterogeneous distributed databases, where the local orderings of all the sub-ETs of a distributed ET may not be the same, and the total inconsistency of a distributed ET may be greater than the sum of that of all its sub-ETs. As a result, in addition to local divergence control in each site, a global mechanism is needed to guarantee ESR.>
Calton Pu, Wenwey Hseush, Gail E. Kaiser, Kun-Lung Wu, Philip S. Yu
ICDCS1
1993 Incremental Partial Evaluation: The Key to High Performance, Modularity and Portability in Operating Systems
abstract
Article Free Access Share on Incremental partial evaluation: the key to high performance, modularity and portability in operating systems Authors: Charles Consel View Profile , Calton Pu View Profile , Jonathan Walpole View Profile Authors Info & Claims PEPM '93: Proceedings of the 1993 ACM SIGPLAN symposium on Partial evaluation and semantics-based program manipulationAugust 1993 Pages 44–46https://doi.org/10.1145/154630.154635Online:01 August 1993Publication History 20citation287DownloadsMetricsTotal Citations20Total Downloads287Last 12 Months9Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Charles Consel, Calton Pu, Jonathan Walpole
PEPM2
1993 Performance comparison of dynamic policies for remote caching
abstract
Abstract In a distributed system, data servers (file systems and databases) can easily become bottlenecks. We propose an approach to offloading data access requests from overloaded data servers to nodes that are Idle or less busy. This approach is referred to asremote caching, and the idle or less busy nodes are calledmutual serversas they help out the busy server nodes on data accesses. In addition to server and client local caches, frequently accessed data are cached in the main memory of mutual servers, thus improving the data access time in the system. We evaluate several data propagation strategics among data servers and mutual servers. These include policies in which senders are active/passive and receivers are active/passive in initiating data propagation. For example, an active sender takes the initiative to offload data onto a passive receiver. Simulation results show that the active‐sender/passive‐receiver policy is the method of choice In most cases. Active‐Sender policies are best able to exploit the main memory of other Idle nodes in the expected normal condition where some nodes are overloaded and others are less loaded. AH active policies perform far better than the policy without remote caching even in the degenerated case where each node is equally loaded.
Calton Pu, Danilo Florissi, Patricia Soares, Philip S. Yu, Kun-Lung Wu
Concurr. Pract. Exp.1
1992 Performance Comparison of Active-Sender and Active-Receiver Policies for Distributed Caching
abstract
The authors propose a distributed caching approach to off-loading data access requests from overloaded data servers in a distributed system to nodes that are idle or less busy. Helping out the busy servers on data accesses, the idle or less busy nodes are called mutual servers. Frequently accessed data are cached in the main memory of mutual servers in addition to server and client local caches. The authors evaluate several data propagation strategies among data servers and mutual servers. Simulation results show that the active-sender passive-receiver policy is the method of choice in most cases. Active-sender policies are best able to exploit the main memory of other idle nodes in the expected normal condition where some nodes are overloaded and others are less loaded. All active policies perform far better than the policy without distributed caching.>
Calton Pu, Danilo Florissi, Patricia Soares, Kun-Lung Wu, Philip S. Yu
HPDC1
1992 Divergence Control for Epsilon-Serializability
abstract
The authors present divergence control methods for epsilon-serializability (ESR) in centralized databases. ESR alleviates the strictness of serializability (SR) in transaction processing by allowing for limited inconsistency. The bounded inconsistency is automatically maintained by divergence control (DC) methods in a way similar to the manner in which SR is maintained by concurrency control mechanisms, but DC for ESR allows more concurrency. Concrete representative instances of divergence-control methods are described based on two-phase locking, timestamp ordering, and optimistic approaches. The applicability of ESR is demonstrated by presenting the designs of DC methods using other most known inconsistency specifications, such as absolute value, age, and total number of nonserializably read data items.>
Kun-Lung Wu, Philip S. Yu, Calton Pu
ICDE3
1991 An Experiment on Measuring Application Performance over the Internet
abstract
The use of wide area networks (WANs) such as the Internet is growing at a tremendous rate.Such networks hold great promise for new types of distributed applications, which will be widely distributed, highly replicated, intensely interactive, and adaptive to many types of network conditions. Developing such applications will require a solid understanding of the performance and availability characteristics of WANs as they evolve. The ability to measure the effect of these conditions will, for example, be important for large-volume applications such as digital libraries, and for near-real-time applications such as collaborative research and teleconferencing.
Calton Pu, Frederick Korz, Robert C. Lehman
SIGMETRICS1
1991 Replica Control in Distributed Systems: An Asynchronous Approach
abstract
Article Free Access Share on Replica control in distributed systems: as asynchronous approach Authors: Calton Pu Department of Computer Science, Columbia University, New York, NY Department of Computer Science, Columbia University, New York, NYView Profile , Avraham Leff Department of Computer Science, Columbia University, New York, NY Department of Computer Science, Columbia University, New York, NYView Profile Authors Info & Claims SIGMOD '91: Proceedings of the 1991 ACM SIGMOD international conference on Management of dataApril 1991 Pages 377–386https://doi.org/10.1145/115790.115856Published:01 April 1991Publication History 153citation1,037DownloadsMetricsTotal Citations153Total Downloads1,037Last 12 Months45Last 6 weeks6 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Calton Pu, Avraham Leff
SIGMOD Conference1
1989 Threads and Input/Output in the Synthesis Kernel
abstract
The Synthesis operating system kernel combines several techniques to provide high performance, including kernel code synthesis, fine-grain scheduling, and optimistic synchronization. Kernel code synthesis reduces the execution path for frequently used kernel calls. Optimistic synchronization increases concurrency within the kernel. Their combination results in significant performance improvement over traditional operating system implementations. Using hardware and software emulating a SUN 3/160 running SUNOS, Synthesis achieves several times to several dozen times speedup for UNIX kernel calls and context switch times of 21 microseconds or faster.
Henry Massalin, Calton Pu
SOSP2
1988 Superdatabases for Composition of Heterogeneous Databases
abstract
Superdatabases, which are designed to compose and extend databases, are discussed. In particular, superdatabases allow consistent update across heterogeneous databases. The general architecture of superdatabases is summarized, and some sufficient conditions to make the element databases composable are given. The design of a superdatabase capable of gluing the element databases together is described, and an implementation plan is sketched. Related work on many different aspects of heterogeneous databases is summarized and compared with the author's study.>
Calton Pu
ICDE1
1988 Split-Transactions for Open-Ended Activities
Calton Pu, Gail E. Kaiser, Norman C. Hutchinson
VLDB1
1988 Regeneration of Replicated Objects: A Technique and Its Eden Implementation
abstract
A replicated directory system based on a method called regeneration is designed and implemented. The directory system allows selection of arbitrary object to be replicated, choice of the number of replicas for each object, and placement of the copies on machines with independent failure modes. Copies can become inaccessible due to node crashes, but as long as a single copy survives, the replication level is restored by automatically replacing lost copies on other active machines. The focus is on a regeneration algorithm for replica replacement and its application to a replicated directory structure in the Eden local area network. A simple probabilistic approach is used to compare the availability provided by the algorithm to three other replication techniques.>
Calton Pu, Jerre D. Noe, Andrew Proudfoot
IEEE Trans. Software Eng.1
1987 Design and Implementation of Nested Transactions in Eden
Calton Pu, Jerre D. Noe
SRDS1
1986 Regeneration of Replicated Objects: A Technique and Its Eden Implementation
abstract
We have designed and implemented a replicated directory system based on a method called Regeneration. The directory system allows selection of arbitrary objects to be replicated, choice of the number of replicas for each object, and placement of the copies on machines with independent failure modes. Copies may become inaccessible due to node crashes, but as long as a single copy survives, the replication level is restored by automatically replacing lost copies on other active machines. The paper focuses on the Regeneration algorithm for replica replacement and on its application to a replicated directory structure in the Eden system. Analytically, we use a simple probabilistic approach to compare the availability provided by the algorithm with other replication techniques. Empirically, we have measured the performance of the implementation.
Calton Pu, Jerre D. Noe, Andrew Proudfoot
ICDE1
1986 On-the-Fly, Incremental, Consistent Reading of Entire Databases
Calton Pu
Algorithmica1
1985 On-the-Fly, Incremental, Consistent Reading of Entire Databases
Calton Pu
VLDB1