Roger S. Barga

dblp:b/RogerSBarga · DBLP profile ↗
← Back
33ranked-venue papers
14as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 17 · 10 first-authorSystems, architecture and hardware · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-authorComputer networks · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Distributed systems · 43% Cloud and datacenter computing · 31% High-performance computing · 26%
Databases, data mining, and information retrieval
7 papers
Transaction processing and concurrency control · 36% Spatial and temporal data management · 34% Data stream processing · 17%
Software engineering, system software, and programming languages
2 papers
Services computing and microservices · 60% Concurrent programming · 40%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 24 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
0.252004
Improving Logging and Recovery Performance in Phoenix/App · ICDE 2004
EOS: Exactly-Once E-Service Middleware · VLDB 2002
Recovery Guarantees for General Multi-Tier Applications · ICDE 2002
Cloud and datacenter computing › big data analytics
cloud data analytics service
0.112012
Project Daytona: Data Analytics as a Cloud Service · ICDE 2012
Distributed systems › distributed scheduling
operator placement
0.112011
Accurate latency estimation in a distributed event processing system · ICDE 2011
Transaction processing and concurrency control › isolation levels
snapshot isolation
0.122006
Transaction Time Support Inside a Database Engine · ICDE 2006
Immortal DB: transaction time support for SQL server · SIGMOD Conference 2005
Spatial and temporal data management › temporal databases
transaction time
0.122006
Recovery from "bad" user transactions · SIGMOD Conference 2006
Immortal DB: transaction time support for SQL server · SIGMOD Conference 2005
Cloud and datacenter computing › cloud applications
cloud-based bioinformatics
0.112010
AzureBlast: a case study of developing science applications on the cloud · HPDC 2010
High-performance computing › scientific computing systems
scientific application portability
0.112010
AzureBlast: a case study of developing science applications on the cloud · HPDC 2010
High-performance computing › scientific workflow
scientific workflow management
0.112010
Versioning for workflow evolution · HPDC 2010
Services computing and microservices
business process management
0.112007
Categorization and Optimization of Synchronization Dependencies in Business Processes · ICDE 2007
Concurrent programming › synchronization
process synchronization
0.112007
Categorization and Optimization of Synchronization Dependencies in Business Processes · ICDE 2007
Transaction processing and concurrency control
recovery
0.122006
Recovery from "bad" user transactions · SIGMOD Conference 2006
Phoenix: Making Applications Robust · SIGMOD Conference 1999
Transaction processing and concurrency control
concurrency control
0.112006
Transaction Time Support Inside a Database Engine · ICDE 2006
Spatial and temporal data management
temporal databases
0.112006
Transaction Time Support Inside a Database Engine · ICDE 2006
Indexing and storage engines
temporal indexing
0.112006
Transaction Time Support Inside a Database Engine · ICDE 2006
Spatial and temporal data management › temporal databases
transaction-time database
0.112006
Transaction Time Support Inside a Database Engine · ICDE 2006
Distributed systems › fault tolerance › failure recovery
logging and recovery
0.012004
Improving Logging and Recovery Performance in Phoenix/App · ICDE 2004
Query processing and optimization › query planning
query plan selection
0.012011
Accurate latency estimation in a distributed event processing system · ICDE 2011
Distributed systems › fault tolerance
exactly-once processing
0.012002
EOS: Exactly-Once E-Service Middleware · VLDB 2002
Bioinformatics and computational biology › sequence alignment › heuristic alignment
BLAST
0.012010
AzureBlast: a case study of developing science applications on the cloud · HPDC 2010
Bioinformatics and computational biology
sequence alignment
0.012010
AzureBlast: a case study of developing science applications on the cloud · HPDC 2010
High-performance computing › distributed computing infrastructure
escience infrastructure
0.012010
Versioning for workflow evolution · HPDC 2010
Spatial and temporal data management
moving objects
0.012005
Immortal DB: transaction time support for SQL server · SIGMOD Conference 2005
Distributed systems › fault tolerance
checkpointing
0.012004
Improving Logging and Recovery Performance in Phoenix/App · ICDE 2004
Transaction processing and concurrency control › transaction models
extended transaction models
0.011995
A Practical and Modular Implementation of Extended Transaction Models · VLDB 1995

Methods — techniques the papers use, named apart from their topics

queueing analysis · 0.2maximum cumulative excess metric · 0.2performance evaluation · 0.2large-scale parallelism · 0.2mapreduce · 0.1version control · 0.1provenance tracking · 0.1logging · 0.1dependency optimization · 0.1dataflow programming · 0.1versioning · 0.1time-split pages · 0.1lazy timestamping · 0.1redo recovery · 0.0log optimization · 0.0ODBC · 0.0
YearPublicationVenuePosition
2016 Processing Big Data in Motion
abstract
Summary form only given. Streaming analytics is about identifying and responding to events happening in your business, in your service or application, and with your customers in near real-time. Sensors, mobile and IoT devices, social networks, and online transactions are all generating data that can be monitored constantly to enable a business to detect and then act on events and insights before they lose their value. The need for large scale, real-time stream processing of big data in motion is more evident than ever before but the potential remains largely untapped by most firms. It's not the size but rather the speed at which this data must be processed that presents the greatest technical challenges. Streaming analytics systems can enable business to inspect, correlate and analyze data in real-time to extract insights in the same manner that traditional analytics tools have allowed them to do with data at rest. In this talk I will draw upon our experience with Amazon Kinesis data streaming services to highlight use cases, discuss technical challenges and approaches, and look ahead to the future of stream data processing and role of cloud computing.
Roger S. Barga
IC2E1
2012 Project Daytona: Data Analytics as a Cloud Service
abstract
Spreadsheets are established data collection and analysis tools in business, technical computing and academic research. Excel, for example, offers an attractive user interface, provides an easy to use data entry model, and offers substantial interactivity for what-if analysis. However, spreadsheets and other common client applications do not offer scalable computation for large scale data analytics and exploration. Increasingly researchers in domains ranging from the social sciences to environmental sciences are faced with a deluge of data, often sitting in spreadsheets such as Excel or other client applications, and they lack a convenient way to explore the data, to find related data sets, or to invoke scalable analytical models over the data. To address these limitations, we have developed a cloud data analytics service based on Daytona, which is an iterative MapReduce runtime optimized for data analytics. In our model, Excel and other existing client applications provide the data entry and user interaction surfaces, Daytona provides a scalable runtime on the cloud for data analytics, and our service seamlessly bridges the gap between the client and cloud. Any analyst can use our data analytics service to discover and import data from the cloud, invoke cloud scale data analytics algorithms to extract information from large datasets, invoke data visualization, and then store the data back to the cloud all through a spreadsheet or other client application they are already familiar with.
Roger S. Barga, Jaliya Ekanayake
ICDE1
2011 A Scalable Communication Runtime for Clouds
abstract
Leveraging cloud computing to acquire the necessary computation resources to scale out parallel applications is becoming common practice. However, many such applications also require communication and synchronization between processes. Although, commercial cloud platforms provide ready access to scalable compute and storage services, implementing communication and synchronization between cooperating processes and efficiently exchanging arbitrary size messages remains a challenge for application developers. In clouds, durable queues provide basic abstractions for communication. However, they are not sufficient for applications that require transferring arbitrary size messages or for applications that require higher level abstractions such as broadcast. Furthermore, direct socket based communication is susceptible to various fluctuations common in data center environments. We envision a solution to this problem that leverages scalable storage services, queues, and direct socket based communication. Publish/subscribe (pub/sub) is a well-known communication pattern that can achieve the above capabilities in a loosely coupled fashion, which is highly desirable in cloud environments where most services are asynchronous. In this paper, we describe the architecture of a pub/sub library implemented on a commercial cloud computing platform, which can be used to develop various parallel applications. We also present an evaluation of our implementation using both micro benchmarks and a real world application. Together, these demonstrate that our approach is both effective and scalable in performing communication and synchronization in cloud scale applications.
Jaliya Ekanayake, Jared Jackson, Roger S. Barga, Atilla Soner Balkir
IEEE CLOUD4
2011 Accurate latency estimation in a distributed event processing system
abstract
A distributed event processing system consists of one or more nodes (machines), and can execute a directed acyclic graph (DAG) of operators called a dataflow (or query), over long-running high-event-rate data sources. An important component of such a system is cost estimation, which predicts or estimates the “goodness” of a given input, i.e., operator graph and/or assignment of individual operators to nodes. Cost estimation is the foundation for solving many problems: optimization (plan selection and distributed operator placement), provisioning, admission control, and user reporting of system misbehavior. Latency is a significant user metric in many commercial real-time applications. Users are usually interested in quantiles of latency, such as worst-case or 99thpercentile. However, existing cost estimation techniques for event-based dataflows use metrics that, while they may have the side-effect of being correlated with latency, do not directly or provably estimate latency. In this paper, we propose a new cost estimation technique using a metric called Mace (Maximum cumulative excess). Mace is provably equivalent to maximum system latency in a (potentially complex, multi-node) distributed event-based system. The close relationship to latency makes Mace ideal for addressing the problems described earlier. Experiments with real-world datasets on Microsoft StreamInsight deployed over 1-13 nodes in a data center validate our ability to closely estimate latency (within 4%), and the use of Mace for plan selection and distributed operator placement.
Badrish Chandramouli, Jonathan Goldstein, Roger S. Barga, Mirek Riedewald, Ivo Santos
ICDE3
2011 Analysis of approaches for supporting the Open Provenance Model: A case study of the Trident workflow workbench
Yogesh L. Simmhan, Roger S. Barga
Future Gener. Comput. Syst.2
2010 Performing Large Science Experiments on Azure: Pitfalls and Solutions
abstract
Carrying out science at extreme scale is the next generational challenge facing the broad field of scientific research. Cloud computing offers to potential for an increasing number of researchers to have ready access to the large scale compute resources required to tackle new challenges in their field. Unfortunately barriers of complexity remain for researchers untrained in cloud programming. In this paper we examine how cloud based architectures can be used to solve large scale research experiments in a manner that is easily accessible for researchers with limited programming experience, using their existing computational tools. We examine the top challenges identified in our own large-scale science experiments running on the Windows Azure platform and then describe a Cloud-based parameter sweep prototype (dubbed Cirrus) which provides a framework of solutions for each challenge.
Jared Jackson, Jaliya Ekanayake, Roger S. Barga, Nelson Araujo
CloudCom4
2010 AzureBlast: a case study of developing science applications on the cloud
abstract
Cloud computing has emerged as a new approach to large scale computing and is attracting a lot of attention from the scientific and research computing communities. Despite its growing popularity, it is still unclear just how well the cloud model of computation will serve scientific applications. In this paper we analyze the applicability of cloud to the sciences by investigating an implementation of a well known and computationally intensive algorithm called BLAST. BLAST is a very popular life sciences algorithm used commonly in bioinformatics research. The BLAST algorithm makes an excellent case study because it is both crucial to many life science applications and its characteristics are representative of many applications important to data intensive scientific research. In our paper we introduce a methodology that we use to study the applicability of cloud platforms to scientific computing and analyze the results from our study. In particular we examine the best practices of handling the large scale parallelism and large volumes of data. While we carry out our performance evaluation on Microsoft's Windows Azure the results readily generalize to other cloud platforms.
Jared Jackson, Roger S. Barga
HPDC3
2010 Versioning for workflow evolution
abstract
Scientists working in eScience environments often use workflows to carry out their computations. Since the workflows evolve as the research itself evolves, these workflows can be a tool for tracking the evolution of the research. Scientists can trace their research and associated results through time or even go back in time to a previous stage and fork to a new branch of research. In this paper we introduce the workflow evolution framework (EVF), which is demonstrated through implementation in the Trident workflow workbench. The primary contribution of the EVF is efficient management of knowledge associated with workflow evolution. Since we believe evolution can be used for workflow attribution, our framework will motivate researchers to share their workflows and get the credit for their contributions.
Eran Chinthaka Withana, Beth Plale, Roger S. Barga, Nelson Araujo
HPDC3
2010 Emerging Trends and Converging Technologies in Data Intensive Scalable Computing
Roger S. Barga
SSDBM1
2010 Client + Cloud: Evaluating Seamless Architectures for Visual Data Analytics in the Ocean Sciences
Keith Grochow, Bill Howe, Mark Stoermer, Roger S. Barga, Edward D. Lazowska
SSDBM4
2009 DryadLINQ for Scientific Analyses
abstract
Applying high level parallel runtimes to data/compute intensive applications is becoming increasingly common. The simplicity of the MapReduce programming model and the availability of open source MapReduce runtimes such as Hadoop, are attracting more users to the MapReduce programming model. Microsoft has released DryadLINQ for academic use, allowing users to experience a new programming model and a runtime that is capable of performing large scale data/compute intensive analyses. In this paper, we present our experience in applying DryadLINQ for a series of scientific data analysis applications, identify their mapping to the DryadLINQ programming model, and compare their performances with Hadoop implementations of the same applications.
Jaliya Ekanayake, Thilina Gunarathne, Geoffrey C. Fox, Atilla Soner Balkir, Christophe Poulain, Nelson Araujo, Roger S. Barga
eScience7
2009 Building Reliable Data Pipelines for Managing Community Data Using Scientific Workflows
abstract
The growing amount of scientific data from sensors and field observations is posing a challenge to ¿data valets¿ responsible for managing them in data repositories. These repositories built on commodity clusters need to reliably ingest data continuously and ensure its availability to a wide user community. Workflows provide several benefits to modeling data-intensive science applications and many of these benefits can help manage the data ingest pipelines too. But using workflows is not panacea in itself and data valets need to consider several issues when designing workflows that behave reliably on fault prone hardware while retaining the consistency of the scientific data. In this paper, we propose workflow designs for reliable data ingest in a distributed environment and identify workflow framework features to support resilience. We illustrate these using the data pipeline for the Pan-STARRS repository, one of the largest digital surveys that accumulates 100TB of data annually to support 300 astronomers.
Yogesh L. Simmhan, Catharine van Ingen, Alex Szalay, Roger S. Barga, Jim Heasley
eScience4
2008 The Trident Scientific Workflow Workbench
abstract
In our demonstration we present Trident, a scientific workflow workbench built on top of a commercial workflow system to leverage existing functionality to the extent possible. Trident is being developed in collaboration with the scientific computing community for use in a number of ongoing eScience projects that make use of scientific workflows, in particular the Pan-STARRS sky survey project and the Ocean Observatory Initiative. In our demonstration of Trident we will illustrate the ability to utilize both local and cloud resources for storage and execution, as well as services such as provenance, monitoring, logging and scheduling workflows over clusters. Our goal is to release Trident in early 2009 as an open source accelerator for others to use for eScience projects and to continue extending with support for new workflow features and services.
Roger S. Barga, Jared Jackson, Nelson Araujo, Dean Guo, Nitin Gautam, Yogesh L. Simmhan
eScience1
2008 On Building Scientific Workflow Systems for Data Management in the Cloud
abstract
Scientific workflows have become an archetype to model in silico experiments in the Cloud by scientists. There is a class of workflows that are used to by "data valets" to prepare raw data from scientific instruments into a science-ready form for use by scientists. These share data-intensive traits with traditional scientific workflows, yet differ significantly, for example, in the required degree of reliability and the type of provenance collected. We compare and contrast science application and data valet workflows through exemplar eScience projects to drive shared and unique requirements for scientific workflows across diverse users in a Science Cloud.
Yogesh L. Simmhan, Roger S. Barga, Catharine van Ingen, Edward D. Lazowska, Alex Szalay
eScience2
2008 Capturing Workflow Event Data for Monitoring, Performance Analysis, and Management of Scientific Workflows
abstract
To effectively support real-time monitoring and performance analysis of scientific workflow execution, varying levels of event data must be captured and made available to interested parties. This paper discusses the creation of an ontology-aware workflow monitoring system for use in the Trident system which utilizes a distributed publish/subscribe event model. The implementation of the publish/subscribe system is discussed and performance results are presented.
Matthew D. Valerio, Satya Sanket Sahoo, Roger S. Barga, Jared Jackson
eScience3
2008 Automatic capture and efficient storage of e-Science experiment provenance
abstract
Abstract For the first provenance challenge, we introduce a layered model to represent workflow provenance that allows navigation from an abstract model of the experiment to instance data collected during a specific experiment run. We outline modest extensions to a commercial workflow engine so it will automatically capture provenance at workflow runtime. We also present an approach to store this provenance data in a relational database. Finally, we demonstrate how core provenance queries in the challenge can be expressed in SQL and discuss the merits of our layered representation. Copyright © 2007 John Wiley & Sons, Ltd.
Roger S. Barga, Luciano A. Digiampietri
Concurr. Comput. Pract. Exp.1
2008 Special Issue: The First Provenance Challenge
abstract
Abstract The first Provenance Challenge was set up in order to provide a forum for the community to understand the capabilities of different provenance systems and the expressiveness of their provenance representations. To this end, a functional magnetic resonance imaging workflow was defined, which participants had to either simulate or run in order to produce some provenance representation, from which a set of identified queries had to be implemented and executed. Sixteen teams responded to the challenge, and submitted their inputs. In this paper, we present the challenge workflow and queries, and summarize the participants' contributions. Copyright © 2007 John Wiley & Sons, Ltd.
Luc Moreau 0001, Bertram Ludäscher, Ilkay Altintas, Roger S. Barga, Shawn Bowers, Steven P. Callahan, George Chin, Ben Clifford, Shirley Cohen, Sarah Cohen Boulakia, Susan B. Davidson, Ewa Deelman, Luciano A. Digiampietri, Ian T. Foster, Juliana Freire, James Frew, Joe Futrelle, Tara Gibson, Yolanda Gil, Carole A. Goble, Jennifer Golbeck, Paul Groth, David A. Holland, Jihie Kim, David Koop, Ales Krenek, Timothy M. McPhillips, Gaurang Mehta, Simon Miles, Dominic Metzger, Steve Munroe, James D. Myers, Beth Plale, Norbert Podhorszki, Varun Ratnakar, Emanuele Santos, Carlos Scheidegger, Karen Schuchardt, Margo I. Seltzer, Yogesh L. Simmhan, Cláudio T. Silva, Peter Slaughter, Eric G. Stephan, Robert Stevens 0001, Daniele Turi, Huy T. Vo, Michael Wilde, Jun Zhao 0003, Yong Zhao 0009
Concurr. Comput. Pract. Exp.4
2007 Consistent Streaming Through Time: A Vision for Event Stream Processing
Roger S. Barga, Jonathan Goldstein, Mohamed H. Ali, Mingsheng Hong
CIDR1
2007 Categorization and Optimization of Synchronization Dependencies in Business Processes
abstract
The current approach for modeling synchronization in business processes relies on sequencing constructs, such as sequence, parallel etc. However, sequencing constructs obfuscate the true source of dependencies in a business process. Moreover, because of the nested structure and scattered code that results from using sequencing constructs, it is hard to add or delete additional constraints without over-specifying necessary constraints or invalidating existing ones. We propose a dataflow programming approach in which dependencies are explicitly modeled to guide activity scheduling. We first give a systematic categorization of dependencies: data, control, service and cooperation. Each dimension models dependency from its own point of view. Then we show that dependencies of various kinds can be first merged and then optimized to generate a minimal dependency set, which guarantees high concurrency and minimal maintenance cost for process execution.
Qinyi Wu, Calton Pu, Akhil Sahai, Roger S. Barga
ICDE4
2006 Transaction Time Support Inside a Database Engine
abstract
Transaction time databases retain and provide access to prior states of a database. An update "inserts" a new record while preserving the old version. Immortal DB builds transaction time database support into a database engine, not in middleware. It supports as of queries returning records current at the specified time. It also supports snapshot isolation concurrency control. Versions are stamped with the "clock times" of their updating transactions. The timestamp order agrees with transaction serialization order. Lazy timestamping propagates timestamps to transaction updates after commit. Versions are kept in an integrated storage structure, with historical versions initially stored with current data. Time-splits of pages permit large histories to be maintained, and enable time based indexing, which is essential for high performance historical queries. Experiments show that Immortal DB introduces little overhead for accessing recent database states while providing access to past states.
David B. Lomet, Roger S. Barga, Mohamed F. Mokbel, German Shegalov, Rui Wang 0002, Yunyue Zhu
ICDE2
2006 DSCWeaver: Synchronization-Constraint Aspect Extension to Procedural Process Specification Languages
abstract
BPEL is emerging as an open-standards language for Web service composition. However, its procedural style can lead to inflexible and tangled code for managing a crosscutting aspect - synchronization constraints that define permissible sequences of execution for activities in a process. In this paper, we present DSCWeaver, a tool that enables a synchronization-aspect extension to BPEL. It uses DSCL, a synchronization expression language, to specify constraints. DSCL has the desirable features of declarative syntax, fine granularity, and validation support. A designer can use DSCL to describe and validate the synchronization behavior and rely on DSCWeaver to generate BPEL code. We demonstrate the advantages of our approach in a service deployment process and evaluate its performance using two metrics: lines of code (LoC) and places to visit (PtV). Evaluation results show that our approach can effectively reduce development effort of process designers while providing performance competitive to un-woven BPEL code
Qinyi Wu, Calton Pu, Akhil Sahai, Roger S. Barga, Gueyoung Jung
ICWS4
2006 Recovery from "bad" user transactions
abstract
User written transaction code is responsible for the "C" in ACID transactions, i.e., taking the database from one consistent state to the next. However, user transactions can be flawed and lead to inconsistent (or invalid) states. Database systems usually correct invalid data using "point in time" recovery, a costly process that installs a backup and rolls it forward. The result is long outages and the "de-commit" of many valid transactions, which must then be re-submitted, frequently manually. We have implemented in our transaction-time database system a technique in which only data tainted by a flawed transaction and transactions dependent upon its updates are "removed". This process identifies and quarantines tainted data despite the complication of determining transactions dependent on data written by the flawed transaction. A further property of our implementation is that no backup needs to be installed for this because the prior transaction-time states provide an online backup.
David B. Lomet, Zografoula Vagena, Roger S. Barga
SIGMOD Conference3
2005 Immortal DB: transaction time support for SQL server
abstract
Immortal DB builds transaction time database support into the SQL Server engine, not in middleware. Transaction time databases retain and provide access to prior states of a database. An update "inserts" a new record while preserving the old version. The system supports as of queries returning records current at the specified time. It also supports snapshot isolation concurrency control. Versions are stamped with the times of their updating transactions. The timestamp order agrees with transaction serialization order. Lazy timestamping propagates timestamps to all updates of a transaction after commit. All versions are kept in an integrated storage structure, with historical versions initially stored with current data. Time-splits of pages permit large histories to be maintained, and enable time based indexing. We demonstrate Immortal DB with a moving objects application that tracks cars in the Seattle area.
David B. Lomet, Roger S. Barga, Mohamed F. Mokbel, German Shegalov, Rui Wang 0002, Yunyue Zhu
SIGMOD Conference2
2004 Improving Logging and Recovery Performance in Phoenix/App
abstract
Phoenix/App supports software components whose states are made persistent across a system crash via redo recovery, replaying logged interactions. Our initial prototype force logged all request/reply events resulting from intercomponent method calls and returns. We describe an enhanced prototype that implements: (i) log optimizations to improve normal execution performance; and (ii) checkpointing to improve recovery performance. Logging is reduced in two ways: (1) we only log information required to remove nondeterminism, and we only force the log when an event "commits" the state of the component to other parts of the system; (2) we introduce new component types that provide our enhanced system with more information, enabling further reduction in logging. To improve recovery performance, we save the values of the fields of a component to the log in an application "checkpoint". We describe the system elements that we exploit for these optimizations, and characterize the performance gains that result.
Roger S. Barga, Shimin Chen, David B. Lomet
ICDE1
2004 Recovery guarantees for Internet applications
abstract
Internet-based e-services require application developers to deal explicitly with failures of the underlying software components, for example web servers, servlets, browser sessions, and so forth. This complicates application programming, and may expose failures to end users. This paper presents a framework for an application-independent infrastructure that provides recovery guarantees and masks almost all system failures, thus relieving the application programmer from having to deal with these failures---by making applications "stateless." The main concept is an interaction contract between two components regarding message and state preservation. The framework provides comprehensive recovery encompassing data, messages, and the states of application components. We describe techniques to reduce logging cost, allow effective log truncation, and permit independent recovery for critical components. We illustrate the framework's utility via web-based e-services scenarios. Its feasibility is demonstrated by our prototype implementation of interaction contracts based on the Apache web server and the PHP servlet engine. Finally, we discuss industrial relevance for middleware architectures such as. Net or J2EE.
Roger S. Barga, David B. Lomet, German Shegalov, Gerhard Weikum
ACM Trans. Internet Techn.1
2003 Persistent Applications via Automatic Recovery
abstract
Building highly available enterprise applications using Web-oriented middleware is hard. Runtime implementations frequently do not address the problems of application state persistence and fault-tolerance, placing the burden of managing session state and, in particular, handling system failures on application programmers. This paper describes Phoenix/APP, a runtime service based on the notion of recovery guarantees. Phoenix/APP transparently masks failures and automatically recovers component-based applications. This both increases application availability and simplifies application development. We demonstrate the feasibility of this approach by describing the design and implementation of Phoenix/APP in Microsoft's .NET runtime and present results on the cost of persisting and recovering component-based applications.
Roger S. Barga, David B. Lomet, Stelios Paparizos, Sirish Chandrasekaran
IDEAS1
2002 Recovery Guarantees for General Multi-Tier Applications
abstract
Database recovery does not mask failures to applications and users. Recovery is needed that considers data, messages and application components. Special cases have been studied, but clear principles for recovery guarantees in general multi-tier applications such as Web-based e-services are missing. We develop a framework for recovery guarantees that masks almost all failures. The main concept is an interaction contract between two components, a pledge as to message and state persistence, and contract release. Contracts are composed into system-wide agreements so that a set of components is provably recoverable with exactly-once message delivery and execution, except perhaps for crash-interrupted user input or output. Our implementation techniques reduce the data logging cost, allow effective log truncation, and provide independent recovery for critical server components. Interaction contracts form the basis for our Phoenix/COM project on persistent components. Our framework's utility is demonstrated with a case study of a web-based e-service.
Roger S. Barga, David B. Lomet, Gerhard Weikum
ICDE1
2002 EOS: Exactly-Once E-Service Middleware
German Shegalov, Gerhard Weikum, Roger S. Barga, David B. Lomet
VLDB3
2001 Measuring and Optimizing a System for Persistent Database Sessions
abstract
High availability for both data and applications is rapidly becoming a business requirement. While database systems support recovery, providing high database availability, applications may still lose work because of server outages. When a server crashes, any volatile state associated with the application's database session is lost and the application may require an operator-assisted restart. This exposes server failures to end-users and always degrades application availability. Our Phoenix/ODBC system supports persistent database sessions that can survive a database crash without the application being aware of the outage, except for possible timing considerations. This improves application availability and eliminates the application programming needed to cope with database crashes. Phoenix/ODBC requires no changes to the database system, data access routines or applications. Hence, it can be deployed in any application that uses ODBC to access a database. Further, our generic approach can be exploited for a variety of data access protocols. In this paper, we describe the design of Phoenix/ODBC and introduce an extension to optimize the response time and to reduce overhead for OLTP workloads. We present a performance evaluation using the TPC-C and TPC-H benchmarks that demonstrate Phoenix/ODBC's extra overhead is modest.
Roger S. Barga, David B. Lomet
ICDE1
2000 Persistent Client-Server Database Sessions
Roger S. Barga, David B. Lomet, Thomas Baby, Sanjay Agrawal 0001
EDBT1
1999 Phoenix: Making Applications Robust
abstract
article Phoenix: making applications robust Share on Authors: Roger Barga Microsoft Corporation, One Microsoft Way, Redmond, WA Microsoft Corporation, One Microsoft Way, Redmond, WAView Profile , David B. Lomet Microsoft Corporation, One Microsoft Way, Redmond, WA Microsoft Corporation, One Microsoft Way, Redmond, WAView Profile Authors Info & Claims ACM SIGMOD RecordVolume 28Issue 2June 1999 pp 562–564https://doi.org/10.1145/304181.304577Online:01 June 1999Publication History 10citation280DownloadsMetricsTotal Citations10Total Downloads280Last 12 Months8Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Roger S. Barga, David B. Lomet
SIGMOD Conference1
1996 Differential Evaluation of Continual Queries
abstract
We define continual queries as a useful tool for monitoring of updated information. Continual queries are standing queries that monitor the source data and notify the users whenever new data matches the query. In addition to periodic refresh, continual queries include Epsilon Transaction concepts to allow users to specify query refresh based on the magnitude of updates. To support efficient processing of continual queries, we propose a differential re-evaluation algorithm (DRA), which exploits the structure and information contained in both the query expressions and the database update operations. The DRA design can be seen as a synthesis of previous research on differential files, incremental view maintenance, and active databases.
Ling Liu 0001, Calton Pu, Roger S. Barga
ICDCS3
1995 A Practical and Modular Implementation of Extended Transaction Models
Roger S. Barga, Calton Pu
VLDB1