Olga Papaemmanouil

dblp:37/3024 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
0since 2021 · last 2019
0000-0003-4526-3595ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 26 · 6 first-authorArtificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
13 papers
Query processing and optimization · 60% Information retrieval · 12% Database system architecture and tuning · 8%
Computer architecture, parallel and distributed computing, and storage systems
10 papers
Cloud and datacenter computing · 69% Distributed systems · 16% Performance modeling and evaluation · 15%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
interactive data exploration
0.942016
AIDE: An Active Learning-Based Approach for Interactive Data Exploration · IEEE Trans. Knowl. Data Eng. 2016
AIDE: An Automatic User Navigation System for Interactive Data Exploration · Proc. VLDB Endow. 2015
Overview of Data Exploration Techniques · SIGMOD Conference 2015
Query processing and optimization
query optimization
0.842019
Neo: A Learned Query Optimizer · Proc. VLDB Endow. 2019
Skew-Aware Join Optimization for Array Databases · SIGMOD Conference 2015
Devel-op: An optimizer development environment · ICDE 2014
Data mining
exploratory data analysis
0.422015
Overview of Data Exploration Techniques · SIGMOD Conference 2015
Explore-by-example: an automatic query steering framework for interactive data exploration · SIGMOD Conference 2014
Query processing and optimization › query optimization › learned query optimization
learned query optimizer
0.412019
Neo: A Learned Query Optimizer · Proc. VLDB Endow. 2019
Information retrieval › evaluation
query performance prediction
0.412019
Plan-Structured Deep Neural Network Models for Query Performance Prediction · Proc. VLDB Endow. 2019
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.312018
NashDB: An End-to-End Economic Method for Elastic Database Fragmentation, Replication, and Provisioning · SIGMOD Conference 2018
Query processing and optimization › query planning
query plan generation
0.322019
Devel-op: An optimizer development environment · ICDE 2014
Neo: A Learned Query Optimizer · Proc. VLDB Endow. 2019
Query processing and optimization
query scheduling
0.312017
A Learning-Based Service for Cost and Performance Management of Cloud Databases · ICDE 2017
Database system architecture and tuning
workload management
0.312017
A Learning-Based Service for Cost and Performance Management of Cloud Databases · ICDE 2017
Cloud and datacenter computing
database-as-a-service
0.312017
A Learning-Based Service for Cost and Performance Management of Cloud Databases · ICDE 2017
Distributed systems › peer-to-peer systems
overlay networks
0.342009
Supporting Generic Cost Models for Wide-Area Stream Processing · ICDE 2009
XPORT: extensible profile-driven overlay routing trees · SIGMOD Conference 2006
Extensible optimization in overlay dissemination trees · SIGMOD Conference 2006
Information retrieval
query prediction
0.212016
AIDE: An Active Learning-Based Approach for Interactive Data Exploration · IEEE Trans. Knowl. Data Eng. 2016
Cloud and datacenter computing › job scheduling
query scheduling
0.212016
WiSeDB: A Learning-based Workload Management Advisor for Cloud Databases · Proc. VLDB Endow. 2016
Cloud and datacenter computing › resource management
workload management
0.212016
WiSeDB: A Learning-based Workload Management Advisor for Cloud Databases · Proc. VLDB Endow. 2016
Query processing and optimization
join processing
0.212015
Skew-Aware Join Optimization for Array Databases · SIGMOD Conference 2015
Query processing and optimization › interactive query processing
query steering
0.212014
Explore-by-example: an automatic query steering framework for interactive data exploration · SIGMOD Conference 2014
Data stream processing
distributed stream processing
0.122009
Supporting Generic Cost Models for Wide-Area Stream Processing · ICDE 2009
Distributed operation in the Borealis stream processing engine · SIGMOD Conference 2005
Performance modeling and evaluation
performance prediction
0.112011
Performance prediction for concurrent database workloads · SIGMOD Conference 2011
Performance modeling and evaluation › performance prediction
query performance prediction
0.112011
Performance prediction for concurrent database workloads · SIGMOD Conference 2011
Performance modeling and evaluation
workload characterization
0.112011
Performance prediction for concurrent database workloads · SIGMOD Conference 2011
Query processing and optimization
query execution
0.112019
Plan-Structured Deep Neural Network Models for Query Performance Prediction · Proc. VLDB Endow. 2019
Cloud and datacenter computing › cluster resource management and scheduling
cluster provisioning
0.112019
NashDB: Fragmentation, Replication, and Provisioning using Economic Methods · Proc. VLDB Endow. 2019
Data stream processing
continuous query processing
0.112008
Simultaneous Equation Systems for Query Processing on Continuous-Time Data Streams · ICDE 2008
Query processing and optimization › query execution
query operator implementation
0.112008
Simultaneous Equation Systems for Query Processing on Continuous-Time Data Streams · ICDE 2008
Machine learning and data management
data management for machine learning
0.112016
AIDE: An Active Learning-Based Approach for Interactive Data Exploration · IEEE Trans. Knowl. Data Eng. 2016
Cloud and datacenter computing
resource provisioning
0.112016
WiSeDB: A Learning-based Workload Management Advisor for Cloud Databases · Proc. VLDB Endow. 2016
Content delivery and video streaming › content sharing
content-based dissemination
0.122006
SemCast: Semantic Multicast for Content-based Data Dissemination · ICDE 2005
XPORT: extensible profile-driven overlay routing trees · SIGMOD Conference 2006
Data models and query languages
multidimensional array data
0.112015
Skew-Aware Join Optimization for Array Databases · SIGMOD Conference 2015
Information retrieval
relevance feedback
0.112015
AIDE: An Automatic User Navigation System for Interactive Data Exploration · Proc. VLDB Endow. 2015
Cloud and datacenter computing › multi-tenancy
multi-tenant database
0.012011
Performance prediction for concurrent database workloads · SIGMOD Conference 2011

Methods — techniques the papers use, named apart from their topics

economic model · 1.0deep neural network · 0.8optimization techniques · 0.7game-theoretic balancing · 0.7machine learning · 0.6supervised learning · 0.2decision tree · 0.2classification algorithms · 0.2active learning · 0.2sampling · 0.2relevance feedback · 0.2cost-based optimization · 0.2queueing model · 0.1profile aggregation · 0.1decentralized optimization · 0.1cost model · 0.1tree transformation protocol · 0.1profile matching · 0.1
YearPublicationVenuePosition
2019 Towards a Hands-Free Query Optimizer through Deep Learning
Ryan Marcus, Olga Papaemmanouil
CIDR2
2019 Neo: A Learned Query Optimizer
abstract
Query optimization is one of the most challenging problems in database systems. Despite the progress made over the past decades, query optimizers remain extremely complex components that require a great deal of hand-tuning for specific workloads and datasets. Motivated by this shortcoming and inspired by recent advances in applying machine learning to data management challenges, we introduce Neo ( Neural Optimizer ), a novel learning-based query optimizer that relies on deep neural networks to generate query executions plans. Neo bootstraps its query optimization model from existing optimizers and continues to learn from incoming queries, building upon its successes and learning from its failures. Furthermore, Neo naturally adapts to underlying data patterns and is robust to estimation errors. Experimental results demonstrate that Neo, even when bootstrapped from a simple optimizer like PostgreSQL, can learn a model that offers similar performance to state-of-the-art commercial optimizers, and in some cases even surpass them.
Ryan Marcus, Parimarjan Negi, Hongzi Mao, Chi Zhang 0068, Mohammad Alizadeh, Tim Kraska, Olga Papaemmanouil, Nesime Tatbul
Proc. VLDB Endow.7
2019 Plan-Structured Deep Neural Network Models for Query Performance Prediction
abstract
Query performance prediction, the task of predicting a query's latency prior to execution, is a challenging problem in database management systems. Existing approaches rely on features and performance models engineered by human experts, but often fail to capture the complex interactions between query operators and input relations, and generally do not adapt naturally to workload characteristics and patterns in query execution plans. In this paper, we argue that deep learning can be applied to the query performance prediction problem, and we introduce a novel neural network architecture for the task: a plan-structured neural network . Our neural network architecture matches the structure of any optimizer-selected query execution plan and predict its latency with high accuracy, while eliminating the need for human-crafted input features. A number of optimizations are also proposed to reduce training overhead without sacrificing effectiveness. We evaluated our techniques on various workloads and we demonstrate that our approach can out-perform the state-of-the-art in query performance prediction.
Ryan Marcus, Olga Papaemmanouil
Proc. VLDB Endow.2
2019 NashDB: Fragmentation, Replication, and Provisioning using Economic Methods
abstract
Modern elastic computing systems allow applications to scale up and down automatically, increasing capacity for workload spikes and ensuring cost savings during lulls in activity. Adapting database management systems to work on top of such elastic infrastructure is not a trivial task, and requires a deep understanding of the sophisticated interplay between data fragmentation, replica allocation, and cluster provisioning. This demonstration showcases NashDB, an end-to-end method for addressing these concerns in an automatic way. NashDB relies on economic models to maximize query performance while staying within a user's budget. This demonstration will (1) allow audience members to see how NashDB handles shifting workloads in an adaptive way, and (2) allow audience members to test NashDB themselves by constructing synthetic workloads and seeing how NashDB adapts a cluster to them in real time.
Ryan Marcus, Chi Zhang 0068, Geoffrey Kao, Olga Papaemmanouil
Proc. VLDB Endow.5
2018 NashDB: An End-to-End Economic Method for Elastic Database Fragmentation, Replication, and Provisioning
abstract
Distributed data management systems often operate on "elastic'' clusters that can scale up or down on demand. These systems face numerous challenges, including data fragmentation, replication, and cluster sizing. Unfortunately, these challenges have traditionally been treated independently, leaving administrators with little insight on how the interplay of these decisions affects query performance. This paper introduces NashDB, an adaptive data distribution framework that relies on an economic model to automatically balance the supply and demand of data fragments, replicas, and cluster nodes. NashDB adapts its decisions to query priorities and shifting workloads, while avoiding underutilized cluster nodes and redundant replicas. This paper introduces and evaluates NashDB's model, as well as a suite of optimization techniques designed to efficiently identify data distribution schemes that match workload demands and transition the system to this new scheme with minimum data transfer overhead. Experimentally, we show that NashDB is often Pareto dominant compared to other solutions.
Ryan Marcus, Olga Papaemmanouil, Sofiya Semenova, Solomon Garber
SIGMOD Conference2
2017 Releasing Cloud Databases for the Chains of Performance Prediction Models
Ryan Marcus, Olga Papaemmanouil
CIDR2
2017 A Learning-Based Service for Cost and Performance Management of Cloud Databases
abstract
Data management applications deployed on IaaS cloud environments must simultaneously strive to minimize cost and provide good performance. Balancing these two goals requires complex decision-making across a number of axes: resource provisioning, query placement, and query scheduling. While previous works have addressed each axis in isolation for specific types of performance goals, this demonstration showcases WiSeDB, a cloud workload management advisor service that uses machine learning techniques to address all dimensions of the problem for customizable performance goals. In our demonstration, attendees will see WiSeDB in action for a variety of workloads and performance goals.
Ryan Marcus, Sofiya Semenova, Olga Papaemmanouil
ICDE3
2016 OptMark: A Toolkit for Benchmarking Query Optimizers
abstract
Query optimizers have long been considered as among the most complex components of a database engine, while the assessment of an optimizer's quality remains a challenging task. Indeed, existing performance benchmarks for database engines (like TPC benchmarks) produce a performance assessment of the query runtime system rather than its query optimizer. To address this challenge, this paper introduces OptMark, a toolkit for evaluating the quality of a query optimizer. OptMark is designed to offer a number of desirable properties. First, it decouples the quality of an optimizer from the quality of its underlying execution engine. Second it evaluates independently both the effectiveness of an optimizer (i.e., quality of the chosen plans) and its efficiency (i.e., optimization time). OptMark includes also a generic benchmarking toolkit that is minimum invasive to the DBMS that wishes to use it. Any DBMS can provide a system-specific implementation of a simple API that allows OptMark to run and generate benchmark scores for the specific system. This paper discusses the metrics we propose for evaluating an optimizer's quality, the benchmark's design and the toolkit's API and functionality. We have implemented OptMark on the open-source MySQL engine as well as two commercial database systems. Using these implementations we are able to assess the quality of the optimizers on these three systems based on the TPC-DS benchmark queries.
Olga Papaemmanouil, Mitch Cherniack
CIKM2
2016 WiSeDB: A Learning-based Workload Management Advisor for Cloud Databases
abstract
Workload management for cloud databases deals with the tasks of resource provisioning, query placement, and query scheduling in a manner that meets the application's performance goals while minimizing the cost of using cloud resources. Existing solutions have approached these three challenges in isolation while aiming to optimize a single performance metric. In this paper, we introduce WiSeDB, a learning-based framework for generating holistic workload management solutions customized to application-defined performance goals and workload characteristics. Our approach relies on supervised learning to train cost-effective decision tree models for guiding query placement, scheduling, and resource provisioning decisions. Applications can use these models for both batch and online scheduling of incoming workloads. A unique feature of our system is that it can adapt its offline model to stricter/looser performance goals with minimal re-training. This allows us to present to the application alternative workload management strategies that address the typical performance vs. cost trade-off of cloud services. Experimental results show that our approach has very low training overhead while offering low cost strategies for a variety of performance metrics and workload characteristics.
Ryan Marcus, Olga Papaemmanouil
Proc. VLDB Endow.2
2016 AIDE: An Active Learning-Based Approach for Interactive Data Exploration
abstract
In this paper, we argue that database systems be augmented with an automated data exploration service that methodically steers users through the data in a meaningful way. Such an automated system is crucial for deriving insights from complex datasets found in many big data applications such as scientific and healthcare applications as well as for reducing the human effort of data exploration. Towards this end, we present AIDE, an Automatic Interactive Data Exploration framework that assists users in discovering new interesting data patterns and eliminate expensive ad-hoc exploratory queries. AIDE relies on a seamless integration of classification algorithms and data management optimization techniques that collectively strive to accurately learn the user interests based on his relevance feedback on strategically collected samples. We present a number of exploration techniques as well as optimizations that minimize the number of samples presented to the user while offering interactive performance. AIDE can deliver highly accurate query predictions for very common conjunctive queries with small user effort while, given a reasonable number of samples, it can predict with high accuracy complex disjunctive queries. It provides interactive performance as it limits the user wait time per iteration of exploration to less than a few seconds.
Kyriaki Dimitriadou, Olga Papaemmanouil, Yanlei Diao
IEEE Trans. Knowl. Data Eng.2
2015 XCloud: Extensible Performance Management for Cloud Data Services
Olga Papaemmanouil
CIDR1
2015 Skew-Aware Join Optimization for Array Databases
abstract
Science applications are accumulating an ever-increasing amount of multidimensional data. Although some of it can be processed in a relational database, much of it is better suited to array-based engines. As such, it is important to optimize the query processing of these systems. This paper focuses on efficient query processing of join operations within an array database. These engines invariably ``chunk'' their data into multidimensional tiles that they use to efficiently process spatial queries. As such, traditional relational algorithms need to be substantially modified to take advantage of array tiles. Moreover, most n-dimensional science data is unevenly distributed in array space because its underlying observations rarely follow a uniform pattern. It is crucial that the optimization of array joins be skew-aware. In addition, owing to the scale of science applications, their query processing usually spans multiple nodes. This further complicates the planning of array joins.
Jennie Rogers, Olga Papaemmanouil, Leilani Battle, Michael Stonebraker
SIGMOD Conference2
2015 Overview of Data Exploration Techniques
abstract
Data exploration is about efficiently extracting knowledge from data even if we do not know exactly what we are looking for. In this tutorial, we survey recent developments in the emerging area of database systems tailored for data exploration. We discuss new ideas on how to store and access data as well as new ideas on how to interact with a data system to enable users and applications to quickly figure out which data parts are of interest. In addition, we discuss how to exploit lessons-learned from past research, the new challenges data exploration crafts, emerging applications and future research directions.
Stratos Idreos, Olga Papaemmanouil, Surajit Chaudhuri
SIGMOD Conference2
2015 AIDE: An Automatic User Navigation System for Interactive Data Exploration
abstract
Data analysts often engage in data exploration tasks to discover interesting data patterns, without knowing exactly what they are looking for. Such exploration tasks can be very labor-intensive because they often require the user to review many results of ad-hoc queries and adjust the predicates of subsequent queries to balance the tradeoff between collecting all interesting information and reducing the size of returned data. In this demonstration we introduce AIDE , a system that automates these exploration tasks. AIDE steers the user towards interesting data areas based on her relevance feedback on database samples, aiming to achieve the goal of identifying all database objects that match the user interest with high efficiency. In our demonstration, conference attendees will see AIDE in action for a variety of exploration tasks on real-world datasets.
Yanlei Diao, Kyriaki Dimitriadou, Wenzhao Liu, Olga Papaemmanouil, Kemi Peng, Liping Peng
Proc. VLDB Endow.5
2014 Contender: A Resource Modeling Approach for Concurrent Query Performance Prediction
abstract
Predicting query performance under concurrency is a difficult task that has many applications in capacity planning, cloud computing, and batch scheduling. We introduce Contender, a new resource-modeling approach for predicting the concurrent query perfor-mance of analytical workloads. Contender’s unique feature is that it can generate effective predictions for both static as well as ad-hoc or dynamic workloads with low training requirements. These characteristics make Contender a practical solution for real-world deployment. Contender relies on models of hardware resource contention to predict concurrent query performance. It introduces two key met-rics, Concurrent Query Intensity (CQI) and Query Sensitivity (QS), to characterize the impact of resource contention on query interac-tions. CQI models how aggressively concurrent queries will use the shared resources. QS defines how a query’s performance changes as a function of the scarcity of resources. Contender integrates these two metrics to effectively estimate a query’s concurrent exe-cution latency using only linear time sampling of the query mixes. Contender learns from sample query executions (based on known query templates) and uses query plan characteristics to gen-erate latency estimates for previously unseen templates. Our ex-perimental results, obtained from PostgreSQL/TPC-DS, show that Contender’s predictions have an error of 19 % for known templates and 25 % for new templates, which is competitive with the state-of-the-art while requiring considerably less training time. 1.
Jennie Rogers, Olga Papaemmanouil, Ugur Çetintemel, Eli Upfal
EDBT2
2014 Devel-op: An optimizer development environment
abstract
Recent advances in the underlying architectures of database management systems (DBMS) have motivated the redesign of key DBMS components such as the query optimizer. Optimizers are inherently difficult to build and maintain, and yet there exists no software engineering tools to facilitate their development. In this paper, we introduce a [Devel]opment Environment for Query [Op]timizers (Devel-Op) designed to facilitate the rapid prototyping, profiling and benchmarking of optimizers. Our current version of the tool permits declarative specification and generation of two key optimizer components (the logical plan enumerator and physical plan generator) as well as debugging and visualization tools for profiling generated components.
Zhibo Peng, Mitch Cherniack, Olga Papaemmanouil
ICDE3
2014 Explore-by-example: an automatic query steering framework for interactive data exploration
abstract
Interactive Data Exploration (IDE) is a key ingredient of a diverse set of discovery-oriented applications, including ones from scientific computing and evidence-based medicine. In these applications, data discovery is a highly ad hoc interactive process where users execute numerous exploration queries using varying predicates aiming to balance the trade-off between collecting all relevant information and reducing the size of returned data. Therefore, there is a strong need to support these human-in-the-loop applications by assisting their navigation in the data to find interesting objects.
Kyriaki Dimitriadou, Olga Papaemmanouil, Yanlei Diao
SIGMOD Conference2
2013 Query Steering for Interactive Data Exploration
Ugur Çetintemel, Mitch Cherniack, Justin A. DeBrabant, Yanlei Diao, Kyriaki Dimitriadou, Alexander Kalinin 0001, Olga Papaemmanouil, Stanley B. Zdonik
CIDR7
2011 Performance prediction for concurrent database workloads
abstract
Current trends in data management systems, such as cloud and multi-tenant databases, are leading to data processing environments that concurrently execute heterogeneous query workloads. At the same time, these systems need to satisfy diverse performance expectations. In these newly-emerging settings, avoiding potential Quality-of-Service (QoS) violations heavily relies on performance predictability, i.e., the ability to estimate the impact of concurrent query execution on the performance of individual queries in a continuously evolving workload.
Jennie Rogers, Ugur Çetintemel, Olga Papaemmanouil, Eli Upfal
SIGMOD Conference3
2009 Supporting Generic Cost Models for Wide-Area Stream Processing
abstract
Existing stream processing systems are optimized for a specific metric, which may limit their applicability to diverse applications and environments. This paper presents XFlow, a generic data stream collection, processing, and dissemination system that addresses this limitation efficiently. XFlow can express and optimize a variety of optimization metrics and constraints by distributing stream processing queries across a wide-area network. It uses metric-independent decentralized algorithms that work on localized, aggregated statistics, while avoiding local optima. To facilitate light-weight dynamic changes on the query deployment, XFlow relies on a loosely-coupled, flexible architecture consisting of multiple publish-subscribe overlay trees that can gracefully scale and adapt to changes to network and workload conditions. Based on the desired performance goals, the system progressively refines the query deployment, the structure of the overlay trees, as well as the statistics collection process. We provide an overview of XFlow's architecture and discuss its decentralized optimization model. We demonstrate its flexibility and the effectiveness using real-world streams and experimental results obtained from XFlow's deployment on PlanetLab. The experiments reveal that XFlow can effectively optimize various performance metrics in the presence of varying network and workload conditions.
Olga Papaemmanouil, Ugur Çetintemel, John Jannotti
ICDE1
2008 Simultaneous Equation Systems for Query Processing on Continuous-Time Data Streams
abstract
We introduce pulse, a framework for processing continuous queries over models of continuous-time data, which can compactly and accurately represent many real-world activities and processes. Pulse implements several query operators, including filters, aggregates and joins, that work by solving simultaneous equation systems, which in many cases is significantly cheaper than processing a stream of tuples. As such, pulse translates regular queries to work on continuous-time inputs, to reduce computational overhead and latency while meeting user-specified error bounds on query results. For error bound checking, pulse uses an approximate query inversion technique that ensures the solver executes infrequently and only in the presence of errors, or no previously known results. We first discuss the high-level design of pulse, which we fully implemented in a stream processing system. We then characterise pulse's behavior through experiments with real data, including financial data from the New York Stock Exchange, and spatial data from the Automatic Identification System for tracking naval vessels. Our results verify that Pulse is practical and demonstrates significant performance gains for a variety of workload and query types.
Yanif Ahmad, Olga Papaemmanouil, Ugur Çetintemel, Jennie Rogers
ICDE2
2006 Extensible optimization in overlay dissemination trees
abstract
We introduce XPORT, a profile-driven distributed data dissemination system that supports an extensible set of data types, profile types, and optimization metrics. XPORT efficiently implements a generic tree-based overlay network, which can be customized per application using a small number of methods that encapsulate application-specific data filtering, profile aggregation, and optimization logic. The clean separation between the "plumbing" and "application" enables the system to uniformly support disparate dissemination-based applications.We first provide an overview of the basic XPORT model and architecture. We then describe in detail an extensible optimization framework, based on a two-level aggregation model, that facilitates easy specification of a wide range of commonly used performance goals. We discuss distributed tree transformation protocols that allow XPORT to iteratively optimize its operation to achieve these goals under changing network and application conditions. Finally, we demonstrate the flexibility and the effectiveness of XPORT using real-world data and experimental results obtained from both prototype-based LAN emulation and deployment on PlanetLab.
Olga Papaemmanouil, Yanif Ahmad, Ugur Çetintemel, John Jannotti, Yenel Yildirim
SIGMOD Conference1
2006 XPORT: extensible profile-driven overlay routing trees
abstract
XPORT is a profile-driven distributed data collection and dissemination system that supports an extensible set of data types, profiles, and optimization metrics. XPORT efficiently builds a generic tree-based overlay network, which can be customized per application using a small number of methods that encapsulate application-specific data-profile matching, aggregation, and optimization logic. The clean separation between the "plumbing" and "application" enables XPORT to uniformly and easily support disparate dissemination-based applications such as content-based feed dissemination and application-level multicast. We propose to demonstrate the basic XPORT system, featuring its extensible optimization framework that facilitates easy specification of a wide range of useful performance goals and a continuous, adaptive optimization model to achieve these goals under changing network and application conditions. We will use two different underlying applications, an RSS feed dissemination application and a multiplayer network game, along with visual system-monitoring tools to illustrate the extensibility and the operational aspects of XPORT.
Olga Papaemmanouil, Yanif Ahmad, Ugur Çetintemel, John Jannotti, Yenel Yildirim
SIGMOD Conference1
2005 SemCast: Semantic Multicast for Content-based Data Dissemination
abstract
We address the problem of content-based dissemination of highly-distributed, high-volume data streams for stream-based monitoring applications and large-scale data delivery. Existing content-based dissemination approaches commonly rely on distributed filtering trees that require filtering at all brokers on the tree. We present a new semantic multicast approach that eliminates the need for content-based filtering at interior brokers and facilitates fine-grained control over the construction of efficient dissemination trees. The central idea is to split the incoming data streams (based on their contents, rates, and destinations) and then spread the pieces across multiple channels, each of which is implemented as an independent dissemination tree. We present the basic design and evaluation of SemCast, an overlay-network based system that implements this semantic multicast approach. Through a detailed simulation study and realistic network topologies, we demonstrate that SemCast significantly improves the efficiency of dissemination compared to traditional approaches.
Olga Papaemmanouil, Ugur Çetintemel
ICDE1
2005 Distributed operation in the Borealis stream processing engine
abstract
Borealis is a distributed stream processing engine that is being developed at Brandeis University, Brown University, and MIT. Borealis inherits core stream processing functionality from Aurora and inter-node communication functionality from Medusa.We propose to demonstrate some of the key aspects of distributed operation in Borealis, using a multi-player network game as the underlying application. The demonstration will illustrate the dynamic resource management, query optimization and high availability mechanisms employed by Borealis, using visual performance-monitoring tools as well as the gaming experience.
Yanif Ahmad, Bradley Berg, Ugur Çetintemel, Mark Humphrey, Jeong-Hyon Hwang, Anjali Jhingran, Anurag Maskey, Olga Papaemmanouil, Alexander Rasin, Nesime Tatbul, Wenjuan Xing, Stanley B. Zdonik
SIGMOD Conference8
2004 Semantic Multicast for Content-based Stream Dissemination
abstract
We consider the problem of content-based routing and dissemination of highly-distributed, fast data streams from multiple sources to multiple receivers. Our target application domain includes real-time, stream-based monitoring applications and large-scale event dissemination. We introduce SemCast, a new semantic multicast approach that, unlike previous approaches, eliminates the need for content-based forwarding at interior brokers and facilitates fine-grained control over the construction of dissemination overlays. We present the initial design of SemCast and provide an outline of the architectural and algorithmic challenges as well as our initial solutions. Preliminary experimental results show that SemCast can significantly reduce overall bandwidth requirements compared to traditional event-dissemination approaches.
Olga Papaemmanouil, Ugur Çetintemel
WebDB1