Garret Swart

dblp:46/4288 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 1 first-authorSystems, architecture and hardware · 3Security and privacy · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Query processing and optimization · 47% Distributed and cloud data management · 34% Data integration and cleaning · 9%
Computer architecture, parallel and distributed computing, and storage systems
7 papers
Parallel and multicore computing · 24% Hardware accelerators and domain-specific architectures · 23% Processor architecture and microarchitecture · 17%

Topics — the 15 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed and cloud data management
distributed query processing
0.212014
Changing Engines in Midstream: A Java Stream Computational Model for Big Data Processing · Proc. VLDB Endow. 2014
Distributed and cloud data management › federated database
federated query processing
0.212014
Changing Engines in Midstream: A Java Stream Computational Model for Big Data Processing · Proc. VLDB Endow. 2014
Data integration and cleaning › interoperability › database interoperability
database management system integration
0.112012
Oracle in-database hadoop: when mapreduce meets RDBMS · SIGMOD Conference 2012
Distributed and cloud data management › mapreduce
mapreduce and parallel database integration
0.112012
Oracle in-database hadoop: when mapreduce meets RDBMS · SIGMOD Conference 2012
Query processing and optimization › query execution › scan processing
sequential scan
0.122008
Constant-Time Query Processing · ICDE 2008
Row-wise parallel predicate evaluation · Proc. VLDB Endow. 2008
Indexing and storage engines
data compression
0.122008
How to barter bits for chronons: compression and bandwidth trade offs for database scans · SIGMOD Conference 2007
Constant-Time Query Processing · ICDE 2008
Query processing and optimization
aggregation
0.112008
Constant-Time Query Processing · ICDE 2008
Query processing and optimization › query execution › expression evaluation
predicate evaluation
0.112008
Row-wise parallel predicate evaluation · Proc. VLDB Endow. 2008
Query processing and optimization
query execution
0.112008
Constant-Time Query Processing · ICDE 2008
Processor architecture and microarchitecture
SIMD
0.112008
Row-wise parallel predicate evaluation · Proc. VLDB Endow. 2008
Storage systems
data compression
0.112006
How to Wring a Table Dry: Entropy Compression of Relations and Querying of Compressed Relations · VLDB 2006
Distributed and cloud data management
mapreduce
0.012012
Oracle in-database hadoop: when mapreduce meets RDBMS · SIGMOD Conference 2012
Parallel and multicore computing › parallel architecture
massively parallel processing
0.012007
How to barter bits for chronons: compression and bandwidth trade offs for database scans · SIGMOD Conference 2007
Memory systems
cache coherence
0.011994
A Coherent Distributed File Cache with Directory Write-Behind · ACM Trans. Comput. Syst. 1994
Storage systems › file systems
distributed file system
0.011994
A Coherent Distributed File Cache with Directory Write-Behind · ACM Trans. Comput. Syst. 1994

Methods — techniques the papers use, named apart from their topics

lambda expressions · 0.4Java Stream API · 0.4bit-packing · 0.2bank layout · 0.2vertical partitioning · 0.1compression · 0.1adaptivity evaluation · 0.1hash-based aggregation · 0.1SIMD predicate evaluation · 0.1cache coherence protocol design · 0.0
YearPublicationVenuePosition
2020 Understanding and Improving Persistent Transactions on Optane™ DC Memory
abstract
Storing data structures in high-capacity byte-addressable persistent memory instead of DRAM or a storage device offers the opportunity to (1) reduce cost and power consumption compared with DRAM, (2) decrease the latency and CPU resources needed for an I/O operation compared with storage, and (3) allow for fast recovery as the data structure remains in memory after a machine failure. The first commercial offering in this space is Intel® Optane™ Direct Connect (Optane™ DC) Persistent Memory. Optane™ DC promises access time within a constant factor of DRAM, with larger capacity, lower energy consumption, and persistence. We present an experimental evaluation of persistent transactional memory performance, and explore how Optane™ DC durability domains affect the overall results. Given that neither of the two available durability domains can deliver performance competitive with DRAM, we introduce and emulate a new durability domain, called PDRAM, in which the memory controller tracks enough information (and has enough reserve power) to make DRAM behave like a persistent cache of Optane™ DC memory.In this paper we compare the performance of these durability domains on several configurations of five persistent transactional memory applications. We find a large throughput difference, which emphasizes the importance of choosing the best durability domain for each application and system. At the same time, our results confirm that recently published persistent transactional memory algorithms are able to scale, and that recent optimizations for these algorithms lead to strong performance, with speedups as high as 6× at 16 threads.
Pantea Zardoshti, Michael F. Spear, Aida Vosoughi, Garret Swart
IPDPS4
2019 A Morsel-Driven Query Execution Engine for Heterogeneous Multi-Cores
abstract
Currently, we face the next major shift in processor designs that arose from the physical limitations known as the "dark silicon effect". Due to thermal limitations and shrinking transistor sizes, multi-core scaling is coming to an end. A major new direction that hardware vendors are currently investigating involves specialized and energy-efficient hardware accelerators (e.g., ASICs) placed on the same die as the normal CPU cores. In this paper, we present a novel query processing engine called SiliconDB that targets such heterogeneous processor environments. We leverage the Sparc M7 platform to develop and test our ideas. Based on the SSB benchmarks, as well as other micro benchmarks, we compare the efficiency of SiliconDB with existing execution strategies that make use of co-processors (e.g., FPGAs, GPUs) and demonstrate speed-up improvements of up to 2x.
Kayhan Dursun, Carsten Binnig, Ugur Çetintemel, Garret Swart, Weiwei Gong
Proc. VLDB Endow.4
2014 Changing Engines in Midstream: A Java Stream Computational Model for Big Data Processing
abstract
With the addition of lambda expressions and the Stream API in Java 8, Java has gained a powerful and expressive query language that operates over in-memory collections of Java objects, making the transformation and analysis of data more convenient, scalable and efficient. In this paper, we build on Java 8 Stream and add a DistributableStream abstraction that supports federated query execution over an extensible set of distributed compute engines. Each query eventually results in the creation of a materialized result that is returned either as a local object or as an engine defined distributed Java Collection that can be saved and/or used as a source for future queries. Distinctively, DistributableStream supports the changing of compute engines both between and within a query, allowing different parts of a computation to be executed on different platforms. At execution time, the query is organized as a sequence of pipelined stages, each stage potentially running on a different engine. Each node that is part of a stage executes its portion of the computation on the data available locally or produced by the previous stage of the computation. This approach allows for computations to be assigned to engines based on pricing, data locality, and resource availability. Coupled with the inherent laziness of stream operations, this brings great flexibility to query planning and separates the semantics of the query from the details of the engine used to execute it. We currently support three engines, Local, Apache Hadoop MapReduce and Oracle Coherence, and we illustrate how new engines and data sources can be added.
Xueyuan Su, Garret Swart, Brian Goetz, Brian Oliver, Paul Sandoz
Proc. VLDB Endow.2
2012 Balancing reducer skew in MapReduce workloads using progressive sampling
abstract
The elapsed time of a parallel job depends on the completion time of its longest running constituent. We present a static load balancing algorithm that distributes work evenly across the reducers in a MapReduce job resulting in significant elapsed time reductions.
Smriti R. Ramakrishnan, Garret Swart, Aleksey Urmanov
SoCC2
2012 Oracle in-database hadoop: when mapreduce meets RDBMS
abstract
Big data is the tar sands of the data world: vast reserves of raw gritty data whose valuable information content can only be extracted at great cost. MapReduce is a popular parallel programming paradigm well suited to the programmatic extraction and analysis of information from these unstructured Big Data reserves. The Apache Hadoop implementation of MapReduce has become an important player in this market due to its ability to exploit large networks of inexpensive servers. The increasing importance of unstructured data has led to the interest in MapReduce and its Apache Hadoop implementation, which has led to the interest of data processing vendors in supporting this programming style.
Xueyuan Su, Garret Swart
SIGMOD Conference2
2009 Configuring storage-area networks using mandatory security
abstract
Storage-area networks are a popular and efficient way of building large storage systems both in an enterprise environment and for multi-domain storage service providers. In both environments the network and the storage has to be configured to ensure that the data is maintained securely and can be delivered efficiently. In this paper, we describe a model of mandatory security for SAN services that incorporates the notion of risk as a measure of the robustness of the SAN's configuration and that formally defines a vulnerability common in systems with mandatory security, i.e. cascaded threats. Our abstract SAN model is flexible enough to reflect the data requirements, tractable for the administrator, and can be implemented as part of an automatic configuration system. The implementation is given as part of a prototype written in OPL.
Benjamin Aziz, Simon N. Foley, John Herbert, Garret Swart
J. Comput. Secur.4
2009 Autonomic query parallelization using non-dedicated computers: an evaluation of adaptivity options
Norman W. Paton, Jorge Buenabad Chávez, Mengsong Chen, Vijayshankar Raman, Garret Swart, Inderpal Narang, Daniel M. Yellin, Alvaro A. A. Fernandes
VLDB J.5
2008 Constant-Time Query Processing
abstract
Query performance in current systems depends significantly on tuning: how well the query matches the available indexes, materialized views etc. Even in a well tuned system, there are always some queries that take much longer than others. This frustrates users who increasingly want consistent response times to ad hoc queries. We argue that query processors should instead aim for constant response times for all queries, with no assumption about tuning. We present Blink, our first attempt at this goal, that runs every query as a table scan over a fully denormalized database, with hash group-by done along the way. To make this scan efficient, Blink uses a novel compression scheme that horizontally partitions tuples by frequency, thereby compressing skewed data almost down to entropy, even while producing long runs of fixed-length, easily-parseable values. We also present a scheme for evaluating a conjunction of range and equality predicates in SIMD fashion over compressed tuples, and different schemes for efficient hash-based aggregation within the L2 cache. A experimental study with a suite of arbitrary single block SQL queries over a TPCH-like schema suggests that constant-time queries can be efficient.
Vijayshankar Raman, Garret Swart, Lin Qiao 0001, Frederick Reiss 0001, Vijay Dialani, Donald Kossmann, Inderpal Narang, Richard Sidle
ICDE2
2008 Row-wise parallel predicate evaluation
abstract
Table scans have become more interesting recently due to greater use of ad-hoc queries and greater availability of multi-core, vector-enabled hardware. Table scan performance is limited by value representation, table layout, and processing techniques. In this paper we propose a new layout and processing technique for efficient one-pass predicate evaluation. Starting with a set of rows with a fixed number of bits per column, we append columns to form a set of banks and then pad each bank to a supported machine word length, typically 16, 32, or 64 bits. We then evaluate partial predicates on the columns of each bank, using a novel evaluation strategy that evaluates column level equality, range tests, IN-list predicates, and conjuncts of these predicates, simultaneously on multiple columns within a bank, and on multiple rows within a machine register. This approach outperforms pure column stores, which must evaluate the partial predicates one column at a time. We evaluate and compare the performance and representation overhead of this new approach and several proposed alternatives.
Ryan Johnson 0001, Vijayshankar Raman, Richard Sidle, Garret Swart
Proc. VLDB Endow.4
2007 Impliance: A Next Generation Information Management Appliance
Bishwaranjan Bhattacharjee, Joseph S. Glider, Richard A. Golding, Guy M. Lohman, Volker Markl, Hamid Pirahesh, Jun Rao, Robert M. Rees, Garret Swart
CIDR9
2007 How to barter bits for chronons: compression and bandwidth trade offs for database scans
abstract
Two trends are converging to make the CPU cost of a table scan a more important component of database performance. First, table scans are becoming a larger fraction of the query processing workload, and second, large memories and compression are making table scans CPU, rather than disk bandwidth, bound. Data warehouse systems have found that they can avoid the unpredictability of joins and indexing and achieve good performance by using massive parallel processing to perform scans over compressed vertical partitions of a denormalized schema.
Allison L. Holloway, Vijayshankar Raman, Garret Swart, David J. DeWitt
SIGMOD Conference3
2006 How to Wring a Table Dry: Entropy Compression of Relations and Querying of Compressed Relations
Vijayshankar Raman, Garret Swart
VLDB2
2005 Trading Off Security in a Service Oriented Architecture
Garret Swart, Benjamin Aziz, Simon N. Foley, John Herbert
DBSec1
2004 Configuring Storage Area Networks for Mandatory Security
abstract
Storage-area networks are a popular and efficient way of building large storage systems both in an enterprise environment and for multi-domain storage service providers. In both environments the network and the storage has to be configured to ensure that the data is maintained securely and can be delivered efficiently. In this paper we describe a model of mandatory security for multi-domain storage services that is flexible enough to reflect the data requirements, tractable for the administrator, and implementable as part of an automatic configuration system. We describe the model abstractly, its implementation as part of a prototype SAN configuration system written in OPL, and illustrate its operation on a set of sample configurations. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Benjamin Aziz, Simon N. Foley, John Herbert, Garret Swart
DBSec4
2003 WS-Workspace: Workspace Versioning for Web Services
Garret Swart
ICSOC1
2003 Collaboration and Undo: The Web Workspace Paradigm
abstract
An informal look at some of the most annoying problems confronting users who attempt to manipulate data stored on the Web and a new update paradigm that handles many of these problems. This new approach, called Web workspaces, allows for collaboration, undo, 'what if' analysis, auditing and a form of long running transaction. The new approach builds on workspace versioning techniques long in use in CASE and CAD systems and on transaction coordination techniques, long used to turn disparate applications into cohesive systems. The ideas can be applied to Web services or to enterprise frameworks to build more robust and usable applications for manipulating data on the Web.
Garret Swart
WISE1
1994 A Coherent Distributed File Cache with Directory Write-Behind
abstract
Extensive caching is a key feature of the Echo distributed file system. Echo client machines maintain coherent caches of file and directory data and properties, with write-behind (delayed write-back) ofallcached information. Echo specifies ordering constraints on this write-behind, enabling applications to store and maintain consistent data structures in the file system even when crashes or network faults prevent some writes from being completed. In this paper we describe the Echo cache's coherence and ordering semantics, show how they can improve the performance and consistency of applications, explain how they are implemented. We also discuss the general problem of reliably notifying applications and users when write-behind is lost; we addressed this problem as part of the Echo design, but did not find a fully satisfactory solution.
Timothy P. Mann, Andrew Birrell, Andy Hisgen, Charles Jerian, Garret Swart
ACM Trans. Comput. Syst.5