Ben Sowell

dblp:s/BenSowell · also Benjamin Sowell · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 3 first-author · 2 since 2021Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Storage systems · 24% Cloud and datacenter computing · 22% Performance modeling and evaluation · 19%
Databases, data mining, and information retrieval
7 papers
Data integration and cleaning · 45% Spatial and temporal data management · 17% Transaction processing and concurrency control · 14%

Topics — the 26 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
extract-transform-load
0.712023
The Story of AWS Glue · Proc. VLDB Endow. 2023
Storage systems
durability and recovery
0.222011
Fast checkpoint recovery algorithms for frequently consistent applications · SIGMOD Conference 2011
BRRL: a recovery library for main-memory applications in the cloud · SIGMOD Conference 2011
Data integration and cleaning
metadata management
0.212023
The Story of AWS Glue · Proc. VLDB Endow. 2023
Cloud and datacenter computing › serverless computing
serverless analytics
0.212023
The Story of AWS Glue · Proc. VLDB Endow. 2023
Cloud and datacenter computing
serverless computing
0.212023
The Story of AWS Glue · Proc. VLDB Endow. 2023
Spatial and temporal data management › spatial query processing › spatial join
in-memory spatial join
0.212013
An Experimental Analysis of Iterated Spatial Joins in Main Memory · Proc. VLDB Endow. 2013
Spatial and temporal data management › spatial query processing
spatial join
0.212013
An Experimental Analysis of Iterated Spatial Joins in Main Memory · Proc. VLDB Endow. 2013
Performance modeling and evaluation
benchmarking
0.212013
An Experimental Analysis of Iterated Spatial Joins in Main Memory · Proc. VLDB Endow. 2013
Database system architecture and tuning
main-memory database
0.112012
Minuet: A Scalable Distributed Multiversion B-Tree · Proc. VLDB Endow. 2012
Transaction processing and concurrency control › concurrency control
multiversion concurrency control
0.112012
Minuet: A Scalable Distributed Multiversion B-Tree · Proc. VLDB Endow. 2012
Transaction processing and concurrency control › OLTP
in-memory transaction processing
0.112011
Fast checkpoint recovery algorithms for frequently consistent applications · SIGMOD Conference 2011
Distributed systems › fault tolerance
checkpointing
0.112011
BRRL: a recovery library for main-memory applications in the cloud · SIGMOD Conference 2011
Distributed systems
fault tolerance
0.112011
BRRL: a recovery library for main-memory applications in the cloud · SIGMOD Conference 2011
Performance modeling and evaluation › simulation
agent-based simulation
0.112010
Behavioral Simulations in MapReduce · Proc. VLDB Endow. 2010
Parallel and multicore computing › data-parallel programming
mapreduce
0.112010
Behavioral Simulations in MapReduce · Proc. VLDB Endow. 2010
Parallel and multicore computing
parallel programming models and runtimes
0.112010
Behavioral Simulations in MapReduce · Proc. VLDB Endow. 2010
Performance modeling and evaluation › simulation › parallel and distributed simulation
parallel simulation
0.112010
Behavioral Simulations in MapReduce · Proc. VLDB Endow. 2010
Hardware reliability and fault tolerance › error recovery
checkpoint recovery
0.112009
An Evaluation of Checkpoint Recovery for Massively Multiplayer Online Games · Proc. VLDB Endow. 2009
Storage systems
database recovery
0.112009
An Evaluation of Checkpoint Recovery for Massively Multiplayer Online Games · Proc. VLDB Endow. 2009
Storage systems › storage reliability
durability
0.112009
An Evaluation of Checkpoint Recovery for Massively Multiplayer Online Games · Proc. VLDB Endow. 2009
Memory systems
main memory database
0.112009
An Evaluation of Checkpoint Recovery for Massively Multiplayer Online Games · Proc. VLDB Endow. 2009
Distributed and cloud data management
distributed data store
0.012012
Minuet: A Scalable Distributed Multiversion B-Tree · Proc. VLDB Endow. 2012
Cloud and datacenter computing
cloud applications
0.012011
BRRL: a recovery library for main-memory applications in the cloud · SIGMOD Conference 2011
Programming languages and type systems
domain-specific languages
0.012010
Behavioral Simulations in MapReduce · Proc. VLDB Endow. 2010
Programming languages and type systems › domain-specific languages
simulation language
0.012010
Behavioral Simulations in MapReduce · Proc. VLDB Endow. 2010
High-performance computing
large-scale simulation
0.012010
Behavioral Simulations in MapReduce · Proc. VLDB Endow. 2010

Methods — techniques the papers use, named apart from their topics

synchronous traversal · 0.3r-tree · 0.3index nested loop · 0.3lock-free checkpointing · 0.2frequent consistency points · 0.2spatial join · 0.2dataflow compilation · 0.2automatic parallelization · 0.2multiversion b-trees · 0.1copy-on-write · 0.1schema-aware checkpointing · 0.1data parallelism · 0.1trace-driven evaluation · 0.1simulation · 0.1declarative scripting language · 0.1
YearPublicationVenuePosition
2025 The Design of an LLM-powered Unstructured Analytics System
Eric Anderson 0003, Jonathan Fritz, Austin Lee, Bohou Li, Mark Lindblad, Henry Lindeman, Alex Meyer, Parth Parmar, Tanvi Ranade, Mehul A. Shah, Ben Sowell, Dan Tecuci, Vinayak Thapliyal, Matt Welsh
CIDR11
2023 The Story of AWS Glue
abstract
AWS Glue is Amazon's serverless data integration cloud service that makes it simple and cost effective to extract, clean, enrich, load, and organize data. Originally launched in August 2017, AWS Glue began as an extract-transform-load (ETL) service designed to relieve developers and data engineers of the undifferentiated heavy lifting needed to load databases, data warehouses, and build data lakes on Amazon S3. Since then, it has evolved to serve a larger audience including ETL specialists and data scientists, and includes a broader suite of data integration capabilities. Today, hundreds of thousands of customers use AWS Glue every month. In this paper, we describe the use cases and challenges cloud customers face in preparing data for analytics and the tenets we chose to drive Glue's design. We chose early on to focus on ease-of-use, scale, and extensibility. At its core, Glue offers serverless Apache Spark and Python engines backed by a purpose-built resource manager for fast startup and auto-scaling. In Spark, it offers a new data structure --- DynamicFrames --- for manipulating messy schema-free semi-structured data such as event logs, a variety of transformations and tooling to simplify data preparation, and a new shuffle plugin to offload to cloud storage. It also includes a Hivemetastore compatible Data Catalog with Glue crawlers to build and manage metadata, e.g. for data lakes on Amazon S3. Finally, Glue Studio is its visual interface for authoring Spark and Python-based ETL jobs. We describe the innovations that differentiate AWS Glue and drive its popularity and how it has evolved over the years.
Mohit Saxena, Ben Sowell, Daiyan Alamgir, Nitin Bahadur, Bijay Bisht, Santosh Chandrachood, Chitti Keswani, G2 Krishnamoorthy, Austin Lee, Bohou Li, Zach Mitchell, Vaibhav Porwal, Maheedhar Reddy Chappidi, Brian Ross, Noritaka Sekiyama, Omer Zaki, Linchi Zhang, Mehul A. Shah
Proc. VLDB Endow.2
2013 An Experimental Analysis of Iterated Spatial Joins in Main Memory
abstract
Many modern applications rely on high-performance processing of spatial data. Examples include location-based services, games, virtual worlds, and scientific simulations such as molecular dynamics and behavioral simulations. These applications deal with large numbers of moving objects that continuously sense their environment, and their data access can often be abstracted as a repeated spatial join. Updates to object positions are interspersed with these join operations, and batched for performance. Even for the most demanding scenarios, the data involved in these joins fits comfortably in the main memory of a cluster of machines, and most applications run completely in main memory for performance reasons. Choosing appropriate spatial join algorithms is challenging due to the large number of techniques in the literature. In this paper, we perform an extensive evaluation of repeated spatial join algorithms for distance (range) queries in main memory. Our study is unique in breadth when compared to previous work: We implement, tune, and compare ten distinct algorithms on several workloads drawn from the simulation and spatial indexing literature. We explore the design space of both index nested loops algorithms and specialized join algorithms, as well as the use of moving object indices that can be incrementally maintained. Surprisingly, we find that when queries and updates can be batched, repeatedly re-computing the join result from scratch outperforms using a moving object index in all but the most extreme cases. This suggests that--given the code complexity of index structures for moving objects -- specialized join strategies over simple index structures, such as Synchronous Traversal over R-Trees, should be the methods of choice for the above applications.
Ben Sowell, Marcos Antonio Vaz Salles, Tuan Cao, Alan J. Demers, Johannes Gehrke
Proc. VLDB Endow.1
2012 Minuet: A Scalable Distributed Multiversion B-Tree
abstract
Data management systems have traditionally been designed to support either long-running analytics queries or short-lived transactions, but an increasing number of applications need both. For example, online games, socio-mobile apps, and e-commerce sites need to not only maintain operational state, but also analyze that data quickly to make predictions and recommendations that improve user experience. In this paper, we present Minuet, a distributed, main-memory B-tree that supports both transactions and copy-on-write snapshots for in-situ analytics. Minuet uses main-memory storage to enable low-latency transactional operations as well as analytics queries without compromising transaction performance. In addition to supporting read-only analytics queries on snapshots, Minuet supports writable clones, so that users can create branching versions of the data. This feature can be quite useful, e.g. to support complex "what-if" analysis or to facilitate wide-area replication. Our experiments show that Minuet outperforms a commercial main-memory database in many ways. It scales to hundreds of cores and TBs of memory, and can process hundreds of thousands of B-tree operations per second while executing long-running scans.
Ben Sowell, Wojciech M. Golab, Mehul A. Shah
Proc. VLDB Endow.1
2011 BRRL: a recovery library for main-memory applications in the cloud
abstract
In this demonstration we present BRRL, a library for making distributed main-memory applications fault tolerant. BRRL is optimized for cloud applications with frequent points of consistency that use data-parallelism to avoid complex concurrency control mechanisms. BRRL differs from existing recovery libraries by providing a simple table abstraction and using schema information to optimize checkpointing. We will demonstrate the utility of BRRL using a distributed transaction processing system and a platform for scientific behavioral simulations.
Tuan Cao, Ben Sowell, Marcos Antonio Vaz Salles, Alan J. Demers, Johannes Gehrke
SIGMOD Conference2
2011 Fast checkpoint recovery algorithms for frequently consistent applications
abstract
Advances in hardware have enabled many long-running applications to execute entirely in main memory. As a result, these applications have increasingly turned to database techniques to ensure durability in the event of a crash. However, many of these applications, such as massively multiplayer online games and mainmemory OLTP systems, must sustain extremely high update rates – often hundreds of thousands of updates per second. Providing durability for these applications without introducing excessive overhead or latency spikes remains a challenge for application developers. In this paper, we take advantage of frequent points of consistency in many of these applications to develop novel checkpoint recovery algorithms that trade additional space in main memory for significantly lower overhead and latency. Compared to previous work, our new algorithms do not require any locking or bulk copies of the application state. Our experimental evaluation shows that one of our new algorithms attains nearly constant latency and reduces overhead by more than an order of magnitude for low to medium update rates. Additionally, in a heavily loaded main-memory transaction processing system, it still reduces overhead by more than a factor of two.
Tuan Cao, Marcos Antonio Vaz Salles, Ben Sowell, Yao Yue, Alan J. Demers, Johannes Gehrke, Walker M. White
SIGMOD Conference3
2010 Behavioral Simulations in MapReduce
abstract
In many scientific domains, researchers are turning to large-scale behavioral simulations to better understand real-world phenomena. While there has been a great deal of work on simulation tools from the high-performance computing community, behavioral simulations remain challenging to program and automatically scale in parallel environments. In this paper we present BRACE (Big Red Agent-based Computation Engine), which extends the MapReduce framework to process these simulations efficiently across a cluster. We can leverage spatial locality to treat behavioral simulations as iterated spatial joins and greatly reduce the communication between nodes. In our experiments we achieve nearly linear scale-up on several realistic simulations. Though processing behavioral simulations in parallel as iterated spatial joins can be very efficient, it can be much simpler for the domain scientists to program the behavior of a single agent. Furthermore, many simulations include a considerable amount of complex computation and message passing between agents, which makes it important to optimize the performance of a single node and the communication across nodes. To address both of these challenges, BRACE includes a high-level language called BRASIL (the Big Red Agent SImulation Language). BRASIL has object-oriented features for programming simulations, but can be compiled to a dataflow representation for automatic parallelization and optimization. We show that by using various optimization techniques, we can achieve both scalability and single-node performance similar to that of a hand-coded simulation.
Guozhang Wang, Marcos Antonio Vaz Salles, Ben Sowell, Tuan Cao, Alan J. Demers, Johannes Gehrke, Walker M. White
Proc. VLDB Endow.3
2009 From Declarative Languages to Declarative Processing in Computer Games
Ben Sowell, Alan J. Demers, Johannes Gehrke, Nitin Gupta 0003, Haoyuan Li 0001, Walker M. White
CIDR1
2009 Database research in computer games
abstract
This tutorial presents an overview of the data management issues faced by computer games today. While many games do not use databases directly, they still have to process large amounts of data, and could benefit from the application of database technology. Other games, such as massively multiplayer online games (MMOs), must communicate with commercial databases and have their own unique challenges. In this tutorial we will present the state-of-the-art of data management in games that we learned from our interaction with various game studios. We will show how the issues involved motivate current research, and illustrate several possibilities for future work.
Alan J. Demers, Johannes Gehrke, Christoph Koch 0001, Ben Sowell, Walker M. White
SIGMOD Conference4
2009 An Evaluation of Checkpoint Recovery for Massively Multiplayer Online Games
abstract
Massively multiplayer online games (MMOs) have emerged as an exciting new class of applications for database technology. MMOs simulate long-lived, interactive virtual worlds, which proceed by applying updates in frames or ticks, typically at 30 or 60 Hz. In order to sustain the resulting high update rates of such games, game state is kept entirely in main memory by the game servers. Nevertheless, durability in MMOs is usually achieved by a standard DBMS implementing ARIES-style recovery. This architecture limits scalability, forcing MMO developers to either invest in high-end hardware or to over-partition their virtual worlds. In this paper, we evaluate the applicability of existing checkpoint recovery techniques developed for main-memory DBMS to MMO workloads. Our thorough experimental evaluation uses a detailed simulation model fed with update traces generated synthetically and from a prototype game server. Based on our results, we recommend MMO developers to adopt a copy-on-update scheme with a double-backup disk organization to checkpoint game state. This scheme outperforms alternatives in terms of the latency introduced in the game as well the time necessary to recover after a crash.
Marcos Antonio Vaz Salles, Tuan Cao, Ben Sowell, Alan J. Demers, Johannes Gehrke, Christoph Koch 0001, Walker M. White
Proc. VLDB Endow.3
2008 SGL: a scalable language for data-driven games
abstract
We propose to demonstrate SGL, a language and system for writing computer games using data management techniques. We will demonstrate a complete game built using the system, and show how complex game behavior can be expressed in a declarative scripting language. The demo will also illustrate the workflow necessary to modify a game and include a visualization of the relational operations that are executed as the game runs.
Robert Albright, Alan J. Demers, Johannes Gehrke, Nitin Gupta 0003, Hooyeon Lee, Rick Keilty, Gregory Sadowski, Ben Sowell, Walker M. White
SIGMOD Conference8
2007 Depth of Field and Cautious-Greedy Routing in Social Networks
David Barbella, George Kachergis, David Liben-Nowell, Anna Sallstrom, Ben Sowell
ISAAC5