VLDB 2026 Research / reviewers in the wild / expert
Ron Barber
dblp:33/3527 · also Ronald Barber
· DBLP profile ↗
16ranked-venue papers
6as first author
0since 2021 · last 2020
0009-0005-2952-7564ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 13 · 5 first-authorArtificial intelligence and machine learning · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
9 papers |
Indexing and storage engines · 40% Query processing and optimization · 31% Database system architecture and tuning · 11% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Cloud and datacenter computing · 67% Storage systems · 33% |
Topics — the 21 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Indexing and storage engines
learned index |
0.8 | 2 | 2019 | HERMIT in Action: Succinct Secondary Indexing Mechanism via Correlation Exploration · Proc. VLDB Endow. 2019 Designing Succinct Secondary Indexing Mechanism by Exploiting Column Correlations · SIGMOD Conference 2019 |
Indexing and storage engines
secondary index |
0.8 | 2 | 2019 | HERMIT in Action: Succinct Secondary Indexing Mechanism via Correlation Exploration · Proc. VLDB Endow. 2019 Designing Succinct Secondary Indexing Mechanism by Exploiting Column Correlations · SIGMOD Conference 2019 |
Query processing and optimization
SQL query processing |
0.4 | 1 | 2020 | Db2 Event Store: A Purpose-Built IoT Database Engine · Proc. VLDB Endow. 2020 |
Indexing and storage engines › column store
main-memory column store |
0.4 | 2 | 2015 | In-memory BLU acceleration in IBM's DB2 and dashDB: Optimized for modern workloads and hardware architectures · ICDE 2015 DB2 with BLU Acceleration: So Much More than Just a Column Store · Proc. VLDB Endow. 2013 |
Information retrieval › indexing › index compression
index pruning |
0.4 | 1 | 2019 | Designing Succinct Secondary Indexing Mechanism by Exploiting Column Correlations · SIGMOD Conference 2019 |
Query processing and optimization
join processing |
0.4 | 2 | 2014 | Joins on Encoded and Partitioned Data · Proc. VLDB Endow. 2014 Memory-Efficient Hash Joins · Proc. VLDB Endow. 2014 |
Query processing and optimization
range query |
0.4 | 1 | 2019 | Designing Succinct Secondary Indexing Mechanism by Exploiting Column Correlations · SIGMOD Conference 2019 |
Database system architecture and tuning
hybrid transactional and analytical processing |
0.2 | 1 | 2016 | Wildfire: Concurrent Blazing Data Ingest and Analytics · SIGMOD Conference 2016 |
Query processing and optimization › join processing › join algorithms
hash join |
0.2 | 2 | 2014 | Memory-Efficient Hash Joins · Proc. VLDB Endow. 2014 Joins on Encoded and Partitioned Data · Proc. VLDB Endow. 2014 |
Query processing and optimization › query execution › hardware-accelerated query processing
SIMD query processing |
0.2 | 2 | 2015 | DB2 with BLU Acceleration: So Much More than Just a Column Store · Proc. VLDB Endow. 2013 In-memory BLU acceleration in IBM's DB2 and dashDB: Optimized for modern workloads and hardware architectures · ICDE 2015 |
Indexing and storage engines
data compression |
0.2 | 1 | 2014 | Joins on Encoded and Partitioned Data · Proc. VLDB Endow. 2014 |
Indexing and storage engines
column store |
0.2 | 1 | 2013 | DB2 with BLU Acceleration: So Much More than Just a Column Store · Proc. VLDB Endow. 2013 |
Query processing and optimization
compressed data processing |
0.2 | 1 | 2013 | DB2 with BLU Acceleration: So Much More than Just a Column Store · Proc. VLDB Endow. 2013 |
Indexing and storage engines › data compression
dictionary compression |
0.2 | 1 | 2013 | DB2 with BLU Acceleration: So Much More than Just a Column Store · Proc. VLDB Endow. 2013 |
Query processing and optimization › query execution
in-memory query processing |
0.2 | 1 | 2013 | DB2 with BLU Acceleration: So Much More than Just a Column Store · Proc. VLDB Endow. 2013 |
Cloud and datacenter computing
database-as-a-service |
0.1 | 1 | 2020 | Db2 Event Store: A Purpose-Built IoT Database Engine · Proc. VLDB Endow. 2020 |
Data stream processing
streaming analytics |
0.1 | 1 | 2016 | Wildfire: Concurrent Blazing Data Ingest and Analytics · SIGMOD Conference 2016 |
Storage systems
data compression |
0.1 | 1 | 2015 | In-memory BLU acceleration in IBM's DB2 and dashDB: Optimized for modern workloads and hardware architectures · ICDE 2015 |
Indexing and storage engines
hash index |
0.1 | 1 | 2014 | Memory-Efficient Hash Joins · Proc. VLDB Endow. 2014 |
Indexing and storage engines
buffer management |
0.0 | 1 | 2013 | DB2 with BLU Acceleration: So Much More than Just a Column Store · Proc. VLDB Endow. 2013 |
Data models and query languages
semistructured data |
0.0 | 1 | 2000 | Evolution of Groupware for Business Applications: A Database Perspective on Lotus Domino/Notes · VLDB 2000 |
Methods — techniques the papers use, named apart from their topics
columnar storage · 1.3SQL compiler reuse · 0.9soft functional dependency · 0.8curve fitting · 0.8SIMD · 0.6shadow tables · 0.4tiered regression search tree · 0.4regression tree · 0.4linear probing · 0.2bloom filter · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Db2 Event Store: A Purpose-Built IoT Database EngineabstractThe requirements of Internet of Things (IoT) workloads are unique in the database space. While significant effort has been spent over the last decade rearchitecting OLTP and Analytics workloads for the public cloud, little has been done to rearchitect IoT workloads for the cloud. In this paper we present IBM Db2 Event Store ™ , a cloud-native database system designed specifically for IoT workloads, which require extremely high-speed ingest, efficient and open data storage, and near real-time analytics. Additionally, by leveraging the Db2 SQL compiler, optimizer and runtime, developed and refined over the last 30 years, we demonstrate that rearchitecting for the public cloud doesn't require rewriting all components. Reusing components that have been built out and optimized for decades dramatically reduced the development effort and immediately provided rich SQL support and excellent run-time query performance. Christian Garcia-Arellano, Adam J. Storm, David Kalmuk, Hamdi Roumani, Ron Barber, Yuanyuan Tian 0001, Richard Sidle, Fatma Özcan 0001, Matt Spilchen, Josh Tiefenbach, Daniel C. Zilio, Lan Pham, Kostas Rakopoulos, Alexander Cheung, Darren Pepper, Imran Sayyid, Gidon Gershinsky, Gal Lushi, Hamid Pirahesh |
Proc. VLDB Endow. | 5 |
| 2019 | WiSer: A Highly Available HTAP DBMS for IoT ApplicationsabstractIn a classic transactional distributed database management system (DBMS), write transactions invariably synchronize with a coordinator before final commitment. While enforcing serializability, this model has long been criticized for not satisfying the applications' availability requirements. When entering the era of Internet of Things (IoT), this problem has become more severe, as an increasing number of applications call for the capability of hybrid transactional and analytical processing (HTAP), where aggregation constraints need to be enforced as part of transactions. Current systems work around this by creating escrows, allowing occasional overshoots of constraints, which are handled via compensating application logic.The WiSer DBMS targets consistency with availability, by splitting the database commit into two steps. First, a PROMISE step that corresponds to what humans are used to as commitment, and runs without talking to a coordinator. Second, a SERIALIZE step, that fixes transactions' positions in the serializable order, via a consensus procedure. We achieve this split via a novel data representation that embeds read-sets into transaction deltas, and serialization sequence numbers into table rows. WiSer does no sharding (all nodes can run transactions that modify the entire database), and yet enforces aggregation constraints. Both read-write conflicts and aggregation constraint violations are resolved lazily in the serialized data. WiSer also covers node joins and departures as database tables, thus simplifying correctness and failure handling. We present the design of WiSer as well as experiments suggesting this approach has promise. Ron Barber, Adam J. Storm, Yuanyuan Tian 0001, Pinar Tözün, Yingjun Wu, Christian Garcia-Arellano, Ronen Grosman, Guy M. Lohman, C. Mohan 0001, René Müller 0001, Hamid Pirahesh, Vijayshankar Raman, Richard Sidle |
IEEE BigData | 1 |
| 2019 | Umzi: Unified Multi-Zone Indexing for Large-Scale HTAPabstractThe rising demands of real-time analytics have emphasized the need for Hybrid Transactional and Analytical Processing (HTAP) systems, which can handle both fast transactions and analytics concurrently. Wildfire is such a large-scale HTAP system prototyped at IBM Research - Almaden, with many techniques developed in this project incorporated into the IBM’s HTAP product offering. To support both workloads efficiently, Wildfire organizes data differently across multiple zones, with more recent data in a more transaction-friendly zone and older data in a more analytics-friendly zone. Data evolve from one zone to another, as they age. In fact, many other HTAP systems have also employed the multi-zone design, including SAP HANA, MemSQL, and SnappyData. Providing a unified index on the large volumes of data across multiple zones is crucial to enable fast point queries and range queries, for both transaction processing and real-time analytics. However, due to the scale and evolving nature of the data, this is a highly challenging task. In this paper, we present Umzi, the multi-version and multi-zone LSM-like indexing method in the Wildfire HTAP system. To the best of our knowledge, Umzi is the first indexing method to support evolving data across multiple zones in an HTAP system, providing a consistent and unified indexing view on the data, despite the constantly on-going changes underneath. Umzi employs a flexible index structure that combines hash and sort techniques together to support both equality and range queries. Moreover, it fully exploits the storage hierarchy in a distributed cluster environment (memory, SSD, and distributed shared storage) for index efficiency. Finally, all index maintenance operations in Umzi are designed to be non-blocking and lock-free for queries to achieve maximum concurrency, while only minimum locking overhead is incurred for concurrent index modifications. Chen Luo 0002, Pinar Tözün, Yuanyuan Tian 0001, Ron Barber, Vijayshankar Raman, Richard Sidle |
EDBT | 4 |
| 2019 | Designing Succinct Secondary Indexing Mechanism by Exploiting Column CorrelationsabstractDatabase administrators construct secondary indexes on data tables to accelerate query processing in relational database management systems (RDBMSs). These indexes are built on top of the most frequently queried columns according to the data statistics. Unfortunately, maintaining multiple secondary indexes in the same database can be extremely space consuming, causing significant performance degradation due to the potential exhaustion of memory space. In this paper, we demonstrate that there exist many opportunities to exploit column correlations for accelerating data access. We propose HERMIT, a succinct secondary indexing mechanism for modern RDBMSs. HERMIT judiciously leverages the rich soft functional dependencies hidden among columns to prune out redundant structures for indexed key access. Instead of building a complete index that stores every single entry in the key columns, HERMIT navigates any incoming key access queries to an existing index built on the correlated columns. This is achieved through the Tiered Regression Search Tree (TRS-Tree), a succinct, ML-enhanced data structure that performs fast curve fitting to adaptively and dynamically capture both column correlations and outliers. Our extensive experimental study in two different RDBMSs have confirmed that HERMIT can significantly reduce space consumption with limited performance overhead, especially when supporting complex range queries. Yingjun Wu, Jia Yu 0001, Yuanyuan Tian 0001, Richard Sidle, Ron Barber |
SIGMOD Conference | 5 |
| 2019 | HERMIT in Action: Succinct Secondary Indexing Mechanism via Correlation ExplorationabstractDatabase administrators construct secondary indexes on data tables to accelerate query processing in relational database management systems (RDBMSs). These indexes are built on top of the most frequently queried columns according to the data statistics. Unfortunately, maintaining multiple secondary indexes in the same database can be extremely space consuming, causing significant performance degradation due to the potential exhaustion of memory space. However, we find that there indeed exist many opportunities to save storage space by exploiting column correlations. We recently introduced Hermit, a succinct secondary indexing mechanism for modern RDBMSs. Hermit judiciously leverages the rich soft functional dependencies hidden among columns to prune out redundant structures for indexed key access. instead of building a complete index that stores every single entry in the key columns, Hermit navigates any incoming key access queries to an existing index built on the correlated columns. This is achieved through the Tiered Regression Search Tree (TRS-Tree), a succinct, ML-enhanced data structure that performs fast curve fitting to adaptively and dynamically capture both column correlations and outliers. In this demonstration, we showcase Hermit's appealing characteristics. we not only demonstrate that Hermit can significantly reduce space consumption with limited performance overhead in terms of query response time and index maintenance time, but also explain in detail the rationale behind Hermit's high efficiency using interactive online query processing examples. Yingjun Wu, Jia Yu 0001, Yuanyuan Tian 0001, Richard Sidle, Ron Barber |
Proc. VLDB Endow. | 5 |
| 2017 | Evolving Databases for New-Gen Big Data Applications
Ron Barber, Christian Garcia-Arellano, Ronen Grosman, René Müller 0001, Vijayshankar Raman, Richard Sidle, Matt Spilchen, Adam J. Storm, Yuanyuan Tian 0001, Pinar Tözün, Daniel C. Zilio, Matt Huras, Guy M. Lohman, C. Mohan 0001, Fatma Özcan 0001, Hamid Pirahesh |
CIDR | 1 |
| 2016 | Wildfire: Concurrent Blazing Data Ingest and AnalyticsabstractWe demonstrate Hybrid Transactional and Analytics Processing (HTAP) on the Spark platform by the Wildfire prototype, which can ingest up to ~6 million inserts per second per node and simultaneously perform complex SQL analytics queries. Here, a simplified mobile application uses Wildfire to recommend advertising to mobile customers based upon their distance from stores and their interest in products sold by these stores, while continuously graphing analytics results as those customers move and respond to the ads with purchases. Ron Barber, Matt Huras, Guy M. Lohman, C. Mohan 0001, René Müller 0001, Fatma Özcan 0001, Hamid Pirahesh, Vijayshankar Raman, Richard Sidle, Oleg Sidorkin, Adam J. Storm, Yuanyuan Tian 0001, Pinar Tözün |
SIGMOD Conference | 1 |
| 2015 | In-memory BLU acceleration in IBM's DB2 and dashDB: Optimized for modern workloads and hardware architecturesabstractAlthough the DRAM for main memories of systems continues to grow exponentially according to Moore's Law and to become less expensive, we argue that memory hierarchies will always exist for many reasons, both economic and practical, and in particular due to concurrent users competing for working memory to perform joins and grouping. We present the in-memory BLU Acceleration used in IBM's DB2 for Linux, UNIX, and Windows, and now also the dashDB cloud offering, which was designed and implemented from the ground up to exploit main memory but is not limited to what fits in memory and does not require manual management of what to retain in memory, as its competitors do. In fact, BLU Acceleration views memory as too slow, and is carefully engineered to work in higher levels of the system cache by keeping the data encoded and packed densely into bit-aligned vectors that can exploit SIMD instructions in processing queries. To achieve scalable multi-core parallelism, BLU assigns to each thread independent data structures, or partitions thereof, designed to have low synchronization costs, and doles out batches of values to threads. On customer workloads, BLU has improved performance on complex analytics queries by 10 to 50 times, compared to the legacy row-organized run-time, while also significantly simplifying database administration, shortening time to value, and improving data compression. UPDATE and DELETE performance was improved by up to 112 times with the new Cancun release of DB2 with BLU Acceleration, which also added Shadow Tables for high performance on mixed OLTP and BI analytics workloads, and extended DB2's High Availability Disaster Recovery (HADR) and SQL compatibility features to BLU's column-organized tables. Ron Barber, Guy M. Lohman, Vijayshankar Raman, Richard Sidle, Sam Lightstone, Berni Schiefer |
ICDE | 1 |
| 2014 | Memory-Efficient Hash JoinsabstractWe present new hash tables for joins, and a hash join based on them, that consumes far less memory and is usually faster than recently published in-memory joins. Our hash join is not restricted to outer tables that fit wholly in memory. Key to this hash join is a new concise hash table (CHT), a linear probing hash table that has 100% fill factor, and uses a sparse bitmap with embedded population counts to almost entirely avoid collisions. This bitmap also serves as a Bloom filter for use in multi-table joins. We study the random access characteristics of hash joins, and renew the case for non-partitioned hash joins. We introduce a variant of partitioned joins in which only the build is partitioned, but the probe is not, as this is more efficient for large outer tables than traditional partitioned joins. This also avoids partitioning costs during the probe, while at the same time allowing parallel build without latching overheads. Additionally, we present a variant of CHT, called a concise array table (CAT), that can be used when the key domain is moderately dense. CAT is collision-free and avoids storing join keys in the hash table. We perform a detailed comparison of CHT and CAT against leading in-memory hash joins. Our experiments show that we can reduce the memory usage by one to three orders of magnitude, while also being competitive in performance. Ron Barber, Guy M. Lohman, Ippokratis Pandis, Vijayshankar Raman, Richard Sidle, Gopi K. Attaluri, Naresh Chainani, Sam Lightstone, David Sharpe |
Proc. VLDB Endow. | 1 |
| 2014 | Joins on Encoded and Partitioned DataabstractCompression has historically been used to reduce the cost of storage, I/Os from that storage, and buffer pool utilization, at the expense of the CPU required to decompress data every time it is queried. However, significant additional CPU efficiencies can be achieved by deferring decompression as late in query processing as possible and performing query processing operations directly on the still-compressed data. In this paper, we investigate the benefits and challenges of performing joins on compressed (or encoded) data. We demonstrate the benefit of independently optimizing the compression scheme of each join column, even though join predicates relating values from multiple columns may require translation of the encoding of one join column into the encoding of the other. We also show the benefit of compressing "payload" data other than the join columns "on the fly," to minimize the size of hash tables used in the join. By partitioning the domain of each column and defining separate dictionaries for each partition, we can achieve even better overall compression as well as increased flexibility in dealing with new values introduced by updates. Instead of decompressing both join columns participating in a join to resolve their different compression schemes, our system performs a light-weight mapping of only qualifying rows from one of the join columns to the encoding space of the other at run time. Consequently, join predicates can be applied directly on the compressed data. We call this procedure encoding translation. Two alternatives of encoding translation are developed and compared in the paper. We provide a comprehensive evaluation of these alternatives using product implementations of each on the TPC-H data set, and demonstrate that performing joins on encoded and partitioned data achieves both superior performance and excellent compression. Jae-Gil Lee 0001, Gopi K. Attaluri, Ron Barber, Naresh Chainani, Oliver Draese, Frederick Ho, Stratos Idreos, Min-Soo Kim 0002, Sam Lightstone, Guy M. Lohman, Konstantinos Morfonios, Keshava Murthy, Ippokratis Pandis, Lin Qiao 0001, Vijayshankar Raman, Vincent KulandaiSamy, Richard Sidle, Knut Stolze |
Proc. VLDB Endow. | 3 |
| 2013 | Go, server, go!: parallel computing with moving serversabstractIn data centers today, servers are stationary and data flows on a hierarchical network of switches and routers. But such static server arrangements require very scalable networks, and many applications are bottlenecked by network bandwidth. In addition, server density is kept low to enable maintenance and upgrades, as well as to increase air flow. In this paper, we propose a design in which servers move physically, and communicate via point-to-point connections (instead of switches). We argue that this allows data transfer bandwidth to scale linearly with the number of servers, and that moving servers is not as expensive as it sounds, at least in terms of power consumption. Moreover, while servers move around, they regularly reach the perimeters of the system, which helps with heat dissipation and with servicing of failed nodes. This design also helps in traditional switch-based networks, to improve density and maintainability. Ron Barber, Guy M. Lohman, René Müller 0001, Ippokratis Pandis, Vijayshankar Raman, Winfried W. Wilcke |
SoCC | 1 |
| 2013 | DB2 with BLU Acceleration: So Much More than Just a Column StoreabstractDB2 with BLU Acceleration deeply integrates innovative new techniques for defining and processing column-organized tables that speed read-mostly Business Intelligence queries by 10 to 50 times and improve compression by 3 to 10 times, compared to traditional row-organized tables, without the complexity of defining indexes or materialized views on those tables. But DB2 BLU is much more than just a column store. Exploiting frequency-based dictionary compression and main-memory query processing technology from the Blink project at IBM Research - Almaden, DB2 BLU performs most SQL operations - predicate application (even range predicates and IN-lists), joins, and grouping - on the compressed values, which can be packed bit-aligned so densely that multiple values fit in a register and can be processed simultaneously via SIMD (single-instruction, multipledata) instructions. Designed and built from the ground up to exploit modern multi-core processors, DB2 BLU's hardware-conscious algorithms are carefully engineered to maximize parallelism by using novel data structures that need little latching, and to minimize data-cache and instruction-cache misses. Though DB2 BLU is optimized for in-memory processing, database size is not limited by the size of main memory. Fine-grained synopses, late materialization, and a new probabilistic buffer pool protocol for scans minimize disk I/Os, while aggressive prefetching reduces I/O stalls. Full integration with DB2 ensures that DB2 with BLU Acceleration benefits from the full functionality and robust utilities of a mature product, while still enjoying order-of-magnitude performance gains from revolutionary technology without even having to change the SQL, and can mix column-organized and row-organized tables in the same tablespace and even within the same query. Vijayshankar Raman, Gopi K. Attaluri, Ron Barber, Naresh Chainani, David Kalmuk, Vincent KulandaiSamy, Jens Leenstra, Sam Lightstone, Shaorong Liu, Guy M. Lohman, Tim Malkemus, René Müller 0001, Ippokratis Pandis, Berni Schiefer, David Sharpe, Richard Sidle, Adam J. Storm |
Proc. VLDB Endow. | 3 |
| 2000 | Evolution of Groupware for Business Applications: A Database Perspective on Lotus Domino/Notes
C. Mohan 0001, Ron Barber, S. Watts, Amit Somani, Markos Zaharioudakis |
VLDB | 2 |
| 1994 | Query by Image Content using Multiple Objects and Multiple Features: Use Interface IssuesabstractOn-line collections of images are growing larger and more common, and tools are needed to efficiently manage, organize, and navigate through them. The authors have developed a prototype system called QBIC which allows complex multi-object and multi-feature queries of large image databases. The queries are based on image content-the colors, textures, shapes, and positions of images and the objects/regions they contain. The system computes numeric features to represent the image properties and uses similarity measures based on these features for image retrieval. The focus of the paper is its user interface which allows a user to graphically pose and refine queries based on multiple visual properties of images and their objects.> Denis Lee 0001, Ron Barber, Wayne Niblack, Myron Flickner, James Lee Hafner, Dragutin Petkovic |
ICIP (2) | 2 |
| 1994 | Indexing for complex queries on a query-by-content image databaseabstractWe describe how the QBIC (Query By Image Content) system handles "multi-*" queries-queries on large image collections involving multifeatures of each image as a whole and of multiple objects within each image. The queries are based on properties of image content-such as colors, textures, shapes, and edges. The system computes a set of features to describe the above properties, uses distance-like measures on the features to provide similarity based retrieval, and has a graphical interface that enable users pose queries visually. In this paper, we present QBIC indexing algorithms that allow these "multi-*" queries to run efficiently. Denis Lee 0001, Ron Barber, Wayne Niblack, Myron Flickner, James Lee Hafner, Dragutin Petkovic |
ICPR (1) | 2 |
| 1994 | Efficient and Effective Querying by Image Content
Christos Faloutsos, Ron Barber, Myron Flickner, James Lee Hafner, Wayne Niblack, Dragutin Petkovic, William Equitz |
J. Intell. Inf. Syst. | 2 |