Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Robert Walzer

dblp:185/4277 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Query processing and optimization · 31% Indexing and storage engines · 19% Information retrieval · 19%

Topics — the 9 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › similarity search
nearest neighbor search
0.812024
SingleStore-V: An Integrated Vector Database System in SingleStore · Proc. VLDB Endow. 2024
Indexing and storage engines
vector database
0.812024
SingleStore-V: An Integrated Vector Database System in SingleStore · Proc. VLDB Endow. 2024
Database system architecture and tuning
hybrid transactional and analytical processing
0.612022
Cloud-Native Transactions and Analytics in SingleStore · SIGMOD Conference 2022
Query processing and optimization › query optimization
distributed query optimization
0.212016
The MemSQL Query Optimizer: A modern optimizer for real-time analytics in a distributed database · Proc. VLDB Endow. 2016
Query processing and optimization
query optimization
0.212016
The MemSQL Query Optimizer: A modern optimizer for real-time analytics in a distributed database · Proc. VLDB Endow. 2016
Distributed and cloud data management › distributed database architecture
distributed relational database
0.212024
SingleStore-V: An Integrated Vector Database System in SingleStore · Proc. VLDB Endow. 2024
Query processing and optimization
SQL query processing
0.212024
SingleStore-V: An Integrated Vector Database System in SingleStore · Proc. VLDB Endow. 2024
Query processing and optimization
analytical query processing
0.212022
Cloud-Native Transactions and Analytics in SingleStore · SIGMOD Conference 2022
Distributed and cloud data management › cloud database
cloud-native database
0.212022
Cloud-Native Transactions and Analytics in SingleStore · SIGMOD Conference 2022

Methods — techniques the papers use, named apart from their topics

operator specialization · 0.3SIMD · 0.3query rewrite · 0.2join enumeration · 0.2bushy join · 0.2
YearPublicationVenuePosition
2024 SingleStore-V: An Integrated Vector Database System in SingleStore
abstract
Vector databases have recently gained significant attention due to the emergence of large language models that produce vector embeddings for text. Existing vector databases can be broadly categorized into two types: specialized and generalized. Specialized vector databases are explicitly designed and optimized for managing vector data, while generalized ones support vector data management within a general purpose database. While specialized vector databases are interesting, there is a substantial customer base interested in generalized vector databases for various reasons, e.g., a reluctance to move data out of relational databases to reduce data silos and costs, the desire to use SQL, and the need for more sophisticated query processing of vector and non-vector data. However, generalized vector databases face two main challenges: performance and interoperability of vector search with SQL, such as combining vector search with filters, joins, or even fulltext search. In this paper, we present SingleStore-V, a full-fledged generalized vector database integrated into SingleStore, a modern distributed relational database optimized for both OLAP and OLTP workloads. SingleStore-V achieves high performance and interoperability via a suite of optimizations. Experiments on standard vector benchmarks show that SingleStore-V performs comparably to Milvus, a highly-optimized specialized vector database, and significantly outperforms pgvector, a popular generalized vector database in PostgreSQL. We believe this paper will shed light on integrating vector search into relational databases in general, as many design concepts and optimizations apply to other databases.
Chenzhe Jin, Sasha Podolsky, Szu-Po Wang, Eric Hanson, Zhou Sun, Robert Walzer, Jianguo Wang 0001
Proc. VLDB Endow.9
2022 Cloud-Native Transactions and Analytics in SingleStore
abstract
The last decade has seen a remarkable rise in specialized database systems. Systems for transaction processing, data warehousing, time series analysis, full-text search, data lakes, in-memory caching, document storage, queuing, graph processing, and geo-replicated operational workloads are now available to developers. A belief has taken hold that a single general-purpose database is not capable of running varied workloads at a reasonable cost with strong performance, at the level of scale and concurrency people demand today. There is value in specialization, but the complexity and cost of using multiple specialized systems in a single application environment is becoming apparent. This realization is driving developers and IT decision makers to seek databases capable of powering a broader set of use cases when looking to adopt a new database. Hybrid transaction and analytical (HTAP) databases have been developed to try to tame some of this chaos.
Adam Prout, Szu-Po Wang, Joseph Victor, Zhou Sun, Yongzhu Li, Jack Chen, Evan Bergeron, Eric N. Hanson, Robert Walzer, Rodrigo Gomes, Nikita Shamgunov
SIGMOD Conference9
2018 BIPie: Fast Selection and Aggregation on Encoded Data using Operator Specialization
abstract
Advances in modern hardware, such as increases in the size of main memory available on computers, have made it possible to analyze data at a much higher rate than before. In this paper, we demonstrate that there is tremendous room for improvement in the processing of analytical queries on modern commodity hardware. We introduce BIPie, an engine for query processing implementing highly efficient decoding, selection, and aggregation for analytical queries executing on a columnar storage engine in MemSQL. We demonstrate that these operations are interdependent, and must be fused and considered together to achieve very high performance. We propose and compare multiple strategies for decoding, selection and aggregation (with GROUP BY), all of which are designed to take advantage of modern CPU architectures, including SIMD. We implemented these approaches in MemSQL, a high performance hybrid transaction and analytical processing database designed for commodity hardware. We thoroughly evaluate the performance of the approach across a range of parameters, and demonstrate a two to four times speedup over previously published TPC-H Query 1 performance.
Michal Nowakiewicz, Eric Boutin, Eric N. Hanson, Robert Walzer, Akash Katipally
SIGMOD Conference4
2016 The MemSQL Query Optimizer: A modern optimizer for real-time analytics in a distributed database
abstract
Real-time analytics on massive datasets has become a very common need in many enterprises. These applications require not only rapid data ingest, but also quick answers to analytical queries operating on the latest data. MemSQL is a distributed SQL database designed to exploit memory-optimized, scale-out architecture to enable real-time transactional and analytical workloads which are fast, highly concurrent, and extremely scalable. Many analytical queries in MemSQL's customer workloads are complex queries involving joins, aggregations, sub-queries, etc. over star and snowflake schemas, often ad-hoc or produced interactively by business intelligence tools. These queries often require latencies of seconds or less, and therefore require the optimizer to not only produce a high quality distributed execution plan, but also produce it fast enough so that optimization time does not become a bottleneck. In this paper, we describe the architecture of the MemSQL Query Optimizer and the design choices and innovations which enable it quickly produce highly efficient execution plans for complex distributed queries. We discuss how query rewrite decisions oblivious of distribution cost can lead to poor distributed execution plans, and argue that to choose high-quality plans in a distributed database, the optimizer needs to be distribution-aware in choosing join plans, applying query rewrites, and costing plans. We discuss methods to make join enumeration faster and more effective, such as a rewrite-based approach to exploit bushy joins in queries involving multiple star schemas without sacrificing optimization time. We demonstrate the effectiveness of the MemSQL optimizer over queries from the TPC-H benchmark and a real customer workload.
Jack Chen, Samir Jindel, Robert Walzer, Rajkumar Sen, Nika Jimsheleishvilli, Michael Andrews
Proc. VLDB Endow.3