VLDB 2026 Research / reviewers in the wild / expert
Stefanos Baziotis
dblp:344/1858
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0001-4061-7094ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Query processing and optimization · 70% Data mining · 30% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
approximate query processing |
0.9 | 1 | 2025 | PilotDB: Database-Agnostic Online Approximate Query Processing with A Priori Error Guarantees · Proc. ACM Manag. Data 2025 |
Query processing and optimization › approximate query processing › sampling-based approximate query processing
block-level sampling |
0.9 | 1 | 2025 | PilotDB: Database-Agnostic Online Approximate Query Processing with A Priori Error Guarantees · Proc. ACM Manag. Data 2025 |
Data mining
sampling |
0.9 | 1 | 2025 | PilotDB: Database-Agnostic Online Approximate Query Processing with A Priori Error Guarantees · Proc. ACM Manag. Data 2025 |
Compilers and program optimization › code generation
instruction selection |
0.8 | 1 | 2024 | Hydride: A Retargetable and Extensible Synthesis-based Compiler for Modern Hardware Architectures · ASPLOS (2) 2024 |
Compilers and program optimization
intermediate representation |
0.8 | 1 | 2024 | Hydride: A Retargetable and Extensible Synthesis-based Compiler for Modern Hardware Architectures · ASPLOS (2) 2024 |
Compilers and program optimization › compiler construction
retargetable compilation |
0.8 | 1 | 2024 | Hydride: A Retargetable and Extensible Synthesis-based Compiler for Modern Hardware Architectures · ASPLOS (2) 2024 |
Data mining
exploratory data analysis |
0.2 | 1 | 2024 | Dias: Dynamic Rewriting of Pandas Code · Proc. ACM Manag. Data 2024 |
Hardware accelerators and domain-specific architectures
domain-specific compilers |
0.2 | 1 | 2024 | Hydride: A Retargetable and Extensible Synthesis-based Compiler for Modern Hardware Architectures · ASPLOS (2) 2024 |
Methods — techniques the papers use, named apart from their topics
synthesis-based compilation · 1.5pseudocode specification · 1.5two-stage online AQP · 0.9statistical guarantees · 0.9dynamic precondition checking · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PilotDB: Database-Agnostic Online Approximate Query Processing with A Priori Error GuaranteesabstractAfter decades of research in approximate query processing (AQP), its adoption in the industry remains limited. Existing methods struggle to simultaneously provide user-specified error guarantees, eliminate maintenance overheads, and avoid modifications to database management systems. To address these challenges, we introduce two novel techniques, TAQA and BSAP. TAQA is a two-stage online AQP algorithm that achieves all three properties for arbitrary queries. However, it can be slower than exact queries if we use standard row-level sampling. BSAP resolves this by enabling block-level sampling with statistical guarantees in TAQA. We implement TAQA and BSAP in a prototype middleware system, PilotDB, that is compatible with all DBMSs supporting efficient block-level sampling. We evaluate PilotDB on PostgreSQL, SQL Server, and DuckDB over real-world benchmarks, demonstrating up to 126X speedups when running with a 5% guaranteed error. Yuxuan Zhu 0003, Tengjun Jin, Stefanos Baziotis, Chengsong Zhang, Charith Mendis, Daniel Kang 0001 |
Proc. ACM Manag. Data | 3 |
| 2024 | Hydride: A Retargetable and Extensible Synthesis-based Compiler for Modern Hardware ArchitecturesabstractAs modern hardware architectures evolve to support increasingly diverse, complex instruction sets for meeting the performance demands of modern workloads in image processing, deep learning, etc., it has become ever more crucial for compilers to provide robust support for evolution of their internal abstractions and retargetable code generation support to keep pace with emerging instruction sets. We propose Hydride, a novel approach to compiling for complex, emerging hardware architectures. Hydride uses vendor-defined pseudocode specifications of multiple hardware ISAs to automatically design retargetable instructions for AutoLLVM IR, an extensible compiler IR which consists of (formally defined) language-independent and target-independent LLVM IR instructions to compile to those ISAs, and automatically generated instruction selection passes to lower AutoLLVM IR to each of the specified hardware ISAs. Hydride also includes a code synthesizer that automatically generates code generation support for schedule-based languages, such as Halide, to optimally generate AutoLLVM IR. Our results show that Hydride is able to represent 3,557 instructions combined in x86, Hexagon, ARM architectures using only 397 AutoLLVM IR instructions, including (Intel) SSE2, SSE4, AVX, AVX2, AVX512, (Qualcomm) Hexagon HVX, and (ARM) NEON vector ISAs. We created a new Halide compiler with Hydride using only a formal semantics of Halide IR, leveraging the auto-generated AutoLLVM IR and back-ends for the three hardware architectures. Across kernels from deep learning and image processing, this compiler is able to perform just as well as the mature, production Halide compiler on Hexagon, and outperform on x86 by 8% and ARM by 3%. Hydride also outperforms the production Halide's LLVM back end by 12% on x86, 100% on HVX, and 26% on ARM across the same kernels. Akash Kothari, Abdul Rafae Noor, Muchen Xu, Hassam Uddin, Dhruv Baronia, Stefanos Baziotis, Vikram S. Adve, Charith Mendis, Sudipta Sengupta |
ASPLOS (2) | 6 |
| 2024 | Dias: Dynamic Rewriting of Pandas CodeabstractIn recent years, dataframe libraries, such as pandas have exploded in popularity. Due to their flexibility, they are increasingly used in ad-hoc exploratory data analysis (EDA) workloads. These workloads are diverse, including custom functions which can span libraries or be written in pure Python. The majority of systems available to accelerate EDA workloads focus on bulk-parallel workloads, which contain vastly different computational patterns, typically within a single library. As a result, they can introduce excessive overheads for ad-hoc EDA workloads due to their expensive optimization techniques. Instead, we identify source-to-source, external program rewriting as a lightweight technique which can optimize across representations, and offer substantial speedups while also avoiding slowdowns. We implemented Dias, which rewrites notebook cells to be more efficient for ad-hoc EDA workloads. We develop techniques for efficient rewrites in Dias, including checking the preconditions under which rewrites are correct, dynamically, at fine-grained program points. We show that Dias can rewrite individual cells to be 57× faster compared to pandas and 1909× faster compared to optimized systems such as modin. Furthermore, Dias can accelerate whole notebooks by up to 3.6× compared to pandas and 27.1× compared to modin. Stefanos Baziotis, Daniel Kang 0001, Charith Mendis |
Proc. ACM Manag. Data | 1 |