Stefanos Baziotis

dblp:344/1858 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0001-4061-7094ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 70% Data mining · 30%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
approximate query processing
0.912025
PilotDB: Database-Agnostic Online Approximate Query Processing with A Priori Error Guarantees · Proc. ACM Manag. Data 2025
Query processing and optimization › approximate query processing › sampling-based approximate query processing
block-level sampling
0.912025
PilotDB: Database-Agnostic Online Approximate Query Processing with A Priori Error Guarantees · Proc. ACM Manag. Data 2025
Data mining
sampling
0.912025
PilotDB: Database-Agnostic Online Approximate Query Processing with A Priori Error Guarantees · Proc. ACM Manag. Data 2025
Compilers and program optimization › code generation
instruction selection
0.812024
Hydride: A Retargetable and Extensible Synthesis-based Compiler for Modern Hardware Architectures · ASPLOS (2) 2024
Compilers and program optimization
intermediate representation
0.812024
Hydride: A Retargetable and Extensible Synthesis-based Compiler for Modern Hardware Architectures · ASPLOS (2) 2024
Compilers and program optimization › compiler construction
retargetable compilation
0.812024
Hydride: A Retargetable and Extensible Synthesis-based Compiler for Modern Hardware Architectures · ASPLOS (2) 2024
Data mining
exploratory data analysis
0.212024
Dias: Dynamic Rewriting of Pandas Code · Proc. ACM Manag. Data 2024
Hardware accelerators and domain-specific architectures
domain-specific compilers
0.212024
Hydride: A Retargetable and Extensible Synthesis-based Compiler for Modern Hardware Architectures · ASPLOS (2) 2024

Methods — techniques the papers use, named apart from their topics

synthesis-based compilation · 1.5pseudocode specification · 1.5two-stage online AQP · 0.9statistical guarantees · 0.9dynamic precondition checking · 0.8
YearPublicationVenuePosition
2025 PilotDB: Database-Agnostic Online Approximate Query Processing with A Priori Error Guarantees
abstract
After decades of research in approximate query processing (AQP), its adoption in the industry remains limited. Existing methods struggle to simultaneously provide user-specified error guarantees, eliminate maintenance overheads, and avoid modifications to database management systems. To address these challenges, we introduce two novel techniques, TAQA and BSAP. TAQA is a two-stage online AQP algorithm that achieves all three properties for arbitrary queries. However, it can be slower than exact queries if we use standard row-level sampling. BSAP resolves this by enabling block-level sampling with statistical guarantees in TAQA. We implement TAQA and BSAP in a prototype middleware system, PilotDB, that is compatible with all DBMSs supporting efficient block-level sampling. We evaluate PilotDB on PostgreSQL, SQL Server, and DuckDB over real-world benchmarks, demonstrating up to 126X speedups when running with a 5% guaranteed error.
Yuxuan Zhu 0003, Tengjun Jin, Stefanos Baziotis, Chengsong Zhang, Charith Mendis, Daniel Kang 0001
Proc. ACM Manag. Data3
2024 Hydride: A Retargetable and Extensible Synthesis-based Compiler for Modern Hardware Architectures
abstract
As modern hardware architectures evolve to support increasingly diverse, complex instruction sets for meeting the performance demands of modern workloads in image processing, deep learning, etc., it has become ever more crucial for compilers to provide robust support for evolution of their internal abstractions and retargetable code generation support to keep pace with emerging instruction sets. We propose Hydride, a novel approach to compiling for complex, emerging hardware architectures. Hydride uses vendor-defined pseudocode specifications of multiple hardware ISAs to automatically design retargetable instructions for AutoLLVM IR, an extensible compiler IR which consists of (formally defined) language-independent and target-independent LLVM IR instructions to compile to those ISAs, and automatically generated instruction selection passes to lower AutoLLVM IR to each of the specified hardware ISAs. Hydride also includes a code synthesizer that automatically generates code generation support for schedule-based languages, such as Halide, to optimally generate AutoLLVM IR. Our results show that Hydride is able to represent 3,557 instructions combined in x86, Hexagon, ARM architectures using only 397 AutoLLVM IR instructions, including (Intel) SSE2, SSE4, AVX, AVX2, AVX512, (Qualcomm) Hexagon HVX, and (ARM) NEON vector ISAs. We created a new Halide compiler with Hydride using only a formal semantics of Halide IR, leveraging the auto-generated AutoLLVM IR and back-ends for the three hardware architectures. Across kernels from deep learning and image processing, this compiler is able to perform just as well as the mature, production Halide compiler on Hexagon, and outperform on x86 by 8% and ARM by 3%. Hydride also outperforms the production Halide's LLVM back end by 12% on x86, 100% on HVX, and 26% on ARM across the same kernels.
Akash Kothari, Abdul Rafae Noor, Muchen Xu, Hassam Uddin, Dhruv Baronia, Stefanos Baziotis, Vikram S. Adve, Charith Mendis, Sudipta Sengupta
ASPLOS (2)6
2024 Dias: Dynamic Rewriting of Pandas Code
abstract
In recent years, dataframe libraries, such as pandas have exploded in popularity. Due to their flexibility, they are increasingly used in ad-hoc exploratory data analysis (EDA) workloads. These workloads are diverse, including custom functions which can span libraries or be written in pure Python. The majority of systems available to accelerate EDA workloads focus on bulk-parallel workloads, which contain vastly different computational patterns, typically within a single library. As a result, they can introduce excessive overheads for ad-hoc EDA workloads due to their expensive optimization techniques. Instead, we identify source-to-source, external program rewriting as a lightweight technique which can optimize across representations, and offer substantial speedups while also avoiding slowdowns. We implemented Dias, which rewrites notebook cells to be more efficient for ad-hoc EDA workloads. We develop techniques for efficient rewrites in Dias, including checking the preconditions under which rewrites are correct, dynamically, at fine-grained program points. We show that Dias can rewrite individual cells to be 57× faster compared to pandas and 1909× faster compared to optimized systems such as modin. Furthermore, Dias can accelerate whole notebooks by up to 3.6× compared to pandas and 27.1× compared to modin.
Stefanos Baziotis, Daniel Kang 0001, Charith Mendis
Proc. ACM Manag. Data1