Ilya Yatsishin

dblp:385/6244 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 56% Database system architecture and tuning · 28% Indexing and storage engines · 8%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
query compilation
0.812024
ClickHouse - Lightning Fast Analytics for Everyone · Proc. VLDB Endow. 2024
Query processing and optimization › query execution › query operator implementation
vectorized query execution
0.812024
ClickHouse - Lightning Fast Analytics for Everyone · Proc. VLDB Endow. 2024
Data mining › data reduction
data pruning
0.212024
ClickHouse - Lightning Fast Analytics for Everyone · Proc. VLDB Endow. 2024
Indexing and storage engines
LSM-tree
0.212024
ClickHouse - Lightning Fast Analytics for Everyone · Proc. VLDB Endow. 2024
YearPublicationVenuePosition
2024 ClickHouse - Lightning Fast Analytics for Everyone
abstract
Over the past several decades, the amount of data being stored and analyzed has increased exponentially. Businesses across industries and sectors have begun relying on this data to improve products, evaluate performance, and make business-critical decisions. However, as data volumes have increasingly become internet-scale, businesses have needed to manage historical and new data in a cost-effective and scalable manner, while analyzing it using a high number of concurrent queries and an expectation of real-time latencies (e.g. less than one second, depending on the use case). This paper presents an overview of ClickHouse, a popular open-source OLAP database designed for high-performance analytics over petabyte-scale data sets with high ingestion rates. Its storage layer combines a data format based on traditional log-structured merge (LSM) trees with novel techniques for continuous transformation (e.g. aggregation, archiving) of historical data in the background. Queries are written in a convenient SQL dialect and processed by a state-of-the-art vectorized query execution engine with optional code compilation. ClickHouse makes aggressive use of pruning techniques to avoid evaluating irrelevant data in queries. Other data management systems can be integrated at the table function, table engine, or database engine level. Real-world benchmarks demonstrate that ClickHouse is amongst the fastest analytical databases on the market.
Robert Schulze, Tom Schreiber, Ilya Yatsishin, Ryadh Dahimene, Alexey Milovidov
Proc. VLDB Endow.3