EDBT 2026 Demo / reviewers in the wild / expert
Reza Sherkat
dblp:02/5500
· DBLP profile ↗
9ranked-venue papers
6as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 6 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
7 papers |
Query processing and optimization · 32% Indexing and storage engines · 31% Database system architecture and tuning · 19% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 100% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › buffer management
buffer cache management |
0.4 | 1 | 2019 | Native Store Extension for SAP HANA · Proc. VLDB Endow. 2019 |
Query processing and optimization › runtime optimization › data skipping
partition pruning |
0.3 | 1 | 2017 | Statisticum: Data Statistics Management in SAP HANA · Proc. VLDB Endow. 2017 |
Query processing and optimization
query optimization |
0.3 | 1 | 2017 | Statisticum: Data Statistics Management in SAP HANA · Proc. VLDB Endow. 2017 |
Indexing and storage engines
columnar storage |
0.2 | 1 | 2016 | Page As You Go: Piecewise Columnar Access In SAP HANA · SIGMOD Conference 2016 |
Indexing and storage engines › column store
main-memory column store |
0.2 | 1 | 2016 | Page As You Go: Piecewise Columnar Access In SAP HANA · SIGMOD Conference 2016 |
Indexing and storage engines
column store |
0.1 | 1 | 2019 | Native Store Extension for SAP HANA · Proc. VLDB Endow. 2019 |
Information retrieval › indexing
index compression |
0.1 | 1 | 2009 | Efficient Index Compression in DB2 LUW · Proc. VLDB Endow. 2009 |
Query processing and optimization › query optimization
statistics management |
0.1 | 1 | 2017 | Statisticum: Data Statistics Management in SAP HANA · Proc. VLDB Endow. 2017 |
Data mining › time series analysis
time series similarity |
0.1 | 1 | 2008 | On efficiently searching trajectories and archival data for historical similarities · Proc. VLDB Endow. 2008 |
Spatial and temporal data management › trajectory analysis
trajectory similarity search |
0.1 | 1 | 2008 | On efficiently searching trajectories and archival data for historical similarities · Proc. VLDB Endow. 2008 |
Indexing and storage engines
spatial index |
0.1 | 1 | 2007 | On MBR Approximation of Histories for Historical Queries: Expectations and Limitations · ICDE 2007 |
Database system architecture and tuning
query history |
0.1 | 1 | 2006 | Efficiently Evaluating Order Preserving Similarity Queries over Historical Market-Basket Data · ICDE 2006 |
Information retrieval
similarity search |
0.1 | 1 | 2006 | Efficiently Evaluating Order Preserving Similarity Queries over Historical Market-Basket Data · ICDE 2006 |
Indexing and storage engines
vector index |
0.0 | 1 | 2008 | On efficiently searching trajectories and archival data for historical similarities · Proc. VLDB Endow. 2008 |
Spatial and temporal data management › spatial query processing
filtering-and-verification |
0.0 | 1 | 2007 | On MBR Approximation of Histories for Historical Queries: Expectations and Limitations · ICDE 2007 |
Data mining › pattern mining › pruning
upper bound pruning |
0.0 | 1 | 2006 | Efficiently Evaluating Order Preserving Similarity Queries over Historical Market-Basket Data · ICDE 2006 |
Methods — techniques the papers use, named apart from their topics
run-length encoding · 1.0prefetching · 0.8dictionary encoding · 0.8implied integrity constraints · 0.3constraint data statistics · 0.3consistency checking · 0.3order-preserving dictionary · 0.2inverted index · 0.2feature extraction · 0.1distance function · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Native Store Extension for SAP HANAabstractWe present an overview of SAP HANA's Native Store Extension (NSE). This extension substantially increases database capacity, allowing to scale far beyond available system memory. NSE is based on a hybrid in-memory and paged column store architecture composed from data access primitives. These primitives enable the processing of hybrid columns using the same algorithms optimized for traditional HANA's in-memory columns. Using only three key primitives, we fabricated byte-compatible counterparts for complex memory resident data structures (e.g. dictionary and hash-index), compressed schemes (e.g. sparse and run-length encoding), and exotic data types (e.g. geo-spatial). We developed a new buffer cache which optimizes the management of paged resources by smart strategies sensitive to page type and access patterns. The buffer cache integrates with HANA's new execution engine that issues pipelined prefetch requests to improve disk access patterns. A novel load unit configuration, along with a unified persistence format, allows the hybrid column store to dynamically switch between in-memory and paged data access to balance performance and storage economy according to application demands while reducing Total Cost of Ownership (TCO). A new partitioning scheme supports load unit specification at table, partition, and column level. Finally, a new advisor recommends optimal load unit configurations. Our experiments illustrate the performance and memory footprint improvements on typical customer scenarios. Reza Sherkat, Colin Florendo, Mihnea Andrei, Rolando Blanco, Adrian Dragusanu, Amit Pathak, Pushkar Khadilkar, Neeraj Kulkarni, Christian Lemke, Sebastian Seifert, Sarika Iyer, Sasikanth Gottapu, Robert Schulze, Chaitanya Gottipati, Nirvik Basak, Vivek Kandiyanallur, Santosh Pendap, Dheren Gala, Rajesh Almeida, Prasanta Ghosh |
Proc. VLDB Endow. | 1 |
| 2018 | Global Range Encoding for Efficient Partition Elimination
Jeremy Chen, Reza Sherkat, Mihnea Andrei, Heiko Gerwens |
EDBT | 2 |
| 2017 | Statisticum: Data Statistics Management in SAP HANAabstractWe introduce a new concept of leveraging traditional data statistics as dynamic data integrity constraints. These data statistics produce transient database constraints, which are valid as long as they can be proven to be consistent with the current data. We denote this type of data statistics by constraint data statistics , their properties needed for consistency checking by consistency metadata , and their implied integrity constraints by implied data statistics constraints ( implied constraints for short). Implied constraints are valid integrity constraints which are powerful query optimization tools employed, just as traditional database constraints, in semantic query transformation (aka query reformulation), partition pruning, runtime optimization, and semi-join reduction, to name a few. To our knowledge, this is the first work introducing this novel and powerful concept of deriving implied integrity constraints from data statistics. We discuss theoretical aspects of the constraint data statistics concept and their integration into query processing. We present the current architecture of data statistics management in SAP HANA and detail how constraint data statistics are designed and integrated into this architecture. As an instantiation of this framework, we consider dynamic partition pruning for data aging scenarios. We discuss our current implementation for constraint data statistics objects in SAP HANA which can be used for dynamic partition pruning. We enumerate their properties and show how consistency checking for implied integrity constraints is supported in the data statistics architecture. Our experimental evaluations on the TPC-H benchmark and a real customer application confirm the effectiveness of the implied integrity constraints; (1) for 59% of TPC-H queries, constraint data statistics utilization results in pruning cold partitions and reducing memory consumption, and (2) we observe up to 3 orders of magnitude speed-up in query processing time, for a real customer running an S/4HANA application. Anisoara Nica, Reza Sherkat, Mihnea Andrei, Martin Heidel, Christian Bensberg, Heiko Gerwens |
Proc. VLDB Endow. | 2 |
| 2016 | Page As You Go: Piecewise Columnar Access In SAP HANAabstractIn-memory columnar databases such as SAP HANA achieve extreme performance by means of vector processing over logical units of main memory resident columns. The core in-memory algorithms can be challenged when the working set of an application does not fit into main memory. To deal with memory pressure, most in-memory columnar databases evict candidate columns (or tables) using a set of heuristics gleaned from recent workload. As an alternative approach, we propose to reduce the unit of load and eviction from column to a contiguous portion of the in-memory columnar representation, which we call a page. In this paper, we adapt the core algorithms to be able to operate with partially loaded columns while preserving the performance benefits of vector processing. Our approach has two key advantages. First, partial column loading reduces the mandatory memory footprint for each column, making more memory available for other purposes. Second, partial eviction extends the in-memory lifetime of partially loaded column. We present a new in-memory columnar implementation for our approach, that we term page loadable column. We design a new persistency layout and access algorithms for the encoded data vector of the column, the order-preserving dictionary, and the inverted index. We compare the performance attributes of page loadable columns with those of regular in-memory columns and present a use-case for page loadable columns for cold data in data aging scenarios. Page loadable columns are completely integrated in SAP HANA, and we present extensive experimental results that quantify the performance overhead and the resource consumption when these columns are deployed. Reza Sherkat, Colin Florendo, Mihnea Andrei, Anil K. Goel, Anisoara Nica, Peter Bumbulis, Ivan Schreter, Günter Radestock, Christian Bensberg, Daniel Booss, Heiko Gerwens |
SIGMOD Conference | 1 |
| 2013 | Efficient Time-Stamped Event Sequence AnonymizationabstractWith the rapid growth of applications which generate timestamped sequences (click streams, GPS trajectories, RFID sequences), sequence anonymization has become an important problem, in that should such data be published or shared. Existing trajectory anonymization techniques disregard the importance of time or the sensitivity of events. This article is the first, to our knowledge, thorough study on time-stamped event sequence anonymization. We propose a novel and tunable generalization framework tailored to event sequences. We generalize time stamps using time intervals and events using a taxonomy which models the domain semantics. We consider two scenarios: (i) sharing the data with a single receiver (the SSR setting), where the receiver’s background knowledge is confined to a set of time stamps and time generalization suffices, and (ii) sharing the data with colluding receivers (the SCR setting), where time generalization should be combined with event generalization. For both cases, we propose appropriate anonymization methods that prevent both user identification and event prediction. To achieve computational efficiency and scalability, we propose optimization techniques for both cases using a utility-based index, compact summaries, fast to compute bounds for utility, and a novel taxonomy-aware distance function. Extensive experiments confirm the effectiveness of our approach compared with state of the art, in terms of information loss, range query distortion, and preserving temporal causality patterns. Furthermore, our experiments demonstrate efficiency and scalability on large-scale real and synthetic datasets. Reza Sherkat, Jing Li 0041, Nikos Mamoulis |
ACM Trans. Web | 1 |
| 2009 | Efficient Index Compression in DB2 LUWabstractIn database systems, the cost of data storage and retrieval are important components of the total cost and response time of the system. A popular mechanism to reduce the storage footprint is by compressing the data residing in tables and indexes. Compressing indexes efficiently, while maintaining response time requirements, is known to be challenging. This is especially true when designing for a workload spectrum covering both data warehousing and transaction processing environments. DB2 Linux, UNIX, Windows (LUW) recently introduced index compression for use in both environments. This uses techniques that are able to compress index data efficiently while incurring virtually no performance penalty for query processing. On the contrary, for certain operations, the performance is actually better. In this paper, we detail the design of index compression in DB2 LUW and discuss the challenges that were encountered in meeting the design goals. We also demonstrate its effectiveness by showing performance results on typical customer scenarios. Bishwaranjan Bhattacharjee, Lipyeow Lim, Timothy Malkemus, George A. Mihaila, Kenneth A. Ross, Sherman Lau, Cathy McCarthur, Zoltan Toth, Reza Sherkat |
Proc. VLDB Endow. | 9 |
| 2008 | On efficiently searching trajectories and archival data for historical similaritiesabstractWe study the problem of efficiently evaluating similarity queries on histories, where a history is a d -dimensional time series for d ≥ 1. While there are some solutions for time-series and spatio-temporal trajectories where typically d ≤ 3, we are not aware of any work that examines the problem for larger values of d. In this paper, we address the problem in its general case and propose a class of summaries for histories with a few interesting properties. First, for commonly used distance functions such as the L p -norm, LCSS, and DTW, the summaries can be used to efficiently prune some of the histories that cannot be in the answer set of the queries. Second, histories can be indexed based on their summaries, hence the qualifying candidates can be efficiently retrieved. To further reduce the number of unnecessary distance computations for false positives, we propose a finer level approximation of histories, and an algorithm to find an approximation with the least maximum distance estimation error. Experimental results confirm that the combination of our feature extraction approaches and the indexability of our summaries can improve upon existing methods and scales up for larger values of d and database sizes, based on our experiments on real and synthetic datasets of 17-dimensional histories. Reza Sherkat, Davood Rafiei |
Proc. VLDB Endow. | 1 |
| 2007 | On MBR Approximation of Histories for Historical Queries: Expectations and LimitationsabstractTraditional approaches for efficiently processing historical queries, where a history is a multidimensional time-series, employ a two step filter-and-refine scheme. In the filter step, an approximation of each history often as a set of minimum bounding hyper-rectangles (MBRs) is organized using a spatial index structure such as R-tree. The index is used to prune redundant disk accesses and to reduce the number of pairwise comparisons required in the refine step. To improve the efficiency of the filtering step, a heuristic is used to decrease the expected number of MBRs that overlap with a query, by reducing the volume of empty space indexed by the index. The heuristic selects, among all possible splitting schemes of a history, the one which results to a set of MBRs with minimum total volume. Although this heuristic is expected to improve the performance of spatial and history based queries with small temporal and spatial extents, in many real settings, the performance of historical queries depends on the extent of the query. Moreover, the optimal approximation of a history is not always the one with minimum total volume. In this paper, we present the limitations of using volume as a criteria for approximating histories, specially in high dimensional cases, where it is not feasible to index MBRs by traditional spatial index structures. Reza Sherkat, Davood Rafiei |
ICDE | 1 |
| 2006 | Efficiently Evaluating Order Preserving Similarity Queries over Historical Market-Basket DataabstractWe introduce a new domain-independent framework for formulating and efficiently evaluating similarity queries over historical data, where given a history as a sequence of timestamped observations and the pair-wise similarity of observations, we want to find similar histories. For instance, given a database of customer transactions and a time period, we can find customers with similar purchasing behaviors over this period. Our work is different from the work on retrieving similar time series; it addresses the general problem in which a history cannot be modeled as a time series, hence the relevant conventional approaches are not applicable. We derive a similarity measure for histories, based on an aggregation of the similarities between the observations of the two histories, and propose efficient algorithms for finding an optimal alignment between two histories. Given the non-metric nature of our measure, we develop some upper bounds and an algorithm that makes use of those bounds to prune histories that are guaranteed not to be in the answer set. Our experimental results on real and synthetic data confirm the effectiveness and efficiency of our approach. For instance, when the minimum length of a match is provided, our algorithm achieves up to an order of magnitude speed-up over alternative methods. Reza Sherkat, Davood Rafiei |
ICDE | 1 |