EDBT 2026 Demo / reviewers in the wild / expert
Hemant Saxena
dblp:167/2106
· DBLP profile ↗
5ranked-venue papers
4as first author
1since 2021 · last 2023
0000-0003-0412-7490ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Data integration and cleaning · 40% Data mining · 36% Database system architecture and tuning · 21% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 100% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › key-value storage
LSM-tree |
0.7 | 1 | 2023 | Real-Time LSM-Trees for HTAP Workloads · ICDE 2023 |
Data mining
clustering |
0.4 | 1 | 2019 | A Semi-Supervised Framework of Clustering Selection for De-Duplication · ICDE 2019 |
Data integration and cleaning › entity resolution
deduplication |
0.4 | 1 | 2019 | A Semi-Supervised Framework of Clustering Selection for De-Duplication · ICDE 2019 |
Data integration and cleaning
dependency discovery |
0.4 | 1 | 2019 | Distributed Implementations of Dependency Discovery Algorithms · Proc. VLDB Endow. 2019 |
Data mining › big data analytics › large-scale data mining
distributed data mining |
0.4 | 1 | 2019 | Distributed Discovery of Functional Dependencies · ICDE 2019 |
Data integration and cleaning › dependency discovery
functional dependency discovery |
0.4 | 1 | 2019 | Distributed Discovery of Functional Dependencies · ICDE 2019 |
Data mining › clustering
semi-supervised clustering |
0.4 | 1 | 2019 | A Semi-Supervised Framework of Clustering Selection for De-Duplication · ICDE 2019 |
Storage systems
key-value storage |
0.2 | 1 | 2023 | Real-Time LSM-Trees for HTAP Workloads · ICDE 2023 |
Data integration and cleaning
data profiling |
0.1 | 1 | 2019 | Distributed Discovery of Functional Dependencies · ICDE 2019 |
Distributed and cloud data management
distributed query processing |
0.1 | 1 | 2019 | Distributed Implementations of Dependency Discovery Algorithms · Proc. VLDB Endow. 2019 |
Algorithms and data structures › data structure design › search structures › hashing
locality-sensitive hashing |
0.1 | 1 | 2019 | A Semi-Supervised Framework of Clustering Selection for De-Duplication · ICDE 2019 |
Algorithms and data structures › randomized algorithms
sampling |
0.1 | 1 | 2019 | A Semi-Supervised Framework of Clustering Selection for De-Duplication · ICDE 2019 |
Methods — techniques the papers use, named apart from their topics
cost model · 1.3sampling · 0.8locality-sensitive hashing · 0.8correlation clustering · 0.8experimental evaluation · 0.4case study · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Real-Time LSM-Trees for HTAP WorkloadsabstractReal-time analytics systems employ hybrid data layouts in which data are stored in different formats throughout their lifecycle. Recent data are stored in a row-oriented format to serve OLTP workloads and support high insert rates, while older data are transformed to a column-oriented format for OLAP access patterns. We observe that a Log-Structured Merge (LSM) Tree is a natural fit for a lifecycle-aware storage engine due to its high write throughput and level-oriented structure, in which records propagate from one level to the next over time. To build a lifecycle-aware storage engine using an LSM-Tree, we make a crucial modification to allow different data layouts in different levels, ranging from purely row-oriented to purely column-oriented, leading to a Real-Time LSM-Tree. We give a cost model and an algorithm to design a Real-Time LSM-Tree that is suitable for a given workload, followed by an experimental evaluation of LASER - a prototype implementation of our idea built on top of the RocksDB key-value store. Hemant Saxena, Lukasz Golab, Stratos Idreos, Ihab F. Ilyas |
ICDE | 1 |
| 2019 | A Semi-Supervised Framework of Clustering Selection for De-DuplicationabstractWe view data de-duplication as a clustering problem. Recently, [1] introduced a framework called restricted correlation clustering (RCC) to model de-duplication problems. Given a set X, an unknown target clustering C* of X and a class F of clusterings of X, the goal is to find a clustering C from the set F which minimizes the correlation loss. The clustering algorithm is allowed to interact with a domain expert by asking whether a pair of records correspond to the same entity or not. Main drawback of the algorithm developed by [1] is that the pre-processing step had a time complexity of theta (|X|2) (where X is the input set). In this paper, we make the following contributions. We develop a sampling procedure (based on locality sensitive hashing) which requires a linear pre-processing time O(|X|). We prove that our sampling procedure can estimate the correlation loss of all clusterings in F using only a small number of labelled examples. In fact, the number of labelled examples is independent of |X| and depends only on the complexity of the class F. Further we show that to sample one pair, with high probability our procedure makes a constant number of queries to the domain expert. We then perform an extensive empirical evaluation of our approach which shows the efficiency of our method. Shrinu Kushagra, Hemant Saxena, Ihab F. Ilyas, Shai Ben-David |
ICDE | 2 |
| 2019 | Distributed Discovery of Functional DependenciesabstractWe address the problem of discovering functional dependencies from distributed big data. Existing (non-distributed) algorithms such as FastFDs focus on minimizing computation. However, distributed algorithms must also optimize data communication costs, especially in shared-nothing settings. We propose a distributed version of FastFDs that is communication-efficient and we experimentally show significant performance improvements over a straightforward distributed implementation. Hemant Saxena, Lukasz Golab, Ihab F. Ilyas |
ICDE | 1 |
| 2019 | Distributed Implementations of Dependency Discovery AlgorithmsabstractWe analyze the problem of discovering dependencies from distributed big data. Existing (non-distributed) algorithms focus on minimizing computation by pruning the search space of possible dependencies. However, distributed algorithms must also optimize communication costs, especially in shared-nothing settings, leading to a more complex optimization space. To understand this space, we introduce six primitives shared by existing dependency discovery algorithms, corresponding to data processing steps separated by communication barriers. Through case studies, we show how the primitives allow us to analyze the design space and develop communication-optimized implementations. Finally, we support our analysis with an experimental evaluation on real datasets. Hemant Saxena, Lukasz Golab, Ihab F. Ilyas |
Proc. VLDB Endow. | 1 |
| 2015 | EdgeX: Edge Replication for Web ApplicationsabstractGlobal Web applications face the problem of high network latency due to their need to communicate with distant data centers. Many applications use edge networks for caching images, CSS, java script, and other static content in order to avoid some of this network latency. However, for updates and for anything other than static content, communication with the data center is still required, and can dominate application request latencies. One way to address this problem is to push more of the web application, as well the database on which it depends, from the remote data center towards the edge of the network. In this paper, we present preliminary work in this direction. Specifically, we present an edge-aware dynamic data replication architecture for relational database systems supporting Web applications. Our objective is to allow dynamic content to be served from the edge of the network, with low latency. Hemant Saxena, Kenneth Salem |
CLOUD | 1 |