EDBT 2026 Demo / reviewers in the wild / expert
Rajkumar Sen
dblp:40/2065
· DBLP profile ↗
6ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
5 papers |
Query processing and optimization · 75% Database system architecture and tuning · 18% Indexing and storage engines · 7% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Embedded and real-time systems · 100% |
Topics — the 9 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization › query optimization
distributed query optimization |
0.2 | 1 | 2016 | The MemSQL Query Optimizer: A modern optimizer for real-time analytics in a distributed database · Proc. VLDB Endow. 2016 |
Database system architecture and tuning
hybrid transactional and analytical processing |
0.2 | 1 | 2016 | Operational Analytics Data Management Systems · Proc. VLDB Endow. 2016 |
Query processing and optimization
query optimization |
0.2 | 1 | 2016 | The MemSQL Query Optimizer: A modern optimizer for real-time analytics in a distributed database · Proc. VLDB Endow. 2016 |
Query processing and optimization › join processing
distributed join |
0.2 | 1 | 2014 | Track join: distributed joins with minimal network traffic · SIGMOD Conference 2014 |
Query processing and optimization › join processing
join algorithms |
0.2 | 1 | 2014 | Track join: distributed joins with minimal network traffic · SIGMOD Conference 2014 |
Query processing and optimization › query optimization › join ordering
join optimization |
0.2 | 1 | 2014 | Of Snowstorms and Bushy Trees · Proc. VLDB Endow. 2014 |
Query processing and optimization › query rewriting
query transformation |
0.2 | 1 | 2014 | Of Snowstorms and Bushy Trees · Proc. VLDB Endow. 2014 |
Indexing and storage engines
access methods |
0.1 | 1 | 2016 | Operational Analytics Data Management Systems · Proc. VLDB Endow. 2016 |
Embedded and real-time systems
resource-constrained computing |
0.0 | 1 | 2005 | Efficient Data Management on Lightweight Computing Device · ICDE 2005 |
Methods — techniques the papers use, named apart from their topics
query rewrite · 0.2join enumeration · 0.2bushy join · 0.2transfer scheduling · 0.2logical query transformation · 0.2join permutation search · 0.2hash join · 0.2heuristic optimization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Operational Analytics Data Management SystemsabstractPrior to mid-2000s, the space of data analytics was mainly confined within the area of decision support systems . It was a long era of isolated enterprise data ware houses curating information from live data sources and of business intelligence software used to query such information. Most data sets were small enough in volume and static enough invelocity to be segregated in warehouses for analysis. Data analysis was not ad-hoc; it required pre-requisite knowledge of underlying data access patterns for the creation of specialized access methods (e.g. covering indexes, materialized views) in order to efficiently execute a set of few focused queries. Alexander Böhm 0002, Jens Dittrich, Niloy Mukherjee, Ippokrantis Pandis, Rajkumar Sen |
Proc. VLDB Endow. | 5 |
| 2016 | The MemSQL Query Optimizer: A modern optimizer for real-time analytics in a distributed databaseabstractReal-time analytics on massive datasets has become a very common need in many enterprises. These applications require not only rapid data ingest, but also quick answers to analytical queries operating on the latest data. MemSQL is a distributed SQL database designed to exploit memory-optimized, scale-out architecture to enable real-time transactional and analytical workloads which are fast, highly concurrent, and extremely scalable. Many analytical queries in MemSQL's customer workloads are complex queries involving joins, aggregations, sub-queries, etc. over star and snowflake schemas, often ad-hoc or produced interactively by business intelligence tools. These queries often require latencies of seconds or less, and therefore require the optimizer to not only produce a high quality distributed execution plan, but also produce it fast enough so that optimization time does not become a bottleneck. In this paper, we describe the architecture of the MemSQL Query Optimizer and the design choices and innovations which enable it quickly produce highly efficient execution plans for complex distributed queries. We discuss how query rewrite decisions oblivious of distribution cost can lead to poor distributed execution plans, and argue that to choose high-quality plans in a distributed database, the optimizer needs to be distribution-aware in choosing join plans, applying query rewrites, and costing plans. We discuss methods to make join enumeration faster and more effective, such as a rewrite-based approach to exploit bushy joins in queries involving multiple star schemas without sacrificing optimization time. We demonstrate the effectiveness of the MemSQL optimizer over queries from the TPC-H benchmark and a real customer workload. Jack Chen, Samir Jindel, Robert Walzer, Rajkumar Sen, Nika Jimsheleishvilli, Michael Andrews |
Proc. VLDB Endow. | 4 |
| 2014 | Track join: distributed joins with minimal network trafficabstractNetwork communication is the slowest component of many operators in distributed parallel databases deployed for large-scale analytics. Whereas considerable work has focused on speeding up databases on modern hardware, communication reduction has received less attention. Existing parallel DBMSs rely on algorithms designed for disks with minor modifications for networks. A more complicated algorithm may burden the CPUs, but could avoid redundant transfers of tuples across the network. We introduce track join, a novel distributed join algorithm that minimizes network traffic by generating an optimal transfer schedule for each distinct join key. Track join extends the trade-off options between CPU and network. Our evaluation based on real and synthetic data shows that track join adapts to diverse cases and degrees of locality. Considering both network traffic and execution time, even with no locality, track join outperforms hash join on the most expensive queries of real workloads. Orestis Polychroniou, Rajkumar Sen, Kenneth A. Ross |
SIGMOD Conference | 2 |
| 2014 | Of Snowstorms and Bushy TreesabstractMany workloads for analytical processing in commercial RDBMSs are dominated by snowstorm queries, which are characterized by references to multiple large fact tables and their associated smaller dimension tables. This paper describes a technique for bushy join tree optimization for snowstorm queries in Oracle database system. This technique generates bushy join trees containing subtrees that produce substantially reduced sets of rows and, therefore, their joins with other subtrees are generally much more efficient than joins in the left-deep trees. The generation of bushy join trees within an existing commercial physical optimizer requires extensive changes to the optimizer. Further, the optimizer will have to consider a large join permutation search space to generate efficient bushy join trees. The novelty of the approach is that bushy join trees can be generated outside the physical optimizer using logical query transformation that explores a considerably pruned search space. The paper describes an algorithm for generating optimal bushy join trees for snowstorm queries using an existing query transformation framework. It also presents performance results for this optimization, which show significant execution time improvements. Rafi Ahmed, Rajkumar Sen, Meikel Pöss, Sunil Chakkappen |
Proc. VLDB Endow. | 2 |
| 2005 | Efficient Data Management on Lightweight Computing DeviceabstractLightweight computing devices are becoming ubiquitous and an increasing number of applications are being developed for these devices. Many of these applications deal with significant amounts of data and involve complex joins and aggregate operations, which necessitate a local database management system on the device. This is a challenge as these devices are constrained by limited stable storage and main memory. Hence new storage models that reduce storage costs are needed and a storage scheme should be selected based on data characteristics, nature of queries, and updates. Also, query execution plan should be chosen depending on the amount of available memory and the underlying storage scheme; memory should be optimally allocated among the database operators involved in the query. To achieve these goals, we utilize a novel storage model, ID based storage, which reduces storage costs considerably. We present an exact algorithm for allocating memory among the database operators. Because of its high complexity, we also propose a heuristic solution based on the benefit of an operator per unit memory allocation. Rajkumar Sen, Krithi Ramamritham |
ICDE | 1 |
| 2004 | DELite: database support for embedded lightweight devicesabstractComputation platforms have extended to small intelligent devices like cellphones, sensors, smartcards, PDAs, etc. As new functionalities and features are being added to these devices, increasing number of applications are being developed, many of them dealing with significant amounts of data leading to the need of embedded database support on these devices [1]. The queries go beyond simple Select-Project-Join queries but still have to be locally executed on the device [4]. Most of the modern day cellphones are being equipped with increasing memory which means more data centric applications are being developed for them. Sensor networks are also proliferating and these collect data from the environment and subject them to various queries. Most of these queries need to be executed on the device itself to reduce communication costs. Applications for PDAs execute complicated join and aggregate queries on the device resident data. Thus, there is an increasing need to facilitate the execution of complex queries locally on a variety of lightweight computing devices. However, scaling down the database footprint poses challenges since these lightweight devices come with very limited computing resources. While the amount of main memory and stable storage available in such devices is relatively small, the devices are not uniformly endowed with resources. For example, the computing capability and main memory of a cellphone differs from that of a PDA. It is essential that the available resources be utilized optimally for a database system that is developed for such devices. The already limited stable storage has to accommodate the operating system as well as the database system code, which means even less storage is available to store the data. Storage Models designed for such database systems should reduce storage cost to a minimum to be able to store more data. Limited stable storage usually precludes the creation and use of any additional index structures, hence the storage models should try to incorporate some index information in the data model itself. Ideally, index structures which can speed up query processing at no additional storage cost should be maintained. Different storage models have different storage and update costs. The selection of the best storage model for a data attribute in a relation depends on the size of the relation, selectivity and length of the attribute, frequency of updates, and the nature of queries. As far as the choice of query processing techniques is concerned, RAM is perhaps the most critical resource in these devices. Existing approaches use minimum memory algorithms for every operator. This can lead to poor performance for complex queries involving several joins and aggregates. Query execution time for complex Krithi Ramamritham, Rajkumar Sen |
EMSOFT | 2 |