Rajkumar Sen

dblp:40/2065 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Query processing and optimization · 75% Database system architecture and tuning · 18% Indexing and storage engines · 7%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Embedded and real-time systems · 100%

Topics — the 9 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › query optimization
distributed query optimization
0.212016
The MemSQL Query Optimizer: A modern optimizer for real-time analytics in a distributed database · Proc. VLDB Endow. 2016
Database system architecture and tuning
hybrid transactional and analytical processing
0.212016
Operational Analytics Data Management Systems · Proc. VLDB Endow. 2016
Query processing and optimization
query optimization
0.212016
The MemSQL Query Optimizer: A modern optimizer for real-time analytics in a distributed database · Proc. VLDB Endow. 2016
Query processing and optimization › join processing
distributed join
0.212014
Track join: distributed joins with minimal network traffic · SIGMOD Conference 2014
Query processing and optimization › join processing
join algorithms
0.212014
Track join: distributed joins with minimal network traffic · SIGMOD Conference 2014
Query processing and optimization › query optimization › join ordering
join optimization
0.212014
Of Snowstorms and Bushy Trees · Proc. VLDB Endow. 2014
Query processing and optimization › query rewriting
query transformation
0.212014
Of Snowstorms and Bushy Trees · Proc. VLDB Endow. 2014
Indexing and storage engines
access methods
0.112016
Operational Analytics Data Management Systems · Proc. VLDB Endow. 2016
Embedded and real-time systems
resource-constrained computing
0.012005
Efficient Data Management on Lightweight Computing Device · ICDE 2005

Methods — techniques the papers use, named apart from their topics

query rewrite · 0.2join enumeration · 0.2bushy join · 0.2transfer scheduling · 0.2logical query transformation · 0.2join permutation search · 0.2hash join · 0.2heuristic optimization · 0.1
YearPublicationVenuePosition
2016 Operational Analytics Data Management Systems
abstract
Prior to mid-2000s, the space of data analytics was mainly confined within the area of decision support systems . It was a long era of isolated enterprise data ware houses curating information from live data sources and of business intelligence software used to query such information. Most data sets were small enough in volume and static enough invelocity to be segregated in warehouses for analysis. Data analysis was not ad-hoc; it required pre-requisite knowledge of underlying data access patterns for the creation of specialized access methods (e.g. covering indexes, materialized views) in order to efficiently execute a set of few focused queries.
Alexander Böhm 0002, Jens Dittrich, Niloy Mukherjee, Ippokrantis Pandis, Rajkumar Sen
Proc. VLDB Endow.5
2016 The MemSQL Query Optimizer: A modern optimizer for real-time analytics in a distributed database
abstract
Real-time analytics on massive datasets has become a very common need in many enterprises. These applications require not only rapid data ingest, but also quick answers to analytical queries operating on the latest data. MemSQL is a distributed SQL database designed to exploit memory-optimized, scale-out architecture to enable real-time transactional and analytical workloads which are fast, highly concurrent, and extremely scalable. Many analytical queries in MemSQL's customer workloads are complex queries involving joins, aggregations, sub-queries, etc. over star and snowflake schemas, often ad-hoc or produced interactively by business intelligence tools. These queries often require latencies of seconds or less, and therefore require the optimizer to not only produce a high quality distributed execution plan, but also produce it fast enough so that optimization time does not become a bottleneck. In this paper, we describe the architecture of the MemSQL Query Optimizer and the design choices and innovations which enable it quickly produce highly efficient execution plans for complex distributed queries. We discuss how query rewrite decisions oblivious of distribution cost can lead to poor distributed execution plans, and argue that to choose high-quality plans in a distributed database, the optimizer needs to be distribution-aware in choosing join plans, applying query rewrites, and costing plans. We discuss methods to make join enumeration faster and more effective, such as a rewrite-based approach to exploit bushy joins in queries involving multiple star schemas without sacrificing optimization time. We demonstrate the effectiveness of the MemSQL optimizer over queries from the TPC-H benchmark and a real customer workload.
Jack Chen, Samir Jindel, Robert Walzer, Rajkumar Sen, Nika Jimsheleishvilli, Michael Andrews
Proc. VLDB Endow.4
2014 Track join: distributed joins with minimal network traffic
abstract
Network communication is the slowest component of many operators in distributed parallel databases deployed for large-scale analytics. Whereas considerable work has focused on speeding up databases on modern hardware, communication reduction has received less attention. Existing parallel DBMSs rely on algorithms designed for disks with minor modifications for networks. A more complicated algorithm may burden the CPUs, but could avoid redundant transfers of tuples across the network. We introduce track join, a novel distributed join algorithm that minimizes network traffic by generating an optimal transfer schedule for each distinct join key. Track join extends the trade-off options between CPU and network. Our evaluation based on real and synthetic data shows that track join adapts to diverse cases and degrees of locality. Considering both network traffic and execution time, even with no locality, track join outperforms hash join on the most expensive queries of real workloads.
Orestis Polychroniou, Rajkumar Sen, Kenneth A. Ross
SIGMOD Conference2
2014 Of Snowstorms and Bushy Trees
abstract
Many workloads for analytical processing in commercial RDBMSs are dominated by snowstorm queries, which are characterized by references to multiple large fact tables and their associated smaller dimension tables. This paper describes a technique for bushy join tree optimization for snowstorm queries in Oracle database system. This technique generates bushy join trees containing subtrees that produce substantially reduced sets of rows and, therefore, their joins with other subtrees are generally much more efficient than joins in the left-deep trees. The generation of bushy join trees within an existing commercial physical optimizer requires extensive changes to the optimizer. Further, the optimizer will have to consider a large join permutation search space to generate efficient bushy join trees. The novelty of the approach is that bushy join trees can be generated outside the physical optimizer using logical query transformation that explores a considerably pruned search space. The paper describes an algorithm for generating optimal bushy join trees for snowstorm queries using an existing query transformation framework. It also presents performance results for this optimization, which show significant execution time improvements.
Rafi Ahmed, Rajkumar Sen, Meikel Pöss, Sunil Chakkappen
Proc. VLDB Endow.2
2005 Efficient Data Management on Lightweight Computing Device
abstract
Lightweight computing devices are becoming ubiquitous and an increasing number of applications are being developed for these devices. Many of these applications deal with significant amounts of data and involve complex joins and aggregate operations, which necessitate a local database management system on the device. This is a challenge as these devices are constrained by limited stable storage and main memory. Hence new storage models that reduce storage costs are needed and a storage scheme should be selected based on data characteristics, nature of queries, and updates. Also, query execution plan should be chosen depending on the amount of available memory and the underlying storage scheme; memory should be optimally allocated among the database operators involved in the query. To achieve these goals, we utilize a novel storage model, ID based storage, which reduces storage costs considerably. We present an exact algorithm for allocating memory among the database operators. Because of its high complexity, we also propose a heuristic solution based on the benefit of an operator per unit memory allocation.
Rajkumar Sen, Krithi Ramamritham
ICDE1
2004 DELite: database support for embedded lightweight devices
abstract
Computation platforms have extended to small intelligent devices like cellphones, sensors, smartcards, PDAs, etc. As new functionalities and features are being added to these devices, increasing number of applications are being developed, many of them dealing with significant amounts of data leading to the need of embedded database support on these devices [1]. The queries go beyond simple Select-Project-Join queries but still have to be locally executed on the device [4]. Most of the modern day cellphones are being equipped with increasing memory which means more data centric applications are being developed for them. Sensor networks are also proliferating and these collect data from the environment and subject them to various queries. Most of these queries need to be executed on the device itself to reduce communication costs. Applications for PDAs execute complicated join and aggregate queries on the device resident data. Thus, there is an increasing need to facilitate the execution of complex queries locally on a variety of lightweight computing devices. However, scaling down the database footprint poses challenges since these lightweight devices come with very limited computing resources. While the amount of main memory and stable storage available in such devices is relatively small, the devices are not uniformly endowed with resources. For example, the computing capability and main memory of a cellphone differs from that of a PDA. It is essential that the available resources be utilized optimally for a database system that is developed for such devices. The already limited stable storage has to accommodate the operating system as well as the database system code, which means even less storage is available to store the data. Storage Models designed for such database systems should reduce storage cost to a minimum to be able to store more data. Limited stable storage usually precludes the creation and use of any additional index structures, hence the storage models should try to incorporate some index information in the data model itself. Ideally, index structures which can speed up query processing at no additional storage cost should be maintained. Different storage models have different storage and update costs. The selection of the best storage model for a data attribute in a relation depends on the size of the relation, selectivity and length of the attribute, frequency of updates, and the nature of queries. As far as the choice of query processing techniques is concerned, RAM is perhaps the most critical resource in these devices. Existing approaches use minimum memory algorithms for every operator. This can lead to poor performance for complex queries involving several joins and aggregates. Query execution time for complex
Krithi Ramamritham, Rajkumar Sen
EMSOFT2