Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sam Idicula

dblp:98/7258 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
0since 2021 · last 2020
0009-0009-4562-0557ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Query processing and optimization · 46% Indexing and storage engines · 23% Database system architecture and tuning · 13%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 40% Hardware accelerators and domain-specific architectures · 27% Cloud and datacenter computing · 20%
Artificial intelligence
1 paper
Efficient and distributed learning · 77% Learning theory · 23%

Topics — the 15 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
automated machine learning
0.412020
Oracle AutoML: A Fast and Predictive AutoML Pipeline · Proc. VLDB Endow. 2020
Indexing and storage engines
columnar storage
0.312018
RAPID: In-Memory Analytical Query Processing Engine with Extreme Performance per Watt · SIGMOD Conference 2018
Memory systems
data movement
0.312017
A many-core architecture for in-memory data processing · MICRO 2017
Cloud and datacenter computing
data movement acceleration
0.312017
A many-core architecture for in-memory data processing · MICRO 2017
Hardware accelerators and domain-specific architectures › domain-specific accelerator
data processing accelerator
0.312017
A many-core architecture for in-memory data processing · MICRO 2017
Memory systems › in-memory computing
in-memory data processing
0.312017
A many-core architecture for in-memory data processing · MICRO 2017
Query processing and optimization › join processing
distributed join
0.212016
Flow-Join: Adaptive skew handling for distributed joins over high-speed networks · ICDE 2016
Distributed and cloud data management
distributed query processing
0.212016
Flow-Join: Adaptive skew handling for distributed joins over high-speed networks · ICDE 2016
Database system architecture and tuning
parallel database system
0.212016
Flow-Join: Adaptive skew handling for distributed joins over high-speed networks · ICDE 2016
Query processing and optimization › parallel query processing
skew handling
0.212016
Flow-Join: Adaptive skew handling for distributed joins over high-speed networks · ICDE 2016
Machine learning › Learning theory
model selection
0.112020
Oracle AutoML: A Fast and Predictive AutoML Pipeline · Proc. VLDB Endow. 2020
Data models and query languages
XML data management
0.112009
Binary XML Storage and Query Processing in Oracle 11g · Proc. VLDB Endow. 2009
Processor architecture and microarchitecture
many-core architecture
0.112017
A many-core architecture for in-memory data processing · MICRO 2017
Datacenter networks
RDMA
0.112016
Flow-Join: Adaptive skew handling for distributed joins over high-speed networks · ICDE 2016
Query processing and optimization
XML query processing
0.012009
Binary XML Storage and Query Processing in Oracle 11g · Proc. VLDB Endow. 2009

Methods — techniques the papers use, named apart from their topics

hardware-software co-design · 0.7runtime load balancing · 0.5approximate histograms · 0.5proxy model prediction · 0.4meta-learning · 0.4hardware RPC · 0.3securefiles · 0.1lightweight navigational index · 0.1NFA-based navigational algorithm · 0.1
YearPublicationVenuePosition
2020 Oracle AutoML: A Fast and Predictive AutoML Pipeline
abstract
Machine learning (ML) is at the forefront of the rising popularity of data-driven software applications. The resulting rapid proliferation of ML technology, explosive data growth, and shortage of data science expertise have caused the industry to face increasingly challenging demands to keep up with fast-paced develop-and-deploy model lifecycles. Recent academic and industrial research efforts have started to address this problem through automated machine learning (AutoML) pipelines and have focused on model performance as the first-order design objective. We present Oracle AutoML, a novel iteration-free AutoML pipeline designed to not only provide accurate models, but also in a shorter runtime. We are able to achieve these objectives by eliminating the need to continuously iterate over various pipeline configurations. In our feed-forward approach, each pipeline stage makes decisions based on metalearned proxy models that can predict candidate pipeline configuration performances before building the full final model. Our approach, which builds and tunes only the best candidate pipeline, achieves better scores at a fraction of the time compared to state-of-the-art open source AutoML tools, such as H2O and Auto-sklearn. This makes Oracle AutoML a prime candidate for addressing current industry challenges.
Anatoly Yakovlev, Hesam Fathi Moghadam, Ali Moharrer, Jingxiao Cai, Nikan Chavoshi, Venkatanathan Varadarajan, Sandeep R. Agrawal, Tomas Karnagel, Sam Idicula, Sanjay Jinturkar, Nipun Agarwal
Proc. VLDB Endow.9
2018 RAPID: In-Memory Analytical Query Processing Engine with Extreme Performance per Watt
abstract
Today, an ever increasing amount of transistors are packed into processor designs with extra features to support a broad range of applications. As a consequence, processors are becoming more and more complex and power hungry. At the same time, they only sustain an average performance for a wide variety of applications while not providing the best performance for specific applications. In this paper, we demonstrate through a carefully designed modern data processing system called RAPID and a simple, low-power processor specially tailored for data processing that at least an order of magnitude performance/power improvement in SQL processing can be achieved over a modern system running on today's complex processors. RAPID is designed from the ground up with hardware/software co-design in mind to provide architecture-conscious extreme performance while consuming less power in comparison to the modern database systems. The paper presents in detail the design and implementation of RAPID, a relational, columnar, in-memory query processing engine supporting analytical query workloads.
Cagri Balkesen, Nitin Kunal, Georgios Giannikis, Pit Fender, Seema Sundara, Felix Schmidt, Jarod Wen, Sandeep R. Agrawal, Arun Raghavan, Venkatanathan Varadarajan, Anand Viswanathan, Balakrishnan Chandrasekaran 0003, Sam Idicula, Nipun Agarwal, Eric Sedlar
SIGMOD Conference13
2017 A many-core architecture for in-memory data processing
abstract
For many years, the highest energy cost in processing has been data movement rather than computation, and energy is the limiting factor in processor design [21]. As the data needed for a single application grows to exabytes [56], there is clearly an opportunity to design a bandwidth-optimized architecture for big data computation by specializing hardware for data movement. We present the Data Processing Unit or DPU, a shared memory many-core that is specifically designed for high bandwidth analytics workloads. The DPU contains a unique Data Movement System (DMS), which provides hardware acceleration for data movement and partitioning operations at the memory controller that is sufficient to keep up with DDR bandwidth. The DPU also provides acceleration for core to core communication via a unique hardware RPC mechanism called the Atomic Transaction Engine. Comparison of a DPU chip fabricated in 40nm with a Xeon processor on a variety of data processing applications shows a 3× - 15× performance per watt advantage.
Sandeep R. Agrawal, Sam Idicula, Arun Raghavan, Evangelos Vlachos, Venkatraman Govindaraju, Venkatanathan Varadarajan, Cagri Balkesen, Georgios Giannikis, Charlie Roth, Nipun Agarwal, Eric Sedlar
MICRO2
2016 Flow-Join: Adaptive skew handling for distributed joins over high-speed networks
abstract
Modern InfiniBand interconnects offer link speeds of several gigabytes per second and a remote direct memory access (RDMA) paradigm for zero-copy network communication. Both are crucial for parallel database systems to achieve scalable distributed query processing where adding a server to the cluster increases performance. However, the scalability of distributed joins is threatened by unexpected data characteristics: Skew can cause a severe load imbalance such that a single server has to process a much larger part of the input than its fair share and by this slows down the entire distributed query. We introduce Flow-Join, a novel distributed join algorithm that handles attribute value skew with minimal overhead. Flow-Join detects heavy hitters at runtime using small approximate histograms and adapts the redistribution scheme to resolve load imbalances before they impact the join performance. Previous approaches often involve expensive analysis phases, which slow down distributed join processing for non-skewed workloads. This is especially the case for modern high-speed interconnects, which are too fast to hide the extra computation. Other skew handling approaches require detailed statistics, which are often not available or overly inaccurate for intermediate results. In contrast, Flow-Join uses our novel lightweight skew handling scheme to execute at the full network speed of more than 6 GB/s for InfiniBand 4×FDR, joining a skewed input at 11.5 billion tuples/s with 32 servers. This is 6.8× faster than a standard distributed hash join using the same hardware. At the same time, Flow-Join does not compromise the join performance for non-skewed workloads.
Wolf Rödiger, Sam Idicula, Alfons Kemper, Thomas Neumann 0001
ICDE2
2009 Binary XML Storage and Query Processing in Oracle 11g
abstract
Oracle RDBMS has supported XML data management for more than six years since version 9i. Prior to 11g, text-centric XML documents can be stored as-is in a CLOB column and schema-based data-centric documents can be shredded and stored in object-relational (OR) tables mapped from their XML Schema. However, both storage formats have intrinsic limitations---XML/CLOB has unacceptable query and update performance, and XML/OR requires XML schema. To tackle this problem, Oracle 11g introduces a native Binary XML storage format and a complete stack of data management operations. Binary XML was designed to address a wide range of real application problems encountered in XML data management---schema flexibility, amenability to XML indexes, update performance, schema evolution, just to name a few. In this paper, we introduce the Binary XML storage format based on Oracle SecureFiles System[21]. We propose a lightweight navigational index on top of the storage and an NFA-based navigational algorithm to provide efficient streaming processing. We further optimize query processing by exploiting XML structural and schema information that are collected in database dictionary. We conducted extensive experiments to demonstrate high performance of the native Binary XML in query processing, update, and space consumption.
Ning Zhang 0002, Nipun Agarwal, Sivasankaran Chandrasekar, Sam Idicula, Vijay Medi, Sabina Petride, Balasubramanyam Sthanikam
Proc. VLDB Endow.4