Abhirup Chakraborty

dblp:58/7437 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
0since 2021 · last 2016
0000-0001-7252-3175ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 first-authorArtificial intelligence and machine learning · 4 · 3 first-authorSystems, architecture and hardware · 4 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Database system architecture and tuning · 50% Query processing and optimization · 50%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 59% Cloud and datacenter computing · 41%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 3 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
approximate nearest neighbor search
0.112010
SONNET: Efficient Approximate Nearest Neighbor Using Multi-core · ICDM 2010
Algorithms and data structures › similarity search
high-dimensional similarity search
0.012010
SONNET: Efficient Approximate Nearest Neighbor Using Multi-core · ICDM 2010
Algorithms and data structures › similarity search
nearest neighbor search
0.012010
SONNET: Efficient Approximate Nearest Neighbor Using Multi-core · ICDM 2010

Methods — techniques the papers use, named apart from their topics

query planning · 0.5mapreduce · 0.5graph processing · 0.5rank aggregation · 0.2multicore parallelization · 0.1multi-core parallelization · 0.1
YearPublicationVenuePosition
2016 SQL-SA for big data discovery polymorphic and parallelizable SQL user-defined scalar and aggregate infrastructure in Teradata Aster 6.20
abstract
There is increasing demand to integrate big data analytic systems using SQL. Given the vast ecosystem of SQL applications, enabling SQL capabilities allows big data platforms to expose their analytic potential to a wide variety of end users, accelerating discovery processes and providing significant business value. Most existing big data frameworks are based on one particular programming model such as MapReduce or Graph. However, data scientists are often forced to manually create adhoc data pipelines to connect various big data tools and platforms to serve their analytic needs. When the analytic tasks change, these data pipelines may be costly to modify and maintain. In this paper we present SQL-SA, a polymorphic and parallelizable SQL scalar and aggregate infrastructure in Aster 6.20. This infrastructure extends Aster 6's MapReduce and Graph capabilities to support polymorphic user-defined scalar and aggregate functions using flexible SQL syntax. The implementation enhances main Aster components including query syntax, API, planning and execution extensively. Integrating these new user-defined scalar and aggregate functions with Aster MapReduce and Graph functions, Aster 6.20 enables data scientists to integrate diverse programming models in a single SQL statement. The statement is automatically converted to an optimal data pipeline and executed in parallel. Using a real world business problem and data, Aster 6.20 demonstrates a significant performance advantage (25%+) over Hadoop Pig and Hive.
Robert M. Wehrmeister, James Shau, Abhirup Chakraborty, Daley Alex, Awny Al Omari, Feven Atnafu, Jeff Davis, Litao Deng, Deepak Jaiswal, Chittaranjan Keswani, Yafeng Lu, Tom Reyes, Kashif Siddiqui, David E. Simmen, Devendra Vidhani, Daniel Yu
ICDE4
2013 Top-K aggregation over a large graph using shared-nothing systems
abstract
Analyzing large graphs is crucial to a variety of application domains, like personalized recommendations in social networks, search engines, communication networks, computational biology, etc. In these domains, there is a need to process aggregation queries over large graphs. Existing approaches for aggregation are not suitable for large graphs, as they involve multi-way relational joins over gigantic tables or repeated multiplications of large matrices. In this paper, we consider top-K aggregation queries that involve identifying top-K nodes with highest aggregate values over their h-hop neighbors. We propose algorithms for processing such queries over large graphs in a shared nothing environment. Using the notion of graph partitioning, we propose an update-based algorithm that minimizes network overhead by propagating updates in the neighborhood information. The algorithm partitions a graph across a number of processing nodes, and uses an iterative join algorithm within each node. We present a hybrid scheme to further reduce the network overhead during a few initial iterations. We develop a baseline algorithm based on distributed joins. Our experimental results validate the effectiveness of the proposed algorithms in reducing the aggregation time and in scaling the aggregation computation over a number of distributed hosts.
Abhirup Chakraborty
IEEE BigData1
2013 Parallelizing windowed stream joins in a shared-nothing cluster
abstract
The availability of a large number of processing nodes in a parallel and distributed computing environment enables sophisticated real time processing over high speed data streams, as required by many emerging applications. Sliding window stream joins are among the most important operators in a stream processing system. In this paper, we consider the issue of parallelizing a sliding window stream join operator over a shared nothing cluster. We propose a Bulk Synchronous Processing (BSP) framework, based on a fixed or predefined sequence of communication, to distribute the join processing loads over a shared-nothing cluster. We consider various processing and communication overheads while scaling over a large number of nodes, and propose solution methodologies to cope with the issues.We implement the algorithm over a cluster using a message passing system, and present the experimental results showing the effectiveness of the join processing algorithm.
Abhirup Chakraborty, Ajit Singh
CLUSTER1
2012 Switching Optically-Connected Memories in a Large-Scale System
abstract
Recent trends in processor and memory systems in large-scale computing systems reveal a new "memory wall" that prompts investigation on alternate main memory organization separating main memory from processors and arranging them in separate ensembles. In this paper, we study the feasibility of transferring data across processors by using the optical interconnection fabric that acts as a bridge between processor and memory ensembles. We propose a memory switching protocol that transfers data across processors without physically moving the data across electrical switches. Such a mechanism allows large-scale data communication across processors through transfer of a few tiny blocks of meta-data. We present detailed techniques for supporting two communication patterns prevalent in any large-scale scientific and data management applications. We present experimental results analyzing the feasibility of memory switching in a wide range of applications, and characterize applications based on the impact of the memory switching on their performance.
Abhirup Chakraborty, Eugen Schenfeld, Dilma Da Silva
IPDPS1
2011 Cost-aware caching schemes in heterogeneous storage systems
Abhirup Chakraborty, Ajit Singh
J. Supercomput.1
2010 A Disk-Based, Adaptive Approach to Memory-Limited Computation of Windowed Stream Joins
Abhirup Chakraborty, Ajit Singh
DEXA (1)1
2010 SONNET: Efficient Approximate Nearest Neighbor Using Multi-core
abstract
Approximate Nearest Neighbor search over high dimensional data is an important problem with a wide range of practical applications. In this paper, we propose SONNET, a simple multi-core friendly approximate nearest neighbor algorithm that is based on rank aggregation. SONNET is particularly suitable for very high dimensional data, its performance gets better as the dimension increases, whereas the majority of the existing algorithms show a reverse trend. Furthermore, most of the existing algorithms are hard to parallelize either due to the sequential nature of the algorithm or due to the inherent complexity of the algorithm. On the other hand, SONNET has inherent parallelism embedded in the core concept of the algorithm, which earns it almost a linear speed-up as the number of cores increases. Finally, SONNET is very easy to implement and it has an approximation parameter which is intuitively simple.
Mohammad Al Hasan, Hilmi Yildirim, Abhirup Chakraborty
ICDM3
2009 Processing Exact Results for Sliding Window Joins over Time-Sequence, Streaming Data Using a Disk Archive
abstract
We consider the problem of processing exact results for sliding window joins over data streams with limited memory. Existing approaches deal with memory limitations by shedding loads, and therefore cannot provide exact or even highly accurate results for sliding window joins over data streams showing time varying rate of data arrivals. We provide an exact window join (EWJ) algorithm incorporating disk storage as an archive. Our algorithm spills window data onto the disk on a periodic basis, refines the output result by properly retrieving the disk resident data, and maximizes output rate by employing techniques to manage the memory blocks. The problem of managing the window blocks in memory-similar in nature to the caching issue-captures both the temporal and frequency related properties of the stream arrivals. At the same, we improve I/O efficiency by amortizing a disk scan over a large number of input tuple. We provide experimental results demonstrating the performance and effectiveness of the proposed algorithm.
Abhirup Chakraborty, Ajit Singh
ACIIDS1
2009 A partition-based approach to support streaming updates over persistent data in an active datawarehouse
abstract
Active warehousing has emerged in order to meet the high user demands for fresh and up-to-date information. Online refreshment of the source updates introduces processing and disk overheads in the implementation of the warehouse transformations. This paper considers a frequently occurring operator in active warehousing which computes the join between a fast, time varying or bursty update stream S and a persistent disk relation R, using a limited memory. Such a join operation is the crux of a number of common transformations (e.g., surrogate key assignment, duplicate detection etc) in an active data warehouse. We propose a partition-based join algorithm that minimizes the processing overhead, disk overhead and the delay in output tuples. The proposed algorithm exploits the spatio-temporal locality within the update stream, and improves the delays in output tuples by exploiting hot-spots in the range or domain of the joining attributes, and at the same time shares the I/O cost of accessing disk data of relation R over a volume of tuples from update stream S. We present experimental results showing the effectiveness of the proposed algorithm.
Abhirup Chakraborty, Ajit Singh
IPDPS1