Rajesh Bordawekar

dblp:28/6851 · also Rajesh R. Bordawekar · DBLP profile ↗
← Back
31ranked-venue papers
16as first author
1since 2021 · last 2022
0000-0002-8063-9220ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 9 first-authorDatabases, data management, data science and information retrieval · 14 · 7 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Machine learning and data management · 62% Web and social media mining · 10% Indexing and storage engines · 10%
Computer architecture, parallel and distributed computing, and storage systems
11 papers
Parallel and multicore computing · 41% Hardware accelerators and domain-specific architectures · 25% High-performance computing · 9%
Theoretical computer science
1 paper
Algorithms and data structures · 100%
Software engineering, system software, and programming languages
3 papers
Runtime systems and virtual machines · 41% Programming languages and type systems · 23% Compilers and program optimization · 21%

Topics — the 30 heaviest of 40, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management
AI for data management
1.022022
aiDM'22: Fifth International Workshop on Exploiting Artificial Intelligence Techniques for Data Management · SIGMOD Conference 2022
Overview of the 2nd International Workshop on Exploiting Artificial Intelligence Techniques for Data Management (aiDM'19) · SIGMOD Conference 2019
Parallel and multicore computing › parallel algorithms › sorting
parallel sorting
0.212015
PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort · Proc. VLDB Endow. 2015
Algorithms and data structures › sequence algorithms › sorting › integer sorting
radix sort
0.212015
PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort · Proc. VLDB Endow. 2015
Algorithms and data structures › sequence algorithms
sorting
0.212015
PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort · Proc. VLDB Endow. 2015
Indexing and storage engines › storage management
data layout
0.212014
Disk-Based Management of Interaction Graphs · IEEE Trans. Knowl. Data Eng. 2014
Web and social media mining › social network analysis
interaction graph
0.212014
Disk-Based Management of Interaction Graphs · IEEE Trans. Knowl. Data Eng. 2014
Processor architecture and microarchitecture › chip multiprocessor
cell processor
0.142009
CellJoin: a parallel stream join operator for the cell processor · VLDB J. 2009
On Efficient Query Processing of Stream Counts on the Cell Processor · ICDE 2009
Executing Stream Joins on the Cell Processor · VLDB 2007
Data stream processing
frequency estimation
0.112009
On Efficient Query Processing of Stream Counts on the Cell Processor · ICDE 2009
Distributed systems › concurrency control
optimistic concurrency control
0.112008
Modeling optimistic concurrency using quantitative dependence analysis · PPoPP 2008
Parallel and multicore computing
speculative parallelization
0.112008
Modeling optimistic concurrency using quantitative dependence analysis · PPoPP 2008
Hardware accelerators and domain-specific architectures › database accelerator
FPGA-based query processing
0.112016
Accelerating database workloads by software-hardware-system co-design · ICDE 2016
GPUs and heterogeneous computing
GPU query processing
0.112016
Accelerating database workloads by software-hardware-system co-design · ICDE 2016
Query processing and optimization
sorting
0.112007
CellSort: High Performance Sorting on the Cell Processor · VLDB 2007
Data stream processing
stream join
0.112007
Executing Stream Joins on the Cell Processor · VLDB 2007
Parallel and multicore computing
load balancing
0.112015
PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort · Proc. VLDB Endow. 2015
Spatial and temporal data management
temporal query processing
0.112014
Disk-Based Management of Interaction Graphs · IEEE Trans. Knowl. Data Eng. 2014
Programming languages and type systems
language design
0.112005
XJ: facilitating XML processing in Java · WWW 2005
Runtime systems and virtual machines
garbage collection
0.012002
Exploiting prolific types for memory management and optimizations · POPL 2002
Operating systems › resource management
memory management
0.012002
Exploiting prolific types for memory management and optimizations · POPL 2002
Runtime systems and virtual machines
object representation
0.012002
Exploiting prolific types for memory management and optimizations · POPL 2002
Parallel and multicore computing › data parallelism
SIMD vectorization
0.012009
On Efficient Query Processing of Stream Counts on the Cell Processor · ICDE 2009
Compilers and program optimization › compiler construction
compilation strategies
0.012000
Quicksilver: a quasi-static compiler for Java · OOPSLA 2000
Runtime systems and virtual machines › virtual machine implementation
java virtual machine
0.012000
Quicksilver: a quasi-static compiler for Java · OOPSLA 2000
Performance modeling and evaluation
workload characterization
0.012000
Quantitative Characterization and Analysis of the I/O Behavior of a Commercial Distributed-Shared-Memory Machine · IEEE Trans. Parallel Distributed Syst. 2000
Memory systems › memory access patterns
irregular memory access
0.012008
Modeling optimistic concurrency using quantitative dependence analysis · PPoPP 2008
Parallel and multicore computing
parallel programming models
0.012008
Modeling optimistic concurrency using quantitative dependence analysis · PPoPP 2008
High-performance computing
parallel i/o
0.021995
A Model and Compilation Strategy for Out-of-Core Data Parallel Programs · PPoPP 1995
Design and Evaluation of primitives for Parallel I/O · SC 1993
High-performance computing › cluster computing
beowulf cluster
0.011998
Scaling of Beowulf-Class Distributed Systems · HPDC 1998
High-performance computing
cluster computing
0.011998
Scaling of Beowulf-Class Distributed Systems · HPDC 1998
High-performance computing › cluster computing
large-scale clusters
0.011998
Scaling of Beowulf-Class Distributed Systems · HPDC 1998

Methods — techniques the papers use, named apart from their topics

SIMD · 0.7software-hardware co-design · 0.5speculative permutation · 0.4distribution-adaptive load balancing · 0.4sketch-based counting · 0.2temporal locality · 0.2spatial locality · 0.2quantitative modeling · 0.1dependence analysis · 0.1prototype compiler · 0.1benchmarking · 0.0static compilation · 0.0performance measurement · 0.0dynamic compilation · 0.0scalability analysis · 0.0
YearPublicationVenuePosition
2022 aiDM'22: Fifth International Workshop on Exploiting Artificial Intelligence Techniques for Data Management
abstract
Recent advances in AI techniques, as well as enabling hardware and infrastructure, has led to the integration of AI in wide-ranging domains and tasks. In particular, AI has been used to handle various types of data (including numerical, textual and graphical data) and has been adopted in large-scale distributed systems. From a data management perspective, this calls for the harnessing of state-of- the-art AI solutions for data management tasks and systems. aiDM is a full-day workshop that offers a stage for innovative interdisciplinary research that studies the interaction between AI and data management and develops new AI technologies for data-related tasks. This year, aiDM'22 particularly focuses on the transparent exploitation of AI techniques in existing enterprise-level data management workloads.
Rajesh Bordawekar, Yael Amsterdamer, Donatella Firmani, Ryan Marcus, Oded Shmueli
SIGMOD Conference1
2019 Exploiting Latent Information in Relational Databases via Word Embedding and Application to Degrees of Disclosure
Rajesh Bordawekar, Oded Shmueli
CIDR1
2019 Overview of the 2nd International Workshop on Exploiting Artificial Intelligence Techniques for Data Management (aiDM'19)
abstract
Recently, the Artificial Intelligence (AI) field has been experiencing a resurgence. AI broadly covers a wide swath of techniques which include logic-based approaches, probabilistic graphical models, and machine learning/deep learning approaches. Advances in hardware capabilities, such as Graphics Processing Units (GPUs), software components (e.g., accelerated libraries, programming frameworks), and systems infrastructures (e.g., GPU-enabled cloud providers) has led to a wide-spread adaptation of AI techniques to a variety of domains. Examples of such domains include image classification, autonomous driving, automatic speech recognition (ASR) and conversational systems (chatbots). AI techniques not only support multiple datatypes (e.g., free text, images, or speech), but are also available in various configurations, from personal devices to large-scale distributed systems.
Rajesh Bordawekar, Oded Shmueli
SIGMOD Conference1
2016 Accelerating database workloads by software-hardware-system co-design
abstract
The key objective of this tutorial is to provide a broad, yet an in-depth survey of the emerging field of co-designing software, hardware, and systems components for accelerating enterprise data management workloads. The overall goal of this tutorial is two-fold. First, we provide a concise system-level characterization of different types of data management technologies, namely, the relational and NoSQL databases and data stream management systems from the perspective of analytical workloads. Using the characterization, we discuss opportunities for accelerating key data management workloads using software and hardware approaches. Second, we dive deeper into the hardware acceleration opportunities using Graphics Processing Units (GPUs) and Field-Programmable Gate Arrays (FPGAs) for the query execution pipeline. Furthermore, we explore other hardware acceleration mechanisms such as single-instruction multiple-data (SIMD) that enables short-vector data parallelism.
Rajesh Bordawekar, Mohammad Sadoghi
ICDE1
2015 Efficient GPU implementation of convolutional neural networks for speech recognition
Ewout van den Berg, Daniel Brand, Rajesh Bordawekar, Leonid Rachevsky, Bhuvana Ramabhadran
INTERSPEECH3
2015 PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort
abstract
In-place radix sort is a popular distribution-based sorting algorithm for short numeric or string keys due to its linear run-time and constant memory complexity. However, efficient parallelization of in-place radix sort is very challenging for two reasons. First, the initial phase of permuting elements into buckets suffers read-write dependency inherent in its in-place nature. Secondly, load balancing of the recursive application of the algorithm to the resulting buckets is difficult when the buckets are of very different sizes, which happens for skewed distributions of the input data. In this paper, we present a novel parallel in-place radix sort algorithm, PARADIS, which addresses both problems: a) "speculative permutation" solves the first problem by assigning multiple non-continuous array stripes to each processor. The resulting shared-nothing scheme achieves full parallelization. Since our speculative permutation is not complete, it is followed by a "repair" phase, which can again be done in parallel without any data sharing among the processors. b) "distribution-adaptive load balancing" solves the second problem. We dynamically allocate processors in the context of radix sort, so as to minimize the overall completion time. Our experimental results show that PARADIS offers excellent performance/scalability on a wide range of input data sets.
Minsik Cho, Daniel Brand, Rajesh Bordawekar, Ulrich Finkler, Vincent KulandaiSamy, Ruchir Puri
Proc. VLDB Endow.3
2014 Disk-Based Management of Interaction Graphs
abstract
In our increasingly connected and instrumented world, live data recording the interactions between people, systems, and the environment is available in various domains, such as telecommunications and social media. This data often takes the form of a temporally evolving graph, where entities are the vertices and the interactions between them are the edges. An important feature of this graph is that the number of edges it has grows continuously, as new interactions take place. We call such graphs interaction graphs. In this paper we study the problem of storing interaction graphs such that temporal queries on them can be answered efficiently. Since interaction graphs are append-only and edges are added continuously, traditional graph layout and storage algorithms that are batch based cannot be applied directly. We present the design and implementation of a system that caches recent interactions in memory, while quickly placing the expired interactions to disk blocks such that those edges that are likely to be accessed together are placed together. We develop live block formation algorithms that are fast, yet can take advantage of temporal and spatial locality among the edges to optimize the storage layout with the goal of improving query performance. We evaluate the system on synthetic as well as real-world interaction graphs, and show that our block formation algorithms are effective for answering temporal neighborhood queries on the graph. Such queries form a foundation for building more complex online and offline temporal analytics on interaction graphs.
Bugra Gedik, Rajesh Bordawekar
IEEE Trans. Knowl. Data Eng.2
2010 Believe it or not!: mult-core CPUs can match GPU performance for a FLOP-intensive application!
abstract
In this paper, we evaluate performance of a real-world image processing application that uses a cross-correlation algorithm to compare a given image with a reference one. We implement this algorithm on a nVidia GTX 285 GPU using CUDA, and also parallelize it for the Intel Xeon (Nehalem) and IBM Power7 processors, using both manual and automatic techniques. Pthreads and OpenMP with SSE and VSX vector intrinsics are used for the manually parallelized version, while a state-of-the-art optimization framework based on the polyhedral model is used for automatic compiler parallelization and optimization. The best performing versions on the Power7, Nehalem, and GTX 285 run in 1.02s, 1.82s, and 1.22s, respectively. The performance of this algorithm on the nVidia GPU suffers from: (1) a smaller shared memory, (2) unaligned device memory access patterns, (3) expensive atomic operations, and (4) weaker single-thread performance. These results conclusively demonstrate that, under certain conditions, it is possible for a FLOP-intensive structured application running on a multi-core processor to match or even beat the performance of an equivalent GPU version.
Rajesh Bordawekar, Uday Bondhugula, Ravi Rao
PACT1
2010 Statistics-based parallelization of XPath queries in shared memory systems
abstract
The wide availability of commodity multi-core systems presents an opportunity to address the latency issues that have plaqued XML query processing. However, simply executing multiple XML queries over multiple cores merely addresses the throughput issue: intra-query parallelization is needed to exploit multiple processing cores for better latency. Toward this effort, this paper investigates the parallelization of individual XPath queries over shared-address space multi-core processors. Much previous work on parallelizing XPath in a distributed setting failed to exploit the shared memory parallelism of multi-core systems. We propose a novel, end-to-end parallelization framework that determines the optimal way of parallelizing an XML query. This decision is based on a statistics-based approach that relies both on the query specifics and the data statistics. At each stage of the parallelization process, we evaluate three alternative approaches, namely, data-, query-, and hybrid-partitioning. For a given XPath query, our parallelization algorithm uses XML statistics to estimate the relative efficiencies of these different alternatives and find an optimal parallel XPath processing plan. Our experiments using well-known XML documents validate our parallel cost model and optimization framework, and demonstrate that it is possible to accelerate XPath processing using commodity multi-core systems.
Rajesh Bordawekar, Lipyeow Lim, Anastasios Kementsietsidis, Bryant Wei-Lun Kok
EDBT1
2009 Parallelization of XPath queries using multi-core processors: challenges and experiences
abstract
In this study, we present experiences of parallelizing XPath queries using the Xalan XPath engine on shared-address space multi-core systems. For our evaluation, we consider a scenario where an XPath processor uses multiple threads to concurrently navigate and execute individual XPath queries on a shared XML document. Given the constraints of the XML execution and data models, we propose three strategies for parallelizing individual XPath queries: Data partitioning, Query partitioning, and Hybrid (query and data) partitioning. We experimentally evaluated these strategies on an x86 Linux multi-core system using a set of XPath queries, invoked on a variety of XML documents using the Xalan XPath APIs. Experimental results demonstrate that the proposed parallelization strategies work very effectively in practice; for a majority of XPath queries under evaluation, the execution performance scaled linearly as the number of threads was increased. Results also revealed the pros and cons of the different parallelization strategies for different XPath query patterns.
Rajesh Bordawekar, Lipyeow Lim, Oded Shmueli
EDBT1
2009 On Efficient Query Processing of Stream Counts on the Cell Processor
abstract
In recent years, the sketch-based technique has been presented as an effective method for counting stream items on processors with limited storage and processing capabilities, such as the network processors. In this paper, we examine the implementation of a sketch-based counting algorithm on the heterogeneous multi-core Cell processor. Like the network processors, the Cell also contains on-chip special processors with limited local memories. These special processors enable parallel processing of stream items using short-vector data-parallel (SIMD) operations. We demonstrate that the inaccuracies of the estimates computed by straightforward adaptations of current sketch-based counting approaches are exacerbated by increased inaccuracies in approximating counts of low frequency items, and by the inherent space limitations of the Cell processor. To address these concerns, we implement a sketch-based counting algorithm, FCM, that is specifically adapted for the Cell processor architecture. FCM incorporates novel capabilities for improving estimation accuracy using limited space by dynamically identifying low- and high-frequency stream items, and using a variable number of hash functions per item as determined by an item's current frequency phase. We experimentshortally demonstrate that with similar space consumption, FCM computes better frequency estimates of both the low- and high-frequency items than a naive parallelization of an existing stream counting algorithm. Using FCM as the kernel, our parallel algorithm is able to scale the over all performance linearly as well as improve the estimate accuracy as the number of processors is increased. Thus, this work demonstrates the importance of adapting the algorithm to the specifics of the underlying architecture.
Dina Thomas, Rajesh Bordawekar, Charu C. Aggarwal, Philip S. Yu
ICDE2
2009 Compiler and runtime techniques for software transactional memory optimization
abstract
Abstract Software transactional memory (STM) systems are an attractive environment to evaluate optimistic concurrency. We describe our experience of supporting and optimizing an STM system at both the managed runtime and compiler levels. We describe the design policies of our STM system and the statistics collected by the runtime to identify performance bottlenecks and guide tuning decisions. We present an initial work on supporting automatic instrumentation of the STM primitives for C/C++ and Java programs in the IBM XL compiler and J9 Java virtual machine. We evaluate and discuss the performance of several transactional programs running on our system. Copyright © 2008 John Wiley & Sons, Ltd.
Peng Wu 0001, Maged M. Michael, Christoph von Praun, Takuya Nakaike, Rajesh Bordawekar, Harold W. Cain, Calin Cascaval, Siddhartha Chatterjee, Stefanie Chiras, Mark F. Mergen, Michael F. Spear, Huayong Wang
Concurr. Comput. Pract. Exp.5
2009 CellJoin: a parallel stream join operator for the cell processor
Bugra Gedik, Rajesh Bordawekar, Philip S. Yu
VLDB J.2
2008 Modeling optimistic concurrency using quantitative dependence analysis
abstract
This work presents a quantitative approach to analyze parallelization opportunities in programs with irregular memory access where potential data dependencies mask available parallelism. The model captures data and causal dependencies among critical sections as algorithmic properties and quantifies them as a density computed over the number of executed instructions. The model abstracts from runtime aspects such as scheduling, the number of threads, and concurrency control used in a particular parallelization.
Christoph von Praun, Rajesh Bordawekar, Calin Cascaval
PPoPP2
2008 An algorithm for partitioning trees augmented with sibling edges
Rajesh Bordawekar, Oded Shmueli
Inf. Process. Lett.1
2007 CellSort: High Performance Sorting on the Cell Processor
Bugra Gedik, Rajesh Bordawekar, Philip S. Yu
VLDB2
2007 Executing Stream Joins on the Cell Processor
Bugra Gedik, Philip S. Yu, Rajesh Bordawekar
VLDB3
2005 XJ: facilitating XML processing in Java
abstract
The increased importance of XML as a data representation format has led to several proposals for facilitating the development of applications that operate on XML data. These proposals range from runtime API-based interfaces to XML-based programming languages. The subject of this paper is XJ, a research language that proposes novel mechanisms for the integration of XML as a first-class construct into Java™. The design goals of XJ distinguish it from past work on integrating XML support into programming languages --- specifically, the XJ design adheres to the XML Schema and XPath standards. Moreover, it supports in-place updates of XML data thereby keeping with the imperative nature of Java. We have built a prototype compiler for XJ, and our preliminary experiments demonstrate that the performance of XJ programs can approach that of traditional low-level API-based interfaces, while providing a higher level of abstraction.
Matthew Harren, Mukund Raghavachari, Oded Shmueli, Michael G. Burke, Rajesh Bordawekar, Igor Pechtchanski, Vivek Sarkar
WWW5
2002 Exploiting prolific types for memory management and optimizations
abstract
In this paper, we introduce the notion of prolific and non-prolific types, based on the number of instantiated objects of those types. We demonstrate that distinguishing between these types enables a new class of techniques for memory management and data locality, and facilitates the deployment of known techniques. Specifically, we first present a new type-based approach to garbage collection that has similar attributes but lower cost than generational collection. Then we describe the short type pointer technique for reducing memory requirements of objects (data) used by the program. We also discuss techniques to facilitate the recycling of prolific objects and to simplify object co-allocation decisions.We evaluate the first two techniques on a standard set of Java benchmarks (SPECjvm98 and SPECjbb2000). An implementation of the type-based collector in the Jalapeño VM shows improved pause times, elimination of unnecessary write barriers, and reduction in garbage collection time (compared to the analogous generational collector) by up to 15%. A study to evaluate the benefits of the short-type pointer technique shows a potential reduction in the heap space requirements of programs by up to 16%.
Yefim Shuf, Manish Gupta 0002, Rajesh Bordawekar, Jaswinder Pal Singh
POPL3
2000 Quicksilver: a quasi-static compiler for Java
abstract
This paper presents the design and implementation of the Quicksilver1 quasi-static compiler for Java. Quasi-static compilation is a new approach that combines the benefits of static and dynamic compilation, while maintaining compliance with the Java standard, including support of its dynamic features. A quasi-static compiler relies on the generation and reuse of persistent code images to reduce the overhead of compilation during program execution, and to provide identical, testable and reliable binaries over different program executions. At runtime, the quasi-static compiler adapts pre-compiled binaries to the current JVM instance, and uses dynamic compilation of the code when necessary to support dynamic Java features. Our system allows interprocedural program optimizations to be performed while maintaining binary compatibility. Experimental data obtained using a preliminary implementation of a quasi-static compiler in the Jalapeño JVM clearly demonstrates the benefits of our approach: we achieve a runtime compilation cost comparable to that of baseline (fast, non-optimizing) compilation, and deliver the runtime program performance of the highest optimization level supported by the Jalapeño optimizing compiler. For the SPECjvm98 benchmark suite, we obtain a factor of 104 to 158 reduction in the runtime compilation overhead relative to the Jalapeño optimizing compiler. Relative to the better of the baseline and the optimizing Jalapeño compilers, the overall performance (taking into account both runtime compilation and execution costs) is increased by 9.2% to 91.4% for the SPECjvm98 benchmarks with size 100, and by 54% to 356% for the (shorter running) SPECjvm98 benchmarks with size 10.
Mauricio J. Serrano, Rajesh Bordawekar, Samuel P. Midkiff, Manish Gupta 0002
OOPSLA2
2000 Quantitative Characterization and Analysis of the I/O Behavior of a Commercial Distributed-Shared-Memory Machine
abstract
This paper presents a unified evaluation of the I/O behavior of a commercial clustered DSM machine, the HP Exemplar. Our study has the following objectives: 1) To evaluate the impact of different interacting system components, namely, architecture, operating system, and programming model, on the overall I/O behavior and identify possible performance bottlenecks, and 2) To provide hints to the users for achieving high out-of-box I/O throughput. We find that for the DSM machines that are built as a cluster of SMP nodes, integrated clustering of computing and I/O resources, both hardware and software, is not advantageous for two reasons. First, within an SMP node, the I/O bandwidth is often restricted by the performance of the peripheral components and cannot match the memory bandwidth. Second, since the I/O resources are shared as a global resource, the file-access costs become nonuniform and the I/O behavior of the entire system, in terms of both scalability and balance, degrades. We observe that the buffered I/O performance is determined not only by the I/O subsystem, but also by the programming model, global-shared memory subsystem, and data-communication mechanism. Moreover, programming-model support can be used effectively to overcome the performance constraints created by the architecture and operating system. For example, on the HP Exemplar, users can achieve high I/O throughput by using features of the programming model that balance the sharing and locality of the user buffers and file systems. Finally, we believe that at present, the I/O subsystems are being designed in isolation, and there is a need for mending the traditional memory-oriented design approach to address this problem.
Rajesh Bordawekar
IEEE Trans. Parallel Distributed Syst.1
1998 Scaling of Beowulf-Class Distributed Systems
abstract
The future of Beowulf-class PC clusters for large-scale scientific and engineering applications will be determined by three issues: system scalability; latency tolerant applications development; and software environments and tools. This paper focuses on the first of these, scalability, which is sensitive to aspects of the remaining two. For the studies and findings described, we consider system designs capable of integrating hundreds or thousands of processors.
John K. Salmon, Thomas L. Sterling, Rajesh Bordawekar, Christopher Stein
HPDC3
1998 Compilation Techniques for Out-of-Core Parallel Computations
abstract
The difficulty of handling out-of-core data limits the performance of supercomputers as well as the potential of the parallel machines. Since writing an efficient out-of-core version of a program is a difficult task and virtual memory systems do not perform well on scientific computations, we believe that there is a clear need for compiler directed explicit I/O approach for out-of-core computations. In this paper, we first present an out-of-core compilation strategy based on a disk storage abstraction. Then, we offer a compiler algorithm to optimize locality of disk accesses in out-of-core codes by choosing a good combination of file layouts on disks and loop transformations. We introduce memory coefficient and processor coefficient concepts to characterize the behavior of out-of-core programs under different memory constraints. We also enhance our algorithm to handle data-parallel programs which contain multiple loop nest. Our initial experimental results obtained on IBM SP-2 and Intel Paragon provide encouraging evidence that our approach is successful at optimizing programs which depend on disk-resident data in distributed-memory machines.
Mahmut T. Kandemir, Alok N. Choudhary, J. Ramanujam, Rajesh Bordawekar
Parallel Comput.4
1997 Implementation of Collective I/O in the Intel Paragon Parallel File System: Initial Experiences
abstract
Article Free Access Share on Implementation of collective I/O in the Intel Paragon parallel file system: initial experiences Author: Rajesh Bordawekar Center for Advanced Computing Research, California Institute of Technology, 1201 E. California Blvd., Pasadena, CA Center for Advanced Computing Research, California Institute of Technology, 1201 E. California Blvd., Pasadena, CAView Profile Authors Info & Claims ICS '97: Proceedings of the 11th international conference on SupercomputingJuly 1997 Pages 20–27https://doi.org/10.1145/263580.263586Published:11 July 1997Publication History 13citation254DownloadsMetricsTotal Citations13Total Downloads254Last 12 Months12Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Rajesh Bordawekar
International Conference on Supercomputing1
1996 Automatic Optimization of Communication in Compiling Out-of-Core Stencil Codes
abstract
In this paper.we describe a technique for optimizing communication for out-of-core distributed memory stencil problems.In these problems, communication may require both inter-processor communication and file 1/0.We show that in certain cases, extra file 1/0 incurred in communication can be completely eliminated by reordering in-core computations.The in-core computation pattern is decided by: (1) how the out-of-core data distributed into in-core slabs (tiling) and (2) how the slabs are accessed.We show that a compiler using the stencil and processor information can choose the tiling parameters and schedule the tile accesses so that theextra file I/O is eliminated and overall performance is improved.1
Rajesh Bordawekar, Alok N. Choudhary, J. Ramanujam
International Conference on Supercomputing1
1996 Compilation and Communication Strategies for Out-of-Core Programs on Distributed Memory Machines
Rajesh Bordawekar, Alok N. Choudhary, J. Ramanujam
J. Parallel Distributed Comput.1
1995 Communication Strategies for Out-of-Core Programs on Distributed Memory Machines
abstract
In this paper, we show that communication in the out-of-core distributed memory problems requires both inter-processor communication and file I/O. Given that primary data structures reside in files, even communication requires I/O. Thus, it is important to optimize the I/O costs associated with a communication step. We present three methods for performing communication in out-of-core distributed memory problems. The first method, termed as the “out-of-core“communication method, follows a loosely synchronous model. Computation and Communication phases in this case are clearly separated, and communication requires permutation of data in files. The second method, termed as”demand-driven-in-core communication” considers only communication required of each in-core data slab individually. The third method, termed as “producer-driven-in-core communication “ goes even one step further and tries to identify the potential (future) use of data while it is in memory. We describe these methods in detail and provide performance results for out-of-core applications: namely, two-dimensional FFT and two-dimensional elliptic solver. Finally, we discuss how “out-of-core” and “in-core” communication methods could be used in virtual memory environments on distributed memory machines.
Rajesh Bordawekar, Alok N. Choudhary
International Conference on Supercomputing1
1995 A Model and Compilation Strategy for Out-of-Core Data Parallel Programs
abstract
It is widely acknowledged in high-performance computing circles that parallel input/output needs substantial improvement in order to make scalable computers truly usable. We present a data storage model that allows processors independent access to their own data and a corresponding compilation strategy that integrates data-parallel computation with data distribution for out-of-core problems. Our results compare several communication methods and I/O optimizations using two out-of-core problems, Jacobi iteration and LU factorization.
Rajesh Bordawekar, Alok N. Choudhary, Ken Kennedy, Charles Koelbel, Michael H. Paleczny
PPoPP1
1994 Compiler and runtime support for out-of-core HPF programs
abstract
This paper describes the design of a compiler which can translate out-of-core programs written in a data parallel language like HPF. Such a compiler is required for compiling large scale scientific applications, such as the Grand Challenge applications, which deal with enormous quantities of data. We propose a framework by which a compiler together with appropriate runtime support can translate an out-of-core HPF program to a message passing node program with explicit parallel I/O. We describe the basic model of the compiler and the various transformations made by the compiler. We also discuss the runtime routines used by the compiler for I/O and communication. In order to minimize I/O, the runtime support system can reuse data already fetched into memory. The working of the compiler is illustrated using two out-of-core applications, namely a Laplace equation solver and LU Decomposition, together with performance results on the Intel Touchstone Delta.
Rajeev Thakur, Rajesh Bordawekar, Alok N. Choudhary
International Conference on Supercomputing2
1993 An Experimental Performance Evaluation of Touchstone Delta Concurrent File System
abstract
For a high-performance parallel machine to be a scal-able system, it must afso have a scalable parallel 1/0 system. This paper presents an experimental evaluation of the Intel Touchstone Delta’s Concurrent File System ( CFS). The main objective of the study is to determine the maximum file read/write rates for various configura-tions of 1/0 and compute nodes. In addition, we study the effects of file access modes, buffer sizes and file sizes on the system performance. In most cases, the result shows that performance of CFS scales as the number of disks is increased, but the sustained performance im-provements are much lower than the system’s peak ca-pacity. If’e observe that the performance of CFS scales with the number of processors in the beginning, how-ever, a plateu a quickly reached due to the 1/0 system bottleneck and enormous software overhead, especially that of synchronization. Finally we also show that the performance of the CFS can greatly vary for various data distributions commonly employed in scientific and engineering applications. 1
Rajesh Bordawekar, Alok N. Choudhary, Juan Miguel del Rosario
International Conference on Supercomputing1
1993 Design and Evaluation of primitives for Parallel I/O
abstract
Article Design and Evaluation of primitives for Parallel I/O Share on Authors: R. Bordawekar Northeast Parallel Architectures Center, 3-201 CST, Syracuse Univ., Syracuse, NY Northeast Parallel Architectures Center, 3-201 CST, Syracuse Univ., Syracuse, NYView Profile , J. M. del Rosario Northeast Parallel Architectures Center, 3-201 CST, Syracuse Univ., Syracuse, NY Northeast Parallel Architectures Center, 3-201 CST, Syracuse Univ., Syracuse, NYView Profile , A. Choudhary Northeast Parallel Architectures Center, 3-201 CST, Syracuse Univ., Syracuse, NY and ECE Dept. Northeast Parallel Architectures Center, 3-201 CST, Syracuse Univ., Syracuse, NY and ECE Dept.View Profile Authors Info & Claims Supercomputing '93: Proceedings of the 1993 ACM/IEEE conference on SupercomputingDecember 1993 Pages 452–461https://doi.org/10.1145/169627.169782Online:01 December 1993Publication History 71citation227DownloadsMetricsTotal Citations71Total Downloads227Last 12 Months5Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Rajesh Bordawekar, Juan Miguel del Rosario, Alok N. Choudhary
SC1