VLDB 2026 Research / reviewers in the wild / expert
Arun Chauhan 0001
dblp:c/ArunChauhan
· DBLP profile ↗
21ranked-venue papers
3as first author
1since 2021 · last 2023
0000-0002-0327-7254ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Hardware accelerators and domain-specific architectures · 43% Memory systems · 43% Parallel and multicore computing · 12% | |
| Software engineering, system software, and programming languages
5 papers |
Operating systems · 50% Compilers and program optimization · 24% Programming languages and type systems · 14% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Operating systems › resource management › memory management
memory allocation |
0.7 | 1 | 2023 | TelaMalloc: Efficient On-Chip Memory Allocation for Production Machine Learning Accelerators · ASPLOS (1) 2023 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.7 | 1 | 2023 | TelaMalloc: Efficient On-Chip Memory Allocation for Production Machine Learning Accelerators · ASPLOS (1) 2023 |
Memory systems › memory management › on-chip memory management
on-chip memory allocation |
0.7 | 1 | 2023 | TelaMalloc: Efficient On-Chip Memory Allocation for Production Machine Learning Accelerators · ASPLOS (1) 2023 |
Programming languages and type systems
type inference |
0.2 | 2 | 2011 | Flow-sensitive type recovery in linear-log time · OOPSLA 2011 Automatic Type-Driven Library Generation for Telescoping Languages · SC 2003 |
Parallel and multicore computing
parallel programming models |
0.2 | 1 | 2013 | Globalizing selectively: shared-memory efficiency with address-space separation · SC 2013 |
Program analysis
control flow analysis |
0.1 | 1 | 2011 | Flow-sensitive type recovery in linear-log time · OOPSLA 2011 |
Compilers and program optimization
domain-specific compilation |
0.1 | 1 | 2005 | Telescoping Languages: A System for Automatic Generation of Domain Languages · Proc. IEEE 2005 |
Compilers and program optimization
dynamic optimization |
0.0 | 1 | 2013 | Globalizing selectively: shared-memory efficiency with address-space separation · SC 2013 |
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation |
0.0 | 1 | 2011 | Flow-sensitive type recovery in linear-log time · OOPSLA 2011 |
Parallel and multicore computing
task scheduling |
0.0 | 1 | 1999 | Scheduling Constrained Dynamic Applications on Clusters · SC 1999 |
Programming languages and type systems › dynamic languages
scripting language |
0.0 | 1 | 2005 | Telescoping Languages: A System for Automatic Generation of Domain Languages · Proc. IEEE 2005 |
High-performance computing
scientific computing systems |
0.0 | 1 | 2003 | Automatic Type-Driven Library Generation for Telescoping Languages · SC 2003 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.0 | 1 | 1999 | Scheduling Constrained Dynamic Applications on Clusters · SC 1999 |
Methods — techniques the papers use, named apart from their topics
memory buffer allocation · 1.3compilation optimization · 1.3zero-copy communication · 0.3static property proving · 0.3compiler transformation · 0.3sub-0CFA · 0.1linear-log-time algorithm · 0.1offline procedure specialization · 0.1library preprocessing · 0.1annotation · 0.1type inference · 0.0optimal scheduling framework · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | TelaMalloc: Efficient On-Chip Memory Allocation for Production Machine Learning AcceleratorsabstractMemory buffer allocation for on-chip memories is a major challenge in modern machine learning systems that target ML accelerators. In interactive systems such as mobile phones, it is on the critical path of launching ML-enabled applications. In data centers, it is part of complex optimization loops that run many times and are the limiting factor for the quality of compilation results. Martin Maas 0001, Ulysse Beaugnon, Arun Chauhan 0001, Berkin Ilbeyi |
ASPLOS (1) | 3 |
| 2014 | Automatic parallelism through macro dataflow in high-level array languagesabstractDataflow computation is a powerful paradigm for parallel computing that is especially attractive on modern machines with multiple avenues for parallelism. However, adopting this model has been challenging as neither hardware- nor language-based approaches have been successful, except, in specialized contexts. We argue that general-purpose array languages, such as MATLAB, are good candidates for automatic translation to macro dataflow-style execution, where each array operation naturally maps to a macro dataflow operation and the model can be efficiently executed on contemporary multicore architecture. We support our argument with a fully automatic compilation technique to translate MATLAB programs to dynamic dataflow graphs that are capable of handling unbounded structured control flow. These graphs can be executed on multicore machines in an event driven fashion with the help of a runtime system built on top of Intel's Threading Building Blocks (TBB). By letting each task itself be data parallel, we are able to leverage existing data-parallel libraries and utilize parallelism at multiple levels. Our experiments on a set of benchmarks show speedups of up to 18x using our approach, over the original data-parallel code on a machine with two 16-core processors. Pushkar Ratnalikar, Arun Chauhan 0001 |
PACT | 2 |
| 2014 | Optimizing LZSS compression on GPGPUs
Adnan Ozsoy, D. Martin Swany, Arun Chauhan 0001 |
Future Gener. Comput. Syst. | 3 |
| 2013 | Co-processing SPMD computation on CPUs and GPUs clusterabstractHeterogeneous parallel systems with multi processors and accelerators are becoming ubiquitous due to better cost-performance and energy-efficiency. These heterogeneous processor architectures have different instruction sets and are optimized for either task-latency or throughput purposes. Challenges occur in regard to programmability and performance when running SPMD tasks on heterogeneous devices. In order to meet these challenges, we implemented a parallel runtime system that used to co-process SPMD computation on CPUs and GPUs clusters. Furthermore, we are proposing an analytic model to automatically schedule SPMD tasks on heterogeneous clusters. Our analytic model is derived from the roofline model, and therefore it can be applied to a wider range of SPMD applications and hardware devices. The experimental results of the C-means, GMM, and GEMV show good speedup in practical heterogeneous cluster environments. Geoffrey C. Fox, Gregor von Laszewski, Arun Chauhan 0001 |
CLUSTER | 4 |
| 2013 | Achieving TeraCUPS on Longest Common Subsequence Problem Using GPGPUsabstractIn this paper, we describe a novel technique to optimize longest common subsequence (LCS) algorithm for one-to-many matching problem on GPUs by transforming the computation into bit-wise operations and a post-processing step. The former can be highly optimized and achieves more than a trillion operations (cell updates) per second (CUPS)-a first for LCS algorithms. The latter is more efficiently done on CPUs, in a fraction of the bit-wise computation time. The bit-wise step promises to be a foundational step and a fundamentally new approach to developing algorithms for increasingly popular heterogeneous environments that could dramatically increase the applicability of hybrid CPU-GPU environments. Adnan Ozsoy, Arun Chauhan 0001, D. Martin Swany |
ICPADS | 2 |
| 2013 | Globalizing selectively: shared-memory efficiency with address-space separationabstractIt has become common for MPI-based applications to run on shared-memory machines. However, MPI semantics do not allow leveraging shared memory fully for communication between processes from within the MPI library. This paper presents an approach that combines compiler transformations with a specialized runtime system to achieve zero-copy communication whenever possible by proving certain properties statically and globalizing data selectively by altering the allocation and deallocation of communication buffers. The runtime system provides dynamic optimization, when such proofs are not possible statically, by copying data only when there are write-write or read-write conflicts. We implemented a prototype compiler, using ROSE, and evaluated it on several benchmarks. Our system produces code that performs better than MPI in most cases and no worse than MPI, tuned for shared memory, in all cases. Nilesh Mahajan, Uday Pitambare, Arun Chauhan 0001 |
SC | 3 |
| 2012 | Pipelined Parallel LZSS for Streaming Data Compression on GPGPUsabstractIn this paper, we present an algorithm and provide design improvements needed to port the serial Lempel-Ziv-Storer-Szymanski (LZSS), lossless data compression algorithm, to a parallelized version suitable for general purpose graphic processor units (GPGPU), specifically for NVIDIA's CUDA Framework. The two main stages of the algorithm, substring matching and encoding, are studied in detail to fit into the GPU architecture. We conducted detailed analysis of our performance results and compared them to serial and parallel CPU implementations of LZSS algorithm. We also benchmarked our algorithm in comparison with well known, widely used programs, GZIP and ZLIB. We achieved up to 34x better throughput than the serial CPU implementation of LZSS algorithm and up to 2.21x better than the parallelized version. Adnan Ozsoy, D. Martin Swany, Arun Chauhan 0001 |
ICPADS | 3 |
| 2011 | Partial globalization of partitioned address spaces for zero-copy communication with shared memoryabstractWe have developed a high-level language, called Kanor, for declaratively specifying communication in parallel programs. Designed as an extension of C++, it serves to coordinate partitioned address space programs written in the bulk synchronous parallel (BSP) style. Kanor's declarative semantics enable the programmers to write correct and maintainable parallel applications. The communication abstraction has been carefully designed to be amenable to compiler optimizations. While partitioned address space programming has several advantages, it needs special compiler optimizations to effectively leverage the shared memory hardware when running on multicore machines. In this paper, we introduce such shared-memory optimizations in the context of Kanor. One major way we achieve these optimizations is by selectively moving some of the variables into a globally shared address space - a process that we term partial globalization. We identify scenarios in which such a transformation is beneficial, and present an algorithm to identify and correctly transform Kanor communication steps into zero-copy communication using hardware shared memory, by introducing minimal synchronization. We then present a runtime strategy that complements the compiler algorithm to eliminate most of the runtime synchronization overheads by using a copy-on-conflict technique. Finally, we show that our solution often performs much better than shared-memory optimized MPI, and ne ver performs significantly worse than MPI even in the presence of dependencies introduced due to buffer sharing. The techniques in this paper demonstrate that it is possible to program in a partitioned address space style, without sacrificing the performance advantages of hardware shared memory. To the best of our knowledge no other automatic compiler techniques have been developed so far that achieve zero-copy communication from a partitioned address space program. We expect out results to be applicable beyond Kanor, to other partitioned address space programming environments, such as MPI. Fangzhou Jiao, Nilesh Mahajan, Jeremiah Willcock, Arun Chauhan 0001, Andrew Lumsdaine |
HiPC | 4 |
| 2011 | Automating GPU computing in MATLABabstractMATLAB is a popular software platform for scientific and engineering software writers. It offers a high level of abstraction for fundamental mathematical operations and extensive highly optimized domain-specific libraries for several scientific and engineering disciplines. With the recent availability of GPU libraries for MATLAB, it has become possible to easily exploit GPGPUs as coprocessors. However, this requires changing the code by carefully declaring variables that would live on the GPU, breaking the simplicity of the MATLAB programming model. Chun-Yu Shei, Pushkar Ratnalikar, Arun Chauhan 0001 |
ICS | 3 |
| 2011 | Flow-sensitive type recovery in linear-log timeabstractThe flexibility of dynamically typed languages such as JavaScript, Python, Ruby, and Scheme comes at the cost of run-time type checks. Some of these checks can be eliminated via control-flow analysis. However, traditional control-flow analysis (CFA) is not ideal for this task as it ignores flow-sensitive information that can be gained from dynamic type predicates, such as JavaScript's 'instanceof' and Scheme's 'pair?', and from type-restricted operators, such as Scheme's 'car'. Yet, adding flow-sensitivity to a traditional CFA worsens the already significant compile-time cost of traditional CFA. This makes it unsuitable for use in just-in-time compilers. In response, we have developed a fast, flow-sensitive type-recovery algorithm based on the linear-time, flow-insensitive sub-0CFA. The algorithm has been implemented as an experimental optimization for the commercial Chez Scheme compiler, where it has proven to be effective, justifying the elimination of about 60% of run-time type checks in a large set of benchmarks. The algorithm processes on average over 100,000 lines of code per second and scales well asymptotically, running in only O(n log n) time. We achieve this compile-time performance and scalability through a novel combination of data structures and algorithms. Michael D. Adams 0001, Andrew W. Keep, Jan Midtgaard, Matthew Might, Arun Chauhan 0001, R. Kent Dybvig |
OOPSLA | 5 |
| 2011 | Kanor - A Declarative Language for Explicit Communication
Eric Holk, William E. Byrd, Jeremiah Willcock, Torsten Hoefler, Arun Chauhan 0001, Andrew Lumsdaine |
PADL | 5 |
| 2010 | Static reuse distances for locality-based optimizations in MATLABabstractThe problem of modeling memory locality of applications to guide compiler optimizations in a systematic manner is an important unsolved problem, made even more significant with the advent of multi-core and many-core architectures. We describe an approach based on a novel source-level metric, called static reuse distance, to model the memory behavior of applications written in matlab. We use matlab as a representative language that lets end-users express their algorithms precisely, but at a relatively high level. Matlab's "high-level" characteristics allow the static analysis to focus on large objects, such as arrays, without losing accuracy due to processor-specific layout of scalar values in memory. We present an efficient algorithm to compute static reuse distances using an extended version of dependence graphs. Our approach differs from earlier similar attempts in three important aspects: it targets high-level programming systems characterized by heavy use of libraries; it works on full programs, instead of being confined to loops; and it integrates practical mechanisms to handle separately compiled procedures as well as pre-compiled library procedures that are only available in binary form. Arun Chauhan 0001, Chun-Yu Shei |
ICS | 1 |
| 2009 | Compile-time disambiguation of MATLAB types through concrete interpretation with automatic run-time fallbackabstractWhile the popularity of MATLAB for scientific and engineering applications is unabated, its poor performance compared to traditional languages, such as Fortran or even C, for a general class of problems continues to impede its deployment in full-scale simulations and data analysis. To ameliorate performance, we have been developing a MATLAB and Octave compiler that leverages the interpreter to implement some of the optimizations as concrete partial evaluations. Specifically, this paper describes constant propagation and type inference, using a high-level tree-transformation tool that has built-in support for solving dataflow problems. The approach allows propagation and folding of constants in cases that would be impractically difficult otherwise. The idea, when extended to infer variable types, provides a natural way to disambiguate types at compile time while leaving the fallback code in place for run-time evaluation. Experimental evaluation on pieces of real MATLAB code demonstrates the effectiveness of the approach. Chun-Yu Shei, Arun Chauhan 0001, Sidney Shaw |
HiPC | 2 |
| 2008 | Concrete Partial Evaluation in RubyabstractModern scientific research is a collaborative process, with researchers from many disciplines and institutions working toward a common goal. Dynamic languages, like Ruby, provide a platform for quickly developing simulation and analysis tools, freeing researchers to focus on research instead of spending time developing infrastructure. Ruby is a particularly good fit, allowing incorporation of existing C libraries, simplifying Domain Specific Language creation, and providing both REST and SOAP web-based API libraries. Ruby also provides RPC-style distributed programming. Concrete partial evaluation of Ruby begins to address Ruby's biggest flaw, performance. The scientific community has already begun to recognize the potential of Ruby. An MPI extension to the language allows quick prototyping of MPI programs. More recently libraries supporting MapReduce have appeared. Web frameworks, such as the popular Ruby on Rails framework, provide tools for producing and consuming REST APIs. Andrew W. Keep, Arun Chauhan 0001 |
eScience | 2 |
| 2008 | Compile-Time Disambiguation of MATLAB Types through Concrete Interpretation with Automatic Run-Time FallbackabstractWhile the popularity of MATLAB for scientific and engineering applications is unabated, its poor performance compared to traditional languages, such as Fortran or even C, for a general class of problems continues to impede its deployment in full-scale simulations and data analysis. To ameliorate performance, we have been developing a MATLAB and Octave compiler that leverages the interpreter to implement some of the optimizations as concrete partial evaluations. Specifically, this poster describes constant propagation and type inference, using a high-level tree-transformation tool that has built-in support for solving dataflow problems. The approach allows propagation and folding of constants in cases that would be impractically difficult otherwise. The idea, when extended to infer variable types, provides a natural way to disambiguate types at compile time while leaving the fallback code in place for run-time evaluation. Experimental evaluation on pieces of real MATLAB code demonstrates the effectiveness of the approach. Chun-Yu Shei, Arun Chauhan 0001 |
eScience | 2 |
| 2008 | A Model for Communication in Clusters of Multi-core MachinesabstractA common paradigm for scientific computing is distributed message-passing systems, and a common approach to these systems is to implement them across clusters of high-performance workstations. As multi-core architectures become increasingly mainstream, these clusters are very likely to include multi-core machines. However, the theoretical models which are currently used to develop communication algorithms across these systems do not take into account the unique properties of processes running on shared- memory architectures, including shared external network connections and communication via shared memory locations. Because of this, existing algorithms are far from optimal for modern clusters. Additionally, recent attempts to adapt these algorithms to multicore systems have proceeded without the introduction of a more accurate formal model and have generally neglected to capitalize on the full power these systems offer. We propose a new model which simply and effectively captures the strengths of multi-core machines in collective communications patterns and suggest how it could be used to properly optimize these patterns. Christine Task, Arun Chauhan 0001 |
eScience | 2 |
| 2007 | Library Function Selection in Compiling OctaveabstractOne way to address the continuing performance problem of high-level domain-specific languages, such as Octave or Matlab, is to compile them to a relatively lower level language for which good compilers are available. As a first step in this direction, specializing the high-level operations in the source, based on operand types, leads to significant gains. However, simple translation of the high-level operations to the underlying libraries can often miss important opportunities to improve performance. This paper presents a global algorithm to select functions from a target library, utilizing the semantics of the operations as well as the platform-specific performance characteristics of the library. Making use of the library properties, the simple and easy-to-implement selection algorithm, is able to achieve as much as three times performance improvement for certain linear algebra kernels, over a straight mapping of operations, which are compiled to the vendor-tuned BLAS. Daniel S. McFarlin, Arun Chauhan 0001 |
IPDPS | 2 |
| 2005 | Telescoping Languages: A System for Automatic Generation of Domain LanguagesabstractThe software gap - the discrepancy between the need for new software and the aggregate capacity of the workforce to produce it - is a serious problem for scientific software. Although users appreciate the convenience (and, thus, improved productivity) of using relatively high-level scripting languages, the slow execution speeds of these languages remain a problem. Lower level languages, such as C and Fortran, provide better performance for production applications, but at the cost of tedious programming and optimization by experts. If applications written in scripting languages could be routinely compiled into highly optimized machine code, a huge productivity advantage would be possible. It is not enough, however, to simply develop excellent compiler technologies for scripting languages (as a number of projects have succeeded in doing for MATLAB). In practice, scientists typically extend these languages with their own domain-centric components, such as the MATLAB signal processing toolbox. Doing so effectively defines a new domain-specific language. If we are to address efficiency problems for such extended languages, we must develop a framework for automatically generating optimizing compilers for them. To accomplish this goal, we have been pursuing an innovative strategy that we call telescoping languages. Our approach calls for using a library-preprocessing phase to extensively analyze and optimize collections of libraries that define an extended language. Results of this analysis are collected into annotated libraries and used to generate a library-aware optimizer. The generated library-aware optimizer uses the knowledge gathered during preprocessing to carry out fast and effective optimization of high-level scripts. This enables script optimization to benefit from the intense analysis performed during preprocessing without repaying its price. Since library preprocessing is performed only at infrequent "language-generation" times, its cost is amortized over many Ken Kennedy, Bradley Broom, Arun Chauhan 0001, Robert J. Fowler, John Garvin, Charles Koelbel, Cheryl McCosh, John M. Mellor-Crummey |
Proc. IEEE | 3 |
| 2003 | Automatic Type-Driven Library Generation for Telescoping LanguagesabstractTelescoping languages is a strategy to automatically generate highly-optimized domain-specific libraries. The key idea is to create specialized variants of library procedures through extensive offline processing. This paper describes a telescoping system, called ARGen, which generates high-performance Fortran or C libraries from prototype Matlab code for the linear algebra library, ARPACK. ARGen uses variable types to guide procedure specializations on possible calling contexts. ARGen needs to infer Matlab types in order to speculate on the possible variants of library procedures, as well as to generate code. This paper shows that our type-inference system is powerful enough to generate all the variants needed for ARPACK automatically from the Matlab development code. The ideas demonstrated here provide a basis for building a more general telescoping system for Matlab. Arun Chauhan 0001, Cheryl McCosh, Ken Kennedy, Richard Hanson |
SC | 1 |
| 2001 | Optimizing strategies for telescoping languages: procedure strength reduction and procedure vectorizationabstractAt Rice University, we have undertaken a project to construct a framework for generating high-level problem solving languages that can achieve high performance on a variety of platforms.The underlying strategy, called telescoping languages, builds problem-solving systems from domain-specific libraries and scripting langauges. To accomplish this it extensively preanalyzes and transforms the library to produce a scripting language precompiler that optimizes library calls within the scripts as if they were primitives in the underlying language. Arun Chauhan 0001, Ken Kennedy |
ICS | 1 |
| 1999 | Scheduling Constrained Dynamic Applications on ClustersabstractThere is an emerging class of computationally demanding multimedia applications involving vision, speech and interaction with the real world (e.g., CRL's Smart Kiosk). These applications are highly parallel and require low latencies for good performance. They are well-suited for implementation on clusters of SMP's, but they require efficient scheduling of application tasks. General purpose schedulers produce high latencies because they lack knowledge of the dependencies between tasks. Previous research in optimal scheduling has been limited to static problems. In contrast, our application is highly dynamic as the optimal schedule depends upon the behavior of the kiosk's customers. We observe that the dynamism of our application class is constrained, in that there are a small number of operating regimes which are determined by the state of the application. We present a framework for optimal scheduling of constrained dynamic applications. The results of an experimental compariso... Kathleen Knobe, James M. Rehg, Arun Chauhan 0001, Rishiyur S. Nikhil, Umakishore Ramachandran |
SC | 3 |