Anwar M. Ghuloum

dblp:56/4059 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Parallel and multicore computing · 56% Processor architecture and microarchitecture · 32% High-performance computing · 12%
Software engineering, system software, and programming languages
4 papers
Compilers and program optimization · 43% Runtime systems and virtual machines · 43% Program analysis · 14%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Runtime systems and virtual machines
language runtime
0.112007
Enabling scalability and performance in a large scale CMP environment · EuroSys 2007
Processor architecture and microarchitecture
chip multiprocessor
0.112007
Enabling scalability and performance in a large scale CMP environment · EuroSys 2007
Parallel and multicore computing
parallel programming runtimes
0.112007
Enabling scalability and performance in a large scale CMP environment · EuroSys 2007
Compilers and program optimization
parallelization
0.021999
SUIF Explorer: An Interactive and Interprocedural Parallelizer · PPoPP 1999
Parallelizing Complex Scans and Reductions · PLDI 1994
Program analysis › static analysis
program slicing
0.011999
SUIF Explorer: An Interactive and Interprocedural Parallelizer · PPoPP 1999
Parallel and multicore computing › parallel programming models
automatic parallelization
0.021999
Parallelizing Complex Scans and Reductions · PLDI 1994
SUIF Explorer: An Interactive and Interprocedural Parallelizer · PPoPP 1999
Parallel and multicore computing
parallel programming models
0.021999
Parallelizing Complex Scans and Reductions · PLDI 1994
SUIF Explorer: An Interactive and Interprocedural Parallelizer · PPoPP 1999
Compilers and program optimization
loop transformation
0.011995
Flattening and Parallelizing Irregular, Recurrent Loop Nests · PPoPP 1995
Parallel and multicore computing
parallelizing compiler
0.011995
Flattening and Parallelizing Irregular, Recurrent Loop Nests · PPoPP 1995
High-performance computing › sparse linear algebra
sparse matrix computation
0.011995
Flattening and Parallelizing Irregular, Recurrent Loop Nests · PPoPP 1995
High-performance computing › sparse linear algebra › sparse matrix computation
sparse matrix-vector multiplication
0.011995
Flattening and Parallelizing Irregular, Recurrent Loop Nests · PPoPP 1995

Methods — techniques the papers use, named apart from their topics

experimental evaluation · 0.1visualization · 0.0program slicing · 0.0dynamic execution analysis · 0.0segmented scan · 0.0segmented reduction · 0.0loop flattening · 0.0functional composition · 0.0closed-form representation · 0.0
YearPublicationVenuePosition
2011 Intel's Array Building Blocks: A retargetable, dynamic compiler and embedded language
abstract
Our ability to create systems with large amount of hardware parallelism is exceeding the average software developer's ability to effectively program them. This is a problem that plagues our industry. Since the vast majority of the world's software developers are not parallel programming experts, making it easy to write, port, and debug applications with sufficient core and vector parallelism is essential to enabling the use of multi- and many-core processor architectures. However, hardware architectures and vector ISAs are also shifting and diversifying quickly, making it difficult for a single binary to run well on all possible targets. Because of this, retargetability and dynamic compilation are of growing relevance. This paper introduces Intel®Array Building Blocks (ArBB), which is a retargetable dynamic compilation framework. This system focuses on making it easier to write and port programs so that they can harvest data and thread parallelism on both multi-core and heterogeneous many-core architectures, while staying within standard C++. ArBB interoperates with other programming models to help meet the demands we hear from customers for a solution with both greater programmer productivity and good performance. This work makes contributions in language features, compiler architecture, code transformations and optimizations. It presents performance data from the current beta release of ArBB and quantitatively shows the impact of some key analyses, enabling transformations and optimizations for a variety of benchmarks that are of interest to our customers.
Chris J. Newburn, Byoungro So, Zhenying Liu, Michael D. McCool, Anwar M. Ghuloum, Stefanus Du Toit, Zhi-Gang Wang, Zhaohui Du, Yongjian Chen, Gansha Wu, Zhanglin Liu
CGO5
2007 Enabling scalability and performance in a large scale CMP environment
abstract
Hardware trends suggest that large-scale CMP architectures, with tens to hundreds of processing cores on a single piece of silicon, are iminent within the next decade. While existing CMP machines have traditionally been handled in the same way as SMPs, this magnitude of parallelism introduces several fundamental challenges at the architectural level and this, in turn, translates to novel challenges in the design of the software stack for these platforms. This paper presents the "Many Core Run Time" (McRT), a software prototype of an integrated language runtime that was designed to explore configurations of the software stack for enabling performance and scalability on large scale CMP platforms. This paper presents the architecture of McRT and discusses our experiences with the system, including experimental evaluation that lead to several interesting, non-intuitive findings, providing key insights about the structure of the system stack at this scale. A key contribution of this paper is to demonstrate how McRT enables near linear improvements in performance and scalability for desktop workloads such as the popular XviD encoder and a set of RMS (recognition, mining, and synthesis) applications. Another key contribution of this work is its use of McRT to explore non-traditional system configurations such as a light-weight executive in which McRT runs on "bare metal" and replaces the traditional OS. Such configurations are becoming an increasingly attractive alternative to leverage heterogeneous computing uints as seen in today's CPU-GPU configurations.
Bratin Saha, Ali-Reza Adl-Tabatabai, Anwar M. Ghuloum, Mohan Rajagopalan, Richard L. Hudson, Leaf Petersen, Vijay Menon 0002, Brian R. Murphy, Tatiana Shpeisman, Eric Sprangle, Anwar Rohillah, Doug Carmean, Jesse Fang
EuroSys3
2007 Compression in cache design
abstract
Increasing cache capacity via compression enables designers to improve performance of existing designs for small incremental cost, further leveraging the large die area invested in last level caches. This paper explores the compressed cache design space with focus on implementation feasibility.
Ali-Reza Adl-Tabatabai, Anwar M. Ghuloum, Shobhit O. Kanaujia
ICS2
1999 SUIF Explorer: An Interactive and Interprocedural Parallelizer
abstract
The SUIF Explorer is an interactive parallelization tool that is more effective than previous systems in minimizing the number of lines of code that require programmer assistance. First, the interprocedural analyses in the SUIF system is successful in parallelizing many coarse-grain loops, thus minimizing the number of spurious dependences requiring attention. Second, the system uses dynamic execution analyzers to identify those important loops that are likely to be parallelizable. Third, the SUIF Explorer is the first to apply program slicing to aid programmers in interactive parallelization. The system guides the programmer in the parallelization process using a set of sophisticated visualization techniques.This paper demonstrates the effectiveness of the SUIF Explorer with three case studies. The programmer was able to speed up all three programs by examining only a small fraction of the program and privatizing a few variables.
Shih-Wei Liao, Amer Diwan, Robert P. Bosch Jr., Anwar M. Ghuloum, Monica S. Lam
PPoPP4
1995 Flattening and Parallelizing Irregular, Recurrent Loop Nests
abstract
Irregular loop nests in which the loop bounds are determined dynamically by indexed arrays are difficult to compile into expressive parallel constructs, such as segmented scans and reductions. In this paper, we describe a suite of transformations to automatically parallelize such irregular loop nests, even in the presence of recurrences. We describe a simple, general loop flattening transformation, along with new optimizations which make it a viable compiler transformation. A robust recurrence parallelization technique is coupled to the loop flattening transformation, allowing parallelization of segmented reductions, scans, and combining-sends over arbitrary associative operators. We discuss the implementation and performance results of the transformations in a parallelizing Fortran 77 compiler for the Cray C90 supercomputer. In particular, we focus on important sparse matrix-vector multiplication kernels, for one of which we are able to automatically derive an algorithm used by one of the fastest library routines available.
Anwar M. Ghuloum, Allan L. Fisher
PPoPP1
1994 Parallelizing Complex Scans and Reductions
abstract
We present a method for automatically extracting parallel prefix programs from sequential loops, even in the presence of complicated conditional statements. Rather than searching for associative operators in the loop body directly, the method rests on the observation that functional composition itself is associative. Accordingly, we model the loop body as a multivalued function of multiple parameters, and look for a closed-form representation of arbitrary compositions of loop body instances. Careful analysis of conditionals allows this search to succeed in cases where existing automatic methods fail. The method has been implemented and used to generate code for the iWarp parallel computer.
Allan L. Fisher, Anwar M. Ghuloum
PLDI2