Walid A. Abu-Sufah

dblp:281/2237 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-authorSoftware engineering, systems software and programming languages · 4 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Performance modeling and evaluation · 34% Parallel and multicore computing · 31% Memory systems · 21%
Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 95% Operating systems · 5%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
parallelizing compiler
0.021987
On Reducing Data Synchronization in Multiprocessed Loops · IEEE Trans. Computers 1987
A Technique for Reducing Synchronization Overhead in Large Scale Multiprocessors · ISCA 1985
Compilers and program optimization › parallel program optimization
synchronization optimization
0.021987
On Reducing Data Synchronization in Multiprocessed Loops · IEEE Trans. Computers 1987
A Technique for Reducing Synchronization Overhead in Large Scale Multiprocessors · ISCA 1985
Memory systems › memory management
virtual memory
0.031982
Some Results on the Working Set Anomalies in Numerical Programs · IEEE Trans. Software Eng. 1982
Experimental Results on the Paging Behavior of Numerical Programs · ICSE 1982
On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981
Performance modeling and evaluation
workload characterization
0.021986
On Input/Output Speedup in Tightly Coupled Multiprocessors · IEEE Trans. Computers 1986
Some Results on the Working Set Anomalies in Numerical Programs · IEEE Trans. Software Eng. 1982
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor
0.021987
On Reducing Data Synchronization in Multiprocessed Loops · IEEE Trans. Computers 1987
A Technique for Reducing Synchronization Overhead in Large Scale Multiprocessors · ISCA 1985
Compilers and program optimization › parallel program optimization › synchronization optimization
synchronization elimination
0.011987
On Reducing Data Synchronization in Multiprocessed Loops · IEEE Trans. Computers 1987
Parallel and multicore computing › loop transformation
loop parallelization
0.011987
On Reducing Data Synchronization in Multiprocessed Loops · IEEE Trans. Computers 1987
Performance modeling and evaluation › parallel system performance
speedup modeling
0.011986
On Input/Output Speedup in Tightly Coupled Multiprocessors · IEEE Trans. Computers 1986
Parallel and multicore computing › synchronization
synchronization overhead reduction
0.011985
A Technique for Reducing Synchronization Overhead in Large Scale Multiprocessors · ISCA 1985
Performance modeling and evaluation
memory access traces
0.011982
Some Results on the Working Set Anomalies in Numerical Programs · IEEE Trans. Software Eng. 1982
Memory systems › virtual memory management
paging behavior
0.011982
Experimental Results on the Paging Behavior of Numerical Programs · ICSE 1982
Compilers and program optimization › memory optimization
data locality optimization
0.011981
On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981
Compilers and program optimization
program transformation
0.011981
On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981
Storage systems
paging performance
0.011981
On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981
Parallel and multicore computing › parallel computing
parallel application performance
0.011986
On Input/Output Speedup in Tightly Coupled Multiprocessors · IEEE Trans. Computers 1986
Memory systems › memory interference
memory bank conflicts
0.011985
Performance Prediction Tools for Cedar: A Multiprocessor Supercomputer · ISCA 1985
High-performance computing › scientific computing
numerical programs
0.011982
Experimental Results on the Paging Behavior of Numerical Programs · ICSE 1982
High-performance computing
scientific computing systems
0.011982
Experimental Results on the Paging Behavior of Numerical Programs · ICSE 1982
Operating systems › resource management
memory management
0.011981
On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981
Operating systems › resource management › memory management
page replacement
0.011981
On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981

Methods — techniques the papers use, named apart from their topics

loop restructuring · 0.0compiler analysis · 0.0empirical measurement · 0.0analytic speedup modeling · 0.0LRU replacement · 0.0simulation · 0.0mean value analysis · 0.0compiler optimization passes · 0.0compiler optimization pass · 0.0trace analysis · 0.0source-to-source transformation · 0.0program analysis · 0.0experimental measurement · 0.0
YearPublicationVenuePosition
2020 CCF: An efficient SpMV storage format for AVX512 platforms
Mohammad Almasri, Walid A. Abu-Sufah
Parallel Comput.2
1987 On Reducing Data Synchronization in Multiprocessed Loops
abstract
In this correspondence we present and prove the correctness of an algorithm for reducing the number of synchronized memory references to shared data elements in multiprocessed loops. Optimizing compilers for shared memory multiprocessors can use this algorithm to reduce synchronization overhead. The algorithm has been implemented as a new module in the multiprocessors version of Parafrase, the restructuring system of the University of Illinois. We present a brief discussion of experiments we performed to asses the effectiveness of this algorithm in reducing the synchronization overhead for the 61 subroutines of EISPACK, a package for computing matrix eigenvectors and eigenvalues.
Zhiyuan Li 0001, Walid A. Abu-Sufah
IEEE Trans. Computers2
1986 Vector Processing on the Alliant FX/8 Multiprocessor
Walid A. Abu-Sufah, Allen D. Malony
ICPP1
1986 On Input/Output Speedup in Tightly Coupled Multiprocessors
abstract
Previous models of program speedup on parallel architectures tend to ignore I/O activity and other important issues. In this paper we derive analytic speedup models including I/O activities. We show that ignoring I/O yields conservative speedup results. We explore the effectiveness of using hardware format conversion units in multiprocessors [33]. We prove that hardware parallel format conversion loses its edge over software parallel format conversion if the ratio of the number of processors to I/O bandwidth increases. For a given number of processors, program speedup is more sensitive to the available I/O bandwidth rather than the format conversion speed. Ninety-one Fortran programs are used in various experiments to verify our models and conclusions. Most of the programs are I/O bound. Our empirical results show that including I/O activity improves the speedup factor for 78 percent of the programs, and 18 percent of the programs are sped up only due to faster I/O activities. For a serial machine, using hardware format conversion units designed in [13] reduces program execution time by an average factor of three. The software format conversion speed used is obtained from direct measurements on an IBM 4341 running CMS and a CDC Cyber 175 running NOS. For multiprocessor systems a factor of eight increase in the processors to I/O bandwidth ratio reduces the effectiveness of hardware format conversion to an average factor of 1.36.
Walid A. Abu-Sufah, Harlan E. Husmann, David J. Kuck
IEEE Trans. Computers1
1985 Performance Prediction Tools for Cedar: A Multiprocessor Supercomputer
abstract
The development of performance prediction tools for highspeed machine organizations has been recognized as a key problem facing the research community in parallel computing [BrAI84].This paper presents a survey of the tools which has been developed for performance prediction of the Cedar multiproce~or supercomputer of the University of Illinois.The system is deterministic, modular, and automatic.The hierarchical organization of the system provides the user with the ability to choose from a set of alternatives for predicting the performance with different levels of accuracy and cost.Using 22 programs, we measure the performance degradation due to conflicts in the shared memory, delay in the Cedar interconnection network, and synchronization overhead.The results confirm that the architecture of Cedar as detailed in [GLYZ84] is balanced.The performance of the Cedar interconnectiou network is very close to a crossbar.Synchronization overhead and shared memory conflicts could degrade performance for some programs considerably.
Walid A. Abu-Sufah, Alex Y. Kwok
ISCA1
1985 A Technique for Reducing Synchronization Overhead in Large Scale Multiprocessors
abstract
The effectiveness of multiprocessing a single subroutine in a tightly coupled muir!processor system is highly dependent (among other factors) on the amount of overhead incurred due to synchronized references to shared data items by active processes.In this paper we present a technique to reduce this overhead.This is achieved by reducing the number of synchronized references issued by different processes.We describe the criteria used to identify the variables whose synchronized accesses will eliminate the need to synchronize references to other variables.Our technique can be implemented in optimizing compilers for mnltiprocessor machines.It can also be used when designing parallel algorithms for multiproeessors.We implemented the technique as a pass in the Parafrase system of the University of Illinois.Relevant measurements on all the subroutines of the EISPACK package are reported (61 subroutines).The results show that our technique is quite effective.For subroutines with interprocess communication, it reduced the number of memory references requiring synchronization by 10% in two thirds of these subroutines and by more than 30% for one third of the subroutines.
Zhiyuan Li 0001, Walid A. Abu-Sufah
ISCA2
1982 Experimental Results on the Paging Behavior of Numerical Programs
Walid A. Abu-Sufah, Mohammad Malkawi, P. Yew
ICSE1
1982 Some Results on the Working Set Anomalies in Numerical Programs
abstract
This paper shows that the working set parameter-real memory and real memory-fault rate anomalies mentioned by Franklin, Graham, and Gupta in [13] do occur in traces generated by real programs. The results of the detailed investigation of this anomalous behavior in four Fortran programs are presented. In some cases a drop of a factor of two in the average real-time memory allotment is observed when the window size is increased. In some instances a bigger real-time memory allotment means an order of magnitude increase in page faults.
Walid A. Abu-Sufah, David A. Padua
IEEE Trans. Software Eng.1
1981 On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations
abstract
It is possible to improve the paging performance of a program by applying transformations to the source program that improve data access locality. We discuss this subject in general terms, including automation of these transformations, and present a number of such transforms. This is followed by experimental results which indicate that these transformations are indeed effective. Use of practical, simple memory management policies like the fixed allocation local lru replacement algorithm leads to average improvements over untransformed programs of a factor of 10 in space-time cost, and a factor of 5 in memory size. Multiprogramming questions are also discussed.
Walid A. Abu-Sufah, David J. Kuck, Duncan H. Lawrie
IEEE Trans. Computers1