EDBT 2026 Demo / reviewers in the wild / expert
Walid A. Abu-Sufah
dblp:281/2237
· DBLP profile ↗
9ranked-venue papers
6as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-authorSoftware engineering, systems software and programming languages · 4 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Performance modeling and evaluation · 34% Parallel and multicore computing · 31% Memory systems · 21% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 95% Operating systems · 5% |
Topics — the 20 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
parallelizing compiler |
0.0 | 2 | 1987 | On Reducing Data Synchronization in Multiprocessed Loops · IEEE Trans. Computers 1987 A Technique for Reducing Synchronization Overhead in Large Scale Multiprocessors · ISCA 1985 |
Compilers and program optimization › parallel program optimization
synchronization optimization |
0.0 | 2 | 1987 | On Reducing Data Synchronization in Multiprocessed Loops · IEEE Trans. Computers 1987 A Technique for Reducing Synchronization Overhead in Large Scale Multiprocessors · ISCA 1985 |
Memory systems › memory management
virtual memory |
0.0 | 3 | 1982 | Some Results on the Working Set Anomalies in Numerical Programs · IEEE Trans. Software Eng. 1982 Experimental Results on the Paging Behavior of Numerical Programs · ICSE 1982 On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981 |
Performance modeling and evaluation
workload characterization |
0.0 | 2 | 1986 | On Input/Output Speedup in Tightly Coupled Multiprocessors · IEEE Trans. Computers 1986 Some Results on the Working Set Anomalies in Numerical Programs · IEEE Trans. Software Eng. 1982 |
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor |
0.0 | 2 | 1987 | On Reducing Data Synchronization in Multiprocessed Loops · IEEE Trans. Computers 1987 A Technique for Reducing Synchronization Overhead in Large Scale Multiprocessors · ISCA 1985 |
Compilers and program optimization › parallel program optimization › synchronization optimization
synchronization elimination |
0.0 | 1 | 1987 | On Reducing Data Synchronization in Multiprocessed Loops · IEEE Trans. Computers 1987 |
Parallel and multicore computing › loop transformation
loop parallelization |
0.0 | 1 | 1987 | On Reducing Data Synchronization in Multiprocessed Loops · IEEE Trans. Computers 1987 |
Performance modeling and evaluation › parallel system performance
speedup modeling |
0.0 | 1 | 1986 | On Input/Output Speedup in Tightly Coupled Multiprocessors · IEEE Trans. Computers 1986 |
Parallel and multicore computing › synchronization
synchronization overhead reduction |
0.0 | 1 | 1985 | A Technique for Reducing Synchronization Overhead in Large Scale Multiprocessors · ISCA 1985 |
Performance modeling and evaluation
memory access traces |
0.0 | 1 | 1982 | Some Results on the Working Set Anomalies in Numerical Programs · IEEE Trans. Software Eng. 1982 |
Memory systems › virtual memory management
paging behavior |
0.0 | 1 | 1982 | Experimental Results on the Paging Behavior of Numerical Programs · ICSE 1982 |
Compilers and program optimization › memory optimization
data locality optimization |
0.0 | 1 | 1981 | On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981 |
Compilers and program optimization
program transformation |
0.0 | 1 | 1981 | On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981 |
Storage systems
paging performance |
0.0 | 1 | 1981 | On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981 |
Parallel and multicore computing › parallel computing
parallel application performance |
0.0 | 1 | 1986 | On Input/Output Speedup in Tightly Coupled Multiprocessors · IEEE Trans. Computers 1986 |
Memory systems › memory interference
memory bank conflicts |
0.0 | 1 | 1985 | Performance Prediction Tools for Cedar: A Multiprocessor Supercomputer · ISCA 1985 |
High-performance computing › scientific computing
numerical programs |
0.0 | 1 | 1982 | Experimental Results on the Paging Behavior of Numerical Programs · ICSE 1982 |
High-performance computing
scientific computing systems |
0.0 | 1 | 1982 | Experimental Results on the Paging Behavior of Numerical Programs · ICSE 1982 |
Operating systems › resource management
memory management |
0.0 | 1 | 1981 | On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981 |
Operating systems › resource management › memory management
page replacement |
0.0 | 1 | 1981 | On the Performance Enhancement of Paging Systems Through Program Analysis and Transformations · IEEE Trans. Computers 1981 |
Methods — techniques the papers use, named apart from their topics
loop restructuring · 0.0compiler analysis · 0.0empirical measurement · 0.0analytic speedup modeling · 0.0LRU replacement · 0.0simulation · 0.0mean value analysis · 0.0compiler optimization passes · 0.0compiler optimization pass · 0.0trace analysis · 0.0source-to-source transformation · 0.0program analysis · 0.0experimental measurement · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | CCF: An efficient SpMV storage format for AVX512 platforms
Mohammad Almasri, Walid A. Abu-Sufah |
Parallel Comput. | 2 |
| 1987 | On Reducing Data Synchronization in Multiprocessed LoopsabstractIn this correspondence we present and prove the correctness of an algorithm for reducing the number of synchronized memory references to shared data elements in multiprocessed loops. Optimizing compilers for shared memory multiprocessors can use this algorithm to reduce synchronization overhead. The algorithm has been implemented as a new module in the multiprocessors version of Parafrase, the restructuring system of the University of Illinois. We present a brief discussion of experiments we performed to asses the effectiveness of this algorithm in reducing the synchronization overhead for the 61 subroutines of EISPACK, a package for computing matrix eigenvectors and eigenvalues. Zhiyuan Li 0001, Walid A. Abu-Sufah |
IEEE Trans. Computers | 2 |
| 1986 | Vector Processing on the Alliant FX/8 Multiprocessor
Walid A. Abu-Sufah, Allen D. Malony |
ICPP | 1 |
| 1986 | On Input/Output Speedup in Tightly Coupled MultiprocessorsabstractPrevious models of program speedup on parallel architectures tend to ignore I/O activity and other important issues. In this paper we derive analytic speedup models including I/O activities. We show that ignoring I/O yields conservative speedup results. We explore the effectiveness of using hardware format conversion units in multiprocessors [33]. We prove that hardware parallel format conversion loses its edge over software parallel format conversion if the ratio of the number of processors to I/O bandwidth increases. For a given number of processors, program speedup is more sensitive to the available I/O bandwidth rather than the format conversion speed. Ninety-one Fortran programs are used in various experiments to verify our models and conclusions. Most of the programs are I/O bound. Our empirical results show that including I/O activity improves the speedup factor for 78 percent of the programs, and 18 percent of the programs are sped up only due to faster I/O activities. For a serial machine, using hardware format conversion units designed in [13] reduces program execution time by an average factor of three. The software format conversion speed used is obtained from direct measurements on an IBM 4341 running CMS and a CDC Cyber 175 running NOS. For multiprocessor systems a factor of eight increase in the processors to I/O bandwidth ratio reduces the effectiveness of hardware format conversion to an average factor of 1.36. Walid A. Abu-Sufah, Harlan E. Husmann, David J. Kuck |
IEEE Trans. Computers | 1 |
| 1985 | Performance Prediction Tools for Cedar: A Multiprocessor SupercomputerabstractThe development of performance prediction tools for highspeed machine organizations has been recognized as a key problem facing the research community in parallel computing [BrAI84].This paper presents a survey of the tools which has been developed for performance prediction of the Cedar multiproce~or supercomputer of the University of Illinois.The system is deterministic, modular, and automatic.The hierarchical organization of the system provides the user with the ability to choose from a set of alternatives for predicting the performance with different levels of accuracy and cost.Using 22 programs, we measure the performance degradation due to conflicts in the shared memory, delay in the Cedar interconnection network, and synchronization overhead.The results confirm that the architecture of Cedar as detailed in [GLYZ84] is balanced.The performance of the Cedar interconnectiou network is very close to a crossbar.Synchronization overhead and shared memory conflicts could degrade performance for some programs considerably. Walid A. Abu-Sufah, Alex Y. Kwok |
ISCA | 1 |
| 1985 | A Technique for Reducing Synchronization Overhead in Large Scale MultiprocessorsabstractThe effectiveness of multiprocessing a single subroutine in a tightly coupled muir!processor system is highly dependent (among other factors) on the amount of overhead incurred due to synchronized references to shared data items by active processes.In this paper we present a technique to reduce this overhead.This is achieved by reducing the number of synchronized references issued by different processes.We describe the criteria used to identify the variables whose synchronized accesses will eliminate the need to synchronize references to other variables.Our technique can be implemented in optimizing compilers for mnltiprocessor machines.It can also be used when designing parallel algorithms for multiproeessors.We implemented the technique as a pass in the Parafrase system of the University of Illinois.Relevant measurements on all the subroutines of the EISPACK package are reported (61 subroutines).The results show that our technique is quite effective.For subroutines with interprocess communication, it reduced the number of memory references requiring synchronization by 10% in two thirds of these subroutines and by more than 30% for one third of the subroutines. Zhiyuan Li 0001, Walid A. Abu-Sufah |
ISCA | 2 |
| 1982 | Experimental Results on the Paging Behavior of Numerical Programs
Walid A. Abu-Sufah, Mohammad Malkawi, P. Yew |
ICSE | 1 |
| 1982 | Some Results on the Working Set Anomalies in Numerical ProgramsabstractThis paper shows that the working set parameter-real memory and real memory-fault rate anomalies mentioned by Franklin, Graham, and Gupta in [13] do occur in traces generated by real programs. The results of the detailed investigation of this anomalous behavior in four Fortran programs are presented. In some cases a drop of a factor of two in the average real-time memory allotment is observed when the window size is increased. In some instances a bigger real-time memory allotment means an order of magnitude increase in page faults. Walid A. Abu-Sufah, David A. Padua |
IEEE Trans. Software Eng. | 1 |
| 1981 | On the Performance Enhancement of Paging Systems Through Program Analysis and TransformationsabstractIt is possible to improve the paging performance of a program by applying transformations to the source program that improve data access locality. We discuss this subject in general terms, including automation of these transformations, and present a number of such transforms. This is followed by experimental results which indicate that these transformations are indeed effective. Use of practical, simple memory management policies like the fixed allocation local lru replacement algorithm leads to average improvements over untransformed programs of a factor of 10 in space-time cost, and a factor of 5 in memory size. Multiprogramming questions are also discussed. Walid A. Abu-Sufah, David J. Kuck, Duncan H. Lawrie |
IEEE Trans. Computers | 1 |