EDBT 2026 Demo / reviewers in the wild / expert
A. P. Wim Böhm
dblp:b/APWBohm · also A. P. Willem Böhm
· DBLP profile ↗
38ranked-venue papers
11as first author
0since 2021 · last 2007
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 10 first-authorSoftware engineering, systems software and programming languages · 7 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Reconfigurable computing and FPGAs · 32% Processor architecture and microarchitecture · 25% Memory systems · 25% | |
| Software engineering, system software, and programming languages
4 papers |
Compilers and program optimization · 74% Programming languages and type systems · 14% Concurrent programming · 12% |
Topics — the 23 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
domain-specific compilation |
0.0 | 1 | 2003 | Accelerated image processing on FPGAs · IEEE Trans. Image Process. 2003 |
Compilers and program optimization › domain-specific compilation
image processing pipeline compilation |
0.0 | 1 | 2003 | Accelerated image processing on FPGAs · IEEE Trans. Image Process. 2003 |
Reconfigurable computing and FPGAs › FPGA-based signal processing
FPGA-based image processing |
0.0 | 1 | 2003 | Accelerated image processing on FPGAs · IEEE Trans. Image Process. 2003 |
Processor architecture and microarchitecture
dataflow architecture |
0.0 | 4 | 1992 | An Analysis of Loop Latency in Dataflow Execution · ISCA 1992 A Quantitative Analysis of Locality in Dataflow Programs · MICRO 1991 Iterative Instructions in the Manchester Dataflow Computer · IEEE Trans. Parallel Distributed Syst. 1990 |
Performance modeling and evaluation
workload characterization |
0.0 | 2 | 1993 | An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993 A Quantitative Analysis of Locality in Dataflow Programs · MICRO 1991 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.0 | 1 | 2003 | Accelerated image processing on FPGAs · IEEE Trans. Image Process. 2003 |
Concurrent programming › concurrency models
nondeterminism |
0.0 | 1 | 1994 | Two Issues in Parallel Language Design · ACM Trans. Program. Lang. Syst. 1994 |
Programming languages and type systems › language design
parallel language design |
0.0 | 1 | 1994 | Two Issues in Parallel Language Design · ACM Trans. Program. Lang. Syst. 1994 |
Memory systems
cache |
0.0 | 1 | 1993 | An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993 |
Memory systems › cache
cache miss |
0.0 | 1 | 1993 | An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993 |
Memory systems
memory referencing behavior |
0.0 | 1 | 1993 | An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993 |
Memory systems › cache
prefetching |
0.0 | 1 | 1993 | An evaluation of bottom-up and top-down thread generation techniques · MICRO 1993 |
Performance modeling and evaluation › workload characterization
locality analysis |
0.0 | 1 | 1991 | A Quantitative Analysis of Locality in Dataflow Programs · MICRO 1991 |
Processor architecture and microarchitecture › dataflow architecture
dataflow machine |
0.0 | 1 | 1990 | Iterative Instructions in the Manchester Dataflow Computer · IEEE Trans. Parallel Distributed Syst. 1990 |
Processor architecture and microarchitecture › microprocessor design › processor core design
instruction execution |
0.0 | 1 | 1990 | Iterative Instructions in the Manchester Dataflow Computer · IEEE Trans. Parallel Distributed Syst. 1990 |
Compilers and program optimization
compiler optimization |
0.0 | 1 | 1989 | Code Optimization for Tagged-Token Dataflow Machines · IEEE Trans. Computers 1989 |
Compilers and program optimization › compiler optimization
dataflow optimization |
0.0 | 1 | 1989 | Code Optimization for Tagged-Token Dataflow Machines · IEEE Trans. Computers 1989 |
Programming languages and type systems › language semantics › formal semantics
denotational semantics |
0.0 | 1 | 1985 | The Denotational Semantics of Dynamic Networks of Processes · ACM Trans. Program. Lang. Syst. 1985 |
Programming languages and type systems
language semantics |
0.0 | 1 | 1985 | The Denotational Semantics of Dynamic Networks of Processes · ACM Trans. Program. Lang. Syst. 1985 |
Concurrent programming › concurrency theory
process calculi |
0.0 | 1 | 1985 | The Denotational Semantics of Dynamic Networks of Processes · ACM Trans. Program. Lang. Syst. 1985 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 1 | 1992 | An Analysis of Loop Latency in Dataflow Execution · ISCA 1992 |
Parallel and multicore computing › parallel programming models › dataflow programming
dataflow parallelism |
0.0 | 1 | 1990 | Iterative Instructions in the Manchester Dataflow Computer · IEEE Trans. Parallel Distributed Syst. 1990 |
Parallel and multicore computing
dataflow computing |
0.0 | 1 | 1989 | Code Optimization for Tagged-Token Dataflow Machines · IEEE Trans. Computers 1989 |
Methods — techniques the papers use, named apart from their topics
single assignment c · 0.1SA-C compiler · 0.1profiling · 0.0cache simulation · 0.0simulation · 0.0benchmarking · 0.0quantitative analysis · 0.0hardware performance measurement · 0.0architectural simulation · 0.0formal semantics · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2007 | The Case for Dynamic Execution on Dynamic HardwareabstractWe present a dynamic dataflow execution model, called the aggregated hierarchical abstract hardware architecture (or AHAHA), for use in FPGA based applications. High level language implementations targetting FPGAs use either handshaking or control-path solutions to schedule computations. Among the handshaking variety are the Compaan VHDL Visitor; the control-path solution uses a simplified form of handshaking which operates only in the same direction as the data flows and is used for example in MAP-C. Charlie Ross, A. P. Wim Böhm |
FCCM | 2 |
| 2006 | A Co-Verification Tool for a High Level Language Compiler for FPGAsabstractThe authors have described a method of testing various implementations of co-designs generated by the SA-C compiler. Each form can be examined using co-simulation. The host code is able to communicate with a FPGA board simulated in ModelSim as if it were physical hardware. The co-simulation approach briefly described in this paper allows us to test and analyze all parts of the complete co-design. In essence, the compiler is able to perform automated co-verification for any SA-C program. At the highest level of simulation, it allows functional verification of the VHDL generated by the compiler. At the lowest level of detail, the FPGA simulation is phase accurate and mimics the hardware behavior down to the individual configurable logic block Charlie Ross, A. P. Wim Böhm |
FCCM | 2 |
| 2004 | Using FIFOs in Hardware-Software Co-Design for FPGA Based Embedded SystemsabstractThis paper deals with the evaluation of FIFO for connecting custom cores to processors in FPGA. Two case studies were done in order to evaluate the use of FIFO in the context of hardware-software co-design. One case study analyzes the rendering images of the Mandelbrot set and the other one deals with the game of Connect Four. In both applications a simple co-processor provides the most speed-up, where a higher quality co-processor implementation was exhibited. Charlie Ross, A. P. Wim Böhm |
FCCM | 2 |
| 2003 | Automatic compilation to a coarse-grained reconfigurable system-opn-chipabstractThe rapid growth of device densities on silicon has made it feasible to deploy reconfigurable hardware as a highly parallel computing platform. However, one of the obstacles to the wider acceptance of this technology is its programmability. The application needs to be programmed in hardware description languages or an assembly equivalent, whereas most application programmers are used to the algorithmic programming paradigm. SA-C has been proposed as an expression-oriented language designed to implicitly express data parallel operations. The Morphosys project proposes an SoC architecture consisting of reconfigurable hardware that supports a data-parallel, SIMD computational model. This paper describes a compiler framework to analyze SA-C programs, perform optimizations, and automatically map the application onto the Morphosys architecture. The mapping process is static and it involves operation scheduling, processor allocation and binding, and register allocation in the context of the Morphosys architecture. The compiler also handles issues concerning data streaming and caching in order to minimize data transfer overhead. We have compiled some important image-processing kernels, and the generated schedules reflect an average speedup in execution times of up to 6× compared to the execution on 800 MHz Pentium III machines. Girish Venkataramani, Walid A. Najjar, Fadi J. Kurdahi, Nader Bagherzadeh, A. P. Wim Böhm, Jeffrey Hammes |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2003 | Accelerated image processing on FPGAsabstractThe Cameron project has developed a language called single assignment C (SA-C), and a compiler for mapping image-based applications written in SA-C to field programmable gate arrays (FPGAs). The paper tests this technology by implementing several applications in SA-C and compiling them to an Annapolis Microsystems (AMS) WildStar board with a Xilinx XV2000E FPGA. The performance of these applications on the FPGA is compared to the performance of the same applications written in assembly code or C for an 800 MHz Pentium III. (Although no comparison across processors is perfect, these chips were the first of their respective classes fabricated at 0.18 microns, and are therefore of comparable ages.) We find that applications written in SA-C and compiled to FPGAs are between 8 and 800 times faster than the equivalent program run on the Pentium III. Bruce A. Draper, J. Ross Beveridge, A. P. Wim Böhm, Charlie Ross, Monica Chawathe |
IEEE Trans. Image Process. | 3 |
| 2002 | Compiling ATR Probing Codes for Execution on FPGA HardwareabstractThis paper describes the implementation of an automatic target recognition (ATR) Probing algorithm on a reconfigurable system, using the SA-C programming language and optimizing compiler. The reconfigurable system is 800 times faster than a comparable Pentium running a C implementation of the same probing task. The reasons for this are analyzed. A. P. Wim Böhm, J. Ross Beveridge, Bruce A. Draper, Charlie Ross, Monica Chawathe, Walid A. Najjar |
FCCM | 1 |
| 2002 | Mapping a Single Assignment Programming Language to Reconfigurable Systems
A. P. Wim Böhm, Jeffrey Hammes, Bruce A. Draper, Monica Chawathe, Charlie Ross, Robert Rinker, Walid A. Najjar |
J. Supercomput. | 1 |
| 2001 | A compiler framework for mapping applications to a coarse-grained reconfigurable computer architectureabstractThe rapid growth of silicon densities has made it feasible to deploy reconfigurable hardware as a highly parallel computing platform. However, in most cases, the application needs to be programmed in hardware description or assembly languages, whereas most application programmers are familiar with the algorithmic programming paradigm. SA-C has been proposed as an expression-oriented language designed to implicitly express data parallel operations. Morphosys is a reconfigurable system-on-chip architecture that supports a data-parallel, SIMD computational model. This paper describes a compiler framework to analyze SA-C programs, perform optimizations, and map the application onto the Morphosys architecture. The mapping process involves operation scheduling, resource allocation and binding and register allocation in the context of the Morphosys architecture. The execution times of some compiled image-processing kernels can achieve up to 42x speed-up over an 800 MHz Pentium III machine. Girish Venkataramani, Walid A. Najjar, Fadi J. Kurdahi, Nader Bagherzadeh, A. P. Wim Böhm |
CASES | 5 |
| 2001 | One-Step Compilation of Image Processing Applications to FPGAs
A. P. Wim Böhm, Bruce A. Draper, Walid A. Najjar, Jeffrey Hammes, Robert Rinker, Monica Chawathe, Charlie Ross |
FCCM | 1 |
| 2001 | Loop fusion and temporal common subexpression elimination in window-based loops
Jeffrey Hammes, A. P. Wim Böhm, Charlie Ross, Monica Chawathe, Bruce A. Draper, Robert Rinker, Walid A. Najjar |
IPDPS | 2 |
| 2001 | Resource Management in Dataflow-Based Multithreaded Execution
Lucas Roh, Bhanu Shankar, A. P. Wim Böhm, Walid A. Najjar |
J. Parallel Distributed Comput. | 3 |
| 2001 | An automated process for compiling dataflow graphs into reconfigurable hardwareabstractWe describe a system, developed as part of the Cameron project, which compiles programs written in a single-assignment subset of C called SA-C into dataflow graphs and then into VHDL. The primary application domain is image processing. The system consists of an optimizing compiler which produces dataflow graphs and a dataflow graph to VHDL translator. The method used for the translation is described here, along with some results on an application. The objective is not to produce yet another design entry tool, but rather to shift the programming paradigm from HDLs to an algorithmic level, thereby extending the realm of hardware design to the application programmer. Robert Rinker, M. Carter, A. Patel, Monica Chawathe, Charlie Ross, Jeffrey Hammes, Walid A. Najjar, A. P. Wim Böhm |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2000 | Compiling Image Processing Applications to Reconfigurable HardwareabstractThis paper describes the compilation of high-level language programs written in a single-assignment language called SA-C into the binary codes used for programming reconfigurable hardware. The primary application domain is image processing. The paper describes the SA-C language, the compiler and the optimizations it performs, the process of converting the intermediate form called dataflow graphs into VHDL, and the generation of hardware configuration codes. Performance data on a typical image processing program, written in SA-C and executed on a reconfigurable computing system, is presented and compared to a hand-written VHDL version and a C version running on conventional processors. Robert Rinker, Jeffrey Hammes, Walid A. Najjar, A. P. Wim Böhm, Bruce A. Draper |
ASAP | 4 |
| 1999 | Sassy: A Language and Optimizing Compiler for Image Processing on Reconfigurable Computing Systems
Jeffrey Hammes, Bruce A. Draper, A. P. Wim Böhm |
ICVS | 3 |
| 1999 | On the performance of pure and impure parallel functional programs
A. P. Wim Böhm, Jeffrey Hammes, Sumit Sur |
Parallel Comput. | 1 |
| 1997 | On the Effectiveness of Functional Language Features: NAS Benchmark FTabstractIn this paper we investigate the effectiveness of functional language features when writing scientific codes. Our programs are written in the purely functional subset of Id and executed on a one node Motorola Monsoon machine, and in Haskell and executed on a Sparc 2. In the application we study – the NAS FT benchmark, a three-dimensional heat equation solver – it is necessary to target and select one-dimensional sub-arrays in three-dimensional arrays. Furthermore, it is important to be able to share computation in array definitions. We compare first order and higher order implementations of this benchmark. The higher order version uses functions to select one-dimensional sub-arrays, or slices, from a three-dimensional object, whereas the first order version creates copies to achieve the same result. We compare various representations of a three-dimensional object, and study the effect of strictness in Haskell. We also study the performance of our codes when employing recursive and iterative implementations of the one-dimensional FFT, which forms the kernel of this benchmark. It turns out that these languages still have quite inefficient implementations, with respect to both space and time. For the largest problem we could run (32 3 ), Haskell is 15 times slower than Fortran and uses three times more space than is absolutely necessary, whereas Id on Monsoon uses nine times more cycles than Fortran on the MIPS R3000, and uses five times more space than is absolutely necessary. This code, and others like it, should inspire compiler writers to improve the performance of functional language implementations. Jeffrey Hammes, Sumit Sur, A. P. Wim Böhm |
J. Funct. Program. | 3 |
| 1996 | Generation, Optimization, and Evaluation of Multithreaded Code
Lucas Roh, Walid A. Najjar, Bhanu Shankar, A. P. Wim Böhm |
J. Parallel Distributed Comput. | 4 |
| 1995 | Control of loop parallelism in multithreaded code
Bhanu Shankar, Lucas Roh, A. P. Wim Böhm, Walid A. Najjar |
PACT | 3 |
| 1995 | Reducing Communication by Honoring Multiple AlignmentsabstractData Decomposition involves the mapping of array to processors of a Distributed Memory Machine goal to obtain the best possible performance of a elements with the program by keeping communication costs-low while exploiting 'parallelism.Data decomposition is typically divided into two subproblems: alignment and partitioning.Alignment deals with the relative allocation of different arrays.Partitioning is concerned with the actual distribution of the array elements among processors.Conflicting alignments may cause communication.This paper presents a technique for reducing communication by honoring multiple alignments and applies this approach in a distributed memory implementation of the strict functional language Sisal..Multiple alignment leads to recomputation and replication of array elements, which is safe in a functional, and hence side effect free, setting.We present performance improvements of up to 80~o for one dimensional arrays, and up to 50% for two dimension al arrays, compared to single alignment implementations on a cluster of workstations. David A. Garza-Salazar, A. P. Wim Böhm |
International Conference on Supercomputing | 2 |
| 1995 | Comparing Id and Haskell in a Monte Carlo Photon Transport CodeabstractAbstract In this paper we present functional Id and Haskell versions of a large Monte Carlo radiation transport code, and compare the two languages with respect to their expressiveness. Monte Carlo transport simulation exercises such abilities as parsing, input/output, recursive data structures and traditional number crunching, which makes it a good test problem for languages and compilers. Using some code examples, we compare the programming styles encouraged by the two languages. In particular, we discuss the effect of laziness on programming style. We point out that resource management problems currently prevent running realistically large problem sizes in the functional versions of the code. Jeffrey Hammes, Olaf M. Lubeck, A. P. Wim Böhm |
J. Funct. Program. | 3 |
| 1995 | Exploiting Data Structure Locality in the Dataflow Model
William Marcus Miller, Walid A. Najjar, A. P. Wim Böhm |
J. Parallel Distributed Comput. | 3 |
| 1994 | A model for dataflow based vector executionabstractAlthough the dataflow model has been shown to allow the exploitation of parallelism at all levels, research of the past decade has revealed several fundamental problems: Synchronization at the instruction level, token matching, coloring and re-labeling operations have a negative impact on performance by significantly increasing the number of non-compute “overhead” cycles. Recently, many novel Hybrid von-Neumann Data Driven machines have been proposed to alleviate some of these problems. The major objective has been to reduce or eliminate unnecesssary synchronization costs through simplified operand matching schemes and increased task granularity. Moreover, the results from recent studies quantifying locality suggest sufficient spatial and temporal locality is present in dataflow execution to merit its exploitation. William Marcus Miller, Walid A. Najjar, A. P. Wim Böhm |
International Conference on Supercomputing | 3 |
| 1994 | Analysis of non-strict functional implementations of the Dongarra-Sorensen eigensolverabstractWe study the producer-consumer parallelism of Eigensolvers composed of a tridiagonalization function, a tridiagonal solver, and a matrix multiplication, written in the non-strict functional programming language Id. We verify the claim that non-strict functional languages allow the natural exploitation of this type of parallelism, in the framework of realistic numerical codes. We compare the standard top-down Dongarra-Sorensen solver with a new, bottom-up version. We show that this bottom-up implementation is much more space efficient than the top-down version. Also, we compare both versions of the Dongarra-Sorensen solver with the more traditional QL algorithm, and verify that the Dongarra-Sorensen solver is much more efficient, even when run in a serial mode. We show that in a non-strict functional execution model, the Dongarra-Sorensen algorithm can run completely in parallel with the Householder function. Moreover, this can be achieved without any change in the code components. We also indicate how the critical path of the complete Eigensolver can be improved. Sumit Sur, A. P. Wim Böhm |
International Conference on Supercomputing | 2 |
| 1994 | Uniqueness and Completeness Analysis of Array Comprehensions
David A. Garza-Salazar, A. P. Wim Böhm |
SAS | 2 |
| 1994 | Two Issues in Parallel Language DesignabstractIn this article, we discuss two programming language features that have value for expressibility and efficiency: nonstrictness and nondeterminism. Our work arose while assessing ways to enhance a currently successful language, SISAL [McGraw et al. 1985]. The questions of how best to include these features, if at all, has led not to conclusions but to an impetus to explore the answers in an objective way. We will retain strictness for efficiency reasons and explore the limits it may impose, and we will experiment with a carefully controlled form of nondeterminism to assess its expressive power. A. P. Wim Böhm, R. R. Oldehoeft |
ACM Trans. Program. Lang. Syst. | 1 |
| 1993 | An evaluation of bottom-up and top-down thread generation techniquesabstractDue to increasing cache-miss latencies, cache control instructions are being implemented for future systems. The authors study the memory referencing behavior of individual machine-level instructions using simulations of fully-associative caches under MIN replacement. Their objective is to obtain a deeper understanding of useful program behavior that can be eventually employed at optimizing programs and to motivate architectural features aimed at improving the efficacy of memory hierarchies. The simulation results show that a very small number of load/store instructions account for a majority of data cache misses. Specifically, fewer than 10 instructions account for half the misses for six out of nine SPEC89 benchmarks. Selectively prefetching data referenced by a small number of instructions identified through profiling can reduce overall miss ratio significantly while only incurring a small number of unnecessary prefetches.> A. P. Wim Böhm, Walid A. Najjar, Bhanu Shankar, Lucas Roh |
MICRO | 1 |
| 1993 | The Dataflow Time and Space Complexity of FFTs
A. P. Wim Böhm, Robert E. Hiromoto |
J. Parallel Distributed Comput. | 1 |
| 1993 | A Quantitative Analysis of Dataflow Program Execution - Preliminaries to a Hybrid Design
Walid A. Najjar, A. P. Wim Böhm, William Marcus Miller |
J. Parallel Distributed Comput. | 2 |
| 1992 | An Analysis of Loop Latency in Dataflow ExecutionabstractRecent evidence indicates that the exploitation of locality in dataflow programs could have a dramatic impact on performance. The current trend in the design of dataflow processors suggest a synthesis of traditional non-strict fine grain instruction execution and a strict coarse grain execution in order to exploit locality. While an increase in instruction granularity will favor the exploitation of locality within a single execution thread, the resulting grain size may increase latency among execution threads. In this paper, the resulting latency incurred through the partitioning of fine grain instructions to quantify coarse grain input and output latencies using a set of numeric benchmarks. The results offer compelling evidence that the inner loops of a significant number of numeric codes would benefit from coarse grain execution. Based on cluster execution times, more than 60% of the measured benchmarks favor a coarse grain execution. IN 64% of the cases the input latency to the cluster is the same in coarse or fine grain execution modes. The results suggest that the effects of increased instruction granularity on latency is minimal for a high percentage of the measured codes, and in large part is offset by available intra-thread locality. Furthermore, simulation results indicate that strict or non-strict data structure access does not change the basic cluster characteristics. Walid A. Najjar, William Marcus Miller, A. P. Wim Böhm |
ISCA | 3 |
| 1992 | Dataflow Parallelism in Genetic Algorithms
V. Scott Gordon 0001, L. Darrell Whitley, A. P. Wim Böhm |
PPSN | 3 |
| 1991 | A Quantitative Analysis of Locality in Dataflow ProgramsabstractArticle A quantitative analysis of locality in dataflow programs Share on Authors: William Marcus Miller Department of Computer Science, Colorado State University, Ft. Collins, CO Department of Computer Science, Colorado State University, Ft. Collins, COView Profile , Walid A. Najjar Department of Computer Science, Colorado State University, Ft. Collins, CO Department of Computer Science, Colorado State University, Ft. Collins, COView Profile , A. P. Wim Böhm Department of Computer Science, Colorado State University, Ft. Collins, CO Department of Computer Science, Colorado State University, Ft. Collins, COView Profile Authors Info & Claims MICRO 24: Proceedings of the 24th annual international symposium on MicroarchitectureSeptember 1991 Pages 12–18https://doi.org/10.1145/123465.123469Online:01 September 1991Publication History 8citation242DownloadsMetricsTotal Citations8Total Downloads242Last 12 Months6Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access William Marcus Miller, Walid A. Najjar, A. P. Wim Böhm |
MICRO | 3 |
| 1990 | Iterative Instructions in the Manchester Dataflow ComputerabstractThe authors investigate the nature and extent of the benefits and adverse effects of iterative instructions in the prototype Manchester Dataflow Computer. Iterative instructions are shown to be highly beneficial in terms of the number of instructions executed and the number of tokens transferred between modules during a program run. This benefit is apparent at hardware level, giving significantly reduced program execution times. However, the full benefits are not realized due to interference between lengthy iterative instructions. It is suggested that restructuring of buffers and the function unit array in the prototype hardware configuration can reduce this interference. Other possibilities for improvement are suggested. For example, the slowdown effect observed in hardware speedup curves could be tackled by treating iterative instructions differently from fine-grain instructions. An alternative structure for the processing element in which certain function units are specialized for executing iterative instructions is being investigated in this connection.> A. P. Wim Böhm, John R. Gurd |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1989 | The Effect of Iterative Instructions in Dataflow Computers
A. P. Wim Böhm, John R. Gurd, Yong Meng Teo |
ICPP (1) | 1 |
| 1989 | Code Optimization for Tagged-Token Dataflow MachinesabstractThe efficiency of dataflow code generated from a high-level language can be considerably improved by both conventional and dataflow-specific optimizations. Such techniques are used in implementing the single-assignment language SISAL on the Manchester Dataflow Machine. The quality of code generated for numeric applications can be measured in terms of the ratio of total number of instructions executed to floating-point operations: the MIPS/MFLOPS ratio. Relevant features of the general-purpose single-assignment language SISAL and the Manchester Dataflow Machine are introduced. An assessment of the initial SISAL implementation shows it to be very expensive. A range of optimizations is then described.> A. P. Wim Böhm, John Sargeant |
IEEE Trans. Computers | 1 |
| 1987 | Tools for Performance Evaluation of Parallel Machines
A. P. Wim Böhm, John R. Gurd |
ICS | 1 |
| 1987 | Performance issues in dataflow machines
John R. Gurd, A. P. Wim Böhm, Yong Meng Teo |
Future Gener. Comput. Syst. | 2 |
| 1985 | The Denotational Semantics of Dynamic Networks of ProcessesabstractDNP (dynamic networks of processes) is a variant of the language introduced by Kahn and MacQueen [11, 12]. In the language it is possible to create new processes dynamically. We present a complete, formal denotational semantics for the language, along the lines sketched by Kahn and MacQueen. An informal explanation of the formal semantics is also given. Arie de Bruin, A. P. Wim Böhm |
ACM Trans. Program. Lang. Syst. | 2 |
| 1978 | Guidelines for Software PortabilityabstractAbstract The areas in which programs are most unlikely to be portable are discussed. Attention is paid to programming languages, operating systems, file systems, I/O device characteristics, machine architecture and documentation. Pitfalls are indicated and in some cases solutions are suggested. Andrew S. Tanenbaum, Paul Klint, A. P. Wim Böhm |
Softw. Pract. Exp. | 3 |