Wim Vanderbauwhede

dblp:11/2815 · DBLP profile ↗
← Back
44ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0001-6768-0037ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 5 first-author · 7 since 2021Software engineering, systems software and programming languages · 8 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2026 A Streaming FPGA Architecture for Sparse Matrix Multiplication with Sparsity-Aware Data Reordering
abstract
Sparse matrix–dense matrix multiplication (SpMM) is a fundamental kernel in graph analytics, graph neural networks, and many sparse deep learning workloads. However, its irregular sparsity patterns make efficient execution on FPGAs challenging. Conventional CSR- or COO-based traversal often leads to irregular control flow, fragmented memory accesses, and limited reuse of dense features, which together reduce pipeline efficiency and make it difficult to sustain high throughput. This work addresses these limitations through a host–FPGA co-design that converts irregular sparse computation into streamable tile-level execution units. By combining sparsity-aware data reordering with an execution-oriented sparse representation, the proposed approach enables more regular data movement and better hardware utilization.
Shuxuan Li, Nikela Papadopoulou, Wim Vanderbauwhede
FCCM3
2026 Modeling Scenarios for Carbon-Aware Geographic Load Shifting of Compute Workloads
abstract
We present an analytical model to evaluate the reductions in emissions resulting from geographic load shifting. This model is optimistic as it ignores issues of grid capacity, demand and curtailment. In other words, real-world reductions will be smaller than the estimates. However, even with these assumptions, the presented scenarios show that the realistic reductions from carbon-aware geographic load shifting are small, of the order of 5\%. This is not enough to compensate the growth in emissions from global data centre expansion.
Wim Vanderbauwhede
IEEE Trans. Sustain. Comput.1
2025 Compiler Support for Speculation in Decoupled Access/Execute Architectures
abstract
Irregular codes are bottlenecked by memory and communication latency. Decoupled access/execute (DAE) is a common technique to tackle this problem. It relies on the compiler to separate memory address generation from the rest of the program, however, such a separation is not always possible due to control and data dependencies between the access and execute slices, resulting in a loss of decoupling. In this paper, we present compiler support for speculation in DAE architectures that preserves decoupling in the face of control dependencies. We speculate memory requests in the access slice and poison mis-speculations in the execute slice without the need for replays or synchronization. Our transformation works on arbitrary, reducible control flow and is proven to preserve sequential consistency. We show that our approach applies to a wide range of architectural work on CPU/GPU prefetchers, CGRAs, and accelerators, enabling DAE on a wider range of codes than before.
Robert Szafarczyk, Syed Waqar Nabi, Wim Vanderbauwhede
CC3
2025 Dynamic Loop Fusion in High-Level Synthesis
abstract
Dynamic High-Level Synthesis (HLS) uses additional hardware to perform memory disambiguation at runtime, increasing loop throughput in irregular codes compared to static HLS. However, most irregular codes consist of multiple sibling loops, which currently have to be executed sequentially by all HLS tools. Static HLS performs loop fusion only on regular codes, while dynamic HLS relies on loops with dependencies to run to completion before the next loop starts.
Robert Szafarczyk, Syed Waqar Nabi, Wim Vanderbauwhede
FPGA3
2023 Wiring Circuits Is Easy as {0, 1, ω}, or Is It
Jan de Muijnck-Hughes, Wim Vanderbauwhede
ECOOP2
2023 Dynamically Scheduled Memory Operations in Static High-Level Synthesis
abstract
Dynamically scheduled high-level synthesis (HLS) achieves higher throughput on codes with unpredictable memory accesses compared to static HLS. However, dynamic scheduling results in circuits that use more resources and have a slower critical path, even if only a small part of the circuit exhibits dynamic behavior. In this extended abstract, we propose to introduce dynamically scheduled memory operations into static HLS. Our goal is to reach the same throughput as dynamic HLS on codes with irregular memory accesses while achieving comparable resource usage and critical paths as static HLS.
Robert Szafarczyk, Syed Waqar Nabi, Wim Vanderbauwhede
FCCM3
2023 Compiler Discovered Dynamic Scheduling of Irregular Code in High-Level Synthesis
abstract
Dynamically scheduled high-level synthesis (HLS) achieves higher throughput than static HLS for codes with unpredictable memory accesses and control flow. However, excessive dataflow scheduling results in circuits that use more resources and have a slower critical path, even when only a part of the circuit exhibits dynamic behavior. Recent work has shown that marking parts of a dataflow circuit for static scheduling can save resources and improve performance (hybrid scheduling), but the dynamic part of the circuit still bottlenecks the critical path. We propose instead to selectively introduce dynamic scheduling into static HLS. This paper presents an algorithm for identifying code regions amenable to dynamic scheduling and shows a methodology for introducing dynamically scheduled basic blocks, loops, and memory operations into static HLS. Our algorithm is informed by modulo-scheduling and can be integrated into any modulo-scheduled HLS tool. On a set of ten benchmarks, we show that our approach achieves on average an up to 3.7× and 3× speedup against dynamic and hybrid scheduling, respectively, with an area overhead of 1.3× and frequency degradation of 0.74× when compared to static HLS.
Robert Szafarczyk, Syed Waqar Nabi, Wim Vanderbauwhede
FPL3
2022 Reducing FPGA Memory Footprint of Stencil Codes through Automatic Extraction of Memory Patterns
abstract
FPGAs are attractive for scientific high-performance computing due to their potential for high performance-per-Watt. Stencil codes in scientific applications are difficult to optimize on FPGAs, because of redundant, non-contiguous memory accesses to relatively low bandwidth DRAM. In this paper, we present an algorithm to aggressively reduce on-chip block RAM (BRAM) and off-chip DRAM utilisation of stencil codes running on FPGAs. The algorithm extracts memory accesses from computational pipelines and removes all redundant intermediate arrays, including those used for stencil buffering, by trading DRAM accesses for computation. The algorithm is based on rewrite-rules on a strict functional representation derived from Fortran code and generates provably correct, optimized code. Typical FPGA implementations store the stencil window in on-chip shift registers implemented in BRAMs; we use only DRAM and optimize the memory accesses instead. Our approach dramatically reduces BRAM usage so that the domain size is only limited by available DRAM. We report a drop of 78% and 18% in BRAM usage in 3-D and 2-D stencil codes compared to a manual implementation using shift registers while staying competitive in performance or even improving performance-per-Watt.
Robert Szafarczyk, Syed Waqar Nabi, Wim Vanderbauwhede
FPL3
2022 Making legacy Fortran code type safe through automated program transformation
abstract
Abstract Fortran is still widely used in scientific computing, and a very large corpus of legacy as well as new code is written in FORTRAN 77. In general this code is not type safe, so that incorrect programs can compile without errors. In this paper, we present a formal approach to ensure type safety of legacy Fortran code through automated program transformation. The objective of this work is to reduce programming errors by guaranteeing type safety. We present the first rigorous analysis of the type safety of FORTRAN 77 and the novel program transformation and type checking algorithms required to convert FORTRAN 77 subroutines and functions into pure, side-effect free subroutines and functions in Fortran 90. We have implemented these algorithms in a source-to-source compiler which type checks and automatically transforms the legacy code. We show that the resulting code is type safe and that the pure, side-effect free and referentially transparent subroutines can readily be offloaded to accelerators.
Wim Vanderbauwhede
J. Supercomput.1
2020 A Framework for Resource Dependent EDSLs in a Dependently Typed Language (Pearl)
abstract
Idris' Effects library demonstrates how to embed resource dependent algebraic effect handlers into a dependently typed host language, providing run-time and compile-time based reasoning on type-level resources. Building upon this work, Resources is a framework for realising Embedded Domain Specific Languages (EDSLs) with type systems that contain domain specific substructural properties. Differing from Effects, Resources allows a language’s substructural properties to be encoded within type-level resources that are associated with language variables. Such an association allows for multiple effect instances to be reasoned about autonomically and without explicit type-level declaration. Type-level predicates are used as proof that the language’s substructural properties hold. Several exemplar EDSLs are presented that illustrates our framework’s operation and how dependent types provide correctness-by-construction guarantees that substructural properties of written programs hold.
Jan de Muijnck-Hughes, Edwin C. Brady, Wim Vanderbauwhede
ECOOP3
2019 A Typing Discipline for Hardware Interfaces
abstract
Modern Systems-on-a-Chip (SoC) are constructed by composition of IP (Intellectual Property) Cores with the communication between these IP Cores being governed by well described interaction protocols. However, there is a disconnect between the machine readable specification of these protocols and the verification of their implementation in known hardware description languages. Although tools can be written to address such separation of concerns, the tooling is often hand written and used to check hardware designs a posteriori. We have developed a dependent type-system and proof-of-concept modelling language to reason about the physical structure of hardware interfaces using user provided descriptions. Our type-system provides correct-by-construction guarantees that the interfaces on an IP Core will be well-typed if they adhere to a specified standard.
Jan de Muijnck-Hughes, Wim Vanderbauwhede
ECOOP2
2019 FPGA design space exploration for scientific HPC applications using a fast and accurate cost model based on roofline analysis
abstract
High-performance computing on heterogeneous platforms in general and those with FPGAs in particular presents a significant programming challenge. We contend that compiler technology has to evolve to automatically optimize applications by transforming a given original program. We are developing a novel methodology based on type transformations on a functional description of a given scientific kernel, for generating correct-by-construction design variants. An associated lightweight costing mechanism for evaluating these variants is a cornerstone of our methodology, and the focus of this paper. We discuss our use of the roofline model to work with our optimizing compiler to enable us to quickly derive accurate estimates of performance from the design’s representation in our custom intermediate language. We show results confirming the accuracy of our cost model by validating it on different scientific kernels. A case study is presented to demonstrate that a solution created from our optimizing framework outperforms commercial high-level synthesis tools both in terms of throughput and power efficiency.
Syed Waqar Nabi, Wim Vanderbauwhede
J. Parallel Distributed Comput.2
2016 Evaluation of the Memory Communication Traffic in a Hierarchical Cache Model for Massively-Manycore Processors
abstract
The scaling of semiconductor technologies is leading to processors with increasing numbers of cores. A key enabler in manycore systems is the use of Networks-on-Chip (NoC) as a global communication mechanism. The adoption of NoCs in manycore systems requires a shift in focus from computation to communication, as communication is fast becoming the dominant factor in processor performance. Many researchers have focused on direct communication between cores in the NoC, however in a manycore processor the communication is actually between the cores and the memory hierarchy. In this work, we investigate the memory communication traffic of shared threads in a hierarchical cache architecture. We argue that the performance scalability for shared-memory applications in a hierarchical cache architecture for systems with thousands of processor cores depends on the distance between threads sharing memory in terms of the cache hierarchy (the "memory distance"). We present latency and throughput results comparing fat quadtree, concentrated mesh and mesh topologies as a function of the "memory distance" between the threads. Our results using the ITRS physical data for 2023 show that the model of thread placement and the distance of placing them significantly affects the NoC performance, and that scale-invariant topologies perform better than flat topologies.
Sharifa Al Khanjari, Wim Vanderbauwhede
PDP2
2016 An analysis of the feasibility and benefits of GPU/multicore acceleration of the Weather Research and Forecasting model
abstract
Summary There is a growing need for ever more accurate climate and weather simulations to be delivered in shorter timescales, in particular, to guard against severe weather events such as hurricanes and heavy rainfall. Due to climate change, the severity and frequency of such events – and thus the economic impact – are set to rise dramatically. Hardware acceleration using graphics processing units (GPUs) or Field‐Programmable Gate Arrays (FPGAs) could potentially result in much reduced run times or higher accuracy simulations. In this paper, we present the results of a study of the Weather Research and Forecasting (WRF) model undertaken in order to assess if GPU and multicore acceleration of this type of numerical weather prediction (NWP) code is both feasible and worthwhile. The focus of this paper is on acceleration of code running on a single compute node through offloading of parts of the code to an accelerator such as a GPU. The governing equations set of the WRF model is based on the compressible, non‐hydrostatic atmospheric motion with multi‐physics processes. We put this work into context by discussing its more general applicability to multi‐physics fluid dynamics codes: in many fluid dynamics codes, the numerical schemes of the advection terms are based on finite differences between neighboring cells, similar to the WRF code. For fluid systems including multi‐physics processes, there are many calls to these advection routines. This class of numerical codes will benefit from hardware acceleration. We studied the performance of the original code of the WRF model and proposed a simple model for comparing multicore CPU and GPU performance. Based on the results of extensive profiling of representative WRF runs, we focused on the acceleration of the scalar advection module. We discuss the implementation of this module as a data‐parallel kernel in both OpenCL and OpenMP. We show that our data‐parallel kernel version of the scalar advection module runs up to seven times faster on the GPU compared with the original code on the CPU. However, as the data transfer cost between GPU and CPU is very high (as shown by our analysis), there is only a small speed‐up (two times) for the fully integrated code. We show that it would be possible to offset the data transfer cost through GPU acceleration of a larger portion of the dynamics code. In order to carry out this research, we also developed an extensible software system for integrating OpenCL code into large Fortran code bases such as WRF. This is one of the main contributions of our work. We discuss the system to show how it allows the replacement of the sections of the original codebase with their OpenCL counterparts with minimal changes – literally only a few lines – to the original code. Our final assessment is that, even with the current system architectures, accelerating WRF – and hence also other, similar types of multi‐physics fluid dynamics codes – with a factor of up to five times is definitely an achievable goal. Accelerating multi‐physics fluid dynamics codes including NWP codes is vital for its application to weather forecasting, environmental pollution warning, and emergency response to the dispersion of hazardous materials. Implementing hardware acceleration capability for fluid dynamics and NWP codes is a prerequisite for up‐to‐date and future computer architectures. Copyright © 2015 John Wiley & Sons, Ltd.
Wim Vanderbauwhede, Tetsuya Takemi
Concurr. Comput. Pract. Exp.1
2015 High Level Programming of Document Classification Systems for Heterogeneous Environments using OpenCL (Abstract Only)
abstract
Document classification is at the heart of several of the applications that have been driving the proliferation of the internet in our daily lives. The ever growing amounts of data and the need for higher throughput, more energy efficient document classification solutions motivated us to investigate alternatives to the traditional homogenous CPU based implementations. We investigate a heterogeneous system where CPUs are combined with FPGAs as system accelerators. Incorporating FPGAs as accelerators in a heterogeneous computing environment allows for the creation of flexible custom hardware solutions that can potentially offer increased power efficiency and performance gains. One of the main issues delaying wide spread adoption of FPGAs as standard heterogeneous system accelerators is the difficulty in programming them. The OpenCL standard offers a unified C programming model for any device that adheres to its standards. An Altera OpenCL FPGA based implementation of a document classification system is investigated in which a stream of HTML documents is scored according to a profile on a document-by-document basis. The results show that the throughput of the document classification application with and without Bloom Filters is 312MB/s and 343MB/s respectively, when running on CPU, and 354MB/s and 452MB/s respectively, when running on an FPGA. Our results also show up to 32% power efficiency improvement for the FPGA implementation over the CPU implementation. We would like to thank Davor Capalija from Altera for his invaluable advice during our work on the FPGA version of the algorithm.
Nasibeh Nasiri, Oren Segal, Martin Margala, Wim Vanderbauwhede, Sai Rahul Chalamalasetti
FPGA4
2015 Number of Tasks, not Threads, is Key
abstract
The concept of task already exists in many parallel programming models. Programmers express parallelism by defining tasks in their applications, and runtime libraries schedule tasks on threads. However, in many task-based parallel programming models, choosing the right number of threads is still key to performance. Hence, the onus is on the programmer to decide not only about the number of tasks, but also about the optimal number of threads in order to get good performance. In this paper, we aim to show that desirable performance can be achieved by only focusing on tasks. For this purpose, we compare a purely task-centric parallel programming model called GPRM with three popular approaches (OpenMP, Intel Cilk Plus, and TBB) on two modern many core systems, the Tilera TILEPro64 and Intel Xeon Phi, which have respectively 64 and 60 physical cores integrated into a single chip. We have chosen three benchmarks with different characteristics to show that a task-centric approach such as GPRM can facilitate parallel programming while it outperforms other models in most cases. It does so by controlling only the number of tasks, rather than having to tune the number of threads.
Ashkan Tousimojarad, Wim Vanderbauwhede
PDP2
2014 A Parallel Task-Based Approach to Linear Algebra
abstract
Processors with large numbers of cores are becoming commonplace. In order to take advantage of the available resources in these systems, the programming paradigm has to move towards increased parallelism. However, increasing the level of concurrency in the program does not necessarily lead to better performance. Parallel programming models have to provide flexible ways of defining parallel tasks and at the same time, efficiently managing the created tasks. OpenMP is a widely accepted programming model for shared-memory architectures. In this paper we highlight some of the drawbacks in the OpenMP tasking approach, and propose an alternative model based on the Glasgow Parallel Reduction Machine (GPRM) programming framework. As the main focus of this study, we deploy our model to solve a fundamental linear algebra problem, LU factorisation of sparse matrices. We have used the SparseLU benchmark from the BOTS benchmark suite, and compared the results obtained from our model to those of the OpenMP tasking approach. The TILEPro64 system has been used to run the experiments. The results are very promising, not only because of the performance improvement for this particular problem, but also because they verify the task management efficiency, stability, and flexibility of our model, which can be applied to solve problems in future many-core systems.
Ashkan Tousimojarad, Wim Vanderbauwhede
ISPDC2
2013 High throughput filtering using FPGA-acceleration
abstract
With the rise in the amount information of being streamed across networks, there is a growing demand to vet the quality, type and content itself for various purposes such as spam, security and search. In this paper, we develop an energy-efficient high performance information filtering system that is capable of classifying a stream of incoming document at high speed. The prototype parses a stream of documents using a multicore CPU and then performs classification using Field-Programmable Gate Arrays (FPGAs). On a large TREC data collection, we implemented a Naive Bayes classifier on our prototype and compared it to an optimized CPU based-baseline. Our empirical findings show that we can classify documents at 10Gb/s which is up to 94 times faster than the CPU baseline (and up to 5 times faster than previous FPGA based implementations). In future work, we aim to increase the throughput by another order of magnitude by implementing both the parser and filter on the FPGA.
Wim Vanderbauwhede, Anton Frolov 0002, Leif Azzopardi, Sai Rahul Chalamalasetti, Martin Margala
CIKM1
2013 Throughput/Resource-Efficient Reconfigurable Processor for Multimedia Applications
abstract
This brief presents the implementation and evaluation of an 8-bit adaptable processor core to be part of the power-throughput-area efficient multimedia oriented reconfigurable architecture reconfigurable array. The design of the processor core was custom implemented in IBM's 90 nm CMOS technology and occupies 0.115 mm2silicon area with approximately 70% area utilized by core circuits. The processor shows a peak throughput performance of 75 MOPS/mW. Benchmarking results show estimated throughputs of 9.5, 21.36, 39.78, 170.88, and 4.54 MSamples/s for variants of 2-D discrete cosine transform (DCT), 4 × 4 H.264 integer transform, and 2-D discrete wavelet transform, respectively. Our analysis shows that the proposed design provides approximately 4-8 times higher throughput for 2-D DCT when compared against popular architectures.
Sohan Purohit, Sai Rahul Chalamalasetti, Martin Margala, Wim Vanderbauwhede
IEEE Trans. Very Large Scale Integr. Syst.4
2013 Design and Evaluation of High-Performance Processing Elements for Reconfigurable Systems
abstract
In this paper, we present the design and evaluation of two new processing elements for reconfigurable computing. We also present a circuit-level implementation of the data paths in static and dynamic design styles to explore the various performance-power tradeoffs involved. When implemented in IBM 90-nm CMOS process, the 8-b data paths achieve operating frequencies ranging over 1 GHz both for static and dynamic implementations, with each data path supporting single-cycle computational capability. A novel single-precision floating point processing element (FPPE) using a 24-b variant of the proposed data paths is also presented. The full dynamic implementation of the FPPE shows that it operates at a frequency of 1 GHz with 6.5-mW average power consumption. Comparison with competing architectures shows that the FPPE provides two orders of magnitude higher throughput. Furthermore, to evaluate its feasibility as a soft-processing solution, we also map the floating point unit onto the Virtex 4 and 5 devices, and observe that the unit requires less than 1% of the total logic slices, while utilizing only around 4% of the DSP blocks available. When compared against popular field-programmable-gate-array-based floating point units, our design on Virtex 5 showed significantly lower resource utilization, while achieving comparable peak operating frequency.
Sohan Purohit, Sai Rahul Chalamalasetti, Martin Margala, Wim Vanderbauwhede
IEEE Trans. Very Large Scale Integr. Syst.4
2012 Evaluating FPGA-acceleration for real-time unstructured search
abstract
Emerging data-centric workloads that operate on and harvest useful insights from large amounts of unstructured data require corresponding new data-centric system architecture optimizations. In particular, with the growing importance of power and cooling costs, a key challenge for such future designs is to achieve increased performance at high energy efficiency. At the same time, recent trends towards better support for reconfigurable logic enable the use of energy-efficient accelerators. Combining these trends, in this paper, we examine the applicability of acceleration in future data-centric system architectures. We focus on an important class of data-centric workloads, real-time unstructured search, or information filtering, where large collections of documents are scored against specific topic profiles, and present an FPGA-based implementation to accelerate such workloads. Our implementation, based on the GiDEL PROCStar IV board using Altera Stratix IV FPGAs, demonstrates excellent performance and energy efficiency, 20 to 40 times better than baseline server systems for typical usage scenarios. Our results also highlight interesting insights for the design of accelerators in future data-centric systems.
Sai Rahul Chalamalasetti, Martin Margala, Wim Vanderbauwhede, Mitch Wright, Parthasarathy Ranganathan
ISPASS3
2012 Impact of Random Dopant Fluctuations on the Timing Characteristics of Flip-Flops
abstract
In this work, we have analyzed the effects of variability, due to random dopant fluctuation (RDF), on the timing characteristics of flip-flops for the future technology generations of 25, 18, and 13 nm, based on extensive Monte Carlo simulations. The results show that RDF has a significant impact on all of the timing parameters and that these parameters do not follow a normal distribution; in particular, they are skewed and exhibit a large tail. Moreover, the dispersion and skewness of the timing parameters increase with technology scaling. The study of the exact shape of these distributions, especially in the tail section, is of fundamental importance in the design and modeling of high-performance, reliable, and economically feasible circuits. In this paper, the distribution tails are estimated based on simulation data, with the aid of statistical nonparametric probability density functions, and it has been found that timing distributions can better be represented by certain nonparametric distributions, in particular Pearson and Johnson systems. The use of these representations during the statistical static timing analysis will provide more accurate results as compared with the normal approximation of distributions and will eventually reduce the probability of yield loss.
Faiz-ul Hassan, Wim Vanderbauwhede, Fernando Rodríguez Salazar
IEEE Trans. Very Large Scale Integr. Syst.2
2011 An analytical model of broadcast in QoS-aware wormhole-routed NoCs
Mahmoud Moadeli, Wim Vanderbauwhede
J. Syst. Softw.2
2010 A C++-embedded Domain-Specific Language for programming the MORA soft processor array
abstract
MORA is a novel platform for high-level FPGA programming of streaming vector and matrix operations, aimed at multimedia applications. It consists of soft array of pipelined low-complexity SIMD processors-in-memory (PIM). We present a Domain-Specific Language (DSL) for high-level programming of the MORA soft processor array. The DSL is embedded in C++, providing designers with a familiar language framework and the ability to compile designs using a standard compiler for functional testing before generating the FPGA bitstream using the MORA toolchain. The paper discusses the MORA-C++ DSL and the compilation route into the assembly for the MORA machine and provides examples to illustrate the programming model and performance.
Wim Vanderbauwhede, Martin Margala, Sai Rahul Chalamalasetti, Sohan Purohit
ASAP1
2010 Search system requirements of patent analysts
abstract
Patent search tasks are difficult and challenging, often requiring expert patent analysts to spend hours, even days, sourcing relevant information. To aid them in this process, analysts use Information Retrieval systems and tools to cope with their retrieval tasks. With the growing interest in patent search, it is important to determine their requirements and expectations of the tools and systems that they employ. In this poster, we report a subset of the findings of a survey of patent analysts conducted to elicit their search requirements.
Leif Azzopardi, Wim Vanderbauwhede, Hideo Joho
SIGIR2
2010 An analytical performance model for the Spidergon NoC with virtual channels
Mahmoud Moadeli, Alireza Shahrabi, Wim Vanderbauwhede, Partha Maji
J. Syst. Archit.3
2010 Communication modeling of multicast in all-port wormhole-routed NoCs
Mahmoud Moadeli, Wim Vanderbauwhede
J. Syst. Softw.2
2009 Quarc: A High-Efficiency Network on-Chip Architecture
abstract
Th novel Quarc NoC architecture, inspired by the Spidergon scheme is introduced as a NoC architecture that is highly efficient in performing collective communication operations including broadcast and multicast. The efficiency of the Quarc architecture is achieved through balancing the traffic which is the result of the modifications applied to the topology and the routing elements of the Spidergon NoC. This paper provides an ASIC implementation of both architectures using UMCpsilas 0.13 mum CMOS technology and demonstrates an analysis and comparison of the cost and performance between the Quarc and the Spidergon NoCs.
Mahmoud Moadeli, Partha Maji, Wim Vanderbauwhede
AINA3
2009 A Communication Model of Broadcast in Wormhole-Routed Networks on-Chip
abstract
This paper presents a novel analytical model to compute communication latency of broadcast as the most fundamental collective communication operation. The novelty of the model lies in its ability to predict the broadcast communication latency in wormhole-routed architectures employing asynchronous multi-port routers scheme. The model is applied to the Quarc NoC and its validity is verified by comparing the model predictions against the results obtained from a discrete-event simulator developed using OMNET++.
Mahmoud Moadeli, Wim Vanderbauwhede
AINA2
2009 Architectural Comparison of Instruments for Transaction Level Monitoring of FPGA-Based Packet Processing Systems
abstract
The fine-grained parallelism inherent in FPGAs has encouraged their use in packet processing systems. To facilitate debugging and performance evaluation, designers require on-chip monitors that provide abstractions of low-level details and a system-level perspective. In this paper, we present five architectures that permit transaction-based communication-centric monitoring of packet processing systems. We compare the resource requirements and filtering functionality of each architecture, demonstrating that sequential matching is more resource efficient than parallel matching. We also show that generic filtering has a low overhead compared to specialised filtering while providing additional flexibility. A scalable architecture is also presented, which is more flexible and adaptable to matching requirements than other architectures. These monitoring architectures permit the implementation of a highly effective test system which provides a system-level perspective and is more resource efficient than conventional RTL debug environments.
Paul Edward McKechnie, Michaela Blott, Wim Vanderbauwhede
FCCM3
2009 A low cost reconfigurable soft processor for multimedia applications: Design synthesis and programming model
abstract
This paper presents an FPGA implementation of a low cost 8 bit reconfigurable processor core for media processing applications. The core is optimized to provide all basic arithmetic and logic functions required by the media processing and other domains, as well as to make it easily integrable into a 2D array. This paper presents an investigation of the feasibility of the core as a potential soft processing architecture for FPGA platforms. The core was synthesized on the entire Virtex FPGA family to evaluate its overall performance, scalability and portability. A special feature of the proposed architecture is its simple programming model which allows low level programming. Throughput results for popular benchmarks coded using the programming model and cycle accurate simulator are presented.
Sai Rahul Chalamalasetti, Wim Vanderbauwhede, Sohan Purohit, Martin Margala
FPL2
2009 FPGA-accelerated Information Retrieval: High-efficiency document filtering
abstract
Power consumption in data centres is a growing issue as the cost of the power for computation and cooling has become dominant. An emerging challenge is the development of ldquoenvironmentally friendlyrdquo systems. In this paper we present a novel application of FPGAs for the acceleration of information retrieval algorithms, specifically, filtering streams/collections of documents against topic profiles. Our results show that FPGA acceleration can result in speed-ups of up to a factor 20 for large profiles.
Wim Vanderbauwhede, Leif Azzopardi, Mahmoud Moadeli
FPL1
2009 Design and implementation of the Quarc Network on-Chip
abstract
Networks-on-Chip (NoC) have emerged as alternative to buses to provide a packet-switched communication medium for modular development of large Systems-on-Chip. However, to successfully replace its predecessor, the NoC has to be able to efficiently exchange all types of traffic including collective communications. The latter is especially important for e.g. cache updates in multicore systems. The Quarc NoC architecture has been introduced as a Networks-on-Chip which is highly efficient in exchanging all types of traffic including broadcast and multicast. In this paper we present the hardware implementation of the switch architecture and the network adapter (transceiver) of the Quarc NoC. Moreover, the paper presents an analysis and comparison of the cost and performance between the Quarc and the Spidergon NoCs implemented in Verilog targeting the Xilinx Virtex FPGA family. We demonstrate a dramatic improvement in performance over the Spidergon especially for broadcast traffic, at no additional hardware cost.
Mahmoud Moadeli, Partha Maji, Wim Vanderbauwhede
IPDPS3
2009 A performance model of multicast communication in wormhole-routed networks on-chip
abstract
Collective communication operations form a part of overall traffic in most applications running on platforms employing direct interconnection networks. This paper presents a novel analytical model to compute communication latency of multicast as a widely used collective communication operation. The novelty of the model lies in its ability to predict the latency of the multicast communication in wormhole-routed architectures employing asynchronous multi-port routers scheme. The model is applied to the Quarc NoC and its validity is verified by comparing the model predictions against the results obtained from a discrete-event simulator developed using OMNET++.
Mahmoud Moadeli, Wim Vanderbauwhede
IPDPS2
2009 Debugging FPGA-based packet processing systems through transaction-level communication-centric monitoring
abstract
The fine-grained parallelism inherent in FPGAs has encouraged their use in packet processing systems. Debugging and performance evaluation of such complex designs can be significantly improved through debug information that provides a system-level perspective and hides the complexity of signal-level debugging. In this paper we present a debugging system that permits transaction-based communication-centric monitoring of packet processing systems. We demonstrate, using two different examples, how this system can improve the debugging information and abstract lower level detail. Furthermore, we demonstrate that transaction monitoring systems require fewer resources than conventional RTL debugging systems and can provide a system-level perspective not permitted by traditional tools.
Paul Edward McKechnie, Michaela Blott, Wim Vanderbauwhede
LCTES3
2009 Developing energy efficient filtering systems
abstract
Processing large volumes of information generally requires massive amounts of computational power, which consumes a significant amount of energy. An emerging challenge is the development of ``environmentally friendly'' systems that are not only efficient in terms of time, but also energy efficient. In this poster, we outline our initial efforts at developing greener filtering systems by employing Field Programmable Gate Arrays (FPGA) to perform the core information processing task. FPGAs enable code to be executed in parallel at a chip level, while consuming only a fraction of the power of a standard (von Neuman style) processor. On a number of test collections, we demonstrate that the FPGA filtering system performs 10-20 times faster than the Itanium based implementation, resulting in considerable energy savings.
Leif Azzopardi, Wim Vanderbauwhede, Mahmoud Moadeli
SIGIR2
2008 Modeling Differentiated Services-Based QoS in Wormhole-Routed NoCs
abstract
The applications running on Networks on-Chip (NoC) typically have to meet a minimum performance requirements to perform successfully. Employing differentiated services is a widely used approach to support QoS (Quality of Service) by relatively prioritizing traffic in the networks employing connection-less communication mechanisms. In this paper we present an analytical evaluation of the average message latency for wormhole-routed interconnect architectures exploiting a traffic prioritization mechanism to achieve differentiated services-based QoS. To verify the validity of the model we apply the method to the Spidergon NoC and compare the model against the results obtained from a discrete-event simulator developed using OMNET.
Mahmoud Moadeli, Alireza Shahrabi, Wim Vanderbauwhede, Mohamed Ould-Khaoua
AINA3
2008 A type system for static typing of a domain-specific language
abstract
With the increase in system complexity, designers are increasingly using IP blocks as a means for filling the designer productivity gap. This has given rise to system level languages which connect IP blocks together. However, these languages have in general not been subject to formalisation. They are considered too trivial to justify the formalisation effort. Unfortunately, the lack of formality in these languages can give rise to errors that are not caught until late in the design cycle. We present a type system for static typing of such a system level language. We argue that the proposed type system will eliminate an important class of errors currently permitted by existing system level languages. A comparison is made against existing tools and we show that the type checker detects errors earlier in the design flow. This reduces synthesis iterations and decreases the time to market
Paul Edward McKechnie, Nathan A. Lindop, Wim Vanderbauwhede
FPGA3
2008 Interface and Reconfiguration Controller for a wireless MAC-oriented dynamically reconfigurable hardware co-processor
abstract
To address the challenges of the consumer wireless device industry, we have designed a dynamically reconfigurable architecture with flexibility limited to address the MAC layer. It is a Software/Hardware partitioned platform in which critical tasks are delegated to a dynamically reconfigurable hardware co-processor. It will handle data streams of multiple (up to 3) different protocol standards, by reconfiguring on a packet-by-packet basis. The Interface and Reconfiguration Controller uses a combination of controllers to dynamically reconfigure the functional units in the architecture and delegate MAC tasks to them. Results of packet transmission on a prototype model indicate that the device handles three transmission requests from different protocol modes in a fraction of the packet durations.
Syed Waqar Nabi, Cade C. Wells, Wim Vanderbauwhede
FPL3
2008 Quarc: A Novel Network-On-Chip Architecture
abstract
This paper introduces the Quarc NoC, a novel NoC architecture inspired by the Spidergon NoC. The Quarc scheme significantly outperforms the Spidergon NoC through balancing the traffic which is the result of the modifications applied to the topology and the routing elements.The proposed architecture is highly efficient in performing collective communication operations including broadcast and multicast. We present the topology, routing discipline and switch architecture for the Quarc NoC and demonstrate the performance with the results obtained from discrete event simulations.
Mahmoud Moadeli, Wim Vanderbauwhede, Alireza Shahrabi
ICPADS2
2008 A Performance Model of Communication in the Quarc NoC
abstract
Networks on-chip (NoC) emerged as a promising communication medium for future MPSoC development. To serve this purpose, the NoCs have to be able to efficiently exchange all types of traffic including the collective communications at a reasonable cost. The Quarc NoC is introduced as a NOC which is highly efficient in performing collective communication operations such as broadcast and multicast. This paper presents an introduction to the Quarc scheme and an analytical model to compute the average message latency in the architecture. To validate the model we compare the model latency prediction against the results obtained from discrete-event simulations.
Mahmoud Moadeli, Wim Vanderbauwhede, Alireza Shahrabi
ICPADS2
2007 An Analytical Performance Model for the Spidergon NoC
abstract
Networks on chip (NoC) emerged as a promising alternative to bus-based interconnect networks to handle the increasing communication requirements of the large systems on chip. Employing an appropriate topology for a NoC is of high importance mainly because it typically trade-offs between cross-cutting concerns such as performance and cost. The spidergon topology is a novel architecture which is proposed recently for NoC domain. The objective of the spidergon NoC has been addressing the need for a fixed and optimized topology to realize cost effective multi-processor SoC (MPSoC) development [7]. In this paper we analyze the traffic behavior in the spidergon scheme and present an analytical evaluation of the average message latency in the architecture. We prove the validity of the analysis by comparing the model against the results produced by a discrete- event simulator.
Mahmoud Moadeli, Alireza Shahrabi, Wim Vanderbauwhede, Mohamed Ould-Khaoua
AINA3
2007 Analytical modelling of communication in the rectangular mesh NoC
abstract
Networks on chip (NoC) emerged as a packets switched, structured communication medium for development of the future systems on chip (SoC). Due to its unique features in terms of scalability and ease of synthesis, the (rectangular) mesh topology is regarded as an appropriate candidate for on-chip network development. This paper presents an analytical model of the average message latency for rectangular mesh topology. The validity of the analysis is verified by comparing the model against the results produced by a discrete-event simulator.
Mahmoud Moadeli, Alireza Shahrabi, Wim Vanderbauwhede
ICPADS3
2007 Communication Modelling of the Spidergon NoC with Virtual Channels
abstract
The spidergon scheme is a commercial NoC (Network On-Chip) proposed recently to address the demand for a fixed and optimized topology to realize cost effective multi-processor SoC (MPSoC) development. The increasing diversity of the applications quality of service requirements may, however, inhibit employing a particular architecture for a wide range of applications, unless the performance it delivers is improved. A traditional approach to enhance the performance of the interconnect networks has been employing the virtual channels. In this paper, we present an analytical model to evaluate the performance of the Spidergon NoC and to study the effect of employing virtual channels. Results obtained through simulation experiments show that the model exhibits a good degree of accuracy in predicting average message latency under various working conditions.
Mahmoud Moadeli, Alireza Shahrabi, Wim Vanderbauwhede, Mohamed Ould-Khaoua
ICPP3