EDBT 2026 Demo / reviewers in the wild / expert
Kees A. Vissers
dblp:31/537
· DBLP profile ↗
50ranked-venue papers
7as first author
5since 2021 · last 2024
0000-0002-6249-315XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 47 · 7 first-author · 4 since 2021Software engineering, systems software and programming languages · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | EcoFlow: Efficient Convolutional Dataflows on Low-Power Neural Network AcceleratorsabstractDilated and transposed convolutions are widely used in modern convolutional neural networks (CNNs). These kernels are used extensively during CNN training and inference of applications such as image segmentation and high-resolution image generation. We find that commonly-used low-power CNN inference accelerators arenotoptimized for both these convolutional kernels. Dilated and transposed convolutions introduce significant zero padding when mapped to the underlying spatial architecture, significantly degrading performance and energy efficiency. Existing approaches that address this issue require significant design changes to the otherwise simple, efficient, and well-adopted architectures used to compute direct convolutions. To address this challenge, we propose EcoFlow, a new set of dataflows and mapping algorithms for dilated and transposed convolutions. These algorithms are tailored to execute efficiently on existing low-cost, small-scale spatial architectures and requires minimal changes to existing accelerators. At its core, EcoFlow eliminates zero padding through careful dataflow orchestration and data mapping tailored to the spatial architecture. We evaluate EcoFlow on CNN training workloads and Generative Adversarial Network (GAN) workloads. Experiments in our new cycle-accurate simulator show that, using a common CNN inference accelerator, EcoFlow 1) reduces end-to-end CNN training time between 7-85%, and 2) improves end-to-end GAN training performance between 29-42%, compared to state-of-the-art CNN dataflows. Lois Orosa 0001, Skanda Koppula, Yaman Umuroglu, Konstantinos Kanellopoulos, Juan Gómez-Luna, Michaela Blott, Kees A. Vissers, Onur Mutlu |
IEEE Trans. Computers | 7 |
| 2022 | Elastic-DF: Scaling Performance of DNN Inference in FPGA Clouds through Automatic PartitioningabstractCustomized compute acceleration in the datacenter is key to the wider roll-out of applications based on deep neural network (DNN) inference. In this article, we investigate how to maximize the performance and scalability of field-programmable gate array (FPGA)-based pipeline dataflow DNN inference accelerators (DFAs) automatically on computing infrastructures consisting of multi-die, network-connected FPGAs. We present Elastic-DF, a novel resource partitioning tool and associated FPGA runtime infrastructure that integrates with the DNN compiler FINN. Elastic-DF allocates FPGA resources to DNN layers and layers to individual FPGA dies to maximize the total performance of the multi-FPGA system. In the resulting Elastic-DF mapping, the accelerator may be instantiated multiple times, and each instance may be segmented across multiple FPGAs transparently, whereby the segments communicate peer-to-peer through 100 Gbps Ethernet FPGA infrastructure, without host involvement. When applied to ResNet-50, Elastic-DF provides a 44% latency decrease on Alveo U280. For MobileNetV1 on Alveo U200 and U280, Elastic-DF enables a 78% throughput increase, eliminating the performance difference between these cards and the larger Alveo U250. Elastic-DF also increases operating frequency in all our experiments, on average by over 20%. Elastic-DF therefore increases performance portability between different sizes of FPGA and increases the critical throughput per cost metric of datacenter inference. Tobias Alonso, Lucian Petrica, Mario Ruiz, Jakoba Petri-Koenig, Yaman Umuroglu, Ioannis Stamelos, Elias Koromilas, Michaela Blott, Kees A. Vissers |
ACM Trans. Reconfigurable Technol. Syst. | 9 |
| 2022 | FPGA HLS Today: Successes, Challenges, and OpportunitiesabstractThe year 2011 marked an important transition for FPGA high-level synthesis (HLS), as it went from prototyping to deployment. A decade later, in this article, we assess the progress of the deployment of HLS technology and highlight the successes in several application domains, including deep learning, video transcoding, graph processing, and genome sequencing. We also discuss the challenges faced by today’s HLS technology and the opportunities for further research and development, especially in the areas of achieving high clock frequency, coping with complex pragmas and system integration, legacy code transformation, building on open source HLS infrastructures, supporting domain-specific languages, and standardization. It is our hope that this article will inspire more research on FPGA HLS and bring it to a new height. Jason Cong, Jason Lau, Gai Liu, Stephen Neuendorffer, Peichen Pan, Kees A. Vissers, Zhiru Zhang |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2021 | ASLR: An Adaptive Scheduler for Learning RateabstractTraining a neural network is a complicated and time-consuming task that involves adjusting and testing different combinations of hyperparameters. One of the essential hyperparameters is the learning rate, which balances the magnitude of changes at each training step. We introduce an Adaptive Scheduler for Learning Rate (ASLR) that significantly lowers the tuning effort since it only has a single hyperparameter. ASLR produces competitive results compared to the state-of-the-art for both hand-optimized learning rate schedulers and line search methods while requiring significantly less tuning effort. Our algorithm's computational cost is trivial and can be used to train various network topologies included quantized networks. Alireza Khodamoradi, Kristof Denolf, Kees A. Vissers, Ryan Kastner |
IJCNN | 3 |
| 2021 | Benchmarking vision kernels and neural network inference accelerators on embedded platforms
Murad Qasaimeh, Kristof Denolf, Alireza Khodamoradi, Michaela Blott, Jack Lo, Lisa Halder, Kees A. Vissers, Joseph Zambreno, Phillip H. Jones |
J. Syst. Archit. | 7 |
| 2020 | FAT: Training Neural Networks for Reliable Inference Under Hardware FaultsabstractDeep neural networks (DNNs) are state-of-the-art algorithms for multiple applications, spanning from image classification to speech recognition. While providing excellent accuracy, they often have enormous compute and memory requirements. As a result of this, quantized neural networks (QNNs) are increasingly being adopted and deployed especially on embedded devices, thanks to their high accuracy, but also since they have significantly lower compute and memory requirements compared to their floating point equivalents. QNN deployment is also being evaluated for safety-critical applications, such as automotive, avionics, medical or industrial. These systems require functional safety, guaranteeing failure-free behaviour even in the presence of hardware faults. In general fault tolerance can be achieved by adding redundancy to the system, which further exacerbates the overall computational demands and makes it difficult to meet the power and performance requirements. In order to decrease the hardware cost for achieving functional safety, it is vital to explore domain-specific solutions which can exploit the inherent features of DNNs. In this work we present a novel methodology called fault-aware training (FAT), which includes error modeling during neural network (NN) training, to make QNNs resilient to specific fault models on the device. Our experiments show that by injecting faults in the convolutional layers during training, highly accurate convolutional neural networks (CNNs) can be trained which exhibits much better error tolerance compared to the original. Furthermore, we show that redundant systems which are built from QNNs trained with FAT achieve higher worse-case accuracy at lower hardware cost. This has been validated for numerous classification tasks including CIFAR10, GTSRB, SVHN and ImageNet. Ussama Zahid, Giulio Gambardella, Nicholas J. Fraser, Michaela Blott, Kees A. Vissers |
ITC | 5 |
| 2019 | Analyzing the Energy-Efficiency of Vision Kernels on Embedded CPU, GPU and FPGA PlatformsabstractThis paper presents a benchmark of the energy efficiency of a wide range of vision kernels on three commonly used hardware accelerators for embedded vision applications: ARM57 CPU, Jetson TX2 GPU and ZCU102 FPGA, using their vendor optimized vision libraries: OpenCV, VisionWorks and xfOpenCV. Our results show that the GPU achieves an energy/frame reduction ratio of 1.1-3.2× compared to CPU and FPGA for simple kernels. While for more complicated kernels, the FPGA outperforms the others with energy/frame reduction ratios of 1.2-22.3×. It is also observed that the FPGA performs increasingly better as a vision kernel's complexity grows. Murad Qasaimeh, Joseph Zambreno, Phillip H. Jones, Kristof Denolf, Jack Lo, Kees A. Vissers |
FCCM | 6 |
| 2019 | Versal: The Xilinx Adaptive Compute Acceleration Platform (ACAP)abstractIn this presentation I will present the new Adaptive Compute Acceleration Platform. I will show the overall system architecture of the family of devices including the Arm cores (scalar engines), the programmable logic (Adaptable Engines) and the new vector processor cores (AI engines). I will focus on the new AI engines in more detail and show the architecture, the integration in the total device, the programming environment and some applications, including Machine Learning and 5G wireless applications. Kees A. Vissers |
FPGA | 1 |
| 2019 | Synetgy: Algorithm-hardware Co-design for ConvNet Accelerators on Embedded FPGAsabstractUsing FPGAs to accelerate ConvNets has attracted significant attention in recent years. However, FPGA accelerator design has not leveraged the latest progress of ConvNets. As a result, the key application characteristics such as frames-per-second (FPS) are ignored in favor of simply counting GOPs, and results on accuracy, which is critical to application success, are often not even reported. In this work, we adopt an algorithm-hardware co-design approach to develop a ConvNet accelerator called Synetgy and a novel ConvNet model called DiracDeltaNet. Both the accelerator and ConvNet are tailored to FPGA requirements. DiracDeltaNet, as the name suggests, is a ConvNet with only $1\times 1$ convolutions while spatial convolutions are replaced by more efficient shift operations. DiracDeltaNet achieves competitive accuracy on ImageNet (89.0% top-5), but with 48× fewer parameters and 65× fewer OPs than VGG16. We further quantize DiracDeltaNet's weights to 1-bit and activations to 4-bits, with less than 1% accuracy loss. These quantizations exploit well the nature of FPGA hardware. In short, DiracDeltaNet's small model size, low computational OP count, ultra-low precision and simplified operators allow us to co-design a highly customized computing unit for an FPGA. We implement the computing units for DiracDeltaNet on an Ultra96 SoC system through high-level synthesis. Our accelerator's final top-5 accuracy of 88.2% on ImageNet, is higher than all the previously reported embedded FPGA accelerators. In addition, the accelerator reaches an inference speed of 96.5 FPS on the ImageNet classification task, surpassing prior works with similar accuracy by at least 16.9×. Qijing Huang 0001, Bichen Wu, Tianjun Zhang, Liang Ma 0003, Giulio Gambardella, Michaela Blott, Luciano Lavagno, Kees A. Vissers, John Wawrzynek, Kurt Keutzer |
FPGA | 9 |
| 2018 | Novel Neural Network Applications on New Python Enabled PlatformsabstractSummary form only given, as follows. The complete presentation was not made available for publication as part of the conference proceedings. Reconfigurable technology is very well suited for novel implementations of neural networks. In this presentation we will show the range of implementations for neural networks on reconfigurable technology. This includes a direct dataflow implementation of networks, and implementations of an 'soft' processor using an array of DSP blocks. We will show the trade-offs in implementation cost and power over a range of implementations with varying bit-precisions. These implementations are leveraging novel python-based abstractions, that support a Jupyter Notebook interface. We will illustrate this with complete implementations on embedded platforms including the Pynq platform and on cloud-based platforms including the AWS F1 platform. Finally, we will project the new possibilities of the new 7nm based Xilinx platforms that contain specialized hardware that is very efficient for these neural network applications. Kees A. Vissers |
FPT | 1 |
| 2018 | FINN-R: An End-to-End Deep-Learning Framework for Fast Exploration of Quantized Neural NetworksabstractConvolutional Neural Networks have rapidly become the most successful machine-learning algorithm, enabling ubiquitous machine vision and intelligent decisions on even embedded computing systems. While the underlying arithmetic is structurally simple, compute and memory requirements are challenging. One of the promising opportunities is leveraging reduced-precision representations for inputs, activations, and model parameters. The resulting scalability in performance, power efficiency, and storage footprint provides interesting design compromises in exchange for a small reduction in accuracy. FPGAs are ideal for exploiting low-precision inference engines leveraging custom precisions to achieve the required numerical accuracy for a given application. In this article, we describe the second generation of the FINN framework, an end-to-end tool that enables design-space exploration and automates the creation of fully customized inference engines on FPGAs. Given a neural network description, the tool optimizes for given platforms, design targets, and a specific precision. We introduce formalizations of resource cost functions and performance predictions and elaborate on the optimization algorithms. Finally, we evaluate a selection of reduced precision neural networks ranging from CIFAR-10 classifiers to YOLO-based object detection on a range of platforms including PYNQ and AWS F1, demonstrating new unprecedented measured throughput at 50 TOp/s on AWS F1 and 5 TOp/s on embedded devices. Michaela Blott, Thomas B. Preußer, Nicholas J. Fraser, Giulio Gambardella, Kenneth O'Brien, Yaman Umuroglu, Miriam Leeser, Kees A. Vissers |
ACM Trans. Reconfigurable Technol. Syst. | 8 |
| 2017 | FINN: A Framework for Fast, Scalable Binarized Neural Network Inference
Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip H. W. Leong, Magnus Jahre, Kees A. Vissers |
FPGA | 7 |
| 2015 | Scalable 10Gbps TCP/IP Stack Architecture for Reconfigurable HardwareabstractTCP/IP is the predominant communication protocol in modern networks but also one of the most demanding. Consequently, TCP/IP offload is becoming increasingly popular with standard network interface cards. TCP/IP Offload Engines have also emerged for FPGAs, and are being offered by vendors such as Intilop, Fraunhofer HHI, PLDA and Dini Group. With the target application being high-frequency trading, these implementations focus on low latency and support a limited session count. However, many more applications beyond high-frequency trading can potentially be accelerated inside an FPGA once TCP with high session count is available inside the fabric. This way, a network-attached FPGA on ingress and egress to a CPU can accelerate functions such as encryption, compression, memcached and many others in addition to running the complete network stack. This paper introduces a novel architecture for a 10Gbps line-rate TCP/IP stack for FPGAs that can scale with the number of sessions and thereby addresses these new applications. We prototyped the design on a VC709 development board, demonstrating compatibility with existing network infrastructure, operating at full 10Gbps throughput full-duplex while supporting 10,000 sessions. Finally, the design has been described primarily using high-level synthesis, which accelerates development time and improves maintainability. David Sidler, Gustavo Alonso, Michaela Blott, Kimon Karras, Kees A. Vissers, Raymond Carley |
FCCM | 5 |
| 2015 | Scaling Out to a Single-Node 80Gbps Memcached Server with 40Terabytes of Memory
Michaela Blott, Kimon Karras, Kees A. Vissers |
HotStorage | 4 |
| 2015 | A Hash Table for Line-Rate Data ProcessingabstractFPGA-based data processing is becoming increasingly relevant in data centers, as the transformation of existing applications into dataflow architectures can bring significant throughput and power benefits. Furthermore, a tighter integration of computing and network is appealing, as it overcomes traditional bottlenecks between CPUs and network interfaces, and dramatically reduces latency. In this article, we present the design of a novel hash table, a fundamental building block used in many applications, to enable data processing on FPGAs close to the network. We present a fully pipelined design capable of sustaining consistent 10Gbps line-rate processing by deploying a concurrent mechanism to handle hash collisions. We address additional design challenges such as support for a broad range of key sizes without stalling the pipeline through careful matching of lookup time with packet reception time. Finally, the design is based on a scalable architecture that can be easily parameterized to work with different memory types operating at different access speeds and latencies. We have tested the proposed hash table in an FPGA-based memcached appliance implementing a main-memory key-value store in hardware. The hash table is used to index 2 million entries in 24GB of external DDR3 DRAM while sustaining 13 million requests per second, the maximum packet rate that can be achieved with UDP packets on a 10Gbps link for this application. Zsolt István, Gustavo Alonso, Michaela Blott, Kees A. Vissers |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2014 | OmpSs@Zynq all-programmable SoC ecosystemabstractOmpSs is an OpenMP-like directive-based programming model that includes heterogeneous execution (MIC, GPU, SMP, etc.) and runtime task dependencies management. Indeed, OmpSs has largely influenced the recently appeared OpenMP 4.0 specification. Zynq All-Programmable SoC combines the features of a SMP and a FPGA and benefits DLP, ILP and TLP parallelisms in order to efficiently exploit the new technology improvements and chip resource capacities. In this paper, we focus on programmability and heterogeneous execution support, presenting a successful combination of the OmpSs programming model and the Zynq All-Programmable SoC platforms. Antonio Filgueras, Eduard Gil, Daniel Jiménez-González, Carlos Álvarez 0001, Xavier Martorell, Jan Langer, Juanjo Noguera, Kees A. Vissers |
FPGA | 8 |
| 2013 | Reconfigurable computing in the era of post-silicon scaling [panel discussion]abstractSummary form only given, as follows. Although transistor densities continue to scale exponentially, the failure of Dennard Scaling prevents us from maximally utilizing die area in future power-constrained multicore processors-a phenomenon referred to as "Dark Silicon". Alternative energy-efficient architectures based on FPGAs, GPGPUs, ASICs, MPPAs, etc. are likely to continue on an exponential scaling trajectory while outperforming conventional architectures by an order-of-magnitude or more. With the impending threat of dark silicon, there is a critical window of opportunity for reconfigurable computing to become a mainstream ingredient and driver of future, scalable computer architectures. Before this can happen, major challenges and opportunities must be addressed: (1) how to gracefully integrate reconfigurable computing into existing software and hardware ecosystems, (2) how to build tools, languages, and compilers for agile application development and debugging, (3) how to identify and exploit emerging applications in datacenters and in energy-constrained form factors, (4) how to train and educate students and practitioners to use these systems in sustainable ways, and (5) how to define new and stable boundaries between software and hardware that make it easier to exploit reconfigurable computing. This panel brings together pioneers and experts in computer architecture and reconfigurable computing to discuss opportunities and challenges in the wake of dark silicon. Eric S. Chung, Doug Burger, Mike Butts, Jan Gray, Charles P. Thacker, Kees A. Vissers, John Wawrzynek |
FCCM | 6 |
| 2013 | A flexible hash table design for 10GBPS key-value stores on FPGASabstractCommon web infrastructure relies on distributed main memory key-value stores to reduce access load on databases, thereby improving both performance and scalability of web sites. As standard cloud servers provide sub-linear scalability and reduced power efficiency to these kinds of scale-out workloads, we have investigated a novel dataflow architecture for key-value stores with the aid of FPGAs which can deliver consistent 10Gbps throughput. In this paper, we present the design of a novel hash table which forms the centre piece of this dataflow architecture. The fully pipelined design can sustain consistent 10Gbps line-rate performance by deploying a concurrent mechanism to handle hash collisions. We address problems such as support for a broad range of key sizes without stalling the pipeline through careful matching of lookup time with packet reception time. Finally, the design is based on a scalable architecture that can be easily parametrized to work with different memory types operating at different access speeds and latencies. We deployed this hash table in a memcached prototype to index 2 million entries in 24GBytes of external DDR3 DRAM while sustaining 13 million requests per second for UDP binary encoded memcached packets which is the maximum packet rate that can be achieved with memcached on a 10Gbps link. Zsolt István, Gustavo Alonso, Michaela Blott, Kees A. Vissers |
FPL | 4 |
| 2013 | Dataflow architectures for 10Gbps line-rate key-value-stores
Michaela Blott, Kees A. Vissers |
Hot Chips Symposium | 2 |
| 2013 | Software-programmable digital pre-distortion on the Zynq SoCabstractDigital pre-distortion (DPD) is an advanced digital signal-processing technique that mitigates the effects of power amplifier (PA) nonlinearity in wireless transmitters. DPD plays a key role in providing efficient radio digital front-end (DFE) solutions for 3G/4G basestations and beyond. Modern FPGAs are a promising target platform for the implementation of flexible wireless DFE solutions, including DPD. In this paper, we present a software programmable design flow that facilitates the implementation and integration of efficient DPD solutions on Xilinx's Zynq All Programmable SoC, combining industry-standard embedded processors and programmable logic fabric into one chip. In addition to software programmability, another key contribution of this design flow is the flexible partitioning of functionality among hardware and software components, depending on the complexity of the DPD parameter estimation algorithm in use. We have applied processor-specific optimizations to the software implementation and used the Vivado High-Level Synthesis (HLS) tool as the design tool for the programmable logic. Using this design flow, we integrated a complete DPD feedback path on the Zynq SoC that achieves up to 7x speed-up from hardware acceleration. Baris Özgül, Jan Langer, Juanjo Noguera, Kees A. Vissers |
VLSI-SoC | 4 |
| 2011 | Building real-time HDTV applications in FPGAs using processors, AXI interfaces and high level synthesis toolsabstractModern FPGAs enable complete system designs that include processors, interconnect systems, memory subsystems and a number of application functions that are implemented using High-Level Synthesis tools. Kees A. Vissers, Stephen Neuendorffer, Juanjo Noguera |
DATE | 1 |
| 2011 | High level synthesis for FPGAs applied to a sphere decoder channel preprocessor (abstract only)abstractDuring the 1990s, High Level Synthesis (HLS) slowly started emerging to allow designers to cope with the ever-increasing complexity of digital signal processing systems such as wireless receivers. Only recently, a new generation of high-quality commercial HLS tools capable of generating decent RTL architectures from algorithmic-style code has become available. Although these tools are now starting to be adopted for actual designs, detailed experiences and results for these tools in significant designs is still lacking. In this work we have analyzed the design process of a significant portion of a wireless communication application using AutoESL's AutoPilot HLS tool. We target a Xilinx Virtex-5 FPGA device and a clock frequency of 225 MHz. We have compared our HLS implementation to a reference implementation obtained using manual RTL design methods. We found that our HLS implementation is competitive to this reference implementation in terms of throughput and resource cost. The time needed to obtain a first HLS implementation matching throughput and resource cost aspects of the reference implementation is similar to the design time of the reference implementation. After obtaining a first HLS implementation, we were able to explore and implement different application architectures with HLS in only a couple of hours, thereby gaining significant savings in design time over manual RTL design, where architectural exploration may take weeks. Sven van Haastregt, Stephen Neuendorffer, Kees A. Vissers, Bart Kienhuis |
FPGA | 3 |
| 2011 | High-Level Synthesis for FPGAs: From Prototyping to DeploymentabstractEscalating system-on-chip design complexity is pushing the design community to raise the level of abstraction beyond register transfer level. Despite the unsuccessful adoptions of early generations of commercial high-level synthesis (HLS) systems, we believe that the tipping point for transitioning to HLS msystem-on-chip design complexityethodology is happening now, especially for field-programmable gate array (FPGA) designs. The latest generation of HLS tools has made significant progress in providing wide language coverage and robust compilation technology, platform-based modeling, advancement in core HLS algorithms, and a domain-specific approach. In this paper, we use AutoESL's AutoPilot HLS tool coupled with domain-specific system-level implementation platforms developed by Xilinx as an example to demonstrate the effectiveness of state-of-art C-to-FPGA synthesis solutions targeting multiple application domains. Complex industrial designs targeting Xilinx FPGAs are also presented as case studies, including comparison of HLS solutions versus optimized manual designs. In particular, the experiment on a sphere decoder shows that the HLS solution can achieve an 11-31% reduction in FPGA resource usage with improved design productivity compared to hand-coded design. Jason Cong, Bin Liu 0006, Stephen Neuendorffer, Juanjo Noguera, Kees A. Vissers, Zhiru Zhang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2010 | Programming high performance signal processing systems in high level languagesabstractNo abstract available. Kees A. Vissers, Devada Varma, Vinod Kathail, Jeff Bier, Don MacMillen, Joseph R. Cavallaro |
FPGA | 1 |
| 2009 | The wild west: conquest of complex hardware-dependent software designabstractEmbedded SW design can be compared to the lawless wild west. With no clear methodology and no standard multi core platform modeling environment every company has to improvise its own solution. The problems facing embedded software users are becoming more complex since: Hiroyuki Yagi, Wolfgang Rosenstiel, Jakob Engblom, Jason Andrews, Kees A. Vissers, Marc Serughetti |
DAC | 5 |
| 2009 | Using C-to-gates to program streaming image processing kernels efficiently on FPGAsabstractEffectively exploiting the variety of computational and storage resources available in common FPGA architectures for complex applications, such as the real-time implementation of vision algorithms, is often difficult in standard HDL design methodologies. Higher-level design tools can enable a design to more quickly explore a range of different architectures. In this paper we apply algorithmic C-to-FPGA synthesis technology in a structured design approach and demonstrate its added value on two relevant vision processing kernels: optical flow and debayering. The impact of the proposed approach on the design time, the FPGA resource consumption and the throughput is measured. Kristof Denolf, Stephen Neuendorffer, Kees A. Vissers |
FPL | 3 |
| 2006 | The next generation 65-nm FPGA
Steve Douglass, Kees A. Vissers, Peter Alfke |
Hot Chips Symposium | 2 |
| 2005 | Optimized Generation of Data-Path from C Codes for FPGAsabstractFPGAs, as computing devices, offer significant speedup over microprocessors. Furthermore, their configurability offers an advantage over traditional ASICs. However, they do not yet enjoy high-level language programmability, as microprocessors do. This has become the main obstacle for their wider acceptance by application designers. ROCCC is a compiler designed to generate circuits from C source code to execute on FPGAs, more specifically on CSoCs. It generates RTL level HDLs from frequently executing kernels in an application. In this paper, we describe the ROCCC's system overview and focus on its data path generation. We compare the performance of ROCCC-generated VHDL code with that of Xilinx IPs. The synthesis result shows that the ROCCC-generated circuit takes around 2/spl times//spl sim/3/spl times/ the area and runs at a comparable clock rate. Zhi Guo, Betul Buyukkurt, Walid A. Najjar, Kees A. Vissers |
DATE | 4 |
| 2005 | Firm-core Virtual FPGA for Just-in-Time FPGA Compilation (abstract only)abstractJust-in-time (JIT) compilation has been used in many applications to enable standard software binaries to execute on different underlying processor architectures, yielding software portability benefits. We previously introduced the concept of a standard hardware binary to achieve similar portability benefits for hardware, using a JIT compiler to compile the hardware binary to an FPGA. Our JIT compiler includes lean versions of technology mapping, placement, and routing algorithms that implement the standard hardware binary on a simple custom FPGA fabric designed specifically for JIT compilation. While directly implementing a custom FPGA fabric on silicon may be feasible for some applications, we investigated the option of implementing the simple FPGA fabric as a circuit mapped to a physical FPGA - a virtual FPGA. We described our simple fabric in structural VHDL, synthesized the fabric onto a Xilinx Spartan-IIE FPGA, and mapped 18 benchmark circuits onto the resulting virtual FPGA. Our results show a 6X decrease in performance and a 100X increase in hardware resource usage for the virtual FPGA approach compared to mapping the circuits directly to the physical FPGA. For applications in which hardware portability is essential, a designer could leverage the large capacity of current commercially available FPGAs to implement a virtual FPGA with tens of thousands of configurable gates, providing about the same amount of configurable logic as FPGAs produced in the mid 1990s. Nevertheless, the large overheads clearly indicate the need to develop a virtual FPGA approach tuned to physical fabrics in order to reduce the overhead. Roman L. Lysecky, Kris Miller, Frank Vahid, Kees A. Vissers |
FPGA | 4 |
| 2005 | Memory Efficient Design of an MPEG-4 Video Encoder for FPGAsabstractThe improving resolutions of new video appliances continuously increase the throughput requirements of video codecs and complicate the challenges encountered during their cost-efficient design. We propose an FPGA implementation of a high-performance MPEG-4 video encoder. The fully dedicated video pipeline is realized using a systematic design approach and exploits the inherent functional parallelism of the compression algorithm. The effect of memory and algorithmic optimizations applied at the high-level are measured on the RTL description. The resulting MPEG-4 video encoder efficiently uses the FPGA blockRAMs, uses burst oriented accesses to external memory and supports real-time processing of 30 4CIF frames per second. Kristof Denolf, Adrian Chirila-Rus, Robert D. Turney, Paul R. Schumacher, Kees A. Vissers |
FPL | 5 |
| 2005 | A scalable, multi-stream MPEG-4 video decoder for conferencing and surveillance applicationsabstractIncreasing resolutions push the throughput requirements of video codecs and complicate the challenges encountered during their cost-efficient implementations. We propose an FPGA implementation of a high-performance MPEG-4 simple profile video decoder, capable of parsing multiple bitstreams from different encoder sources. Its video pipeline architecture exploits the inherent functional parallelism and enables multi-stream support at a limited FPGA resource cost compared to a single stream version. The design is scalable with a number of added compile-time parameters - including maximum frame size and number of input bitstreams - which can be set by the user to suit his application. Paul R. Schumacher, Kristof Denolf, Adrian Chirila-Rus, Robert D. Turney, Nick Fedele, Kees A. Vissers, Jan Bormans |
ICIP (2) | 6 |
| 2004 | Programming models and architectures for FPGA platformsabstractModern Platform FPGAs contain a combination of processors, embedded memory, programmable interconnect, dedicated DSP elements, and conventional lookup tables. On top of that they have multiple clock domains, very high speed Serial I/Os and a large number of pins.This talk will focus on the programming of these systems. Conventional general purpose processors have been successfully programmed with sequential, control dominated programming languages like C and C++. This programming style has been succesfull for high-end DSPs, only after the architecture of the DSPs were tuned. This tuning required the interaction between compiler writers and processor architects.Todays conventional FPGAs are programmed with VHDL or Verilog. The semantics of these languages enable significant current systems at the expense of relatively low-level, explicit timing oriented programming.The existing notion of Dataflow models should be very good for expressing concurrency in high-end DSP systems. One of the first programming environments for FPGAs that exploit this is the Matlab/Simulink environment.This talk will show how the various dataflow models can be exploited to build more powerfull programming environments for todays FPGAs. In particular the combination of programming models and the systematic design of architecture elements is extremely powerfull. The talk will finish by showing the interesting research directions that can have a significant impact on the design of future FPGA platforms. The various approaches will be illustrated with facts and figures from actual designs, including JPEG2000 and MPEG4 encoders and decoders. Kees A. Vissers |
CASES | 1 |
| 2004 | A quantitative analysis of the speedup factors of FPGAs over processorsabstractThe speedup over a microprocessor that can be achieved by implementing some programs on an FPGA has been extensively reported. This paper presents an analysis, both quantitative and qualitative, at the architecture level of the components of this speedup. Obviously, the spatial parallelism that can be exploited on the FPGA is a big component. By itself, however, it does not account for the whole speedup. In this paper we experimentally analyze the remaining components of the speedup. We compare the performance of image processing application programs executing in hardware on a Xilinx Virtex E2000 FPGA to that on three general-purpose processor platforms: MIPS, Pentium III and VLIW. The question we set out to answer is what is the inherent advantage of a hardware implementation over a von Neumann platform. On the one hand, the clock frequency of general-purpose processors is about 20 times that of typical FPGA implementations. On the other hand, the iteration level parallelism on the FPGA is one to two orders of magnitude that on the CPUs. In addition to these two factors, we identify the efficiency advantage of FPGAs as an important factor and show that it ranges from 6 to 47 on our test benchmarks. We also identify some of the components of this factor: the streaming of data from memory, the overlap of control and data flow and the elimination of some instruction on the FPGA. The results provide a deeper understanding of the tradeoff between system complexity and performance when designing Configurable SoC as well as designing software for CSoC. They also help understand the one to two orders of magnitude in speedup of FPGAs over CPU after accounting for clock frequencies. Zhi Guo, Walid A. Najjar, Frank Vahid, Kees A. Vissers |
FPGA | 4 |
| 2004 | Pel reconstruction on FPGA-augmented TriMediaabstractThis paper presents a TriMedia processor extended with three reconfigurable designs for entropy decoding (ED), inverse quantization (IQ), and two-dimensional (2-D) inverse discrete cosine transform (IDCT), and assesses the performance gain that is provided by such extensions when performing MPEG2-compliant pel reconstruction. We first describe an extension of the TriMedia architecture, which consists of a multiple-context field programmable gate array (FPGA)-based reconfigurable functional unit (RFU), a configuration unit managing the reconfiguration of the RFU, and their associated instructions. Then, we address the computation of the ED, IQ, and 2-D IDCT tasks, and propose to provide reconfigurable hardware support for a variable-length decoder that can decode two symbols per call (VLD-2), an inverse quantizer that can dequantize four coefficients per call (IQ-4), and an 1-D IDCT (1-D IDCT). The most important aspects concerning the implementation of the FPGA-mapped VLD-2, IQ-4, and 1-D IDCT units, as well as the organization of the software routines calling these FPGA-mapped computing units are outlined. Experimental results indicate that by configuring each of the VLD-2, IQ-4, and 1-D IDCT units on a different FPGA context, and by activating the contexts as needed, the FPGA-augmented TriMedia can perform MPEG2-compliant pel reconstruction with an average speed-up of 1.4/spl times/ over the standard TriMedia. Mihai Sima, Sorin Cotofana, Stamatis Vassiliadis, Jos T. J. van Eijndhoven, Kees A. Vissers |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2003 | Panel Title: Reconfigurable Computing - Different Perspectives
Wolfgang Rosenstiel, Rudy Lauwereins, Ivo Bolsens, Chris Rowen, Yankin Tanurhan, Kees A. Vissers |
DATE | 6 |
| 2003 | Parallel Processing Architectures for Reconfigurable Systems
Kees A. Vissers |
DATE | 1 |
| 2002 | MPEG-Compliant Entropy Decoding on FPGA-Augmented TriMedia/CPU64abstractThe paper presents a Design Space Exploration (DSE) experiment which has been carried out in order to determine the optimum FPGA-based Variable-Length Decoder (VLD) computing resource and its associated instructions, with respect to an entropy decoding task which is to be executed on the FPGA-augmented TriMedia/CPU64 processor We first outline the extension of the TriMedia/CPU64 architecture, which consists of an FPGA-based Reconfigurable Functional Unit (RFU) and the associated generic instructions. Then we address entropy decoding and propose a strategy to partially break the data dependency related to variable-length decoding. Three VLDs (VLD-1, VLD-2, VLD-3) instructions which can return 1, 2, or 3 symbols, respectively, are subsequently analyzed. After completing the DSE, we determined that VLD-2 instruction leads to the most efficient entropy decoding in terms of instruction cycles and FPGA area. The FPGA-based implementation of the computing resource associated to VLD-2 instruction is subsequently presented. When mapped on an ACEX EP1K100 FPGA from Altera, VLD-2 exhibits a latency of 8 TriMedia cycles, and uses all the Electronic Array Blocks and 51% of the logic cells of the device. The simulation results indicate that the VLD-2-based entropy decoder is 43% faster than its pure software counterpart. Mihai Sima, Sorin Cotofana, Stamatis Vassiliadis, Jos T. J. van Eijndhoven, Kees A. Vissers |
FCCM | 5 |
| 2002 | Field-Programmable Custom Computing Machines - A Taxonomy -
Mihai Sima, Stamatis Vassiliadis, Sorin Cotofana, Jos T. J. van Eijndhoven, Kees A. Vissers |
FPL | 5 |
| 2001 | An 8x8 IDCT Implementation on an FPGA-Augmented TriMedia
Mihai Sima, Sorin Cotofana, Jos T. J. van Eijndhoven, Stamatis Vassiliadis, Kees A. Vissers |
FCCM | 5 |
| 2001 | MPEG Macroblock Parsing and Pel Reconstruction On An FPGA-Augmented TriMedia ProcessorabstractThis paper describes an experiment which aims to reveal the potential impact on performance yielded by augmenting a TriMedia-CPU64 processor with a multiple-context FPGA core. We first propose an extension of the TriMedia CPU64 architecture, which consists of a reconfigurable functional unit and its associated instructions. Then, we address the decoding of variable-length codes on such extended TriMedia and describe the architecture and FPGA-implementation of a variable-length decoder (VLD) computing facility. When mapped on an ACEX EP1K100 FPGA, the proposed VLD exhibits a latency of 7 cycles. Preliminary results indicate that by configuring each of the VLD and 1-D IDCT (which is described elsewhere) facilities on a different FPGA context, and by activating the contexts as needed, the augmented TriMedia can perform macroblock parsing followed up by pel reconstruction with an improvement of 20 - 25% over the standard TriMedia. Mihai Sima, Sorin Cotofana, Stamatis Vassiliadis, Jos T. J. van Eijndhoven, Kees A. Vissers |
ICCD | 5 |
| 2000 | YAPI: application modeling for signal processing systemsabstractWe present a programming interface called YAPI to model signal processing applications as process networks. The purpose of YAPI is to enable the reuse of signal processing applications and the mapping of signal processing applications onto heterogeneous systems that contain hardware and software components. To this end, YAPI separates the concerns of the application programmer, who determines the functionality of the system, and the system designer, who determines the implementation of the functionality. The proposed model of computation extends the existing model of Kahn process networks with channel selection to support non-deterministic events. We provide an efficient implementation of YAPI in the form of a C++ run-time library to execute the applications on a workstation. Subsequently, the applications are used by the system designer as input for mapping and performance analysis in the design of complex signal processing systems. We evaluate this methodology on the design of a digital video broadcast system-on-chip. Erwin A. de Kock, W. J. M. Smits, Pieter van der Wolf, Jean-Yves Brunel, Wido Kruijtzer, Paul Lieverse, Kees A. Vissers, Gerben Essink |
DAC | 7 |
| 2000 | The Future of Flexible HW Platform Architectures Panel Discussion
Rolf Ernst, Grant Martin, Oz Levia, Pierre G. Paulin, Stamatis Vassiliadis, Kees A. Vissers |
DATE | 6 |
| 1999 | System level design and debug of high-performance embedded media systems (tutorial)
Rolf Ernst, Kees A. Vissers, Pieter van der Wolf, Gert-Jan van Rootselaar |
ICCAD | 2 |
| 1999 | TriMedia CPU64 ArchitectureabstractWe present a new VLIW core as a successor to the TriMedia TM1000. The processor is targeted for embedded use in media-processing devices like DTVs and set-top boxes. Intended as a core, its design must be supplemented with on-chip co-processors to obtain a cost-effective system. Good performance is obtained through a uniform 64-bit 5 issue-slot VLIW design, supporting subword parallelism with an extensive instruction set optimized with respect to media-processing. Multi-slot 'super-ops' allow powerful multi-argument and multi-result operations. As an example, the IDCT algorithm shows a very low instruction count in comparison with other processors. To achieve good performance, critical sections in the application program source code need to be rewritten with vector data types and function calls for media operations. Benchmarking with several media applications was used to tune the instruction set and study cache behaviour. This resulted in a VLIW architecture with wide data paths and relatively simple CPU control. Jos T. J. van Eijndhoven, Kees A. Vissers, Evert-Jan D. Pol, P. Struik, R. H. J. Bloks, Pieter van der Wolf, Harald P. E. Vranken, Frans Sijstermans, M. J. A. Tromp, Andy D. Pimentel |
ICCD | 2 |
| 1999 | TriMedia CPU64 Application Domain and Benchmark SuiteabstractAt Philips Research Labs, we are investigating the 64-bit VLIW core (also called CPU64) for future TriMedia processors. This processor is targeted towards embedded multimedia applications. In order to be able to perform a quantitative design space exploration, a set of benchmark applications has been developed which is representative of the application domain. This article describes the way the benchmark set was developed as well as providing a more detailed description of one of the benchmarks. As a result, it is shown that this new CPU has a high performance for the application domain while the programming interface is very convenient. Furthermore, such a programmable media processor may offer new options and challenges to algorithm designers. A. K. Riemens, Kees A. Vissers, R. J. Schutten, Gerben J. Hekstra, G. D. La Hei, Frans Sijstermans |
ICCD | 2 |
| 1998 | Design space exploration for future TriMedia CPUsabstractIt is widely recognized that fine-grain parallelism can greatly enhance a processor's performance for signal processing applications. For this reason, future generation TriMedias will combine VLIW and subword parallelism in a single CPU. We present a snapshot of the new CPUs design process: the outlines are clear but fine tuning is still ongoing. We present the design flow and 'workbench' that the designers use for further tuning. Frans Sijstermans, Evert-Jan D. Pol, Bram Riemens, Kees A. Vissers, Selliah Rathnam, Gert Slavenburg |
ICASSP | 4 |
| 1997 | An Approach for Quantitative Analysis of Application-Specific Dataflow ArchitecturesabstractIn this paper we present an approach for quantitative analysis of application-specific dataflow architectures. The approach allows the designer to rate design alternatives in a quantitative way and therefore supports him in the design process to find better performing architectures. The context of our work is video signal processing algorithms which are mapped onto weakly-programmable, coarse-grain dataflow architectures. The algorithms are represented as Kahn graphs with the functionality of the nodes being coarse-grain functions. We have implemented an architecture simulation environment that permits the definition of dataflow architectures as a composition of architecture elements, such as functional units, buffer elements and communication structures. The abstract, clock-cycle accurate simulator has been built using a multi-threading package and employs object oriented principles. This results in a configurable and efficient simulator. Algorithms can subsequently be executed on the architecture model producing quantitative information for selected performance metrics. Results are presented for the simulation of a realistic application on several dataflow architecture alternatives, showing that many different architectures can be simulated in modest time on a modern workstation. Bart Kienhuis, Ed F. Deprettere, Kees A. Vissers, Pieter van der Wolf |
ASAP | 3 |
| 1995 | Architecture and programming of two generations video signal processors
Kees A. Vissers, Gerben Essink, Piet J. van Gerwen, P. J. M. Janssen, O. Popp, E. Riddersma, W. J. M. Smits, Harry J. M. Veendrick |
Microprocess. Microprogramming | 1 |
| 1991 | Scheduling in Programmable Video Signal ProcessorsabstractThe authors discuss the problem of mapping algorithms for real-time processing of digital video signals onto a fixed configuration of identical programmable video signal processors. Due to the periodic nature of the algorithms and the small periods that are involved, successive executions of the algorithm have to be interleaved in time. The resulting scheduling problem is mathematically modeled and examined. The authors present a novel solution approach that is based on a divide-and-conquer strategy using phase assignment as the central part. This approach has been implemented and it gives good results for industrially significant video applications. Specifically, the proposed approach has been implemented in only 1300 lines of C and has been applied to a number of problem instances, whose signal flow graphs originate from industrially relevant algorithms, including contour enhancement and progressive scan, noise reduction, 4:3 to 16:9 screen format conversion, and a very elaborate progressive scan algorithm.> Gerben Essink, Emile H. L. Aarts, R. van Dongen, Piet J. van Gerwen, Jan H. M. Korst, Kees A. Vissers |
ICCAD | 6 |
| 1991 | Architecture and Programming of a VLIW Style Programmable Video Signal ProcessorabstractThe architecture and programming aspects of a programmable video signal processor are discussed.The processor is an integrated circuit that has a modular architecture with a number of programmable, pipelined processing elements.Networks of these processors can be programmed conveniently with the aid of dedicated programming tools.In this paper the emphasis is on the scheduling of video algorithms and the micro code generation for a network of video signal processors.Due to the periodic nature of the video algorithms and the small periods that are involved, successive executions of the video algorithm have to be interleaved in time.We present a novel solution approach to the scheduling problem using phase assignment as the central part.Results of this approach are presented for industrially significant video applications. Gerben Essink, Emile H. L. Aarts, R. van Dongen, Piet J. van Gerwen, Jan H. M. Korst, Kees A. Vissers |
MICRO | 6 |