EDBT 2026 Demo / reviewers in the wild / expert
Diederik Verkest
dblp:93/763
· DBLP profile ↗
66ranked-venue papers
3as first author
5since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 56 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 13 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6Theory of computation · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Tiny ci-SAR A/D Converter for Deep Neural Networks in Analog in-Memory ComputationabstractThis paper presents a tiny charge injection-Successive Approximation (ci-SAR) A/D converter (ADC) to be integrated at the periphery of analog Matrix Vector Multiplication (MVM) accelerators for Deep Neural Network (DNN) inference. Derived from the ci-SAR ADC, this converter exploits a single charge injecting cell to minimize area and energy consumption. The ADC exhibits a signal-to-noise and distortion ratio of 30.5 dB, at 5 bits of nominal resolution. The energy per conversion is 86 fJ, running at 34 MS/s, with a silicon area of $75 \mu m ^{2}$, in 22 nm technology node. From the results of our analytical framework, an SRAM-based Analog in-Memory Compute (AiMC) array, including the proposed ADC at 5 bits of resolution, can achieve an energy efficiency of 1650 TOPs/W. Michele Caselli, Debjyoti Bhattacharjee, Arindam Mallik, Peter Debacker, Diederik Verkest |
ISCAS | 5 |
| 2022 | Write-Verify Scheme for IGZO DRAM in Analog in-Memory ComputingabstractLarge weight variations cause significant degradation in Deep Neural Networks (DNNs) accuracy in machine learning (ML) context. Indium-Gallium-Zinc-Oxide (IGZO) DRAM compute cell is a promising option for Analog in-Memory Computing (AiMC) accelerators, but its applicability requires the compensation of large variations affecting the voltage threshold of the readout device. This paper proposes a write-verify scheme for IGZO-based AiMC accelerator, designed in 22-nm technology, based on a compensation loop operating on the analog weight value stored in the IGZO DRAM cell. After the compensation routine, the IGZO IONnormalized variation drops from $\pm \mathbf{2 7} \%$ to $\pm 3 \%$ in simulation. With sufficiently large weight reuse, the additional energy spent for the compensation of the entire arrays becomes negligible, recovering the 2000 TOPS/W baseline performance of an ideal IGZO array without the write-verify. Michele Caselli, Subhali Subhechha, Peter Debacker, Arindam Mallik, Diederik Verkest |
ISCAS | 5 |
| 2022 | Dynamic Quantization Range Control for Analog-in-Memory Neural Networks AccelerationabstractAnalog in Memory Computing (AiMC) based neural network acceleration is a promising solution to increase the energy efficiency of deep neural networks deployment. However, the quantization requirements of these analog systems are not compatible with state-of-the-art neural network quantization techniques. Indeed, while the quantization of the weights and activations is considered by modern deep neural network quantization techniques, AiMC accelerators also impose the quantization of each Matrix Vector Multiplication (MVM) result. In most demonstrated AiMC implementations, the quantization range of MVM results is considered a fixed parameter of the accelerator. This work demonstrates that dynamic control over this quantization range is possible but also desirable for analog neural networks acceleration. An AiMC compatible quantization flow coupled with a hardware aware quantization range driving technique is introduced to fully exploit these dynamic ranges. Using CIFAR-10 and ImageNet as benchmarks, the proposed solution results in networks that are both more accurate and more robust to the inherent vulnerability of analog circuits than fixed quantization range based approaches. Nathan Laubeuf, Jonas Doevenspeck, Ioannis A. Papistas, Michele Caselli, Stefan Cosemans, Peter Vrancx, Debjyoti Bhattacharjee, Arindam Mallik, Peter Debacker, Diederik Verkest, Francky Catthoor, Rudy Lauwereins |
ACM Trans. Design Autom. Electr. Syst. | 10 |
| 2021 | Noise tolerant ternary weight deep neural networks for analog in-memory inferenceabstractAnalog in memory computing (AiMC) is a promising hardware solution to efficiently perform inference with deep neural networks (DNNs). Similar to digital DNN accelerators, AiMC systems benefit from aggressively quantized DNNs. In addition, AiMC systems also suffer from noise on activations and weights. Training strategies to condition DNNs against weight noise can increase the efficiency of AiMC systems by enabling the use of more compact but more noisy weight memory devices. In this work, we utilize noise-aware training and introduce gradual noise training and network width scaling to increase the tolerance of DNNs against weight noise. Our results show that noise-aware training and gradual noise training drastically lowers the impact of weight noise without changing the network size. By utilizing network width scaling, the weight noise tolerance is increased even more with the penalty of more network parameters. Jonas Doevenspeck, Peter Vrancx, Nathan Laubeuf, Arindam Mallik, Peter Debacker, Diederik Verkest, Rudy Lauwereins, Wim Dehaene |
IJCNN | 6 |
| 2021 | Design-Technology Space Exploration for Energy Efficient AiMC-Based Inference AccelerationabstractExtremely energy-efficient convolutional neural network inference (CNN) is recently enabled by analog in-memory compute (AiMC). The integration of AiMC in a primarily digital inference system brings new challenges ranging from device specifications to defining novel system architectures. A novel framework to evaluate the impact of AiMC array at the system level is presented. The framework is used to model a SRAM-based 1024 × 512 prototype AiMC array, capable of energy efficiency of upto 675 TMACs/W. The proposed framework allows modelling different compute cells, array dimensions, operating voltage, and activation buffer energy models which can be used to determine the overall energy efficiency for various CNN workloads. Debjyoti Bhattacharjee, Nathan Laubeuf, Stefan Cosemans, Ioannis A. Papistas, Arindam Mallik, Peter Debacker, Myung Hee Na, Diederik Verkest |
ISCAS | 8 |
| 2017 | Statistical Timing Analysis Considering Device and Interconnect Variability for BEOL Requirements in the 5-nm Node and BeyondabstractIn an increasing interconnect resistance era and aggressive metal pitch scaling, the elevating RC delay could significantly shadow the improvements from advanced device architectures and become a severe design issue. This paper will holistically analyze the interplay between transistors and interconnect delay and the variability induced by back-end-ofline (BEOL) process for the 5-nm node. A global sensitivity analysis using Monte Carlo simulation is employed as a powerful tool for understanding the significance of different variation sources and propagating these process uncertainties to circuit performance and parametric yield. For the BEOL integration process, our results show that dielectric κ-value is the most sensitive parameter. Regarding the patterning options, the BEOL process using self-aligned quadruple pattering with positive tone process requires more than a 4× process margin and suffers from 50% parametric yield loss. The required guardband for lithoetch litho-etch becomes as critical as for the self-aligned double patterning process when the overlay control is 6× higher than the critical dimension control. For trench patterning using spacerdefined techniques, a negative tone process is required to achieve a large process window. From a design perspective, the wire length in SoC can be optimized using a disruptive architecture as a vertical FET, which could potentially reduce the average wire length by 11%. Trong Huynh Bao, Julien Ryckaert, Zsolt Tokei, Abdelkarim Mercha, Diederik Verkest, Aaron Thean, Piet Wambacq |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2016 | Physical Design Solutions to Tackle FEOL/BEOL Degradation in Gate-level Monolithic 3D ICsabstractIn this paper, we develop physical design tools and methodologies to tackle the inter-tier performance variations caused by low temperature manufacturing in 2-tier gate-level monolithic 3D ICs (M3D). First, we model the top tier front-end-of-line (FEOL) device mobility degradation and its impact on cell delay/power values. Next, we quantify the impact of tungsten interconnect and cost-driven metal layer saving in the back-end-of-line (BEOL) of the bottom tier. These device and interconnect degradation models are used in our new full-chip M3D physical design flow named Derated 2D. This flow overcomes the well-known drawback of the state-of-the-art Shrunk 2D that requires shrinking of layout objects and RC parasitics. Also, Derated 2D performs low-temperature process-aware tier partitioning to effectively keep timing-critical components in the bottom tier. Moreover, Derated 2D conducts timing-driven monolithic inter-tier via (MIV) planning to cope with the resistivity increase in tungsten BEOL. Lastly, Derated 2D offers an effective timing closure solution through a post-route optimization. Experiments based on a foundry-grade 7nm FinFET process design kit (PDK) show that Derated 2D achieves up to 36% performance improvement and 10% energy saving compared with Shrunk 2D. Using a post-route optimization, Derated 2D further improves timing under various FEOL/BEOL degradation settings at a minimum energy overhead. Bon Woong Ku, Peter Debacker, Dragomir Milojevic, Praveen Raghavan, Diederik Verkest, Aaron Thean, Sung Kyu Lim |
ISLPED | 5 |
| 2015 | Impact of interconnect multiple-patterning variability on SRAMs
Ioannis Karageorgos, Michele Stucchi, Praveen Raghavan, Julien Ryckaert, Zsolt Tokei, Diederik Verkest, Rogier Baert, Sushil Sakhare, Wim Dehaene |
DATE | 6 |
| 2013 | Design issues in heterogeneous 3D/2.5D integrationabstractEfficient processing of fine-pitched Through Silicon Vias, micro-bumps and back-side re-distribution layers enable face-to-back or face-to-face integration of heterogeneous ICs using 3D stacking and/or Silicon Interposers. While these technology features are extremely compelling, they considerably stress the existing design practices and EDA tool flows typically conceived for 2D systems. With all system, technology and implementation level options brought with these features, the design space increases to an extent where traditional 2D tools cannot be used any more for efficient exploration. Therefore, the cost-effective design of future 3D ICs products will require new planning and co-optimisation techniques and tools that are fast and accurate enough to cope with these challenges. In this paper we present design methodology and the practical EDA tool chain that covers different aspects of the design flow and is specific to efficient design of 3D-ICs. Flow features include: fast synthesis and 3D design partitioning at gate level, TSV/micro-bump array planning, 3D floor planning, placement and routing, congestion analysis, fast thermal and mechanical modeling, easy technology vs. implementation trade-off analysis, 3D device models generations and Design-for-Test (DfT). The application of the tool chain is illustrated using concrete example of a real-world design, showing not only the applicability of the tool chain, but also the benefits of heterogeneous 2.5 and 3D integration technologies. Dragomir Milojevic, Paul Marchal, Erik Jan Marinissen, Geert Van der Plas, Diederik Verkest, Eric Beyne |
ASP-DAC | 5 |
| 2013 | TEASE: a systematic analysis framework for early evaluation of FinFET-based advanced technology nodesabstractThis paper proposes TEASE (Technology Exploration and Analysis for SoC-level Evaluation), a framework to systematically analyze and evaluate system design in finFET-based technology node. The proposed framework combines both lithography and electrical constraints of a particular technology node to optimize the standard cell library performance. Growing complexity of logic design at nodes below 20nm causes to adopt a design style that can embrace the simplicity required to enable manufacturing, along with a process technology that can be finely tuned to the desired performance constraints. Additionally, the introduction of finFET based devices poses a new challenge for the designers to come up with an efficient standard cell template. The proposed framework can be used to detect the technology constraints that act as the bottleneck for the enablement of design at these advanced nodes. Results presented in this paper show by optimizing these bottlenecks we can improve the performance of a standard cell library significantly. Furthermore, adapting to such an analysis framework at an early stage of technology development helps to take the design constraints into the decision loop for realization of technology research into real products. Arindam Mallik, Paul Zuber, Tsung-Te Liu, Bharani Chava, Bhavana Ballal, Pablo Royer Del Bario, Rogier Baert, Kris Croes, Julien Ryckaert, Mustafa Badaroglu, Abdelkarim Mercha, Diederik Verkest |
DAC | 12 |
| 2011 | Real-time high-definition stereo matching on FPGAabstractAlthough many fast stereo matching designs have been proposed in the past decades, it is still very challenging to achieve real-time speed at high definition resolution while maintaining high matching accuracy. In this paper, we propose a real-time high definition stereo matching design on FPGA. By using the Mini-Census transform and the Cross-based cost aggregation, the proposed algorithm is robust to radiometric differences and produces accurate disparity maps. The algorithm modules have been optimized for efficient hardware implementations and instantiated in an SoC environment. Implemented on a single EP3SL150 FPGA, our design achieves 60 frames per second for 1024 × 768 stereo images. Evaluated with the Middlebury stereo benchmark, the proposed design also delivers leading stereo matching accuracy among prior related work. Lu Zhang 0018, Ke Zhang 0012, Tian-Sheuan Chang, Gauthier Lafruit, Georgi Kuzmanov, Diederik Verkest |
FPGA | 6 |
| 2010 | Modeling and exploiting spatial locality trade-offs in wavelet-based applications under varying resource requirementsabstractFuture dynamic applications will require new mapping strategies to deliver power-efficient performance. Fully static design-time mappings will not be able to optimally address the unpredictably varying application characteristics and system resource requirements. Instead, the platforms will not only need to be programmable in terms of instruction set processors, but also at least partial reconfigurability will be required, while the applications themselves will need to exploit this increased freedom at runtime to adapt to the dynamism. In this context, it is important for applications to optimally exploit the memory hierarchy under varying memory availability. This article presents an analysis of spatial locality trade-offs in wavelet-based applications, to be used in dynamic execution environments: Depending on the encountered runtime conditions, the execution switches to different memory optimized instantiations or localizations, optimally exploiting temporal and spatial locality under these conditions. This is enabled by systematic mapping guidelines, indicating how the miss-rate behavior of a localization is influenced by a specific execution condition, under which conditions a certain localization is optimal and which miss-rate gains may be obtained by switching to that localization. Bert Geelen, Vissarion Ferentinos, Francky Catthoor, Gauthier Lafruit, Diederik Verkest, Rudy Lauwereins, Thanos Stouraitis |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2009 | 3-D Technology Assessment: Path-Finding the Technology/Design Sweet-SpotabstractIt is widely acknowledged that three-dimensional (3-D) technologies offer numerous opportunities for system design. In recent years, significant progress has been made on these 3-D technologies, and they have become probably the best hope for carrying the semiconductor industry beyond the path of Moore's law. However, a clear roadmap is missing to successfully introduce this 3-D technology onto the market. Today, a plurality of 3-D technology options exists, which requires different design and test strategies. To crystallize the many technology options in a few mainstream technologies, it is mandatory to coexplore both technology and design options. The contribution of this paper is to introduce a novel path finding methodology to untangle the many intertwined design/technology options. This holistic approach will be applied on a representative 3-D case study. Initial results demonstrate the benefits of the proposed path-finding methodology to steer the technology development and fine-tune design strategies. Paul Marchal, Bruno Bougard, Guruprasad Katti, Michele Stucchi, Wim Dehaene, Antonis Papanikolaou, Diederik Verkest, Bart Swinnen, Eric Beyne |
Proc. IEEE | 7 |
| 2009 | Distributed Loop Controller for Multithreading in Unithreaded ILP ArchitecturesabstractReduced energy consumption is one of the most important design goals for embedded application domains like wireless communication, multimedia and biomedical applications. The instruction memory hierarchy has been proven to be one of the most power hungry parts of the system. This paper introduces an architectural enhancement for the instruction memory to reduce energy consumption and improve performance. The proposed distributed instruction memory organization requires minimal hardware overhead and supports the execution of multiple incompatible loops in parallel in a uni-processor system. We present different methods to implement the loop controller architecture, compare them, and show that distributing the instruction memory helps to reduce the interconnect cost as well. This architecture enhancement can reduce the energy consumed in the instruction memory hierarchy by 59% and improve the performance by 22% compared to hardware based enhanced SMT based architectures. Praveen Raghavan, Andy Lambrechts, Murali Jayapala, Francky Catthoor, Diederik Verkest |
IEEE Trans. Computers | 5 |
| 2009 | Spatial locality exploitation for runtime reordering of JPEG2000 wavelet data layoutsabstractExploitation of spatial locality is essential for memories to increase the access bandwidth and to reduce the access-related latency and energy per word. Spatial locality exploitation of a kernel can be improved by modifying placement of data in memory, but this may be felt not only by the kernel itself, but also in other application components accessing the same data. Thus care is needed to avoid global miss-rate improvements are thwarted by miss-rate increases in other application components. This article examines application-level miss-rate increases due to handling modified Wavelet Transform data layouts by explicitly reordering at runtime, exploiting the execution order freedom within a reordering buffer when the layout of surrounding components is known. For the JPEG2000 application, taking into account the reordering costs still results in 80% net WT miss-rate gains. Bert Geelen, Vissarion Ferentinos, Francky Catthoor, Gauthier Lafruit, Diederik Verkest, Rudy Lauwereins, Thanos Stouraitis |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2009 | Interconnect Exploration for Energy Versus Performance Tradeoffs for Coarse Grained Reconfigurable ArchitecturesabstractModern portable embedded devices require processors that can provide sufficient performance for demanding multimedia and wireless applications. At the same time they have to be flexible to support a wide range of products and extremely energy efficient to provide a long battery life. Coarse grained reconfigurable architectures (CGRAs) potentially meet these constraints by providing a mix of flexible computational resources and large amounts of programmable interconnect. The vast design space of CGRAs complicates the development of optimized processors. Most effort has been spent on improving the performance. However, the energy cost of the programmable interconnect is becoming more expensive and this cost can no longer be neglected. In this work we present an energy- and performance-aware exploration for the interconnect of a CGRA and show that important tradeoffs can be made for those metrics. This will enable designers to develop more efficient architectures, tuned to a targeted application domain. Andy Lambrechts, Praveen Raghavan, Murali Jayapala, Bingfeng Mei, Francky Catthoor, Diederik Verkest |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2008 | Coffee: COmpiler Framework for Energy-Aware Exploration
Praveen Raghavan, Andy Lambrechts, Javed Absar, Murali Jayapala, Francky Catthoor, Diederik Verkest |
HiPEAC | 6 |
| 2008 | Spatial locality trade-offs of wavelet-based applications in dynamic execution environmentsabstractFuture dynamic applications will require new mapping strategies to deliver power-efficient performance. Fully static design-time mappings will not address the unpredictably varying application characteristics and resource requirements. Instead, the platforms will not only need to be programmable in terms of instruction set processors, but also at least partial reconfigurability will be required, while the applications themselves will need to exploit this freedom at run-time to adapt to the dynamism. In this context, it is important for applications to exploit the memory hierarchy under varying memory availability. This paper presents a mapping strategy for wavelet-based applications: depending on the run-time conditions, it switches to different memory optimized instantiations, optimally exploiting temporal and spatial locality under these conditions. A comparison is performed between the gains of fully in-placed lifting-based wavelet transforms, and non in-placed versions with higher spatial locality. Bert Geelen, Aris Ferentinos, Francky Catthoor, Gauthier Lafruit, Diederik Verkest |
ICASSP | 5 |
| 2008 | Joint hardware-software leakage minimization approach for the register file of VLIW embedded architectures
David Atienza 0001, Praveen Raghavan, José Luis Ayala, Giovanni De Micheli, Francky Catthoor, Diederik Verkest, Marisa López-Vallejo |
Integr. | 6 |
| 2008 | Concepts and Implementation of Spatial Division Multiplexing for Guaranteed Throughput in Networks-on-ChipabstractTo ensure low power consumption while maintaining flexibility and performance, future Systems-on-Chip (SoC) will combine several types of processor cores and data memory units of widely different sizes. To interconnect the IPs of these heterogeneous platforms, Networks-on-Chip (NoC) have been proposed as an efficient and scalable alternative to shared buses. NoCs can provide throughput and latency guarantees by establishing virtual circuits between source and destination. State-of-the-art NoCs currently exploit Time-Division Multiplexing (TDM) to share network resources among virtual circuits, but this typically results in high network area and energy overhead with long circuit set-up time. We propose an alternative solution based on Spatial Division Multiplexing (SDM). This paper describes our design of an SDM-based network, discusses design alternatives for network implementation and shows why SDM can be better adapted to NoCs than TDM in a specific context. Our case study clearly illustrates the advantages of our technique over TDM in terms of energy consumption, area overhead, and flexibility. A comparison is also performed with a State-of-the-Art industrial reference NoC: Arteris. Anthony Leroy, Dragomir Milojevic, Diederik Verkest, Frédéric Robert, Francky Catthoor |
IEEE Trans. Computers | 3 |
| 2008 | Run-Time Management of a MPSoC Containing FPGA Fabric TilesabstractMultimedia applications, like, e.g., 3-D games and video decoders, are typically composed of communicating tasks. Their target embedded computing platforms (e.g., TI OMAP3, IBM Cell) contain multiple heterogeneous processing elements. At application design-time, it is often unknown which applications will execute simultaneously. Hence, resource assignment decisions need to be made by a run-time manager. Run-time assignment of these communicating tasks onto the communication and computation resources of such a multiprocessor platform is a challenging task. In the presence of fine-grain reconfigurable hardware processing elements, the run-time manager also needs to consider the creation of a so-called configuration hierarchy. Instead of executing a dedicated hardware task, the fine-grain reconfigurable hardware fabric hosts a programmable softcore block that, in turn, executes the task functionality. Hence, the next challenge for run-time management is to efficiently handle a configuration hierarchy. This paper details a run-time task assignment heuristic that performs fast and efficient task assignment in a multiprocessor system-on-chip containing fine-grain reconfigurable hardware tiles. In addition, this algorithm is capable of managing a configuration hierarchy. We show that being capable of handling a configuration hierarchy significantly improves the task assignment performance (i.e., success rate and assignment quality). In several cases, adding a configuration hierarchy improves the assignment success rate of the assignment heuristic by 20%. Vincent Nollet, Prabhat Avasare, Hendrik Eeckhaut, Diederik Verkest, Henk Corporaal |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2007 | Very wide register: an asymmetric register file organization for low power embedded processorsabstractIn current embedded systems processors, multi-ported register files are one of the most power hungry parts of the processor, even when they are clustered. This paper presents a novel register file architecture, which has single ported cells and asymmetric interfaces to the memory and to the datapath. Several realistic kernels from the TI DSP benchmark and from software defined radio (SDR) are mapped on the architecture. A complete physical design of the architecture is done in TSMC 90nm technology. The novel architecture presented is shown to obtain energy gains of up to 10times with respect to conventional multi-ported register file over the different benchmarks Praveen Raghavan, Andy Lambrechts, Murali Jayapala, Francky Catthoor, Diederik Verkest, Henk Corporaal |
DATE | 5 |
| 2006 | Distributed loop controller architecture for multi-threading in uni-threaded VLIW processorsabstractReduced energy consumption is one of the most important design goals for embedded application domains like wireless, multimedia and biomedical. Instruction memory hierarchy has been proven to be one of the most power hungry parts of the system. This paper introduces an architectural enhancement for the instruction memory to reduce energy and improve performance. The proposed distributed instruction memory organization requires minimal hardware overhead and allows execution of multiple loops in parallel in a uni-processor system. This architecture enhancement can reduce the energy consumed in the instruction and data memory hierarchy by 70.01 % and improve the performance by 32.89% compared to enhanced SMT based architectures Praveen Raghavan, Andy Lambrechts, Murali Jayapala, Francky Catthoor, Diederik Verkest |
DATE | 5 |
| 2005 | Power Breakdown Analysis for a Heterogeneous NoC Platform Running a Video ApplicationabstractUsers expect future handheld devices to provide extended multimedia functionality and have long battery life. This type of application imposes heavy constraints on performance and power consumption and forces designers to optimize all parts of their platform. Evaluating the overall platform power breakdown is therefore critical to determine where to spend the efforts on power optimization. Surprisingly, few studies exist on that topic and decisions generally rely on common belief. We have realized a complete power breakdown for a realistic platform to identify the major power bottlenecks. This paper presents this power assessment of a realistic heterogeneous network on chip platform including processors, network and data/instruction memory hierarchy, running a video processing chain from camera to display. Our power breakdown identifies the main bottlenecks in the memory hierarchy and the foreground memory, and shows that global interconnect is not that critical for a well-optimized application mapping. Andy Lambrechts, Praveen Raghavan, Anthony Leroy, Guillermo Talavera, Tom Vander Aa, Murali Jayapala, Francky Catthoor, Diederik Verkest, Geert Deconinck, Henk Corporaal, Frédéric Robert, Jordi Carrabina |
ASAP | 8 |
| 2005 | Low Cost Task Migration Initiation in a Heterogeneous MP-SoCabstractRun-time task migration in a heterogeneous multiprocessor system-on-chip (MP-SoC) is a challenge that requires cooperation between the task and the operating system. In task migration, minimization of the overhead during normal task execution (i.e., when not migrating) and the minimization of the migration reaction time are important. We introduce a novel technique that reuses the processor's debug registers in order to minimize the overhead during normal execution. This paper explains our task migration proof-of-concept setup and compares it to the state-of-the art. By reusing existing hardware and software functionality our approach reduces the run time overhead. Vincent Nollet, Prabhat Avasare, Jean-Yves Mignolet, Diederik Verkest |
DATE | 4 |
| 2005 | Centralized end-to-end flow control in a best-effort network-on-chipabstractRun-time communication management in a Network-on-Chip (NoC) is a challenging task. On one hand, the NoC needs to satisfy the communication requirements (e.g. throughput) of running applications competing for NoC resources. On the other hand, the NoC resources should be managed efficiently while keeping additional management functionalities minimal. This paper details a NoC communication management scheme based on a centralized, end-to-end flow control mechanism deployed in a best-effort NoC. This scheme comes at a very low resource (i.e. limited hardware and run-time overhead) cost. We show that by using a flow control mechanism it is possible, even in a best-effort NoC, to provide sufficient communication guarantees with respect to the application requirements. Finally, we illustrate the applicability of our approach for real-life multimedia applications. Prabhat Avasare, Vincent Nollet, Jean-Yves Mignolet, Diederik Verkest, Henk Corporaal |
EMSOFT | 4 |
| 2004 | Operating-system controlled network on chipabstractManaging a Network-on-Chip (NoC) in an efficient way is a challenging task. To succeed, the operating system (OS) needs to be tuned to the capabilities and the needs of the NoC. Only by creating a tight interaction can we combine the necessary flexibility with the required efficiency. This paper illustrates such an interaction by detailing the management of communication resources in a system containing a packet-switched NoC and a closely integrated OS. Our NoC system is emulated by linking an FPGA to a PDA. We show that, with the right NoC support, the OS is able to optimize communication resource usage. Additionally, the OS is able to diminish or remove the interference between independent applications sharing a common NoC communication resource. Vincent Nollet, Théodore Marescaux, Diederik Verkest, Jean-Yves Mignolet, Serge Vernalde |
DAC | 3 |
| 2004 | Design Methodology for a Tightly Coupled VLIW/Reconfigurable Matrix Architecture: A Case StudyabstractCoarse-grained reconfigurable architectures have seen growing importance recently. Design tools and methodology are essential to their success. Based on our previous work on modulo scheduling algorithms and a novel architecture with tightly coupled VLIW/reconfigurable matrix, we present a C-based design flow using an MPEG-2 decoder as a design example. The application is mapped to the architecture in less than one person-week starting from a software implementation. The kernel and overall speedup over the reference VLIW are 4.84 and 3.05 respectively. The case study shows that our methodology and architecture can deliver a competitive package in terms of design efforts and performance over other programmable architectures. Bingfeng Mei, Serge Vernalde, Diederik Verkest, Rudy Lauwereins |
DATE | 3 |
| 2004 | Design-Time Data-Access Analysis for Parallel Java Programs with Shared-Memory Communication Model
Richard Stahl, Francky Catthoor, Rudy Lauwereins, Diederik Verkest |
Euro-Par | 4 |
| 2004 | Network-on-Chip for Reconfigurable Systems: From High-Level Design Down to Implementation
T. Andrei Bartic, Dirk Desmet, Jean-Yves Mignolet, Théodore Marescaux, Diederik Verkest, Serge Vernalde, Rudy Lauwereins, Frédéric Robert |
FPL | 5 |
| 2004 | High-Level Data-Access Analysis for Characterisation of (Sub)task-Level Parallelism in JavaabstractIn the era of future embedded systems the designer is confronted with multi-processor systems both for performance and energy reasons. Exploiting (sub)task-level parallelism is becoming crucial because the instruction-level parallelism alone is insufficient. The challenge is to build compiler tools that support the exploration of the task-level parallelism in the programs. To achieve this goal, we have designed an analysis framework to evaluate the potential parallelism from sequential object-oriented programs. Parallel-performance and data-access analysis are the crucial techniques for estimation of the transformation effects. We have implemented support for platform-independent data-access analysis and profiling of Java programs, which is an extension to our earlier parallel-performance analysis framework. The toolkit comprises automated design-time analysis for performance and data-access characterisation, program instrumentation, program-profiling support and post-processing analysis. We demonstrate the usability of our approach on a number of realistic Java applications. Richard Stahl, Robert Pasko, Francky Catthoor, Rudy Lauwereins, Diederik Verkest |
HIPS | 5 |
| 2004 | Design Style Case Study for Embedded Multi Media Compute NodesabstractUsers expect future handheld devices to provide extended multimedia functionality and have long battery life. This type of application imposes heavy constraints on both (realtime) performance and energy consumption and forces designers to optimise all parts of their platform. In this experiment we focus on the different processor core design options for embedded platforms, including the effect of instruction memory hierarchy on the energy consumption. The results show that significant improvements for energy efficiency and/or performance over currently used RISC or VLIW processors can be achieved. We conclude, based on concrete data for a realistic application, that different styles, including both configurable hardware and instruction set processors, find their way into heterogeneous platforms and designers need to be aware of the trade-offs. Secondly, we show for the same application task that a heavily optimised instruction/configuration memory hierarchy can significantly reduce the energy consumption of this part, so it forms a crucial part of every energy aware design. Andy Lambrechts, Tom Vander Aa, Murali Jayapala, Guillermo Talavera, Anthony Leroy, Adelina Shickova, Francisco Barat, Bingfeng Mei, Francky Catthoor, Diederik Verkest, Geert Deconinck, Henk Corporaal, Frédéric Robert, Jordi Carrabina |
RTSS | 10 |
| 2004 | Run-time support for heterogeneous multitasking on reconfigurable SoCs
Théodore Marescaux, Vincent Nollet, Jean-Yves Mignolet, T. Andrei Bartic, W. Moffat, Prabhat Avasare, Paul Coene, Diederik Verkest, Serge Vernalde, Rudy Lauwereins |
Integr. | 8 |
| 2003 | Exploiting Loop-Level Parallelism on Coarse-Grained Reconfigurable Architectures Using Modulo Scheduling
Bingfeng Mei, Serge Vernalde, Diederik Verkest, Hugo De Man, Rudy Lauwereins |
DATE | 3 |
| 2003 | Infrastructure for Design and Management of Relocatable Tasks in a Heterogeneous Reconfigurable System-on-Chip
Jean-Yves Mignolet, Vincent Nollet, Paul Coene, Diederik Verkest, Serge Vernalde, Rudy Lauwereins |
DATE | 4 |
| 2003 | Application of Task Concurrency Management on Dynamically Reconfigurable Hardware PlatformsabstractDynamically reconfigurable hardware (DRHW) can take advantage of its reconfiguration capability to adapt at run-time its performance and its power consumption. However, due to the lack of programming support for dynamic task placement on these platforms, no previous work has been presented studying the performance/power trade-offs. To cope with the task placement problem in a straight way that allows us to go one step further, we have adopted an interconnection-network-based DRHW mode, which includes operating system support to reallocate tasks at run-time. On top of this model we have applied an emerging task concurrency management (TCM) methodology initially developed for multiprocessor platforms with promising results. Moreover, we have identified the next step needed to create a specific TCM support for DRHW platforms. Javier Resano, Diederik Verkest, Daniel Mozos, Serge Vernalde, Francky Catthoor |
FCCM | 2 |
| 2003 | Networks on Chip as Hardware Components of an OS for Reconfigurable Systems
Théodore Marescaux, Jean-Yves Mignolet, T. Andrei Bartic, W. Moffat, Diederik Verkest, Serge Vernalde, Rudy Lauwereins |
FPL | 5 |
| 2003 | ADRES: An Architecture with Tightly Coupled VLIW Processor and Coarse-Grained Reconfigurable Matrix
Bingfeng Mei, Serge Vernalde, Diederik Verkest, Hugo De Man, Rudy Lauwereins |
FPL | 3 |
| 2003 | Run-Time Minimization of Reconfiguration Overhead in Dynamically Reconfigurable Systems
Javier Resano, Daniel Mozos, Diederik Verkest, Serge Vernalde, Francky Catthoor |
FPL | 3 |
| 2003 | Performance Analysis for Identification of (Sub-)Task-Level Parallelism in Java
Richard Stahl, Robert Pasko, Luc Rijnders, Diederik Verkest, Serge Vernalde, Rudy Lauwereins, Francky Catthoor |
SCOPES | 4 |
| 2003 | Power-efficient flexible processor architecture for embedded applicationsabstractIn the design of embedded systems, a processor architecture is a tradeoff between energy consumption, area, speed, design time, and flexibility to cope with future design changes. New versions in a product generation may require small design changes in any part of the design. We propose a novel processor architecture concept, which provides the flexibility needed in practice at a reduced power and performance cost compared to a fully programmable processor. The crucial element is a novel protocol combining an efficient, customized component with a flexible processor into a hybrid architecture. Frederik Vermeulen, Francky Catthoor, Lode Nachtergaele, Diederik Verkest, Hugo De Man |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2002 | System-level performance optimization of the data queueing memory management in high-speed network processorsabstractIn high-speed network processors, data queueing has to allow real-time memory (de)allocation, buffering, retrieving, and forwarding of incoming data packets. Its implementation must be highly optimized to combine high speed, low power, large data storage, and high memory bandwidth. In this paper, such data queueing is used as case study to demonstrate the effectiveness of a new system-level exploration method for optimizing the memory performance in dynamic memory management. Assuming that a multi-bank memory architecture is used for data storage, the method trades off bank conflicts against memory accesses during real-time memory (de)allocation. It has been applied to the data queueing module of the PRO3 system [8]. Compared with the conventional memory management technique for embedded systems, our exploration method can save up to 90% of the bank conflicts, which allows to improve worst-case memory performance of data queueing operations by 50% too. Chantal Ykman-Couvreur, Jurgen Lambrecht, Diederik Verkest, Francky Catthoor, Aristides Nikologiannis, George E. Konstantoulakis |
DAC | 3 |
| 2002 | Design Technology for Networked Reconfigurable FPGA PlatformsabstractFuture networked appliances should be able to download new services or upgrades from the network and execute them locally. This flexibility is typically achieved by processors that can download new software over the network, using JAVA technology. The paper demonstrates that FPGAs are a realistic implementation platform for thin server or client applications. FPGAs can offer the same end-user experience as software based systems, combined with more computational power and lower cost. Steven A. Guccione, Diederik Verkest, Ivo Bolsens |
DATE | 2 |
| 2002 | Adding Hardware Support to the HotSpot Virtual Machine for Domain Specific Applications
Yajun Ha, Radovan Hipik, Serge Vernalde, Diederik Verkest, Marc Engels, Rudy Lauwereins, Hugo De Man |
FPL | 4 |
| 2002 | Interconnection Networks Enable Fine-Grain Dynamic Multi-tasking on FPGAs
Théodore Marescaux, T. Andrei Bartic, Diederik Verkest, Serge Vernalde, Rudy Lauwereins |
FPL | 3 |
| 2002 | DRESC: a retargetable compiler for coarse-grained reconfigurable architecturesabstractCoarse-grained reconfigurable architectures have become increasingly important in recent years. Automatic design or compiling tools are essential to their success. In this paper, we present a retargetable compiler for a family of coarse-grained reconfigurable architectures. Several key issues are addressed. Program analysis and transformation prepare dataflow for scheduling. Architecture abstraction generates an internal graph representation from a concrete architecture description. A modulo scheduling algorithm is key to exploit parallelism and achieve high performance. The experimental results show up to 28.7 instructions per cycle (IPC) over tested kernels. Bingfeng Mei, Serge Vernalde, Diederik Verkest, Hugo De Man, Rudy Lauwereins |
FPT | 3 |
| 2002 | Dynamic memory management methodology applied to embedded telecom network systemsabstractPresents a new methodology for dynamic memory management of embedded telecom network systems. This methodology enables the designer to further raise the abstraction level of the initial system specification and to achieve optimized embedded system designs. This methodology is well suited for systems characterized by a set of concurrent and dynamic processes, very high-bit-rate data streams, and intensive data transfer and storage, as encountered in telecom network applications. Up to now, it has been successfully applied to four telecom network systems. This methodology can be easily integrated into any C++-based system synthesis approach that bridges the gap between a concurrent process-level system specification and an optimized (for area, performance, or power) embedded implementation of communicating hardware/software processors. This is in contrast to current system design practice, where VHDL/C is derived without room for exploration, refinement, and verification, leading to expensive late design iterations. In this paper, the main focus lies on the system-level specification model and the dynamic memory management applied to two real-life telecom network systems. Chantal Ykman-Couvreur, Jurgen Lambrecht, Diederik Verkest, Francky Catthoor, Bengt Svantesson, Ahmed Hemani, F. Wolf |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2001 | Task concurrency management methodology summaryabstractThis paper summarizes a new methodology for the design of concurrent dynamic real-time embedded systems. An embedded system can be specified at a grey-box abstraction level in a combined MTG-CDFG model. The authors believe that task concurrency management can be implemented in four major steps. Firstly, the grey box model is built, including the necessary concurrency extraction. Then transformations are applied on the specified MTG-CDFG to increase the opportunities for concurrency exploration and cost minimization. Then static scheduling will be applied on the design time analyzable parts of the grey-box model, including processor assignment in the multiple processor context. Finally, a dynamic scheduler will schedule the dynamic and coarse-grain constructs at run time on the given platform while making trade-offs based on Pareto curves. Chun Wong, Paul Marchal, Francky Catthoor, Hugo De Man, Aggeliki S. Prayati, Nathalie Cossement, Rudy Lauwereins, Diederik Verkest |
DATE | 9 |
| 2001 | Optimisation Problems for Dynamic Concurrent Task-Based SystemsabstractOne of the most critical bottlenecks in many novel multi-media applications is their very dynamic concurrent behaviour. This is especially true because of the quality-of-service (QoS) aspects of these applications. Prominent examples of this can be found in MPEG4 and JPEG2000 standards and especially the new MPEG21 standard. In order to deal with these dynamic issues where tasks and complex data types are created and deleted at run-time based on non-deterministic events, a novel system design paradigm is required. Because of the high computation and communicating requirement from these kind of applications, multi-processor SoC (system on chip) is accepted as a solution (e.g., TriMedia), which is also promised by the advance of the processing technology. It is different from traditional general-purpose computers because it is application specific and it is an embedded system, which means costs like energy consumption are of major concern. This paper focused on the new requirements in system-level synthesis. In particular a "task concurrency management" problem formulation is proposed, with special focus on the formal definition of the most crucial optimization problems. The concept of Pareto curve based exploration is crucial in these formulations. Diederik Verkest, Chun Wong, Paul Marchal |
ICCAD | 1 |
| 2000 | Dynamic scheduling of concurrent tasks with cost performance trade-offabstractThis paper addresses the run time task scheduling problem on a multiprocessor platform for embedded systems, where energy consumption is a major concern, as opposed to the traditional static and dynamic scheduling approaches.Our approach i n tends to combine the advantages of the low run time complexity o f the static scheduler and the exibility of the dynamic scheduler and to optimize the system energy consumption at run time based on precomputed costperformance Pareto curves.We have applied our method to an ADSL modem application and the result shows the eectiveness of our method.Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page.To copy otherwise, to republish, to post on servers or to redistribute to lists, requires prior specific Dirk Desmet, Francky Catthoor, Diederik Verkest |
CASES | 4 |
| 2000 | Operating system based software generation for systems-on-chipabstractIn this paper we propose a system-level design environment, aimed at System-on-Chip (SOC) designs, including real-time embedded software. While many SOC modeling languages originate from hardware description languages, and thus tend to describe statical architectures, we observe that embedded software makes SOC designs essentially dynamic, and so a SOC modeling environment must include dynamic behavior. Such behavior is analogous to the services an Operating System offers in the software world, hence the term System-on-Chip Operating System (SoCOS). Dirk Desmet, Diederik Verkest, Hugo De Man |
DAC | 2 |
| 2000 | System Level Design Using C++abstractThis paper discusses the use of C++ for the design of digital systems. The paper distinguishes a number of different approaches towards the use of programming languages for digital system design and will discuss in more detail how C++ can be used for system modeling and refinement, for simulation, and for architecture design. Diederik Verkest, Joachim Kunkel, Frank Schirrmeister |
DATE | 1 |
| 2000 | Formalized Three-Layer System-Level Reuse Model and Methodology for Embedded Data-Dominated ApplicationsabstractIn embedded data-dominated applications a global system-level data transfer and storage exploration phase is crucial in obtaining an efficient solution. We have developed a novel formalism to describe reusable blocks such that the essential part of the design exploration freedom is retained. This formalism is the basis for a system-level reuse methodology which allows to reuse large parts of the design as structural VHDL and describes the costly data access related constructs at higher levels in the code hierarchy. Compared to a reuse approach based on fixed blocks, considerable power and area savings can be obtained as demonstrated on real-life video and modem applications. Frederik Vermeulen, Francky Catthoor, Hugo De Man, Diederik Verkest |
DATE | 4 |
| 2000 | Flexible hardware acceleration for multimedia oriented microprocessorsabstractThe execution of multimedia applications on a microprocessor greatly benefits from hardware acceleration, both in terms of speed and energy consumption. While the basic functionality implemented in these accelerators remains constant over different product versions, small changes are still often required. With the proposed architecture and protocol, the accelerator hardware has the performance and cost benefits of a hardwired solution, while featuring all the flexibility needed in practice. From a user point of view, the entire application is still programmable. Frederik Vermeulen, Lode Nachtergaele, Francky Catthoor, Diederik Verkest, Hugo De Man |
MICRO | 4 |
| 2000 | Formalized three-layer system-level model and reuse methodology for embedded data-dominated applicationsabstractIn embedded data-dominated applications, a global system-level data transfer and storage exploration phase is crucial in obtaining a cost- and performance-efficient solution. We have developed a novel formalism to describe reusable blocks such that the essential part of the design exploration freedom is retained. This formalism is the basis for a system-level reuse methodology which allows reusing large parts of the design as heavily optimized structural VHDL or assembly code and describes the costly data access-related constructs at higher levels in the code hierarchy. Compared to a reuse approach based on fixed blocks, considerable power and area savings can be obtained, as demonstrated on real-life video and modem applications. Frederik Vermeulen, Francky Catthoor, Diederik Verkest, Hugo De Man |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1999 | Global Multimedia System Design Exploration Using Accurate Memory Organization FeedbackabstractSuccessful exploration of system-level design decisions is impossible without fast and accurate estimation of the impact on the system cost.In most multimedia applications, the dominant cost factor is related to the organization of the memory architecture.This paper presents a systematic approach which allows effective system-level exploration of memory organization design alternatives, based on accurate feedback by using our earlier developed tools.The effectiveness of this approach is illustrated on an industrial application.Applying our approach, a substantial part of the design search space has been explored in a very short time, resulting in a cost-efficient solution which meets all design constraints. Arnout Vandecappelle, Miguel Corbalan, Erik Brockmeyer, Francky Catthoor, Diederik Verkest |
DAC | 5 |
| 1999 | Combining Software Synthesis and Hardware/Software Interface Generation to Meet Hard Real-Time ConstraintsabstractThis paper presents an orchestrated combination of software synthesis and automatic hardware/software interface generation to meet hard real-time constraints. The target applications are digital communication systems, which are specified as concurrent communicating processes. The overall approach is theoretically founded and is demonstrated on an industrial strength design, with promising results. Steven Vercauteren, Jan van der Steen, Diederik Verkest |
DATE | 3 |
| 1999 | Guest Editorial
Gaetano Borriello, Diederik Verkest, Francky Catthoor |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | Efficient System Exploration and Synthesis of Applications with Dynamic Data Storage and Intensive Data TransferabstractMatisse is a design flow intended for developing embedded systems characterize dby tight inter action b etwe encontrol and data-flow behavior, intensive data storage and tr ansfer, dynamic creation of data, and stringent real-time requirements. Matisse bridges the gap from a system specification, using a cocurr ent obje ct-oriented language, to an optimize d embedded single-chip HW/SW implementation. Matisse supp orts stepwise system-level exploration and refinement, memory architecture exploration, and gradualincorporation of timing constr aints b efore going to tr aditional tools for HW synthesis, SW compilation, and HW/SW interprocessor communication synthesis. Application of Matisse on telecom protocol processing systems shows significant improvements in area usage and power c onsumption. Julio Leao da Silva Jr., Chantal Ykman-Couvreur, Miguel Corbalan, Kris Croes, Sven Wuytack, Gjalt G. de Jong, Francky Catthoor, Diederik Verkest, Paul Six, Hugo De Man |
DAC | 8 |
| 1998 | Efficient Verification using Generalized Partial Order AnalysisabstractThis paper presents a new formal method for the efficient verification of concurrent systems that are modeled using a safe Petri net representation. Our method generalizes upon partial-order methods to explore concurrently enabled conflicting paths simultaneously. We show that our method can achieve an exponential reduction in algorithmic complexity without resorting to an implicit enumeration approach. Steven Vercauteren, Diederik Verkest, Gjalt G. de Jong, Bill Lin 0001 |
DATE | 2 |
| 1998 | Power exploration for dynamic data types through virtual memory management refinementabstractIn this p ap er we pr esent our novelpower exploration methodolo gy for applic ations with dynamic data types. Our methodolo gy is crucial to obtain effe ctive solutions in an emb edded (HW or SW) processor context. The c ontributionsare twofold. First we define the complete search sp ace for Virtual Memory Management (VMM) mechanisms in a structured way with orthogonal de cision tr eesSe condly we present our systematic methodolo gy for explor ation of the maximal power that takes into account characteristics of the application to he avily prune the search space guiding the choies of a VMM mechanism. Finally we demonstrate for two industrial examples that power can vary consider ably dep ending on the VMM chosen. Moreover these exp eriments show the effe ctiveness of our exploration methodolo gy. Julio Leao da Silva Jr., Francky Catthoor, Diederik Verkest, Hugo De Man |
ISLPED | 3 |
| 1997 | Hardware/software co-design of digital telecommunication systemsabstractWe reflect on the nature of digital telecommunication systems. We argue that these systems require, by nature, a heterogeneous specification and an implementation with heterogeneous architectural styles. CoWare is a hardware/software co-design environment based on a data model that allows to specify, simulate, and synthesize heterogeneous hardware/software architectures from a heterogeneous specification. CoWare is based on the principle of encapsulation of existing hardware and software compilers and special attention is paid to the interactive synthesis of hardware/software and hardware/hardware interfaces. The principles of CoWare are illustrated by the design process of a spread-spectrum receiver for a pager system. Ivo Bolsens, Hugo De Man, Bill Lin 0001, Karl van Rompaey, Steven Vercauteren, Diederik Verkest |
Proc. IEEE | 6 |
| 1994 | A Proof of the Nonrestoring Division Algorithm and its Implementation on an ALU
Diederik Verkest, Luc Claesen, Hugo De Man |
Formal Methods Syst. Des. | 1 |
| 1993 | On the Comparison of HOL and Boyer-Moore for Formal Hardware Verification
Catia M. Angelo, Diederik Verkest, Luc Claesen, Hugo De Man |
Formal Methods Syst. Des. | 2 |
| 1989 | Application of system semantics to VLSI for the transformational design of a parameterized booth multiplier module - a case study
Luc Claesen, Raymond T. Boute, Jozef De Man, W. Ploegaerts, Marc Seutter, Johan Vanslembrouck, Diederik Verkest |
Microprocessing and Microprogramming | 7 |
| 1989 | Description and verification of more-dimensional regular and non-homogeneous structures using a functional hardware description language
W. Ploegaerts, Diederik Verkest, Luc Claesen, Hugo De Man |
Microprocessing and Microprogramming | 2 |