EDBT 2026 Demo / reviewers in the wild / expert
Dirk Stroobandt
dblp:14/2499
· DBLP profile ↗
108ranked-venue papers
10as first author
4since 2021 · last 2025
0000-0002-4477-5313ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 84 · 10 first-author · 4 since 2021Artificial intelligence and machine learning · 18Software engineering, systems software and programming languages · 7 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reducing FPGA Placement Runtime by Clustering of Netlist BlocksabstractThe physical design phase of FPGA programming includes the placement step, where logic blocks of the netlist are mapped to hardware primitives. This work explores an hierarchical strategy for initial placement that reduces complexity of the placement problem by first grouping related blocks into clusters. Such reduced netlist could further be used during placement optimization. Here, a placement is made in two stages: an intercluster placement to determine coarse locations, followed by intra-cluster refinement. Experimental results demonstrate that the approach achieves comparable timing and wirelength quality to existing methods, although runtime performance remains an area for further tweaks. Markus Rein, Dirk Stroobandt |
FPL | 2 |
| 2025 | Interconnection-Aware Resynthesis for Improving FPGA Physical DesignabstractThe complexity of interconnects is a critical, yet often underestimated, bottleneck in FPGA physical design. FPGA compilation is increasingly dominated by interconnect complexity, which limits routability and slows down placement and routing. We propose an interconnection-aware logic resynthesis flow that explicitly targets netlist structures with high fanout and interconnect congestion. We leverage local interconnection complexity metrics to guide resynthesis decisions. By allowing controlled node duplication and restructuring, our method improves logic locality and wireability at the expense of minor area overhead. Preliminary results show that this strategy can reduce routed wirelength and improve timing without changing the logic function, suggesting the potential to complement existing logic synthesis pipelines. Dirk Stroobandt |
FPL | 2 |
| 2022 | RWRoute: An Open-source Timing-driven Router for Commercial FPGAsabstractOne of the key obstacles to pervasive deployment of FPGA accelerators in data centers is their cumbersome programming model. Open source tooling is suggested as a way to develop alternative EDA tools to remedy this issue. Open source FPGA CAD tools have traditionally targeted academic hypothetical architectures, making them impractical for commercial devices. Recently, there have been efforts to develop open source back-end tools targeting commercial devices. These tools claim to follow an alternate data-driven approach that allows them to be more adaptable to the domain requirements such as faster compile time. In this paper, we present RWRoute, the first open source timing-driven router for UltraScale+ devices. RWRoute is built on the RapidWright framework and includes the essential and pragmatic features found in commercial FPGA routers that are often missing from open source tools. Another valuable contribution of this work is an open-source lightweight timing model with high fidelity timing approximations. By leveraging a combination of architectural knowledge, repeating patterns, and extensive analysis of Vivado timing reports, we obtain a slightly pessimistic, lumped delay model within 2% average accuracy of Vivado for UltraScale+ devices. Compared to Vivado, RWRoute results in a 4.9× compile time improvement at the expense of 10% Quality of Results (QoR) loss for 665 synthetic and six real designs. A main benefit of our router is enabling fast partial routing at the back-end of a domain-specific flow. Our initial results indicate that more than 9× compile time improvement is achievable for partial routing. The results of this paper show how such a router can be beneficial for a low touch flow to reduce dependency on commercial tools. Pongstorn Maidee, Chris Lavin, Alireza Kaviani, Dirk Stroobandt |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2021 | cREAtIve: reconfigurable embedded artificial intelligenceabstractcREAtIve targets the development of novel highly-adaptable embedded deep learning solutions for automotive and traffic monitoring applications, including position sensor processing, scene interpretation based on LiDAR, and object detection and classification in thermal images for traffic camera systems. These applications share the need for deep learning solutions tailored for deployment on embedded devices with limited resources and featuring high adaptability and robustness to changing environmental conditions. cREAtIve develops knowledge, tools and methods that enable hardware-efficient, adaptable, and robust deep learning. Poona Bahrebar, Leon Denis, Maxim Bonnaerens, Kristof Coddens, Joni Dambre, Wouter Favoreel, Illia Khvastunov, Adrian Munteanu 0001, Hung Nguyen-Duc, Stefan Schulte 0001, Dirk Stroobandt, Ramses Valvekens, Nick Van den Broeck, Geert Verbruggen |
CF | 11 |
| 2020 | On the Exploration of Connection-aware Partitioning for Parallel FPGA RoutingabstractRouting is one of the most time-consuming steps in the FPGA synthesis flow. Existing works have described several ways to accelerate the routing process. The partitioning-based parallel routing technique that leverages the high-performance computing of multi-core processors are gaining popularity recently. Specifically, those parallel routers partition nets to regions by nets' bounding boxes, followed by a parallel routing procedure. Nets can be split up into source-sink connections that share wire segments as much as possible. In order to exploit more parallelism by a finer granularity in both spatial partitioning and routing, a connection-aware routing bounding box model is introduced in this work. We first explore in detail to show that connection-aware partitioning using the new routing bounding boxes enables the parallel routing to perform better runtime efficiency than the existing net-based partitioning by analyzing the workloads of parallel routers. It reduces the connections spanning more than one region and exploits more parallelism. The large heterogeneous Titan23 designs and a detailed representation of the Stratix IV FPGA are used for benchmarking. Experimental results show that the parallel FPGA router is faster when using our connection-aware partitioning than using the existing net-based partitioning, while achieving similar quality of routing results in terms of the wirelength and critical path delay. The connection-aware routing bounding box model is easy to be embedded into other existing parallel routers and further enables them to be faster. Dries Vercruyce, Dirk Stroobandt |
FPGA | 3 |
| 2020 | In-Circuit Debugging with Dynamic Reconfiguration of FPGA InterconnectsabstractIn this work, a novel method for in-circuit debugging on FPGAs is introduced that allows the insertion of low-overhead debugging infrastructure by exploiting the technique of parameterized configurations. This allows the parameterization of the LUTs and the routing infrastructure to create a virtual network of debugging multiplexers. It aims to facilitate debugging, to increase the internal signal observability, and to reduce the debugging (area and reconfiguration) overhead. Signal ranking techniques are also introduced that classify signals that can be traced during debug. Finally, the results of the method are presented and compared with a commercial tool. The area and time results and the tradeoffs between internal signal observability and area and reconfiguration overhead are also explored. Alexandra Kourfali, Dirk Stroobandt |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2020 | Accelerating FPGA Routing Through Algorithmic Enhancements and Connection-aware ParallelizationabstractRouting is a crucial step in Field Programmable Gate Array (FPGA) physical design, as it determines the routes of signals in the circuit, which impacts the design implementation quality significantly. It can be very time-consuming to successfully route all the signals of large circuits that utilize many FPGA resources. Attempts have been made to shorten the routing runtime for efficient design exploration while expecting high-quality implementations. In this work, we elaborate on the connection-based routing strategy and algorithmic enhancements to improve the serial FPGA routing. We also explore a recursive partitioning-based parallelization technique to further accelerate the routing process. To exploit more parallelism by a finer granularity in both spatial partitioning and routing, a connection-aware routing bounding box model is proposed for the source-sink connections of the nets. It is built upon the location information of each connection’s source, sink, and the geometric center of the net that the connection belongs to, different from the existing net-based routing bounding box that covers all the pins of the entire net. We present that the proposed connection-aware routing bounding box is more beneficial for parallel routing than the existing net-based routing bounding box. The quality and runtime of the serial and multi-threaded routers are compared to the router in VPR 7.0.7. The large heterogeneous Titan23 designs that are targeted to a detailed representation of the Stratix IV FPGA are used for benchmarking. With eight threads, the parallel router using the connection-aware routing bounding box model reaches a speedup of 6.1× over the serial router in VPR 7.0.7, which is 1.24× faster than the one using the existing net-based routing bounding box model, while reducing the total wire-length by 10% and the critical path delay by 7%. Dries Vercruyce, Dirk Stroobandt |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2019 | CRoute: A Fast High-Quality Timing-Driven Connection-Based FPGA RouterabstractFPGA routing is an important part of physical design as the programmable interconnection network requires the majority of the total silicon area and the connections largely contribute to delay and power. It should also occur with minimum runtime to enable efficient design exploration. In this work we elaborate on the concept of the connection-based routing principle. The algorithm is improved and a timing-driven version is introduced. The router, called CRoute, is implemented in an easy to adapt FPGA CAD framework written in Java, which is publicly available on GitHub. Quality and runtime are compared to the state-of-the-art router in VPR 7.0.7. Benchmarking is done with the Titan23 design suite, which consists of large heterogeneous designs targeted to a detailed representation of the Stratix IV FPGA. CRoute gains in both the total wire-length and maximum clock frequency while reducing the routing runtime. The total wire-length reduces by 11% and the maximum clock frequency increases by 6%. These high-quality results are obtained in 3.4x less routing runtime. Dries Vercruyce, Elias Vansteenkiste, Dirk Stroobandt |
FCCM | 3 |
| 2019 | MODA-PSO: Towards Fast Hard Block Legalization for Analytical FPGA PlacementabstractPlacement is a crucial step in the FPGA design tool flow, as it determines the overall performance of the circuits. Unfortunately, it is a time-consuming task. Analytical placers have been shown to be the most time-efficient while retaining good quality. One way of implementing analytical placement is to use an iterative technique that consists of optimization and look-ahead legalization, followed by an optional refinement step. In this work, with the aim towards fast hard block legalization for further accelerating analytical placement, a novel optimizer is proposed based on the modified discrete adaptive particle swarm optimization. The proposed optimizer is embedded into the publicly available analytical placer Liquid. When compared to its version using simulated annealing for hard block legalization, this approach results in a 30% reduction in hard block legalization time and a consequent 5% runtime reduction for the analytical placement, at the cost of only a 1% increase in post-routed wirelength and critical path delay. The results indicate that the nature-inspired particle swarm optimization is promising for tackling such a problem with new learning strategies and adaptation. Dries Vercruyce, Dirk Stroobandt |
FPGA | 3 |
| 2019 | An Integrated on-Silicon Verification Method for FPGA Overlays
Alexandra Kourfali, Florian Fricke, Michael Hübner 0001, Dirk Stroobandt |
J. Electron. Test. | 4 |
| 2019 | Reco-Pi: A reconfigurable Cryptoprocessor for π-Cipher
Mohamed El-Hadedy 0001, Amit Kulkarni 0002, Dirk Stroobandt, Kevin Skadron |
J. Parallel Distributed Comput. | 3 |
| 2018 | Reconfigurable FPGA Implementation of the AVC Quantiser and De-quantiser Blocks
Vijaykumar Guddad, Amit Kulkarni 0002, Dirk Stroobandt |
ACIVS | 3 |
| 2018 | Hierarchical Force-Based Block Spreading for Analytical FPGA PlacementabstractEnabling efficient FPGA application development requires fast design compilation to high quality FPGA configurations. Placement and routing are the most challenging compilation steps. A promising trend to accelerate the placement problem is the use of iterative analytical placement techniques. Each iteration in these placers consists of an optimization and a legalization phase. The legalization phase is an important part of the technique, reducing block overlap in the optimized placements. In this work we propose a hierarchical pushing force-based block spreader that improves the legalization of CLBs. Overlap is reduced as the blocks push each other away under the influence of gravity, like a drop of water flowing over a surface. The algorithm is designed with parallelism in mind, making it suitable for GPU acceleration. The regularity of FPGAs is exploited to accelerate the calculations. Moreover, it further builds on a hierarchical FPGA CAD tool flow to enhance quality of results and runtime scalability. The spreading legalizer is embedded in the Liquid placement technique. It replaces the default partitioning-based legalizer for sparse designs. The result is a reduction of 7.4% in post-routing total wire-length, without giving in on runtime. The source code is publicly available in the FPGA CAD framework on GitHub. Dries Vercruyce, Elias Vansteenkiste, Dirk Stroobandt |
FPL | 3 |
| 2018 | How Preserving Circuit Design Hierarchy During FPGA Packing Leads to Better PerformanceabstractGenerating a configuration for a field-programmable gate array (FPGA) starting from a high level description of a design is a time consuming task. The resulting configuration should have a high quality so that the FPGA resources are used in an efficient way while being able to run at high clock frequencies and having a low power consumption. In this paper, we present MultiPart, a new hierarchical packing algorithm that obtains better quality and faster runtimes when compared to the frequently used AAPack packer in VPR. MultiPart combines the benefits of partitioning-based and seed-based packing approaches. It tries to preserve the design hierarchy during packing. This results in a gain of 32% in total wirelength and a gain of 10% in critical path delay. The partitioning-based methodology allows us to exploit multithreading, leading to $9.3 {\times }$ faster packing runtimes on a CPU with 10 cores. We also gain in the total routing runtime because MultiPart reduces congestion problems on a higher level. The subcircuits in the partitioned circuit are clustered with a seed-based packer. This allows MultiPart to deal with the constraints of complex heterogeneous architectures. In short, MultiPart targets heterogeneous commercial FPGAs with a lower runtime while increasing the quality of the configuration. The source code of MultiPart is available in our FPGA CAD framework on Github. Dries Vercruyce, Elias Vansteenkiste, Dirk Stroobandt |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Data Reuse Buffer Synthesis Using the Polyhedral Model
Wim Meeus, Dirk Stroobandt |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | An open reconfigurable research platform as stepping stone to exascale high-performance computingabstractTo handle the stringent performance and power requirements of future exascale-class applications, High Performance Computing (HPC) systems need ultra-efficient heterogeneous compute nodes and hardware accelerators with a high degree of specialization. Ideally, dynamic reconfiguration will be an intrinsic feature, so that specific HPC application features can be optimally accelerated, even if they regularly change over time. We create a new and flexible exploration platform for developing reconfigurable architectures, design tools and HPC applications with run-time reconfiguration built-in as a core fundamental feature instead of an add-on. Our project proposes an open research platform that covers the entire stack from architecture up to the application, focusing on the fundamental building blocks for run-time reconfigurable exascale HPC systems: new chip architectures with very low reconfiguration overhead, new tools that truly take reconfiguration as a central design concept, and applications that are tuned to maximally benefit from the proposed run-time reconfiguration techniques. Ultimately, this open platform will enable groundbreaking research towards new exascale computing platforms. Dirk Stroobandt, Catalin Bogdan Ciobanu, Marco D. Santambrogio, Gabriel Figueiredo, Andreas Brokalakis, Dionisios N. Pnevmatikatos, Michael Hübner 0001, Tobias Becker, Alex J. W. Thom |
DATE | 1 |
| 2017 | A NoC-based custom FPGA configuration memory architecture for ultra-fast micro-reconfigurationabstractRun-time reconfiguration in FPGAs is an important feature that offers design flexibility under low-cost silicon area and power budgets, at the cost of reconfiguration overhead. The reconfiguration time overhead produced by the conventional configuration ports (such as ICAP) is too high for the reconfiguration technology to be embraced as a standard. Furthermore, the current FPGA configuration memory architecture restricts the access of configuration data to the frame level; this significantly delays the reconfiguration process. The work presented in this paper explores the design space of the configuration memory architecture that fits the design of large FPGA's and is suitable to accomplish needs for ultra-fast reconfiguration. Therefore, the proposed method could be a stepping stone for next generation FPGA configuration memory architecture. Our simulation results show a reconfiguration speed gain of a factor of at least 1000 for substantially big parameterized applications that come with the cost of extra auxiliary hardware used on top of the column-based FPGA architecture. Amit Kulkarni 0002, Poona Bahrebar, Dirk Stroobandt, Giulio Stramondo, Catalin Bogdan Ciobanu, Ana Lucia Varbanescu |
FPT | 3 |
| 2017 | Liquid: High quality scalable placement for large heterogeneous FPGAsabstractGenerating a configuration for an FPGA is a time consuming task. Most time is required for placement and routing. Placing one of the large Titan23 designs can take more than an hour with the placer in VPR. This is too long to allow efficient turnaround times. New placement techniques are proposed to speed up the process. LIQUID is a new fast placement prototyping technique that is based on analytical placement but without exactly solving the linear system of equations. Instead, the blocks are moved in several small steps in the direction that reduces the placement cost the most. In this work we introduce improvements to the LIQUID placement technique so that it can be used as a high quality placer. The main contribution is a new legalization method for the hard blocks in heterogeneous designs. We achieve a gain of 14% in wire-length cost and 15% smaller critical path delays when compared to the original version of LIQUID. Most analytical placement techniques produce a placement in two steps. First a global placement prototype is generated. The prototype is then further optimized in a refinement step. With the improved version of LIQUID, the same quality of results is obtained when compared to the placer in VPR without the need for a refinement step, leading to 23.7× faster runtimes. Dries Vercruyce, Elias Vansteenkiste, Dirk Stroobandt |
FPT | 3 |
| 2017 | Dynamically Reconfigurable Architecture for Fault-Tolerant 2D Networks-on-ChipabstractWith the increasing device scaling in the semiconductor technology, the necessity for designing robust and efficient Networks-on-Chip (NoCs) is more pronounced. The rerouting approach which is employed in most of the fault-tolerant methods causes the network performance to degrade considerably due to taking longer paths and creating hotspots around the faults. In this paper, a dynamically reconfigurable technique is proposed to target fault-tolerance and minimal routing in a unified manner. To accomplish this goal, the router architecture is modified to enable the frequently communicating nodes to bypass the faulty router and communicate through shorter paths. Thus, not only the rerouting is minimized, the connectivity of the network is maintained in the vicinity of faults. The experimental results validate the performance and reliability of the proposed technique with a small hardware overhead. Poona Bahrebar, Azarakhsh Jalalvand, Dirk Stroobandt |
ICCCN | 3 |
| 2017 | SICTA: A superimposed in-circuit fault tolerant architecture for SRAM-based FPGAsabstractReassuring fault tolerance in computing systems that contain FPGA devices is the most important problem for mission critical space components. With the rise in interest of commercial SRAM-based FPGAs, it is crucial to provide runtime reconfigurable recovery from a failure. In this paper, we propose a superimposed virtual coarse-grained reconfigurable architecture, embedded with on-demand three level fault-mitigation technique. The proposed method performs run-time recovery via discrete microscrubbing. This approach can provide up to 3× faster runtime recovery with 10.2× less resources in FPGA devices, by providing integrated layers of fault mitigation. Alexandra Kourfali, Amit Kulkarni 0002, Dirk Stroobandt |
IOLTS | 3 |
| 2016 | Liquid: Fast placement prototyping through steepest gradient descent movementabstractFPGA design compilation takes too much time to allow efficient design turnaround times. The largest runtime consuming steps of the compilation are placement and routing. To speed up the FPGA placement process, analytical placement techniques have become more popular in the last decade. Analytical techniques produce a placement in two steps, a placement prototyping step and a refinement step. In this work we focus on fast FPGA placement prototyping. Placement prototypes are also used to obtain fast accurate timing estimations and speed up the design cycle. In conventional analytical placement prototyping techniques the placement problem is formulated as a linear system which is solved several times to find a good legal placement. The most time consuming step of that process is solving the linear system. We show that it is not necessary to exactly solve this system, but that it is sufficient to optimize the placement following the steepest gradient descent in between legalization phases. This technique is implemented in our new placement tool called Liquid. In Liquid each block's position is updated several times following an accelerated gradient simulation. We compare this new technique with conventional analytical placement. The Titan23 designs and the Stratix IV FPGA are used for benchmarking. The net effect is that the runtime can be reduced by 2× on average compared to conventional analytical placement without losing any quality. Elias Vansteenkiste, Seppe Lenders, Dirk Stroobandt |
FPL | 3 |
| 2016 | Runtime-quality tradeoff in partitioning based multithreaded packingabstractIt takes a long time to generate a configuration for an FPGA starting from a description of a digital circuit in a hardware design language. This configuration should have a high quality so that the FPGA resources are used in an efficient way with the maximum clock frequency and minimizing the power consumption. In this work we present two new packing algorithms that obtain better quality and faster runtimes when compared to the frequently used AAPack packer. The partitioning based methodology allows us to exploit the advantage of multithreading on commodity hardware. Firstly we demonstrate the benefits of our fully partitioning based PartSA packer. Existing packers with a partitioning based approach have problems with the cluster size and bandwidth constraint of the functional blocks. We added a fast simulated annealing step after partitioning to solve these problems. A gain of 26% in total wirelength is obtained while reaching up to 2.3× faster packing runtimes for large circuits on a CPU with four cores. Unfortunately the PartSA packer can not be used for architectures without a complete crossbar in the functional blocks. Therefore a second packer is proposed that combines the benefits of partitioning based and seed based packing. MultiPart has up to 4× faster packing runtimes while still having a gain of 20% in total wirelength. Dries Vercruyce, Elias Vansteenkiste, Dirk Stroobandt |
FPL | 3 |
| 2015 | Hamiltonian Path Strategy for Deadlock-Free and Adaptive Routing in Diametrical 2D Mesh NoCsabstractThe overall performance of Network-on-Chip (NoC) is strongly affected by the efficiency of the on-chip routing algorithm. Among the factors associated with the design of a high-performance routing method, adaptivity is an important one. Moreover, deadlock-and live lock-freedom are necessary for a functional routing method. Despite the advantages that the diametrical mesh can bring to NoCs compared with the classical mesh topology, the literature records little research efforts to design pertinent routing methods for such networks. Using the available routing algorithms, the network performance degrades drastically not only due to the deterministic paths, but also to the deadlocks created between the packets. In this paper, we take advantage of the Hamiltonian routing strategy to adaptively route the packets through deadlock-free paths in a diametrical 2D mesh network. The simulation results demonstrate the efficiency of the proposed approach in decreasing the likelihood of congestion and smoothly distributing the traffic across the network. Poona Bahrebar, Dirk Stroobandt |
CCGRID | 2 |
| 2015 | Avoiding transitional effects in dynamic circuit specialisation on FPGAsabstractDynamic Circuit Specialisation (DCS) is a technique that uses the reconfigurability of an FPGA to optimise a circuit during run-time, thus achieving higher performance and lower resource cost. However, run-time reconfiguration causes transitional effects that form an important problem for DCS. Because of these, the DCS circuit cannot be used while it is being reconfigured. This limits the usability of DCS for streaming applications and other applications that cannot tolerate downtime. For other applications, this results in a loss of performance. Karel Heyse, Dirk Stroobandt |
DAC | 2 |
| 2015 | Logic Gates in the routing network of FPGAs (Abstract Only)abstractWe propose a new kind of FPGA architecture with a routing network that not only provides interconnections between the functional blocks but also performs some logic operation. More specifically we replaced the routing multiplexer node in the conventional architecture with an element that can be used as both AND gate and multiplexer. A conventional routing multiplexer node consists of a multiplexer and a two stage buffer. In our new architecture a NAND gate replaces the first inverter stage of the buffer and two multiplexers half the size of the original multiplexer replace the original multiplexer. The aim of this study is to determine if this kind of architecture is feasible and if it is worth to implement pack, placement and routing tools in the future. We developed a new technology-mapping algorithm and sized the transistors in this new architecture to evaluate the area and delay. Preliminary results indicate that the gain in logic depth and area achieved by mapping to not only LUTs but also to AND gates outweighs the overhead of introducing AND gates in the routing network with a net reduction in area-delay product of 5.6. Designs implemented on the proposed architecture would require 11.2 % more area, but they will have a 14 % decreased logic depth and the architecture has a slightly faster representative critical path. These results are preliminary because the pack, place and route routines are not implemented yet. Elias Vansteenkiste, Berg Severens, Dirk Stroobandt |
FPGA | 3 |
| 2015 | Estimating circuit delays in FPGAs after technology mappingabstractAn FPGA implementation requires a significant effort of the hardware designer, who optimizes FPGA designs by going through many time-consuming CAD flow iterations. These iterations provide two types of feedback: (1) the FPGA performance and (2) the identification of the parts having the highest impact on the FPGA performance. Both depend on the wirelength behavior. Studies have been dedicated to the estimation of local [5] and global [4] wirelengths, but to our knowledge both performance estimations and identification of the critical zone are not present in literature. Therefore this paper, firstly, presents a comparison of three performance estimation techniques: logic depth, Monte Carlo simulation and fast placement (ordered from low to high accuracy and runtime). Secondly, four methods identifying the critical zone are compared. Results show that Monte Carlo simulations provide a good identification of the parts having the highest impact on the performance. We conclude that Monte Carlo simulations provide useful feedback within a short runtime (about 30 times faster than placement), reducing the time-to-market of FPGA implementations. Berg Severens, Elias Vansteenkiste, Karel Heyse, Dirk Stroobandt |
FPL | 4 |
| 2015 | TCONMAP: Technology Mapping for Parameterised FPGA ConfigurationsabstractParameterised configurations are FPGA configuration bitstreams in which the bits are defined as functions of user-defined parameters. From a parameterised configuration, it is possible to quickly and efficiently derive specialised, regular configuration bitstreams by evaluating these functions. The specialised bitstreams have different properties and functionality depending on the chosen values of the parameters. The most important application of parameterised configurations is the generation of specialised configuration bitstreams for Dynamic Circuit Specialisation, a technique for optimising circuits at runtime using partial reconfiguration of the FPGA. Generating and using parameterised configurations requires a new FPGA tool flow. In this article, we present a new technology mapping algorithm for parameterised designs, called TCONMAP, that can be used to produce parameterised configurations in which both the configuration of the logic blocks and routing is a function of the parameters. In our experiments, we demonstrate that in using TCONMAP, the depth and area of the mapped circuit is close to the minimal depth and area attainable. Both Dynamic Circuit Specialisation and fine-grained modular reconfiguration are extracted by TCONMAP from the HDL description of the design requiring only simple parameter annotations. Karel Heyse, Brahim Al Farisi, Karel Bruneel, Dirk Stroobandt |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2015 | Identification of Dynamic Circuit Specialization Opportunities in RTL CodeabstractDynamic Circuit Specialization (DCS) optimizes a Field-Programmable Gate Array (FPGA) design by assuming a set of its input signals are constant for a reasonable amount of time, leading to a smaller and faster FPGA circuit. When the signals actually change, a new circuit is loaded into the FPGA through runtime reconfiguration. The signals the design is specialized for are called parameters. For certain designs, parameters can be selected so the DCS implementation is both smaller and faster than the original implementation. However, DCS also introduces an overhead that is difficult for the designer to take into account, making it hard to determine whether a design is improved by DCS or not. This article presents extensive results on a profiling methodology that analyses Register-Transfer Level (RTL) implementations of applications to check if DCS would be beneficial. It proposes to use the functional density as a measure for the area efficiency of an implementation, as this measure contains both the overhead and the gains of a DCS implementation. The first step of the methodology is to analyse the dynamic behaviour of signals in the design, to find good parameter candidates. The overhead of DCS is highly dependent on this dynamic behaviour. A second stage calculates the functional density for each candidate and compares it to the functional density of the original design. The profiling methodology resulted in three implementations of a profiling tool, the DCS-RTL profiler. The execution time, accuracy, and the quality of each implementation is assessed based on data from 10 RTL designs. All designs, except for the two 16-bit adaptable Finite Impulse Response (FIR) filters, are analysed in 1 hour or less. Tom Davidson, Elias Vansteenkiste, Karel Heyse, Karel Bruneel, Dirk Stroobandt |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2015 | On the Impact of Replacing Low-Speed Configuration Buses on FPGAs with the Chip's Internal Configuration InfrastructureabstractIt is common for large hardware designs to have a number of registers or memories whose contents have to be changed very seldom (e.g., only at startup). The conventional way of accessing these memories is through a low-speed memory bus. This bus uses valuable hardware resources, introduces long global connections, and contributes to routing congestion. Hence, it has an impact on the overall design even though it is only rarely used. A Field-Programmable Gate Array (FPGA) already contains a global communication mechanism in the form of its configuration infrastructure. In this article, we evaluate the use of the configuration infrastructure as a replacement for a low-speed memory bus on the Maxeler HPC platform. We find that by removing the conventional low-speed memory bus, the maximum clock frequency of some applications can be improved by 8%. Improvements by 25% and more are also attainable, but constraints of the Xilinx reconfiguration infrastructure prevent fully exploiting these benefits at the moment. We present a number of possible changes to the Xilinx reconfiguration infrastructure and tools that would solve this and make these results more widely applicable. Karel Heyse, Jente Basteleus, Brahim Al Farisi, Dirk Stroobandt, Oliver Kadlcek, Oliver Pell |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2014 | Improving hamiltonian-based routing methods for on-chip networks: A turn model approachabstractThe overall performance of Multi-Processor System-on-Chip (MPSoC) platforms depends highly on the efficient communication among their cores in the Network-on-Chip (NoC). Routing algorithms are responsible for the on-chip communication and traffic distribution through the network. Hence, designing efficient and high-performance routing algorithms is of significant importance. In this paper, a deadlock-free and highly adaptive path-based routing method is proposed without using virtual channels. This method strives to exploit the maximum number of minimal paths between any source and destination pair. The simulation results in terms of performance and power consumption demonstrate that the proposed method significantly outperforms the other adaptive and non-adaptive schemes. This efficiency is achieved by reducing the number of hotspots and smoothly distributing the traffic across the network. Poona Bahrebar, Dirk Stroobandt |
DATE | 2 |
| 2014 | Automating data reuse in High-Level SynthesisabstractCurrent High-Level Synthesis (HLS) tools perform excellently for the synthesis of computation kernels, but they often don't optimize memory bandwidth. As memory access is a bottleneck in many algorithms, the performance of the generated circuit will benefit substantially from memory access optimization. In this paper we present an automated method and a toolchain to detect reuse of array data in loop nests and to build hardware that exploits this data reuse. This saves memory bandwidth and improves circuit performance. We make use of the polyhedral representation of the source program, which makes our method computationally easy. Our software complements the existing HLS flows. Starting from a loop nest written in C, our tool generates a reuse buffer and a loop controller, and preprocesses the loop body for synthesis with an existing HLS tool. Our automated tool produces designs from unoptimized source code that are as efficient as those generated by a commercial HLS tool from manually-optimized source code. Wim Meeus, Dirk Stroobandt |
DATE | 2 |
| 2014 | Reducing the overhead of dynamic partial reconfiguration for multi-mode circuitsabstractA multi-mode circuit implements the functionality of a limited number of circuits, called modes, of which at any given time only one needs to be realised. Using dynamic partial reconfiguration of an FPGA, all the modes can be implemented on the same reconfigurable region, requiring only an area that can contain the biggest mode. This can save considerable chip area. Conventional dynamic partial reconfiguration techniques generate a configuration for every mode separately. As a result, to switch between modes the complete reconfigurable region is rewritten, which often leads to long reconfiguration times. In this paper we give an overview of research we conducted to reduce this overhead of dynamic partial reconfiguration for multi-mode circuits. In this research we explored several joint optimization strategies at different stages of the tool flow. Brahim Al Farisi, Karel Heyse, Dirk Stroobandt |
FPT | 3 |
| 2014 | FPGA-Based Design Using the FASTER Toolchain: The Case of STM Spear Development BoardabstractEven though FPGAs are becoming more and more popular as they are used in many different scenarios like communications and HPC, the steep learning curve needed to work with this technology is still the major limiting factor to their full success. Many works proposed to mitigate this problem by creating a companion of tools to support the designer during the development phase for this technology. The EU FASTER Project aims at realizing an integrated toolchain that assists the designer in the steps of the design flow that are necessary to port a given application onto an FPGA device. The novelty of the framework relies in the fact that the partial dynamic reconfiguration, which FPGA devices can exploit, is seen as a first class citizen throughout the whole design flow. This work reports a case study in which the FASTER toolchain has been used to port a raytracer application onto the STM Spear prototyping embedded platform. The paper discusses the steps done for the realization of the prototype and the results obtained on the target device. It finally reports some improvements that can be exploited to improve the performance of the hardware implementation that has been realized. Fabrizio Spada, Alberto Scolari, Gianluca Durelli, Riccardo Cattaneo, Marco D. Santambrogio, Donatella Sciuto, Dionisios N. Pnevmatikatos, Georgi Gaydadjiev, Oliver Pell, Andreas Brokalakis, Wayne Luk, Dirk Stroobandt, Danilo Pau |
ISPA | 12 |
| 2014 | TPaR: Place and Route Tools for the Dynamic Reconfiguration of the FPGA's Interconnect NetworkabstractDynamic partial reconfiguration of FPGAs enables the dynamic specialization of the circuit for the runtime needs of the application. Previously a tool flow, called the TLUT tool flow, was developed to aid the designer in applying dynamic circuit specialization (DCS) for their designs. The TLUT tool flow generates an implementation in which the lookup tables (LUTs) can be specialized during runtime. In this paper, place and route algorithms are described for the TCON tool flow. The TCON tool flow generates implementations in which not only the logic infrastructure (LUTs) is dynamically specialized, but also the routing infrastructure of the FPGA. Exploiting the reconfigurability of the FPGA interconnection network further improves area (50% to 92% less LUTs and 36% to 81% less wiring), logic depth (a 63% to 80% reduction) and power consumption. To achieve this, major changes were needed, not only in the mapping, but also in the place and route steps. This work describes the altered place and route algorithms, called TPlace and Troute. Elias Vansteenkiste, Brahim Al Farisi, Karel Bruneel, Dirk Stroobandt |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2013 | An automatic tool flow for the combined implementation of multi-mode circuitsabstractA multi-mode circuit implements the functionality of a limited number of circuits, called modes, of which at any given time only one needs to be realised. Using run-time reconfiguration of an FPGA, all the modes can be implemented on the same reconfigurable region, requiring only an area that can contain the biggest mode. Typically, conventional run-time reconfiguration techniques generate a configuration for every mode separately. To switch between modes the complete reconfigurable region is rewritten, which often leads to very long reconfiguration times. In this paper we present a novel, fully automated tool flow that exploits similarities between the modes and uses Dynamic Circuit Specialization to drastically reduce reconfiguration time. Experimental results show that the number of bits that is rewritten in the configuration memory reduces with a factor from 4.6× to 5.1× without significant performance penalties. Brahim Al Farisi, Karel Bruneel, João M. P. Cardoso, Dirk Stroobandt |
DATE | 4 |
| 2013 | Staticroute: A novel router for the Dynamic Partial Reconfiguration of FPGASabstractUsing Dynamic Partial Reconfiguration (DPR) of FPGAs, several circuits can be time-multiplexed on the same chip region, saving considerable area. However, the long reconfiguration time when switching between circuits remains a large problem with DPR. In this paper we show it is possible to significantly reduce reconfiguration time when the number of circuits is limited. We tackle the problem by reducing the time needed to reconfigure the FPGA's routing. We divide the configuration memory of the FPGA's routing in a static and a dynamic portion. A novel router, called StaticRoute, is presented that is able to route the nets of the different circuits in such a way that the static portion is shared and only the dynamic portion needs to be reconfigured. The static portion of the configuration memory does not need to be rewritten during run-time. In the experiments we show it is possible to reach a 2× speed-up of the reconfiguration process, while the increase in wire length per circuit is limited. Brahim Al Farisi, Karel Bruneel, Dirk Stroobandt |
FPL | 3 |
| 2013 | Efficient implementation of Virtual Coarse Grained Reconfigurable Arrays on FPGASabstractFine grained Field Programmable Gate Arrays (FPGA) are complex to program and therefore suffer from high development costs. To solve this problem, Virtual Coarse Grained Reconfigurable Arrays (Virtual CGRA), or CGRAs implemented on FPGAs, have been proposed. Conventional implementations of VCGRAs use functional FPGA resources, such as LookUp Tables, to implement the virtual switch blocks, registers and other components that make the VCGRA configurable. We show that this is a large overhead that can often be avoided by mapping these components directly on lower level FPGA resources such as physical switch blocks and configuration memory. We show how this can be achieved using the tool flow for parameterised FPGA configurations and illustrate the advantages of this method by showing that an area reduction of 50% is attainable for a VCGRA aimed at regular expression matching. Karel Heyse, Tom Davidson, Elias Vansteenkiste, Karel Bruneel, Dirk Stroobandt |
FPL | 5 |
| 2013 | A connection-based router for FPGAsabstractThe FPGA's interconnection network not only requires the larger portion of the total silicon area in comparison to the logic available on the FPGA, it also contributes to the majority of the delay and power consumption. Therefore it is essential that routing algorithms are as efficient as possible. In this work the connection router is introduced. It is capable of partially ripping up and rerouting the routing trees of nets. To achieve this, the main congestion loop rips up and reroutes connections instead of nets, which allows the connection router to converge much faster to a solution. The connection router is compared with the VPR directed search router on the basis of VTR benchmarks on a modern commercial FPGA architecture. It is able to find routing solutions 4.4% faster for a relaxed routing problem and 84.3% faster for hard instances of the routing problem. And given the same amount of time as the VPR directed search, the connection router is able to find routing solutions with 5.8% less tracks per channel. Elias Vansteenkiste, Karel Bruneel, Dirk Stroobandt |
FPT | 3 |
| 2013 | Bidirectional truncated recurrent neural networks for efficient speech denoisingabstractWe propose a bidirectional truncated recurrent neural network architecture for speech denoising. Recent work showed that deep recurrent neural networks perform well at speech denoising tasks and outperform feed forward architectures [1]. However, recurrent neural networks are difficult to train and their simulation does not allow for much parallelization. Given the increasing availability of parallel computing architectures like GPUs this is disadvantageous. The architecture we propose aims to retain the positive properties of recurrent neural networks and deep learning while remaining highly parallelizable. Unlike a standard recurrent neural network, it processes information from both past and future time steps. We evaluate two variants of this architecture on the Aurora2 task for robust ASR where they show promising results. The models outperform the ETSI2 advanced front end and the SPLICE algorithm under matching noise conditions. Philemon Brakel, Dirk Stroobandt, Benjamin Schrauwen |
INTERSPEECH | 2 |
| 2013 | Making Communication a First-Class Citizen in Multicore PartitioningabstractComputation-intensive image processing applications need to be implemented on multicore architectures. If they are to be executed efficiently on such platforms, the underlying data and/or functions should be partitioned and distributed among the processors. The optimal partitioning approach is the one which aims to minimize the inter-processor communication while maximizing the load balance. With the continuously increasing number of cores which exacerbates the demand for more complex memory hierarchies, non-uniform memory access, etc., on-chip communication has gained a significant role in taking advantage of the multicore chips. Therefore, making partitioning decisions just based on conventional performance results and without communication profiling is suboptimal. In this paper, we explore the behavior of a mesh decoder as a case study in terms of communication and computation, and propose models that allow early prediction of the application's behavior. Using these models, profiling the application for all of the input samples is not necessary anymore. As a result, communication- and computation-aware parallelization could be performed faster and easier. Poona Bahrebar, Ruxandra-Marina Florea, Wim Heirman, Leon Denis, Adrian Munteanu 0001, Dirk Stroobandt |
PDP | 6 |
| 2013 | Training energy-based models for time-series imputation
Philemon Brakel, Dirk Stroobandt, Benjamin Schrauwen |
J. Mach. Learn. Res. | 2 |
| 2013 | How to efficiently implement dynamic circuit specialization systemsabstractDynamic circuit specialization (DCS) is a technique used to implement FPGA applications where some of the input data, called parameters, change slowly compared to other inputs. Each time the parameter values change, the FPGA is reconfigured by a configuration that is specialized for those new parameter values. This specialized configuration is much smaller and faster than a regular configuration. However, the overhead associated with the specialization process should be minimized to achieve the desired benefits of using the DCS technique. This overhead is represented by both the FPGA resources needed to specialize the FPGA at runtime and by the specialization time. The introduction of parameterized configurations [Bruneel and Stroobandt 2008] has improved the efficiency of DCS implementations. However, the specialization overhead still takes a considerable amount of resources and time. In this article, we explore how to efficiently build DCS systems by presenting a variety of possible solutions for the specialization process and the overhead associated with each of them. We split the specialization process into two main phases: the evaluation and the configuration phase. The PowerPC embedded processor, the MicroBlaze, and a customized processor (CP) are used as alternatives in the evaluation phase. In the configuration phase, the ICAP and a custom configuration interface (SRL configuration) are used as alternatives. Each solution is used to implement a DCS system for three applications: an adaptive finite impulse response (FIR) filter, a ternary content-addressable memory (TCAM), and a regular expression matcher (RegEx). The experiments show that the use of our CP along with the SRL configuration achieves minimum overhead in terms of resources and time. Our CP is 1.8 and 3.5 times smaller than the PowerPC and the area-optimized implementation of the MicroBlaze, respectively. Moreover, the use of the CP enables a more compact representation for the parameterized configuration in comparison to both the PowerPC and the MicroBlaze processors. For instance, in the FIR, the parameterized configuration compiled for our CP is 6--7 times smaller than that for the embedded processors. Fatma Abouelella, Tom Davidson, Wim Meeus, Karel Bruneel, Dirk Stroobandt |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2012 | FASTER: Facilitating Analysis and Synthesis Technologies for Effective ReconfigurationabstractThe FASTER project aims to ease the definition, implementation and use of dynamically changing hardware systems. Our motivation stems from the promise reconfigurable systems hold for achieving better performance and extending product functionality and lifetime via the addition of new features that work at hardware speed. This is a clear advantage over the more straightforward software component adaptivity. However, designing a changing hardware system is both challenging and time consuming. The FASTER project will facilitate the use of reconfigurable technology by providing a complete methodology that enables designers to easily specify, analyse, implement and verify applications on platforms with general-purpose processors and acceleration modules implemented in the latest reconfigurable technology. To better adapt to different application requirements, the tool-chain will support both region-based and micro-reconfiguration and provide a flexible run-time system that will efficiently manage the reconfigurable resources. We will use applications from the embedded, high performance computing, and desktop domains to demonstrate the potential benefits of the FASTER tools on metrics such as performance, power consumption and total ownership cost. Dionisios N. Pnevmatikatos, Tobias Becker, Andreas Brokalakis, Karel Bruneel, Georgi Gaydadjiev, Wayne Luk, Kyprianos Papademetriou, Ioannis Papaefstathiou, Oliver Pell, Christian Pilato, M. Robart, Marco D. Santambrogio, Donatella Sciuto, Dirk Stroobandt, Tim Todman |
DSD | 14 |
| 2012 | Automatically exploiting regularity in applications to reduce reconfiguration memory requirementsabstractPartial reconfiguration (PR) of FPGAs is a very promising technique. Applications implemented with PR are smaller and faster than applications that are not reconfigured. However, the overhead emerging from the reconfiguration process can nullify the benefits of PR. Moreover, the lack of automatic tools hinders the widespread use of the PR technique. In previous work, the PR barriers have been tackled by introducing parameterized configurations and a tool flow that exploits these configurations. For regularly structured applications mapped through this tool flow, the memory resources needed to store the parameterized configuration can be significantly reduced when regularity is exploited. In this paper, we propose a front-end to the tool flow that automatically detects regular structures at the HDL level and transfers those regularities into the reconfiguration process. The results show that a reduction factor of 76, 10 and 167 is achieved in the memory resources needed to store the parameterized configuration when the regularity is exploited for an adaptive FIR, a regular expression matcher and a Ternary Content Addressable Memory (TCAM) respectively. The reduction factor will be further increased when applications scale. Fatma Abouelella, Karel Bruneel, Dirk Stroobandt |
FPL | 3 |
| 2012 | Mapping logic to reconfigurable FPGA routingabstractParameterised configurations for FPGAs are configuration bitstreams of which part of the bits are defined as Boolean functions of parameters. By evaluating these Boolean functions using different parameter values, it is possible to quickly and efficiently derive specialised configuration bitstreams with different properties. An important application of parameterised configurations is the generation of specialised configuration bitstreams for Dynamic Circuit Specialisation. Generating and using parameterised configurations requires a new FPGA tool flow. In this paper we present an algorithm for technology mapping of parameterised designs that can exploit the reconfigurability of the logic blocks and routing of the FPGA. This algorithm, called TCONMAP, is based on “Cut enumeration, cut ranking, node selection”. As part of it, a new method to calculate the feasibility of cuts based on the Binary Decision Diagrams (BDD) of their local function is proposed. Karel Heyse, Karel Bruneel, Dirk Stroobandt |
FPL | 3 |
| 2012 | Maximizing the reuse of routing resources in a reconfiguration-aware connection routerabstractParameterised configurations for FPGAs are configuration bitstreams of which some of the bits are defined as Boolean functions of parameters. By evaluating these Boolean functions using different parameter values, it is possible to quickly and efficiently derive specialised configuration bitstreams with different properties. Generating and using parameterized configurations requires a new tool flow. In this paper we propose a novel algorithm for the routing step of this tool flow. This new router, called the connection bundle router, is able to route a circuit with parameterized interconnections. It produces routing solutions in less time (up to a factor 5,2) and with a better quality in terms of number of wires (up to 38%) and minimum track width (up to 25%) than its predecessors. The connection bundle router is fully automated and uses a scalable connection-based representation for the parameterized interconnections in a tunable circuit. Elias Vansteenkiste, Karel Bruneel, Dirk Stroobandt |
FPL | 3 |
| 2011 | Memory-Efficient and Fast Run-Time Reconfiguration of Regularly Structured DesignsabstractPrevious work has shown that run-time reconfiguration of FPGAs benefits greatly from the use of Tunable LUT (TLUT) circuits. These can be rapidly transformed into a specialized LUT circuit and are also very memory efficient when representing regularly structured designs, where the same hardware module is instantiated many times. However, the memory requirements and reconfiguration time of a run-time reconfigurable application are also dependent on the reconfiguration mechanism. In this paper, we will show that the memory requirements of conventional ICAP reconfiguration grow very fast with the number of modules, resulting in excessive memory usage. We propose to use Shift-Register-LUT (SRL) reconfiguration which is faster and results in a memory usage that is independent of the number of modules. Brahim Al Farisi, Karel Heyse, Karel Bruneel, Dirk Stroobandt |
FPL | 4 |
| 2011 | Automatic detection of epileptic seizures on the intra-cranial electroencephalogram of rats using reservoir computing
Pieter Buteneers, David Verstraeten, Pieter van Mierlo, Tine Wyckhuys, Dirk Stroobandt, Robrecht Raedt, Hans Hallez, Benjamin Schrauwen |
Artif. Intell. Medicine | 5 |
| 2011 | Dynamic data folding with parameterizable FPGA configurationsabstractIn many applications, subsequent data manipulations differ only in a small set of parameter values. Because of their reconfigurability, FPGAs (field programmable gate arrays) can be configured with a specialized circuit each time the parameter values change. This technique is called dynamic data folding. The specialized circuits are smaller and faster than their generic counterparts. However, the overhead involved in generating the configurations for the specialized circuits at runtime is very large when conventional tools are used, and this overhead will in many cases negate the benefit of using optimized configurations. This article introduces an automatic method for generating runtime parameterizable configurations from arbitrary Boolean circuits. These configurations, in which some of the configuration bits are expressed as a closed-form Boolean expression of a set of parameters, enable very fast run-time specialization, since specialization only involves evaluating these expressions. Our approach is validated on a ternary content-addressable memory (TCAM). We show that the specialized configurations, produced by our method use 2.82 times fewer LUTs than the generic configuration, and even 1.41 times fewer LUTs than the implementation generated by Xilinx Coregen. Moreover, while Coregen needs hand-crafted generators for each type of circuit, our toolflow can be applied to any VHDL design. Using our automatic and generally applicable method, run-time hardware optimization suddenly becomes feasible for a large class of applications. Karel Bruneel, Wim Heirman, Dirk Stroobandt |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2010 | A Parallel for Loop Memory Template for a High Level Synthesis CompilerabstractWe propose a parametrized memory template for applications with parallel for loops. The template's parameters reflect important trade-offs made during system design. The template is incorporated in our high level synthesis (HLS) compiler, where the template's parameters are adjusted to the application. The template fits parallel for loops with no loop dependencies and sequential bodies. We found two alternative template implementations using our compiler. In the future, we will develop templates for other types of for loops. These will be added to the compiler and it will identify the template that works best for the application it is compiling. Once a template is selected, the compiler will use design space exploration to select the best combination of template parameters for the targeted hardware and application. Craig Moore, Wim Meeus, Harald Devos, Dirk Stroobandt |
DSD | 4 |
| 2010 | Automatic tool flow for shift-register-LUT reconfiguration: making run-time reconfiguration fast and easy (abstract only)abstractThe Shift-Register-Lut (SRL) functionality is a powerful extension of Xilinx FPGA architectures and has been used successfully in many applications. If routing is kept fixed, these SRLs can also be used for run-time reconfiguration. So far, this technique has mainly been used to reconfigure specialized functions. In contrast, we propose a generic tool flow that uses SRLs for fast run-time reconfiguration of general data folding applications. We show that, in such an automatic toolflow, SRL reconfiguration is over two orders of magnitude faster than run-time reconfiguration using the ICAP. It thus makes run-time reconfiguration viable for applications with a more dynamic behaviour. Our generic tool flow is also very easy to use since the designer only has to annotate slowly varying signals in an RTL HDL description, while the tool flow takes care of all the rest. Brahim Al Farisi, Karel Bruneel, Harald Devos, Dirk Stroobandt |
FPGA | 4 |
| 2010 | Efficiently Generating FPGA Configurations through a Stack MachineabstractParameterizable configurations are regular FPGA configurations in which some of the configuration bits are expressed as Boolean functions of a set of parameters. These configurations can be rapidly transformed to a specialized configuration by evaluating the Boolean functions for a specific set of parameter values and are therefore ideal for use in run-time reconfigurable systems. To accommodate the use of parameterizable configurations in commercial FPGAs, the concept of the Parameterizable Bitstream was introduced. In this paper, we introduce a hardware implementation that evaluates the Parameterizable Bitstream based on a stack machine architecture. We enabled fast generation of specialized configurations with a significant reduction in resources (80%) in comparison to the MicroBlaze soft processor when it is used as a configuration generation engine. Fatma Abouelella, Karel Bruneel, Dirk Stroobandt |
FPL | 3 |
| 2010 | PinComm: Characterizing Intra-application Communication for the Many-Core EraabstractAs the number of cores in both embedded Multi-Processor Systems-on-Chip and general purpose processors keeps rising, on-chip communication becomes more and more important. In order to write efficient programs for these architectures it is therefore necessary to have a good idea of the communication behavior of an application. We present a communication profiler that extracts this behavior from compiled, sequential or parallel C/C++ programs, and constructs a dynamic data-flow graph at the level of major functional blocks. In contrast to existing methods of measuring inter-program communication, our tool automatically generates the program's data-flow graph and is less demanding for the developer. It can also be used to view differences between program phases (such as different video frames), which allows both input- and phase-specific optimizations to be made. We will also describe briefly how this information can subsequently be used to guide the effort of parallelizing the application, to co-design the software, memory hierarchy and communication hardware, and to provide new sources of communication-related runtime optimizations. Wim Heirman, Dirk Stroobandt, Narasinga Rao Miniskar, Roel Wuyts, Francky Catthoor |
ICPADS | 2 |
| 2009 | Automatically mapping applications to a self-reconfiguring platformabstractThe inherent reconfigurability of FPGAs enables us to optimize an FPGA implementation in different time intervals by generating new optimized FPGA configurations and reconfiguring the FPGA at the interval boundaries. With conventional methods, generating a configuration at run-time requires an unacceptable amount of resources. In this paper, we describe a tool flow that can automatically map a large set of applications to a self-reconfiguring platform, without an excessive need for resources at run-time. The self-reconfiguring platform is implemented on a Xilinx Virtex-II Pro FPGA and uses the FPGA's PowerPC as configuration manager. This configuration manager generates optimized configurations on-the-fly and writes them to the configuration memory using the ICAP. We successfully used our approach to implement an adaptive 32-tap FIR filter on a Xilinx XUP board. This resulted in a 40% reduction in FPGA resources compared to a conventional implementation and a manageable reconfiguration overhead. Karel Bruneel, Fatma Abouelella, Dirk Stroobandt |
DATE | 3 |
| 2009 | Pruning and regularization in reservoir computing
Xavier Dutoit, Benjamin Schrauwen, Jan M. Van Campenhout, Dirk Stroobandt, Hendrik Van Brussel, Marnix Nuttin |
Neurocomputing | 4 |
| 2009 | Accelerating Event-Driven Simulation of Spiking Neurons with Multiple Synaptic Time ConstantsabstractThe simulation of spiking neural networks (SNNs) is known to be a very time-consuming task. This limits the size of SNN that can be simulated in reasonable time or forces users to overly limit the complexity of the neuron models. This is one of the driving forces behind much of the recent research on event-driven simulation strategies. Although event-driven simulation allows precise and efficient simulation of certain spiking neuron models, it is not straightforward to generalize the technique to more complex neuron models, mostly because the firing time of these neuron models is computationally expensive to evaluate. Most solutions proposed in literature concentrate on algorithms that can solve this problem efficiently. However, these solutions do not scale well when more state variables are involved in the neuron model, which is, for example, the case when multiple synaptic time constants for each neuron are used. In this letter, we show that an exact prediction of the firing time is not required in order to guarantee exact simulation results. Several techniques are presented that try to do the least possible amount of work to predict the firing times. We propose an elegant algorithm for the simulation of leaky integrate-and-fire (LIF) neurons with an arbitrary number of (unconstrained) synaptic time constants, which is able to combine these algorithmic techniques efficiently, resulting in very high simulation speed. Moreover, our algorithm is highly independent of the complexity (i.e., number of synaptic time constants) of the underlying neuron model. Michiel D'Haene, Benjamin Schrauwen, Jan M. Van Campenhout, Dirk Stroobandt |
Neural Comput. | 4 |
| 2009 | Efficient memory management for hardware accelerated Java Virtual MachinesabstractApplication-specific hardware accelerators can significantly improve a system's performance. In a Java-based system, we then have to consider a hybrid architecture that consists of a Java Virtual Machine running on a general-purpose processor connected to the hardware accelerator. In such a hybrid architecture, data communication between the accelerator and the general-purpose processor can incur a significant cost, which may even annihilate the original performance improvement of adding the accelerator. A careful layout of the data in the memory structure is therefore of major importance to maintain the acceleration performance benefits. This article addresses the reduction of the communication cost in a distributed shared memory consisting of the main memory of the processor and the accelerator's local memory, which are unified in the Java heap. Since memory access times are highly nonuniform, a suitable allocation of objects in either main memory or the accelerator's local memory can significantly reduce the communication cost. We propose several techniques for finding the optimal location for each Java object's data, either statically through profiling or dynamically at runtime. We show how we can reduce communication cost by up to 86% for the SPECjvm and DaCapo benchmarks. We also show that the best strategy is application dependent and also depends on the relative cost of remote versus local accesses. For a relative cost higher than 10, a self-learning dynamic approach often results in the best performance. Peter Bertels, Wim Heirman, Erik H. D'Hollander, Dirk Stroobandt |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2008 | Java and the Power of Multi-Core ProcessingabstractThe new era of multi-core processing challenges software designers to efficiently exploit the parallelism that is now massively available. Programmers have to exchange the conventional sequential programming paradigm for parallel programming: single-threaded designs must be decomposed into dependent, interacting tasks. The Java programming language has built-in thread support and is therefore suitable for the development of parallel software, but programming multi-threaded applications is a tedious task. Therefore we are working on a framework and tool support to alleviate the burden of threads, synchronisation and locking, based on process networks. This paper describes our initial ideas for this new programming model. Peter Bertels, Dirk Stroobandt |
CISIS | 2 |
| 2008 | Pruning and Regularisation in Reservoir Computing: a First Insight
Xavier Dutoit, Benjamin Schrauwen, Jan M. Van Campenhout, Dirk Stroobandt, Hendrik Van Brussel, Marnix Nuttin |
ESANN | 4 |
| 2008 | Automatic generation of run-time parameterizable configurationsabstractIn many applications, subsequent data manipulations differ only in a small set of parameter values. Because of their reconfigurability, FPGAs (field programmable gate arrays) can be configured with an optimized configuration every time the parameter values change. These optimized configurations are smaller and faster than their generic counterparts. However, the overhead involved in generating the configurations at run-time with conventional tools is very large. This paper introduces an automatic method for generating runtime parameterizable configurations from arbitrary Boolean circuits. These configurations in which some of the configuration bits are expressed as a function of a set of parameters enable very fast run-time specialization since specialization only involves evaluating these functions. Our approach is validated on adaptive filtering. We show that the specialized filter configurations produced by our method are 2.3 times smaller and 36% faster than a generic filter configuration and that these configurations can be generated in on average 166 mus. Being a generic method, run-time hardware optimization suddenly becomes feasible for a large class of applications. Karel Bruneel, Dirk Stroobandt |
FPL | 2 |
| 2008 | Stable Output Feedback in Reservoir Computing Using Ridge Regression
Francis Wyffels, Benjamin Schrauwen, Dirk Stroobandt |
ICANN (1) | 3 |
| 2008 | Real-Time Epileptic Seizure Detection on Intra-cranial Rat Data Using Reservoir Computing
Pieter Buteneers, Benjamin Schrauwen, David Verstraeten, Dirk Stroobandt |
ICONIP (1) | 4 |
| 2008 | Mobile robot control in the road sign problem using Reservoir Computing networksabstractIn this work we tackle the road sign problem with reservoir computing (RC) networks. The T-maze task (a particular form of the road sign problem) consists of a robot in a T-shaped environment that must reach the correct goal (left or right arm of the T-maze) depending on a previously received input sign. It is a control task in which the delay period between the sign received and the required response (e.g., turn right or left) is a crucial factor. Delayed response tasks like this one form a temporal problem that can be handled very well by RC networks. Reservoir computing is a biologically plausible technique which overcomes the problems of previous algorithms such as backpropagation through time - which exhibits slow (or non-) convergence on training. RC is a new concept that includes a fast and efficient training algorithm. We show that this simple approach can solve the T-maze task efficiently. Eric A. Antonelo, Benjamin Schrauwen, Dirk Stroobandt |
ICRA | 3 |
| 2008 | Band-pass Reservoir ComputingabstractMany applications of Reservoir Computing (and other signal processing techniques) have to deal with information processing of signals with multiple time-scales. Classical Reservoir Computing approaches can only cope with multiple frequencies to a limited degree. In this work we investigate reservoirs build of band-pass filter neurons which can be made sensitive to a specified frequency band. We demonstrate that many currently difficult tasks for reservoirs can be handled much better by a band-pass filter reservoir. Francis Wyffels, Benjamin Schrauwen, David Verstraeten, Dirk Stroobandt |
IJCNN | 4 |
| 2008 | Modeling multiple autonomous robot behaviors and behavior switching with a single reservoir computing networkabstractReservoir computing (RC) uses a randomly created Recurrent Neural Network as a reservoir of rich dynamics which projects the input to a high dimensional space. These projections are mapped to the desired output using a linear output layer, which is the only part being trained by standard linear regression. In this work, RC is used for imitation learning of multiple behaviors which are generated by different controllers using an intelligent navigation system for mobile robots previously published in literature. Target seeking and exploration behaviors are conflicting behaviors which are modeled with a single RC network. The switching between the learned behaviors is implemented by an extra input which is able to change the dynamics of the reservoir, and in this way, change the behavior of the system. Experiments show the capabilities of Reservoir Computing for modeling multiple behaviors and behavior switching. Eric A. Antonelo, Benjamin Schrauwen, Dirk Stroobandt |
SMC | 3 |
| 2008 | Improving reservoirs using intrinsic plasticity
Benjamin Schrauwen, Marion Wardermann, David Verstraeten, Jochen J. Steil, Dirk Stroobandt |
Neurocomputing | 5 |
| 2008 | Event detection and localization for small mobile robots using reservoir computing
Eric A. Antonelo, Benjamin Schrauwen, Dirk Stroobandt |
Neural Networks | 3 |
| 2007 | Adapting reservoir states to get Gaussian distributions
David Verstraeten, Benjamin Schrauwen, Dirk Stroobandt |
ESANN | 3 |
| 2007 | A Method for Fast Hardware Specialization at run-timeabstractDynamic hardware generation is a powerful technique that can substantially reduce both the required hardware resources and the time needed to perform a calculation, reflected in an improved functional density. This performance improvement is a result of additional run-time optimizations enabled by the knowledge of values at certain inputs at runtime. However, due to the large overhead conventional hardware generation tools incur, the usability of dynamic hardware generation is limited. We present a dual approach that combines compile-time generation of generic hardware and run-time specialization. This drastically decreases the dynamic generation overhead. Our approach is used for dynamic generation of FIR filters and compared to a static and a conventional dynamic implementation. The experiments clearly show that the dual approach improves the usability of dynamic hardware generation. Karel Bruneel, Peter Bertels, Dirk Stroobandt |
FPL | 3 |
| 2007 | Improving External Memory Access for Avalon Systems on Programmable ChipsabstractIn this paper we present a new hardware design pattern for improving memory transfers to external dynamic memory in Altera's SOPC-builder tool by reusing the standard DMA IP core for all bulk memory transfers without the need for a CPU. The presented approach doubles the data throughput without the need for extra system resources. In addition it is more effective for choosing optimal clock settings for the different components of the system on a programmable chip. The benefits and limitations of this new approach are illustrated with a real world example: a bitplane assembler for scalable wavelet based video. The new design is 2.3 times faster with the same clock settings as the original design and uses about 100 logic elements less. Applying our new approach also has a positive impact on energy consumption. Hendrik Eeckhaut, Mark Christiaens, Dirk Stroobandt |
FPL | 3 |
| 2007 | Event Detection and Localization in Mobile Robot Navigation Using Reservoir Computing
Eric A. Antonelo, Benjamin Schrauwen, Xavier Dutoit, Dirk Stroobandt, Marnix Nuttin |
ICANN (2) | 4 |
| 2007 | Mobility of Data in Distributed Hybrid Computing SystemsabstractIn distributed hybrid computing systems, traditional sequential processors are loosely coupled with reconfigurable hardware for optimal performance. This loose coupling proves to be a communication challenge; the processor units cannot efficiently share a physical memory. This paper proposes a distributed shared memory architecture and a method for effective data migration within that shared memory. Data is moved using a novel garbage collection scheme, the dual semispace collector. The new garbage collector and the distributed memory prove to be an effective means of data migration in distributed hybrid computing systems. Philippe Faes, Mark Christiaens, Dirk Stroobandt |
IPDPS | 3 |
| 2007 | Predicting reconfigurable interconnect performance in distributed shared-memory systems
Wim Heirman, Joni Dambre, Iñigo Artundo, Christof Debaes, Hugo Thienpont, Dirk Stroobandt, Jan M. Van Campenhout |
Integr. | 6 |
| 2007 | Special issue on System-Level Interconnect Prediction
Igor L. Markov, Louis K. Scheffer, Dirk Stroobandt |
Integr. | 3 |
| 2007 | An experimental unification of reservoir computing methods
David Verstraeten, Benjamin Schrauwen, Michiel D'Haene, Dirk Stroobandt |
Neural Networks | 4 |
| 2007 | Scalable, Wavelet-Based Video: From Server to Hardware-Accelerated ClientabstractVideo source, carrier and client diversification have led the video coding community to develop scalable video codecs supporting efficient decoding at varying resolution, frame rate and quality. Scalable video has several advantages over a nonscalable approach, but a large scale deployment is far from trivial and a lot of open questions remain. To resolve these, we developed a complete video delivery chain for scalable wavelet-based video. This includes a video server, a negotiation framework, a video scaling infrastructure and two scalable video clients, one pure software client and one real-time, hardware accelerated client. This paper describes the complete chain and identifies and quantifies the impact of using scalable video in every link of this chain. Hendrik Eeckhaut, Harald Devos, Peter Lambert, Davy De Schrijver, Wim Van Lancker, Vincent Nollet, Prabhat Avasare, Tom Clerckx, Fabio Verdicchio, Mark Christiaens, Peter Schelkens, Rik Van de Walle, Dirk Stroobandt |
IEEE Trans. Multim. | 13 |
| 2007 | Systematic Simulation-Based Predictive Synthesis of Integrated Optical InterconnectabstractIntegrated optical interconnect has been identified by the ITRS as a potential solution to overcome predicted interconnect limitations in future systems-on-chip. However, the multiphysics nature of the design problem and the lack of a mature integrated photonic technology have contributed to severe difficulties in assessing its suitability. This paper describes a systematic, fully automated synthesis method for integrated microsource-based optical interconnect capable of optimally sizing the interface circuits based on system specifications, CMOS technology data, and optical device characteristics. The simulation-based nature of the design method means that its results are relatively accurate, even though the generation of each data point requires only 5 min on a 1.3-GHz processor. This method has been used to extract typical performance metrics (delay, power, interconnect density) for optical interconnect of length 2.5-20 mm in three predictive technologies at 65-, 45-, and 32-nm gate length. Ian O'Connor, Faress Tissafi-Drissi, Frédéric Gaffiot, Joni Dambre, Michiel De Wilde, Jan M. Van Campenhout, Dries Van Thourhout, Dirk Stroobandt |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2006 | Optimizing the critical loop in the H.264/AVC CABAC decoderabstractThis paper presents an innovative hardware implementation of the H.264/AVC CABAC binary arithmetic decoder and context modeler capable of decoding one symbol per clock cycle at high clock frequencies while maintaining a slim hardware footprint. This was achieved by substantially decreasing the latency of the central feedback loop through extensive use of speculative prefetching and aggressive pipelining. Actual synthesis results targeted at the state-of-the-art FPGA families show that our approach results in a fast and compact IP core, ideal for a SoC H.264/AVC implementation Hendrik Eeckhaut, Mark Christiaens, Dirk Stroobandt, Vincent Nollet |
FPT | 3 |
| 2006 | Accelerating Event Based Simulation for Multi-synapse Spiking Neural Networks
Michiel D'Haene, Benjamin Schrauwen, Dirk Stroobandt |
ICANN (1) | 3 |
| 2006 | Reservoir-based techniques for speech recognitionabstractA solution for the slow convergence of most learning rules for Recurrent Neural Networks (RNN) has been proposed under the terms Liquid State Machines (LSM) and Echo State Networks (ESN). These methods use a RNN as a reservoir that is not trained. For this article we build upon previous work, where we used reservoir-based techniques to solve the task of isolated digit recognition. We present a straightforward improvement of our previous LSM-based implementation that results in an outperformance of a state-of-the-art Hidden Markov Model (HMM) based recognizer. Also, we apply the Echo State approach to the problem, which allows us to investigate the impact of several interconnection parameters on the performance of our speech recognizer. David Verstraeten, Benjamin Schrauwen, Dirk Stroobandt |
IJCNN | 3 |
| 2005 | A Hardware-Friendly Wavelet Entropy Codec for Scalable VideoabstractA scalable video codec provides the ability to produce a smaller video stream with reduced frame rate, resolution or image quality starting from the original encoded video stream with almost no additional computation. This is important for portable devices that have different quality of service (QoS) requirements and power restrictions. Conventional video codecs do not possess this property; reduced quality is obtained through the arduous process of decoding the encoded video stream and recoding it at a lower quality. Producing such a smaller stream has therefore a very high computational cost. In this article, we present the results of our investigation into the hardware implementation of such a scalable video codec. In particular, we found that the implementation of the entropy codec is a significant bottleneck. We present an alternative, hardware friendly algorithm for entropy coding with superior data locality (both temporal and spatial), with a smaller memory footprint and superior compression while maintaining all required scalability properties. Hendrik Eeckhaut, Harald Devos, Benjamin Schrauwen, Mark Christiaens, Dirk Stroobandt |
DATE | 5 |
| 2005 | Isolated word recognition using a Liquid State Machine
David Verstraeten, Benjamin Schrauwen, Dirk Stroobandt |
ESANN | 3 |
| 2005 | FPGA-Aware Garbage Collection in JavaabstractDuring codesign of a system, one still runs into the impedance mismatch between the software and hardware worlds. This paper identifies the different levels of abstraction of hardware and software as a major culprit of this mismatch. For example, when programming in high-level object-oriented languages like Java, one has disposal of objects, methods, memory management, that facilitates development but these have to be largely abandoned when moving the same functionality into hardware. As a solution, this paper presents a virtual machine, based on the Jikes Research Virtual Machine, that is able to bridge the gap by providing the same capabilities to hardware components as to software components. This seamless integration is achieved by introducing an architecture and protocol that allow reconfigurable hardware and software to communicate with each other in a transparent manner i.e. no component of the design needs to be aware whether other components are implemented in hardware or in software. Further, in this paper we present a novel technique that allows reconfigurable hardware to manage dynamically allocated memory. This is achieved by allowing the hardware to hold references to objects and by modifying the garbage collector of the virtual machine to be aware of these references in hardware. We present benchmark results that show, for four different, well-known garbage collectors and for a wide range of applications, that a hardware-aware garbage collector results in a marginal overhead and is therefore a worthwhile addition to the developer's toolbox. Philippe Faes, Mark Christiaens, Dries Buytaert, Dirk Stroobandt |
FPL | 4 |
| 2005 | Isolated word recognition with the Liquid State Machine: a case study
David Verstraeten, Benjamin Schrauwen, Dirk Stroobandt, Jan M. Van Campenhout |
Inf. Process. Lett. | 3 |
| 2004 | Toward the accurate prediction of placement wire length distributions in VLSI circuitsabstractSince its introduction, Donath's technique for predicting placement wire length distributions has become one of the most popular techniques for a priori wire length estimation. However, in its original form, it was heavily constrained by the underlying circuit and architecture models. In this paper, we show how a careful relaxation of those constraints results in very high correlations between predicted and experimentally measured average wire lengths as well as, a much improved accuracy in predicting wire length distributions. Because the availability of the Rent characteristic is crucial for the quality of our model, we investigate how the prediction quality degrades when only an estimated characteristic is available. Such an estimation can be required to save computation time or when the complete netlist is not yet available (partial use of typical values). It turns out that a fitted /spl beta/-model, based on only a few partitioning levels, can still result in a relatively high-prediction quality. In particular, with respect to the wire length distribution, the results are considerably better than when Rent's Rule is used. Joni Dambre, Dirk Stroobandt, Jan M. Van Campenhout |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Hardware implementation of an EAN-13 bar code decoderabstractThis paper presents an FPGA hardware implementation of a bar code decoder using the EAN-13 standard. The design was implemented on an FPGA situated on an ATMEL FPLSIC (field programmable system level integrated circuit, AT94K40-25DQC). The design has comparable performance and is in many ways more robust than commercially available devices. However, to the authors' knowledge, these all use a microprocessor while our design is purely dedicated hardware. Jeroen De Maeyer, Harald Devos, Wim Meeus, Peter Verplaetse, Dirk Stroobandt |
ASP-DAC | 5 |
| 2003 | Improved a priori interconnect predictions and technology extrapolation in the GTX systemabstractA priori interconnect prediction and technology extrapolation are closely intertwined. Interconnect predictions are at the core of technology extrapolation models of achievable system power, area density, and speed. Technology extrapolation, in turn, informs a priori interconnect prediction via models of interconnect technology and interconnect optimizations. In this paper, we address the linkage between a priori interconnect prediction and technology extrapolation in two ways. First, we describe how rapid changes in technology, as well as rapid evolution of prediction methods, require a dynamic and flexible framework for technology extrapolation. We then develop a new tool, the GSRC technology extrapolation system (GTX), which allows capture of such knowledge and rapid development of new studies. Second, we identify several "nontraditional" facets of interconnect prediction and quantify their impact on key technology extrapolations. In particular, we explore the effects of interconnect design optimizations such as shield insertion, repeater sizing and repeater staggering, as well as modeling choices for RLC interconnects. Yu Cao 0001, Chenming Hu, Xuejue Huang, Andrew B. Kahng, Igor L. Markov, Michael Oliver, Dirk Stroobandt, Dennis Sylvester |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2003 | A comparison of various terminal-gate relationships for interconnect prediction in VLSI circuitsabstractOver the years, different interpretations of Rent's rule and different ways of estimating the Rent parameters have emerged. In general, these parameters are extracted from the average terminal-gate relationship for a set of circuit modules. We show that this relationship (the Rent characteristic) strongly depends on the definition of the circuit modules. These can be generated in many different ways, either from the topology of the circuit graph or, in a geometric way, by cutting regions from a circuit layout. The resulting Rent parameters can be quite far apart. This paper discusses the fundamental differences between the topological and the two geometric interpretations of the Rent characteristic that are expected to be most appropriate for current wirelength estimation techniques. Our discussion is based on experimental data, as well as on a theoretical model that can be used to estimate certain geometric Rent characteristics from the topological Rent parameters. Using this model, we derive a theoretical lower limit to the value of the average geometric Rent exponent. We also study the impact of the placement approach and placement quality on the geometric Rent characteristics. Joni Dambre, Peter Verplaetse, Dirk Stroobandt, Jan M. Van Campenhout |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2003 | A priori wire length distribution models with multiterminal netsabstractInterconnections are quickly becoming a dominant factor in the design of computer chips. Techniques to estimate interconnection lengths a priori (very early in the design flow) therefore gain attention and will become important for making the right design decisions when one still has the freedom to do so. However, at that time, one also knows least about the possible results of subsequent design steps. Conventional models for a priori estimation of wire lengths in computer chips use Rent's rule to estimate the number of terminals needed for communication between sets of gates. The number of interconnections then follows by taking into account that most nets are point-to-point connections. In this paper, we apply our previously introduced model for multiterminal nets to show that such nets have a fundamentally different influence on the wire length estimations than point-to-point nets. We then estimate the wire length distribution of Steiner tree lengths for applications related to routing resource estimation. Experiments show that the new estimated Steiner-length distributions capture the multiterminal effects much better than the previous point-to-point length distributions. The accuracy of the estimated values is still too low, as for the conventional point-to-point models, because we are still lacking a good model for placement optimization. However, the new results are a step closer to the application of wire length estimation techniques in real-world situations. Dirk Stroobandt |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2002 | AQUASUN: adaptive window query processing in CAD applications for physical design and verificationabstractCAD applications for physical design and verification very often require enumerating all layout objects whose bounding box intersects an axis-aligned rectangular area. A number of multidimensional access methods exist to process such window queries. The performance of some important design and verification algorithms heavily depends on the processing speed of the used access method. For complex layouts, these methods require huge amounts of resident memory to attain this speed.In this paper, we present a new access method called AQUASUN, which brings a significant query processing performance improvement over other adaptive methods---methods which can cope with a continuously changing layout. These methods generally descend from the database world and are designed to perform the equivalent query in n-dimensional space. Our method is specifically tailored to two dimensions, exploiting 2D optimisations that significantly accelerate window queries within oblong objects like PCB tracks. Furthermore, AQUASUN makes use of an efficient compression technique which greatly cuts down on memory usage. Michiel De Wilde, Dirk Stroobandt, Jan M. Van Campenhout |
ACM Great Lakes Symposium on VLSI | 2 |
| 2002 | Toward better wireload models in the presence of obstaclesabstractWirelength estimation techniques typically contain a site density function that enumerates all possible path sites for each wirelength in an architecture and an occupation probability function that assigns a probability to each of these paths to be occupied by a wire. In this paper, we apply a generating polynomial technique to derive complete expressions for site density functions which take effects of layout region aspect ratio and the presence of obstacles into account. The effect of an obstacle is separated into two parts: the terminal redistribution effect and the blockage effect. The layout region aspect ratio and the obstacle area are observed to have a much larger effect on the wirelength distribution than the obstacle's aspect ratio and location. Accordingly, we suggest that these two parameters be included as indices of lookup tables in wireload models. Our results apply to a priori wirelength estimation schemes in chip planning tools to improve parasitic estimation accuracy and timing closure; this is particularly relevant for system-on-chip designs where IP blocks are combined with row-based layout. Chung-Kuan Cheng, Andrew B. Kahng, Bao Liu 0001, Dirk Stroobandt |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2002 | Guest editorial - system-level interconnect prediction
Dirk Stroobandt |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2001 | Toward better wireload models in the presence of obstaclesabstractEfficient and accurate interconnect estimation is crucial to design convergence. With System-on-Chip design, IP blocks form routing obstacles that cannot be accounted for by existing a priori wirelength estimations. In this paper, we identify two distinct effects of obstacles on interconnection length: (i) changes due to the redistribution of interconnect terminals and (ii) detours that have to be made around the obstacles. Theoretical expressions of both effects for point-to-point nets with a single obstacle are derived and compared to experimental observations. We also experimentally assess these effects for multi-terminal interconnections and in the presence of multiple obstacles. We single out cases where the effects are additive, which suggests the use of lookup tables and equivalent blockage relations. Our results are applicable in chip planning tools, where they enable improved accounting for obstacles in a priori wirelength estimation schemes. Chung-Kuan Cheng, Andrew B. Kahng, Bao Liu 0001, Dirk Stroobandt |
ASP-DAC | 4 |
| 2001 | Toward accurate models of achievable routingabstractModels of achievable routing, i.e., chip wireability, rely on estimates of available and required routing resources. Required routing resources are estimated from placement or (a priori) using wire length estimation models. Available routing resources are estimated by calculating a nominal "supply" then take into account such factors as the efficiency of the router and the impact of vias. Models of achievable routing can be used to optimize interconnect process parameters for future designs or to supply objectives that guide layout tools to promising solutions. Such models must be accurate in order to be useful and must support empirical verification and calibration by actual routing results. In this paper, we discuss the validation of such models and we apply our validation process to three existing models. We find notable inaccuracies in the existing models when matched against real data. We then present a thorough analysis of the assumptions underlying these models. Based on this analysis, we discuss requirements for predictors of routing resources and make suggestions for a new model of achievable routing. Andrew B. Kahng, Stefanus Mantik, Dirk Stroobandt |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2001 | A stochastic model for the interconnection topology of digital circuitsabstractRent's rule has been successfully applied to a priori estimation of wire length distributions. However, this approach is very restrictive: the circuits are assumed to be homogeneous. In this paper, recursive clustering is described as a more advanced model for the partitioning behavior of digital circuits. It is applied to predict the variance of the terminal count distribution. First, the impact of the block degree distribution is analyzed with a simple model. A more refined model incorporates the effect of stochastic self similarity. Finally, the model is further extended to describe the effects of heterogeneity. This model is a promising candidate for more accurate a priori estimation tools. Peter Verplaetse, Dirk Stroobandt, Jan M. Van Campenhout |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | GTX: the MARCO GSRC technology extrapolation systemabstractTechnology extrapolation — the calibration and prediction of achievable design in future technology generations — drives the evolution of VLSI system architectures, design methodologies, and design tools. This paper describes initial experiences with development and use of GTX, the MARCO GSRC Technology Extrapolation system. GTX provides a robust, portable framework for interactive specification and comparison of modeling choices, e.g., for predicting system cycle time, die size and power dissipation. We use GTX to reveal surprising levels of uncertainty (modeling and parameter sensitivity) in widely-cited cycle-time models that drive recent roadmaps. We also describe new SOI and bulk device models that have been developed for GTX, as well as studies of power dissipation and delay uncertainty under various implementation assumptions for global interconnects. Andrew E. Caldwell, Yu Cao 0001, Andrew B. Kahng, Farinaz Koushanfar, Hua Lu 0004, Igor L. Markov, Michael Oliver, Dirk Stroobandt, Dennis Sylvester |
DAC | 8 |
| 2000 | Effects of Global Interconnect Optimizations on Performance Estimation of Deep Submicron DesignabstractIn this paper, we quantify the impact of global interconnect optimization techniques that address such design objectives as delay, peak noise, delay uncertainty due to noise, power, and cost. In doing so, we develop a new system-performance simulation model as a set of studies within the MARCO GSRC Technology Extrapolation (GTX) system. We model a typical point-to-point global interconnect and focus on accurate assessment of both circuit and design technology with respect to such issues as inductance, signal line shielding, dynamic delay, buffer placement uncertainty and repeater staggering. We demonstrate, for example, that optimal wire sizing models need to consider inductive effects-and that use of more accurate (-1,3) worst-case capacitive coupling noise switch factors substantially increases peak noise estimates compared to traditional (0,2) bounds. We also find that optimal repeater sizes are significantly smaller than conventional models would suggest, especially when considering energy-delay issues. Yu Cao 0001, Chenming Hu, Xuejue Huang, Andrew B. Kahng, Sudhakar Muddu, Dirk Stroobandt, Dennis Sylvester |
ICCAD | 6 |
| 2000 | On synthetic benchmark generation methodsabstractIn the process of designing complex chips and systems, the use of benchmark designs is often necessary. However, the existing benchmark suites are not sufficient for the evaluation of new architectures and EDA tools; synthetic benchmark circuits are a viable alternative. In this paper, a systematic approach for the generation and evaluation of synthetic benchmark circuits is presented. A number of existing benchmark generation methods are examined using direct validation of size and topological parameters. This exposes certain features and drawbacks of the different methods. Peter Verplaetse, Jan M. Van Campenhout, Dirk Stroobandt |
ISCAS | 3 |
| 2000 | Requirements for models of achievable routingabstractABSTRACT General Terms Models of achievable routing, i.e., chip wireability, rely on estimates of available and required routing resources. Re-quired routing resources are estimated from placement, or (a priori) using wirelength estimation models. Available rout-ing resources are estimated by calculating a nominal “sup-ply”, then taking into account such factors as the efficiency of the router and the impact of vias. Models of achievable routing can be used to optimize inter-connect process parameters for future designs or to supply objectives that guide layout tools to promising solutions. Such models must be accurate in order to be useful, and must support empirical verification and calibration by ac-tual routing results. In this paper, we discuss the validation of such models and we apply our validation process to three existing models. We find notable inaccuracies in the existing models when matched against real data. We then present a thorough analysis of the assumptions underlying these models; based on this analysis, we discuss requirements for predictors of routing resources within models of achievable routing. Andrew B. Kahng, Stefanus Mantik, Dirk Stroobandt |
ISPD | 3 |
| 2000 | Generating synthetic benchmark circuits for evaluating CAD toolsabstractFor the development and evaluation of computer-aided design tools for partitioning, floorplanning, placement, and routing of digital circuits, a huge amount of benchmark circuits with suitable characteristic parameters is required. Observing the lack of industrial benchmark circuits available for use in evaluation tools, one could consider to actually generate synthetic circuits. In this paper, we extend a graph-based benchmark generation method to include functional information. The use of a user-specified component library, together with the restriction that no combinational loops are introduced, now broadens the scope to timing-driven and logic optimizer applications. Experiments show that the resemblance between the characteristic Rent curve and the net degree distribution of real versus synthetic benchmark circuits is hardly influenced by the suggested extensions and that the resulting circuits are more realistic than before. An indirect validation verifies that existing partitioning programs have comparable behavior for both real and synthetic circuits. The problems of accounting for timing-aware characteristics in synthetic benchmarks are addressed in detail and suggestions for extensions are included. Dirk Stroobandt, Peter Verplaetse, Jan M. Van Campenhout |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2000 | The interpretation and application of Rent's ruleabstractThis paper provides a review of both Rent's rule and the placement models derived from it. It is proposed that the power-law form of Rent's rule, which predicts the number of terminals required by a group of gates for communication with the rest of the circuit, is a consequence of a statistically homogeneous circuit topology and gate placement. The term "homogeneous" is used to imply that quantities such as the average wire length per gate and the average number of terminals per gate are independent of the position within the circuit. Rent's rule is used to derive a variety of net length distribution models and the approach adopted in this paper is to factor the distribution function into the product of an occupancy probability distribution and a function which represents the number of valid net placement sites. This approach places diverse placement models under a common framework and allows the errors introduced by the modeling process to be isolated and evaluated. Models for both planar and hierarchical gate placement are presented. Phillip Christie, Dirk Stroobandt |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | On an Efficient Method for Estimating the Interconnection Complexity of Designs and on the Existence of Region III in Rent's RuleabstractThe interconnection complexity of digital designs can be captured by the well-known Rent exponent, described by Landman and Russo [1971]. In this paper we present an efficient method for obtaining the Rent exponent of a design through a hierarchical partitioning algorithm. Experimental results not only confirm the Landman and Russo observations of a region I and region II but also show a hitherto unknown region 12%. Dirk Stroobandt |
Great Lakes Symposium on VLSI | 1 |
| 1999 | Towards synthetic benchmark circuits for evaluating timing-driven CAD toolsabstractFor the development and evaluation of CAD-tools for partitioning, floorplanning, placement, and routing of digital circuits, a huge amount of benchmark circuits with suitable characteristic parameters is required. Observing the lack of industrial benchmark circuits for use in evaluation tools, one could consider to actually generate such circuits. In this paper, we extend a graph-based benchmark generation method to include functional information. The use of a user-specified component library, together with the restriction that no combinational loops are introduced, now broadens the scope to timing-driven and logic optimizer applications. Experiments show that the resemblance between the characteristic Rent curve and the net degree distribution of real versus synthetic benchmark circuits is hardly influenced by the suggested extensions and that the resulting circuits are more realistic than before. However, the synthetic benchmark circuits are still very redundant, compared to existing sets of real benchmarks. It is shown that a correlation exists between the degree of redundancy and key circuit parameters. Dirk Stroobandt, Peter Verplaetse, Jan M. Van Campenhout |
ISPD | 1 |
| 1999 | Generating new benchmark designs using a multi-terminal net model
Dirk Stroobandt, Jo Depreitere, Jan M. Van Campenhout |
Integr. | 1 |
| 1998 | A Quantitative Study of the Benefits of Area-I/O in FPGAsabstractDesigns targeted for FPGAs are becoming increasingly larger and more complex. The need for I/O often surpasses the number of I/O pads that can be provided at the perimeter of the FPGA chip. As a result, these designs have to be implemented in larger FPGAs, the size of which is fired by the number of I/O pads and not by the logic needed, reducing the performance of the implementation. Providing FPGA chips with I/O pads that are spread out across the whole chip area drastically reduces this problem. In this paper, we present a quantitative analysis of the impact of area-I/O in FPGAs. Herwig Van Marck, Jo Depreitere, Dirk Stroobandt, Jan M. Van Campenhout |
Great Lakes Symposium on VLSI | 3 |
| 1998 | On the Characterization of Multi-Point Nets in Electronic DesignsabstractImportant layout properties of electronic designs include interconnection length values, clock speed, area requirements, and power dissipation. A reliable estimation of those properties is essential for improving placement and routing techniques for digital circuits. Previous work on estimating design properties failed to take multi-point nets into account. All nets were assumed to be 2-point nets (especially for estimating the number of nets). In this paper we aim at characterizing multi-point nets in electronic designs. We develop a model for the behaviour of multi-point nets during the partitioning process. The resulting distribution of nets over their net degree is validated through comparison with benchmark data. Dirk Stroobandt, Fadi J. Kurdahi |
Great Lakes Symposium on VLSI | 1 |
| 1996 | Hierarchical Test Generation with Built-In Fault DiagnosisabstractA hierarchical test generation method is presented that uses the inherent hierarchical structure of the circuit under test and takes fault diagnosability into account right from the start. An efficient test compaction method leads to a very compact test set, while retaining a maximum of diagnostic power and a 100% fault coverage for non-fanout circuits. An extension for fanout circuits is also presented. Dirk Stroobandt, Jan M. Van Campenhout |
Asian Test Symposium | 1 |
| 1996 | An Accurate Interconnection Length Estimation for Computer LogicabstractImportant layout properties of electronic designs include space requirements and interconnection lengths. A reliable interconnection length estimation is essential for improving placement and routing techniques. Donath found an upper bound for the average interconnection length that follows the trends of experimentally obtained average lengths. Yet, this upper bound deviates from the experimentally obtained value by a factor /spl delta//spl ap/2, which is not sufficiently accurate for some applications. We show that we obtain a significantly more accurate estimate by taking into account the inherent features of the optimal placement process. Dirk Stroobandt, Herwig Van Marck, Jan M. Van Campenhout |
Great Lakes Symposium on VLSI | 1 |