Martin Danek

dblp:43/4274 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
0since 2021 · last 2012
0000-0002-3030-7375ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 4 first-authorArtificial intelligence and machine learning · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Electronic design automation · 49% Reconfigurable computing and FPGAs · 43% Integrated circuit design · 4%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
dynamic reconfiguration
0.122005
Figaro: an automatic tool flow for designs with dynamic reconfiguration (abstract only) · FPGA 2005
Dynamic reconfiguration in FPGA-based SoC designs (abstract only) · FPGA 2005
Electronic design automation
physical design
0.122005
Figaro: an automatic tool flow for designs with dynamic reconfiguration (abstract only) · FPGA 2005
FPGA modelling for high-performance algorithms · FPGA 2004
Electronic design automation › physical design › placement and routing
FPGA placement and routing
0.112005
Figaro: an automatic tool flow for designs with dynamic reconfiguration (abstract only) · FPGA 2005
Reconfigurable computing and FPGAs
FPGA SoC
0.112005
Dynamic reconfiguration in FPGA-based SoC designs (abstract only) · FPGA 2005
Electronic design automation › physical design
placement and routing
0.012004
FPGA modelling for high-performance algorithms · FPGA 2004
Integrated circuit design
system-on-chip
0.012005
Dynamic reconfiguration in FPGA-based SoC designs (abstract only) · FPGA 2005
Reconfigurable computing and FPGAs
FPGA architecture
0.012004
FPGA modelling for high-performance algorithms · FPGA 2004

Methods — techniques the papers use, named apart from their topics

floating-point hardware implementation · 0.1dynamic reconfiguration · 0.1topological modeling · 0.0elmore delay model · 0.0
YearPublicationVenuePosition
2012 The architecture and the technology characterization of an FPGA-based customizable Application-Specific Vector Processor
abstract
The traditional approach to IP core design is to use simulations with test vectors. This is not feasible when dealing with complex function cores such as the Image Segmentation case-study algorithm in this paper. An algorithm developer needs to carry out experiments on large real-world data sets, with fast turn-around times, and in real time to facilitate performance tuning and incremental development. We propose a methodology called Application-Specific Vector Processor (ASVP). The ASVP approach first constructs a programmable architecture customized for a given application, then employs software techniques to develop firmware that implements the algorithm. Our sample implementation that supports the Image Segmentation kernel is capable of 332 MFLOPs, 400 MFLOPs, and 250 MFLOPs per coprocessor core in Virtex 5, Virtex 6 and Spartan 6 technologies, respectively. The core size is roughly 1500 slices, depending on the configuration and technology.
Jaroslav Sykora, Lukas Kohout, Roman Bartosinski, Leos Kafka, Martin Danek, Petr Honzík
DDECS5
2012 Reducing Instruction Issue Overheads in Application-Specific Vector Processors
abstract
The traditional approach to IP core design is to use simulations with test vectors. This is not feasible when dealing with complex function cores such as the Image Segmentation case-study algorithm in this paper. An algorithm developer needs to carry out experiments on large real-world data sets, with fast turn-around times, and in real time to facilitate performance tuning and incremental development. Previously we proposed a methodology called Application-Specific Vector Processor (ASVP). The ASVP approach first constructs a programmable architecture customized for a given application, then employs software techniques to develop firmware that implements the algorithm. In our setting we employ an embedded simple scalar CPU (8-bit PicoBlaze 3) to control a floating-point vector processing unit (VPU) by issuing wide (horizontally encoded) instructions to it. In this work we dramatically reduce the overhead of the wide-instruction issue (in one case by 13x) by implementing a new two-level configuration table. The table stores frequently used vector definitions (in Level 1) and vector instructions (in Level 2), pre-loading them quickly into the issue buffer. A configuration in the issue buffer can be further modified before being sent to the processing unit. This ensures the architecture stays general and fully customizable.
Jaroslav Sykora, Roman Bartosinski, Lukas Kohout, Martin Danek, Petr Honzík
DSD4
2011 Microthreading as a Novel Method for Close Coupling of Custom Hardware Accelerators to SVP Processors
abstract
We present a new low-level interfacing scheme for connecting custom accelerators to processors that tolerates latencies that usually occur when accessing hardware accelerators from software. The scheme is based on the Self-adaptive Virtual Processor (SVP) architecture and on the micro-threading concept. Our presentation is based on a sample implementation of the SVP architecture in an extended version of the LEON3 processor called UTLEON3. The SVP concurrency paradigm makes data dependencies explicit in the dynamic tree of threads. This enables a system to execute threads concurrently in different processor cores. Previous SVP work presumed the cores are homogeneous, for example an array of micro threaded processors sharing a dynamic pool of micro threads. In this work we propose a heterogeneous system of general-purpose processor cores and custom hardware accelerators. The accelerators dynamically pick families of threads from the pool and execute them concurrently. We introduce the Thread Mapping Table (TMT) hardware unit that couples the software and hardware implementations of the user computations. The TMT unit allows to realize the coupling scheme seamlessly without modifications of the processor ISA. The advantage of the described scheme is in decoupling application programming from specific details of the hardware accelerator architecture (identical behaviour of a software create and hardware create), and in eliminating the influence of hardware access latencies. Our simulation and FPGA implementation results prove that the additional hardware access latencies in the processor are tolerated by the SVP architecture.
Jaroslav Sykora, Leos Kafka, Martin Danek, Lukas Kohout
DSD3
2011 SMECY: smart multi-core embedded systems
abstract
SMECY project is an ambitious European initiative involving 29 partners across 9 countries to enable Europe to have a leader role in multi-core domain by developing new programming technologies enabling the exploitation of architectures offering hundreds of cores. Multi-core technologies will rapidly provide to the parallel computing field improved performance, energy saving and cost reduction and will become of strategic value in winning market share in all areas of embedded systems. Given the need, SMECY lays the focus on targeting programming multi-core architecture for consumer electronics with efficient resources management. The first presentation describes the overall project while the two others are respectively dedicated to the multi-core platforms targeted in the project and the description of the tools constituting the bricks of the tool chains.
François Pacull, Koen Bertels, Martin Danek, Giulio Urlini
ACM Great Lakes Symposium on VLSI3
2010 Instruction set extensions for multi-threading in LEON3
abstract
This paper describes instruction set extensions for a variant of multi-threading called micro-threading for the LEON3 SPARCv8 processor. We show an architecture of the developed processor and its key blocks - cache controller, register file, thread scheduler. The processor has been implemented in a Xilinx Virtex2Pro FPGA. The extensions are evaluated in terms of extra resources needed, and the overall performance of the developed processor is evaluated on a simple DSP computation typical for embedded systems.
Martin Danek, Leos Kafka, Lukas Kohout, Jaroslav Sykora
DDECS1
2010 Reconfigurable hardware objects for image processing on FPGAs
abstract
Embedded systems are getting more complex; that is why the high level of abstraction is required during the development process. High abstraction methods simplify implementation of complex computation systems and shorten the time to market. This paper presents an implementation of a graphic computing element (GCE) which can be used as a runtime parametrized building block in image processing applications in FPGAs. In terms of the object oriented model GCE encapsulates its internal data representation and rules for their manipulation. Several basic image processing operations have been implemented (Sobel edge detection, Gauss, mean, etc. filtering). These operations are called as GCE methods. Because of high spatial dependency of image data in image processing, an efficient image data reuse method has been implemented.
Jan Kloub, Petr Honzík, Martin Danek
DDECS3
2008 Increasing the level of abstraction in FPGA-based designs
abstract
Traditional design techniques for FPGAs are based on using hardware description languages, with functional and post-place-and-route simulation as a means to check design correctness and remove detected errors. With large complexity of things to be designed it is necessary to introduce new design approaches that will increase the level of abstraction while maintaining the necessary efficiency of a computation performed in hardware that we are used to today. This paper presents one such methodology that builds upon existing research in multithreading, object composability and encapsulation, partial runtime reconfiguration, and self adaptation. The methodology is based on currently available FPGA design tools. The efficiency of the methodology is evaluated on basic vector and matrix operations.
Martin Danek, Jirí Kadlec, Roman Bartosinski, Lukas Kohout
FPL1
2007 Accelerating Microblaze Floating Point Operations
abstract
The MicroBlaze processor serves in many FPGA designs as the central 32 bit CPU with access to the global off chip memory and peripherals. MicroBlaze provides FSL links for up to 8 coprocessors. We present two MicroBlaze designs. The first design works with 8 PicoBlaze-based accelerators for pipelined, single-precision floating point vector-oriented operations, and delivers over 1.2 GFLOPs. The second design uses 4 similar double precision accelerators and delivers 600 MFLOPs. The acceleration results are documented on batch computation of a finite impulse response filter. Each PicoBlaze soft core can be re-programmed by MicroBlaze. This provides a framework for a partial dynamic change of the functionality of accelerators. This program change can be done via the FSL link in parallel with the current computation of the accelerator.
Jirí Kadlec, Roman Bartosinski, Martin Danek
FPL3
2005 Dynamic reconfiguration in FPGA-based SoC designs (abstract only)
abstract
This paper discusses architectural issues arising from the use of dynamic reconfiguration and shows a possible use of dynamic reconfiguration to extend and accelerate a computation performed in system-on-a-chip designs with microprocessors with fixed instruction sets. Further a sample application is discussed that uses a dynamically reconfigurable FPGA to implement different floating-point calculations in hardware, reconfigured as required by the execution of the user code. The implementation data for two dynamically reconfigurable platforms available on the market - the Xilinx Virtex2 family FPGAs and the Atmel FPSLIC family FPGAs - is compared in terms of resource requirements, operating frequency, and power consumption.
Roman Bartosinski, Martin Danek, Petr Honzík, Rudolf Matousek
FPGA2
2005 Figaro: an automatic tool flow for designs with dynamic reconfiguration (abstract only)
abstract
Although runtime dynamic reconfiguration of the FPGA devices has been an issue of the last decade, it has yet to achieve general recognition by the design community. The reasons for this are clear; there exists no straightforward design methodology, and the partitioning and CAD tool support is poor. This paper presents general concepts implemented in a placement and routing tool that provides an environment where designs that are partially and dynamically reconfigurable can be processed in order to be implemented on FPGAs that support this technology, such as the Atmel AT40K and AT94K series. The function of the tool is demonstrated on a simple real-world example.
Kelly Nasi, Martin Danek, Theodoros Karoubalis, Zdenek Pohl
FPGA2
2005 Figaro - An Automatic Tool Flow for Designs with Dynamic Reconfiguration
abstract
Although runtime partial dynamic reconfiguration of FPGAs has been researched for many years and there have been a few FPGAs equipped with the required architectural features, it has yet to achieve general recognition by the commercial design community. This is mainly due to the lack of a professional CAD tool support. This paper presents extended concepts from E. L. Horta et al. (2002), Xilinx Application Note 290 (2004) and I. Robertson et al. (2002) implemented in a placement and routing tool. The tool supports creation of partially dynamically reconfigurable designs from input EDIF files and user-specified reconfiguration schedule down to bitstream generation for FPGAs that support this technology, such as the Atmel AT40K and AT94K series.
Kelly Nasi, Martin Danek, Theodoros Karoubalis, Zdenek Pohl
FPL2
2004 FPGA modelling for high-performance algorithms
abstract
The poster deals with topological modelling of FPGA circuits for timing-driven algorithms. It presents a method for analysing and deriving topological placement/routing models from architectural description of existing FPGAs. A metric is introduced that reflects information loss in more abstract global routing models. The metric captures both the loss of topological information and the decrease in precision of possible signal delay estimation based on the model. The practical use of the metric is demonstrated for a wire-type model and a global routing model derived from the Xilinx XC4000 family and for two linear and one Elmore-based signal delay estimation models. This research has been partially supported by the Grant Agency of the Czech Republic under Project No. 102/04/2137.
Martin Danek, Josef Kolár
FPGA1
2002 Integrated Iterative Approach to FPGA Placement
Martin Danek, Zdenek Muzikár
FPL1
2002 XCS Applied To Mapping FPGA Architectures
Martin Danek, Robert E. Smith 0001
GECCO1
2001 A Generalisable Measure of Self-Organisation and Emergence
W. Andy Wright, Robert E. Smith 0001, Martin Danek, Pillip Greenway
ICANN3