EDBT 2026 Demo / reviewers in the wild / expert
Tomasz S. Czajkowski
dblp:95/3345 · also Tomasz Czajkowski
· DBLP profile ↗
20ranked-venue papers
8as first author
3since 2021 · last 2026
0009-0008-3294-6198ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 8 first-author · 3 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | InterStellar 2.0: Fine-grained stream-guided HW/SW co-design for multi-channel DRAM performance steeringabstractThe gap between processor speed and memory latency limits system scalability, especially in data-intensive and artificial intelligence workloads where memory-level parallelism and bandwidth efficiency are critical. Prior work ( InterStellar ) showed that hardware/software (HW/SW) co-design can expose program-level access streams to the memory system, enabling more informed memory-controller (MC) scheduling. This work extends that approach to high-bandwidth, multi-channel platforms. We present InterStellar 2.0 , a scalable HW/SW co-design that: (1) supports multi-channel dynamic random-access memory (DRAM) by partitioning stream batches across channels, allowing each channel to operate independently without cross-channel coordination; and (2) introduces fine-grained stream descriptors so software can distinguish distinct access patterns, even within the same data structure. These capabilities improve DRAM locality management and allow the MC to issue future requests efficiently. We evaluate InterStellar 2.0 on an 8-core RISC-V platform across DRAM configurations from 1 to 32 channels. At 32 channels, InterStellar 2.0 improves performance by up to 2.92 × and increases memory bandwidth by up to 2.83 × over a commercial off-the-shelf (COTS) controller. Fine-grained stream tracking alone improves performance by up to 1 . 35 × . Overall, InterStellar 2.0 shows that stream-aware HW/SW co-design is practical, compatible, and scalable for multi-channel memory systems without ISA changes or inter-channel communication. Abdelrhman Mohamed Abotaleb, Maziar Goudarzi, Tomasz S. Czajkowski, Mohamed Hassan 0002 |
J. Syst. Archit. | 3 |
| 2024 | Work-in-Progress:ACPO: An AI-Enabled Compiler FrameworkabstractThis paper presents ACPO: An AI-Enabled Compiler Framework; a novel framework that provides LLVM with simple and comprehensive tools to enable employing ML models for different optimization passes. We showcase a couple of use cases of ACPO by ML-enabling the Loop Unroll (LU) and Function Inlining (FI) passes and experimental results reveal that by including both models, ACPO can provide a combined speedup of 2.4% on Cbench when compared with LLVM’s O3. Amir H. Ashouri, Muhammad Asif Manzoor, Raymond Zhang, Angel Zhang, Tomasz S. Czajkowski, Yaoqing Gao |
CASES | 8 |
| 2024 | RoDMap: A Reserve-on-Demand Mapper for Spatially-Configured Coarse-Grained Reconfigurable ArraysabstractWe propose, implement, and evaluate a novel approach for mapping dataflow graphs (DFGs) onto spatially configured Coarse-Grained Reconfigurable Arrays (CGRAs). The approach tackles mapping failure due to the congestion that arises when more than one routing path uses the same CGRA link. Heuristics are used to identify congestion patterns and “reserve” CGRA processing elements (PEs) around the congestion by preventing them from being used for DFG nodes in a mapping re-attempt. The reserved PEs effectively increases routing resources around the congestion, thereby increasing the likelihood of mapping success. This approach is referred to as reserve-on-demand mapping since PEs are reserved only when congestion exists and is driven by its patterns. Kyle Zhao Bin Chen, Tarek S. Abdelrahman, Tomasz S. Czajkowski, Maziar Goudarzi |
ICPP | 4 |
| 2018 | High-level synthesis of software-customizable floating-point coresabstractParameterized cores with fixed capabilities are typically used for floating-point (FP) operations on FPGAs. However, such standard cores can be over provisioned or lack specific specializations as required by applications. We consider FP cores described in the C language, synthesized to hardware using the LegUp high-level synthesis (HLS) tool [1]. Their software specification permits straightforward customization to non-compliant variants having superior area and performance characteristics, such as reduced-precision floating point, or cores without full IEEE 754 exceptions support. We create and evaluate the IEEE 754 FP standard cores for the key operations of addition, subtraction, division and multiplication, targeted to an FPGA and compare with widely used optimized RTL FP cores from Altera [7] and FloPoCo [3]. The software-specified HLS-generated cores are surprisingly close to the optimized RTL cores in terms of area/performance, and superior in certain cases, such as FP division. Samridhi Bansal, Hsuan Hsiao, Tomasz S. Czajkowski, Jason Helge Anderson |
DATE | 3 |
| 2015 | Silicon Verification using High-Level Design Tools (Abstract Only)abstractModern FPGAs comprise ever more complex blocks to enable a wide variety of customer applications. Verification of the complex blocks can be a time consuming process, especially at the late stages of the release cycle. A key challenge is the time it takes to create circuits that can run on a target device to test a given block. This paper demonstrates how High-Level Design tools, such as Altera SDK for OpenCL, can be utilized to aid in this work to verify the operation of complex hardened blocks. As a proof of concept, we present the methodology used to verify the correctness of hardened single-precision floating point adder, subtractor and multiplier units on Altera Arria 10 FPGA in a single day. Each design comprised an instance of a hardened floating point unit, either an adder, subtractor or a multiplier, and a functional equivalent there of implemented purely using Lookup Tables (LUTs). Both the hardened module instance and the LUT implementation were generated from OpenCL description using Altera SDK for OpenCL. The results for each computation were compared between the two implementations and any single discrepancy constituted a test failure. To simplify the test, the I/O for each design comprised LEDs (for pass/fail/running/done status) and two switches -- start and reset. Tomasz S. Czajkowski |
FPGA | 1 |
| 2015 | High-Level Design Tools for Floating Point FPGAsabstractThis tutorial describes tools for efficiently implementing floating point applications on FPGAs. We present both the SDK for OpenCL and DSP Builder Advanced Blockset and show that they can be effectively used to implement many floating point applications. The methods for optimizing application performance are also described. Deshanand P. Singh, Bogdan Pasca 0001, Tomasz S. Czajkowski |
FPGA | 3 |
| 2014 | Automating the Design of Processor/Accelerator Embedded Systems with LegUp High-Level SynthesisabstractLegUp [1] is an open-source high-level synthesis (HLS) tool that accepts a C program as input and automatically synthesizes it into a hybrid system. The hybrid system comprises an embedded processor and custom accelerators that realize user-designated compute-intensive parts of the program with improved throughput and energy efficiency. In this paper, we overview the LegUp framework and describe several recent developments: 1) support for an embedded ARM processor, as is available on Altera's recently released SoC FPGA, 2) HLS support for software parallelization schemes -- pthreads and OpenMP, 3) enhancements to LegUp's core HLS algorithms that raise the quality of the auto-generated hardware, and, 4) a preliminary debugging and verification framework providing C source-level debugging of HLS hardware. Since its first release in 2011, LegUp has been downloaded over 1000 times by groups around the world, providing a powerful platform for new research in high-level synthesis algorithms and embedded systems design. Blair Fort, Andrew Canis, Jongsok Choi, Nazanin Calagar, Ruolong Lian, Stefan Hadjis, Yu Ting Chen, Mathew Hall, Bain Syrowik, Tomasz S. Czajkowski, Stephen Brown 0003, Jason Helge Anderson |
EUC | 10 |
| 2013 | From software to accelerators with LegUp high-level synthesisabstractEmbedded system designers can achieve energy and performance benefits by using dedicated hardware accelerators. However, implementing custom hardware accelerators for an application can be difficult and time intensive. LegUp is an open-source high-level synthesis framework that simplifies the hardware accelerator design process [8]. With LegUp, a designer can start from an embedded application running on a processor and incrementally migrate portions of the program to hardware accelerators implemented on an FPGA. The final application then executes on an automatically-generated software/hardware coprocessor system. This paper presents on overview of the LegUp design methodology and system architecture, and discusses ongoing work on profiling, hardware/software partitioning, hardware accelerator quality improvements, Pthreads/OpenMP support, visualization tools, and debugging support. Andrew Canis, Jongsok Choi, Blair Fort, Ruolong Lian, Qijing Huang 0001, Nazanin Calagar, Marcel Gort, Jia Jun Qin, Mark Aldham, Tomasz S. Czajkowski, Stephen Brown 0003, Jason Helge Anderson |
CASES | 10 |
| 2013 | Harnessing the power of FPGAs using altera's OpenCL compilerabstractIn recent years, Field-Programmable Gate Arrays have become extremely powerful computational platforms that can efficiently solve many complex problems. The most modern FPGAs comprise effectively millions of programmable elements, signal processing elements and high-speed interfaces, all of which are necessary to deliver a complete solution. The power of FPGAs is unlocked via low-level programming languages such as VHDL and Verilog, which allow designers to explicitly specify the behavior of each programmable element. While these languages provide a means to create highly efficient logic circuits, they are akin to "assembly language" programming for modern processors. This is a serious limiting factor for both productivity and the adoption of FPGAs on a wider scale. In this talk, we use the OpenCL language to explore techniques that allow us to program FPGAs at a level of abstraction closer to traditional software-centric approaches. OpenCL is an industry standard parallel language based on 'C' that offers numerous advantages that enable designers to take full advantage of the capabilities offered by FPGAs, while providing a high-level design entry language that is familiar to a wide range of programmers. Deshanand P. Singh, Tomasz S. Czajkowski, Andrew C. Ling |
FPGA | 2 |
| 2013 | LegUp: An open-source high-level synthesis tool for FPGA-based processor/accelerator systemsabstractIt is generally accepted that a custom hardware implementation of a set of computations will provide superior speed and energy efficiency relative to a software implementation. However, the cost and difficulty of hardware design is often prohibitive, and consequently, a software approach is used for most applications. In this article, we introduce a new high-level synthesis tool called LegUp that allows software techniques to be used for hardware design. LegUp accepts a standard C program as input and automatically compiles the program to a hybrid architecture containing an FPGA-based MIPS soft processor and custom hardware accelerators that communicate through a standard bus interface. In the hybrid processor/accelerator architecture, program segments that are unsuitable for hardware implementation can execute in software on the processor. LegUp can synthesize most of the C language to hardware, including fixed-sized multidimensional arrays, structs, global variables, and pointer arithmetic. Results show that the tool produces hardware solutions of comparable quality to a commercial high-level synthesis tool. We also give results demonstrating the ability of the tool to explore the hardware/software codesign space by varying the amount of a program that runs in software versus hardware. LegUp, along with a set of benchmark C programs, is open source and freely downloadable, providing a powerful platform that can be leveraged for new research on a wide range of high-level synthesis topics. Andrew Canis, Jongsok Choi, Mark Aldham, Victor Zhang, Ahmed Kammoona, Tomasz S. Czajkowski, Stephen Brown 0003, Jason Helge Anderson |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2012 | Impact of Cache Architecture and Interface on Performance and Area of FPGA-Based Processor/Parallel-Accelerator SystemsabstractWe describe new multi-ported cache designs suitable for use in FPGA-based processor/parallel-accelerator systems, and evaluate their impact on application performance and area. The baseline system comprises a MIPS soft processor and custom hardware accelerators with a shared memory architecture: on-FPGA L1 cache backed by off-chip DDR2 SDRAM. Within this general system model, we evaluate traditional cache design parameters (cache size, line size, associativity). In the parallel accelerator context, we examine the impact of the cache design and its interface. Specifically, we look at how the number of cache ports affects performance when multiple hardware accelerators operate (and access memory) in parallel, and evaluate two different hardware implementations of multi-ported caches using: 1) multi-pumping, and 2) a recently-published approach based on the concept of a live-value table. Results show that application performance depends strongly on the cache interface and architecture: for a system with 6 accelerators, depending on the cache design, speed up swings from 0.73× to 6.14×, on average, relative to a baseline sequential system (with a single accelerator and a direct-mapped, 2KB cache with 32B lines). Considering both performance and area, the best architecture is found to be a 4-port multi-pump direct-mapped cache with a 16KB cache size and a 128B line size. Jongsok Choi, Kevin Nam, Andrew Canis, Jason Helge Anderson, Stephen Brown 0003, Tomasz S. Czajkowski |
FCCM | 6 |
| 2012 | Impact of FPGA architecture on resource sharing in high-level synthesisabstractResource sharing is a key area-reduction approach in high-level synthesis (HLS) in which a single hardware functional unit is used to implement multiple operations in the high-level circuit specification. We show that the utility of sharing depends on the underlying FPGA logic element architecture and that different sharing trade-offs exist when 4-LUTs vs. 6-LUTs are used. We further show that certain multi-operator patterns occur multiple times in programs, creating additional opportunities for sharing larger composite functional units comprised of patterns of interconnected operators. A sharing cost/benefit analysis is used to inform decisions made in the binding phase of an HLS tool, whose RTL output is targeted to Altera commercial FPGA families: Stratix IV (dual-output 6-LUTs) and Cyclone II (4-LUTs). Stefan Hadjis, Andrew Canis, Jason Helge Anderson, Jongsok Choi, Kevin Nam, Stephen Brown 0003, Tomasz S. Czajkowski |
FPGA | 7 |
| 2012 | From opencl to high-performance hardware on FPGASabstractWe present an OpenCL compilation framework to generate high-performance hardware for FPGAs. For an OpenCL application comprising a host program and a set of kernels, it compiles the host program, generates Verilog HDL for each kernel, compiles the circuit using Altera Complete Design Suite 12.0, and downloads the compiled design onto an FPGA.We can then run the application by executing the host program on a Windows(tm)-based machine, which communicates with kernels on an FPGA using a PCIe interface. We implement four applications on an Altera Stratix IV and present the throughput and area results for each application. We show that we can achieve a clock frequency in excess of 160MHz on our benchmarks, and that OpenCL computing paradigm is a viable design entry method for high-performance computing applications on FPGAs. Tomasz S. Czajkowski, Utku Aydonat, Dmitry Denisenko, John Freeman, Michael Kinsner, David Neto, Peter Yiannacouras, Deshanand P. Singh |
FPL | 1 |
| 2011 | LegUp: high-level synthesis for FPGA-based processor/accelerator systemsabstractIn this paper, we introduce a new open source high-level synthesis tool called LegUp that allows software techniques to be used for hardware design. LegUp accepts a standard C program as input and automatically compiles the program to a hybrid architecture containing an FPGA-based MIPS soft processor and custom hardware accelerators that communicate through a standard bus interface. Results show that the tool produces hardware solutions of comparable quality to a commercial high-level synthesis tool. Andrew Canis, Jongsok Choi, Mark Aldham, Victor Zhang, Ahmed Kammoona, Jason Helge Anderson, Stephen Brown 0003, Tomasz S. Czajkowski |
FPGA | 8 |
| 2010 | Decomposition-Based Vectorless Toggle Rate Computation for FPGA CircuitsabstractThis paper presents a novel and accurate method of estimating the toggle rates of signals in field-programmable gate array (FPGA)-based logic circuits without the use of simulation vectors. Compared to previous vectorless techniques, our approach provides improved accuracy-of-results, especially for individual signals, which could be leveraged by computer-aided design (CAD) tools for performing power optimization of logic circuits. Increased accuracy is achieved by using stochastic methods that estimate the transition densities at FPGA logic elements while accounting for both spatial and temporal correlation of logic signals. Spatial correlation is calculated by leveraging a unique XOR-based decomposition technique that provides both accurate results and fast computation times. We also consider the delay information of implemented circuits, providing for a comprehensive treatment of glitches, including the effects of inertial limits on power dissipation. Our toggle-rate estimation approach has been tested on a commonly used set of Microelectronic Center of North Carolina circuits, as well as a set of industrial circuits targeted to Altera Stratix II FPGAs. Results show that our techniques provide a three times lower percent error, while maintaining a low processing time, when compared to two existing techniques: the vectorless estimation tool shipped with the commercial Quartus II 8.0 CAD tool, and the ACE v2.0 academic tool produced from the University of British Columbia, Vancouver, BC, Canada. Tomasz S. Czajkowski, Stephen Brown 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2008 | Functionally linear decomposition and synthesis of logic circuits for FPGAsabstractThis paper presents a novel logic synthesis method to reduce the area of XOR-based logic functions. The idea behind the synthesis method is to exploit linear dependency between logic sub-functions to create an implementation based on an XOR relationship with a lower area overhead. Experiments conducted on a set of 99 MCNC benchmark (25 XOR based, 74 non-XOR) circuits show that this approach provides an average of 18.8% area reduction as compared to BDS-PGA 2.0 and 25% area reduction as compared to ABC for XOR-based logic circuits. Tomasz S. Czajkowski, Stephen Brown 0003 |
DAC | 1 |
| 2008 | Fast toggle rate computation for FPGA circuitsabstractThis paper presents a fast and scalable method of computing signal toggle rate in FPGA-based circuits. Our technique is a vectorless estimation technique, which can be used in a CAD tool to identify the parts of the circuit that can benefit from power optimization. A key advantage of our approach is its ability to efficiently account for spatial correlation of related logic cones, which is accomplished using a novel XOR-based decomposition. In addition, our approach uses post-routing circuit delays to account for glitches in a logic circuit. The proposed approach was tested on 14 MCNC benchmark circuits compiled for the Altera Stratix II devices. The results indicate that our method improves the vectorless estimation technique available in the latest version of Alterapsilas Quartus II commercial CAD tool, reducing the average error by 37% and standard deviation by 59%. Tomasz S. Czajkowski, Stephen Brown 0003 |
FPL | 1 |
| 2008 | Functionally Linear Decomposition and Synthesis of Logic Circuits for FPGAsabstractThis paper presents a novel XOR-based logic synthesis approach called functionally linear decomposition and synthesis (FLDS). This approach decomposes a logic function to expose an XOR relationship by using Gaussian elimination. It is fundamentally different from the traditional approaches to this problem, which are based on the work of Ashenhurst and Curtis. FLDS utilizes binary decision diagrams to efficiently represent logic functions, making it fast and scalable. This technique was tested on a set of 99 MCNC benchmarks, mapping each design into a network of four input lookup tables. On the 25 of the benchmarks, which have been classified by previous researchers as XOR-based logic circuits, our approach provides significant area savings. In comparison to the leading logic synthesis tools, ABC and BDS-PGA 2.0, FLDS produces XOR-based circuits with 25.3% and 18.8% smaller area, respectively. The logic circuit depth is also improved by 7.7% and 14.5%, respectively. Tomasz S. Czajkowski, Stephen Brown 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2007 | Using Negative Edge Triggered FFs to Reduce Glitching Power in FPGA CircuitsabstractThis paper presents an algorithm for reducing dynamic power dissipated by Field-Programmable Gate Array (FPGA) circuits. The algorithm uses a fast probability based model to estimate glitches on wires in a circuit and then inserts negative edge triggered FFs at outputs of Lookup Tables (LUTs) that produce glitches. A negative edge triggered FF maintains the logic value produced by the LUT in the previous cycle for the first half of the clock period, filtering glitches that occur at the output of the LUT. The power dissipation is lowered by reducing the number of transitions that propagate to the general routing network. Tomasz S. Czajkowski, Stephen Brown 0003 |
DAC | 1 |
| 2004 | A synthesis oriented omniscient manual editorabstractThe cost functions used to evaluate logic synthesis transformations for FPGAs are far removed from the final speed and routability determined after placement, routing and timing analysis. This distance has given rise to the field of physical synthesis, which attempts to improve logic synthesis by employing cost functions that contain placement, routing and/or timing analysis information.In this work we take this notion to an extreme that we call omniscience, in which post-routing timing analysis is provided in the context of a manual editor in which the user selects logical and physical transformations. After each incremental circuit modification, the user is informed of the circuit performance after routing and timing analysis. Since the computations involved in providing this level of information are large, we restrict the application to relatively small circuits, no larger than 1000 logic elements.Using this approach on a commercial FPGA, we propose a set of logic transformations specific to the logic and routing architecture of the Xilinx Virtex-E device. On a set of 10 circuits we have achieved an average performance improvement of 10% when both logical and physical changes are used. Another value of the editor is that it reveals new types of automatable physical-synthesis transformations and optimization strategies that arise from architectural properties of the target device. Tomasz S. Czajkowski, Jonathan Rose |
FPGA | 1 |