Steven Colleman

dblp:286/8788 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-4199-2926ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Novel Depth-First Scheduling for Spatially Dynamic Neural Networks
Steven Colleman, Andrea Nardi-Dei, Marc Geilen, Sander Stuijk, Toon Goedemé
ICAART (4)1
2025 Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
abstract
As the landscape of deep neural networks evolves, heterogeneous dataflow accelerators, in the form of multi-core architectures or chiplet-based designs, promise more flexibility and higher inference performance through scalability. So far, these systems exploit the increased parallelism by coarsely mapping a single layer at a time across cores, which incurs frequent costly off-chip memory accesses, or by pipelining batches of inputs, which falls short in meeting the demands of latency-critical applications. To alleviate these bottlenecks, this work explores a new fine-grain mapping paradigm, referred to as layer fusion, on heterogeneous dataflow accelerators through a novel design space exploration framework called Stream . Stream captures a wide variety of heterogeneous dataflow architectures and mapping granularities, and implements a memory and communication-aware latency and energy analysis validated with three distinct state-of-the-art hardware implementations. As such, it facilitates a holistic exploration of architecture and mapping, by strategically allocating the workload through constraint optimization. The findings demonstrate that the integration of layer fusion with heterogeneous dataflow accelerators yields up to 2.2× lower energy-delay product in inference efficiency, addressing both energy consumption and latency concerns.
Arne Symons, Linyan Mei, Steven Colleman, Pouya Houshmand, Sebastian Karl, Marian Verhelst
IEEE Trans. Computers3
2023 Stream: A Modeling Framework for Fine-grained Layer Fusion on Multi-core DNN Accelerators
abstract
To keep up with the ever-growing performance demand of DNN processing, specialized hardware (HW) accelerators are shifting towards multi-core architectures. Stream is the first open-source design space exploration (DSE) framework for co-optimization of HW architecture and fine-grained scheduling of such multi-core DNN accelerators. Stream supports finegrained layer fusion, to optimally trade-off energy, latency, and/or on-chip memory footprint for constrained edge devices. Validation against three SotA chips, together with a case study on seven HW architectures with different scheduling granularity, demonstrate the reliability and capabilities of Stream. Results show that high-level architectural decisions greatly impact HW efficiency under the fine-grained scheduling paradigm, reducing the energy-delay product from $2.4 \times$ for single-core architectures to up to $30 \times$ for heterogeneous multi-core architectures compared to traditional scheduling at layer granularity. Stream is open-source at github.com/ZigZag-Project/stream.
Arne Symons, Linyan Mei, Steven Colleman, Pouya Houshmand, Sebastian Karl, Marian Verhelst
ISPASS3
2023 COAC: Cross-Layer Optimization of Accelerator Configurability for Efficient CNN Processing
abstract
To achieve high accuracy, convolutional neural networks (CNNs) are increasingly growing in complexity and diversity in layer types and topologies. This makes it very challenging to efficiently deploy such networks on custom processor architectures for resource-scarce edge devices. Existing mapping exploration frameworks enable searching for the optimal execution schedules or hardware mappings of individual network layers, by optimizing each layer’s spatial (dataflow parallelization) and temporal unrolling (TU, execution order). However, these tools fail to take into account the overhead of supporting different unrolling schemes within a common hardware architecture. Using a fixed unrolling scheme across all layers is also not ideal, as this misses significant opportunities for energy and latency savings from optimizing the mapping of diverse layer types. A balanced approach assesses the right amount of mapping flexibility needed across target neural networks, while taking into account the overhead to support multiple unrollings. This article, therefore, presents cross-layer optimization of accelerator configurability (COAC), a cross-layer design space exploration and mapping framework to optimize the flexibility of neural processing architectures by balancing configurability overhead against resulting energy and latency savings for end-to-end inference. COAC does not only provide a systematical analysis of the architectural overhead in function of the supported spatial unrollings (SUs), but also builds an automated flow to find the best unrolling combination(s) for efficient end-to-end inference with limited hardware overhead. Results demonstrate that architectures with carefully optimized flexibility can achieve up to 38% energy-delay-product (EDP) savings for a set of six neural networks at the expense of a relative area increase of 9.5%.
Steven Colleman, Man Shi, Marian Verhelst
IEEE Trans. Very Large Scale Integr. Syst.1
2021 Processor Architecture Optimization for Spatially Dynamic Neural Networks
abstract
Spatially dynamic neural networks adjust network execution based on the input data, saving computations by skipping non-important image regions. Yet, GPU implementations fail to achieve speedups from these spatially dynamic execution patterns for most neural network architectures. This paper investigates hardware constraints preventing such speedup and proposes and compares novel processor architectures and dataflows enabling latency improvements due to the dynamic execution with minimal loss of utilization. The presented architectures flexibly support spatial execution of a broad range of networks. For the derived architectures, the spatial unrolling for each layer type is optimized and validated making use of the ZigZag design space exploration framework where appropriate. This allows to benchmark and compare the hardware architectures on NNs for classification and human pose estimation, increasing throughput up to $\times 1.9$ and $\times 2.3$ compared to their static executions, respectively. This is the same order of magnitude as other dynamic execution methods, while being complementary to those.
Steven Colleman, Thomas Verelst, Linyan Mei, Tinne Tuytelaars, Marian Verhelst
VLSI-SoC1
2021 High-Utilization, High-Flexibility Depth-First CNN Coprocessor for Image Pixel Processing on FPGA
abstract
Recently, CNNs are increasingly exploited for pixel processing tasks, such as denoising, which opens up new challenges due to the increased activation and operation count. This article presents a CNN coprocessor architecture to solve these challenges on field-programmable gate array (FPGA) through four main contributions. First, the I/O communication between the host processor and the FPGA is reduced to a minimum using a depth-first (DF) principle. Three new DF approaches are presented. Second, to ensure high throughput, the increased parallelization opportunities of the proposed line-based DF operation are analyzed. Third, introducing programmability to the compute array is introduced to enable a broad deployment while maintaining high utilization of the available multipliers digital signal processings (DSPs), independently of the kernel dimensions and without control of the host processor. This is in contrast with many state-of-the-art FPGA implementations, focusing on only one algorithm and/or one kernel topology. Fourth, a model is built to investigate the influence of architecture parameters and show the benefits of DF. The scalable design can be deployed on a wide range of FPGAs, maintaining 78%-93% DSP utilization across all algorithms (denoising, optical flow, depth estimation, segmentation, and super-resolution) and FPGA platforms. Up to 695 GOPS is achieved on a Zynq XCZU9EG board, matching state-of-the-art performance with a more flexible design. The throughput is compared with other pixel processing architectures on FPGA.
Steven Colleman, Marian Verhelst
IEEE Trans. Very Large Scale Integr. Syst.1