EDBT 2026 Demo / reviewers in the wild / expert
Dionysios Filippas
dblp:313/3274
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2023
0000-0002-4729-3336ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | ArrayFlex: A Systolic Array Architecture with Configurable Transparent PipeliningabstractConvolutional Neural Networks (CNNs) are the state-of-the-art solution for many deep learning applications. For maximum scalability, their computation should combine high performance and energy efficiency. In practice, the convolutions of each CNN layer are mapped to a matrix multiplication that includes all input features and kernels of each layer and is computed using a systolic array. In this work, we focus on the design of a systolic array with configurable pipeline with the goal to select an optimal pipeline configuration for each CNN layer. The proposed systolic array, called ArrayFlex, can operate in normal, or in shallow pipeline mode, thus balancing the execution time in cycles and the operating clock frequency. By selecting the appropriate pipeline configuration per CNN layer, ArrayFlex reduces the inference latency of state-of-the-art CNNs by 11 %, on average, as compared to a traditional fixed-pipeline systolic array. Most importantly, this result is achieved while using 13 %-23 % less power, for the same applications, thus offering a combined energy-delay-product efficiency between$1.4\times$and$1.8\times$. Christodoulos Peltekis, Dionysios Filippas, Giorgos Dimitrakopoulos, Chrysostomos Nicopoulos, Dionisios N. Pnevmatikatos |
DATE | 2 |
| 2023 | Streaming Dilated Convolution EngineabstractConvolution is one of the most critical operations in various application domains and its computation should combine high performance with energy efficiency. This requirement is critical both for standard convolution and for its other spatial variants, such as dilated, strided, or transposed convolutions. In this work, we focus on the design of a streaming convolution engine, called LazyDCstream, that is tuned for dilated convolution. LazyDCstream utilizes a sliding-window architecture for input data reuse and leverages the already-known decomposition of dilated convolution to: (a) maximize window buffer sharing and (b) enable “lazy” data movement that keeps data transfers per clock cycle as few as possible, and, most importantly, independent of the dilation rate. These two architectural features reduce the power consumption relative to efficient streaming convolution engines without introducing any throughput or area penalty. Dionysios Filippas, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | Synthesis of Approximate Parallel-Prefix AddersabstractApproximate computation has evolved recently as a viable alternative for maximizing energy efficiency. One aspect of approximate computing involves the design of hardware units that return a sufficiently accurate result for the examined occasion, rather than computing an accurate result. As long as the hardware units are allowed to compute approximately, they can be designed with multiple new ways. In this work, we focus on the synthesis of approximate parallel-prefix adders. Instead of exploring specific architectures, as done by state-of-the-art approaches, the introduced synthesizer can produce every solution that meets the designer’s criteria, resulting in adders with various delay, area, and error tradeoffs. This automatic design space exploration allows approaching, in several cases, optimal solutions that could have not been designed with any other known parallel-prefix architecture. The synthesized adders, when compared with state-of-the-art adders, achieve 27%–36% better error frequency (EF) on average for random inputs and improve image quality metrics by 8%–42% for image filtering. These results are achieved with the proposed adders requiring the same or marginally more hardware area or energy. On the contrary, in split-accuracy configurations, more than 30% of hardware area/energy can be saved for the same classification accuracy for a neural network application. Apostolos Stefanidis, Ioanna Zoumpoulidou, Dionysios Filippas, Giorgos Dimitrakopoulos, Georgios Ch. Sirakoulis |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | FusedGCN: A Systolic Three-Matrix Multiplication Architecture for Graph Convolutional NetworksabstractMachine-learning applications have garnered widespread adoption over the last several years. Graph Neural Networks have been proposed as an extension of machine-learning models to graph-structured data. The training and inference tasks on graph neural networks involve graph convolution operations that can be equivalently expressed as three-matrix multiplications. In this work, we propose FusedGCN, a custom systolic architecture that computes in a fused, i.e., combined, manner the product of three matrices. FusedGCN supports compressed sparse representations and tiled computations, which allow the design to adapt to the available input/output bandwidth without losing the regularity of a systolic architecture. The experimental results show that FusedGCN achieves lower execution times than the current best-performing state-of-the-art architecture for computing representative GCN applications. Most importantly, this result is achieved by consuming only marginally more area/power than a traditional systolic array used for two-matrix multiplications. Christodoulos Peltekis, Dionysios Filippas, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos |
ASAP | 2 |
| 2022 | Low-Cost Online Convolution Checksum CheckerabstractManaging random hardware faults requires the faults to be detected online, thus simplifying recovery. Algorithm-based fault tolerance has been proposed as a low-cost mechanism to check online the result of computations against random hardware failures. In this case, the checksum of the actual result is checked against a predicted checksum computed in parallel by a hardware checker. In this work, we target the design of such checkers for convolution engines that are currently the most critical building block in image processing and computer vision applications. The proposed convolution checksum checker, named ConvGuard, utilizes a newly introduced invariance condition of convolution to predictimplicitlythe output checksum using only the pixels at the border of the input image. In this way, ConvGuard reduces the power required for accumulating the input pixels without requiring large buffers to hold intermediate checksum results. The design of ConvGuard is generic and can be configured for different output sizes and strides. The experimental results show that ConvGuard utilizes only a small percentage of the area/power of an efficient convolution engine while being significantly smaller and more power efficient than a state-of-the-art checksum checker for various practical cases. Dionysios Filippas, Nikolaos Margomenos, Nikolaos Mitianoudis, Chrysostomos Nicopoulos, Giorgos Dimitrakopoulos |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |