Petros Toupas

dblp:250/6596 · DBLP profile ↗
← Back
6ranked-venue papers
6as first author
4since 2021 · last 2024
0000-0002-3836-071XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2024 SMOF: Streaming Modern CNNs on FPGAs with Smart Off-Chip Eviction
abstract
Convolutional Neural Networks (CNNs) have demonstrated their effectiveness in numerous vision tasks. However, their high processing requirements necessitate efficient hardware acceleration to meet the application's performance targets. In the space of FPGAs, streaming-based dataflow architectures are often adopted by users, as significant performance gains can be achieved through layer-wise pipelining and reduced off-chip memory access by retaining data on-chip. However, modern topologies, such as the UNet, YOLO, and X3D models, utilise long skip connections, requiring significant on-chip storage and thus limiting the performance achieved by such system architectures. The paper addresses the above limitation by introducing weight and activation eviction mechanisms to off-chip memory along the computational pipeline, taking into account the available compute and memory resources. The proposed mechanism is incorporated into an existing toolflow, expanding the design space by utilising off-chip memory as a buffer. This enables the mapping of such modern CNNs to devices with limited on-chip memory, under the streaming architecture design approach. SMOF has demonstrated the capacity to deliver competitive and, in some cases, state-of-the-art performance across a spectrum of computer vision tasks, achieving up to 10.65 × throughput improvement compared to previous works. The tool is available at https://github.com/ICIdsl/smof.git.
Petros Toupas, Zhewen Yu, Christos-Savvas Bouganis, Dimitrios Tzovaras
FCCM1
2023 FMM-X3D: FPGA-Based Modeling and Mapping of X3D for Human Action Recognition
abstract
3D Convolutional Neural Networks are gaining increasing attention from researchers and practitioners and have found applications in many domains, such as surveillance systems, autonomous vehicles, human monitoring systems, and video retrieval. However, their widespread adoption is hindered by their high computational and memory requirements, especially when resource-constrained systems are targeted. This paper addresses the problem of mapping X3D, a state-of-the-art model in Human Action Recognition that achieves accuracy of 95.5% in the UCF101 benchmark, onto any FPGA device. The proposed toolflow generates an optimised stream-based hardware system, taking into account the available resources and off-chip memory characteristics of the FPGA device. The generated designs push further the current performance-accuracy pareto front, and enable for the first time the targeting of such complex model architectures for the Human Action Recognition task.
Petros Toupas, Christos-Savvas Bouganis, Dimitrios Tzovaras
ASAP1
2023 HARFLOW3D: A Latency-Oriented 3D-CNN Accelerator Toolflow for HAR on FPGA Devices
abstract
For Human Action Recognition tasks (HAR), 3D Convolutional Neural Networks have proven to be highly effective, achieving state-of-the-art results. This study introduces a novel streaming architecture-based toolflow for mapping such models onto FPGAs considering the model's inherent characteristics and the features of the targeted FPGA device. The HARFLOW3D toolflow takes as input a 3D CNN in ONNX format and a description of the FPGA characteristics, generating a design that minimises the latency of the computation. The toolflow is comprised of a number of parts, including (i) a 3D CNN parser, (ii) a performance and resource model, (iii) a scheduling algorithm for executing 3D models on the generated hardware, (iv) a resource-aware optimisation engine tailored for 3D models, (v) an automated mapping to synthesizable code for FPGAs. The ability of the toolflow to support a broad range of models and devices is shown through a number of experiments on various 3D CNN and FPGA system pairs. Furthermore, the toolflow has produced high-performing results for 3D CNN models that have not been mapped to FPGAs before, demonstrating the potential of FPGA-based systems in this space. Overall, HARFLOW3D has demonstrated its ability to deliver competitive latency compared to a range of state-of-the-art hand-tuned approaches, being able to achieve up to 5× better performance compared to some of the existing works. The tool is available at https://github.com/ptoupas/harflow3d.
Petros Toupas, Alexander Montgomerie-Corcoran, Christos-Savvas Bouganis, Dimitrios Tzovaras
FCCM1
2023 fpgaHART: A Toolflow for Throughput-Oriented Acceleration of 3D CNNs for HAR onto FPGAs
abstract
Surveillance systems, autonomous vehicles, human monitoring systems, and video retrieval are just few of the many applications in which 3D Convolutional Neural Networks are exploited. However, their extensive use is restricted by their high computational and memory requirements, especially when integrated into systems with limited resources. This study proposes a toolflow that optimises the mapping of 3D CNN models for Human Action Recognition onto FPGA devices, taking into account FPGA resources and off-chip memory characteristics. The proposed system employs Synchronous Dataflow (SDF) graphs to model the designs and introduces transformations to expand and explore the design space, resulting in high-throughput designs. A variety of 3D CNN models were evaluated using the proposed toolflow on multiple FPGA devices, demonstrating its potential to deliver competitive performance compared to earlier hand-tuned and model-specific designs.
Petros Toupas, Christos-Savvas Bouganis, Dimitrios Tzovaras
FPL1
2019 Accelerating Physics Engine Components with Embedded FPGAs
abstract
In recent years there has been a steady increase in the use of physics engines, deployed in applications such as video games, scientific simulations, computer graphics and film productions. Their main purpose is to simulate the motions of objects based on real-world physics rules. As the complexity of the simulated scenes increases with the use of multiple objects and desirable effects, the computational cost of the physics-related calculations explodes. Typically, physics engines make use of the general-purpose computational capabilities of modern GPUs in order to take advantage of their massively parallel resources. In this paper, we consider the use of FPGAs to accelerate certain demanding components of the physics simulation pipeline aiming to provide better performing solutions at significantly lower energy cost. The results of our work demonstrate that by employing Zynq UltraScale+ devices featuring embedded ARM cores and FPGA fabric, we can accelerate physics computations of the popular Bullet library on highly demanding scenes up to 2.2x compared to high-end GPUs at a fraction of the energy required (up to 44x better energy efficiency).
Petros Toupas, Andreas Brokalakis, Ioannis Papaefstathiou
FPL1
2019 An Intrusion Detection System for Multi-class Classification Based on Deep Neural Networks
abstract
Intrusion Detection Systems (IDSs) are considered as one of the fundamental elements in the network security of an organisation since they form the first line of defence against cyber threats, and they are responsible to detect effectively a potential intrusion in the network. Many IDS implementations use flow-based network traffic analysis to detect potential threats. Network security research is an ever-evolving field and IDSs in particular have been the focus of recent years with many innovative methods proposed and developed. In this paper, we propose a deep learning model, more specifically a neural network consisting of multiple stacked Fully-Connected layers, in order to implement a flow-based anomaly detection IDS for multi-class classification. We used the updated CICIDS2017 dataset for training and evaluation purposes. The experimental outcome using MLP for intrusion detection system, showed that the proposed model can achieve promising results on multi-class classification with respect to accuracy, recall (detection rate), and false positive rate (false alarm rate) on this specific dataset.
Petros Toupas, Dimitra Chamou, Konstantinos M. Giannoutakis, Anastasios Drosou, Dimitrios Tzovaras
ICMLA1