EDBT 2026 Demo / reviewers in the wild / expert
Othon Tomoutzoglou
dblp:154/4391
· DBLP profile ↗
8ranked-venue papers
1as first author
1since 2021 · last 2022
0000-0003-4708-2923ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
GPUs and heterogeneous computing · 26% Parallel and multicore computing · 21% Reconfigurable computing and FPGAs · 16% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
parallel programming runtimes |
0.4 | 1 | 2020 | Efficient Job Offloading in Heterogeneous Systems Through Hardware-Assisted Packet-Based Dispatching and User-Level Runtime Infrastructure · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator |
0.3 | 1 | 2018 | Energy-Performance Considerations for Data Offloading to FPGA-Based Accelerators Over PCIe · ACM Trans. Archit. Code Optim. 2018 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.3 | 1 | 2018 | Energy-Performance Considerations for Data Offloading to FPGA-Based Accelerators Over PCIe · ACM Trans. Archit. Code Optim. 2018 |
Interconnection networks and networks-on-chip
network interface |
0.2 | 1 | 2015 | Security in MPSoCs: A NoC Firewall and an Evaluation Framework · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Processor architecture and microarchitecture
security enforcement |
0.2 | 1 | 2015 | Security in MPSoCs: A NoC Firewall and an Evaluation Framework · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Methods — techniques the papers use, named apart from their topics
unified virtual memory · 0.4packet-based dispatching · 0.4zero-copy transfer · 0.3scatter-gather DMA · 0.3rule-based filtering · 0.2gem5 simulation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Poster: Secure Multi-tenant Provisioning of IoT Devices by Combining On-chip Cortex-M TrustZone with Secure Element
Dimitrios Bakoyiannis, Othon Tomoutzoglou, Marcello Coppola |
EWSN | 2 |
| 2020 | Towards holistic secure networking in connected vehicles through securing CAN-bus communication and firmware-over-the-air updating
Othon Tomoutzoglou, Dimitrios Mbakoyiannis, Nikolas Karadimitriou, Marcello Coppola, Eleonora Montanari, Ioannis A. Deligiannis, Giovanni Gherardi |
J. Syst. Archit. | 2 |
| 2020 | Efficient Job Offloading in Heterogeneous Systems Through Hardware-Assisted Packet-Based Dispatching and User-Level Runtime InfrastructureabstractEmerging heterogeneous systems architectures increasingly integrate general-purpose processors, GPUs, and other specialized computational units to provide both power and performance benefits. While the motivations for developing systems with accelerators are clear, it is important to design efficient dispatching mechanisms in terms of performance and energy while leveraging programmability and orchestration of the diverse computational components. In this paper, we present an infrastructure composed of a hardware, general, packet-based processing-dispatching unit, named generic packet processing unit (GPPU), and of an associated runtime that facilitates userlevel access to GPPU objects, such as packets, queues, and contexts. Hence, we remove drawbacks of traditional costly userto-kernel-level operations, low-level accelerator subtleties that hinder programming productivity, along with architectural obstacles such as handling accelerators' unified virtual address space. We present the design and evaluation of our framework by integrating the GPPU infrastructure with data streaming type accelerators, image filtering, and matrix multiplication, tightly coupled to ARMv8 architecture via unified virtual memory. Under scaling workload our proposed dispatching methods can deliver 3.7× performance improvement over baseline offloading, and up to 4.7× better energy efficiency. Othon Tomoutzoglou, Dimitrios Mbakoyiannis, Marcello Coppola |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Energy-Performance Considerations for Data Offloading to FPGA-Based Accelerators Over PCIeabstractModern data centers increasingly employ FPGA-based heterogeneous acceleration platforms as a result of their great potential for continued performance and energy efficiency. Today, FPGAs provide more hardware parallelism than is possible with GPUs or CPUs, whereas C-like programming environments facilitate shorter development time, even close to software cycles. In this work, we address limitations and overheads in access and transfer of data to accelerators over common CPU-accelerator interconnects such as PCIe. We present three different FPGA accelerator dispatching methods for streaming applications (e.g., multimedia, vision computing). The first uses zero-copy data transfers and on-chip scratchpad memory (SPM) for energy efficiency, and the second uses also zero-copy but shared copy engines among different accelerator instances and local external memory. The third uses the processor’s memory management unit to acquire the physical address of user pages and uses scatter-gather data transfers with SPM. Even though all techniques exhibit advantages in terms of scalability and relieve the processor from control overheads through using integrated schedulers, the first method presents the best energy-efficient acceleration in streaming applications. Dimitrios Mbakoyiannis, Othon Tomoutzoglou |
ACM Trans. Archit. Code Optim. | 2 |
| 2015 | Hardware Support for Cost-Effective System-Level Protection in Multi-core SoCsabstractThe increasing adoption of multi-core Systems-on-Chip (SoC) in critical systems has turned security into an important design requirement. In addition to making a SoC tamper-resistant by embedding cryptographic solutions, in order to make a system robust, we need to control the level of access to the critical functions and capabilities. We propose a hardware protection architecture to enhance a traditional SoC platform in terms of protection. These hardware enhancements focus on isolating physical memory compartments by applying access rules, thus we allow dynamic security policies to be enforced at the hardware for protection against untrustworthy hardware or software components. We present and analyze an implementation of a prototype that allows sixteen concurrently active protection domains at a system cost of less that three percent and negligible operational overhead. Ioannis Christoforakis, Othon Tomoutzoglou, Dimitrios Bakoyiannis, Kallia Vazakopoulou, Miltos D. Grammatikakis, Antonis Papagrigoriou |
DSD | 3 |
| 2015 | Dithering-Based Power and Thermal Management on FPGA-Based Multi-core Embedded SystemsabstractIn this paper, we describe the design of a heterogeneous island-based network-on-chip to achieve a power-and thermal-aware coherent system. To this end we utilize different management techniques which employ dynamic frequency scaling circuitry and continuous monitoring through power and temperature sensors per node for dynamic control of workloads. Both monitoring functions and response mechanisms can be engaged in distributed and in centralized mode. The developed multi-core architecture on a multi-FPGA platform employes a hierarchical memory model and supports a multi-threaded general-purpose processor together with many soft-core accelerators per node with independent dynamic frequency scaling per core. Utilizing on-line monitoring we propose a novel response mechanism using a distributed power management algorithm to evenly reduce and normalize power transients. Ioannis Christoforakis, Othon Tomoutzoglou, Dimitrios Bakoyiannis |
EUC | 2 |
| 2015 | Security in MPSoCs: A NoC Firewall and an Evaluation FrameworkabstractIn multiprocessor system-on-chip (MPSoC), a CPU can access physical resources, such as on-chip memory or I/O devices. Along with normal requests, malevolent ones, generated by malicious processes running in one or more CPUs, could occur. A protection mechanism is therefore required to prevent injection of malicious instructions or data across the system. We propose a self-contained Network-on-Chip (NoC) firewall at the network interface (NI) layer which, by checking the physical address against a set of rules, rejects untrusted CPU requests to the on-chip memory, thus protecting all legitimate processes running in a multicore SoC. To sustain high performance, we implement the firewall in hardware, with rule-checking performed at segment-level based on deny rules. Furthermore, to evaluate its impact, we develop a novel framework on top of gem5 simulation environment, coupling ARM technology and an instance of a commercial point-to-point interconnect from STMicroelectronics (STNoC). Simulation tests include scenarios in which legitimate and malicious processes, running in different CPUs, request access to shared memory. Our results indicate that a firewall implementation at the NI can have a positive effect on network performance by reducing both end-to-end network delay and power consumption. We also show that our coarse-grain firewall can prevent saturation of the on-chip network and performs better than fine-grain alternatives that perform rule checking at page-level. Simulation results are accompanied with field measurements performed on a Zedboard platform running Linux, whereas the NoC Firewall is implemented as a reconfigurable, memory-mapped device on top of AMBA AXI4 interconnect fabric. Miltos D. Grammatikakis, Kyprianos Papademetriou, Polydoros Petrakis, Antonis Papagrigoriou, Ioannis Christoforakis, Othon Tomoutzoglou, George Tsamis, Marcello Coppola |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2014 | Runtime Adaptation of Embedded Tasks with A-Priori Known Timing Behavior Utilizing On-Line Partner-Core Monitoring and RecoveryabstractAs the development of heterogeneous embedded Systems-on-Chip with a multitude of hardware accelerator coprocessors creates new possibilities for evolution in aerospace, medicine, communications and consumer eras, improving reliable performance of systems is therefore increasingly important and challenging. Our contributions pertaining to this context are two-fold. We focus on enhancing reliability in the execution of coprocessor tasks with a priori known execution times by allowing an embedded system to identify anomalous software behaviors and additionally to provide rapid online reconfiguration and re-execution in run-time. We present an innovative methodology that combines hardware and software techniques for flexibility, through essentially employing low-cost on-line monitoring, debugging and real-time replacement of the failing sections of software algorithms in an embedded multi-core system. The proposed mechanisms introduce negligible performance degradation, reduced hardware cost and require minimum code instrumentation. Ioannis Christoforakis, Othon Tomoutzoglou, Dimitrios Bakoyiannis |
EUC | 2 |