Olivier Franza

dblp:56/1551 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0007-0803-9064ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware reliability and fault tolerance · 30% High-performance computing · 30% Distributed systems · 30%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › supercomputing
exascale computing
0.912025
Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated Systems · SC 2025
Distributed systems
fault tolerance
0.912025
Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated Systems · SC 2025
GPUs and heterogeneous computing › GPU computing
GPU-accelerated systems
0.312025
Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated Systems · SC 2025

Methods — techniques the papers use, named apart from their topics

failure analysis · 0.9
YearPublicationVenuePosition
2025 Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated Systems
abstract
As high-performance computing (HPC) systems scale in size, system wide hardware failure rates increase. Historical data from previous large-scale HPC installations illustrate this trend, with the mean time between failures (MTBF) decreasing steadily over the past decade. Recent studies from artificial intelligence and machine-learning (AI/ML) training extrapolate MTBF declining even further for future GPU accelerated systems. As MTBF decreases, mean time to repair (MTTR) becomes more pronounced, highlighting the need for efficient recovery strategies.
Yonatan Levitt, Richard Barella, Sam Zeltner, Thomas Musta, Lance C. Cheney, Gustavo Espinosa, Olivier Franza, Balazs Gerofi
SC7
2021 Autonomic Architectural Framework for Internet of Brain Controlled Things (IoBCT)
abstract
Brain Machine Interface (BMI) involves the acquisition and analysis of brainwave signals and translating them into commands that can be relayed to control devices in order to achieve a desired action. With the wide adoption of smart objects realized by the Internet of Things (IoT) and wearable technology across applications and industries, one can envision how BMI can empower users to control such objects (i.e.: home appliances, alert systems, or assistive robots), via their thoughts. However, the implementation of a practical BMI-based IoT system is faced with several challenges, most importantly being the issue of accurately translating the user’s intention which is derived from the raw brainwave signals. Such translations are often enabled by computationally heavy algorithms which require a larger computing system. Thus, a new strategy and process flow for a concrete and resilient framework over which a BMI-IoT solution is built will be needed. In this paper, a novel system architecture based on Edge computing is proposed for controlling BMI-IoT enabled environments. The proposed simplified and versatile architecture is aimed at helping the designer articulate the key functions and elements of BMI-based IoT system which we call “Internet of Brain Controlled Things (IoBCT).
Haider Khaleel Raad, Farah Fargo, Olivier Franza
ISNCC3
2019 Autonomic Resource Management for Power, Performance, and Security in Cloud Environment
abstract
High performance computing is widely used for large-scale simulations, designs and analysis of critical problems especially through the use of cloud computing systems nowadays because cloud computing provides ubiquitous, on-demand computing capabilities with large variety of hardware configurations including GPUs and FPGAs that are highly used for high performance computing. However, it is well known that inefficient management of such systems results in excessive power consumption affecting the budget, cooling challenges, as well as reducing reliability due to the overheating and hotspots. Furthermore, considering the latest trends in the attack scenarios and crypto-currency based intrusions, security has become a major problem for high performance computing. Therefore, to address both challenges, in this paper we present an autonomic management methodology for both security and power/performance. Our proposed approach first builds knowledge of the environment in terms of power consumption and the security tools' deployment. Next, it provisions virtual resources so that the power consumption can be reduced while maintaining the required performance and deploy the security tools based on the system behavior. Using this approach, we can utilize a wide range of secure resources efficiently in HPC system, cloud computing systems, servers, embedded systems, etc.
Farah Fargo, Olivier Franza, Cihan Tunc, Salim Hariri
AICCSA2
2018 Autonomic Secure HPC Architecture Against Power Attacks
abstract
High Performance Computing (HPC) systems are enabling broad computing capabilities across scientific simulations, data analytics, and machine learning. Such HPC systems are composed of high-end compute elements (CPUs, accelerators, co-processors...) to handle high throughput and high parallelism, which results in significant power consumption. However, the power infrastructures built for such systems are proportionately older and may not be built to sustain system's peak power for extensive periods of time. Therefore, innovative power management schemes are used to allocate power across all system's components within a given power limit. HPC power management systems regulate resources' power within budget constraints and the overall system's power limit. Failing not to limit system's power consumption can result in operational failures that can bring down the whole system. Additionally, system reliability would be compromised under extensive computational system power and temperature. In addition to power consumption, possible attack scenarios can cause reliability issues where the attackers can manipulate the reported data of the measured power that is used for power management, which in turn can misguide the schedulers or can assign tasks that will result in exceeding the power limits, leading to system failures. Hence, intrusion detection systems (IDS) for power attacks are extremely important. This paper presents a framework for power attack IDS and describe each steps to be taken.
Olivier Franza, Farah Fargo
AICCSA1
2012 Compact Expressions for Supply Noise Induced Period Jitter of Global Binary Clock Trees
abstract
Period jitter plays a critical role in global clock distribution design, because it directly impacts the time available for logic operations between sequential elements. Moreover, time-varying supply noise injected in global clock drivers can worsen the timing margin of critical paths by modulating period jitter. In the planning stage of global clock distribution for a high-end microprocessor, it is very critical to differentiate and understand the impacts of different design parameters on period jitter. However, it is hard to achieve due to complex relationship among different independent/dependent design parameters: supply noise amplitude, supply noise frequency, clock driver size, physical structures of interconnects, number of clock stages, temperature, process corners, etc.
Jinwook Jang, Olivier Franza, Wayne P. Burleson
IEEE Trans. Very Large Scale Integr. Syst.2