Amal Gueroudji

dblp:312/3944 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0004-4830-3139ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 StreamGuard: Low-Overhead Resilience for Real-time HPC Data Streams
abstract
Real-time scientific workflows operate on continuous data streams and must produce timely, high-quality results despite executing on complex, failure-prone infrastructure. Hardware faults, network disruptions, and performance anomalies caused by resource contention or system heterogeneity can severely degrade performance and violate real-time constraints. We focus on strengthening the resilience of the producer–consumer streaming pattern, a fundamental building block of scientific streaming workflows. We present two complementary techniques: (i) a dynamic, asynchronous, non-blocking checkpointing mechanism that preserves progress without interrupting computation, and (ii) a progress-aware load redistribution strategy that detects slow workers and proactively rebalances tasks. Together, these mechanisms maintain forward progress and balanced execution even in highly error-prone environments. Experimental results show that our approach reduces the impact of failures and performance anomalies by up to 6 ×, while introducing less than 1% overhead in failure-free execution.
Hai Nguyen 0005, Bogdan Nicolae, Tekin Bicer, Amal Gueroudji, Matthieu Dorier, Kyle Chard, Ian T. Foster
ICS4
2025 PROV-AGENT: Unified Provenance for Tracking AI Agent Interactions in Agentic Workflows
abstract
Large Language Models (LLMs) and other foundation models are increasingly used as the core of AI agents. In agentic workflows, these agents plan tasks, interact with humans and peers, and influence scientific outcomes across federated and heterogeneous environments. However, agents can hallucinate or reason incorrectly, propagating errors when one agent’s output becomes another’s input. Thus, assuring that agents’ actions are transparent, traceable, reproducible, and reliable is critical to assess hallucination risks and mitigate their workflow impacts. While provenance techniques have long supported these principles, existing methods fail to capture and relate agent-centric metadata such as prompts, responses, and decisions with the broader workflow context and downstream outcomes. In this paper, we introduce PROV-AGENT, a provenance model that extends W3C PROV and leverages the Model Context Protocol (MCP) and data observability to integrate agent interactions into end-to-end workflow provenance. Our contributions include: (1) a provenance model tailored for agentic workflows, (2) a near real-time, open-source system for capturing agentic provenance, and (3) a cross-facility evaluation spanning edge, cloud, and HPC environments, demonstrating support for critical provenance queries and agent reliability analysis.
Renan Souza 0001, Amal Gueroudji, Stephen DeWitt, Daniel Rosendo, Tirthankar Ghosal, Robert B. Ross, Prasanna Balaprakash, Rafael Ferreira da Silva
eScience2
2025 ControlA: Agentic Workflow Control Mechanisms for Reliable Science
abstract
AI-driven scientific discovery has emerged as a transformative fifth paradigm in research, with agentic AI playing an increasingly prominent role across scientific domains. Agentic AI can enable collaborative AI-human or even fully autonomous decision-making, but it also introduces significant reliability challenges due to the dynamic and evolutionary nature of the AI agents. Specifically, foundation model-powered agents are prone to generating hallucinated, misleading, or adversarial outputs that can propagate silently through workflows and corrupt downstream results. In this paper we present a conceptual framework for a unified approach that integrates agentic workflow-level instrumentation and agent-level safeguards to enhance the reliability of the wider system, particularly critical in science. Embedding these mechanisms into a provenance-augmented infrastructure enables early detection, containment, and recovery from erroneous behavior, ultimately enhancing reliability and reproducibility in AI-assisted scientific workflows.
Amal Gueroudji, Tanwi Mallick, Renan Souza 0001, Rafael Ferreira da Silva, Robert B. Ross, Matthieu Dorier, Philip H. Carns, Kyle Chard, Ian T. Foster
eScience1
2025 PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training
abstract
Spatiotemporal graph neural networks (ST-GNNs) are powerful tools for modeling spatial and temporal data dependencies. However, their applications have been limited primarily to small-scale datasets because of memory constraints. While distributed training offers a solution, current frameworks lack support for spatiotemporal models and overlook the properties of spatiotemporal data. Informed by a scaling study on a large-scale workload, we present PyTorch Geometric Temporal Index (PGT-I), an extension to PyTorch Geometric Temporal that integrates distributed data parallel training and two novel strategies: index-batching and distributed-index-batching. Our index techniques exploit spatiotemporal structure to construct snapshots dynamically at runtime, significantly reducing memory overhead, while distributed-index-batching extends this approach by enabling scalable processing across multiple GPUs. Our techniques enable the first-ever training of an ST-GNN on the entire PeMS dataset without graph partitioning, reducing peak memory usage by up to 89% and achieving up to a 11.78x speedup over standard DDP with 128 GPUs.
Seth Ockerman, Amal Gueroudji, Tanwi Mallick, Yixuan He 0001, Line C. Pouchard, Robert B. Ross, Shivaram Venkataraman
SC2
2024 Diaspora: Resilience-Enabling Services for Real-Time Distributed Workflows
abstract
The need for real-time processing to enable automated decision making and experimental steering has driven a shift from high-performance computing workflows on a centralized system to a distributed approach that integrates remote data sources, edge devices, and diverse compute facilities. Under this paradigm, data can be processed close to the source where it is generated, thus reducing latency and bandwidth usage. System resilience is thus a key challenge, requiring distributed workflows to survive component failures and to meet stringent quality-of-service requirements, which results in the need to mitigate anomalies such as congestion and low availability of resources. To address these challenges, we propose Diaspora, a unified resilience framework that is inspired by event-driven communication patterns used in public clouds. Specifically, we propose an event fabric that extends across sites, facilities, and computations to provide timely, reliable, and accurate information about data, application, and resource status. On top of the event fabric, we build resilience-enabling services that combine QoS-aware data streaming, resilient data views, resilient compute and data resources, and anomaly detection and prediction, all of which collectively enhance workflow resilience for these scientific cases.
Bogdan Nicolae, Justin M. Wozniak, Tekin Bicer, Hai Nguyen 0005, Haochen Pan, Amal Gueroudji, Maxime Gonthier, Valérie Hayot-Sasson, Eliu A. Huerta, Kyle Chard, Ryan Chard, Matthieu Dorier, Nageswara S. V. Rao, Anees Al-Najjar, Alessandra Corsi, Ian T. Foster
e-Science7
2024 Extreme-scale workflows: A perspective from the JLESC international community
Orcun Yildiz, Amal Gueroudji, Julien Bigot, Bruno Raffin, Rosa M. Badia, Tom Peterka
Future Gener. Comput. Syst.2
2021 DEISA: Dask-Enabled In Situ Analytics
abstract
A widening performance gap is separating CPU performance and IO bandwidth on large scale systems. In some fields such as weather forecast and nuclear fusion, numerical models generate such amounts of data that classical post hoc processing is not feasible anymore due to the limits in both storage capacity and IO performance. In situ approaches are attractive to bypass disk accesses in these cases and fully leverage the HPC platform. They are however often complex to set up and can require to re-develop parallel versions of the analysis from scratch. In this paper we propose a hybrid model that is well suited for in situ workflows that combine regular simulations and irregular analytics. Our model couples the bulk synchronous parallel paradigm for simulation with a distributed task-based one for analysis. This reduces complexity and leverages the best of each of these two powerful paradigms. We validate the model with a prototype, called DEISA, that supports coupling MPI parallel codes with analyses written using Dask. This implementation requires minimal modifications of both the simulation and analysis codes compared to their post hoc counterpart. It give access to an already existing rich ecosystem to be used in situ such as the parallel versions of Numpy, Pandas and scikit-learn. Experiments in configurations up to 1024 cores show that DEISA can improve the simulation wallclock time (excluding analysis) by a factor up to 3 and the total experiment (including analysis) hour. core cost by a factor of up to 5 compared to parallel post hoc with plain Dask while requiring the modification of only two lines of python code, three of YAML, and none at all in a C simulation code already instrumented with PDI Data Interface.
Amal Gueroudji, Julien Bigot, Bruno Raffin
HiPC1